UNet diffusion model in pure CUDA
UNet diffusion model in pure CUDA
Hi HN! I was inspired by Andrej Karpathy's llm.c ( https://github.com/karpathy/llm.c ), and wrote a full diffusion model training loop in CUDA. I learnt a lot about CUDA from Simon Boehm's Matmul blog ( https://siboehm.com/articles/22/CUDA-MMM ). Currently there is still a lot of room for optimization: the model is running at 45% speed of PyTorch with torch.compile. I'm curious about any thoughts or CUDA tips for convolutions.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
CUDA and cuBLAS matrices in Clojure
Raytracer with both C++ and CUDA back ends
One Billion Rows in CUDA
CUDA Fractal Renderer
Pure CUDA C Inference for Qwen3 0.6B in One File, No Dependencies
A pure Tcl JPEG decoder
A pure-Ruby implementation of systemd's sd_notify(3)
Go-osc – OSC Packet Implementation for Golang. Implemented in Pure Go
Pure functional lenses in Racket