Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU
Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU
I wanted to know how fast a 26B mixture-of-experts model could run on a desktop CPU with no GPU. Got ~40 tok/s single-stream (lossless) and ~124 batched. The surprising part was the byte budget: for this model you compress the output head (32% of per-token bytes), not the experts (16%). The writeup has the bandwidth roofline and the dead-ends; the repo has the reproducible recipe. Happy to answer questions. Repo: https://github.com/arun-prasath2005/gemma4-cpu-moe
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Metaballs Implementation in the CPU and GPU
wc-GPU: The Unix util `wc` running on a GPU
Checkpoint K8s pods transparently (plain CPU or GPU accelerated) [video]
Find out if your CPU or GPU is holding your PC back
Harvard CPU in Verilog and Assembler in Go
H2 Forth CPU
CPU Databse, CPU, Processors, Mobile CPU, CPU List
Tilery-VM – run Nvidia cuTile GPU kernels on a CPU, no GPU required
Video about the CPU vulnerability Zenbleed (CVE-2023-20593)
Online CPU profiler for Golang