Warp – Run the 313B GLM-5.3-Flash on a MacBook with 8GB RAM
Warp – Run the 313B GLM-5.3-Flash on a MacBook with 8GB RAM
A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well. It requires as little as 5.14 GB of RAM to run, and on a 64 GB MacBook Pro M5 Pro it reaches about 3.32 tok/s, or 3.86 tok/s on longer runs. More memory means a larger expert cache, while higher storage and memory bandwidth can further improve performance. The project is completely open-source and free to use: https://github.com/sqliteai/warp
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
GLM-5.3 744B at 4 tok/s on a MacBook Pro, experts streamed from 4 SSDs
Our GLM-5.3 Flash Switchless recipe is now out for 4x DGX Sparks
MACBOOK SKINS
Slap your MacBook. It screams back. That's it.
Slap your MacBook. It screams back. That's it.
Abliterated GLM-5.3 API (84.5% CyberGym, FP8)
Alpaca.cpp – Run an Instruction-Tuned Chat-Style LLM on a MacBook
Thermal assistant for fanless MacBook Airs
I was able to run tiberian sun on cncnet on an M4 MacBook Pro
Flashover - online flash decompiler