I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp
I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp
Democratisation of local AI is key. I've been working on pushing the limits of commercial hardware, squeezing any extra bit possible. My Scientific Agentic AI hareness helped me to reallocate every single bit of it. I rewrote the Kernel, I went down the CUDA rabbit hole until I have been able to explain any bit and any ms of computational power involved in the process pushing the Qwen 30B-A3B from 8 tok7s to 19 tok/s with llama.cpp up to 22.2 tok/s with my project and 109 tok/s on not novel content and speeding up the prefill by 5-9X
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Run Llama 3.1 8B in the browser
Decrypting Rita, a graphic novel
American Apocalypse (Novel) (2004)
Nevertwenty a novel approach to chord synthesis
Llama or Alpaca?
Llama 3.2 Interpretability with Sparse Autoencoders
Run any Llama model finetune and more, instantly
4 yrs ago, I wrote an AI novel, now I published & made it free
TypistStories, new Gothic novel released
(gr)album – graphic novel + music album