A minimalist proxy for your local LLM cluster (~1100 lines)
A minimalist proxy for your local LLM cluster (~1100 lines)
Hi all! I made this small proxy when I built a small GPU cluster at home and wanted to share it with friends while keeping speed limits and token accounting. I also put a lot of effort in making it work as quick as possible, and IMO the result is worth sharing! It is not trying to be LiteLLM: in fact, I wanted to make it opposite, as small as possible, with a small amount of dependencies and without cloud connections by default. Will be happy to hear your feedback :) PS also installable with pip install smol-llm-proxy
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Minimalist Tor-to-Web Proxy
Zero dependency LLM standby, rollouts, no proxy
PGT-Proxy – A PostgreSQL Proxy in 277 Lines of Rust
A TCP Proxy in 30 lines of Rust
WyPyPlus is a minimalist wiki in 23 lines of code
Spinning Cube in 45 lines, 45 Chars Each
Flappy Bird in 128 Lines of CoffeeScript
A CORS proxy in a container
Cors-container – A CORS proxy in a container
A CORS proxy in a container