Local LLM AIME benchmarking tool
Local LLM AIME benchmarking tool
I made this simple tool to compare local LLMs. Any provider that supports OpenAI-like APIs can be used (LMStudio, Llama.cpp, Ollama) but you can also use Openrouter/OpenAI if you change the base URL accordingly. In my opinion it is not particularly useful for comparing different models from different companies since some models are optimized heavily on math or even trained on AIME problems. However it is really useful for testing different quantizations of the same model or the same quantization from different providers. Let me know what you think about it! Also check the README to see some examples of the results you will get from it.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Redis-LLM – Redis module integrates LLM with Redis
LitLLM the Spiciest LLM Wrapper
LLM Reasonsers
Resilient LLM
Hegelion – Force your LLM to argue with itself before answering
I Stopped Hoping My LLM Would Cooperate
Module for LLM Homeostasis (PoC)
Doom Compiled into an LLM
The Smallest LLM