LLM Council – Run multiple LLMs with critique and consensus eval
LLM Council – Run multiple LLMs with critique and consensus eval
Building reliable LLM systems often means not trusting a single model. We open-sourced LLM Council: https://github.com/abhishekgandhi-neo/llm_council It’s a small framework we internally built with Neo to run multiple LLMs on the same task, let them critique each other, and produce a structured final answer. Useful for tasks like: • Comparing local vs API models on your own dataset • Validating RAG outputs • Prompt regression testing • Dataset labeling with model-as-judge • Catching hallucinations in code or research summaries A few practical details: • Async parallel calls so latency stays close to one model • Structured outputs with each model’s answer and critiques • Provider-agnostic configs for local + hosted models • Built to plug into evaluation pipelines, not just demos We built this using Neo. We’ve been experimenting with similar council setups to catch silent failures in ML workflows, and this repo is a cleaned-up version of that idea. If you’ve built multi-LLM evaluation pipelines, would love to hear what aggregation or critique strategies worked well for you.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Call Multiple LLMs with GraphQL and AI Chainer
ModelMashup – Chat with Multiple LLMs Simultaneously
Run LLMs on the Browser
LLM Litmus Test – compare multiple LLMs for coding tasks with context
Distributed Llama – Run LLMs on multiple devices in parallel
Determines the best answer for you across multiple LLMs
YamChat – Chat with Multiple LLMs from one location
Multiple invitations on top of devise_invitable (RoR)
Multiple Imputation with Lightgbm
Synchronize OTP credentials across multiple Yubikeys