LL

LLM Council – Run multiple LLMs with critique and consensus eval

Hacker News

LLM Council – Run multiple LLMs with critique and consensus eval

Building reliable LLM systems often means not trusting a single model. We open-sourced LLM Council: https://github.com/abhishekgandhi-neo/llm_council It’s a small framework we internally built with Neo to run multiple LLMs on the same task, let them critique each other, and produce a structured final answer. Useful for tasks like: • Comparing local vs API models on your own dataset • Validating RAG outputs • Prompt regression testing • Dataset labeling with model-as-judge • Catching hallucinations in code or research summaries A few practical details: • Async parallel calls so latency stays close to one model • Structured outputs with each model’s answer and critiques • Provider-agnostic configs for local + hosted models • Built to plug into evaluation pipelines, not just demos We built this using Neo. We’ve been experimenting with similar council setups to catch silent failures in ML workflows, and this repo is a cleaned-up version of that idea. If you’ve built multi-LLM evaluation pipelines, would love to hear what aggregation or critique strategies worked well for you.

Share card

Actual performance

4points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, single · Missing: mac, agents, macos
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para · Missing: supports, reddit linkedin, podcasting
60%60% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, pipe · Missing: https docs, excited, just released
52%52% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
36%36% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host, calls · Missing: plus, platform, intuitive
32%32% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ca
Call Multiple LLMs with GraphQL and AI Chainer42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Call Multiple LLMs with GraphQL and AI Chainer

Hacker News2
Mo
ModelMashup – Chat with Multiple LLMs Simultaneously32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ModelMashup – Chat with Multiple LLMs Simultaneously

Hacker News1
Ru
Run LLMs on the Browser74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run LLMs on the Browser

Hacker News6
LL
LLM Litmus Test – compare multiple LLMs for coding tasks with context43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Litmus Test – compare multiple LLMs for coding tasks with context

Hacker News1
Di
Distributed Llama – Run LLMs on multiple devices in parallel76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Distributed Llama – Run LLMs on multiple devices in parallel

Hacker News12
bandleader.ai
bandleader.ai38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Determines the best answer for you across multiple LLMs

Product Hunt+6
Ya
YamChat – Chat with Multiple LLMs from one location45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

YamChat – Chat with Multiple LLMs from one location

Hacker News1
Mu
Multiple invitations on top of devise_invitable (RoR)38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multiple invitations on top of devise_invitable (RoR)

Hacker News1
Mu
Multiple Imputation with Lightgbm38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multiple Imputation with Lightgbm

Hacker News2
Sy
Synchronize OTP credentials across multiple Yubikeys40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synchronize OTP credentials across multiple Yubikeys

Hacker News5