Li

Lightless Labs Refinery – multi-model consensus and synthesis

Hacker News

Lightless Labs Refinery – multi-model consensus and synthesis

Hi! In the past few weeks I (mostly Claude) cobbled together a Rust library + cli to run the same prompt across multiple models, through multiple rounds of iterative consensus. Each model is fed the same initial prompt, produces an answer, then every model individually reviews and scores each of the other model's answers independently. The original prompt, previous answer, and the reviews, are then fed back to the models for the next round, until either one model "wins" two rounds in a row or a limit is reached. It did quite well on the car wash test ( https://github.com/Lightless-Labs/refinery?tab=readme-ov-fil... ). Most models answer badly initially, but it just takes one for all of them to quickly converge towards better answers. Although, to my initial surprise, adding more models quickly breaks the current voting+threshold selection strategy. I also recently added a synthesis mode, which does the same thing but with an additional synthesis round at the end where each model produces a synthesis of all the answers that scored above the threshold in the last round, followed by one last review round. The total number of calls quickly blows up with rounds and model count, but it's been fun! Currently, I'm racking my brain trying to figure out a way to select for both diversity and quality, for a "brainstorm" process. If you have any ideas either on that or other features, let me know!

Share card

Actual performance

2points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
81%81% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
66%66% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: answers, way · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
46%46% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: reviews, calls · Missing: plus, platform, intuitive
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Goombay Labs
Goombay Labs74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

#HowDoYouFeelToday?

Indie Hackers1$500/moe-commerce
Re
RelatedAI – Multi Model Chat45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RelatedAI – Multi Model Chat

Hacker News3
Mu
Multi-Probe LSH and LSH Forest in Golang35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi-Probe LSH and LSH Forest in Golang

Hacker News1
Mu
Multi-threaded Opus and AAC encoding32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi-threaded Opus and AAC encoding

Hacker News2
hy
hybridcontents – A Multi ContentsManager Wrapper For Jupyter28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

hybridcontents – A Multi ContentsManager Wrapper For Jupyter

Hacker News1
Pelladio
Pelladio41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi Model Collaboration

Indie Hackers1ai
Op
Openfactor – A multi factor risk model46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Openfactor – A multi factor risk model

Hacker News2
Be
Bestow – flexible, multi-model labelling gem for Rails36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bestow – flexible, multi-model labelling gem for Rails

Hacker News2
Mu
Multi Agent Protocol for AI Scientist by Hexo Labs37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi Agent Protocol for AI Scientist by Hexo Labs

Hacker News5
Kenmare Labs
Kenmare Labs70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi-armed bandit A/B testing via link shortener

Indie Hackers1analytics