10

100% LLM accuracy–no fine-tuning, JSON only

Hacker News

100% LLM accuracy–no fine-tuning, JSON only

# Show HN: 100% LLM accuracy—no fine-tuning, JSON only *GitHub:* https://github.com/Mysticbirdie/hallucination-elimination-be... *Paper:* https://github.com/Mysticbirdie/hallucination-elimination-be... --- *Model-agnostic Triad Engine*: JSON domain guide → 100% accuracy across Mistral 7B/Claude/GPT—no fine-tuning. --- ## Results (222 questions, adversarial domain, Gemini 2.0 Flash judge) | Model | Raw | + Triad | ∆ | |-------|-----|---------|---| | Mistral 7B (local) | 22.5% | *99.5%* | +77pp | | Bielik 11B (local) | 21.6% | *88.7%* | +67.1pp | | GPT-5.2 | 26.1% | *100%* | +73.9pp | | Gemini 2.5 Pro | 42.3% | *95%* | +52.7pp | | Claude 4.6 | 45.0% | *100%* | +55pp | | Perplexity Sonar (RAG) | 64.4% | *93.7%* | +29.3pp | Perplexity has live web search—Triad still adds 29.3pp. Gemini 2.0 judge (cross-model). Claude Opus (stricter): raw Claude 14.9% → Triad 95.9%. Zero regressions. *Beyond accuracy:* - Adversarial: Raw Claude accepts false premises 25% → Triad 5% - Consistency: Raw Claude 0% agreement across personas → Triad 100% - Concise: 2.1× shorter responses (473 vs 1,015 chars) *Real-world (Windsurf live codebase):* | Phase | Context | Score | |-------|---------|-------| | No context | — | 40% | | Unstructured docs | — | 40% | | Triad JSON guide | — | *100%* | --- ## Triad Engine Multi-voice layer above any LLM: - λ: Character voice from domain guide - μ: Truth/false enforcement - ν: User calibration - ω: Voice compositor *Only input*: JSON domain guide (what exists/doesn't, agents, norms). No weight changes. Works with any LLM. Applies to medical, legal, compliance domains. --- ## Open source 222 questions, runners (Claude/GPT/Gemini/Mistral), JSON results, guide schema in repo. Happy to answer technical questions.

Share card

Actual performance

2points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
59%59% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, open source, ide · Missing: https docs, excited, just released
42%42% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Predibase Reinforcement Fine-Tuning
Predibase Reinforcement Fine-Tuning60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM reinforcement fine-tuning platform to improve LLM output

Product Hunt+172SaaS
GP
GPT Fine-Tuning Notebook Boosts Classification Accuracy from 69% to 94%39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPT Fine-Tuning Notebook Boosts Classification Accuracy from 69% to 94%

Hacker News1
Fi
Fine tuning and RLHF mistralai 7B using DeepSpeed38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fine tuning and RLHF mistralai 7B using DeepSpeed

Hacker News1
Sh
ShadowPEFT – Centralized and Detachable Parameter-Efficient Fine-Tuning44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ShadowPEFT – Centralized and Detachable Parameter-Efficient Fine-Tuning

Hacker News6
Lumino AI
Lumino AI75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Serverless LLM Fine-Tuning SDK

Product Hunt+12
A
A 3 step no-code process for LLM Fine-tuning44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A 3 step no-code process for LLM Fine-tuning

Hacker News2
Te
Terracotta – Platform for fine-tuning and evaluating LLMs40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Terracotta – Platform for fine-tuning and evaluating LLMs

Hacker News1
Op
Open Source Reinforcement Fine-Tuning for Your Agents49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source Reinforcement Fine-Tuning for Your Agents

Hacker News5
Py
Pykoi – a Python library for LLM data collection and fine tuning56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pykoi – a Python library for LLM data collection and fine tuning

Hacker News119
In
Interactive synthetic data generation for LLM fine-tuning46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Interactive synthetic data generation for LLM fine-tuning

Hacker News2