A

A new benchmark for testing LLMs for deterministic outputs

Hacker News

A new benchmark for testing LLMs for deterministic outputs

When building workflows that rely on LLMs, we commonly use structured output for programmatic use cases like converting an invoice into rows or meeting transcripts into tickets or even complex PDFs into database entries. The model may return the schema you want, but with hallucinated values like `invoice_date` being off by 2 months or the transcript array ordered wrongly. The JSON is valid, but the values are not. Structured output today is a big part of using LLMs, especially when building deterministic workflows. Current structured output benchmarks (e.g., JSONSchemaBench) only validate the pass rate for JSON schema and types, and not the actual values within the produced JSON. So we designed the Structured Output Benchmark (SOB) that fixes this by measuring both the JSON schema pass rate, types, and the value accuracy across all three modalities, text, image, and audio. For our test set, every record is paired with a JSON Schema and a ground-truth answer that was verified against the source context manually by a human and an LLM cross-check, so a missing or hallucinated value will be considered to be wrong. Open source is doing pretty well with GLM 4.7 coming in number 2 right after GPT 5.4. We noticed the rankings shift across modalities: GLM-4.7 leads text, Gemma-4-31B leads images, Gemini-2.5-Flash leads audio. For example, GPT-5.4 ranks 3rd on text but 9th on images. Model size is not a predictor, either: Qwen3.5-35B and GLM-4.7 beat GPT-5 and Claude-Sonnet-4.6 on Value Accuracy. Phi-4 (14B) beats GPT-5 and GPT-5-mini on text. Structured hallucinations are the hardest bug. Such values are type-correct, schema-valid, and plausible, so they slip through most guardrails. For example, in one audio record, the ground truth is "target_market_age": "15 to 35 years", and a model returns "25 to 35". This is invisible without field-level checks. Our goal is to be the best general model for deterministic tasks, and a key aspect of determinism is a controllable and consistent output structure. The first step to making structured output better is to measure it and hold ourselves against the best.

Share card

Actual performance

60points
30comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, new · Missing: mac, agents, macos
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
72%72% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: month · Missing: mobile apps, ios, personal
50%50% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: open source, ide, io · Missing: https docs, excited, just released
41%41% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

A
A New Implementation of the Seven GUIs Benchmark59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A New Implementation of the Seven GUIs Benchmark

Hacker News4
De
DeepTeam – Penetration Testing for LLMs51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepTeam – Penetration Testing for LLMs

Hacker News3
Kr
KraspAI Kompass – keep up with new LLMs53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

KraspAI Kompass – keep up with new LLMs

Hacker News1
LO
LOL Bench – a benchmark for whether LLMs get jokes48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LOL Bench – a benchmark for whether LLMs get jokes

Hacker News1
Snappy - LLMs Speed Test
Snappy - LLMs Speed Test52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmark your LLMs in Seconds ⚡

Product Hunt+5
A
A new android benchmark35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A new android benchmark

Hacker News1
We
WebGL Sprites Benchmark58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WebGL Sprites Benchmark

Hacker News38
NA
NAB – The Numenta Anomaly Benchmark42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NAB – The Numenta Anomaly Benchmark

Hacker News17
NA
NAB – The Numenta Anomaly Benchmark42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NAB – The Numenta Anomaly Benchmark

Hacker News15
I
I built an open-source benchmark that evaluates LLMs through gameplay55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source benchmark that evaluates LLMs through gameplay

Hacker News2