Te

Tested 12 LLMs with few-shot examples

Hacker News

Tested 12 LLMs with few-shot examples

I evaluated 12 models (6 cloud, 6 local) across 5 tasks at shot counts 0, 1, 2, 4, and 8, with 3 trials each. 60 model-task pairs, 27k+ evaluations. Three patterns stood out: 1. Few-shot can cause collapse: Gemini 3 Flash scored 93% at zero-shot on route optimization, then crashed to 30% at 8-shot. Same model family (Gemma 3 27B, local) stayed stable at 90%. 2. Most models benefit from few-shot: On classification, all models scored 0-20% at zero-shot. At 8-shot, scores spread from 27% to 80%. Zero-shot benchmarks would have led to the wrong model choice. 3. Task mismatch ≠ collapse: Reasoning-specialized models scored low on summarization regardless of shot count. They're not "collapsing" — they're just not suited for the task. A 27B local model (Gemma 3) matched Claude Haiku's adaptation efficiency (AUC 0.814 vs 0.815). The 12-model results are included as default demo data — explore the patterns without API keys. Article: https://dev.to/shuntarookuma/i-tested-12-llms-with-few-shot-... GitHub (MIT): https://github.com/ShuntaroOkuma/adapt-gauge-core

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
80%80% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
67%67% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, io · Missing: https docs, excited, just released
57%57% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
35%35% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
31%31% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
6%6% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ad
AdaptGauge – I found that adding few-shot examples can make LLMs worse48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AdaptGauge – I found that adding few-shot examples can make LLMs worse

Hacker News1
Re
RegexGo, get regex by providing examples44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RegexGo, get regex by providing examples

Hacker News1
PJ
PJON 12.0 Communications bus system for Arduino47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PJON 12.0 Communications bus system for Arduino

Hacker News1
Be
Beating GCC 12 – 118x Speedup for Jensen Shannon D. Via AVX-512FP1640%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Beating GCC 12 – 118x Speedup for Jensen Shannon D. Via AVX-512FP16

Hacker News8
Do
DocsGPT 0.12.040%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DocsGPT 0.12.0

Hacker News1
UM
UML model and code examples of GoF design patterns in 12 languages57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

UML model and code examples of GoF design patterns in 12 languages

Hacker News1
Re
RegexGo – Regex Generator from Examples40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RegexGo – Regex Generator from Examples

Hacker News8
Si
Simple XState Examples53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Simple XState Examples

Hacker News4
One Shot Keto
One Shot Keto33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One Shot Keto is a dietary supplement that naturally stimula

Indie Hackerscommitment-full-time
Fr
Freeter dashboard examples58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Freeter dashboard examples

Hacker News1