Ad

AdaptGauge – I found that adding few-shot examples can make LLMs worse

Hacker News

AdaptGauge – I found that adding few-shot examples can make LLMs worse

I tested 8 LLMs across 4 tasks at different few-shot counts (0, 1, 2, 4, 8) and found three patterns where adding examples actively degrades performance: 1. Peak regression: Gemini 3 Flash scored 64% at 4-shot, then crashed back to 33% at 8-shot 2. Ranking reversal: The zero-shot leader dropped to third once examples were added 3. Selection method matters: Switching from hand-picked to TF-IDF examples collapsed a model from 50%+ to 35% This aligns with recent research (Tang et al. 2025 "over-prompting", NDSS 2025 vulnerability detection drops, Chroma Research "context rot"). I built AdaptGauge to detect these patterns automatically. It tracks learning curves across shot counts and flags collapse with pattern classification (immediate, gradual, peak regression). Open source, MIT licensed. Pre-computed demo results included so you can see the patterns without API keys. Article with full results: https://shuntaro-okuma.medium.com/when-more-examples-make-yo... Repo: https://github.com/ShuntaroOkuma/adapt-gauge-core

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, context, tasks · Missing: mac, agents, macos
82%82% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
61%61% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: open source, io · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
48%48% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
3%3% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Te
Tested 12 LLMs with few-shot examples57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tested 12 LLMs with few-shot examples

Hacker News2
Re
RegexGo, get regex by providing examples44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RegexGo, get regex by providing examples

Hacker News1
Re
RegexGo – Regex Generator from Examples40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RegexGo – Regex Generator from Examples

Hacker News8
Si
Simple XState Examples53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Simple XState Examples

Hacker News4
One Shot Keto
One Shot Keto33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One Shot Keto is a dietary supplement that naturally stimula

Indie Hackerscommitment-full-time
Fr
Freeter dashboard examples58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Freeter dashboard examples

Hacker News1
GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2
Ma
Make your own end2end platform for LLMs in under 4 minutes62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Make your own end2end platform for LLMs in under 4 minutes

Hacker News2
Co
Counting 1121 BigData and MachineLearning Framework Toolset and Examples49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Counting 1121 BigData and MachineLearning Framework Toolset and Examples

Hacker News3