An

An AI eval based on a silly joke from an underrepresented language

Hacker News

An AI eval based on a silly joke from an underrepresented language

Marathi is an Indian language with 83 million speakers, but it's underrepresented as text online. There's a silly joke every Marathi-speaking kid learns: kapus kondyachi goshta (the story of the kapus konda). Jokes like this spread orally, not through text. It's not a real joke. There's no punchline. It's pure infinite-loop trolling—the kind of thing kids use to annoy each other or adults use to tease children. Someone asks: "Can I tell you the story of the kapus konda?" You say yes, no, whatever. Doesn't matter. There is no story. Your answer gets echoed back, and the question repeats. Forever. "No." "What do you mean 'no'? Can I tell you the story of the kapus konda?" "Fine, tell me." "What do you mean 'fine, tell me'? Can I tell you the story of the kapus konda?" That's it. That's the whole joke. I turned this into an AI eval: https://kapuskonda.vercel.app The words "kapus konda" mean nothing coherent, at least AFAIK, although kapus = cotton, konda = bran. So models that don't know the joke try to make sense of it. They hallucinate elaborate stories. I tested 31 models two ways: recognizing the joke when someone initiates it, and performing the joke themselves. None of them got it. Bonus: with web search enabled, Claude Opus 4.5 (on Claude.ai) passed. The gap is real, but retrieval helps. All prompts, responses, and scoring visible on the site. Feedback welcome. This is my first eval and I'm sure there's stuff I got wrong. Also curious: does your language/culture have a something like this that would make for a good eval?

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
58%58% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
48%48% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: io · Missing: https docs, excited, just released
44%44% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
33%33% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Sh
Shade, a language based on C, Go, and Haskell77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Shade, a language based on C, Go, and Haskell

Hacker News15
Pi
Pico, a small register-based scripting language I wrote in C49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pico, a small register-based scripting language I wrote in C

Hacker News2
Em
Emmetp a template language based on emmet49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Emmetp a template language based on emmet

Hacker News1
AI
AI Simulaionen Based on FEP51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Simulaionen Based on FEP

Hacker News4
La
Language proposal - dave60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Language proposal - dave

Hacker News24
Th
The Termcat Markup Language60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Termcat Markup Language

Hacker News1
20
2048 in PDP-8 Assembly language56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

2048 in PDP-8 Assembly language

Hacker News2
C#
C# language kernel for IPython55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

C# language kernel for IPython

Hacker News2
Je
JedLang, my first attempt at a language57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

JedLang, my first attempt at a language

Hacker News59
A
A brainfuck language kernel for IPython and Jupyter45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A brainfuck language kernel for IPython and Jupyter

Hacker News4