Wi

Will Jev pull the lever in the trolley problem?

Hacker News

Will Jev pull the lever in the trolley problem?

The Jev benchmark you've all been waiting for: Will Jev pull the lever in the trolley problem? Don't think too much about it :) Place whatever you'd like on each side of the tracks, and GPU (my startup's AI system, whee!) will have Jev judge the result. As a bonus, two really good looking SVGs are rendered with gpt-oss-120b for what you place on the tracks. I've made a really high quality thing here. My first impression is that Jev does well! I understand it's not exactly designed for this kind of "classification" -- I think -- but I haven't been able to get it to pull the lever away from any villains or geese yet. It's public against my better judgement. No account necessary, have fun. Side presentation: I've put together a more "serious" harness here if you're interested in evaluating Jev for yourself. I've been interested in seeing how well it does if you augment an existing smaller model with it: https://gpu.studio/jev

Share card

Actual performance

8points
6comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
75%75% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, lua, existing · Missing: https docs, excited, just released
59%59% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Go
Go-dutchflag – An implementation of the Dutch flag problem in Golang37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go-dutchflag – An implementation of the Dutch flag problem in Golang

Hacker News3
Th
The AZ Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The AZ Problem

Hacker News2
So
Someone Else's Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Someone Else's Problem

Hacker News1
Th
The farmer, wolf, goat and cabbage problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The farmer, wolf, goat and cabbage problem

Hacker News1
Mo
Monty Hall Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Monty Hall Problem

Hacker News1
Wo
Wolf, Goat and Cabbage Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wolf, Goat and Cabbage Problem

Hacker News1
Gu
Guiderail – How I (partly) solved my procrastination problem46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Guiderail – How I (partly) solved my procrastination problem

Hacker News3
Ar
Are the Riemann Hypothesis and Navier-Stokes the Same Problem?56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Are the Riemann Hypothesis and Navier-Stokes the Same Problem?

Hacker News7
Th
The problem with the epsilon greedy method33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The problem with the epsilon greedy method

Hacker News22
Grubl
Grubl65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The "What's for dinner" problem solved!

Indie Hackerscommitment-full-time