Back to Analyze

Prediction methodology

How we estimate launch success probability and expected engagement from a product name, tagline, and description.

Training data

Each platform has its own independent model, because success signals differ significantly across them. All data is collected directly from each platform's public listings.

150,764
Hacker News posts
36,099
Indie Hackers products
25,636
Product Hunt posts
3,300
TrustMRR startups
2,459
AppSumo deals
1,400
BetaList startups
1,193
Acquire.com listings
256,000+
Total labeled outcomes

Positive labels vary by platform:

  • Product Hunt - made the daily leaderboard (featured)
  • Indie Hackers - generates revenue (MRR > $0)
  • AppSumo - above-median review count
  • Hacker News - above-median Show HN points
  • TrustMRR - generates revenue (MRR > $0)
  • Acquire.com - has verified annual revenue
  • BetaList - accepted and listed by BetaList

What we analyze

The only inputs are the product name, tagline, description, and category. No metadata, images, pricing, founder history, or external signals are used - the model evaluates the pitch on its own merits.

Text is converted into a numeric representation that captures which words and phrases are most informative across the training corpus. The vocabulary is built from training data and frozen; at inference time your input is projected into that same space.

Success classifier

Each platform has an independent ML classifier trained to predict whether a launch will reach its platform's success threshold. Models are trained with balanced class weights to handle the natural imbalance between successful and unsuccessful launches.

The output is a calibrated probability between 0 and 1. Raw model scores are post-processed to align with empirical success rates - a score of 70% is calibrated to reflect that roughly 70% of similar products historically succeeded on that platform.

A secondary universal model trained across all platforms helps generalize predictions, particularly for platforms with smaller training sets. Its signal is blended with the platform-specific score.

Top signals shown in the UI are words and phrases the model learned correlate most strongly with successful launches on that specific platform.

Engagement estimates

For platforms with quantitative outcomes (Product Hunt votes and comments, Indie Hackers MRR, TrustMRR MRR and customer count), separate regression models estimate expected magnitude. These are trained only on the successful subset of each platform.

TrustMRR uses a two-stage approach: first predicting whether a product will generate revenue at all, then estimating how much - conditioned on the first stage being positive.

Estimates are presented as data-driven ranges based on historical percentiles, not point predictions.

Inference

Models run entirely in-browser. No data is sent to external APIs, third-party ML services, or our servers - your product description never leaves your device.

Limitations

The model only sees text. Product quality, timing, founder reputation, pricing, and distribution are not captured. A product with a mediocre description can still succeed, and a well-worded one does not guarantee it.

Training data covers multi-year archives for all sources. Language trends evolve; model accuracy will degrade over time without retraining.

Similarity search uses cosine similarity over the same TF-IDF space. It finds textually similar products, not market competitors. A product in a different domain with similar vocabulary may rank higher than a closer competitor with different phrasing.