We

We scored 50k PRs with AI – what we learned about code complexity

Hacker News

We scored 50k PRs with AI – what we learned about code complexity

I'm a CTO with a ~16-person engineering team. Last year I wanted real data on what was actually shipping, not guesswork or story point theater. So we built GitVelocity. Every merged PR gets scored 0–100 by Claude across six dimensions: scope (0–20), architecture (0–20), implementation (0–20), risk (0–20), quality (0–15), perf/security (0–5). Six dimensions added up, then scaled by change size — a 10-line fix scores lower than a 500-line refactor even at the same complexity. Full formula at gitvelocity.dev/scoring-guide. After scoring 50,000+ PRs across TypeScript, Python, Rust, Go, Java, Elixir, and more, some things surprised us: Big PRs don't automatically score high. An 800-line migration with low complexity scores worse than a 200-line architectural change. Size gets you the full multiplier, but the base score still has to earn it. You can't score well without tests. The quality dimension (0–15) won't give you points without test coverage. At similar experience levels, this was the clearest separator between engineers. Juniors started outscoring some seniors. They adopted AI tools faster and took on harder problems. Once they could see their own scores, they aimed higher. We score AI-generated code the same as human-written code. Code is code. An engineer who uses AI to ship more complex work faster is more productive, and their scores reflect that. Scoring consistency was the hardest technical problem. Without reference examples anchoring each dimension, Claude's scores drifted 15+ points between runs. With 18 calibrated anchors (three per dimension at low/mid/high), we got it down to 2–4 points on the same PR. The thing we didn't expect was behavioral. We call it the Fitbit effect — the tool doesn't make you ship better code, but seeing the score does. Engineers started referencing their own scores in 1:1s unprompted, because the numbers matched what they already felt about their work. A junior who shipped a tricky concurrency fix could point to a score that proved it wasn't "just a small PR." We recently added team benchmarks (gitvelocity.dev/demo/benchmarks). Once you're scoring PRs, you can see how your team compares to others across the dataset — about 1,000 engineers on 60 teams so far. Headline's team ships faster than roughly 95% of them, which was nice to confirm but also made us wonder who the other 5% are. The competitive angle surprised us: teams that were skeptical about individual scores got genuinely curious once they could measure themselves against the field. Every score is fully visible to the engineer who wrote the PR, with per-dimension breakdowns and reasoning. There's no hidden dashboard that management sees and engineers don't. Free, BYOK (your Anthropic API key). We default to Sonnet 4.6, which scores nearly as well as Opus 4.6 at a fraction of the cost — but you can switch models if you want. Pennies per PR either way. No source code stored, diffs analyzed and discarded. Works with GitHub, GitLab, and Bitbucket. Ask me anything about the scoring methodology, how we solved calibration, or what it was actually like rolling this out to a team.

Share card

Actual performance

11points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, models · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started, para · Missing: supports, reddit linkedin, podcasting
86%86% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, 000, io · Missing: https docs, excited, just released
53%53% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
35%35% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way, para · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
22%22% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

De
Decrease system complexity by separating Code from Data58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Decrease system complexity by separating Code from Data

Hacker News1
AIUR
AIUR73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Automation without the complexity.

Indie Hackers1ai
Fa
Fallow – Find unused code, duplication, and complexity in TS/JS (Rust)51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fallow – Find unused code, duplication, and complexity in TS/JS (Rust)

Hacker News4
Ta
Taking the complexity out of timezone coordination53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Taking the complexity out of timezone coordination

Hacker News2
Me
Measure code by Kolmogrov complexity, not lines52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Measure code by Kolmogrov complexity, not lines

Hacker News6
SeekSimpler
SeekSimpler35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Complexity into simplicity

Product Hunt+9
Py
Pylon – Sentry Errors to PRs via Claude Code, with Telegram Approval21%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pylon – Sentry Errors to PRs via Claude Code, with Telegram Approval

Hacker News2
Sl
Sloc Cloc & Code a fast code counter w complexity calculations46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Sloc Cloc & Code a fast code counter w complexity calculations

Hacker News11
Ty
TypeScript Complexity Tracer55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TypeScript Complexity Tracer

Hacker News3
Syncaut
Syncaut30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Automate your agency workflows in minutes, no code, no complexity.

AppSumo