Be

Benzi – Code Intelligence Infrastructure for Frontier AI Models

Hacker News

Benzi – Code Intelligence Infrastructure for Frontier AI Models

Roughly speaking, the way current AI coding agents/harnesses work is by either: a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or b) Parsing code to make high dimensional embeddings to approximate a symptom map, and hand that to the agent. Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping. Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check. Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (Benchmark details: https://benzi.fly.dev/benchmark ) "But what if the compiler isn't doing its job right! Wouldn't you mislead the AI model?" - Absolutely. Benzi meticulously takes care of this by having 3 truth tiers. RESOLVED has definite evidence, CANDIDATE is what couldn't be resolved by the static analysis, and OBSERVED is what actually happened during an execution. The artificial intelligence and the determinstic intelligence layers coordinate to reduce source hits where possible, without producing incorrect results for the sake of efficiency. It also has several bonus features such as a runtime tracer, self-aware model upgrade mid task if it thinks the job is over its pay grade, context aware model written repro, and SEVERAL more. It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first. On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is notable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities. If you're curious to learn more, click https://benzi.fly.dev/about and check out StallionSwipe. probably the best thing i ever made. It's a Fireship inspired horse tinder app greenfielded entirely in Benzi Opus 4-8 and a little bit v4 flash. and lastly, please star on github if you like where this is headed!

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports · Missing: reddit linkedin, podcasting, created
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
57%57% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
40%40% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
22%22% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ailin¹
Ailin¹40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Collective Intelligence: The Next Frontier of AI

Indie Hackers1ai
Re
Replicover – Find the hottest AI models on Replicate31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Replicover – Find the hottest AI models on Replicate

Hacker News1
WisGate
WisGate62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Models PAI

Indie Hackerscommitment-full-time
Ti
Tinx.ai – Prebuilt AI Models for You33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tinx.ai – Prebuilt AI Models for You

Hacker News7
Un
Unobin compiles Infrastructure as Code to one binary62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unobin compiles Infrastructure as Code to one binary

Hacker News20
CometAPI
CometAPI81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

All AI Models in One API

Product Hunt+6
Mo
ModelAtlas – Find AI models that HuggingFace search can't48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ModelAtlas – Find AI models that HuggingFace search can't

Hacker News1
Ai
Aisir – AI models deliberate and critique each other like a council34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Aisir – AI models deliberate and critique each other like a council

Hacker News3
Ye
Yelp for AI Models/Products24%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Yelp for AI Models/Products

Hacker News2
Test AI Models
Test AI Models72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Compare AI models side-by-side on same prompt

Indie Hackers1$9/moai