Be

Benchmarking VLMs vs. Traditional OCR

Hacker News

Benchmarking VLMs vs. Traditional OCR

Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year ( https://github.com/getomni-ai/zerox ). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https://getomni.ai/ocr-benchmark Github: https://github.com/getomni-ai/benchmark Huggingface: https://huggingface.co/datasets/getomni-ai/ocr-benchmark Couple notes on the methodology: 1. We are using JSON accuracy as our primary metric. The end goal is to evaluate how well each OCR provider can prepare the data for LLM ingestion. 2. This methodology differs from a lot of OCR benchmarks, because it doesn't rely on text similarity. We believe text similarity measurements are heavily biased towards the exact layout of the ground truth text, and penalize correct OCR that has slight layout differences. 3. Every document goes Image => OCR => Predicted JSON. And we compare the predicted JSON against the annotated ground truth JSON. The VLMs are capable of Image => JSON directly, we are primarily trying to measure OCR accuracy here. Planning to release a separate report on direct JSON accuracy next week. This is a continuous work in progress! There are at least 10 additional providers we plan to add to the list. The next big roadmap items are: - Comparing OCR vs. direct extraction. Early results here show a slight accuracy improvement, but it’s highly variable on page length. - A multilingual comparison. Right now the evaluation data is english only. - A breakdown of the data by type (best model for handwriting, tables, charts, photos, etc.)

Share card

Actual performance

146points
40comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: para, gemini · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, gemini · Missing: mac, agents, macos
63%63% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: para · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
52%52% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

40
401K Traditional vs. Roth Calculator42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

401K Traditional vs. Roth Calculator

Hacker News2
Cast I Ching
Cast I Ching23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Faithfully reproducing traditional yarrow stalk probabilitie

Indie Hackers1education
Ho
Homoglyph Attack Prevention with OCR59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Homoglyph Attack Prevention with OCR

Hacker News5
Traditional TEnt
Traditional TEnt30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Traditional Tent, where heritage craftsmanship meets modern

Indie Hackers
BB
BBC vs. Fox vs. CNN44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BBC vs. Fox vs. CNN

Hacker News10
We
WebTransport vs. WebRTC vs. WebSocket68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WebTransport vs. WebRTC vs. WebSocket

Hacker News33
Pr
Predator vs. Boids46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Predator vs. Boids

Hacker News2
Da
Darth Vader VS Disney pwning Vader OR Meh.44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Darth Vader VS Disney pwning Vader OR Meh.

Hacker News1
Om
Omegle vs Cleverbot44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Omegle vs Cleverbot

Hacker News2
In
Infographic Timmmmeee: DogVacay vs. Rover44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Infographic Timmmmeee: DogVacay vs. Rover

Hacker News2