Be

BenchFlow – run AI benchmarks as an API

Hacker News

BenchFlow – run AI benchmarks as an API

I built BenchFlow, an open-source framework that lets you integrate and evaluate AI tasks using Docker-based benchmarks. You can try it out right now by cloning the repo and running a benchmark in minutes. As an AI researcher, I was frustrated with how much time my team spent setting up benchmark environments rather than actually improving our models. We'd spend weeks configuring environments, only to find inconsistencies when comparing results with other teams. BenchFlow started as an internal tool to standardize our evaluation process, and we decided to open-source it after seeing how much time it saved us. Unlike other benchmarking tools that focus on specific domains, BenchFlow provides a unified interface for any AI task. The Docker-based approach ensures consistent environments across different machines and teams. You don't need to worry about dependency conflicts or environment setup - just implement a simple interface and you're ready to go. How to try it out? check our link but here's a preview of that 1. pip install benchflow 2. load a benchmark and define how to call your agents/models 3. run it and get the result Available benchmarks you can try today: - MMLU-PRO: Test your model's knowledge across 57 subjects - Bird: Evaluate business intelligence reasoning capabilities - WebArena: See how your agent performs on web-based tasks - MedQA-CS: Test medical question answering abilities The framework handles all the containerization, task distribution, and result collection, so you can focus on improving your models rather than managing infrastructure. I'd love to hear your feedback and see how you use it. What benchmarks would you like to see added next? Please give us a star if you can, thanks! GitHub: https://github.com/benchflow-ai/benchflow Website: https://benchflow.ai/ Benchmark Hub: https://benchflow.ai/benchmarks Inspo: https://github.com/ServiceNow/BrowserGym

Share card

Actual performance

24points
1comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, agents, agent · Missing: macos, cursor, claude
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
83%83% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
50%50% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: interface · Missing: plus, platform, intuitive
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Evoke
Evoke45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run AI on the cloud with our API

Indie Hackers18ai
Be
Benchmarks of UUID bintext codecs in Go34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarks of UUID bintext codecs in Go

Hacker News1
Cr
Crowdfunding forecasts and benchmarks31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowdfunding forecasts and benchmarks

Hacker News1
Ashera AI
Ashera AI70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GTM, Run by AI

Product Hunt+98
Re
Reproducible open-source STT API benchmarks with full methodology39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Reproducible open-source STT API benchmarks with full methodology

Hacker News1
Gr
GraphQL Server Benchmarks56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GraphQL Server Benchmarks

Hacker News1
LocalMode
LocalMode31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run AI entirely in your browser, no servers or API keys

Indie Hackerscommitment-side-project
EMAIL ROYALE
EMAIL ROYALE76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

War game, fought by email, run by an AI gamekeeper

Product Hunt+1
NobodyWho
NobodyWho74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run AI models on any device

Product Hunt+107Android
Ru
Run AI apps and dev environments with dstack68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run AI apps and dev environments with dstack

Hacker News1