He

Helicone (YC W23) – OSS LLM Observability and Development Platform

Hacker News

Helicone (YC W23) – OSS LLM Observability and Development Platform

Hey HN, we're Justin and Cole, the founders of Helicone ( https://helicone.ai ). Helicone is an open-source platform that helps teams build better LLM applications through a complete development lifecycle of logging, evaluation, experimentation, and release. You can try our free demo by signing up ( https://helicone.ai/signup ) or self-deploy with our new fully open-source helm chart ( https://helicone.ai/selfhost ). When we first launched 22 months ago, we focused on providing visibility into LLM applications. With just a single line of code, teams could trace requests and responses, track token usage, and debug production issues. That simple integration has since processed over 2.1B requests and 2.6T tokens, working with teams ranging from startups to Fortune 500 companies. However, as we scaled and our customers matured, it became clear that logging alone wasn’t enough to manage production-grade applications. Teams like Cursor and V0 have shown what peak AI application performance looks like and it's our goal to help teams achieve that quality. From speaking with users, we realized our platform was missing the necessary tools to create an iterative improvement loop - prompt management, evaluations, and experimentation. Helicone V1: Log → Review → Release (Hope it works) From talking with our users, we noticed a pattern: while many successfully launch their MVP quickly, the teams that achieve peak performance take a systematic approach to improvement. They identify inconsistent behaviors through evaluation, experiment methodically with prompts, and measure the impact of each change. This observation shaped our new workflow: Helicone V2: Log → Evaluate → Experiment → Review → Release It begins with comprehensive logging, capturing the entire context of an LLM application. Not just prompts and responses, but variables, chain steps, embeddings, tool calls, and vector DB interactions ( https://docs.helicone.ai/features/sessions ). Yet even with detailed traces, probabilistic systems are notoriously hard to debug at scale. So, we released evaluators (either via LLM-as-judge or custom Python evaluators leveraging the CodeSandbox SDK - https://codesandbox.io/docs/sdk/sandboxes ). From there, our users were able to more easily monitor performance and investigate what went wrong. Did the embedding search return poor results? Did a tool call fail? Did the prompt mishandle an edge case? But teams would still edit prompts in a playground, run a few test cases, and deploy based on intuition. This lacked the systematic testing we’re used to in traditional software development. That’s why we built experiments (similar to Anthropic's workbench but model-agnostic) ( https://docs.helicone.ai/features/experiments ). For instance, when a prompt generates occasional rude support responses, you can test prompt variations against historical conversations. Each variant runs through your production evaluators, measuring real improvement before deployment. Once deployed, the cycle begins again. We recognize that Helicone can’t solve all of the problems you might face when building an LLM application, but we hope that we can help you bring a better product to your customers through our new workflow. If you're curious how our infrastructure handled our growth: Our initial architecture struggled - synchronous log processing overwhelmed our database and query times went from milliseconds to minutes. We've completely rebuilt our infrastructure with two key changes: 1) using Kafka to decouple log ingestion from processing, and 2) splitting storage by access pattern across S3, Kafka, and ClickHouse. This was a long journey but resulted in zero data loss and fast query times even at billions of records. You can read about that here: https://upstash.com/blog/implementing-upstash-kafka-with-clo... We'd love your feedback and questions - join us in this HN thread or on Discord ( https://discord.gg/2TkeWdXNPQ ). If you're interested in contributing to what we build next, check out our GitHub.

Share card

Actual performance

29points
7comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: cursor, model, user · Missing: mac, agents, macos
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
94%94% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, clickhouse · Missing: https docs, excited, just released
78%78% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host, occasional · Missing: plus, intuitive, reviews
47%47% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: month, users · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: growth · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLM Observability Platform65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Observability Platform

Hacker News1
Ho
HolmesGPT – OSS AI Agent for On-Call and Observability46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

HolmesGPT – OSS AI Agent for On-Call and Observability

Hacker News2
Op
OpenMetadata – OSS platform for data discovery observability governance54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetadata – OSS platform for data discovery observability governance

Hacker News19
Th
Thredded forums are minimalistic and OSS62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Thredded forums are minimalistic and OSS

Hacker News5
Dy
Dynmgrm – Operate DynamoDB with GORM (Golang OSS)49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dynmgrm – Operate DynamoDB with GORM (Golang OSS)

Hacker News1
We
We Put Chromium on a Unikernel (OSS Apache 2.0)63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We Put Chromium on a Unikernel (OSS Apache 2.0)

Hacker News132
My
My First OSS as a Teen44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

My First OSS as a Teen

Hacker News4
Al
All you need is Prometheus and Jaeger for LLM Observability58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

All you need is Prometheus and Jaeger for LLM Observability

Hacker News3
La
Langtrace – OpenTelemetry-Based LLM App Observability62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Langtrace – OpenTelemetry-Based LLM App Observability

Hacker News9
La
Langtrace – OpenTelemetry Based LLM App Observability63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Langtrace – OpenTelemetry Based LLM App Observability

Hacker News2