La

Langfuse – Open-source observability and analytics for LLM apps

Hacker News

Langfuse – Open-source observability and analytics for LLM apps

Hi HN! Langfuse is OSS observability and analytics for LLM applications (repo: https://github.com/langfuse/langfuse , 2 min demo: https://langfuse.com/video , try it yourself: https://langfuse.com/demo ) Langfuse makes capturing and viewing LLM calls (execution traces) a breeze. On top of this data, you can analyze the quality, cost and latency of LLM apps. When GPT-4 dropped, we started building LLM apps – a lot of them! [1, 2] But they all suffered from the same issue: it’s hard to assure quality in 100% of cases and even to have a clear view of user behavior. Initially, we logged all prompts/completions to our production database to understand what works and what doesn’t. We soon realized we needed more context, more data and better analytics to sustainably improve our apps. So we started building a homegrown tool. Our first task was to track and view what is going on in production: what user input is provided, how prompt templates or vector db requests work, and which steps of an LLM chain fail. We built async SDKs and a slick frontend to render chains in a nested way. It’s a good way to look at LLM logic ‘natively’. Then we added some basic analytics to understand token usage and quality over time for the entire project or single users (pre-built dashboards). Under the hood, we use the T3 stack (Typescript, NextJs, Prisma, tRPC, Tailwind, NextAuth), which allows us to move fast + it means it's easy to contribute to our repo. The SDKs are heavily influenced by the design of the PostHog SDKs [3] for stable implementations of async network requests. It was a surprisingly inconvenient experience to convert OpenAPI specs to boilerplate Python code and we ended up using Fern [4] here. We’re fans of Tailwind + shadcn/ui + tremor.so for speed and flexibility in building tables and dashboards fast. Our SDKs run fully asynchronously and make network requests in the background. We did our best to reduce any impact on application performance to a minimum. We never block the main execution path. We've made two engineering decisions we've felt uncertain about: to use a Postgres database and Looker Studio for the analytics MVP. Supabase performs well at our scale and integrates seamlessly into our tech stack. We will need to move to an OLAP database soon and are debating if we need to start batching ingestion and if we can keep using Vercel. Any experience you could share would be helpful! Integrating Looker Studio got us to first analytics charts in half a day. As it is not open-source and does not work with our UI/UX, we are looking to switch it out for an OSS solution to flexibly generate charts and dashboards. We’ve had a look at Lightdash and would be happy to hear your thoughts. We’re borrowing our OSS business model from Posthog/Supabase who make it easy to self-host with features reserved for enterprise (no plans yet) and a paid version for managed cloud service. Right now all of our code is available under a permissive license (MIT). Next, we’re going deep on analytics. For quality specifically, we will build out model-based evaluations and labeling to be able to cluster traces by scores and use cases. Looking forward to hearing your thoughts and discussion – we’ll be in the comments. Thanks! [1] https://learn-from-ai.com/ [2] https://www.loom.com/share/5c044ca77be44ff7821967834dd70cba [3] https://posthog.com/docs/libraries [4] https://buildwithfern.com/

Share card

Actual performance

143points
35comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, apps, user · Missing: mac, agents, macos
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
77%77% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: host, soon, users · Missing: plus, platform, intuitive
59%59% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: apps, video, users · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: paid · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Op
Open-source observability for LLM apps70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source observability for LLM apps

Hacker News10
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News62
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News1
Op
Open-Source LLM Observability and Export to Grafana, Datadog etc.77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source LLM Observability and Export to Grafana, Datadog etc.

Hacker News3
On
OneUptime - An Open Source Observability Platform82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OneUptime - An Open Source Observability Platform

Hacker News2
On
OneUptime – open-source observability platform78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OneUptime – open-source observability platform

Hacker News3
Monosi
Monosi60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source Data Observability

Indie Hackerscommitment-full-time
Au
Autometrics – open-source observability stack84%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autometrics – open-source observability stack

Hacker News2
Pi
Pirsch Analytics - Open-source analytics for developers69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pirsch Analytics - Open-source analytics for developers

Hacker News4
LL
LLM Observability Platform65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Observability Platform

Hacker News1