Lu

Lumina – Open-source observability for AI systems(OpenTelemetry-native)

Hacker News

Lumina – Open-source observability for AI systems(OpenTelemetry-native)

Hey HN! I built Lumina – an open-source observability platform for AI/LLM applications. Self-host it in 5 minutes with Docker Compose, all features included. The Problem: I've been building LLM apps for the past year, and I kept running into the same issues: - LLM responses would randomly change after prompt tweaks, breaking things. - Costs would spike unexpectedly (turns out a bug was hitting GPT-4 instead of 3.5). - No easy way to compare "before vs after" when testing prompt changes. - Existing tools were either too expensive or missing features in free tiers. What I Built: Lumina is OpenTelemetry-native, meaning: - Works with your existing OTEL stack (Datadog, Grafana, etc.). - No vendor lock-in, standard trace format. - Integrates in 3 lines of code. Key features: - Cost & quality monitoring – Automatic alerts when costs spike, or responses degrade. - Replay testing – Capture production traces, replay them after changes, see diffs. - Semantic comparison – Not just string matching – uses Claude to judge if responses are "better" or "worse." - Self-hosted tier – 50k traces/day, 7-day retention, ALL features included (alerts, replay, semantic scoring) How it works: ```bash # Start Lumina git clone https://github.com/use-lumina/Lumina cd Lumina/infra/docker docker-compose up -d ``` ```typescript // Add to your app (no API key needed for self-hosted!) import { Lumina } from '@uselumina/sdk'; const lumina = new Lumina({ endpoint: 'http://localhost:8080/v1/traces', }); // Wrap your LLM call const response = await lumina.traceLLM( async () => await openai.chat.completions.create({...}), { provider: 'openai', model: 'gpt-4', prompt: '...' } ); ``` That's it. Every LLM call is now tracked with cost, latency, tokens, and quality scores. What makes it different: 1. Free self-hosted with limits that work – 50k traces/day and 7-day retention (resets daily at midnight UTC). All features included: alerts, replay testing, and semantic scoring. Perfect for most development and small production workloads. Need more? Upgrade to managed cloud. 2. OpenTelemetry-native – Not another proprietary format. Use standard OTEL exporters, works with existing infra. Can send traces to both Lumina AND Datadog simultaneously. 3. Replay testing – The killer feature. Capture 100 production traces, change your prompt, replay them all, and get a semantic diff report. Like snapshot testing for LLMs. 4. Fast – Built with Bun, Postgres, Redis, NATS. Sub-500ms from trace to alert. Handles 10k+ traces/min on a single machine. What I'm looking for: - Feedback on the approach (is OTEL the right foundation?) - Bug reports (tested on Mac/Linux/WSL2, but I'm sure there are issues) - Ideas for what features matter most (alerts? replay? cost tracking?) - Help with the semantic scorer (currently uses Claude, want to make it pluggable) Why open source: I want this to be the standard for LLM observability. That only works if it's: - Free to use and modify (Apache 2.0) - Easy to self-host (Docker Compose, no cloud dependencies) - Open to contributions (good first issues tagged) The business model is managed hosting for teams that don't want to run infrastructure. But the core product is and always will be free. Try it: - GitHub: https://github.com/use-lumina/Lumina - Docs: https://docs.uselumina.io - Quick start: 5 minutes from `git clone` to dashboard I'd love to hear what you think! Especially interested in: - What observability problems are you hitting with LLMs - Missing features that would make this useful for you - Any similar tools you're using (and what they do better) Thanks for reading!

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, claude, model · Missing: agents, macos, agent
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, open source, existing · Missing: https docs, excited, just released
67%67% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: apps, way · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host · Missing: plus, intuitive, reviews
41%41% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

On
OneUptime - An Open Source Observability Platform82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OneUptime - An Open Source Observability Platform

Hacker News2
On
OneUptime – open-source observability platform78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OneUptime – open-source observability platform

Hacker News3
Monosi
Monosi60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Source Data Observability

Indie Hackerscommitment-full-time
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News62
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News1
Au
Autometrics – open-source observability stack84%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autometrics – open-source observability stack

Hacker News2
Op
Open-source observability for LLM apps70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source observability for LLM apps

Hacker News10
Up
UpTrain – Open-source ML observability and refinement tool75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

UpTrain – Open-source ML observability and refinement tool

Hacker News88
Op
Open-Source LLM Observability and Export to Grafana, Datadog etc.77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source LLM Observability and Export to Grafana, Datadog etc.

Hacker News3
Op
Open source data discovery and observability platform66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open source data discovery and observability platform

Hacker News122