Us

Using LLMs and Embeddings to classify application errors

Hacker News

Using LLMs and Embeddings to classify application errors

Hi Hacker News! We’re Vadim and Chris from Highlight.io [1]. We do web app monitoring and are working on using LLMs/embeddings to add new functionality to our error monitoring product. Given that there’s a lot of founders/engineers using LLMs in their products, we figured we’d share how we built the new functionality, their impact on our workflows, and how you can try it out. Our goal was to build two features: (1) tagging errors (e.g. deeming an error as “authentication error” or a “database error”); and (2) grouping similar errors together (e.g. two errors that have a different stacktrace and body, but are semantically not very different). Each of these rely heavily on comparing text across our application. After some experimentation with the OpenAI embeddings API [3], we went ahead and hosted a private model instance of thenlper/gte-large (an open-source MIT licensed model), which is a 1024-dimension model running on an Intel Ice Lake 2 vCPU machine on Hugging face [4]. Our general approach for classifying/comparing text is as follows. As each set of tokens (i.e a string) comes in, our backend makes a request to an inference endpoint and receives a 1024-dimension float vector as a response (see the code here [5]). We then store that vector using pgvector [6]. To compare any two sets for similarity, we simply look at the Euclidian distance between their respective embeddings using the ivfflat index implemented by pgvector (example code here [7]). To tag errors, we assign an error its most relevant tag from a predetermined set decided by us. For example, if we tag an error as an "authentication error" or a "database error", we can allow developers to have a starting point before inspecting an issue.(see the logic here [8]). Anecdotally, this approach seems to work very well. For example, here are two authentication errors that got tagged as “Authentication Error”: * Firebase: A network AuthError has occurred * Error retrieving user from firebase api for email verification: cannot find user from uid. We also use these error embeddings to group similar errors. To decide whether an error joins a group or starts a new one, we decide on a distance threshold (using the euclidean distance) ahead of time. An interesting thing about this approach, compared to using a text-based heuristic, is that two errors with different stack traces can still be grouped together. Here’s an example: * github.com/highlight-run/highlight/backend/worker.(*Worker).ReportStripeUsage * github.com/highlight-run/highlight/backend/private-graph/graph.(*Resolver).GetSlackChannelsFromSlack.func1 Both reported as `integration api error` as they involve the Stripe and Slack integrations respectively. The neat thing is that the LLM can use the full context of an error and match based on the most relevant details about the error. We have rolled out a first version of the error grouping logic to our cloud product [9], and there’s a demo of all the functionality at [2]. Long-term, if the HN community has other ideas of what we could build with LLM tooling in observability, we’re all ears. Let us know what you think! Links [1] https://news.ycombinator.com/item?id=36774611 [2] https://app.highlight.io/error-tags [3] https://platform.openai.com/docs/guides/embeddings [4] https://huggingface.co/thenlper/gte-large [5] https://github.com/highlight/highlight/blob/main/backend/emb... [6] https://github.com/highlight/highlight/blob/main/backend/mod... [7] https://github.com/highlight/highlight/blob/main/backend/pub... [8] https://github.com/highlight/highlight/blob/main/backend/pri... [9] https://app.highlight.io

Share card

Actual performance

65points
10comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, slack · Missing: agents, macos, agent
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
80%80% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: hacker news, ide, io · Missing: https docs, excited, just released
80%80% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host · Missing: plus, intuitive, reviews
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Wi
Witness – A Web Application Using IPFS and Ethereum60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Witness – A Web Application Using IPFS and Ethereum

Hacker News10
Go
Golang Application monitoring using Prometheus27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Golang Application monitoring using Prometheus

Hacker News2
Mo
Monitor LLMs Using LangChain and Infino47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Monitor LLMs Using LangChain and Infino

Hacker News1
Ra
Raink – Document ranker using LLMs60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Raink – Document ranker using LLMs

Hacker News2
Ne
Neurooo – DeepL clone using LLMs for translations40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neurooo – DeepL clone using LLMs for translations

Hacker News18
St
StockTracker Application42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

StockTracker Application

Hacker News2
Ch
CheerpJ 1.1 – Compile any Java application to HTML557%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CheerpJ 1.1 – Compile any Java application to HTML5

Hacker News1
Cl
Clojure Cup Application - What The Fn49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Clojure Cup Application - What The Fn

Hacker News1
I
I opensourced my first Django/Angular Application35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I opensourced my first Django/Angular Application

Hacker News1
my
my gem to modularize Sinatra application37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

my gem to modularize Sinatra application

Hacker News3