Kr

KraspAI Kompass – keep up with new LLMs

Hacker News

KraspAI Kompass – keep up with new LLMs

Hello HN! I want share KraspAI Kompass, an LLM comparison tool I have been building. Things around LLMs are moving fast, and new models are constantly coming out. It's easy to get disoriented and I personally have had a hard time keeping up. Because I wanted to quickly understand the capabilities of any new model, I built Kompass: a tool that lets you create your own test suite (or suites!) of prompts to evaluate new models with. It allows you to examine and compare performance across models and to figure out which of them are any good for your desired task. For instance, Kompass makes it easy to compare - open-source vs. closed-source (e.g. https://app.krasp.ai/c/b/helpme-software-deve?models=openai-... ) - or US vs. rest of the world (e.g. https://app.krasp.ai/c/b/helpme-office-worker?models=openai-... ) Keen for any feedback! Are there features missing? Models you would like to be added? What use-cases you would interested in? I'd love to hear from you. Thanks a lot for taking the time, I appreciate your help and support!

Share card

Actual performance

1points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
94%94% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua · Missing: https docs, excited, just released
54%54% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
45%45% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: personal · Missing: mobile apps, ios, entrepreneurs
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2
Th
Thought Forgery, a new technique for jailbreaking LLMs44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Thought Forgery, a new technique for jailbreaking LLMs

Hacker News2
LL
LLMs can be susceptible to a new kind of malware65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLMs can be susceptible to a new kind of malware

Hacker News17
Ne
New.email – Building Emails with LLMs64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

New.email – Building Emails with LLMs

Hacker News11
A
A new benchmark for testing LLMs for deterministic outputs41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A new benchmark for testing LLMs for deterministic outputs

Hacker News60
My
My new homepage32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

My new homepage

Hacker News103
th
the new HN Kansai Meetup homepage39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

the new HN Kansai Meetup homepage

Hacker News2
Th
The new Square45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The new Square

Hacker News2
Ne
New from Ruby Jokes: taint_aliases40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

New from Ruby Jokes: taint_aliases

Hacker News1