KraspAI Kompass – keep up with new LLMs
KraspAI Kompass – keep up with new LLMs
Hello HN! I want share KraspAI Kompass, an LLM comparison tool I have been building. Things around LLMs are moving fast, and new models are constantly coming out. It's easy to get disoriented and I personally have had a hard time keeping up. Because I wanted to quickly understand the capabilities of any new model, I built Kompass: a tool that lets you create your own test suite (or suites!) of prompts to evaluate new models with. It allows you to examine and compare performance across models and to figure out which of them are any good for your desired task. For instance, Kompass makes it easy to compare - open-source vs. closed-source (e.g. https://app.krasp.ai/c/b/helpme-software-deve?models=openai-... ) - or US vs. rest of the world (e.g. https://app.krasp.ai/c/b/helpme-office-worker?models=openai-... ) Keen for any feedback! Are there features missing? Models you would like to be added? What use-cases you would interested in? I'd love to hear from you. Thanks a lot for taking the time, I appreciate your help and support!
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
GPTCache – Redis for LLMs
prompttest – pytest for LLMs
Thought Forgery, a new technique for jailbreaking LLMs
LLMs can be susceptible to a new kind of malware
New.email – Building Emails with LLMs
A new benchmark for testing LLMs for deterministic outputs
My new homepage
the new HN Kansai Meetup homepage
The new Square
New from Ruby Jokes: taint_aliases