Phare: A Safety Probe for Large Language Models
Phare: A Safety Probe for Large Language Models
We've just published a benchmark and accompanying paper on arXiv that challenges conventional leaderboard-driven LLM evaluation. Phare focuses on factual reliability, prompt sensitivity, multilingual support, and how models handle false premises like issues that actually matter when you're building serious applications. Some insights: - Preference scores ≠ factual correctness. - Framing effects can cause models to miss obvious falsehoods. - Safety metrics like sycophancy and stereotype reproduction show surprising results across popular models. Would love feedback from the community.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
BadSeek – How to backdoor large language models
Explore large language models with 512MB of RAM
TEG, a linguistic game powered by large language models
Chat with private and local large language models
A tool to give large language models better memory
Observability for teams building with large language models.
Tidbits, use large language models to filter through news
An all-in-one blog for learning Large Language Models (LLMs)
Open 3B language models from AMD
The rise of open source large language models