SycoFact 4B: Open model detecting sycophancy and delusion confirmation
SycoFact 4B: Open model detecting sycophancy and delusion confirmation
I published a model you can use now to help detect sycophantic AI responses before they harm users. It rejects 100% of the sycophantic delusion affirming responses from psychosis-bench. It also does well on the AISI Harmful Advice, PKU-SafeRLHF, and safety subsets of RewardBench. It's small enough it can run on a gaming GPU locally. It's got a GGUF checkpoint on hugging face and is available on ollama. You can pull it and run scenarios against it in minutes: https://ollama.com/izzie/sycofact The synthetic training data is also public, you can train other models over the data or reproduce my results. The labels were all generated by Gemma 3 27B with activation steering based on generated contrastive data. A write-up is planned at a later date, feel free to get in touch if curious.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
7B open model mixing transformers and linear RNNs
Unifying agentic capabilities in one open model
Assist: an open-model-first work surface for agents
The first open model to beat Sonnet made for productivity
Tencent’s 770B open model for long-horizon work
The 1T parameter open model for agentic intelligence
A powerful open model for agentic coding tasks
GLM-4.7 is a open model as your coding partner.
A System Model of Western Civilisation
javscript model of Ackermann steering