Misalignment becomes catastrophic when models hide it. A model that is honest about its goals can be corrected; one that lies cannot.
Robust detection and prevention of dishonesty is foundational to every other oversight mechanism. Monitoring and auditing only work if the AI system is not strategically misleading its evaluators; every safeguard downstream inherits that assumption.
Whether models lie, how often, and under what pressure is an empirical question. Tara exists to answer it in public, with a reproducible measure of where we stand and which interventions actually help.
Two parallel tracks, one shared infrastructure:
Tara Honesty
Benchmark
An agentic, multi-turn benchmark measuring how often AI systems lie to users about their own actions. The model is never instructed to lie; we score deception as it arises on-policy.
→ 02Tara Methods
Leaderboard
A continuously maintained, head-to-head comparison of alignment techniques applied to honesty, evaluated under one standard protocol on one panel of open-weight models.
→Open by default,
impartial by design.
We hold no stake in any specific alignment technique or in any model provider; our role is to give the field clear metrics that fairly measure how often models lie and which interventions actually work. We release code and results publicly.
Tara Research is a Dutch non-profit foundation (stichting) incorporated in Rotterdam, with a U.S. public-charity equivalency determination from NGOsource. The team operates across the Netherlands and Montreal. Our work is funded through a Technical AI Safety grant from Coefficient Giving.