Core work
COLM 2026
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence A systematic comparison of three steering-vector methods on Llama-3.3-70B and Qwen3.6-27B, across two threat models (dishonesty, dismissiveness). Introduces two new conditional methods that recover honesty without the capability collapse seen in unconditional steering. A single honesty direction generalizes out of distribution: it raises MASK scores, suppresses deception in multi-agent play (Among Us), doubles hidden-behavior discovery on AuditBench, and restores honesty in an emergently misaligned model.
arXiv →
Collaborations
ICML 2026
Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions Position paper showing that reasoning agents in simulated pricing markets drift toward tacit collusion, a tendency that persists even when the agents are explicitly prompted not to collude. It argues that behavioral certification should be required before AI systems make real market decisions.
OpenReview →
AAAI-26
Persistent Instability in LLMs' Personality Measurements: Effects of Scale, Reasoning, and Conversation History PERSIST, an evaluation framework spanning 25 open-weight models (1B–685B parameters) and over two million responses, showing that personality measurements shift under perturbations as small as question reordering, and that neither scale nor reasoning fixes it.
AAAI →
2026
We Built a Game Where Lying Has an Advantage. The Most Honest AI Won Anyway. A Minecraft deception game in which the informed agent has a payoff incentive to lie about which of four bridges is deadly. Frontier models diverged widely, with uninstructed deception rates ranging from 5% to 90%. An environment that reliably surfaces deception propensity is also a testbed for interventions: methods such as activation steering can be evaluated by whether the deception rate actually falls.
kradle.ai →