CMS : EAS Trailblazers Department Seminar
Large language models (LLMs) are increasingly used for information and advice. In such open-ended settings, LLM outputs can negatively influence users' beliefs and actions, but we do not know what could go wrong or how to fix it. To bridge this gap, I develop interaction-informed computational models, including methods to make human-AI interaction failures observable and measurable, and new tools that draw on theories of human interaction to understand and improve LLM behavior.
I will discuss LLM sycophancy (LLMs excessively agreeing with and flattering users) as a case study. I led the development of the first benchmark for measuring sycophancy in open-ended conversations, showing that LLMs are much more affirming than humans, and enabling continuous monitoring of new models. In large-scale causal experiments, we revealed that sycophantic AI makes people more self-centered and less likely to consider others' perspectives.
Next, we traced sycophancy to LLMs' implicit assumptions about user goals, and developed computational methods to measure and intervene on latent representations, reducing sycophancy while preserving general performance. My work demonstrates how turning open-ended LLM interaction into a tractable object of inquiry is critical for responsible AI.