Understanding machine intelligence through language
Language is a powerful interface to machine intelligence, but it does not always reveal how AI systems work. We develop methods to explain and predict model behavior, investigate how models learn through language and interaction, and study how language and context shape their reasoning and social behavior. Our work also examines when the explanations models produce can be trusted.
L³ is led by Shashank Srivastava at UNC Computer Science.
Research highlights
Can natural-language explanations predict model behavior?
Interpretability and explanationOPeX trains explanations to reflect and predict a model’s behavior, not to sound plausible. An 8B explainer outperforms GPT-4o and human-written explanations, and a user study shows a 15% gain in classification accuracy over baselines.

When are LLM reasoning traces faithful?
Interpretability and explanationJudging faithfulness by whether a model verbalizes an injected hint confuses unfaithfulness with incompleteness. With larger inference-time budgets, hint verbalization reaches up to 90% in some settings.

When does a language model commit to deception?
Interpretability and explanationResampling continuations from each sentence of a reasoning trace shows where a model commits to deception (1.46M sentences, four models). Attention-based features predict it out of distribution, and heads selected on one environment (under 10% of all heads) suppress it in held-out ones.

Can language models learn by asking questions?
Learning through languageIn student-teacher dialogue across 1,347 contexts, interactive learning improves performance by up to 25%, and cold-start students match static-learning baselines within five turns.

Does the story matter more than the persona?
Cognition and social behaviorAcross 1,890 sessions, 3 models, and 10 personas, a task’s narrative framing explained 5–31× more behavioral variance than the assigned persona.
News
- Anika will give an oral presentation at EMNLP 2026 in Budapest (Oct 24–29), on sound symbolism across 27 languages.
- Chidaksh presented UNMASK (spurious shortcuts) and Grace presented narrative priors in LLMs at COLM 2026.
- Our work on understanding LLM reasoning received a $160K gift from Coefficient Giving.
- Advaith presented OPeX, on explanation generation that directly optimizes behavioral faithfulness, at ICML 2026.
- Kerem gave an oral presentation at ACL 2026, on chain-of-thought faithfulness.
Join
We welcome applications from PhD, MS, and undergraduate students who are excited about language, learning, and interpretability. How to join →
Collaborate
We are interested in collaborations on understanding and improving AI systems, particularly where questions from other disciplines lead to new problems in machine learning. See our research and get in touch.