Learning from Language Lab

Understanding machine intelligence through language

Language is a powerful interface to machine intelligence, but it does not always reveal how AI systems work. We develop methods to explain and predict model behavior, investigate how models learn through language and interaction, and study how language and context shape their reasoning and social behavior. Our work also examines when the explanations models produce can be trusted.

L³ is led by Shashank Srivastava at UNC Computer Science.

Explore our research →

Research highlights

  • Can natural-language explanations predict model behavior?

    Interpretability and explanation

    OPeX trains explanations to reflect and predict a model’s behavior, not to sound plausible. An 8B explainer outperforms GPT-4o and human-written explanations, and a user study shows a 15% gain in classification accuracy over baselines.

    OPeX · ICML 2026 Read the paper →

  • Figure from the paper: CoT faithfulness

    When are LLM reasoning traces faithful?

    Interpretability and explanation

    Judging faithfulness by whether a model verbalizes an injected hint confuses unfaithfulness with incompleteness. With larger inference-time budgets, hint verbalization reaches up to 90% in some settings.

    CoT faithfulness · ACL 2026Oral Read the paper →

  • Figure from the paper: Deceptive commitment

    When does a language model commit to deception?

    Interpretability and explanation

    Resampling continuations from each sentence of a reasoning trace shows where a model commits to deception (1.46M sentences, four models). Attention-based features predict it out of distribution, and heads selected on one environment (under 10% of all heads) suppress it in held-out ones.

    Deceptive commitment · arXiv 2026Preprint Read the paper →

  • Figure from the paper: INTERACT

    Can language models learn by asking questions?

    Learning through language

    In student-teacher dialogue across 1,347 contexts, interactive learning improves performance by up to 25%, and cold-start students match static-learning baselines within five turns.

    INTERACT · ACL 2025Oral Read the paper →

  • Figure from the paper: Narrative priors

    Does the story matter more than the persona?

    Cognition and social behavior

    Across 1,890 sessions, 3 models, and 10 personas, a task’s narrative framing explained 5–31× more behavioral variance than the assigned persona.

    Narrative priors · COLM 2026 Read the paper →

All publications →

News

  • Anika will give an oral presentation at EMNLP 2026 in Budapest (Oct 24–29), on sound symbolism across 27 languages.
  • Chidaksh presented UNMASK (spurious shortcuts) and Grace presented narrative priors in LLMs at COLM 2026.
  • Our work on understanding LLM reasoning received a $160K gift from Coefficient Giving.
  • Advaith presented OPeX, on explanation generation that directly optimizes behavioral faithfulness, at ICML 2026.
  • Kerem gave an oral presentation at ACL 2026, on chain-of-thought faithfulness.

All news →

Join

We welcome applications from PhD, MS, and undergraduate students who are excited about language, learning, and interpretability. How to join →

Collaborate

We are interested in collaborations on understanding and improving AI systems, particularly where questions from other disciplines lead to new problems in machine learning. See our research and get in touch.