AI Interpretability in the Era of LLMs: Architecture, Behavior and Beyond, Section 1 (Post-hoc Interpretability)

Date: Spring 2026

Tutorial at University of Illinois Urbana-Champaign, Urbana, IL, USA

In this section, we examine post-hoc interpretability methods for large language models, covering techniques that explain already-trained models’ architectures and behavior after the fact.

Slides

Recording