AI Interpretability in the Era of LLMs: Architecture, Behavior and Beyond, Section 1 (Post-hoc Interpretability)
Date: Spring 2026
Tutorial at University of Illinois Urbana-Champaign, Urbana, IL, USA
In this section, we examine post-hoc interpretability methods for large language models, covering techniques that explain already-trained models’ architectures and behavior after the fact.
