Talks

AI Interpretability in the Era of LLMs: Architecture, Behavior and Beyond, Section 2 (Generative Interpretability)

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA · Spring 2026

In this section, we turn to generative interpretability, exploring how large language models can be understood and steered through the lens of their generative behavior rather than solely their internal structure.

AI Interpretability in the Era of LLMs: Architecture, Behavior and Beyond, Section 1 (Post-hoc Interpretability)

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA · Spring 2026

In this section, we examine post-hoc interpretability methods for large language models, covering techniques that explain already-trained models’ architectures and behavior after the fact.

A Tutorial of Interpretable and Biologically Plausible LLMs, Section 4

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA ·

In this section, we shift our focus from structural interpretability to behavioral interpretability by introducing Test-Time Compute (System 2 Thinking), Neuro-Symbolic Reasoning, and Competitive Market-Based Learning, and illustrating how structural and behavioral interpretability form a coherent and complementary framework.

A Tutorial of Interpretable and Biologically Plausible LLMs, Section 3

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA ·

In this section, we explore model-level interpretability by introducing Transformer Circuit Theories, outlining their mathematical foundations, pinpointing how current models benefit from these principles, and illustrating how such insights inspire the development of modular and sparse next-generation AI architectures.

A Tutorial of Interpretable and Biologically Plausible LLMs, Section 2

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA ·

In this section, we explore neuron-level interpretability by introducing a biologically plausible neural model — the Spiking Neural Network (SNN) — illustrating how the incorporation of temporal dynamics enhances the expressiveness of individual neurons, and highlighting the delicate balance between biological plausibility and computational efficiency that enables the scalability of spiking neurons in large language models.

A Tutorial of Interpretable and Biologically Plausible LLMs, Section 1

Tutorial · University of Illinois Urbana-Champaign, Urbana, IL, USA ·

In this section, we introduce a conceptual overview of AI interpretability and argue for its vital role in shaping next-generation LLM architectures, drawing insights from research in artificial neural networks, complex systems, and philosophy.