Portfolio item number 1
Short description of portfolio item number 1
Short description of portfolio item number 1
Short description of portfolio item number 2 
A comprehensive, interdisplinary survey of interpretable and biologically plausible LLMs based on my tutorial delivered at CS 591 BAI seminar. See Talks for more information.
Memory is central to autonomous, self-evolving agents. Leveraging Integrated Information Theory (IIT), we design time-variant memory that integrates past experience via event-driven updates gated by reward, effectively treating consolidation as learning in a standard RL framework.
Learn the LLM’s computational topology instead of hand-engineering it: We propose a structure predictor adjoint to a LLM, which generates a context-dependent computational graph applied to the LLM. This allows the modular structure to emerge with minimal contraints during training.
Published:
In this section, we introduce a conceptual overview of AI interpretability and argue for its vital role in shaping next-generation LLM architectures, drawing insights from research in artificial neural networks, complex systems, and philosophy.
Published:
In this section, we explore neuron-level interpretability by introducing a biologically plausible neural model — the Spiking Neural Network (SNN) — illustrating how the incorporation of temporal dynamics enhances the expressiveness of individual neurons, and highlighting the delicate balance between biological plausibility and computational efficiency that enables the scalability of spiking neurons in large language models.
Published:
In this section, we explore model-level interpretability by introducing Transformer Circuit Theories, outlining their mathematical foundations, pinpointing how current models benefit from these principles, and illustrating how such insights inspire the development of modular and sparse next-generation AI architectures.
Published:
In this section, we shift our focus from structural interpretability to behavioral interpretability by introducing Test-Time Compute (System 2 Thinking), Neuro-Symbolic Reasoning, and Competitive Market-Based Learning, and illustrating how structural and behavioral interpretability form a coherent and complementary framework.
Published:
In this section, we examine post-hoc interpretability methods for large language models, covering techniques that explain already-trained models’ architectures and behavior after the fact.
Published:
In this section, we turn to generative interpretability, exploring how large language models can be understood and steered through the lens of their generative behavior rather than solely their internal structure.
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.