<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://xiaocong-yang.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://xiaocong-yang.github.io/" rel="alternate" type="text/html" /><updated>2026-08-24T13:19:24+00:00</updated><id>https://xiaocong-yang.github.io/feed.xml</id><title type="html">Xiaocong Yang</title><subtitle>CS PhD student at UIUC, advised by Prof. ChengXiang Zhai. Team leader of AI Interpretability @ Illinois.</subtitle><author><name>Xiaocong Yang</name><email>xy51@illinois.edu</email></author><entry><title type="html">New Year, Deep Reflection</title><link href="https://xiaocong-yang.github.io/posts/2026/01/new-year-reflection/" rel="alternate" type="text/html" title="New Year, Deep Reflection" /><published>2026-01-01T00:00:00+00:00</published><updated>2026-01-01T00:00:00+00:00</updated><id>https://xiaocong-yang.github.io/posts/2026/01/new-year-reflection</id><content type="html" xml:base="https://xiaocong-yang.github.io/posts/2026/01/new-year-reflection/"><![CDATA[<p>Happy New Year, folks! First, I sincerely wish everyone reading this post – no matter what you do or where you come from – a wonderful and prosperous new journey starting today.</p>

<p>We have experienced a great deal over the past year. Unprecedented attention from governments, enterprises, and the general public has been directed toward AI, accompanied by numerous national policies and development plans, massive new investments, and intense public discussions. These forces have created several eye-catching rags-to-riches stories. Much like the technological transformation during the Industrial Revolution, AI offers this generation an opportunity to achieve social and financial mobility through intellectual capability, serving as a potential channel for breaking rigid social strata. However, as a public-facing technology, AI also demands caution and vigilance regarding several fundamental issues in current systems.</p>

<p>Let us start from the very beginning: how we define intelligence. To date, nearly all progress in AI has been defined exclusively in terms of improvements in instrumental capacity—the quality with which AI systems complete tasks. We see ever-higher scores on various leaderboards. This is certainly necessary for AI development, but it should not be the sole objective. Human history offers a clear lesson: the exclusive pursuit of instrumental rationality (“Zweckrationalität”) without value rationality (“Wertrationalität”) often leads to undesirable outcomes, including immoral behavior and negative externalities (London fog is not only the name of a latte, but also a historical environmental disaster caused by unchecked industrial production).</p>

<p>Indeed, we have already observed early signs of similar behavior in AI systems, such as reward hacking and overly flattering interactions. It is time for AI researchers to reconsider what additional factors, beyond instrumental capability, should be included in our definition of intelligence.</p>

<p>At the core of this question lies the interpretability of AI systems. This is precisely why I founded AI Interpretability @ Illinois: to bring together researchers from diverse backgrounds to develop models that operate in an interpretable and accountable manner. We do not interpret black-box models in hindsight—doing so does not make them safer or more trustworthy the next time they are deployed. Nor do we cherry-pick individual safety concerns and patch models accordingly—it is impossible to enumerate all potential risks. Instead, we pursue what we call “generative interpretability”: the property that an AI system’s internal mechanisms are human-understandable during deployment and real-world use, rather than only in laboratory settings with post-hoc tools. Such generative interpretability allows risks to be detected before an AI system’s actions are executed in the environment and cause real-world consequences.</p>

<p>Decentralization also plays a critical role. As I discussed earlier in a TDS blog post, decentralized computation is a defining feature of many of the most powerful systems on Earth, including human brains, financial markets, and biological swarms. Modern deep neural networks partially inherit this property, which is one reason they are substantially more powerful than centralized statistical learning models. What remains missing, however, is the decentralization of training objectives—the incentives that drive an AI system’s learning dynamics. More fundamentally, decentralization contributes to the democratization of AI and makes the formation of monopoly power technically more difficult.</p>

<p>Finally, we must reconsider the relationship between AI and humans. One of the most popular AI product forms today is the AI agent, which aims to automate entire workflows with minimal human involvement. Many ambitious teams seek to build general agents capable of performing arbitrary tasks. With current black-box foundation models, this goal is neither safe nor practical. Unlike chatbots—the dominant AI product today—agents operate in an action space whose outputs are executed directly and have real effects on the environment, rather than merely producing text. In domains where unsafe or irreversible actions are possible, such as autonomous driving or medical decision-making, deploying AI agents is unacceptable unless unsafe intentions can be detected internally before actions are taken—once again pointing to the necessity of generative interpretability. At the same time, truly general agents that can perform any task remain technically infeasible, given the vast action space of the real world.</p>

<p>This dilemma forces us to rethink the human–AI relationship. An inequality mentioned by my advisor, Prof. ChengXiang Zhai, captures two possible paradigms: an agentic model, where Intelligence (AI) ≤ Intelligence (Human), and a collaborative model, where Intelligence (AI + Human) ≥ Intelligence (Human). This inequality carries rich implications. From my perspective, it resembles an envelope curve or Pareto frontier. Each individual possesses unique advantages that AI cannot fully learn, as AI systems compress information by smoothing out edge cases. At best, an AI agent may become an “extraordinary generic person”—capable across many domains without obvious weaknesses, yet lacking rare and outstanding skills. Human–AI collaboration, by contrast, offers the opportunity to break through this envelope when uniquely human intelligence is brought to bear.</p>

<p>At this moment, we are both witnessing history and writing it. Some skeptics point out that AI has thus far made only limited contributions to global GDP. Ironically, this echoes famous remarks from earlier technological eras: two centuries ago, “What use is electricity?”—“Sir, what use is a newborn baby?”; and in the 1970s, “You can see the computer age everywhere but in the productivity statistics.” The intrinsic lag between technological innovation and measurable productivity gains has been well documented in economic growth theory and validated by decades of empirical evidence. The real question we should be asking is this: now that the bus is already accelerating, how do we steer it onto the right path at this critical crossroads?</p>]]></content><author><name>Xiaocong Yang</name></author><category term="AI" /><category term="Philosophy" /><summary type="html"><![CDATA[New Year reflections by Xiaocong Yang on AI progress, value rationality, generative interpretability, decentralization of training incentives, and the future of human–AI collaboration.]]></summary></entry><entry><title type="html">Neuro-Symbolic Systems: The Art of Compromise</title><link href="https://xiaocong-yang.github.io/posts/2025/11/symbolic-system/" rel="alternate" type="text/html" title="Neuro-Symbolic Systems: The Art of Compromise" /><published>2025-11-15T00:00:00+00:00</published><updated>2025-11-15T00:00:00+00:00</updated><id>https://xiaocong-yang.github.io/posts/2025/11/symbolic-system</id><content type="html" xml:base="https://xiaocong-yang.github.io/posts/2025/11/symbolic-system/"><![CDATA[<p>Long before we built computers and Artificial Intelligence, we had established institutions designed to reason systematically about human behavior – the court. The legal system is one of humanity’s oldest reasoning engines, where facts and evidence are taken as input, relevant <em>laws</em> are used as reasoning rules and verdicts are the system’s output. The laws, however, have been consistently evolving from the very beginning of human civilization. The earliest <strong>Codified Law</strong> - the <em>Code of Hammurabi</em> (circa 1750 BCE) - represents one of the first large-scale attempts to formalize moral and social reasoning into explicit symbolic rules. Its beauty lies in clarity and uniformity — yet it is also rigid, incapable of adaptation to context. Centuries later, <strong>Common Law</strong> traditions like those shaped by the <em>Case of Donoghue v Stevenson (1932)</em>, introduced the opposite philosophy: reasoning based on precedential experience and cases. Today’s legal systems, as we know, are usually a combination of both, while the proportions vary across different countries.</p>

<p>In contrast to the cohesive combination in legal systems, a similar paradigm pair in AI – <strong>Symbolism</strong> and <strong>Connectionism</strong> – seem to be significantly harder to unite. The latter has dominated the surge of AI development in recent years, where everything is implicitly learned with enormous amounts of data and computing resources and encoded across parameters in neural networks. And this direction, indeed, has been proven very effective in terms of benchmark performance. So, do we really need a symbolic component in our AI systems?</p>

<h2 id="symbolic-systems-vs-neural-networks-a-perspective-of-information-compression">Symbolic Systems v.s. Neural Networks: A Perspective of Information Compression</h2>

<p>To answer the question above, we need to take a closer look at both systems. From a computational standpoint, both symbolic systems and neural networks can be seen as machines of compression — they reduce the vast complexity of the world into compact representations that enable reasoning, prediction, and control. Yet they do so through fundamentally different mechanisms, guided by opposite philosophies of what it means to “understand”.</p>

<p>In essence, both paradigms can be imagined as filters applied to raw reality. Given input $X$, each learns or defines a transformation $H(\cdot)$ that yields a compressed representation $Y = H(X)$, preserving information that it considers meaningful and discarding the rest. But the shape of this filtering is different. Generally speaking, symbolic systems behave like <strong>high-pass filters</strong> — they extract the sharp, rule-defining contours of the world while ignoring its smooth gradients. Neural networks, by contrast, resemble <strong>low-pass filters</strong>, smoothing local fluctuations to capture global structure. The difference is not in what they see, but in what they choose to forget.</p>

<p>Symbolic systems compress by <strong>discretization</strong>. They carve the continuous fabric of experience into distinct categories, relations, and rules: a legal code, a grammar or an ontology. Each symbol acts as a <em>crisp boundary</em>, a handle for manipulation within a pre-defined schema. The process resembles projecting a noisy signal onto a set of human-designed basis vectors — a space spanned by concepts such as Entity and Relation. A knowledge graph, for instance, might read the sentence “UIUC is an extraordinary university and I love it”, and retain only <em>(UIUC, is_a, Institution)</em>, discarding everything that falls outside its schema. The result is clarity and composability, but also rigidity: meaning outside the ontological frame simply evaporates.</p>

<p>Neural networks, in contrast, compress by <strong>smoothing</strong>. They forgo discrete categories in favor of smooth manifolds where nearby inputs yield similar activations (usually bounded by some Lipschitz constant in modern LLMs). Rather than mapping data to predefined coordinates, they learn a latent geometry that encodes correlations implicitly. The world, in this view, is not a set of rules but a field of gradients. This makes neural representations remarkably adaptive: they can interpolate, analogize, and generalize across unseen examples. But the same smoothness that grants flexibility also breeds opacity. Information is entangled, semantics become distributed, and interpretability is lost in the very act of generalization.</p>

<table>
  <thead>
    <tr>
      <th><strong>Property</strong></th>
      <th><strong>Symbolic Systems</strong></th>
      <th><strong>Neural Networks</strong></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Survived Information</strong></td>
      <td>Discrete, schema-defined facts</td>
      <td>Frequent, continuous statistical patterns</td>
    </tr>
    <tr>
      <td><strong>Source of Abstraction</strong></td>
      <td>Human-defined ontology</td>
      <td>Data-driven manifold</td>
    </tr>
    <tr>
      <td><strong>Robustness</strong></td>
      <td>Brittle at rule edges</td>
      <td>Locally robust but globally fuzzy</td>
    </tr>
    <tr>
      <td><strong>Error Mode</strong></td>
      <td>Missed facts <em>(coverage gaps)</em></td>
      <td>Smoothed facts <em>(hallucinations)</em></td>
    </tr>
    <tr>
      <td><strong>Interpretability</strong></td>
      <td>High</td>
      <td>Low</td>
    </tr>
  </tbody>
</table>

<p>In conclusion, we can summarize the difference between the two systems from the information compression perspective in one sentence: “<strong><em>Neural Networks are blurry images of the world, while symbolic systems are high-resolution pictures with missing patches.</em></strong>” This actually indicates the reason why neuro-symbolic systems are an art of compromise: they can harness knowledge from both paradigms by using them collaboratively at different scales, with neural networks providing a global, low-resolution backbone and symbolic components supplying high-resolution local details.</p>

<figure style="margin:2rem auto;text-align:center;">
  <img src="/images/compression.png" alt="Counterfactual intervention diagram" style="display:block;width:100%;max-width:960px;margin:0 auto;border:none;box-shadow:none;" />
</figure>

<h2 id="the-challenge-of-scalability">The Challenge of Scalability</h2>

<p>Though it is very tempting to add symbolic components into neural networks to harness benefits from both, scalability is a big problem getting in the way of our attempts, especially in the era of Foundation Models. Traditional neuro-symbolic systems rely on a set of expert-defined ontology / schema / symbols, which is assumed to be able to cover all possible input cases. This is acceptable for domain-specific systems (for example, a pizza order chatbot); however, you cannot apply similar approaches to open-domain systems, where you will need experts to construct trillions of symbols and their relations.</p>

<p>A natural reaction is to go fully data-driven: instead of asking humans to handcraft an ontology, we let the model induce its own “symbols” from internal activations. <strong>Sparse autoencoders (SAEs)</strong> are a prominent incarnation of this idea. By factorizing hidden states into a large set of sparse features, they appear to give us a <strong>dictionary of neural concepts</strong>: each feature fires on a particular pattern, is (often) human-interpretable, and behaves like a discrete unit that can be turned on or off. At first glance, this looks like a perfect escape from the expert bottleneck: we no longer design the symbol set; we learn it.</p>

\[L_{\text{SAE}} = \|h - D\,\mathrm{ReLU}(Wh + b)\|_2^2 + \lambda \|\mathrm{ReLU}(Wh + b)\|_1\]

<p>Here $D$ is called the <em>dictionary matrix</em> where each column stores a semantically meaningful concept; the first term is the <em>reconstruction loss</em> of the hidden state $h$, while the second is a <em>sparsity penalty</em> encouraging minimal activated neurons in the code. See the <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">post</a> from Anthropic for details about SAE training and usage.</p>

<figure style="margin:2rem auto;text-align:center;">
  <img src="/images/SAE.png" alt="SAE diagram" style="display:block;width:80%;max-width:720px;margin:0 auto;border:none;box-shadow:none;" />
</figure>

<p>However, an SAE-only approach runs into two fundamental issues. The first is computational: using SAEs as a live symbolic layer would require multiplying every hidden state by an enormous dictionary matrix, paying a dense computation cost even if the resulting code is sparse. This makes them impossible for deployment at Foundation Model scales. The second is conceptual: SAE features are symbol-like representations, but they are not a symbolic system – they lack an explicit formal language, compositional operators, and executable rules. <strong>They tell us what concepts exist in the model’s latent space, but not how to reason with them.</strong></p>

<p>This does not mean we should abandon SAEs altogether — they provide ingredients, not a finished meal. Rather than asking SAEs to be the symbolic system, we can treat them as a bridge between the model’s internal concept space and the many symbolic artefacts we already have: knowledge graphs, ontologies, rule bases, taxonomies, where reasoning can happen by definition. And a high-quality SAE trained on a large model’s hidden states then becomes a shared “concept coordinate system”: different symbolic systems can then be aligned within this coordinate system by associating their symbols with the SAE features that are consistently activated when those symbols are invoked in context.</p>

<p>Doing this has several advantages over simply placing symbolic systems side by side and querying them independently. First, it enables <strong>symbol merging and aliasing across systems</strong>: if two symbols from different formalisms repeatedly light up almost the same set of SAE features, we have strong evidence that they correspond to the same underlying neural concept, and can be linked or even unified. Second, it supports <strong>cross-system relation discovery</strong>: symbols that are far apart in our hand-designed schemas but consistently close in SAE space point to bridges we failed to encode — new relations, abstractions, or mappings between domains. Third, SAE activations give us a model-centric notion of salience: symbols that never find a clear counterpart in the neural concept space are candidates for pruning or refactoring, while strong SAE features with no matching symbol in any system highlight blind spots shared by all of our current abstractions.</p>

<p>Crucially, this use of SAEs remains scalable. The expensive SAE is trained offline, and the symbolic systems themselves do not need to grow to “Foundation Model size” — they can remain as small or as large as their respective tasks require. At inference time, the neural network continues to do the heavy lifting in its continuous latent space; the symbolic artefacts only shape, constrain, or audit behaviour at the points where explicit structure and accountability are most valuable. SAEs help by tying all these heterogeneous symbolic views back to a single learned conceptual map of the model, making it possible to compare, merge, and improve them without ever constructing a monolithic, expert-designed symbolic twin.</p>

<h2 id="when-can-an-sae-serve-as-a-symbolic-bridge">When Can an SAE Serve as a Symbolic Bridge?</h2>
<p>The picture above quietly assumes that our SAE is “good enough” to serve as a meaningful coordinate system. What does that actually require? We do not need perfection, nor do we need the SAE to outperform human symbolic systems on every axis. Instead, we need a few more modest but crucial properties:</p>

<ul>
  <li><strong>Semantic Continuity</strong>: Inputs that express the same underlying concept should induce similar <em>support patterns</em> in the sparse code: the same subset of SAE features should tend to be non-zero, rather than flickering on and off under small paraphrases or context shifts. In other words, semantic equivalence should be reflected in a stable pattern of active concepts.</li>
  <li><strong>Partial Interpretability</strong>: We do not have to understand every feature, but a nontrivial fraction of them should admit robust human descriptions, so that merging and debugging are possible at the concept level.</li>
  <li><strong>Behavioral Relevance.</strong> The features that the SAE discovers must actually matter for the model’s outputs: intervening on them, or conditioning on their presence, should change or predict the model’s decisions in systematic ways.</li>
  <li><strong>Capacity and Grounding.</strong> An SAE can only refactor whatever structure already exists in the base model; it cannot conjure rich concepts out of a weak backbone. For the “concept coordinate system” picture to make sense, the base model itself has to be large and well-trained enough that its hidden states already encode a diverse, non-trivial set of abstractions. Meanwhile, the SAE must have sufficient dimensionality and overcompleteness: if the code space is too small, many distinct concepts will be forced to share the same features, leading to entangled and unstable representations.</li>
</ul>

<p>Now we discuss the first three properties in detail.</p>

<h3 id="semantic-continuity">Semantic Continuity</h3>

<p>At the level of pure function approximation, a deep neural network with ReLU- or GELU-type activations implements a Lipschitz-continuous map: small perturbations in the input cannot cause arbitrarily unbounded jumps in the output logits. But this kind of continuity is very different from what we need in a sparse autoencoder. For the base model, a few neurons flipping on or off can easily be absorbed by downstream layers and redundancy; as long as the final logits change smoothly, we are satisfied.</p>

<p>In an SAE, by contrast, we are no longer just looking at a smooth output — <strong>we are treating the support pattern of the sparse code reconstructed over the residual stream as a proto-symbolic object</strong>. A “concept” is identified with a particular code subset being active. That makes the geometry much more brittle: if a small change in the underlying representation pushes a pre-activation across the ReLU threshold in the SAE layer, a neuron in the code will suddenly flip from off to on (or vice versa), and from the symbolic point of view the concept has appeared or disappeared. There is no downstream network to average this out; the code itself is the representation we care about.</p>

<p>Sparsity penalty in constructing the SAE even exacerbates this. The usual SAE objective combines a reconstruction loss with an $\ell_1$ penalty on the activations, which explicitly encourages most neuron values to be as close to zero as possible. As a result, even many useful neurons end up sitting near the activation boundary: just above zero when they are needed, just below zero when they are not – this is known as “activation shrinkage” in SAEs. This is bad for semantic continuity at the support pattern level: tiny perturbations in the input can change non-zero neurons, even if the underlying meaning has barely changed. Therefore, Lipschitz continuity of the base model does not automatically give us a stable non-zero subset of code in the SAE space, and support-level stability has to be treated as a separate design target and evaluated explicitly.</p>

<h3 id="partial-interpretability">Partial Interpretability</h3>

<p>SAE defines an <em>overcomplete</em> (redundant) dictionary to store possible features learned from data. Therefore, <em>we only need a subset of these dictionary entries to be interpretable features</em>. Even for that subset, <em>meanings of the features are only required to be approximately accurate.</em> When we align existing symbols to the SAE space, it is the <strong>activation patterns</strong> in the SAE layer that we rely upon: we probe the model in contexts where a symbol is “in play”, record the resulting sparse codes, and use the aggregated code as an embedding for that symbol. Symbols from different systems whose embeddings are close can be linked or merged, even if we never assign human-readable semantics to every individual feature.</p>

<p>Interpretable features then play a more focused role: they provide <strong>human-facing anchors</strong> inside this activation geometry. If a particular feature has a reasonably accurate description, all symbols that load heavily on it inherit a shared semantic hint (e.g. “these are all duty-of-care-like things”), making it easier to inspect, debug, and organize the merged symbolic space. In other words, we do not need a perfect, fully named dictionary. We need (i) enough capacity so that important concepts can get their own directions, and (ii) a sizeable, behaviorally relevant subset of features whose approximate meanings are stable enough to serve as anchors. The rest of the overcomplete code can remain as anonymous background; it still contributes to distances and clusters in the SAE space, even if we never name it.</p>

<h3 id="behavioral-relevance-via-counterfactuals">Behavioral Relevance via Counterfactuals</h3>

<p>A feature is only interesting, as part of a bridge, if it actually influences the model’s behavior — not just if it correlates with a pattern in the data. In causal terms, we care about whether the feature lies on a <em>causal path</em> in the network’s computation from input to output: if we perturb the feature while holding everything else fixed, does the model’s behaviour change in the way that its believed meaning would predict?</p>

<p>Formally, changing a feature is similar to an intervention of the form $\text{do}(z = c)$ in the causal sense, where we overwrite that internal variable and rerun the computation. But unlike classical causal inference modeling, we do not really need Pearl’s <em>do-calculus</em> to identify $P(y \mid \text{do}(z))$. The neural network is a <strong>fully observable and intervenable system</strong>, so we can simply execute the intervention on the internal nodes and observe the new output. In this sense, neural networks give us the luxury of performing idealized interventions that are impossible in most real-world social or economic systems.</p>

<p>Intervening on SAE features is conceptually similar but implemented differently. We typically do not know the meaning of an arbitrary value in the feature space, so the <em>hard intervention</em> mentioned above may not be meaningful. Instead, we amplify or suppress the magnitude of an existing feature, which behaves more like a <em>soft intervention</em>: the structural graph is left untouched, but the feature’s effective influence is modified. Because SAE reconstructs hidden activations as a linear combination of a small number of semantically meaningful features, we can change the coefficients of those features to implement meaningful, localized interventions without affecting other features.  See the <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">post</a> from Anthropic about their implementation.</p>

<figure style="margin:2rem auto;text-align:center;">
  <img src="/images/causal.png" alt="Counterfactual intervention diagram" style="display:block;width:100%;max-width:960px;margin:0 auto;border:none;box-shadow:none;" />
</figure>

<h2 id="symbolic-system-based-compression-as-an-alignment-process">Symbolic-System Based Compression as an Alignment Process</h2>

<p>Now let’s take a slightly different view. While neural networks compress the world into some highly abstract, continuous manifolds, <em>symbolic systems compress it into a human-defined space with semantically meaningful axes along which the system’s behaviors can be judged</em>. From this perspective, compressing information into the symbolic space is an <strong>alignment process</strong>, where a messy, high-dimensional world is projected onto a space whose coordinates reflect human concepts, interests, and values.</p>

<p>When we introduce symbols like “duty of care”, “threat of violence”, or “protected attribute” into a symbolic system, we are not just inventing labels. This compression process does three things at once:</p>

<ul>
  <li>It selects which aspects of the world the system is <strong>obliged to care about</strong> (and which it is supposed to ignore).</li>
  <li>It creates a <strong>shared vocabulary</strong> so that different stakeholders can reliably point to “the same thing” in disputes and audits.</li>
  <li>It turns those symbols into <strong>commitment points</strong>: once written down, they can be cited, challenged, and reinterpreted, but not quietly erased.</li>
</ul>

<p>By contrast, a purely neural compression lives entirely inside the model. Its latent axes are unnamed, its geometry is private, and its content can drift as training data or fine-tuning objectives change. Such a representation is excellent for generalization, but poor as a locus of obligation. It is hard to say, in that space alone, what the system <em>owes</em> to anyone, or which distinctions it is supposed to treat as invariant. In other words, <strong>neural compression serves prediction, while symbolic compression serves alignment with a human normative frame</strong>.</p>

<p>Once you see symbolic systems as alignment maps rather than mere rule lists, the connection to <em>accountability</em> becomes direct. To say “the model must not discriminate on protected attributes”, or “the model must apply a duty-of-care standard”, is to insist that certain symbolic distinctions be reflected, in a stable way, inside its internal concept space — and that we be able to locate, probe, and, if necessary, correct those reflections. And this accountability is usually desired, even at the cost of compromising part of the model capability.</p>

<h2 id="from-hidden-law-to-shared-symbols--on-the-ethics-of-transparency">From Hidden Law to Shared Symbols — On the Ethics of Transparency</h2>

<p>In <em>Zuo Zhuan</em>, the Jin statesman Shu-Xiang once wrote to Zi-Chan of Zheng: “<em>When punishment is unknown, deterrence becomes unfathomable.</em>” For centuries, the ruling class maintained order through secrecy, believing that fear thrived where understanding ended. That’s why it became a milestone in ancient Chinese history when Zi-Chan shattered that tradition, cast the criminal code onto bronze tripods and displayed it publicly in 536 BCE. Now AI systems are facing a similar problem. Who will be the next Zi-Chan?</p>]]></content><author><name>Xiaocong Yang</name></author><category term="AI" /><category term="Philosophy" /><summary type="html"><![CDATA[Examines neuro-symbolic systems as a bridge between neural networks and symbolic reasoning, highlighting interpretability trade-offs.]]></summary></entry><entry><title type="html">Decentralized Computation: The Hidden Principle Behind Deep Learning</title><link href="https://xiaocong-yang.github.io/posts/2025/11/decentralization/" rel="alternate" type="text/html" title="Decentralized Computation: The Hidden Principle Behind Deep Learning" /><published>2025-11-05T00:00:00+00:00</published><updated>2025-11-05T00:00:00+00:00</updated><id>https://xiaocong-yang.github.io/posts/2025/11/decentralization</id><content type="html" xml:base="https://xiaocong-yang.github.io/posts/2025/11/decentralization/"><![CDATA[<p>Most breakthroughs in deep learning — from simple neural networks to large language models — are built upon a principle that is much older than AI itself: <strong>decentralization</strong>. Instead of relying on a powerful “central planner” coordinating and commanding the behaviors of other components, modern deep-learning-based AI models succeed because many simple units interact locally and collectively to produce intelligent global behaviors.</p>

<p>This article explains why decentralization is such a powerful design principle for modern AI models, by putting them in the context of general <strong>Complex Systems</strong>. If you have ever wondered:</p>
<ul>
  <li><em>Why internally chaotic neural networks perform much better than most statistical ML models that are analytically clear?</em></li>
  <li><em>Is it possible to establish a unified view among AI models and other natural intelligent systems (e.g. insect colonies, human brains, financial market, etc.)?</em></li>
  <li><em>How to borrow key features from natural intelligent systems to help design next-generation AI systems?</em></li>
</ul>

<p>… then the theories of Complex Systems where decentralization is a key property provides a surprisingly useful perspective.</p>

<h2 id="decentralization-in-natural-complex-systems">Decentralization in Natural Complex Systems</h2>

<p>A Complex System can be very roughly defined as a system composed of many interacting parts, such that <strong>the collective behavior of those parts together is more than the sum of their individual behaviors</strong>. Across nature and human society, many of the most intelligent and adaptive systems belong to the Complex System family and operate without a central controller. Whether we look at human collectives, insect colonies, or mammalian brains, we consistently see the same phenomenon: complicated, coherent behavior emerging from simple units following local rules.</p>

<p>Human collectives provide one of the earliest documented examples. Aristotle observed that <em>“many individuals, though each imperfect, may collectively judge better than the best man alone”</em> (<em>Politics</em>, 1281a). Modern instances — from juries to prediction markets — confirm that decentralized aggregation can outperform centralized expertise. The natural world offers even more striking demonstrations: a single ant has almost no global knowledge, yet an ant colony can discover the shortest route to a food source or reorganize itself when the environment changes. The human brain represents this principle at its most sophisticated scale. Roughly 86 billion neurons operate with no master neuron in charge; each neuron simply responds to its inputs from very few other neurons. Still, memory, perception, and reasoning arise from distributed patterns of activity that no individual neuron encodes.</p>

<p>Across these domains, the common message is clear: <strong>intelligence often emerges not from top-down control, but from bottom-up coordination</strong>. And we’ll see the principle provides a powerful lens for understanding not only natural systems but also the design and behavior of modern AI architectures.</p>

<h2 id="ais-journey-from-centralized-learning-to-distributed-intelligence">AI’s Journey: From Centralized Learning to Distributed Intelligence</h2>

<p>One of the most striking shifts in AI world in the past years, is the transition from a mostly centralized, hand-designed approach to a more distributed, self-organizing approach. Early statistical learning methods often resembled a top-down design: human experts would carefully craft features or rules, and algorithms would then optimize a single model, usually with <em>strong structural assumptions</em>, against a small set of data. Whereas today’s most successful AI systems – Deep Neural Networks – look very different. They involve lots of simple computational units (“artificial neurons”) connected in networks, learning collaboratively from a large amount of data with minimal human intervention in feature and structural design. In a sense, AI has moved from a paradigm of “let’s have one smart algorithm figure it all out” to “let’s have many simple units learn together, and let the solution emerge.”</p>

<h3 id="ensemble-learning">Ensemble Learning</h3>

<p>One bridge between traditional statistical learning and modern deep learning approaches in AI is the rise of ensemble learning. Ensemble methods combine the predictions of multiple models (“base learners”) to make a final decision. Instead of relying on a single classifier or regressor, we train a collection of models and then aggregate their outputs – for example, by voting or averaging. The idea is straightforward: even if each individual model is imperfect, their errors may be uncorrelated and can be cancelled. Ensemble algorithms like Random Forest and XGBoost have leveraged this insight to win many machine learning competitions since the late 2000s, and they remain competitive in some areas even today.</p>

<h3 id="statistical-learning-vs-deep-learning-a-battle-between-centralization-and-decentralization">Statistical Learning v.s. Deep Learning: A Battle between Centralization and Decentralization</h3>

<p>Now let’s look at both sides of this bridge. Traditional statistical learning theory, as formalized by Vapnik, Fisher, and others, explicitly targets at <strong>analytical tractability</strong> — both in the model and in its optimization. In these models, parameters are analytically separable: they interact directly with the loss function, not through one another; models such as Linear Regression, SVM, or LDA admit closed-form parameter estimators that can be written down in the form of $ \widehat{\theta} = \arg\min_{\theta} L(\theta) $. Even when closed forms are not available, as in Logistic Regression or CRF, the optimization usually remains convex and thus theoretically well-characterized.</p>

<p>In contrast, Deep Neural Networks admit no analytically tractable relationship between input and output. The mapping from input to output is a deep composition of nonlinear transformations where parameters are sequentially coupled; to understand the model’s behavior, one must perform a <em>full forward simulation</em> of the entire network. In the meantime, the learning dynamics of such networks are governed by iterative, non-convex optimization processes that lack analytical guarantees. In this dual sense, deep networks exhibit <strong>computational irreducibility</strong> — their behavior can only be revealed through computation itself, not derived through analytical expressions.</p>

<p>If we explore the root cause of the difference above, you’ll find it’s due to the model structures – as we might well expect to see. In statistical learning methods, the computational graphs are single-layer: $\theta \longrightarrow f(x;\theta) \longrightarrow L$ without any intermediate variables, and a “central planner” (the optimizer) passes the global information directly to each parameter. However, in Deep Neural Networks, parameters are organized in layers which are stacked on top of each other. For example, an MLP network without bias terms can be expressed as $y = f_L(W_L f_{L-1}(W_{L-1} \dots f_1(W_1 x)))$ where each $W_l$ affects the next layer’s activation. When calculating the gradient to update parameters $\theta = \lbrace W_i \rbrace_{i=1}^L$, it is inevitable that you’ll rely on <strong>backpropagation</strong> to update parameters layer by layer:</p>

\[\nabla_{W_l} L = \frac{\partial L}{\partial h^{(L)}} \frac{\partial h^{(L)}}{\partial h^{(L-1)}} \dots \frac{\partial h^{(l)}}{\partial W_l}\]

<p>This structural coupling makes direct, centralized optimization infeasible — <strong>information must propagate along the network’s topology, forming a non-factorizable dependency graph that must be traversed both forward and backward during training</strong>.</p>

<p>It’s worth noticing that most real-world Complex Systems, such as those we mentioned above, are decentralized and computationally irreducible, as solidly supported in Stephen Wolfram’s book <em>A New Kind of Science</em>.</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>Statistical Learning</th>
      <th>Deep Learning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Decision-Making</strong></td>
      <td>Centralized</td>
      <td>Distributed</td>
    </tr>
    <tr>
      <td><strong>Information Flow</strong></td>
      <td>Global feedback; all parameters get informed simultaneously</td>
      <td>Local feedback; signals propagate layer-by-layer</td>
    </tr>
    <tr>
      <td><strong>Parameter Dependence</strong></td>
      <td>Computationally separable</td>
      <td>Dynamically interdependent</td>
    </tr>
    <tr>
      <td><strong>Inference Nature</strong></td>
      <td>Evaluate explicit formula</td>
      <td>Simulate the dynamics of the network</td>
    </tr>
    <tr>
      <td><strong>Interpretability</strong></td>
      <td>High — parameters have global, often linear meaning</td>
      <td>Low — distributed representations</td>
    </tr>
  </tbody>
</table>

<h2 id="signal-propagation-the-invisible-hand-of-coordination">Signal Propagation: The Invisible Hand of Coordination</h2>

<p>A natural question about decentralized systems is: how do these systems coordinate the behavior of their inner components? Well, as we showed above, in Deep Neural Networks it’s via the propagation of gradients (gradient flow). In an ant colony, it’s via the spread of pheromone. And you must have heard the famous “<em>Invisible Hand</em>” coined by Adam Smith: price is the key to coordinating the agents in an economic system. These are all specific cases of <strong>signal propagation</strong>.</p>

<p>Signal propagation lies at the heart of Complex Systems. <strong>A signal proxy compress the landscape of the system, and is taken by each agent in this system to determine its optimal behavior</strong>. Take the competitive economic system as an example. In such an economic system, the price dynamics $p(t)$ of a commodity is used as the signal proxy and transmitted to the agents in this system to coordinate their behaviors. The price dynamics $p(t)$ <em>compresses and encapsulates key information of other agents</em>, such as their marginal believes of value and cost on the commodity, to impact the decision of each agent. Compared to spreading the full information of all agents, there are two major advantages:</p>

<ul>
  <li><strong>Better Propagation Efficiency.</strong> Instead of transmitting high-dimensional information variable — such as each agent’s willingness-to-pay function — only a scalar is propagated at a time. This drastic reduction in information bandwidth makes decentralized convergence to a market-clearing equilibrium feasible and stable.</li>
  <li><strong>Proper Signal Fidelity.</strong> Price provides a proxy with a just-right fidelity level of the raw information that could lead to a <em>Pareto Optimum</em> state at the system level in a competitive market, formalized and proven in the <a href="https://www.jstor.org/stable/1907353?seq=1">foundational work</a> by Arrow &amp; Debreu (1954). The magic behind is that, with this public signal being the <em>only</em> one available, each agent regards itself as a <strong>price-taker</strong> at the current price level, not an influencer, so that there’s no room for <em>strategic behavior</em>.</li>
</ul>

<p>It’s surprising that access to full information of all agents won’t result in a better state for the market system, even without the consideration of propagation efficiency. It introduces <strong>strategic coupling</strong>: each agent’s optimal action depends on others’ actions, which is observable under full information. From the perspective of each agent, it is no longer solving an <em>optimization problem</em> with the form of</p>

\[\max_{a_i \in A_i(p, e_i)} \; u_i(a_i), \qquad A_i(p, e_i) = \{ a_i : Cost(a_i, p) \le e_i \}\]

<p>Instead, its behavior is guided by the following <em>strategy</em>:</p>

\[\max_{a_i \in A_i(e_i)} u_i(a_i, a_{-i}),\qquad A_i(e_i) = \{ a_i : \text{Feasible}(a_i; e_i)\}\]

<p>here $a_i$ and $e_i$ are action and endowment of agent $i$ respectively, $a_{-i}$ are the actions of other agents, $p$ is the price of a commodity independent of the action of any single agent, and $u_i$ is the utility of agent $i$ to be maximized. With full information accessible, each agent is able to speculate the behaviors of other agents and so $a_{-i}$ enters the utility of agent $i$, creating strategic coupling. The economic system, therefore, eventually converges to a <em>Nash equilibrium</em> and suffers from inefficiencies inherent in non-cooperative behaviors (e.g. The Prisoner’s dilemma).</p>

<p>Technically, the signal propagation mechanism in markets is structurally equivalent to a <strong>Mean-Field model</strong>. Its steady-state corresponds to a Mean-Field equilibrium, and the framework can be interpreted as a special instance of a Mean-Field Game. Many Complex Systems in nature can be described with a specific Mean-Field model, such as <em>Volume Transmission</em> in brains and <em>Pheromone Field Model</em> in insect colonies.</p>

<h2 id="the-missing-part-in-neural-networks">The Missing Part in Neural Networks</h2>

<p>Similar to the above areas, the dynamics of neural network training are well characterized by mean field models in many previous works. However, there’s a vital difference between the training of neural networks and the evolution of most other Complex Systems: <strong>the structure of objectives</strong>. In Deep Neural Networks, the update dynamics of all modules is driven by a centralized, global loss $L(\theta)$; while in other complex systems, system updates are usually driven by <em>heterogeneous, local</em> objectives. For example, in economic systems, agents change their behaviors to maximize their own utility functions, and there’s no such “global utility” covering all agents that plays a role.</p>

<p>The direct consequence of this difference is the missing of <strong>competition</strong> in a trained Deep Neural Network. Different modules in a model form a <em>production network</em> that contributes to a single final product — the next token, in which the relationship between different modules is purely upstream-downstream collaboration (proposed in <a href="https://arxiv.org/abs/2503.05828">Market-based Architectures in RL and Beyond</a>; refer to <a href="https://interpretability.web.illinois.edu/tutorial-materials/#sec4">Section 4</a> of my lecture slides for a simplified derivation). However, as we know, competitive pressures induce <strong>functional specialization</strong> for agents in an economic system, which further gives the potential for a <em>Pareto Improvement</em> for the system via well-functioning exchanges. Similar logics has also been found when manually introducing competition in neural networks: <strong>a sparsity penalty induces local competition among units for being activated, which suppresses redundant activations, drives functional specialization, and empirically improves representation quality</strong>, as demonstrated in Rozell et al. (2008) where competitive LCAs produce more accurate representations than non-competitive baselines. <em>Intra-modular competition modeling, in this sense, would be an important direction for the design of next-generation AI systems</em>.</p>

<h2 id="decentralization-contributes-to-ai-democracy">Decentralization Contributes to AI Democracy</h2>

<p>At the end of this article, one more thing to talk about is the ethical meaning of decentralization. Decentralized structure of Deep Neural Networks provides a technical foundation for collaboration between models. When intelligence is distributed across many components, it becomes possible to assemble, merge or coordinate different models to build a more powerful system. Such an architecture naturally supports a more democratic form of AI, where ideally no single model monopolizes influence. This is surprisingly consistent with the belief from Aristotle that “<em>every human, though imperfect, is capable of reason</em>”, though the “humans” here are built from silicon.</p>]]></content><author><name>Xiaocong Yang</name></author><category term="Complex System" /><category term="AI" /><category term="Philosophy" /><summary type="html"><![CDATA[Explores how decentralized information processing shapes complex systems in biology, economics, and modern AI.]]></summary></entry></feed>