The Explanation Engine Redux
In my 2010 book, The Learning Layer, I made the case that explanatory capabilities in AI systems are an imperative, but I also described why providing such a capability that is sufficiently informative is highly challenging. Over the next few months after The Learning Layer was published, I published a series of blogs that discussed this topic in more detail, describing how a separate function is actually required, an “explanation engine,” to provide intelligible explanations for the more inscrutable, mathematical-based capabilities that in reality make the decisions in advanced AI systems. The specific example I gave as follows was with respect to recommender systems, but it applies more broadly for any sufficiently complex AI-based system.
items to recommend are going to be the product of various high-powered mathematical evaluations and manipulations of vectors and matrices. How can they possibly be compactly conveyed in an explanation to the recommendation recipient? After a while we realized that they really cannot—explanations necessarily have to be an approximation of the actual thought processes of the recommendation engine. A very useful approximation, but an approximation nevertheless. We also realized that to do them right, it was an architectural necessity to have a dedicated engine for explanations complementary to, but separate from, the powerful but inarticulate recommendation engine.
I further made the point that this two-part architecture of, 1) a complex and inscrutable decision-making function coupled with, 2) a separate explanatory function to explain the decisions of the inscrutable decision-making function to others (or even to itself, which I described as self-inception) basically recapitulates the evolution of the analogous two-part structure of human cognitive processes.
Turns out we are no different—we humans have a language-based explanation engine that explains to others, as well as ourselves, why we do the things we do. And it has become increasingly apparent from psychological studies over the past decade or two that our explanations are really only approximations of our underlying, unconscious decision making. In fact, it has been confirmed by recent brain imaging studies that although we tend to believe that our conscious and logical explanation engine is making the decisions, in reality it is just providing an after-the-fact explanation for a decision already unconsciously made elsewhere in our brain. . . our brain most fundamentally is nothing more than a vast, weighted network of neurons. Decisions are assessed and made by inscrutable interactions among this network. We humans, fairly recently in our evolutionary history, have developed a language-based explanation engine that has essentially been grafted on to this underlying network that enables us to effectively communicate with other humans. To achieve a reasonable compaction, the explanations we give must necessarily typically be extreme simplifications, and quite often will also have some degree of fabrication woven in.
Later, I was delighted to read Daniel Kahneman’s book, Thinking, Fast and Slow, which was published a little over a year after The Learning Layer, as well as after my follow-on blogs. Kahneman explained that human decision-making processes can be categorized into two fundamental processes: a first process that is fast, intuitive, pattern matching, and heuristic in nature, which Kahneman termed, “System 1,” and a second slower, more logical, reasoning and explanatory capability, which Kahneman termed, “System 2.” It is readily apparent, of course, that System 1 maps directly to the inscrutable decision-making function and System 2 maps to the explanation engine that I described with respect to AI systems. And, in fact, Kahneman’s book was based on earlier insights from cognitive research such as the those that I mentioned in The Learning Layer and in my blogs. Kahneman’s book greatly popularized these insights from the field of cognitive science, and for me, provided further credence that there was a strong parallel between the way AI would necessarily need to evolve and the way the human cognitive processes had evolved.
I also emphasized that we tend to think of ourselves as our explanation engine rather than the underlying, inscrutable processes that studies had shown actually drive our decision making.
And our explanation engine continuously explains itself to us—so that we come to believe that we are our explanation engine, that there is a clear logic to what we do, that we have an explicit freedom to decide as we wish, and a variety of other explanatory conceits.
We humans believe that we are our language-based explanation engine, and we therefore literally tell and convince ourselves that true meaning is solely a product of the conscious, language-based reasoning capacity of our explanation engines.
And because we think of ourselves as our explanation engine, it is not surprising that our initial focus in developing AI was almost solely directed toward embedding this type of symbolic, rule-based, logical capability within our AI, and more broadly, within our computing systems in general. But that’s not the way nature proceeded, and if we wanted to develop truly intelligent systems, I advocated that we needed to take lessons from nature in this regard!
Inarticulate inferencing and decision making capabilities evolved over the course of billions of years and work quite well, thank you. Only very recently did we humans, apparently uniquely, become endowed with a very powerful explanation engine that provides a rationale (and often a rationalization!) for the decisions already made by our unconscious intelligence—an endowment most probably for the primary purpose of delivering compact communications to others rather than for the purpose of improving our individual decision making. So to focus first on the explanation engine is getting things exactly backward in trying to develop machine intelligence. To recapitulate evolution, we first need to build intelligent systems that generate good inferences and decisions from large amounts of data, just like we humans continuously and unconsciously do. And like it or not, we can only do so by applying those inscrutable, inarticulate, complex, messy, math-based methods. With this essential foundation in place, we can then (mostly for the sake of our own conceit of consciousness!) build useful explanatory engines on top of the highly intelligent unconsciousness.
Well, since The Learning Layer and these subsequent blogs, my advice has indeed been heeded (although correlation should certainly not be confused with causation!), and the significant AI advances, particularly in the past decade, have come from advancing those “inscrutable, inarticulate, complex, messy, math-based methods” I mention in the above passage, primarily in the form of neural network-based technology, specifically deep learning. And in particular, with the advent of transformer-based deep learning technology in 2017, we have achieved remarkable facility with language in the form of large language models (LLMs) and their conversational interfaces that are profoundly changing our world.
And now that we’ve walked down memory lane, I want to get to the real point of this piece. We now have even more evidence that the path to artificial general intelligence (AGI) is by recapitulating the evolution of the explanation engine (or System 2) in humans. And that further evidence derives from the advances in the application of Chain of Thought (CoT)-based techniques in improving LLM-based performance. Chain of thought is a technique of prompting LLMs to explain their reasoning steps in deciding on a conclusion or an action, often performed in an iteratively interactive manner. It has been demonstrated to greatly improve the resulting output from LLMs. This prompting method is classically performed by a user, and in that case the explanation engine is partially external to the LLM-based system. But that need not be the case—the chain of thought prompting can be fully performed automatically by another system or sub-system of the LLM. And such an autonomous architecture should sound familiar—it’s basically the explanation engine, whether embodied in minds or machines!
This automated, built-in chain of thought capability is now a key direction of all the major LLM providers that is designed to achieve a more System 2-like capability to complement the inherently System 1-like capabilities of LLMs. These explanation engine-based architectures therefore promise to be ubiquitous. That such an architecture seems required to flexibly achieve truly human-level capabilities adaptable to many different domains, and the fact that nature has evolved a similar architecture, leads me to the following generalized hypothesis.
Hypothesis:
- Only large scale, connectionist-based correlative learning systems are flexible enough to provide the base support for effective real-world decision making and agency (e.g., human System 1, AI LLMs/foundation models)
- Every such connectionist, correlative learning system requires a complementary chain of thought-type function to achieve arbitrarily good inferences and decisions (e.g., human System 2, AI chain of thought)
The first point we have learned the hard way over many decades of trying to build AI. The second point I outlined nearly a decade and a half ago and is further confirmed by the essential nature of CoT techniques in augmenting LLMs. This hypothesis suggests that even if the human language-based explanation engine/System 2 initially evolved primarily for social communicative purposes as I speculated in the blog passage above, it undoubtedly also is applied internally to facilitate logical reasoning that was unattainable by our System 1 alone, and thereby improving our decision making.
The most fundamental open technical question right now in AI is whether an explanation engine/System 2 functionality can be based solely on neural network-based systems or does it need to be augmented by symbolic-based systems. The human brain suggests it can be done with just connectionist models, so my bet is AI architected the same way will be able to achieve human level capabilities, but that we will necessarily want to augment these capabilities with specialized systems that may be more symbolic in nature, just as we currently employ to augment our own capabilities.

