Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Explanations

This post describes what explanations are, and forms a basis for understanding how to explain artificial intelligent systems.


Definition

What is an explanation? There are many definitions. Here is one, from "How People Explain Action (and Autonomous Intelligent Systems Should Too)" (2017):

Explanation is arguably a three-value predicate: someone, a communicator, explains something to someone, an audience. The success of an explanation therefore depends on several critical audience factors—assumptions, knowledge, and interests that an audience has when decoding the explanation.

The definition of "Explanation" given above is unclear regarding what "explains" means. We could just as well define "Interview" as "Someone interviews someone about something". Though even in this vague form, it still highlights an exchange between two agents.

Inspired by Explanation in artificial intelligence: insights from the social sciences, this post defines "explaining" broadly as:

Explaining

A two-step process involving 1. the generation of explanatory hypotheses (cognitive process) and 2. the communication to an audience (social process).

Usually, one best hypothesis may be selected until contradicted by experience, superseded by a simpler one, or shown to be inconsistent with prior knowledge.

The process may repeat and update during the interaction. Sometimes it's during an explanation that we find errors in the understanding. Hence, explanations can provide understanding! Also, the explanandum (that which is to be explained) may be refined just as a photo-camera may gain focus with increased exposure.

Our understanding is reflected in the hypothesis formed in the cognitive process: Understanding is having a theory, hypothesis or model about how something works (cognitive process). It's also frequent that there is an illusion of understanding, and explaining or forcing predictions may make the illusion more evident.

Note

This post is mostly jargon-free. Technical articles about topics such as abductive inference are linked in the "Sources" at the bottom of this post. A brief discussion of logic inference in AI is given in this article by Gordon Brander.

Cognitive Process

Contrastive Questions

Research has shown that why-questions are usually contrastive. That is, they are phrased as Why P rather than Q? instead of simply Why P? Usually P is the real case (or fact) and Q the expected case (or foil), which may also be implicit.

As the paper Beware of Inmates Running the Asylum states:

For example, explaining "Why did Mr. Jones open the window?" with the response "Because he was hot" is not useful if the implied foil is Mr. Jones turning on the air conditioner, as this explains both the fact and the foil; or if the implied foil was why Ms. Smith, who was sitting closer to the window, did not open it instead, as the cited cause does not refer to a cause of Ms. Smith's lack of action.

The foil focuses the explanation on the differences between the two cases (ignoring similarities). This is usually easier to explain than the standalone fact. It can also reduce confusion.

Hesslow states this idea in a concise way:

What I want to suggest, then, is that the explanandum should be construed as a relation which involves three things: an object a, an object of comparison b and an explanandum property E which a has and b does not have.

The complexity, of course, lies on knowing which differences matter.

Attributing Causes

As Miller et al. state:

Attribution theory is the study of how people attribute causes to events; something that is necessary to provide explanations.

We never provide a full causal chain (it is endless), but a short-enough one that explains the event in question (this is the causal selection problem).

Researchers have pointed out many heuristics used by humans to favour some candidate causes (causal hypotheses) over others: proximal over distal events (in the causal chain of events); abnormal or unexpected events; controllable events, deviation from theoretical ideals, model, predictive power, responsibility, and so forth.

But those are taken care of by contrastive why-questions which compare the event to be explained to a reference case (particular instance or general case). In this regard, Hesslow states (bold is mine):

Many of the selection criteria listed in Section 3 can be construed as the result of choosing different objects of comparison or reference classes. Let us consider again the fire in the barn, and let us suppose that we have in the back of our minds the picture of a normal barn. (...) the normal barn has not caught fire, it follows that an explanatorily relevant condition for this barn's catching fire must be abnormal. Thus, selection of abnormal conditions can be viewed as the result of comparing the explanandum object with a normal object.

And also most other causal selections are contained:

(...) the difference between this barn now and this barn yesterday, i.e. we would be selecting a precipitating cause [proximal in the list above]. Selection of the unexpected may be viewed as the result of explaining the difference between an expected and an actual outcome. Selection according to responsibility follows from a comparison between actual and morally ideal behaviour. Selection of conditions which cause a deviation from a theoretical ideal involves a comparison between an actual and a theoretically ideal situation, and so on (cf. Hesslow, 1983).

Here is yet another illustration by Hesslow, of how contrasts cases narrow down possible causes:

For instance, if we want to explain why the fly Ml has shorter wings than Nl, then the temperature in which the flies were raised is explanatorily irrelevant, since the temperature was the same in both cases. The mutated gene on the other hand was present in one case and absent in the other.It is, therefore, explanatorily relevant.

Pragmatism

Notably, accuracy may not be preferred in an explanation; rather, usefulness, simplicity, generality and consistency with prior knowledge are.

Many of these results come from work by Tania Lombrozo. (This section will eventually be expanded.)

Social Process (Communication)

We have gone through the cognitive process and how contrastive questions can aid the generation and selection of a hypothesis or a cause. The second process is that of commucation.

The communication can be aided by the gricean maxims: rules of effective communication.

  • Informative (Quantity): right amount of context and details,
  • Truthful (Quality, or Fidelity): the explanation should be true,
  • Relevance (Relation): avoid presumed-known or superfluous details, focus on what provides insight,
    • One example given earlier is to focus on unexpected events, whilst ignoring what is presumed to be known by the listener.
  • Manner (clarity): express it in elegant terms.

In some cases, humans also tend to prefer concrete over abstract explanations, so "concreteness" could be added to the list.

Relevance is primarily related to the causal selection problem in relation to an audience, as Malle et. al., state:

How do people solve this problem? They determine what exact question the audience is interested in (McClure and Hilton 1998); they take into account what their audience member already knows (Slugoski et al. 1993); and they offer elements of explanations that build bridges between presumed knowledge and novel information (Korman and Malle 2016). In short, they offer explanations that generate coherence in a knowledge structure of old and new information (Thagard 1989).

Contrastive explanations can also take care of many of these aspects automatically, by selecting a contrast that is relevant or understood by the audience.

Metaphors: The Machine and The Agent

Humans often use "explanatory stances" to explain events, as noted by Daniel Dennett. There are three common ones:

  1. Mechanical stance (which I call "Machine Metaphor/Model"),
    • Explain outcomes by considering the parts of a system, what they do and how they interact (that is, a mechanism).
  2. Design stance, this has different interpretations. One is of the perceived purpose of something (applied to things created by humans such as tools, but also those hypothesised to be created by universal designer or god).
  3. Intentional stance (which I call "Agent Metaphor/Model").
    • Explanation uses goals, motives, feelings, intent to explain actions and/or behaviour.
    • Unintentional behaviour/events is usually explained using the machine metaphor (see the following paper, section Ordinary Behavior Explanation).

They can be complementary when applied to the same phenomena or as Ruth Byrne puts it:

Notably, each explanatory stance can be applied to explain the same device or action, but they have different consequences for understanding it. Each stance can lead to different kinds of insights, and to different kinds of erroneous inferences. The atypical application of a particular stance, say, a mechanical stance to explain an action more typically understood from an intentional stance, such as explaining travelers in a crowded airport as like pinballs careening around a pinball machine, may be interpreted analogically to yield new inferences [Keil, 2006].

In technical fields, many complex systems are conceptualised as machines: composed of parts, each with a function, a role. Many are also conceptualised as graphs.

Ordinary people conceptualise certain kinds of complex systems as humans or agents (wholly or in part). This may happen with systems using human language or behaving autonomously, but other times it is due to pragmatic reasons. They would use and expect the kind of explanation a human would give, if there were one.

What seems here most fundamental than the particular stances is the selection of a metaphor to structure thinking and obtaining insights.

Other metaphors and analogies could be proposed for specific problems.

Similar ideas can be found in "How People Explain Action (and Autonomous Intelligent Systems Should Too)":

For those intentional agents, we hypothesize, people will apply the same conceptual framework of behavior explanation that they apply to humans (...) a subset of AIS that people do not regard as intentional agents; and for those, they may apply a purely mechanical explanatory framework.

And more recently, in Good Explanations in XAI:

People may tend to adopt multiple stances in their preferred explanations of an AI decision support system and its decisions, not unlike their tendencies in interacting with social robots [Clark and Fischer, 2023]. People are aware that a social robot is a machine, but interpret it as a depiction of a character, not unlike a ventriloquist dummy, and engage with it in pretense of interacting with the depicted character [Clark and Fischer, 2023]. Similarly, they may be aware that an AI decision support system is an algorithm but they may interpret its decisions as a depiction of those provided by a human, e.g., a bank loan assessor, or the organization the human represents, a bank. Hence, an intentional stance and a design stance may both be useful in different contexts for explaining how automated agents behave [Veit and Browning, 2023].

We can summarise some of these ideas (including a standard audience) in a brief table:

PerspectiveModel is a…Preferred Explanation styleAudience
ScientificMachineMechanistic, causal, formalExperts
Human-facingAgent/PersonIntentional, narrativeUsers, stakeholders

The post on explanatory stances continues this line of reasoning and connects them with how we explain humans and deep learning models.


Sources
  1. Studies in the logic of explanation (1948), Their logically deductive model, and the related covariation model (Kelley, 1967) isn't how human explanations are considered in social and cognitive sciences any more. However, these are important historical background.
  2. Explanations, Predictions and Laws (1948),
  3. On the mechanization of abductive logic (1973). The first page is quite interesting.
  4. The Problem of Causal Selection (1988) fascinating and easy-to-read article.
  5. Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences (2017): Section 1 describes what the wrong approach is: building explanation models with an idea of explanation that only applies to experts. Section 2 surveys papers and notes almost none uses insights from social science of explanation to build their XAI algorithms, and even less evaluate them on humans. Section 3 is the most useful, and describes which insights from social sciences could be used (and points to research).
  6. How People Explain Action (and Autonomous Intelligent Systems Should Too) (2017). Argues that Agents will necessarily have initiative, planning, decision making and people will regard them as intentional agents. They will explain them (and expect the system to do so) as if it were a human.
  7. Blog Posts: What is Explainable AI? (2022) and from IBM.
  8. Good Explanations in Explainable Artificial Intelligence (XAI): Evidence from Human Explanatory Reasoning (2023). This paper discusses certain aspects of human explanations and understanding. For example: the illusion of understanding, thinking fast (intuitive, heuristic) and slow (deliberate, methodical), and explanatory stances. It also discusses counterfactual and causal explanations.