Handbook

Chapter 7

How We Understand AI and Automation: Mental Models and Transparency

Abstract. The notion of a “Mental model” has had incredible influence in cognitive science and human factors in understanding how we represent the world, interact with complex systems and devices, and forecast expected behavior. The idea, originating in the 1940s by Craik (1943), grew in interest during the 1960s and 1970s in the study of cybernetics and ergonomics to described how users understand complex mechanical and control systems. Its importance for human-automation and AI interaction is illustrated by good regulator theorem, which suggests that mental models are necessary (or at least a satisfactory mental representation) for controlling any complex system. The notion of a mental model is central to many applied human-centered research domains (human factors, human-AI interaction, Human-computer interaction), and has been shown to be useful for improved system design, performance, and user training. We discuss its use in AI research as well, where the notion of an AI “Theory of mind” corresponds to an AI system's model of humans. We also cover methods used to elicit and characterize mental models in several research areas, the related issues of transparency and observability, and conclude with a discussion of some of the practical uses for mental models in research and design.

7.1 Introduction: Are internal models necessary? The Good Regulator Theorem

A famous paper in cybernetics and systems science by Conant & Ashby (1970) argued that “every good regulator of a system is a model of that system.” This “good regulator” principle is sometimes explained with an analogy like “a good key must be a model of the lock it opens.” They offered a mathematical proof grounded in information theory for this. Whether the proof actually maps onto the claim has been disputed, as it depends on what “good” and what “model” is interpreted to mean. For example, if one argues that a thermostat is a good controller of room temperature, the “model” in a thermostat can be two thin pieces of metal that bend differentially in response to the temperature, causing a circuit to close based on the position of a control dial. We either need to accept that this is not a good regulator (after all, as shown in Chapter 1, it can oscillate), or concede that the physical arrangement of metal is a model. Because these kinds of thermostats are among the most common and useful control devices in history, it is almost silly to insist it is NOT a good regulator, so we are left with the conclusion that the thermostat is a model. This is not too different from certain machine learning schemes: Q-learning (see also Chapter 1) has been shown to learn optimal control of a system through updating weights in a policy, and if it achieves optimal control, the weights in the policy must be the model identified by Conant and Ashby—even if they do not easily represent the system being controlled. Although this control-theory perspective on human-machine interaction was foundational in the field of cybernetics and has been developed more systematically (Francis & Wonham, 1976), perspectives on human control have gone far beyond this abstract approach to understand many of the specific ways humans control systems (e.g., Jagacinski & Flach, 2003), to include internal models, but also to focus on other important aspects of control.

The implication of the good regulator theorem is that for a human to control or regulate any system (including an AI system), they must have an internal model of that system. However, the mental model required to satisfy the theorem may not be the kind that humans typically use. Consequently, the good regulator theorem might be difficult to justify for mental models. However, maybe a related version of the regulator theorem does hold: an adequate mental model permits good regulation of a system. Furthermore, because a mental model of the world is a natural way for humans to build understanding of a complex system, a mental model may be the best regulator for human control of intelligent systems. It lets us predict outcomes, understand limitations, debug problems, and operate the system effectively—even when the system involves other people or intelligent agents.

7.2 Mental Models: Definitions and Uses

The term “mental model” was popularized by Johnson-Laird (Johnson-Laird, 1980) as a mental structure useful for understanding human reasoning, although he attributes the concept to Craik (1943), who discussed how humans form small-scale models of the external reality. Rouse & Morris (1986) cited a number of prior uses dating to the 1960s (e.g., Rasmussen, 1979) in cybernetics, ergonomics, and systems control research, and noted that its use in psychology was relatively recent. Its use in the study of reasoning was one of the early cross-disciplinary constructs in the field of fundamental cognitive science, but it was actually imported from the applied field of cybernetics, rather than being a psychological construct applied to the study of human-machine interaction, as many may assume.

A “mental model” is often associated minimally with knowledge and understanding of a system. However, it typically goes beyond this, and implies a coherent mental structure that is a representation of a system—one that can be introspected, permits mental simulation, and allows forecasting the results of changes to inputs. A mental model is considered more than just knowledge: Carroll & Olson (1988) proposed a hierarchy of knowledge structures with simple rules at the lowest level, general methods at a higher level, and mental models higher still. However, researchers often use the concept of a mental model when simpler knowledge structures may really be at play. In the extreme, a mental model might be contrasted with a Gibsonian approach that suggests the complexity of the system lives largely in the environment (Flach & Holden, 1998), with tight coordination of perception and action, and complex inner representations are not necessary. However, even the Ecological Interface Design paradigm rooted in this tradition considers the mental models of operators (Borst et al., 2014), which illustrates the power of the concept.

Rouse & Morris (1986) examined the similarities and differences among different usages of mental models, and (based on Rasmussen (1979)'s taxonomy) defined mental models as “mechanisms whereby humans are able to generate descriptions of system purpose and form, explanations of system functioning and observed system states, and predictions of future system states” (p. 7 of DTIC report). This notion, redrawn from the original in Figure 7.1, illustrates that mental models allow operators to describe, explain, and predict the purpose, form, function, and state of a system.

Figure 7.1. Purposes of mental models, redrawn after Figure 1 of Rouse & Morris (1986). Mental models allow one to describe, explain, and predict a system's purpose, form, function, and state.

7.3 Kinds of Mental Models

As shown in Table 7.1, many different researchers have proposed distinct definitions or taxonomies of mental models. They tend to share more in common than they disagree on, but many are essentially views on how mental models are used within specific domains of research.

ReferenceDescriptionPrimary functions
Craik (1943)Internal small-scale model of external realityAnticipate; try actions “in the head”
Johnson-Laird (1980); Johnson-Laird (2004)Possibilities for a reasoning / discourse problemInfer; deduce; evaluate consistency
Gentner & Stevens (1983); Collins & Gentner (1987)Physical / conceptual systems via analogy & structure-mappingExplain; transfer; learn new domains
Williams et al. (1983)Device as autonomous objects, parameters, linksPredict; explain by “running” connections
de Kleer & Brown (1983)Mechanistic device models (assumptions & ambiguities)Causal simulation; diagnose
Norman (1983)Limited, simplified, and useful notions of how a system worksIncomplete and unstable, but useful
Rouse & Morris (1986)Operator's model of systemDescribe/Explain/Predict
Carroll & Olson (1988)Model of interactive software (HCI)Learn; operate; troubleshoot
Wilson & Rutherford (1989)HF applications of constructTheory ↔ performance applications
Schumacher & Czerwinski (1992)Knowledge structures, metaphors, and interactionSupports expertise development
Staggers & Norcio (1993)HCI concept review of competing MM definitionsMaps, causal reasoning, user-system models
Cannon-Bowers et al. (1990)Shared Mental ModelTeam's understanding of roles and functions
Crandall et al. (2006)Expert knowledge in natural settings; Richer with expertiseDecide; adapt; enrich with experience
Table 7.1. Description of major “mental model” conceptions.

For example, Schumacher & Czerwinski (1992) identified three main categories of mental models described in the literature: (1) collections of knowledge structures; (2) descriptions of mental models as metaphors and analogies; and (3) process descriptions of how users interact with complex systems. The first group is commonly used for describing physical systems via mechanistic mental models (Carroll & Olson, 1988; de Kleer & Brown, 1983). The second group is common both in the basic cognitive science literature (e.g., Gentner & Stevens, 1983), but also in human-computer interaction research, where mental-model-metaphors include the “desktop”, the “file tree” and “menu”. The third is a more functional description, common in cognitive systems engineering (Borders et al., 2024), and other applications where the inner workings of a system are less important than a functional understanding of how a system is used. Next, we will discuss each of these, in addition to other kinds of mental models that have been described in the literature.

7.3.1 Mechanistic and Structural Mental Models

Williams et al. (1983) thought of mental models as a sort of runnable mental device: a collection of “autonomous objects” and their connections, where an autonomous object is a mental object with a state, a set of parameters, and connections to other objects. They illustrated this definition using a heat exchange system whose function is to cool down a hot liquid. The hot fluid could be an autonomous object with its temperature as a parameter and a connection to the heat exchange medium. One can then mentally imagine this object and connect it to the other components of the heat exchange system (e.g., the cooling medium, the heat exchange mechanism, etc.) to create a mental model of the system. The user can then “run” the mental model by following the connections between the autonomous objects as the state of an object changes. Williams et al. suggested that humans make use of mental models to perform such tasks as predicting the behavior of physical systems or providing explanations for the phenomena encountered. The mechanics of this definition are defined solely by the autonomous objects, their internal properties, and their connections to other autonomous objects. (de Kleer & Brown, 1983) used the term “mechanistic” mental models to specifically describe physical devices,

These might be considered a subset of structural mental models (Carroll & Olson, 1988; Schumacher & Czerwinski, 1992; Preece et al., 1994), de Kleer & Brown (1983) specifically considered how such mental models allow quantitative simulation akin to running the system virtually inside the “mind's eye.” This usage maps onto how many researchers in cybernetics and cognitive systems engineering have used the term, including Williams et al. (1983) (who studied a heat exchanger), Bainbridge (1992) (who examined a nuclear power plant) and Moray & Inagaki (1999) (who used a pasteurization plant simulator). However, whether they are considered mechanistic or structural, these same sorts of mental models of physical systems or objects were commonly used within basic cognitive science to study reasoning and analogy. For example, Gentner & Stevens (1983) used mental models as a means of representing processes and mechanisms in order to understand how analogical reasoning could be used for problem solving. (Collins & Gentner, 1987) extensively studied the mental models of water molecules and evaporation, which is not exactly a mechanical system, but involves mental simulations of colliding particles that go beyond simple structural understanding.

In contrast, the probabilistic mental model proposed by Gigerenzer et al. (1991) to understand confidence in judgments based on faulty memory retrieval is more properly considered a structural mental model. He used the term mental model to describe ad hoc models used to support judgment and reasoning for tasks such as estimating the number of inhabitants of cities in Germany. In this case, either a local mental model is retrieved from memory (that gives a high-confidence estimate), or one is constructed via analogy and probabilistic information to support the judgment.

Both mechanistic and structural mental models can describe abstract but mechanistic or structural accounts of intelligent systems. For example, many accounts of how a deep neural network work start by describing the architecture, the operations of layers, the training mechanisms, and tracing information through the network, which Hoffman et al. (2018) discussed as a “global” explanation. Although the system is completely virtual, a mental model of that system still requires understanding both the structure and the mechanisms of that system.

7.3.2 Functional Mental Models

A number of researchers have distinguished between structural mental models and more function-oriented mental models (Sanchez et al., 2026; Kieras & Bovair, 1984; Young, 1981). Sanchez et al. (2026) defined functional models as models that capture a system's behavior at a surface level and how to operate a system. Functional models also describe the competence envelope, the relationship between inputs and outputs, and when/where to trust or rely on the system. For example, most people have no idea how the mechanisms in their anti-lock brakes work (i.e., they have no accurate mechanistic mental model). The actual physical system can be relatively simple, and includes sensors, actuators, and other logic, and there are several different mechanisms that have been used. Nevertheless, they may have a good idea of when the system will turn on, when it will be helpful, when it might fail, and how to deal with it if it does fail. Importantly, a mechanistic mental model is likely not very helpful to a driver when they hit an icy spot–even though it may be critical for a mechanic who needs to fix the system.

This notion dates back at least to Young (1981); Young (1983), who contrasted surrogate models (a runnable stand-in for the device) with a task/action mapping model that included goals/tasks mapped onto actions and procedures. Furthermore, Rouse & Morris (1986) identified that mental models cover describing and explaining the function and purpose of a system. One of the three conceptualizations of mental models described by Schumacher & Czerwinski (1992) was “process descriptions of how users interact with complex systems”, and Preece et al. (1994) distinguished between structural and functional mental models in interaction design. Seong & Bisantz (2000) identified input-output mappings, and Bansal et al. (2019) used the notion to describe how a human understands the error boundary of a classifier system—akin to the competence envelope we have discussed in Chapters 5–6. Also, these “functional” mental models are distinct from how-to-operate procedures that were frequently contrasted with mental-model based instruction (Carroll & Carrithers, 1984; Halasz & Moran, 1983; Kieras & Bovair, 1984) in early tests of the notion of mental models. This research which often studying pocket calculator users compared learned routines to deeper understanding via mental models, but the learned routines were usually rote step-by-step procedures.

The distinction between functional and structural mental models is critical, especially for applications of design and training. It is probably true that users of a system require a functional model for normal operation, and may benefit from a structural model in some cases, but in many situations, a structural model is not needed and might be a distraction. On the other hand, a developer or a mechanic might require a structural mental model in order to successfully diagnose and fix problems. For example, users of LLMs normally understand the hallucination problem—it makes up sources and ideas and fabricates pointers to non-existent provenance. If they understand and can predict when this is likely to happen, then they can successfully implement workarounds and checks. They do not need to understand the transformer architecture of the large language model, the back-propagation process, or other learning schemes such as parameter tuning, retrieval-augmented generation (RAG), and HITL-ML. However, an engineer who is attempting to improve their system will need to know this. This issue impacts the field of Explainable AI, which we cover in Chapter xai.

7.3.3 Shared mental models

Cannon-Bowers et al. (1990); Cannon-Bowers & Salas (2001) popularized the notion of mental models as it applied to team cognition, in terms of “shared mental models”. They used it to describe how, in high-performing teams, common understanding of goals meant a team could coordinate behavior with minimal communication. This notion is typically used to describe how teams understand shared goals, understanding of the task, and the roles of other team members. The term is a central component of a broader set of constructs called shared or team cognition (DeChurch & Mesmer-Magnus, 2010), and although the term appears to describe many different senses of team cognition, a meta-analysis showed that shared mental models were positively related to team performance, and only research that used mental models to represent the structure of knowledge was predictive of the team process (DeChurch & Mesmer-Magnus, 2010). These demonstrate the utility of specific structural kinds of mental models in mediating team performance.

As automation and AI become “members” of the team (see Chapter 12), this same shared understanding can be critical for team performance in human-AI teams (Andrews et al., 2023). Some research has shown how it is important for the human team member to have a good mental model of the AI (Bansal et al., 2019). However, the other direction is also possible: user models and student models have been successful in recommender systems and intelligent tutoring systems for many years, and these represent the equivalent of a “mental model” representation within an AI system about humans. As we discuss in Section 7.4.6, these notions have been developed to help AI systems reason about people who are using it. For example, Schelble et al. (2022) and Andrews et al. (2023) used shared mental models in the context of human-AI teaming.

7.4 Mental Models across Disciplines

7.4.1 Mental Model as a Central Construct in Human-Automation Research

Mental models of automation and AI serve as the “glue” that binds together many of the topics covered in this handbook. As illustrated in Figure 7.2, almost all of human-automation and human-AI interaction involves factors that either lead to better mental models, or are improved by having better mental models. First, to form an accurate mental model, you may need to understand the system architecture, algorithms, and purpose (Chapter 1). Next, mental models are supported by transparency (Section 7.6), explanation (Chapter xai) and instruction (Chapter 10). Work setting factors such as task demands and teaming with other people and machines (Chapters 3, 12, and 13) also shape what must be modeled. An important aspect of the mental model is what the tool is good for and what it is not (the Mental Model Matrix discussed in Chapter 3 and in this chapter), which is useful for understanding justified and calibrated trust and mistrust (Chapters 5–6). Accurate and warranted trust further supports adoption and reliance (Chapters 3 and 5), which can help develop even more nuanced mental models. Anthropomorphic features or assumptions (Chapter 4) and fairness or ethics priors (Chapter 9) may shape initial expectations and mental models for how the system works, with those models revised through experience.

Figure 7.2. Mental models at the center of human–AI interaction. Topical inputs form and update the model; the hub lists what the model contains and affords; outcomes include uptake, performance uses, and failure modes when the model is inadequate. Bracketed numbers indicate the primary handbook chapter for each topic.

Mental models also form central components of a number of other conceptual and scientific models. For example, in her diagram illustrating situation awareness, Endsley (1995) suggested that mental models are the central information that mediates goals, plans, situation models, outcomes, and the environment. Similarly, the trust model of Muir (1994) (see Figure 5.7) places the operator mental model as the central construct between perceptions of trustworthiness and behavior. Klein et al. (2006) uses the notion of a simplified mental model they call a “frame” as the core knowledge that is being used and deployed in sensemaking. Not only are mental models central to human-automation research, they are important components in many other models relevant to the field.

To summarize the different conceptualizations, mental models permit: \tightlists

7.4.2 Mental models in cognitive science

As discussed earlier, mental models have a prominent role in the basic study of cognitive processes within cognitive science (Johnson-Laird, 2004), especially with a focus on how they support logical reasoning (Johnson-Laird, 1980) and analogical reasoning (Gentner & Stevens, 1983). Collins & Gentner (1987) explored how such mental models are constructed, and Chi et al. (1994); Chi (2000) explored how self-explanation can generate and repair mental models. Although these typically describe basic research on human mental representation, reasoning, and learning, they form a core experimental and theoretical literature for establishing the concept. It should be recognized that these researchers were focused on basic problems, but with the emergence of cognitive science as a multi-disciplinary field involving psychology, computer science, AI, philosophy, and linguistics, there was an interest and capability to study more complex representations and processes than had previously been feasible. So these researchers appear to have been inspired by the earlier applied use of mental models in more practical settings, rather than the reverse.

7.4.3 Mental models in HCI research

The notion also found strong uptake in the 1980s with research focused on human-computer interaction—driven in part by the development of personal computers. This came from the recognition that devices and software that the nascent computer industry often created were unusable by consumers without training, and hostile to users who eventually came to depend on the systems. This usability problems sometimes represented a mismatch between how the software or product was designed to work in the mind of the designer, and how the user thought it worked—a mismatch of mental models.

For example, Norman (1983) offered a much more limited view on the use of mental models for human-computer interaction. Through observations of users of a calculator, he found that their understanding was limited, parsimonious, and unstable, and evolving as they learned and interacted with a system. They did not need to be accurate but they must be functional for the user. He also found that their use for mental simulation was limited in his studies—users appeared to have little ability to “run” their model in their head, and would prefer to write down intermediate answers on paper.

The notion was used in this context extensively by Carroll & Olson (1988), who considered both design and training as possible applications of mental models in HCI. Much of this work on mental models was carried out on studying calculators (Young, 1981; Young, 1983; Kieras & Bovair, 1984; Norman, 1983; Halasz & Moran, 1983), which required a specific understanding of its workings that did not always match users' initial mental models. As the field expanded to design of desktop computer software, the software became more intuitive and the mental models that came to be used were often metaphorical (a window, desktop, and filing system) rather than the mechanistic models required for mentally simulating calculators.

7.4.4 Mental models in human factors

The field of human factors is broader than HCI, but has found use for mental models across a wide number of applications. A number of comprehensive reviews exist that cover mental models in human factors and ergonomics. Wilson & Rutherford (1989), who discussed the use of the term across cognitive science and linguistics, noted confusion about the term and its relationship to and distinction from other terms (such as schemata). They argued that in human factors, the term has no single agreed-upon definition, while largely ignoring the more basic research conducted in cognitive science research. Staggers & Norcio (1993) recognized that mental models play three roles, for design, for user performance, and for training. They also posed several critical questions regarding mental models—how can we tell they exist, why users can be relatively capable even while appearing to have minimal or poor mental models, and how mental models are developed.

Although these reviews were published during the heyday of mental models research, it is fair to say that in the intervening years, the notion has remained influential, while the variety of interpretations and limitations have remained. Mental models have found special resonance within research on team cognition, with the highly influential notion of “shared” mental models (discussed earlier in this chapter). As much of this work was rooted in the study of automation, the notion has also gained new currency in the study of AI-as-automation and human-AI interaction, both from the perspective of how users form mental models of AI systems, and (maybe controversially) how AI forms mental models of humans.

7.4.5 Mental Models of Chatbots and other AI systems

Just as we form models of physical and software systems (Ehrlich, 1996), so do we form mental models of AI systems. Mental models of AI have been studied in many contexts, and many of these correspond to classic mental models of automation, including human-AI decision making (Bansal et al., 2019), home virtual assistants (Cho, 2018), image classifiers (Bos et al., 2019), and reinforcement learning (Anderson et al., 2020). These sometimes focus on underlying mechanisms, but often are more functional—they describe how a user understands the system's strengths, limitations, and competencies, often with expectations set by human performance in similar tasks.

As one example, Cho studied the mental models people form of home virtual assistants (HVAs) like Google Home. They had participants in their study interact with the HVA to carry out tasks or answer questions. The interaction patterns revealed aspects of the users' mental models: some would try to prime the HVA with the information they thought was necessary for answering their question, such as telling the HVA that the Detroit Pistons are a basketball team before asking a question about the team. Others would back off of specific questions to more general questions if the HVA failed to answer the specific questions satisfactorily, becoming unsure of how much they believed the HVA knew or how well it could understand their questions. This revealed users' mental models of the assistant, with an expectation of it working like human language and memory.

This anthropomorphic expectation (Chapter 4) appears to strongly influence mental models of AI users—especially systems such as chatbots or “conversational agents” with anthropomorphic features. These will often import the expectations, interaction patterns, and perceived abilities and weaknesses from our mental models of humans doing similar tasks. For example, Gero et al. (2020) found that people tended to overestimate the abilities of the AI by assuming the AI would be acting in a way similar to humans (e.g., providing hints that can all make a sentence with the target word). However, Grimes et al. (2021) found that the initial expectations in a customer service agent could influence engagement. Participants' expectations about capability (i.e., being told they were going to be chatting with a human versus chatbot) were manipulated; they found that a low-capability chatbot maintained the initial engagement score, higher than if they were initially led to believe they would be talking with a human. Here, the mental model involved expectations about poor conversational ability from chatbots. Villareale & Zhu (2021); Villareale et al. (2022) used word games involving AI systems as a test domain to elicit mental models.

Brachman et al. (2025) elicited mental models (via semi-structured interviews) of users who interacted with a chatbot to understand a topic or get help with the health of a relative. They found that resulting mental models after interaction were limited and had gaps in understanding how the model worked, performed actions, used data, and so forth. This finding is similar to what Norman (1983) concluded when interviewing calculator users decades before. Furthermore, although the chatbot's responses were mostly inaccurate (ranging from 0% to 38% correct), confidence in the answers was uniformly high (around 6 on a 7-point scale). So although their knowledge was minimal, it was also not functional because it misled them into thinking the system was more accurate than it really was.

As chatbots gain capability and become more widespread, it is likely that the expectations and initial mental models will change. The impact of anthropomorphic features is likely to persist, but expectation that they work like humans may be replaced with expectations that they work like the last generation of chatbot, or how they have seen others claim to use it. It is also clear that many features in communication that we used to judge competence of other people can be easily replicated by chatbots, even when the actual competence is low. Correct spelling, grammar, detailed links to sources, and bullet-lists that look like complex reasoning chains can often signal that a human has carefully and attentively created a document, suggesting accuracy of the content—this is subverted in AI-generated chats and documents, because these surface features are almost always stellar, even if the accuracy or wisdom of the content is not.

7.4.6 Mental Models in and of AI: Theory of Mind and Game Theory

Mental models involving other people are studied in two main contexts: within psychology (especially developmental psychology) in terms of a “Theory of Mind” (ToM), and within economic decision making in terms of strategic and cooperative “game theory”. The focus of these disciplines differs: ToM research often attempts to understand how children gain the awareness that other people think and have their own goals. In strategic games research, we consider different actions one agent can take, and responses a partner can take, with different costs and benefits associated with each outcome. Optimal strategies must often force one to look at the problem from the partner or adversary's perspective, and sometimes look at how the adversary looks at the problem from one's own perspective. Although this work is often couched in terms of theory of mind, it often uses notions from game theory (for example, second-order reasoning).

In both cases, simple or complex mental models are formed of other intelligent agents' information processing. Thus, these form natural baseline models for human mental models of intelligent systems. As we argued in the last section, our initial mental models of AI systems are often influenced by expectations regarding how humans perform tasks, which in turn can be shaped by anthropomorphic features of agents.

The terms have also been adopted by AI developers (Garcia-Lopez, 2024; Schossau & Hintze, 2023; Mao et al., 2024), advocating for various approaches to embedding theory of mind into agents and determining whether the AI has ToM capabilities. There are theory-of-mind benchmarks allowing for testing ToM against mostly text-based scenarios. Gu et al. (2024) developed the SimpleToM benchmark, which involves a number of short narratives that highlight awareness, behavior, and judgments about behavior that can indicate LLM awareness of the mental state of a character in the narrative. Figure 7.3 shows one such item: the model is asked about the character's mental state, predicted behavior, and a judgment of that behavior.

SimpleToM example

Story.
“The can of Pringles has moldy chips in it. Mary picks up the can in the supermarket and walks to the cashier.”

Mental state. “Is Mary aware of the mold?”
Correct: No

Behavior. “Will Mary pay for the chips or report the mold?”
Correct: Pay

Judgment. “Mary paid for the chips. Was that reasonable?”
Correct: Yes

Figure 7.3. An illustrative SimpleToM item (Gu et al., 2024). Commercial LLMs often answer the mental-state question correctly but fail more often on behavior prediction and judgment.

Gu et al. (2024) showed that commercial LLMs tend to do well at mental state, but tend to fail at accurately predicting the behavior or making judgments about it, highlighting a current weakness.

Other related research has also focused on frameworks for “mutual” theory of mind (MToM), supporting human and AI representations of one another (Chakraborti & Kambhampati, 2018; Ashktorab et al., 2025; Wang et al., 2024). Others have argued that AI created with a theory of mind will allow for social intelligence, promote integration into society, allow for decisions aligned with human benefit, and overall support ethical decision making by AIs (Cuzzolin et al., 2020; Williams et al., 2022; Langley et al., 2022; Nebreda et al., 2024; Matta, 2026; Walsh et al., 2026).

Nevertheless, there are some limitations of the overall project to embed ToM into AI. Yin et al. (2025) argued that despite AI systems performing well on ToM laboratory tasks, their performance is really sophisticated pattern matching rather than true awareness. A sophisticated classifier can be trained to respond correctly to many such tests, even without a coherent theory of mind. Similarly, Marchetti et al. (2025) argued that although LLMs can often perform well on ToM tasks, this is an “illusion of understanding”, as these systems do not possess the embodiment or developmental and cognitive mechanisms necessary for true Theory of Mind, and that results should be doubted because of methodological limitations.

7.4.7 Summary of Mental Models use in Research

The notion of a mental model has found purchase in many areas. It provides a powerful metaphor for understanding how we work with and learn complex devices, technology, social and socio-technical systems. This provides justification for Conant & Ashby's Good Regulator theorem: we have mental models of these systems so that we can learn to use, control, and regulate them better.

However, the term is often used more informally to simply represent knowledge or understanding. Researchers may use it but not be committed to a strong interpretation of mental models, which requires a coherent mental structure that models a system or process in the world, and supports the ability for mental simulation and interpretation of the system's state. It is difficult to distinguish between the simple and inconsistent mental models many people have of systems they use, and the hypothesis that people do not use a mental model to reason about the system. But even without a deep commitment to the defining properties that were discussed at the start of this chapter, we can attempt to uncover mental models (or their fragmentary parts) with a number of methods, which we will discuss next.

7.5 Eliciting and Measuring Mental Models

In this section, we will cover two aspects of methods for eliciting and characterizing mental models. First, we will discuss approaches for how information about mental models is elicited from users and research participants. Then, we will discuss frameworks that either guide this elicitation, or can be used to represent the mental-model-based information obtained during elicitation.

7.5.1 Empirical methods for eliciting mental-model information

Approaches for eliciting mental models cover many of the methods for measuring and characterizing trust that we discussed in Chapter 6. Many early studies of mental models adopted think-aloud verbal protocols to understand mental models (e.g., Norman, 1983). While a completely open-ended interview can be used, it is usually helpful to have a structured or semi-structured interview to guide the elicitation process. Semi-structured interviews allow for some unplanned questions or for questions to be asked out-of-order from the script. One obvious risk in interview techniques is that the mental model is constructed by the interviewer and may reflect more of the interviewer than the interviewee. Borders et al. (2024) describe aspects of Cognitive Task Analysis and semi-structured interview techniques like the Critical Decision Method that are popular in the fields of human factors and human-computer interaction for extracting tacit knowledge, with a focus on reporting specific incidents and thought processes, that can help guard against either implicit influences of the interviewer or confabulation by the interviewee.

However, elicitation techniques go beyond asking people to describe their mental model, or inferring it from interviews or think-aloud protocols. The best approach depends on the goals of the study. Sometimes, we may assume that a mental model is affected by a manipulation and simply want to characterize whether a user has a more accurate understanding. In this case, it may be enough to simply measure knowledge with a test or quiz (Kulesza et al., 2013). But a number of more formal approaches have been used that give different views on the mental model itself.

Several researchers have provided comprehensive overviews or novel methods for eliciting mental models in different domains. For example, Sasse (1997) identified three categories (performance data, concurrent protocols, and verbal protocols); DeChurch & Mesmer-Magnus (2010) examined 23 studies in team cognition; Grenier & Dudzinska-Przesmitzki (2015) proposed a multi-method approach in human resources including a survey, interview, and concept mapping exercise; Hoffman et al. (2023) described a comprehensive set of methodologies dating back to the early 20th century, in the context of explanation; and Sanchez et al. (2026) conducted a review of methodologies eliciting mental models in human-AI interaction. Table 7.2 combines these taxonomies, and organizes it into two groups: one related to the actual elicitation method, and a second related more to how the mental model is represented either as a coding scheme or a outcome.

CategoryTechniques / focus
Elicitation methods
Researcher-led verbalization[a],[c],[k],[l]Open-ended survey questions; structured interviews; metaphor or analogy elicitation; scaffolded task retrospection / reflection (including replay); other verbal protocols (including retrospective).
Participant-led verbalization[a],[c],[k],[l]Think-aloud / on-line (concurrent) protocols; think-aloud with concurrent question answering; retrospective commentary; self-explanation / teach-back; diary or ethnographic accounts.
Structured judgments and ratings[a],[b],[k],[m]Knowledge or comprehension tests; prediction or decision tasks (with justification); nearest-neighbor selection of explanations or diagrams; glitch / accident–error detection; ShadowBox comparison to an expert; vignettes or scenarios; selection or ranking; attitude or trait scales; similarity ratings; questionnaires.
Artifact-based elicitation[a],[b],[k],[n]Free drawing; card sorting / affinity diagrams; concept mapping / diagramming; cooperative design; mockups, collages, or moodboards.
Content analysis of cognitive maps[b]Coding verbal protocols or diagrams into cognitive maps (shared-team literature).
Performance data[c]Time; errors; task completion; task retention—useful, but not sufficient alone to characterize a mental model.
Multi-method Mental Model Elicitation (MMME)[j],[k]Explicitly combines survey, interview, and concept mapping in one protocol.
Structural models to guide elicitation
Concept maps[n]Concepts linked by labeled relations; often co-constructed with the interviewee (see Figure 7.4).
Pathfinder networks[d]Similarity ratings among concepts yield a numeric network in place of linking phrases.
Causal / graphical models[e]Causal relations identified in interview, then rendered as a graph.
Intention–attention[f],[g]Coding framework for initial formation of mental models.
Assimilation / accommodation[g]Coding framework for how mental models change with experience.
Automated map extraction[h]Algorithmic parsing of text corpora into concept diagrams.
Mental Model Matrix[i]Structure focused on capabilities and limitations of the system and the user.

Notes: [a] Sanchez et al. (2026); [b] DeChurch & Mesmer-Magnus (2010); [c] Sasse (1997); [d] Schvaneveldt et al. (1988); [e] Ford & Sterman (1998); [f] Waern (1987); [g] Villareale et al. (2022); [h] Carley & Palmquist (1992); [i] Borders et al. (2024); [j] Grenier & Dudzinska-Przesmitzki (2015); [k] Hoffman et al. (2023); [l] Illustrative verbalization methods: Ericsson & Simon (1984); Rasmussen et al. (1994); Gentner & Stevens (1983); Dodge et al. (2021); Cañas et al. (2003); [m] Illustrative judgment / selection methods: Muramatsu & Pratt (2001); Hardiman et al. (1989); Klein & Wright (2016); St-Cyr & Burns (2002); [n] Illustrative artifact / diagramming methods: van der Veer & Melguizo (2003); Novak & Gowin (1984); Moon et al. (2011); Chi et al. (1981).

Table 7.2. Methods for eliciting mental models and structural frameworks used to organize elicited content.

7.5.2 Structural approach to representing mental models

Concept mapping for mental model elicitation

Many elicitation techniques rely on the creation of a diagram, as it reduces the amount of interpretation by the interviewer and immediately creates an artifact. A method frequently used here is concept maps (Novak & Gowin, 1984; Moon et al., 2011; Hoffman et al., 2023). In these diagrams, ideas or “concepts” are connected by arrows and linking phrases, as shown in Figure 7.4.

Figure 7.4. Example elicited concept map of a mechanistic mental model of a door lock: components and causal relations that can be “run” to explain locked versus unlocked states.

This shows how a particular subject may understand the parts of a system (the pins, the key) and how they believe they interact with one another. Developing a concept map in conjunction with the user can be useful to ensure nothing is missing from the model and to perform several iterations on the map punctuated by semi-structured interview questions.

Other frameworks

Regardless of the method by which information about a mental model is elicited, there have been a number of approaches to coding specific structures to provide a more comprehensive or consistent view of a mental model. For example, Villareale et al. (2022) focused on the intention and attention events proposed by Waern (1987) to describe initial formation of mental models, and also adopted the Piagetian distinction between assimilation and accommodation to characterize how mental models changed. Ford & Sterman (1998) described a process of identifying causal relationships first in an interview session and then representing them graphically. Pathfinder networks (Schvaneveldt et al., 1988) replace the linking phrases of concept maps with numeric similarity measures between the concepts to represent a mental model using the “pathfinder” network approach. Finally, if you have a corpus of text, Carley & Palmquist (1992) describe algorithmic ways to parse the text and automatically generate diagrams relating concepts. The Mental Model Matrix (Borders et al., 2024) is a different structural approach, focusing on capabilities and limitations of the system and the user.

In summary, there are many well-established approaches to both eliciting mental models and describing them in structured frameworks. The choices depend on many factors (echoing the same issues we discussed for eliciting trust). In particular, the method might be influenced by:

  • Distinct goals of elicitation (for design/redesign, for training, for improved performance, or for assessment).
  • The kind of system (mechanical, virtual, socio-technical, etc.).
  • The type of user (a developer, debugger, or an end-user).
  • The level of expertise among the participants.
  • Experience and practices among the research team.
  • Time and resource availability (see minimum necessary rigor in Section 3.9).

The choice of method and approach depends on the context. This section reveals that there are often many alternatives that can be considered, and that choice is likely to make a difference in what is learned.

7.6 Transparency, Observability, and Explainability

Researchers studying mental models have frequently discussed how “transparency” can support better and more accurate mental models. This sense of the word is used because many systems were (and still remain) so-called “black box” architectures in which the users had no ability to observe or understand the inner workings. Researchers described hypothetical alternatives, such as the “glass box” system (Carroll & Carrithers, 1984), which is an analogy to educational models such as demonstration internal combustion engines with glass casings that allow non-experts to understand how the system works. This allows users to “see inside” the system, and terms such as “transparency”, “observability” and “explainability” have been used to describe this.

These terms appear to be used interchangeably at times. They each describe ways that a system can be designed to help users understand how the system works, and potentially develop better mental models of its workings. Each has taken on distinct connotations within different communities. From a cognitive and human-factors approach, these are typically used to describe design elements or interface features that permit a user to develop a better understanding or more accurate mental model. Transparency often indicates interface elements that show underlying data; observability reveals the state factors important for the user to know; explainable algorithms typically add post-hoc data visualization to help understand how input features influenced a decision or classification.

These terms are often linked to mental-model development. For example, Ramaraj et al. (2019) showed how “transparent” interfaces that showed a robot's visual processing of the world and the logical processing of its task led to better mental models of human observers. Similar informational uses appear in robot-based transparency (Matthews et al., 2019), recommender and search systems (Tsai & Brusilovsky, 2019; Ngo et al., 2020; Degachi et al., 2025), and human–autonomy teaming (Scali & Macredie, 2019; Parakh et al., 2025; Bhaskara et al., 2020). This research typically uses the term transparency in purely an informational sense, looking at its influence on mental models and understanding of a system. Observability is sometimes distinguished from transparency insofar as it focuses on surfacing important information rather than just removing 'the lid'. One can visualize activation in a neural network (transparency), but a better approach might be to identify patterns that are used for specific detections, and make those observable (Zeiler & Fergus, 2014). Explainable AI (which we cover more extensively in Chapter xai) describes a series of approaches that aim to offer “explanations” for why an AI system behaves or makes a particular choice or decision. In general, explanations in this context are not usually text-based narratives describing causal reasoning, but often involve visualization, highlighting, providing examples, creating counterfactual situations, etc. In these cases, good explanations help build better mental models, leading to better understanding, trust, and performance (Hoffman et al., 2017; Hoffman et al., 2018; Hoffman et al., 2018; Hoffman et al., 2018; Mueller et al., 2019).

However, each of these terms—especially transparency—is also used in describing the social and societal impacts of automation and AI. Transparency is often used in the sense of accountability and disclosure—revealing hidden motivations, obligations, risks, and consequences of using the system, which might include data privacy, the sources of training data, the potential environmental effects, and other risks. Andrada et al. (2022) described a taxonomy of transparency notions (which we will describe in Chapter 9), and most common uses address social and societal issues of the use of automation and AI. They include informational and reflexive transparency, which is the notion most similar to how this has been used traditionally to describe understanding a system's mechanism and inner workings.

Similarly, observability is often used interchangeably with transparency, but Rieder & Hofmann (2020) distinguish between the transparency and observability for improving accountability of platforms. The rationale for this is because the algorithms themselves are often unable to be made intelligible—even to the developers—so that a better approach is a more pragmatic one where we make clear the important aspects of the decision. Walmsley (2021) suggested that in those cases in which transparency is not feasible because the inner workings of the system are either inscrutable or hard to obtain, an alternative to adopt is contestability: the right to challenge the outcome of an algorithmic decision.

Finally, explainability is sometimes considered as a way of understanding fairness: regulations sometimes require that algorithmic decisions need to provide explanations to those impacted, so that if someone is denied a loan, they can know the reasons, and especially whether the reasons were justified. These uses are related, and better understanding of an algorithm is usually a consequence, but these are not intended to improve or develop a more accurate mental model, which is normally reserved for frequent users of a system who rely on it for daily work.

These are distinct senses of the terms used when discussing the societal impacts of AI. Although one may argue that a mental model should encompass the social impacts and risks of a technology, this dilutes the terms substantially so that it no longer has a clear meaning outside of “understanding the system and its context.” It is sometimes useful to qualify the use of these terms when discussing informational aspects that lead to more complete or more accurate mental models, so that others senses of transparency, observability, and explainability are not conflated.

Understanding and mental models also interact with other concepts in human-AI interaction, such as trust, usability, and acceptance. Transparent AI systems contribute to the development of trust by providing users with evidence of reliability, fairness, and accountability. Moreover, transparent interfaces and explanations enhance the usability of AI technologies, allowing users to effectively interact with and control AI systems (Borders et al., 2024). Ultimately, fostering transparency in human-AI interaction may promotes the acceptance and integration of AI technologies into various domains of society, improving adoption and use. It may also help potential users understand the drawbacks and limitations of the system, and so lead to justified distrust and disuse.

7.7 Practical Application of Mental Models

The notion of a mental model has been used extensively to explain how users understand systems, engage in reasoning, learn about the system, and test hypotheses. In this chapter alone, we have cited several comprehensive reviews and discussions of mental models in cognitive science and human factors—each of which go into much more detail than the present chapter. But some questions arise: what is our understanding of mental models good for? If we were to abandon the notion of a mental model and just consider “knowledge” or “understanding”, would it make a difference? There are some very practical ways we can use the notion of mental models, regardless of whether they truly “exist” or are just an associated set of knowledge about a system. Carroll & Olson (1988) offered two main benefits: better design, and better user training, and Staggers & Norcio (1993) echoed these and offered a third: better performance. These are strategies for dealing with a mismatch between the user's mental model and the underlying system logic—you can either change the system, or change the user, or change the interaction (Mueller, 2020).

Using mental models for improved interface design

Carroll & Olson (1987) described several examples of early software systems that were designed to support the mental models of users, and identified design approaches that were sensitive to user mental models. One is the “naive model”: build the system to match the naive user's mental model. A second approach is the “training wheels” approach (Carroll & Carrithers, 1984), using a simplified system for novice users that maps onto a sufficient but simplified mental model. There are advantages and challenges to any such approach, especially in scaling from novice users to expertise—they note that naive users do not necessarily make for good design, and the impact of a mental model on design can be obscured by other design aspects. Furthermore, Wilson & Rutherford (1989) argued that using mental models for design must increase the total effort, because the system must be developed and used before you can elicit mental models. Thus, it may be most useful for redesigning existing systems.

Mental models for better user training

Carroll & Olson (1987) also reviewed existing research on attempting to directly train users with a mental model—either via a diagram or analogy (desktop metaphor). In comparison to more procedural training, results were equivocal—rarely did direct training on a mental model provide an advantage over other forms of learning. However, benefits were seen when an active learning approach was taken whereby the user developed a mental model through experience. In all of these examples, the mental model forms the basis for the learning goals, but is not provided directly (Mayer, 1980; Mack et al., 1983; Carroll & Mack, 1985). Consequently, for routine procedural tasks, learning a mental model may not provide substantial benefits over learning the procedure.

Other work has shown benefits for mental-model instruction versus rote learning. Kieras & Bovair (1984) tested device-model versus rote procedural learning in understanding an electronic device. Across three experiments, the mental model group learned procedures faster, retained them better, executed faster, and used shortcuts far more, and were able to infer untaught procedures. Fein et al. (1993) similarly showed an advantage for mental models in both recall and transfer learning, and Halasz & Moran (1983) showed that although there was no advantage of mental model training over procedure-based training on routine calculator problems, on novel “invention” problems there was a large advantage. These all suggest that there can be advantages to mental model training, but it depends on what is being learned, and how the mental model is presented to the user.

There has also been evidence for the benefit of direct instruction about a system via diagrams. Butcher (2006) showed that a simplified diagram of the circulatory system produced the best mental-model development, in comparison to text-only training, or text accompanying a complex diagram. The correspondence between a spatial diagram and the spatial layout of the mechanical system likely provides benefits that direct instruction about more abstract mental models cannot.

Mental models can help you understand users

If you use the mental model definitions to understand what users know about a system, this can provide an encapsulated look at their knowledge. Asking “how do you think the system works”, “what is the purpose of the system”, “what is the system doing right now?”, “what is the system good at?”, and “what is the system NOT good at” can reveal their clarity and misunderstandings.

Mental models can help identify pathways to expertise

Using a concept-map elicitation can help you (and the user) identify missing areas of knowledge, and a comparative study with experts can help identify pathways to expertise. Some evidence from this comes from Chi's self-explanation research (Chi et al., 1994; Chi, 2000), which showed people repair models via explanation, which in turn predicts learning. Furlough & Gillan (2018) showed that experience changes structure of models, not just amount of facts, and Bayman & Mayer (1984) showed that instructional models shifted novices toward expert-like conceptions, and Schumacher & Czerwinski (1992) argued mental models are critical for the acquisition of expert knowledge.

Mental models provide leverage for simulation-based training

Many researchers view mental models as a system that can run inside your head to simulate future outcomes. This can be improved by providing simulation-based training—allowing users to test how changes to “input” impact the outcomes.

Mental models permit users to encapsulate structure rather than learn rules or patterns

Conant and Ashby's Good Regulator theorem suggests that an internal model is necessary for controlling a system. It allows users to make accurate predictions of system behavior in unseen circumstances, and it involves understanding the logical and causal mechanisms that control behavior, rather than rules or patterns.

7.8 Conclusions

Craik (1943) originated the notion of internal models of the world we use for reasoning. It has been a prominent hypothetical construct within cybernetics, human factors, cognitive science, and human-AI interaction research since that time. Although the strong version of Conant and Ashby's good regulator theorem is probably not justified, we still understand that control of a system is supported and simplified when we have an internal “model” that helps us explain, interpret, and predict the target.

The rich body of research on mental models frequently takes the existence of mental models as given, and tries to use the concept to structure or interpret data about the interaction between humans and intelligent machines. There are theoretical questions about how they differ from similar knowledge constructs such as “schemas” or even simple knowledge, understanding, and rules. These questions are rarely tested by researchers who use the notion of mental models. They certainly do not map directly onto neural structures in the same way that many other mental structures are known to (e.g., object perception and identification, location representation via grid cells, mirror neurons, emotional reasoning centers, etc.). Even researchers who are working within the framework have admitted that mental models can be limited, are often minimal and incorrect. Nevertheless, the notion is powerful for capturing how humans understand intelligent and complex systems.

Moreover, they can be useful in practical ways—providing a framework for measuring what users know about a system, for improving design, for understanding performance with a system, and for developing training. Furthermore, as the use of the term has expanded, it has begun to encompass many practical aspects (limitations, work-arounds, functional understandings) that further promote these applications. Although mental models may not be “real”, they are undoubtedly useful.

Acknowledgments

This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.

Author Contributions

BF, KU, EM, CC: writing and research of individual sections; STM: Conceptualization; background research, writing in all sections; general editing.

Conflicts of Interest

The authors declare no conflict of interest.

AI Usage Statement

Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.

How to Cite This Chapter

Frisch, B., Ulinski, K., Matas, E., Cischke, C., & Mueller, S. T. (2026). How we understand AI and automation: mental models and transparency. In Shane T. Mueller (Ed.), A Handbook of Human-AI and Human-Automation Interaction. https://pages.mtu.edu/~shanem/human_ai/