Chapter 5
Trust in Automation: Definitions, Models, and Related Concepts
Abstract. Trust is a central construct in human-AI relations. Research on trust in automation has been the focus of research for decades, and has led to a number of theories and models of trust and trustworthiness. In this chapter, we discuss how trust has been defined, some of the important theories of trust, the underpinning of human-machine trust in interpersonal (human-human) trust, and related concepts. We discuss how trust has many definitions, models, theories, and components. We argue that trust in automation and AI is best understood as contextual, multi-dimensional, componential construct. The primary reason for understanding trust is to appropriately manage how new automation systems should be embedded or replace existing systems, in order to anticipate problems, but also improve processes.
5.1 Introduction
One of the most central topics in the study of automation and human-machine interaction is trust. Trust is an important topic because the fundamental exchange made in automation is gaining efficiency by allowing a machine to take over work, but in some sense the human is still being responsible for the outcome. Trust can be thought of as the thing being demonstrated when a person or organization is willing to turn over work but remain responsible. Trust (and its partner trustworthiness) has been discussed and defined broadly within the human factors and human-machine teaming literature, and in this chapter we will discuss some of the definitions of trust, popular models of trust, and theoretical approaches to understanding notions of conditional trust such as calibrated and appropriate (e.g. Lee & See, 2004), justified or unjustified (Hoffman, 2018), or warranted or unwarranted (Jacovi et al., 2021) trust. Finally, we will discuss some of the particular ways in which trust in AI systems is both similar and different to trust in other automation systems.
Our aim in this chapter is not to pick one true definition of trust, but to show how definitions, interpersonal analogies, and formal models converge on a practical problem: when should a human rely on a machine, for what, and on what grounds? We use that question to connect classic automation trust research to contemporary concerns about generative AI, AI services, and regulatory concerns of trustworthy AI.
5.2 Definitions of Trust and Trustworthiness
5.2.1 Defining Trust
Trust is not a simple term to define, although many similar definitions exist in the literature (Figure 5.1). Definitions have been proposed in many domains, including technology (Taddeo, 2009), business and commerce (Li & Betts, 2003), nursing (Meize-Grochowski, 1984), information systems (McKnight & Chervany, 2001), economics (Arai, 2009), education (Noonan & others, 2008), and philosophy (Simpson, 2012; de Fine Licht & Brülde, 2021). The meaning of trust can vary depending on the context of the term and the perspective of the person using the word. Different researchers tend to focus on different meanings of the word. Nevertheless, most definitions share properties and allow us to identify some of the common approaches to thinking about trust (Adams et al., 2003). To begin, we will examine some of the common ways trust has been defined.
Trust as an expectation
One early framing of trust frames it as an expectation about a multi-dimensional construct. Here, the constructs (persistence of the natural order, technical competence, and fiduciary responsibility) are quite broad and go far outside human-machine trust, but there is a sense in which trust is a prospective attitude about expected behavior.
Trust as an expectation has been defined as:
Trust is the expectation (E), held by a member of a system (i), of persistence (P) of the natural (n) and moral social (m) orders, and of technically competent performance (TCP), and of fiduciary responsibility (FR), from a member (j) of the system, and is related to, but is not necessarily isomorphic with, objective measures of these properties (Barber, 1983)
Trust as an Attitude
Trust may be defined as an attitude, that is a way of thinking or feeling about a particular situation (Jones, 1996). Under this view, trust may be more focused on how an individual's feelings or thoughts impact their trust in someone or something.
Trust as an attitude can be defined as:
“The attitude that an agent will help achieve an individual's goals in a situation characterized by uncertainty and vulnerability” (Lee & See, 2004).
“the attitude of a user to be willing to be vulnerable to the actions of an automation based on the expectation that it will perform a particular action important to the user, irrespective of the ability to monitor or to intervene.” (Körber et al., 2018)
Trust as a Belief
Trust as a belief involves the expectation that an agent will behave in the desired way (Reiersen, 2017). Beliefs are normally an indication that one holds something to be true (whether justified or not), so in this sense, trust accepts as true that something will behave in an expected way.
Trust as a belief has been defined as:
“… an attitude which includes the belief that the collaborator will perform as expected, and can, within the limits of its designers' intentions, be relied on to achieve the design goals” (Moray & Inagaki, 1999).
Trust as a State
Trust as a state involves the condition of the agent at a particular time. This includes how the trustee (another person, a system, etc.) will perform or act at the necessary time for the trustor (Adams et al., 2003).
Trust as a state can be defined as:
“A state involving confident predictions about another's motives with respect to oneself in situations entailing risk” (Boon & Holmes, 1991).
Trust as an exchange of risk for benefit
Many notions of trust have a sense of a willingness to put one's own consequences under the control of another, often for some benefit. When one trusts a bank, an automated vehicle, or a stranger's directions, they are accepting one sort of risk in exchange for a benefit or reduction in another risk.
Trust as an exchange has been defined as:
“The willingness of a party to be vulnerable to the actions of another party based on the expectation that the other will perform a particular action important to the trustor, irrespective of the ability to monitor or control that other party” (Mayer et al., 1995).
Trust as a measure
Many researchers use trust as a measure to assess a person's potential willingness to use or rely on something. The focus on knowing whether some intervention changes (at a subjective or behavioral level) how much someone “trusts” a system. This is an extension of thinking about trust as a state (or at least, a continuous set of states). Practically speaking, it typically involves either an explicit multi-dimensional measure of trust via agreement responses on a Likert-style scale, or a related measure of reliance–whether a user either is more willing to use a system, or at least more willing to agree that they would use a system.
There are many questionnaire-based assessments of trust. Many include a question such as item 9 of the TIA:
“On a scale from Strongly disagree to Strongly agree, rate the following statement: 9. I trust the system” (Körber, 2019).
Here, we expect a user to understand what is meant by trust (even though researchers cannot agree). This scale (and most others) typically involves asking a set of additional questions, including about a system's capabilities, its developers, the user's knowledge of it and similar systems, etc. We will discuss a number of these specific instruments in Chapter 6.
Trust as a Verb or Process (Trusting)
Trust as a verb involves the active process of trusting a person or system. Under this view, trusting can be independent or dependent of the context and situation, meaning that multiple trust relationships can occur with the same agents at the same time (Hoffman, 2017).
Trust as a verb has been defined as:
“… a continual process of active exploration and evaluation of trustworthiness and reliability, within the envelope of the ever-changing work and the ever-changing work system” (Hoffman, 2017).
Trust as a behavior
Although trust is often thought of as an internal attitude, that attitude can be assessed via behavior (Adams et al., 2003). That is, one can infer trust if one behaves as if one trusts. This could be as an economic behavior (Zheng et al., 2002) or within a game theory context such as prisoner's dilemma (Juvina et al., 2015), as investing or cooperating indicates you trust your partner, although in many human-machine settings a behavior of using the system is considered reliance. Trust as a verb has been defined as:
“the essential features of a situation confronting the individual with a choice to trust or not in the behavior of another person are.... the individual is confronted with an ambigous path.... If he choses to take an ambigouus path with such properties I shall say that he makes a trusting choice; if he chooses not to take the path, he makes a distrustful choice.“ (Deutsch, 1960).
Commonalities of Trust Definitions
This exercise illustrates that trust has many distinct senses. Nevertheless, these senses overlap, and many of these definitions have similar components that help shape the meaning of trust. Table 5.1 shows common words used in definitions of trust that may provide a better understanding of what is involved in the meaning of trust.
| Component | Role in the definition of trust |
|---|---|
| Relationship between trustor & trustee | Trust is supported by past relationship establishing a track record warranting trust. |
| Goals/Motives | Trust does not require shared goals or motives, but involves an understanding that the trustor's goals/motives are likely to be advanced through the relationship. |
| Expectations | Trust involves an expectation that the trusted party will do something. |
| Willingness | Trust is a willing relationship. |
| Risk | Trust involves risking money, reputation, or other valued outcomes. |
| Vulnerability | Trust involves making oneself vulnerable to the actions of another. |
| Non-zero gain | Often, a trust relationship involves non-zero-sum gains in which both parties benefit. |
First, trust always involves some type of relationship between the trustor and the trustee, whether it be a relationship between two humans, a human and a machine, an organization, or multiple actors. In some cases, a trust relationship may be between a present and future party, as is often the case in a legal trust used to manage assets. Trust is often second-hand: a builder or plumber may be trusted based on the recommendation of a friend. Trust involves some form of goal or motive to reach or some expectation of an outcome, normally with some sort of exchange of risk, vulnerability, and benefit. This does not mean that both parties have shared goals, but rather that the trustor's goals can be advanced through the relationship, such that putting trust in something involves the expectation that the desired outcome will be achieved. Similarly, there has to be a willingness to trust the trustee in order for any trust to be established. This willingness entails the potential for risk, and the outcome may not work out in the way intended, leading the trustor to be vulnerable to the behavior of the trustee. However, there is often the prospect of a non-zero-sum outcome, where both parties may benefit from the trust relationship.
We will not pick or create a new definition for trust. We simply recognize that when people talk about trust, they can mean different things, and an important initial question to ask is “what do you mean by trust?”
5.2.2 Defining Trustworthiness
In the most basic of definitions, the term “trustworthy” can be defined as the ability to be trusted, or being worthy of trust, but this is almost a circular definition. Several of the models we discuss in the next section characterize properties of trustworthiness. For something or someone to be trustworthy, the characteristics of trust must be earned (Jones, 2012). This involves the acceptance of the aspects of trust, such as risk and vulnerability that come with trusting in order to get to the expected outcome (Hardin, 1996). Thus, trust often is a behavior or attitude of the trustor, whereas trustworthiness is often thought of as a property of the trusted entity.
In terms of the relationship between the trustor and the trustee, trustworthiness can be defined as the trusted acting in the way the trustor expects (Hardin, 1996), perhaps in line with an agreement or contract (Jacovi et al., 2021). If a system acts in the way it is supposed to, it will likely act the same way again, leading to the conclusion that the system is trustworthy. The perceived trustworthiness of a system, especially a new system, may also be influenced by external factors such as institutional backing, independent oversight structures, etc. (Hardin, 1996; Shneiderman, 2020). For example, if a new in-home AI system is released from a reputable company known for protecting user privacy and data, consumers may find it more trustworthy than if the same system were released by a different company. This all is involved in how people come to trust a system and deem it trustworthy and thus may adopt it into their lives.
Interestingly, although trust and trustworthiness go hand-in-hand, many theories of trust in automation appear to focus on trustworthiness–things that make an automated system able to be trusted, such as its reliability or predictability. In a later section of this chapter, we will discuss some of these models, which often are intended to predict how both trustworthiness and other factors lead trust to change.
5.3 Human–Human (Interpersonal) vs. Human–Machine Trust
One of the common themes of research on trust in automation is its basis in human-human (interpersonal) trust. As discussed in Chapter 4, one consequence of social robotics and anthropomorphic design is that it tends to enhance trust in the system (although the human-like features may only impact perceived trustworthiness and not actual trustworthiness). Consequently, it is useful to understand some of the existing theories on interpersonal trust such as Mayer et al. (1995)'s theory of organizational trust. While a number of researchers have based their approaches to understanding trust in machines on interpersonal trust (e.g. Muir, 1994; Adams et al., 2003; Lee & See, 2004; Hancock et al., 2015; Hoffman, 2017), it is important to note that almost all argue that this is a good starting point but insufficient, and have elaborated additional aspects for understanding trust and trustworthiness in machines (e.g. Jacovi et al., 2021). Thus understanding interpersonal trust is a useful starting point for understanding human-machine and human-AI trust.
5.3.1 Theories of interpersonal trust
There are many different beliefs on why humans form trusting relationships. Mayer's theory of organizational trust suggests human-human trust develops for two distinct reasons: because we have the disposition to rely on others (probably passed down via genetics and culture) and because others demonstrate that they can be relied on (Mayer et al., 1995; Baer & Colquitt, 2018). On their own, each of these concepts have a limited ability to explain trust, but together to form a more complete picture of why humans trust. This theory posits a cognitive and affective basis for trust, both of which benefit from time knowing the trustee. The theory does not address trusting when in short-encounters, such as when trusting a stranger.
Other theories of why we trust arise from an evolutionary perspective. For example, Montague et al. (2015) identified the Three R's of trust: reaping, regarding, and recursive modelling. By reaping, they mean trust is responsive to reward and punishment. Regarding involves pro-social mechanisms that encourage trust. Recursive modelling represents a theory-of-mind or game-theoretical aspect, where our trust is reciprocal if it is reciprocated. It supports coordination when each party can reason about the other's reasoning, allowing temporary alliances even when short-term payoffs favor defection. From an evolutionary perspective, trust can look irrational from a short-term self-interest view, but (Fehr & Gächter, 2002) demonstrated that cooperation can be enforced by altruistic punishment–where punishment is costly and yields no material gain.
Others have developed more biologically- and neurologically-focused theories of why people trust. One theory argues that the release of oxytocin in the brain leads to trusting relationships (Kosfeld et al., 2005; Nave et al., 2015). Oxytocin is a neuropeptide produced in the thalamus known to promote social behavior, including bonding, maternal behaviors, and sexual behaviors. Since so much of trust is based on social behavior, oxytocin is an important factor in creating trust, and it has been shown to help in overcoming aversions to the uncertainty of the behavior of others. And the motivational and emotional brain circuits that are implicated have consequences in emotional responses that can drive behavior. For example, one may feel guilty about not trusting someone (Crockett et al., 2010), so trusting is favored because those negative feelings are avoided.
These three levels of analysis are not necessarily at odds: evolutionary benefits have accrued to those who both have a willingness to trust but are sensitive to reward and punishment based on the outcome of the trust; this behavior may rely on specific neural signaling that was adapted from core caring and bonding behaviors in animals, but is still reflected even at the organizational level. Interestingly, unlike Mayer's theory, the evolutionary and biological theories focus primarily on the trustor. This has two implications for trust in automation. First, to the extent that our trust is derived from neural computations related to trusting other people, it is likely that behaviors that signal trustworthiness co-evolved with how we detect trust, and so aspects of interpersonal trust may not translate to human-machine contexts. It is also likely that, to the extent that trust has evolved, machines and AI can exploit the signals of trust. As discussed in the previous chapter on anthropomorphism, anthropomorphic aspects of robotics and AI such as cuteness have been argued to be “Dark Patterns” (e.g. Lacey & Caudwell, 2019). Of course, these patterns merely reflect other ways in which other humans exploit trust, and so it may lead to an overall distrustful stance toward novelty. Overall, even if trust mechanisms are formed through evolution, they are also likely to be influenced by social learning and demonstrations of (un)trustworthiness, and so many aspects of interpersonal trust are likely to translate to human-machine contexts.
5.3.2 Forming Trusting Relationships
From the evolutionary perspective, infant animals who appropriately trust and distrust are certainly more likely to survive. In humans, the disposition towards trust is thought to begin in children 6 to 18 months old (Rotter, 1971). As a parent meets the needs of an infant, the infant may be able to generalize trust within and outside the family unit (Sakai, 2010). While a disposition towards trust can be formed, trust in interpersonal relationships is based on each relationship–suggesting we are well equipped to compartmentalize our trust and not simply be trusting or distrustful overall. One theory suggests that for every relationship, there exists a “zero baseline” of trust where a person neither trusts or distrusts another (Rempel et al., 1985), and through experience trust can move from the baseline. However, (Borum, 2010) argues that people enter a situation with a higher or lower starting level depending on multiple factors including their psychological and cultural context. This suggests that there may be strong individual differences in a baseline trust stance toward others (and presumably toward machines).
5.3.3 Implications for Human–Machine Trust
Humans are driven by social interactions and social expectations. Even with technology, people may perceive technology as more human-like than system-like (Lankton et al., 2015), and these different aspects of trust may operate differently for different kinds of systems (e.g., human-focused facebook vs. a system-focused database). Therefore, people may interact with technology as if they are interacting with another human, as opposed to a machine (Reeves & Nass, 1996). However, trying to understand human-machine trust through the lens of interpersonal trust may also lead to incorrect assumptions about how trust with machines form (Hoffman, 2018). Aspects of interpersonal trust, such as humility and leadership, are not considered in the literature of human-machine trust. The type of machine being used needs to be considered when studying how the trusting relationship is formed and works.
Humans can also specifically trust or distrust the algorithms that are being used by AI and ML. Pariser (2011) coined the term “filter bubble” to describe how algorithms prominent on social media are providing each of us an individually-tailored view of the world, but as shown by Klug & Strang (2019), the most intense users of social media recognized this more often and have a more negative perception of the algorithms. Cabiddu et al. (2022) presented six main determinants of trust in algorithms: users' propensity to trust; IT acceptance levers; human-like characteristics the AI presents; social influence; familiarity; and the system characteristics of the AI. Similarly to human-human trust, some of these dimensions can change over time. Since AI appears to learn and problem-solve, people may be more likely to trust AI over other technologies. AI's ability to be humanized lends itself well to fostering trust. However, when an algorithm has made a mistake, people may experience algorithm aversion. Algorithm aversion is the reluctance of humans to use algorithms, which are superior but imperfect, in favor of their own judgment (Burton et al., 2020). This suggests that trust in AI and algorithms can be destroyed, perhaps more easily than it can be restored.
5.4 Theories and Models of Trust in Automation
Numerous authors have proposed models or theories that attempt to explain trust, at a number of different levels of abstraction. Many of these were reviewed by Adams et al. (2003), which remains the most comprehensive scholarship on trust in automation. This section will summarize some of those models, and introduce some other prominent models that have emerged since that report, with the goal of providing the main motivations and approaches of these different models rather than identifying particular commonalities or differences between models.
5.4.1 Conceptual Models of the Predictors of Trust(worthiness)
Muir (1994) proposed maybe the first framework or model used to describe trust in automation, based on previous social notions of trust by Rempel et al. (1985) who identified dynamic properties related to predictability, dependability, faith, and by Barber (1983), who identified expectations such as persistence, technical competence, and fiduciary responsibility. This model might be thought of as a causal models of the predictors of trust, which is useful because it can help system designers understand how they might change automation or AI to improve trust. For example, recent advances in large language models have been widely criticized for being unpredictable and undependable insofar as they can produce different output to the same prompts, high-confidence hallucinations, and bad advice based on low-quality training sources. These models would likely predict this to lead to a loss of trust.
More generally, the goal of creating understanding levels of trust as predicted by aspects of a system and aspects of the user or operator has been the basis for many other models. For example, Llinas et al. (1998) and Seong & Bisantz (2000) adapted this basic notion, but extended it using the Lens model (Brunswik, 1943). The Lens model has been used in human factors research to understand judgment in socio-technical systems about multiple cues of differing validity (e.g., Mosier & Kirlik, 2004). In the context of trustworthiness of a system, it is used to examine how trust judgments are supported by a number of characteristics of a system. The lens notion is useful because these characteristics need to be observed by a user, and the cues of those characteristics might be weak or unreliable signals. As shown in Figure 5.2, the system on the left has a number of observable signals of its trustworthiness. An operator or user (on the right) may not clearly know these signals, or not have experienced them. If the weights on the right match the signals on the left, the model considers trust to be 'calibrated'. However, a trustworthy system may not always provide observable behavior to help the operator develop this trust. Presumably, many of the observable characteristics overlap with those proposed by Muir (dependability, reliability, etc.), but the key insight of this model is that trust is developed through the interaction of information signaled by the automation and interpreted by the user, but they may not accurately use these signals as a basis for their trust.
5.4.2 Models incorporating the user's influence on trust
The specific approach of Muir was influential, insofar as it framed the problem of understanding trust as understanding the factors that influence trust by a user of automation. Both this specific model and the general approach has been adopted and extended to incorporate aspects of the user as well. For example, Lee & Moray (1992) incorporated a third dimension attributed to Zuboff (1988), who found that understanding and trial-and-error experience is related to predictability and reliability (see Table 5.2).
| Barber (1983) | Rempel et al. (1985) | Zuboff (1988) | |
|---|---|---|---|
| Purpose | Fiduciary Responsibility | Faith | Leap of faith |
| Process | Dependability | Understanding | |
| Performance | Technical competence | Predictability | Trial-and-error experience |
| Foundation | Natural Laws |
Others have expanded on Muir's basic notion by identifying other kinds of taxonomies. For example, Madsen & Gregor (2000) proposed a model of human-computer trust and included two main themes: cognition-based trust (incorporating understandability, competence, and reliability), and emotional/affective responses of the user (personal attachment and faith). Kelly (2001) proposed a model with three main themes: one incorporating the so-called 'competence' of the automation reminiscent of Muir dimensions (dependability, reliability, usefulness, and robustness), and two additional themes related to the user: understanding (predictability, intention, and familiarity) and self-confidence (faith, reputation, skills, training, experience). These last two frameworks importantly incorporate elements of both the automation and aspects of the user.
Consequently, many models of trust incorporate a 'user' or 'trustor', as well as the system or 'trustee'. These entities form the basis for more sophisticated models of trust, but there are several that can be considered systems models attempting to capture a bigger picture.
5.4.3 Systems models of trust
A number of models have been proposed that might most accurately be described as systems models: accounting for a variety of aspects of trust, including the precursors, consequences, interactions, decisions, users, and technology. These typically focus on three major components: aspects of the trustworthiness of the system, aspects of the trustor, and consequences in terms of decision, uses, or reliance.
Lee & See (2004) model of Trust in Automation.
Lee & See (2004) proposed a system model of trust that incorporates individual, organizational, cultural, and environmental factors (Figure 5.3). It is structured as a general control loop involving automation, display, belief formation and reliance, with different factors influencing different components. This loop can be thought of as controlling the 'appropriateness' or 'calibration' of trust: the operator decides to rely on the automation for some purpose, and the automation acts and displays results, which then lead the operator to evaluate whether it behaved appropriately and adjust their willingness to trust the system later.
Adams et al. (2003) Automation trust model in military context.
In their extensive report on trust theories, models, and measures, Adams et al. (2003) propose a preliminary model of trust in automation for military contexts (Figure 5.4). The model is organized around a trust development process: perceived risk is weighed against a contextual need to trust, producing a trust decision that may change psychological state, support reliance, yield an outcome, and feed back through trust calibration. Antecedents enter that process from three sides—qualities of the trustor (current state), qualities of the automated system (including predictability and dependability that emerge with experience), and operational/organizational context.
This model was developed to characterize common themes of trust in technology in military contexts. Although this is true for other industries as well, in military settings, there is a constant stream of new systems including new automations that are often not developed with the warfighter's direct involvement or feedback. These systems often fail to be adopted, often because the capability is lacking, but also because the end-users lack trust in the systems in comparison to their older, less capable but proven systems. Even though in theory, adoption can be dictated from higher-level command, the individual units and warfighters often have the opportunity to weigh whether adoption and use is warranted in their own use context, and may simply work around the new system. The model highlights how there are many ways technology adoption can go wrong, and also many vectors to improve trust and adoption of systems that are genuinely useful or necessary to handle new environments.
Although this was designed for understanding trust in a military context, it is likely to be broadly generalizeable to other settings, including industrial automation and consumer systems.
Mayer-Hancock Trust Framework
Hancock and colleagues (Hancock et al., 2011; Schaefer, 2016; Hancock et al., 2021) adapted Mayer et al. (1995)'s model of organizational trust to understand the elements of human-machine trust. This model involves the perceived trustworthiness of the technology, characteristics and perceptions of the trustee and trustor, and examines outcomes of trust. Consequently, the model has high overlap with many other models of trust. One substantial advance made recently (Hancock et al., 2021) is that they used this model as a basis of a large meta-analysis including 338 studies and over 2000 effect sizes. This model enabled elaborating and validating many of the putative precursors and consequences of trust identified by other researchers. Furthermore, this elaboration helped focus the authors on three main aspects of trust: those related to the trustee, the trustor, and the context. Hancock et al. (2021) also examined how these factors impacted distinguished between directional trust (i.e., supervisor/supervisee relationships); see Figures 5.5 and 5.6.
As with other models, this model organizes the trustee and trustor; it identifies influences of trustworthiness and trust, and captures larger-scale contextual concerns. Overall, the model was intended to promote 'appropriate' levels of trust in robots, and so even if the notion of trust calibration is not central to the model, it is the overall goal.
5.4.4 Process/cognitive models
These models typically focus on the cognitive processes that underlie development and revision of trust, decision processes that lead to use and reliance, the kind of information used to inform trust, and the way in which trust gets adjusted or calibrated based on evidence. In fact, it appears that the common thread of process and cognitive models of trust is that they focus on how information and evidence is used to adapt (or "calibrate") trust to an automation system, context, or application.
Muir model of the relationship between automation, trust and predictability
Muir (1994) proposed what might be considered a process model of trust, in terms of a flowchart of processes and decision (Figure 5.7). Here, an operator forms a mental model of a system related to the three principal factors identified by Barber (1983), as discussed further in Chapter 7 (Mental Models and Transparency), which work together to form a basic degree of trust in the automation. One might consider whether this internal model is calibrated with the actual trustworthiness of the system (the red boxes at the top), but more importantly, the mental models enable a prediction of the system's behavior, which can be compared to observed behavior (tempered by confidence in one's own predictions) to understand whether confidence in the system is calibrated. This system essentially suggests an adaptive model by which behavior and mismatch between expectations and observations are used to adjust trust. This model also incorporates the good regulator theorem (Conant & Ashby, 1970) discussed in Chapter 2, insofar as it posits that the operator/user incorporates a mental model of the system. Presumably, although the model is focused on trust, it could also be adapted to understand the operator's ability to use and “regulate” the system.
Cohen et al. (1997) proposed the Argument-based Probabilistic Trust (APT) model: a conceptual account of trust as the outcome of an argumentation process under uncertainty, intended to be specific enough to support training design (see Figure 5.8). APT is inspired by Toulmin (1958)'s argumentation model, using it to structure a formal argumentation process with a claim (the putative conclusion), grounds (facts in support of the claim), warrant (beliefs supporting causal conclusions), backing (assumpitons supporting warrant), and rebuttals (counter-arguments). In APT, this is used to consider a probabilistic process, where the claim gains qualified or probabilistic support. The specific aspects of the argument structure are linked to trust in a decision aid, so that grounds indicates awareness of current system and situation features that support a qualified claim (the chance of correct system action over a period \(t\)), via a warrant (belief that those features are generally tied to performance) that is itself supported by backing (assumptions, experience, and design knowledge). Rebuttals are ways that estimate could be wrong. Furthermore, they identify (via dashed lines) many ways that uncertainty, completeness, and reliability can impact support for the claim.
Together, these characterize uncertainty about system quality, but it enabled Cohen et al. (1997) to formally define probabilistic models in event trees, to understand how trust in a decision aid can unfold over time. Furthermore, they also used it to identify specific training and decision aiding strategies: uncertainty of grounds can be improved through aids that help monitor appropriate information; improved training of mental models can impact backing and warrant; improving critical thinking skills to improve rebuttals or counter-arguments about the trust estimate, etc. As Adams et al. (2003) note, APT's main value is charting how trust varies across users, decisions, situations, and phases of decision-aid use—not treating trust as a single static score.
An ACT-R Computational model of trust
A series of papers by Juvina and colleagues (Juvina et al., 2015; Juvina et al., 2019; Collins & Juvina, 2021) described a computational cognitive model of trust during repeated game-theory 'strategic interaction' games (Prisoner's Dilemma, Chicken, and a multi-arm trust game).
They developed a computational cognitive model using instant-based learning (IBL), which maintained memory for past interactions of the user and their partner, and a accumulator model that tracks elements of the game. Unlike any of the other models of trust we review here, this model makes specific predictions about behavior that were generally consistent with human research. A related advantage is that, by the nature of the game, data involve hundreds of repeated game outcomes (e.g., cooperate vs defect) that can be used as a behavioral proxies for trust. Of course, the limits of this model are that it is narrowly-focused on a handful of simple and repeated strategic games, and so it does not need to incorporate many of the elements of trust and trustworthiness that are important in more natural contexts
5.4.5 Statistical and Predictive models of trust
Researchers have developed more formal statistical and machine learning models that have been used to either test or discover some of the conceptual elements of trust. For example, Muir (1994) conceptualized her initial model on experiments using regression analysis, and Lee & Moray (1992) developed similar regression and autocorrelation models to examine the dynamics of trust.
Following Barber's (Barber, 1983) conceptual framing of trust as an equation
where \(i\) refers to the trustor, \(j\) refers to the trustee, \(T\) is trust, \(P_n\) and \(P_m\) is moral and natural persistence, \(TCP\) is technically competent performance, and \(FR\) is fiduciary responsibility. She expanded this to a linear model as follows:
where \(B\) are all parameters, \(X_1\) is persistence, \(X_2\) is TCP, and \(X_3\) is FR, which captures all the main effects and interactions.
Furthermore, Muir & Moray (1996) proposed a new trust equation based on a stage model (Rempel et al., 1985; Muir, 1994):
which they tested in a regression model on subjective ratings of each of these constructs predicting ratings of trust in different targets in the simulation (three properties of three different pumps), during a simulator experiment in which users controlled a pasteurization process.
They found that these new predictors did not improve the prediction of trust. They also used similar modeling to establish that trust ratings were related to the time spent in “auto-pump” mode, indicating a close link between stated trust and reliance; and showed that the variability in the control error of the pumps was a strong predictor of trust.
This regression modeling approach gains traction and by collecting many subjective ratings during a simulation, and shows how these are related. Even though the assumptions of this approach are not justified (e.g., linear combinations, independent predictors, and that trust is a value) Lee (2013) used a more sophisticated but similar machine learning approach to examine the factors in a human-robot interaction that led to greater trusting behavior, specifically predicting the outcome of an economic game (rather than subjective trust.) These trust equations are useful insofar as they represent tacit assumptions in many of the more complex trust and trustworthiness models, and thus allow empirical testing of those assumptions. However, it may be a blunder to take them as a serious theory of trust per se, but rather treat them as a interpretive model of the processes and influences of trust.
5.5 Contributors to trust and trustworthiness
As we have seen, many of the models of trust are really catalogs of factors that impact trust (if they are about the trustor) or trustworthiness (if they are about the trustee/automation or AI), as well as external/contextual contributors ((Hancock et al., 2021); see discussion in previous section). These factors, such as perceived trustworthiness, the AI system's ability to perform tasks predictably, the benefits and costs of trusting these systems, and legal regulations, are essential in facilitating the design, implementation, and assessment of trustworthy AI systems.
Perceived trustworthiness is a significant contributor to both social and human-machine relationships. In social relationships, individuals rely on various cues, such as past experiences, communication style, and reputation, to assess another person's trustworthiness. Similarly, in human-machine relationships, users assess the perceived trustworthiness of AI systems based on their experience with the technology. Trust in machines is frequently shaped by considerations such as accuracy, consistency, reliability, and the system's ability to consistently deliver expected outcomes (Adams et al., 2003). When users perceive an AI system as trustworthy, they are more likely to rely on it and maintain a positive relationship with it.
The AI system's ability to perform tasks predictably is another critical factor influencing trust in human-machine relationships. Humans value predictability in their interactions with AI systems, as it allows them to understand and anticipate the behavior of the technology. When an AI system consistently performs tasks as expected, users develop a sense of trust in its capabilities (Lee & See, 2004). On the other hand, uncertainty and unpredictability can erode trust. Therefore, designing AI systems that prioritize predictability by setting clear expectations and providing transparent feedback enhances trust in these relationships.
Trustor-related contributors, specifically the benefits and costs of trusting AI systems, also affect trust in human-machine interactions. Trustors evaluate the potential advantages and disadvantages associated with relying on an AI system. Benefits such as efficiency, accuracy, and convenience can increase trust as users perceive that the technology improves their overall experience. However, if the costs of trusting AI systems, such as privacy concerns or the potential for errors, outweigh the benefits, trust can be undermined. Therefore, a thorough understanding of the trade-offs involved in trusting AI systems is vital to fostering trustworthiness (Lee & See, 2004).
Contextual contributors (Hancock et al., 2021), including legal regulation, also influence trust in human-machine relationships. Laws and regulations provide a framework that governs the behavior and actions of AI systems, ensuring accountability and transparency (Adams et al., 2003). Legal safeguards can enhance users' trust in AI systems by assuring them that their rights and interests are protected. The presence of legal regulations can establish a strong foundation for building trust in the technology, as users perceive it to be governed by ethical and responsible practices.
5.6 The flavors of trust
5.6.1 Mistaken assumptions about trust and trustworthiness
Hoffman (2017) argued that trust, trustworthiness, and reliance are often mistakenly assumed to be an end-point stable state, which is a mistake. Furthermore, he criticized the assumption in many trust calibration models that it is a single value measured on a single scale, arguing trusting is a process, there are different kinds of trust, and it depends on context. These mistaken assumptions work together to form what we will refer to as a control-theory of trust–trust is a unidimensional continuous value that is a holistic property of a system that applies in general to its use. This involves four specific assumptions that are unwarranted:
Trust is assumed to be unidimensional. Instead, trust can be pluralistic.
One common faulty assumption is that trust is unidimensional–that trust is a single property of the relationship between two entities (Hoffman, 2017). As we will discuss in Chapter 6, most measures of trust assume trust is pluralistic–many different aspects of trust exist, so that trust, reliance, understanding, adoption, and other aspects all make up one's trust state (Hoffman, 2017).
Trust is assumed to be continuous and metrical. Instead, trust can be categorical.
Many models of trust implicitly or explicitly take trust to be a continuously-varying concept–you can trust something to a larger or lesser extent (also Hoffman, 2017). This is implicit in the notion of “trust calibration” and implies a control-theory view of the trust process. It is true that people can rate their level of trust at some finer level, but it may not accurately describe a behavioral version of trust where trust is indicated by cooperating or defecting. When he offered Jasmine his hand to get on the magic carpet, Aladdin famously asked Jasmine “Do you trust me?”, not “How much do you trust me?”. By encouraging one to consider trust as finely-graded, it may avoid thinking about the other ways in which trust is not all-or-none (contextual and componential trust).
Trust is assumed to be global. Instead, it can be contextual.
Another mistaken assumption is that trust (or perhaps trustworthiness) is universal. Instead, it is bound to a context—it can depend on when and where and the environment the trust relationship occurs. For example, new drivers are trusted to operate a vehicle during daylight hours but not at night; automated drivers may be trusted in low-traffic and well-mapped areas but not on snowy roads or in off-road settings. Here, trust is a three-part relationship: A trusts B to do C (Baier, 1986; Hardin, 2002). You (A) trust your vehicle autopilot (B) to drive on snow (C).
Trust is assumed to be holistic. Instead, it is componential or mereological.
Finally, trust and trustworthiness are assumed to be holistic: your trust state applies to the entire entity you are trusting. Instead, trust is componential or mereological–you might trust one part, function, or property of a system differently from another. Models such as Muir (1994) explicitly suggest that the trustworthiness of some of these different properties compose the trustworthiness of the entire system, but it can be much more useful to understand what someone trusts, and what they do not trust about a system. Furthermore, users may be forgiving for errors they can explain if they know that the untrustworthiness of a system is limited to specific functions.
Altogether, the control-theoretic trust calibration model implies that an appropriate level of trust is an end-state of system design and learning. However, an alternative perspective does not ask “how much should I trust the system”, but rather asks “what aspects, how, when/where and for what should you trust the system”. This is illustrated by the comments of one Tesla FSD driver interviewed for the research described by (Ibne Mamun, 2023; Ibne Mamun & Mueller, 2024).1 He said that there were things he could trust the car to do 100% of the time, and these are things he would allow it to do whenever he operated the vehicle. But there were things that it would do correctly less often and he would take over or avoid those situations. And there were things that were even less reliable, and those he would only try when his wife was not in the car. Here, trust is about different things (context, multiple dimensions, and different components), and levels of trust are really categories about probabilities of success, (categorical and not numerical). Furthermore, for him, lack of trust did not mean he would not rely on it, but reliance depended on the context of his passenger's risk tolerance.
This contingent notion of trust has been considered extensively in the past, as we will cover in the next section.
5.6.2 Appropriate, calibrated, justified, and warranted trust
“Appropriate trust” Lee & See (2004); Hancock et al. (2021) might be the most general way of discussing how our trust in a system should somehow correspond to the trustworthiness of that system. However, the term trust calibration has been used for decades (Muir, 1994; Llinas et al., 1998; Seong & Bisantz, 2000; Adams et al., 2003; Lee & See, 2004) to describe how trust might be matched to trustworthiness. This notion arguably incorporates the tacit control-theory assumptions that trust is holistic, uni-dimensional, continuous, and global, which we argued against in the previous section. As an abstraction, this can be helpful to consider processes by which trust and trustworthiness align. It was used by Lee & See (2004) to help guide their model of trust, along with appropriate trust, overtrust and distrust to describe different trust calibration regions, with the prediction that overtrust can lead to misuse, and distrust can lead to disuse. However, they also distinguish between calibrated trust and appropriate reliance–the appropriate use, adoption, or deployment of a system based on its trustworthiness. Ultimately, reliance is more important–trust is a stance about a system, but reliance is actual use.
Nevertheless, the control-theory metaphor begins to fail when we consider contextual, componential, multidimensional and categorical aspects of trust. Hoffman et al. (2013); Hoffman (2018) proposed an alternative approach: justified or unjustified trust and mistrust. The similar notion of warranted or unwarranted trust was used by others, such as Jacovi et al. (2021). In this framing, trustworthiness is not a single value that needs to be adjusted to match. Moreover, a single automation or tool can be both trusted (about some things) and distrusted (about others). This suggests a more qualitative analysis in which one examine which aspects are trusted and distrusted, and whether these done for good or bad reasons. This can helps focus human factors researchers on contextual, componential, and multidimensional aspects of trust. Instead of asking “how much do you trust the system”, you may discover what parts of the system are trusted, at what times and in what contexts, and why.
5.6.3 Hoffman (2017) Taxonomy of trusting
Along with justified and unjustified trusting, Hoffman (2017) identified many other kinds of “positive” trusting relationships, including absolute trusting, tentative trusting, skeptical trusting, contingent trusting, stable trusting, progressive trusting (trust improves with experience), digressive trusting (trusting deteriorates with experience) swift trusting, and default trusting, faith-based trusting, authority-based trusting, and over-trusting (i.e., automation bias). However, there are also negative varieties of trusting: mistrusting, distrusting, anti-trusting (belief that the machine will do things not in the human's interest), and counter-trusting (machine is providing evidence that it cannot be trusted). These underscore the variety of trust situations that can occur, and especially highlight how trust is often dynamic, both in how it begins, and how it changes with experience and time.
Many of these might just be thought of as dynamic patterns that might emerge in under the control-theory perspective on trust: initial set points that occur for many of the reasons the trust calibration models cite for impacting trust or trustworthiness (e.g., as Hancock et al., 2021), along with dynamics over time as new information is learned, systems exhibit strengths and weaknesses.
5.6.4 Ways to impact trust
Default trust.
As with humans, some researchers have suggested interacting with AI starts with some amount of default trust. Unlike human-human interactions, human-machine default trust does not consider that trustworthiness hinges on the partnering agent being a computer (Smith & others, 2017). This can be a degree of trust or distrust depending on the individual. At default, the user may have zero expectations of a system and have little-to-no reason to expect it to work against expectations. To this end, the user may begin using the system with blind trust until there is a rapport built over time. With experience in dealing with intelligent systems, a person can also start with some default distrust in the system and expect it to under-perform the user's expectations. As previously mentioned, trust is dynamic and so is default trust; not all users start interacting with a system and expect it to perform perfectly.
Changes in trust through experience.
Most models of trust suggest that justification (or calibration) of trust comes in part the form of repeated interactions that build a sense of trust/distrust. To make a proper justification, the user must provide a series of test cases to determine if a system can provide adequate output.
Expertise in trust.
In the extreme, extended experience can build expertise. Expertise with a device should increase trust in the system though understanding of how it works (Ferrario et al., 2020). As the operator spends time working alongside an autonomous system, they should gain an understanding of its capabilities and modify trust accordingly. A system that cannot provide consistent and predictable output will likely not be trustable.
Transparency and explainability.
Explainable AI (Chapter 8) can also increase trust in a system and provide a framework for justifying trust or distrust (Akula & others, 2019). If a system can provide justifications of its actions and the user can interpret these explanations the user might trust the system in similar scenarios. As with humans, prolonged use can increase or decrease trust and one extraneous event should not ruin the user's expectations of the system; one failure would not destroy the foundation of trust in a system.
5.7 Trust in “Artificial Intelligence”
As Chapter 1 emphasizes, AI systems vary widely in capability and role. Much of the research and models of trust were developed on older automation systems such as power control systems, air-traffic control automation, commercial air pilot automation, health care applications including warning/alert systems, and factory automation. AI can often serve the role as automation–albeit with more complex algorithms. In those cases, many of the same concerns apply. Furthermore, even when AI is not being used to automate a process and rather augment cognitive work, many of these same issues exist. The nature and complexity of modern AI systems (including generative AI, large-language models, medical diagnostic systems, autonomous control systems, etc) means they can have many more functions than most simpler automation systems involved. Determining the important aspects of trust in AI is critical as it becomes more prominent in society and new AI models and systems continue to emerge. We will first describe the fast-growing domain of research focusing on “trust in AI” in particular, without much of the theoretical or methodological traditions we have explored in the trust in automation domain. Then, we will examine how the fields that emerged to study trust in automation embedded into work processes (e.g., cognitive systems engineering, human-centered AI) have assimilated modern AI systems into their research. Finally, we will briefly review the emerging legal, ethical, and societal concerns about AI that go by the term “trustworthy AI”.
5.7.1 Trust-in-AI research in comparison to trust-in-automation
Although AI systems of various types have been in the field for almost 50 years, the recent growth in capability of AI has generated research on AI specifically as a target of trust. Prior to the wide adoption of large language models, Glikson & Woolley (2020) reviewed empirical research on human trust in AI across management, HCI, and engineering. They organized their findings along two dimensions: how AI is represented (embedded in software, virtually embodied as an agent, or physically present as a robot) and how much autonomy it exercises (as a tool, assistant, or manager). Trust requirements and antecedents differ across these: embedded systems involve perceived performance, transparency, and control; virtual agents require calibration between anthropomorphic presentation and actual capability; and physical systems on depend on responsiveness and shared situational awareness. They argue that expectation–performance calibration matters more than a single trust score, and that lab studies often show optimistic initial trust fails to replicate when fielded, such as for embedded and robotic AI. Their analysis is relevant to modern generative AI systems, as they fall within the embedded-to-virtual range and the tool-to-assistant autonomy region. One prediction of this framework is that, for these systems, higher autonomy and less visible failure modes will require justified, contextual, and componential trust rather than blanket holistic reliance.
The literature specifically labeled “trust in AI” has grown quickly since that review. Some of this research has discarded or been ignorant of the theories, models, or measures developed for automation and focused on new and unique aspects of trust-in-AI. Afroogh et al. (2024) reviewed definitions, challenges, and open questions and concluded that the field still lacks a shared construct: papers mix trust, trustworthiness, acceptance, and ethical “trustworthy AI.”; and these are probably clearer in the trust-in-automation domain.
Another shift in focus in this field has been in research on general trust in AI in comparison to interpersonal trust. In light of the discussion in Chapter 4 about anthropomorphism, it may be the case that trust in AI chatbots becomes more like interpersonal trust than has been found for trust in general automation. However, Montag et al. (2023); Montag et al. (2024) found that self-reported general trust toward humans (i.e., as a personality trait) and trust toward AI (both generally and with respect to specific classes of products) were essentially uncorrelated, and even found different brain activation patterns for these cases. This suggests that default trust stance toward other people does not carry over to AI systems, although Gillath et al. (2021) showed that attachment-related individual differences can influence trust in AI, and that experimentally priming attachment security can increase rated trust in AI. In any event, these results neither support nor refute the assumption that trust in AI is built on the same sorts of cognitive and psychological mechanisms and thought patterns as interpersonal trust. Consistent with many of the models of trust we reviewed here, trust is not a blanket stance toward the target, but is about distinct functions and activities, within a context, in service of a goal. These results merely show that default stances toward others and toward AI are unrelated, even though there appears to be an attachment-related common pathway.
A second shift is that a large fraction of “trust in AI” research is really about acceptance of AI as a product. Even though one of Muir's (Muir, 1994) primary concerns was about whether the “supervisor” (i.e., operator) would accept or reject the automation, this was normally in the context of a system that had automation embedded within it, but the user had autonomy to use it or do the work themselves. Much of the recent work on trust in AI is framed within a technology acceptance model, which normally is used to follow products through initial contact where potential users become aware of the product and may recognize its perceived usefulness, which may lead to eventual adoption. For example, Choung et al. (2023) had users rate trust into in facets: human-like facet (benevolence and integrity) and functionality (competence and reliability), and examined these facets as precursors to in a technology-acceptance model. Both facets predicted intention to use, although functionality was a more important predictor. Gillespie et al. (2023) asked how much people trust “AI” in the abstract (often by sector: health, hiring, banking). Those items captured general willingness to trust hypothetical systems, helping understand the public's willingness to consider these applications in the abstract, and were not direct assessments of AI tools.
A final recent result about trust in AI showed that by 2025, users had begun to develop a distrustful stance toward GenAI Li & Aral (2025). Participants saw identical answer text to queries framed either as a traditional search result or as generative AI. The GenAI label lowered trust and willingness to share on average. They also investigated several interface features. Adding reference links raised trust and sharing (for both valid and invalid/hallucinated references), and color highlighting indicating model uncertainty lowered trust regardless of whether the highlights indicated low confidence or a mix of high and low confidence. Finally, a message that “users found this helpful” raised trust. These results show that people were influenced by many superficial and even fake cues of reliability (citations, social proof, AI-rated certainty), indicating weakly-justified and unjustified trust and distrust appear to influence use of AI.
However, many researchers in this domain do recognize that trust in AI is really a new name for the same issues that emerged with trust in automation. The main difference is a shift in focus toward more cognitive/informational aspects of trust that modern AI systems entail, and a focus on new settings and AI applications (chatbots, generative search, medical advice) rather than completely new theories. For example, Kaplan et al. (2023) conducted a meta-analysis empirical studies similar to Hancock et al. (2021), examining factors related to the human, the AI, and the context—a common set of divisions in automation-trust models. Also, Everett et al. (2026) framed AI trust in the same vocabulary as interpersonal and psychological principles (risk, vulnerability, evidence, and social inference), which may be an evolution of some of the factors influencing trust and trustworthiness of automation, but is not a novel theory of trust.
Consequently, because of its more commercial nature, its cognitive focus, and its information-rich setting, and the more human-like intelligence embodied by modern AI, research on trust in AI has explored new questions that did not interest researchers involved in trust in automation. However, the basic frameworks developed in automation contexts are largely applicable.
5.7.2 Trust in AI in cognitive systems engineering and human-centered AI
Expert users of automation and cognitive systems have long understood that trust is not all-or-none. In high-stakes environments, operators must rely on systems they do not fully trust, because experience teaches them what each system is good at and what it cannot be trusted to accomplish (Lee & See, 2004; Cohen et al., 1997; Hoffman, 2018; Ferrario et al., 2020). Consequently, many of the factors involved in human-automation trust (e.g., Parasuraman et al., 2000; Sheridan & Parasuraman, 2005) are likely to recur for modern AI (Glikson & Woolley, 2020).
However, new challenges are emerging. General-purpose AI systems are increasingly embedded in organizational and cognitive work, including legal research, programming, accounting, documentation, planning, design, analysis, and trading (Brynjolfsson et al., 2023; García-Peñalvo & Vázquez-Ingelmo, 2023; AbuMusab, 2023). In many cases, sensitive or proprietary data must be sent to external services, so users must trust not only the model but also the vendor: that data will not be used to train models that will be used by competitors, it will not be resold, or exposed through weak security (Jacovi et al., 2021; Shneiderman, 2020; Shneiderman, 2022; High-Level Expert Group on Artificial Intelligence, 2019; Tabassi, 2023; Smuha, 2019).
Another unique challenge of large language models is that they cannot be trusted to match reality without supervision. Hallucinations are well documented (Zhang et al., 2023), and visible errors can rapidly erode reliance even when the system is otherwise useful (Burton et al., 2020; Schemmer et al., 2023; Lee & See, 2004). The hallucination problem is also pernicious because it can impact any aspect of an LLM's output (although they appear more frequently in some conditions than others), unlike automation where the trust issues are often encapsulated and relevant to certain contexts. Moreover, their training is broad but may be out of date, domain-incomplete, or poorly aligned with local facts, producing confident but false outputs (Zhang et al., 2023; Endsley, 2023). Trust in LLMs is also literally contextual. Models operate within a finite token context window (Brown et al., 2020), and performance degrades when relevant information falls in the middle of long inputs rather than at the beginning or end (Liu et al., 2024). Instructions given early in a session may therefore “fall out” of context, and the system will no longer follow them—another reason trust must be scoped to task, session, and domain rather than attributed holistically to “the AI.”
Complex AI systems are increasingly responsible for ground and air vehicles, and in some settings for identifying objects and controlling weapons (Parasuraman et al., 2000; Tate et al., 2016). Although autopilots and guidance systems are not new, extending AI across the full sense–decide–act chain raises the stakes when trust is miscalibrated (Waytz et al., 2014; Linja et al., 2022; Hancock et al., 2011). Users may trust automation for routine driving yet distrust it in edge cases they have learned to anticipate (Linja et al., 2022). More generally, historical automation was often narrow and tightly governed. Integrating more general AI into complex work systems can disrupt processes in unanticipated ways (Klein et al., 2023; Endsley, 2023; Brynjolfsson et al., 2023). Adoption therefore requires justified, contextual trust in what each system is trustworthy for—and appropriate reliance rather than blanket trust or blanket refusal (Lee & See, 2004; Schemmer et al., 2023; Wickens et al., 2015; Shneiderman, 2022). Furthermore, the challenges of trust in AI are a consequence of performance issues, rather than the cause. Hallucination and over-stepping capabilities requires human workers to shift from creating knowledge to supervising and checking knowledge creation, but it is not because of distrust specifically–it is because of demonstrated limitations of the system.
5.7.3 “Trustworthy AI” as a legal standard
So far, our notion of trust and trustworthiness has focused on human factors and system design. With the rise of highly-capable AI systems, standards and governmental bodies have increasingly used the term “Trustworthy AI” to describe properties AI systems should exhibit in context of legal, societal, and ethical concerns. We will cover these issues more completely in Chapter 9, but cover it briefly here as well.
In its 2019 plan for federal engagement in AI standards, the National Institute of Standards and Technology identified seven core characteristics of trustworthy AI technologies: accuracy, reliability, objectivity, resiliency, security, explainability, and safety (National Institute of Standards and Technology, 2019). Stanton & Jensen (2021) extended this taxonomy to nine characteristics by adding accountability and privacy, arguing that these properties are necessary (though not sufficient) for perceived trustworthiness. similarly, the EU High-Level Expert Group on AI developed a framework for trustworthy AI (High-Level Expert Group on Artificial Intelligence, 2019; Smuha, 2019). It involves three general areas (lawful, ethical, and robust AI), with specific properties including: respect for human autonomy, prevention of harm, fairness, explicability, the role of human agency and oversight, technical robustness and safety, privacy and data governance, transparency, diversity and fairness, societal well-being, and accountability. All of these characteristics involve some aspect of the human-AI relationship, as the level of interaction that the human has with the AI system impacts how much they trust using the system and the development of other AI systems in the future. Many of these notions were also incorporated by Shneiderman (2022) in his Human-centered AI Trustworthiness Scale. This proposal allows a system to be 'scored' based on factors related to the development process, such as whether the software was validated, internal/external review was completed, and fairness was tested.
Jacovi et al. (2021) attempted to formalize the nature of trust between humans and AI. They argue that although interpersonal trust can be used as a basis for understanding trust in machines, machine trust relies on understanding risk, and entails the need for mitigating those risks, and the ability to anticipate the behavior of the AI is critical. As such, they develop the notion of contractual trust. This notion prescribes that trust in AI should be framed as trust in a commitment to a contract, which also require explicitly identifying the performance expectations of the contract. This notion is similar to Tate et al. (2016) notion of licensure, although the contracts in licensure frameworks may be backed by legal or organizational policy.
Nevertheless, the notion of trustworthy AI has its critics. For example, Ryan (2020) provides a extensive discussion defining the ethical and legal perspectives on trust in th AI context. Although he mostly ignores the extensive research on trust in automation we have covered here, one important insight calls into question whether trust can really apply to artificial intelligence at all. He argues that according to a normative account of trust (i.e., what one “ought to” or “should” do in such a trusting relationship), the trustee is held responsibility for its actions. However, legal discussion has rightly challenged whether an artifact can bear responsibility for an behavior at all–this falls on its designer, developer, or vendor. For example, according to a normative notion of trust, whoever is operating the vehicle ought to bear responsibility for causing an accident. For automated vehicles, the vehicle itself can not bear responsibility, and it should fall on the designer or the operator who is not actually driving (but may have elected to allow it to operate). In this sort of case, it was probably predictable that when Tesla Autopilot and FSD-enabled vehicles caused high-profile accidents, Tesla (who likely bears responsibility based on a normative model) attempted to shift the blame onto the vehicle drivers, who were not driving, as if they were saying “They weren't not driving the car correctly.” Because of this, Ryan (2020) argued that what was often mean by trustworthiness was really reliability, which indeed should be the focus, and the field should emphasize the trustworthiness of the organizations and people using and creating AI.
One possible resolution to this is the contractual and licensure notions discussed earlier (Tate et al., 2016; Jacovi et al., 2021). Since responsibility entails identifying blame for violations, and ultimately punishment, we already manage this with licensing, indemnification through insurance contracts in many situations. You cannot put a vehicle in jail for violating laws, but just as with human drivers you can (1) remove its license to operate; and (2) require financial payments to victims.
5.8 Conclusions
Trust is among the most frequently used constructs in human–automation and human–AI research, yet it is rarely a single thing. Across domains, definitions converge on a common set of ideas: a trustor accepts vulnerability in a relationship with a trustee under uncertainty, in pursuit of some goal that will produce a (normally non-zero-sum) benefit. The chapter began by surveying how trust has been framed—as expectation, attitude, belief, state, willingness to accept risk, ongoing process, or observable reliance—and contrasted trust with trustworthiness as a property of the trusted party (or the organization behind it). These distinctions matter because researchers and practitioners often measure one while meaning the other.
Much of the automation literature has drawn on interpersonal trust as a starting analogy, and for good reason: social cognition, anthropomorphism, and default dispositions shape how people encounter new systems. But the analogy is incomplete. Human–machine trust must also be understood through system cues, mental models, feedback from use, and the organizational and legal context in which a system is deployed. The models reviewed here—from cue-based and lens accounts to system loops, argument-based trust, and computational games—differ in emphasis, but address the same practical question: how evidence about a system's capabilities and limitations comes to support or undermine trust and reliance.
A recurring mistake in both research and design is to treat trust as a holistic, continuous quantity to be maximized. Instead, trust is often pluralistic, contextual, and componential: one may trust a system for some functions and not others, in some environments and not others, in some ways and not others, and for good or bad reasons. Concepts such as appropriate or calibrated trust, justified or unjustified trust, and warranted or unwarranted trust are therefore normative standards on reliance: whether a person's willingness to depend on a system matches what the system (and its operators) can defensibly be counted on to do. Transparency, experience, and expertise can shape that judgment, but they do not automatically “improve” trust. The goal of transparent system design should be to help someone trust a system only when they should.
Finally, contemporary AI systems, and especially general-purpose AI systems such as ChatGPT, implicate many of the concerns observed in research on trust in automation, and add new ones. Broader competence claims, opaque failure modes, vendor-mediated data use, and regulatory frameworks for “trustworthy AI” all widen the gap between perceived capability, organizational promise, and justified reliance. For human factors work, the practical implication is to understand the scope of trust explicitly: who is trusted, for what task, under what conditions, and with what accountability when reliance fails. Measurement and assessment are the next step: Chapter 6 examines how trust and related constructs have been operationalized, and what those instruments imply—and sometimes assume—about the nature of trust itself.
Acknowledgments
This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.
Conflicts of Interest
The authors declare no conflict of interest.
AI Usage Statement
Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.
How to Cite This Chapter
Walker, S., Matas, E., Ulinski, K., Woolman, B., & Mueller, S. T. (2026). Trust in automation: Definitions, models, and related concepts. In Shane T. Mueller (Ed.), A Handbook of Human-AI and Human-Automation Interaction. https://pages.mtu.edu/~shanem/human_ai/
- ↩ This quote was secondary to that research and did not appear there directly