Handbook

Chapter 6

Measuring Trust in Automation and AI: Key Points, Methods, and Scale Applications

Abstract. In automation and AI trust plays a pivotal role in ensuring smooth interaction and effective use of technology. This chapter covers many different ways trust has been assessed and measured. First, we discuss different contexts in which trust is measured. including field studies of real or simulated environments, simulators, microworlds, wizard-of-oz simulations, and experimental research. Second, we describe empirical methods for characterizing or quantifying trust in these contexts, including interviews, self-report methods, behavioral methods, and physiological indices methods. We then identify many of the existing self-report questionnaires and scales that have been used, looking at a one directly to understand how it measures trust. Finally, we examine new scales and measures that have been developed to assess trust and trustworthiness of AI applications, both as a specific attitude toward a system, and as a general disposition to different technologies. This provides a practical foundation for researchers who are interested in evaluating trust in automation or AI, and want to identify their best approach.

6.1 Introduction

Chapter 5 discussed definitions, theories, models of trust in AI and automation, with the perspectives that trust is multi-dimensional, contextual, componential, and can be both categorical and dynamic. Assessing or measuring “trust” can thus be complex and difficult to carry out. Often, the go-to solution amounts to asking a user a few questions, such as “On a 1 to 10 scale, how much do you trust the system.” Although this can be useful, we hope the take-away from this chapter is that–due to its nature–there may be many distinct ways to measure trust, and some of these may be more effective at understanding the human-machine relationship than a simple rating scale.

In fact, there are many methods and strategies for evaluating and assessing trust–some are alternative measures, others are ways of embedding these measures within measurement contexts. These range from quantitative surveys assessing subjective perceptions to experimental setups analyzing behavioral and neuropsychological response. Of course, even when one simply wishes to assess trust with a questionnaire, there are many options to choose from, many of which are rooted in research on trust in automation. Recently, scales measuring general and specific trust in AI have emerged, and we will discuss many of these, and how they are either rooted in or distinct from previous measures of trust in automation.

6.2 Methods for measuring trust in automation

Chapter 10 of Adams et al. (2003)'s comprehensive technical report on trust provided an exhaustive picture of how trust in automation had been measured up until that point. They first identified research approaches—the context in which measurement happens (field, simulator, microworld, interview, experiment)—then covered measures, examining trust as a psychological state (ratings and scales) and trust as behaviour (defensive monitoring and use of/reliance on automation). Kohn et al. (2021) more recently reviewed trust-in-automation research and similarly identified self-report, behavioral, and physiological measures, each mapped onto a step in Mayer et al. (1995)'s process (perceived trustworthiness, propensity, trust attitude, or risk-taking), providing new cases of trust assessment that had become popular over the previous 20 years.

Table 6.1 summarizes these measurement contexts, which we augment with two other contexts that are perhaps unique special cases. We will discuss several of these contexts in more detail in the next section. After that, we cover many of the measures used to characterize or quantify trust in Table 6.2, which is a combination of both reviews (Adams et al., 2003; Kohn et al., 2021), and incorporates and expands on these to some additional approaches.

6.3 Contexts of trust measurement

Table 6.1 identifies a number of contexts in which trust can be assessed. These range across a spectrum of fidelity to the real work context, with trade-offs as you move from live study of a real system in the field to scenario-based surveys. Study context is often not a choice one makes, or is very constrained by access and resources. It may seem that ideally, you should understand trust in the context of the actual work-system. However, this may not be possible if the system is in a planning phase, is government classified or proprietary, dangerous to operate, requires specialized expertise, or is only used in unique circumstances or infrequently. In these conditions, a simulator, a scenario, or studying another existing surrogate system may be a good alternative. Even with access to an existing system, alternatives may be preferred—one may not be able to do live questionnaires of an aircraft pilot during use of a combat system, but could do so either in a lower-fidelity simulation, or in a video scenario where they watch a maneuver and engage in a think-aloud protocol.

ContextDescriptionStrength / weakness
Field study of real or simulated systems *Operators in aircraft, plants, or other natural settings or surrogate systems.High realism; little control of extraneous variables.
Simulators *High-fidelity physical and dynamic replicas (aviation, navigation, process control).Realism with more control than the field; expensive, and still not the job.
Microworlds *Small closed-loop worlds that capture some complexity of process control without simulating one plant.A trade-off between laboratory control and field realism.
Wizard-of-Oz (WoZ) studiesUseful in microworlds, simulation, or low-fidelity scenarios; involves a human playing the role of the AI or automation (Green & Wei-Haas, 1985; Dahlbäck et al., 1993).Can simulate complex behavior not yet achievable by current AI.
Low-fidelity scenariosNormally, written or video scenarios describing/showing situations with trust elicited via questionnaire or interviewEasy to implement but low fidelity means potentially low validity.
Laboratory experiments*Simple games or gambling tasks; experimental manipulation of risk, reliability, or other context.High control but easy to lose the operational meaning of “trust.”
Table 6.1. Context of measurement for trust in automation and AI. Those marked with * covered in Chapter 10 of Adams et al. (2003).

Most of these contexts are self-explanatory. They each represent a different point on the fidelity spectrum, which offers trade-offs for studying trust. This trade-space not only represents verisimilude, it also can represent different levels of risk and responsibility that is an important part of most notions of trust. That is, you can use a simulator or a scenario to understand when someone would allow an automation to operate, but if the consequences of failure in the real-world involve real danger, financial loss/gain, reputational damage, health consequences, or equipment destruction, the low-stakes environment may not accurately measure what trust in the real situation would be.

Another aspect of the context is the level of experience of the user or research subject being studied. From an experimental research perspective, the user should be sampled from the same population that one hopes to generalize to, but often the trust relationships that are most interesting are for experienced users who have expertise and know-how about a domain, and are determining whether to trust a new system. Training-to-criterion and simplification of systems can be used if only inexperienced novices are available for testing, but this can also limit validity. A compromise might be to study expert users of surrogate system, gaining realism in the research subjects but losing direct application to the target system. For example, one may not be able to test military UAV operators because of access and security reasons, but a civilian hobbyist may suffice for some evaluations.

Vignettes are a very lightweight and flexible approach to understanding trust. Using a vignette, you describe a situation and have the participant consider what choice or action, or assessment they would make. This can be framed in the third person ("What should the pilot do?"), or first-person ("what will you do?"). This can be used to test out possible interaction modes. For example, Alam et al. (2025) used vignettes to examine how participants would react if they were patients getting diagnosed by an AI system, and tested the effects of embedding elements of both emotional and cognitive empathy. Similarly, Zondag et al. (2024) used vignettes to study patient-physician trust. These are akin to “Wizard of Oz” (WoZ) approaches, but a true WoZ study would have a user interact with an AI controlled by a human confederate, and not a use a clearly hypothetical scenario.

Altogether, these different contexts and settings for studying trust permit assessing trust with different fidelity to real-world systems. However, when the setting and context of the assessment is determined, the actual assessment, measure or metric of trust still needs to be determined. We will cover many ways of doing this next.

6.4 Measures and metrics for characterizing trust in automation

There are also a number of measures and assessments of trust that have been used to characterize or quantify trust and trustworthiness. These are summarized in Table 6.2. The suitability of these are also often constrained by the application, experimental hypotheses, domain, and work context. Adams et al. (2003) described many of these, and Kohn et al. (2021) described a distinct but overlapping set. We have added several specific approaches that are either distinct or important special cases of other measurement paradigms. The table shows that there are MANY ways to characterize or quantify trust. Many developers or researchers simply use an existing subjective self-report scale (or develop/modify one for their own purposes). However, there are many additional measures that can be used, many of which might be more informative. Depending on the critical phase, it may be more important to characterize the issues related to trust (using interviews or open-ended feedback) than it is to quantify a subjective trust level (using behaviors, neuropsychological measures, or surveys).

Assessment / measureDescriptionStrength / weakness
Interviews (and CTA)*Semi-structured interviews, often early in a program (e.g., naval data-fusion work reviewed in Aldern, 1995). A CTA of the work itself can expose trust issues indirectly (Matthews et al., 1999).Explains why and what is trusted; non-numeric.
Think-aloud / verbal protocolConcurrent or retrospective spoken thoughts while the person works with the system (Crandall et al., 2006).Although a useful direct indicator of user thought process while in the situation, aspects related to trust may be rare.
Questionnaires and scales*†Ranges from informal ratings to validated scales with subscales: (see Table 6.3).Repeatable and easy to field. Risk that an easy measure replaces in-depth inquiry into system.
Neuropsychological measures†Secondary indices: electrodermal activity, heart rate / HRV, eye gaze and pupil, and neural measures (EEG, fNIRS, fMRI) (Kohn et al., 2021).Objective and precisely timed. Does not measure trust directly, and may incorporate load, arousal, risk, and other unrelated constructs.
Defensive monitoring / verification*†Examining user monitoring frequency while in supervisory control.Continuous and behavioral, but coding audio/video logs is labor-intensive.
Use of / reliance on automation*†Frequency, duration, and automatic vs. manual control (Adams et al., 2003). Also discussed by Kohn et al. (2021) in terms of reliance, compliance, intervention, delegation, and decision time.Objective, but is at best a proxy for trust, because many users rely on tools they do not trust.
Economic or strategic games / markets†Money, tokens, points, or contract value placed at risk by cooperating or investing. Examples include the Give-Some game after a conversation with a robot (DeSteno et al., 2012; Lee et al., 2013), investment and social-dilemma payoffs in computer-mediated teams (Zheng et al., 2002), loans to a robot partner (Jessup et al., 2019), and repeated Prisoner's Dilemma, Chicken, and multi-arm trust games (Juvina et al., 2015; Juvina et al., 2019; Collins & Juvina, 2021).Incentivized behavior adds realism; often limited to understanding interpersonal trust with a machine as the partner.
Table 6.2. Trust assessment and measures. Items marked * appear in Chapter 10 of Adams et al. (2003); items marked † are described by Kohn et al. (2021). Interviews/CTA and think-aloud protocols are included here even though neither review treats them as a developed measurement method.

6.4.1 Characterizing Trust Using Semi-Structured Interview Methods

We know of no dedicated “trust interview” or commonly-used Cognitive Task Analysis (CTA) method Crandall et al. (2006) for exploring trust in work systems, automation, or AI. CTA approaches are not usually used for quantifying trust (providing a number that you can compare or track), but are useful for characterizing trust; providing detail about what people trust, the context of trust, how they trust, and why they trust. Because these methods tend to focus on specific incidents, issues related to trust frequently emerge (deliberately or by happenstance), especially if the incidents are selected carefully. This can be done by choosing incidents in which human-machine trust or reliance was either needed or fell apart. For example, one might start by asking, “Tell me about a time in which you needed to trust SYSTEM X, and it let you down”. The responses from this approach can provide much better insights than simply a numeric value you might get from a survey. Although Adams et al. (2003) mentions interviews as a measurement context rather than measure, they made reference to Matthews et al. (1999), who used a CTA of Halifax-class operations-room officers as the source of their measurement framework, and discussed semi-structured interviews around a naval data-fusion prototype (Aldern, 1995).

Interview methods are commonly used in medical and health care research as well, for a wide range of purposes. For example, Hanna et al. (2023) used semistructured interviews and focus groups across three groups (medical cancer advisors, volunteers from cancer support groups, and people from the general public) to evaluate the trust people have in a medical chatbot, finding that willingness to use a chatbot would depend on personal need, such that they would prefer it over searching online if a physician was not available. Kostick-Quenet et al. (2023) use in-depth interviews of medical professionals and patients who participated in high-risk procedures to broadly evaluate trust in AI, and determined that high accuracy was a critical desire among physicians. King et al. (2022) used interviews of pathologists to evaluate the trust of AI in pathology specifically, and found there was a desire for explanation from AI, while recognizing that current pathology practice can be a black box without explanation. Fernandes et al. (2023) used interviews of weight loss professionals as part of their development process of a weight loss tool. Their participants noted that having limited variables and an overly predictable output that does not take into account all of the factors the experts believed would impede their trust in the system.

We consider interviews as the elicitation method rather than a context, as they can refer retrospectively to events that happened either in the field or in any simplified simulation of that system. Trust-related information frequently emerges from interviews, and several published studies serve as useful examples and potential models for new interview methods. For example, Capiola et al. (2020) used CTA with intelligence operators in multi-domain command and control to characterize swift trust in ad hoc teams. They used fairly standard CTA methods, targeting the process by which ad hoc teams quickly developed trust in one another, and attempting to understand the precursors that led to its development. Although this research dealt with interpersonal trust, a similar approach could be used for human-machine trust by changing the basic incident selection and targets of the prompts.

Task analysis assessments of trust can also be used to improve design by improving system trustworthiness. Given a typical trust process and a trust-cue taxonomy, de Visser et al. (2014) give a procedure for identifying trust cues for a specific case:

  1. Select a scenario that involves trust in a cognitive agent.
  2. Conduct a task analysis to identify critical trust-related tasks.
  3. Identify key pieces of information for operator decision making.
  4. Verify pieces of information against the trust-cue taxonomy.
  5. Construct a visual display representing this information.

This procedure provides a potential avenue for designing or redesigning an information display that is able to capture the important factors that lead to trust in users.

6.4.1.1 Using Critical Decision Method for Understanding Trust

The Critical Decision Method (CDM) Klein et al. (1989) is a popular CTA approach commonly used to understand the challenges for important decisions, typically focusing on expert decisions. It works by selecting one or more incidents that serve as `critical' decisions, establishing a timeline of events prior to and following the incident, and then engaging in several passes and retellings of the incident that focus on different factors, each time deepening the understanding of the incident. Klein et al. (1989) originally described a number of probes for these passes (e.g., cues, knowledge, aiding, time pressure, situation assessment, etc.), and trust issues often emerge during these additional passes. For example, focusing on knowledge and cues might uncover unreliability that would signal either trust or distrust, and the proper interpretation might depend on background knowledge. Klein et al. (2010) used this method to study developers of a complex city-wide simulation model that incorporated social, infrastructure, and economic factors. The simulation showed periodic dips in economic activity that would recover after a few days. When novices saw those dips, they inferred it as a cue that the model was broken and maybe should to be trusted. Experts who understood the modeling system better knew these dips were linked to simulated power outages that were caused by sabotage or disrepair. The same cue that appeared to novices to indicate unreliability demonstrated to experts that the system was working properly.

Although trust can emerge in some of the common CDM deepening passes, Brown et al. (2025) suggested using “How did you know to trust the information?” as a probe that would focus on trust specifically. This could be phrased in many ways depending on the person and the system being interviewed. To tailor the CDM to eliciting trust, we recommend the following as a starting point:

  1. The incident should be selected to focus on a case where trust was involved, perhaps within routine operation of an automation system, intelligent machine, or AI.
  2. Following the initial timeline, at least one pass can focus on how trust or reliance impacted the decision or outcome.
  3. Rather than simply probing about `trust', consider incorporating relevant factors from the trust models in Chapter 5 that are known to impact trust of the interviewee (past experience, expertise, personality, etc.), trustworthiness of the trustee (e.g., reliability, transparency, interactivity, capability, etc.), and the context (e.g., the team, organization, policies, etc.).
  4. Rather than trying to elicit a degree of trust, recognize the contextual, componential, categorical, and multidimensional nature of trust. Ask “What about it did you trust?”, “What did you trust it for?”, “What part or function were you trusting?”, and “When would you have trusted or not trusted it?”.
  5. To consider how the trust or distrust was justified, ask “Why did you trust/distrust the system?”.

6.4.1.2 Lighter-weight CTA methods

The CDM can be quite time-intensive to use, and it produce long transcripts that require substantial resources to organize, code, and extract themes. Other methods are lighter weight, including Applied Cognitive Task Analysis (ACTA) (Militello & Hutton, 1998), which incorporates an approach called the Knowledge Audit. One of the standard probes of a knowledge audit is Equipment difficulties which is intended to target “Not trusting misleading instruments” by asking “Has the equipment pointed one way while your judgment said another?”. This is a fairly narrow focus on trust, and it would be feasible to incorporate several additional trust-related probes here, such as “Did you consider not using the system and doing it yourself?”, “What part of the system could you trust completely, and what would you be unsure about?”; “In what ways would you trust the system at this point?”. If trust is critical, a comprehensive “Trust audit” might be developed along those lines

Overall, these interview methods can be powerful ways to understand how and why, and in what context people trust the systems they use. They are not without drawbacks. They require access to specialized users, trained interviewers, and can require manual data coding. Also, they typically involve small numbers of observations, and so can be subject to sampling issues and interviewee biases. There are several other verbal-based methods for analyzing trust that bear discussion. These have some of the same strengths and weaknesses, but also offer distinct approaches.

6.4.2 Other Verbal Measures

Interview methods and many CTA approaches typically involve verbal information that is analyzed and processed for clues regarding trust.

Think-aloud protocols

A think-aloud protocol is a self-report method oftentimes considered a task-analysis interview method, but it is somewhat distinct in that it asks a user to verbalize their thoughts, goals, and actions as they perform a task. It can be revealing for many aspects of performance, although there are also limits into what people are really able to verbalize about their performance (Nisbett & Wilson, 1977; Ericsson & Simon, 1980). Although this is a potentially useful method, it must (like most behavioral measures) be accompanied by appropriate tasks that are likely to influence trust. Furthermore, to use such an approach, it is probably useful to provide instructions and prompts for users to focus on their feelings of trust or trustworthiness in a system, to help surface those attitudes.

This method sheds light on the person's thought process and could be used to identify leverage points of trust in real-time as the user encounters them. A retrospective protocol analysis can also be used, where participants complete a task and explain their thought process afterwards, possibly prompted by a video recording of the event. However, just like interviews, think-aloud protocols can generate a large amount of qualitative data that need to be analyzed and summarized—typically with similar qualitative research methods used in interviews—to identify issues related to trust.

This approach was used to evaluate theories of trust within a risk management setting by Earle (2004), as well as by Gautam et al. (2026) to investigate researcher trust in LLMs to perform different early-stage research tasks.

Discourse analysis

There is a body of work in the domain of pragmatic linguistics that has used discourse analysis to study trust in written or spoken language. Rather than examining transcripts or responses generated by users of a tool, discourse analysis often works on written public documents, and it describes a number of related qualitative research approaches where linguistic form and usage in those documents is examined for insights into a construct (e.g., trust). For example, Fuoli & Paradis (2014) examined CEO letters following the BP oil spill accident that were intended to repair trust in BP after the spill. Saragih et al. (2026) used this approach to study how AI was represented in online media, and identified themes related to trust. This approach can be used for both archival transcripts (e.g, cockpit communications logs, after-action review reports) as well as newly-designed studies where communication is logged incidentally or as the purpose of the study (for example, in a focus group, unstructured user interview, or user feedback).

Natural Language Processing and Sentiment Analysis

In contrast to discourse analysis which is typically performed by trained experts, there are a number of natural language processing approaches that use classification and clustering to extract attitudes and sentiment. Sentiment analysis is a technique that tries to characterize positive or negative attitudes within text, by counting key words, using latent semantic analysis, or other machine learning or AI tools. Although this does not assess trust per se, it can provide an understanding of whether social media posts, user reviews, traditional media, or small-scale communication involves positive or negative opinions, and further language classification might be used to narrow down trust and distrust specifically. For example, Guo et al. (2024) used sentiment analysis to infer trust relationships from online review data, and Niu et al. (2020) used a similar approach to identify trust in lenders in online comments.

6.4.3 Behavioral Measures

Human behavior can be impacted by their trust in a system in a number of ways—both direct and indirect. Here, the term appears as sort of a catch-all that is contrasted with non-behavioral measures (neurophysiological measures, text analysis) as well as verbal behaviors (interviews, questionnaires are technically behaviors). The most obvious behavior indicating trust is whether a person uses a system—which might be called adoption or reliance instead of trust, but there are other less direct behavioral indicators of trust.

Observational methods

A number of observational methods can be used, in which behavior is observed in real or simulated environments, and potentially without the awareness of the observed party. Although this class of methods is very broad, it includes some very useful techniques such as observing a first-time user of a tool (Mueller et al., 2009), or an observational usability study. Observational methods often occur in healthcare settings to understand many aspects of trust, including patient trust in processes, advice, physicians (Zondag et al., 2024; George et al., 2019; Paul et al., 2022; Riedl et al., 2023; Jongerius et al., 2022; Aiken et al., 2021; Orrange et al., 2021), AI patient-oriented tools (Stein & Brooks, 2017), and AI decision support (Sakamoto et al., 2024).

Defensive Monitoring

Muir & Moray (1996) argued that trust is negatively related to the need for monitoring automation. This “defensive” monitoring refers to a trust-related behavior that is undertaken to reduce the likelihood of negative outcomes if automation deviates from expectations and predictions. When automation is highly reliable, the need for monitoring reduces, but unreliable systems may require constant monitoring. For example, one might use an eyetracker to count eye movements to the roadway for an operator of a vehicle in autonomous mode, while the operator is doing some other task such as reading. Even if the operator never takes over driving from the autonomy, when a person is looking at traffic or the road ahead, we might infer that they trust the system less than periods when they are not looking at traffic. In other settings, this can be examined more directly through recording specific sensors displays that need to be activated by the user, or other means.

If using a monitoring behavioral paradigm, the “optimal sampling rate” normally needs to be established. Most forms of automation need a certain level of monitoring, and deviations from that might indicate overtrust or undertrust of the automation. Therefore, self-report scales and behavioral measures of defensive monitoring, such as the number of checks on automation, can be used to assess the required frequency of monitoring. Analyzing video logs or digital records of instrumented systems can also be effective ways to measure defensive monitoring (Adams et al., 2003).

Use or reliance

The use of automation can be assessed based on the frequency of automation use, the length of each period of automation use, or the total proportion of time that the system is under automatic control versus manual control (Adams et al., 2003). In many situations, the use of automation may signify trust in the automated system. The method for measuring use of automation is through system logs that collect data on automation vs. manual control. Additionally, an alternative method is to analyze audio/visual logs to decide the use of automation. By observing and analyzing them, researchers will know whether individuals trust automation.

However, use and reliance should be interpreted carefully as a measure of trust. First, there can be many reasons people use or rely on a tool they do not trust. They may have no other choice, or the consequences of failure may be low, or they may simply trust the system more than themselves (the “authority hypothesis” proposed by Mosier & Skitka (1996)). Moreover, reliance or adoption is often understood as distinct from trust, which is often considered more prospective and about future behavior, not immediate performance.

Non-verbal communication, social cues, other gestures

Nonverbal communication and nonverbal social cues can include facial expressions, a set of expressions, gestures, and mimicry, posture, and other behaviors. For example, an eye roll in conjunction with a large grin may more accurately convey humor. In the context of measuring trust, these can be both deliberate (a gesture communicating to a team-mate) and subconscious (posture, agitation, expressions of anger). Although analysis of these cues have been used in many contexts, they have occasionally been used to specifically study trust, such as interpersonal ratings of trust ((Wood, 2006)), who found that non-verbal cues impacted trustworthiness ratings of salespersons. Similarly, Burgoon et al. (2021) provided a very comprehensive examination of nonverbal signals (including glare, voice, facial expression, body movement, speaking tempo, turn-taking, interruption), and found that longer and slower turn-taking was positively associated with perceptions of trust, but many non-verbal behaviors were not.

These cues have also been used to assess trust in social robotics. Lee et al. (2013) found certain nonverbal gestures like face touching, crossed arms, leaning backward, and hand touching were associated with lower levels of trust in a robot partner during a social game.

Interaction and social graph analysis

Another way to infer trust between entities is through interactions, or explicit surveys about trust relationships within a team (possibly including AI/technology entities). Interactions could be measured from communication logs, video logs or location sensors, and equipment usage records. For example, Krackhardt (1987) inferred `cognitive social structures' by asking individuals within a team who they went to for advice; Zhou et al. (2023) used an interactive hybrid task analysis approach to create flowchart diagrams describing different team members in a human-autonomy team in a remotely piloted aircraft scenario. Mohammadi & Hashemi Golpayegani (2021) proposed a social-network–based model that infers trust. By using these task and social representations characterize interaction, a number of well-established network metrics can be deployed to understand team effectiveness, interdependence, and reliance.

6.4.4 Physiological and Neuropsychological Methods

Kohn et al. (2021) identified four groups of neuropsychological trust measures—electrodermal activity, heart-rate change and variability, eye-gaze tracking, and neural measures. They must be interpreted in the context of the task being performed, because the same signal also tracks load, arousal, and surprise, but they can be useful real-time or post-event measures that could either indicate timepoints where trust breaks down, or physiological precursors to trust or distrust. Typically, they should be correlated with other direct measures of trust to establish their validity in the context of study. Often, these measures are used together, along with both behavioral paradigms and subjective self-report scales of trust, alongside stress, workload, and mental effort. Gupta et al. (2020) described more than 10 studies that have used combinations of heart rate, GSR, eyegaze, EEG, and other physiological measures as proxies for trust. They note that cognitive load and trust are often related experimentally, and in many of the tasks that they reviewed (VR and other games), this may make sense–immediate cognitive load may prevent assessing the trustworthiness of a partner and so lead the two measures to co-vary.

Eye Tracking.

Eye-tracking has been used to establish trust in automation as shown by Lu & Sarter (2019) in ways similar to defensive monitoring discussed earlier. They found that fixations increased in both duration and length in a low-reliability condition when compared to a high-reliability condition. Furthermore, participants who were primed to expect the system was high or low reliability altered their eyegaze behavior, which demonstrates that eyegaze can track not only trust level, but change in trust or trust calibration as experience is gained.

Heart monitoring.

Heart rate changes are known to be sensitive to many aspects of skilled performance, as can the secondary measure of heart rate variability (HRV). These measures have also been linked to trust states, but these are often correlational results, and work must be done to establish both sensitivity (the measure is related) and specificity (it is NOT related to other explanatory factors). For example, Leichtenstern et al. (2011) found that when a website was not designed to elicit trust in the user, heart rate increased from baseline and post-test subjective surveys found that users were stressed when interacting with the low trust website. Gupta et al. (2020) used both heart rate and HRV as measures of trust and cognitive load, and established that a higher cognitive load while engaged with the VR indicated lower user trust in a voice-assist agent. This sort of research normally relies on manipulations of trustworthiness and subjective measures to establish sensitivity of heart rate. They likely cannot be used alone as a measure of trust, but need to be incorporated alongside other validating measures.

Galvanic skin response (GSR).

GSR is another physiological measure that has been used to measure levels of trust in AI. But like heart rate monitoring, it also measures arousal, anxiety, stress and cognitive load, so specific features of GSR may be needed to distinguish among these psychological states. For example, Akash et al. (2018) used GSR together with EEG to find associations with human trust in machines. They found that mean GSR as well as GSR peak value were significantly affected by the amount of cognitive load required to interact with the machine and trust in the machine. However, the patterns of load and trust differed, even though higher load was associated with greater levels of trust. High cognitive load was associated with higher peaks in GSR while low amounts of human trust are associated with a higher mean GSR. Together with EEG, GSR could estimate a user's reported level of trust in the machine. Similarly, Hald et al. (2020) determined that GSR could be used to measure trust in a robot during human/robot team collaboration. This research found that GSR spiked when robots made unexpected movements and sounds and when measured alongside user's changes in physical activity to the robot, indicated loss of trust in the robot.

EEG and fNIRS.

EEG (electroencephalograph) and fNIRS (functional near-infrared spectroscopy) are relatively non-intrusive methods for measuring brain activity. EEG uses sensors attached to the scalp and measures electrical activity, whereas fNIRS attaches a light sensor to the skull to measure blood flow. Research conducted by Eloy et al. (2022) found that using fNIRS to examine the human prefrontal cortex can elicit information about trust in human-agent teams (HAT) to such a degree that decisions in relation to trust can be predicted up to 15 seconds beforehand. This research found that the dorsal medial prefrontal cortex, which is responsible for social impressions, activates when the user is determining default trust in the agent. During a condition designed to elicit suspicion in the human team member there was fronto-polar area (strategy), temporoparietal junction (moral decision making), and dorsolateral prefrontal cortex (planning and abstract reasoning) activation. Palmer et al. (2020) elicit mistrust from a human user in an automated driving scenario in which the autonomous system made poor decisions. An fNIRS sensor detected ventrolateral prefrontal cortex activation, which is associated with decision uncertainty.

Finally, EEG has been utilized to determine human controller trust in a situation where they are monitoring an automated system. In their condition designed to elicit mistrust, Oh et al. (2020) found increased gamma wave power, associated with increased focus and concentration, when interacting with the automated system. They also found that when the system is designed to increase trust, EEG detected increased alpha, associated with relaxation and contentment, and beta waves, associated with active thinking and wakefulness. These techniques can take advantage of spatial specialization of human brain regions to provide a more complex signal for detecting trust.

Overall, these physiological measures have advantages–they can record brain state without interfering in a user's activity, they can have high temporal accuracy to pinpoint exactly when trust changes occur, and they can provide additional evidence of different brain subsystems responsible for trust (e.g., emotional vs cognitive). However, they have limitations, including the cost and hassle of using them in comparison to other methods, as well as the typical lack of specificity in measuring trust. The measures developed may also be closely linked to experimental scenarios and manipulations, and may not easily generalize to other paradigms, or field evaluation of an actual work system.

6.4.5 Assessing Trust with Surveys or Questionnaires

Questionnaires are used extensively to measure trust in decision aids, automation, AI assistants, and the like. They are also often used in conjunction with other methods. As we saw earlier, they can be used to help validate behavioral and physiological measures. They are also sometimes coupled with interviews or focus groups to provide more detailed assessments. Although most trust scales involve repeated Likert-scale (1 to 7) responses, they can also include open-ended items that provide information more like one would obtain from interviews.

Serious attempts to measure trust as a subjective rating emerged by the 1990s, as research on trust in automation became popular and started to investigate the emerging problems and accidents occurring because of automation. These were sometimes just direct ratings of specific constructs, such as Muir & Moray (1996), who used self-report assessments of trust and related constructs on a 1-100 scale to test a regression-style model of trust during a pasteurization simulator.

Kohn et al. (2021) noted that custom single-item “how much do you trust…” scales are the most common method in practice. Moreover, many studies develop ad hoc surveys targeting their own specific research questions. For example, Xiang et al. (2020) studied the attitude of the public towards AI in the medical domain using a questionnaire, designed to elicit different attitudes toward the use of AI in medical settings, including perceptions (efficiency, safety, validity, trust, and expectations), receptivity (willingness to use AI services), and demands (which health care functions need medical AI).

However, systematic psychometrically-sound scales of trust have been developed that are intended to be used generally (see Table 6.3 and 6.5). Distinct scales are often intended for different purposes (e.g., prospective, dispositional, or system evaluation), or are targeting different kinds of systems (technology in general, robots, AI). In Section 6.5, we will discuss these in much greater detail.

6.4.6 Summary of Trust Measurement Approaches

Many researchers and developers appear to view trust as simply a quantifiable subjective report that is obtained via a questionnaire. As we saw in Chapter 5, there are many perspectives on what trust is, and so it is not surprising that there are also many distinct ways of measuring trust. Questionnaire-based trust can be helpful, but other ways of evaluating the nature of trust may be more helpful for system design and evaluation. Furthermore, many studies have benefited from using complementary approaches to establish coherence and validity of measures. In all cases, understanding basic psychometrics is critical–is the measure valid (does it actually measure trust) and reliable (is it likely to measure the same thing twice), and also important to understand both its sensitivity (how well can it distinguish between trust and non-trustworthy states) and specificity (is it measuring global performance or other related constructs and not trust specifically). Scales that measure trust have often established these psychometric properties of their measures, and we will cover many of the existing scales in the next sections.

6.5 Questionnaires and Scales Used to Measure Trust in Automation

The methods above treat trust as something that can be inferred from interviews, behavior, or physiology. In practice, many empirical studies operationalize trust with a questionnaire: participants rate agreement with statements about a system, a class of systems, or people in general. Chapter 5 used one such item—“I trust the system,” item 9 of Körber (2019)'s Trust in Automation (TiA) questionnaire—to illustrate how one perspective on trust is that it IS the outcome of a questionnaire. It also shows a common problem faced by trust scale developers—to answer the question, we must assume a user already knows what trust is. Many questionnaires include items such as this, but also other items with related constructs that are either related to trust definitions (e.g., faith, reliability, accuracy), or predictors or consequences of trust (reliance, understanding). But these related concepts are not identical to trust—a user can rely on a system they do not trust, or they can have faith in a system for unjustified reasons. Consequently, it may be important to be careful when interpreting a global trust score, as it may be composed of a user's own interpretation, and related but not identical concepts.

6.5.1 Scales that Measure General Trust and Trust in Automation or Technology

A number of trust scales and questionnaires appear in Table 6.3, and we will discuss some of the background research for several of the scales next.

InstrumentSourceFocus
General interpersonal trust
Mayer and Davis trust and trustworthinessMayer & Davis (1999)Interpersonal items for ability, benevolence, integrity, and trust. Sometimes adapted to automation.
Merritt trust scaleMerritt (2011)Six trust items plus ability and benevolence; often adapted for automation.
ESS / Breyer Social Trust ScaleBreyer (2015)Three items on people in general: whether most people can be trusted.
ESS Institutional Trust Scale (ESS-ITS)European Social Survey (2023)Seven 0–10 ratings of trust in legal and political systems—not trust in machines.
Trust in automation: component and general scales
Trust vs. self-confidence itemsLee & Moray (1992); Lee & Moray (1994)Parallel ratings of trust in automation and confidence in manual control.
Complacency-Potential Rating Scale (CPRS)Singh et al. (1993)20 items on attitudes toward automation that includes a trust.
Component ratings (competence, predictability, etc.)Muir & Moray (1996)Subjective ratings of trust's predictors / components.
Trust in Automated Systems Survey (TIAS)Jian et al. (2000)12 empirically clustered items; A later 3-item short form (S-TIAS) was validated for AI (McGrath et al., 2025).
Human–Computer Trust (HCT)Madsen & Gregor (2000)25 items; reliability, competence, understandability, faith, personal attachment.
Dynamic reporting of trustDesai (2012)Repeated increase / same / decrease reports relative to the last rating; low interruption, for tracking change over a trial.
Propensity to Trust Automation (PTA)Jessup et al. (2019)Six pre-task items with “automated agent” as referent. Used for initial trust propensity, not in-task trust.
Trust in Automation (TiA)Körber (2019)19 items, six subscales; CC BY-SA 4.0 (Table 6.4).
TOASTWojton et al. (2020)9 items; Understanding and Performance scored separately.
Propensity to Trust in Automated Technology (PTT-A)Scholz et al. (2025)Validation of scale for propensity to trust automated technology, compared with propensity to trust humans (PTT-H).
Special-purpose measures (robots, HRI, XAI)
SHAPE Automation Trust Index (SATI)Goillau et al. (2001)General questions (trust, usefulness, etc.) tailored to air-traffic management simulator
Trust Perception Scale–HRISchaefer (2016)Robot-directed trustworthiness ratings.
MDMTMalle & Ullman (2021)Multidimensional human–robot trust (performance and moral facets).
Table 6.3. Selected questionnaire measures of trust and related constructs, excluding scales and questionnaires used for measuring trust in AI which are covered in Table 6.5.

These instruments do not all measure the same construct. The first section of the table covers commonly used interpersonal trust measures (e.g., Mayer & Davis, 1999) or social trust measures (Breyer, 2015). These generally are not suitable as a drop-in measure for AI or automation, but they are sometimes adapted for this purpose. Furthermore, two scales that appear in the European Social Survey (ESS) characterize general trust toward others and institutions, which may be useful in understanding how an individual views AI within society or their organization. The next section includes ten scales intended to be used for measuring trust in automation, computers, or technology. Instruments such as Human–Computer Trust (HCT) sit alongside system-state and propensity measures. Some provide ratings of the trust state in a specific system during or after use. Some rate a propensity or disposition to trust automation (much like the ESS survey items). Many of these have short version (i.e., the 12-item TIAS vs the 3-item S-TIAS) or subscales (TiA includes six subscales) that can allow for more efficient measurement or narrowing in on specific aspects of trust. The last section includes trust scales developed for specific kinds of automation domains or applications (air traffic control, robotics).

There are several caveats to be considered when selecting a scale. Adams et al. (2003) warned that many early measures were theoretically rather than empirically derived, which suggests they have unknown or may have poor psychometric properties. Furthermore, using a single numerical rating derived from one or more items falls prey to all of the mistaken assumptions about trust discussed in Chapter 5 (a numeric, uni-dimensional, holistic, and universal concept). This is partly borne out empirically—Jessup et al. (2019) found poor reliability and no prediction of the Jian ratings (a state-based measure) for predicting robot-investment behavior, while propensity to trust was predictive. Also, asking trust items during a session can influence how people rely on the recommendations of a system (Schrills et al., 2026), as it presumably encourages them to focus on the trustworthiness of the system in ways they wouldn't otherwise consider. Finally, Perrig et al. (2023) showed that trust in automation scales can fail to transfer to AI contexts, and so new measures and new validations may be needed. We will examine a number of trust-in-AI scales in a later section of this chapter.

6.5.2 Körber's Trust in Automation (TiA) Questionnaire

The dozen or so scales in Table 6.3 overlap, but are often targeting different things–disposition versus current state, or targeting automation versus robotics versus humans. It is useful to examine one scale to understand what it specifically measures. Körber (2019) released his TiA scale as a CC-BY-SA, which allows us to reproduce it here in Table 6.4. Körber (2019) built TiA from Mayer and colleagues' organizational-trust model and Lee & See (2004)'s account of trust in automation. An initial pool of 57 items was reduced to 19 items on six subscales: Reliability/Competence, Understanding/Predictability, Familiarity, Intention of Developers, Propensity to Trust, and Trust in Automation. Responses use a five-point agreement scale (plus “no response”), and several items (5, 7, 10, 15, and 16) are reverse-framed so that higher values indicate more distrust.

#ItemSubscale
1The system is capable of interpreting situations correctly.Reliability/Competence
2The system state was always clear to me.Understanding/Predictability
3I already know similar systems.Familiarity
4The developers are trustworthy.Intention of Developers
5*One should be careful with unfamiliar automated systems.Propensity to Trust
6The system works reliably.Reliability/Competence
7*The system reacts unpredictably.Understanding/Predictability
8The developers take my well-being seriously.Intention of Developers
9I trust the system.Trust in Automation
10*A system malfunction is likely.Reliability/Competence
11I was able to understand why things happened.Understanding/Predictability
12I rather trust a system than I mistrust it.Propensity to Trust
13The system is capable of taking over complicated tasks.Reliability/Competence
14I can rely on the system.Trust in Automation
15*The system might make sporadic errors.Reliability/Competence
16*It is difficult to identify what the system will do next.Understanding/Predictability
17I have already used similar systems.Familiarity
18Automated systems generally work well.Propensity to Trust
19I am confident about the system's capabilities.Reliability/Competence
Table 6.4. Trust in Automation (TiA) questionnaire (Körber, 2019). Items marked * are reverse-coded. Response scale: 1 = strongly disagree to 5 = strongly agree. Reproduced under CC BY-SA 4.0.

Like many of the other trust scales, this includes questions about reliability, predictability, clarity, and user knowledge. This illustrates one of the assertions in Chapter 5 that trust is multidimensional. Like most scales, it treats trust as a value that can be quantified, and refers to `the system' as a whole, ignoring the componential nature of trust. Finally, the questions tend to ignore context–they neither ask nor qualify ratings to understand when trust is appropriate. Nevertheless, this can easily provide a user assessment of trust that is comprehensive and has strong psychometric properties.

Körber suggests individual subscales should be examined separately and not as a full scale. It is feasible to collect individual subscales as well, such as the Trust in Automation items (9 and 14) which include “I trust the system”. Körber (2019) also discusses when a single trust item is enough. Note that the 3-item S-TIAS (McGrath et al., 2025) is a short form of Jian's scale, not of TiA.

6.5.3 Measures of Trust in Artificial Intelligence

Most scales of trust in automation were adapted from trust theories rooted in interpersonal trust, in the context of automation. These often found that the principles used to understand interpersonal trust do not translate directly to automation. So it is also important to know whether constructs related to trust in automation translate to trust-in-AI. We might suspect that the additional autonomy, risks, increased cognitive domain, and anthropomorphic properties will make measuring trust in AI more like measuring interpersonal trust–but it may also involve other new factors. Perrig et al. (2023) showed that several popular trust questionnaires do not easily translate to AI– their factor structures change when used to evaluate AI. Scharowski et al. (2025) ran a pre-registered test of Jian et al.'s Trust between People and Automation (TPA) scale and of the Hoffman Trust Scale for the XAI Context (TXAI). TXAI proved useful as a short state measure, but TPA did not support Jian's original single factor: items 1–5 behaved as distrust and items 6–12 as trust. They advised scoring those composites separately without reverse-coding the negative items into one total. McGrath et al. (2025) similarly revalidated TPA for AI and published a 3-item short form (S-TIAS).

InstrumentSourceNature and lineage
Trust in XAI Context (TXAI)Hoffman et al. (2021); Hoffman et al. (2023)8 Likert items plus an open reason for a named tool after use (Table ch06-xai-trust). Independent validation: Scharowski et al. (2025).
Jian TPA used for AI; S-TIASJian et al. (2000); McGrath et al. (2025); Scharowski et al. (2025)Same 12 items as the automation scale, now with an AI referent. Scharowski recommends using two scores (distrust 1–5, trust 6–12). Adaptation of Jian.
TAIS (Trust in AI Scale)Wischnewski et al. (2026)30 items; bifactor: global trust plus ability, integrity, transparency, unbiasedness, and vigilance. New scale adapted from Muir, Schaefer, MDMT, TOAST, and others.
Semantic-differential AI trustShang et al. (2024)New adjective-pair scale for an AI agent, scoring emotional and cognitive trust separately.
Human-like vs. functionality trustChoung et al. (2023)11 items (6 human-like benevolence/integrity; 5 functionality/competence). Built for a TAM path model of AI acceptance, starting with Mayer-style interpersonal trust constructs.
TrAAITStevens & Stetson (2023)Clinician trust and acceptance of AI in medical/health-care contexts.
Big Six AI trustDickens et al. (2026)20-item scale with 6 factors for knowledge-work settings.
Gulati HCT (HCI)Gulati et al. (2019); Pinto et al. (2022)Human–computer trust for interactive systems; adapted by Pinto et al. (2022) for robotics.
GAAISSchepman & Rodway (2023)General attitudes toward AI (positive/negative), extends beyond trust.
TAIGHA / TAIGHA-sKopka et al. (2026)New scale (and short form) for trust in AI-generated health advice. Domain-specific.
Human–generative AI trustWang et al. (2026); Alon & Levkovich (2026)Validation and adaptation of multidimensional trust items for generative AI / “black box” tools.
Feeling trusted by AIXie et al. (2025)New construct measuring perception that an agent trusts respondent. Not a measure of trust in the AI.
Table 6.5. Questionnaire measures developed or validated specifically for trust in AI.

The measures in Table 6.5 focus on trust in AI in general or in broad contexts (XAI, health advice, generative AI). There are a number of other surveys and questionnaires that have an even narrower focus, including: distance-learning trust in AI (Üstün et al., 2026), hospital follow-up systems (Xie et al., 2025), social-service robots (Cai et al., 2024), a child-centered scale (Ragone et al., 2026), and a human-and-AI trust-attitude scale (Larasati et al., 2023; Larasati, 2025).

There are also a number of attitude measures (single-item AI attitude; AI-interaction positivity scales) that have been developed, and these have been shown to correlate with ChatGPT trust, but are not measures of trust alone (Montag & Ali, 2025; Montag & Elhai, 2025). Finally, dispositional propensity to trust technology (akin to the PTA and PTT-A discussed earlier) was found to mediate AI fear and acceptance (Montag et al., 2023).

Conclusions

Trust has been measured in many ways. The first concern is not the metric, but the context and paradigm trust is measured under. These vary from real-world settings to simulations, wizard-of-oz studies, and lightweight scenarios.

Once a setting is determined, measuring trust frequently starts and ends with short questionnaires or surveys. However, there are many other measures, including interview and linguistic approaches, behavioral measures and neurophysiological systems. Many of these are useful in ways and contexts that questionnaires are not.

Within scale-based trust measures, there are also many options, measuring different aspects and senses of trust. Research has shown that automation-based trust measures do not necessarily translate to AI systems, so care should be used when selecting which instruments, and how they are used.

Acknowledgments

This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.

Author Contributions

YW, BF, JW, LS: Background research and writing for specific subtopics; STM: Conceptualization; background research, additional writing in all sections; general editing.

Conflicts of Interest

The authors declare no conflict of interest.

AI Usage Statement

Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.