Handbook

Chapter 3

Stages of Technology Adoption and Development for Human-Centered AI Tools and Methods

Abstract. In this chapter, we discuss the development life cycle of an AI or automated system using frameworks developed from several perspectives, including user acceptance/adoption, product development, and research and development acquisition, and post-deployment adoption. Each of these frameworks provide specific user-centered ways that human factors and usability researchers can influence, evaluate, and mitigate potential problems. Synthesizing these different perspectives, we propose that there are critical periods at which individual human factors methods and human-centered AI design can have the most leverage. Using this as an organizing principle, we describe a variety of tools, methods, and evaluations that might be applied at different points during a system's life cycle to develop human-centered AI systems and support human-AI interactions.

3.1 Introduction

There are numerous measures, interventions, tools, and methods used by human factors practitioners and cognitive systems engineers to improve the usability and usefulness of AI systems, and adopt a human-centered design approach for AI. Human-centered designers often decry that they need to be involved in the initial development stages in order to have the greatest influence, and build the most useful tool. Although this is undeniably true, it also ignores important practicalities and constraints in technology development. Furthermore, the different methods and tools used for human-centered design each have a critical period—a point during the development of technology in which they will have a maximum cost-benefit.

In this chapter, we examine several existing stage-based approaches to understanding the adoption and development lifespan of technology in general, and specifically for intelligent software tools and AI. Some arise from the user perspective, others from the developer perspective, and others from the acquisition perspective. We will then suggest a new framework loosely inspired by research on human language development: critical periods. This approach acknowledges that human-centered researchers can have different influences in different stages of technology development and adoption, and have developed different tools and approaches for these different critical stages.

3.2 Technology Acceptance Model (TAM) and AI-specific revisions

Any technology can be examined in terms of “stages” in its lifecycle or lifespan. One perspective is to examine adoption from the user's perspective. Perhaps the most popular such approach is the Technology Acceptance Model (TAM) (Davis, 1989), shown in Figure 3.1. This model was initially developed to characterize how users were adapting to new innovations such as personal computers at home and in the workplace. The approach treats the system itself as external to the model. A user forms cognitive judgments of usefulness and ease of use, an affective attitude toward using, and a behavioral response of actual use.

Figure 3.1. The Technology Acceptance Model (Davis, 1989). External design features shape perceived usefulness and ease of use; those judgments form an attitude toward using, which predicts actual system use.

Here, initially, potential users may perceive how useful a technology is (by watching others, through advertisements, testimonials, etc.), and how easy it is to use. Ease-of-use feeds into usefulness, and both impact their attitude toward actually using it, which was considered an affective response. This is the ultimate predictor of actual system use or adoption—the behavioral use. One of the implications of the model is that perceptions about ease-of-use for novices can play an important role in technology adoption, which suggests that good human-centered design principles at early stages are likely to pay off. However, the model itself is not tailored to AI in particular, and does not incorporate many of the additional concerns users have about AI systems (e.g., trustworthiness, alignment, safety, to name a few.). It is also limited to tracing a user's attitude from a fixed “input” toward adoption; it does not address what happens after adoption, or how changes to the system may require additional assessments by the user.

More recent research has examined the TAM in the context of AI and automation. These have typically incorporated additional stages in the TAM, or specified different influence in specific stages. For example, Baroni et al. (2022) proposed an AI-TAM model incorporating AI function, trustworthiness, and output collaboration into the TAM path diagram. They suggested explainable AI (XAI) can have positive influences at many of these stages. Choung et al. (2023) provided support for this by showing trust-in-AI impacted perceived usefulness leading to more positive attitudes about AI. Researchers have discussed fairness, bias, transparency and integrity (Borenstein & Howard, 2021), and Mustofa et al. (2025) suggested that trust and ethics are moderators of the relationship between attitude and adoption. Overall, these adaptations to TAM for understanding AI include:

Ultimately, TAM and its extensions describe whether people will use an existing system. The specific adaptations to TAM suggest many different ways human-centered researchers can gain purchase on the problem, and potentially influence initial adoption. Many of these approaches are covered in greater detail in later chapters (including trust/trustworthiness, explainable AI, fairness/bias, transparency, etc.), and so this handbook covers many possible ways to improve these systems from a human-centered approach. However, TAM focuses on initial adoption and influence of users. There are other ways to understand how human-centered approaches to AI and automation can impact systems, by examining the broader product development timeline.

3.3 Product Development Lifecycles

An alternative approach is to examine how a product is built. Perhaps the most classic model is referred to as a “cascade”, or “waterfall” development approach. It proceeds once through requirements, design, implementation, test, and operation (Royce, 1970). Royce (1970) potentially suggested this model as a warning: without feedback between stages, the process is risky and likely to fail. Boehm (1988)'s response is a “spiral” development model that repeats a cycle: first determine objectives, next analyze risk, then develop and verify, and finally plan the next loop. In theory this helps the system grow in capability and fidelity to the original vision at each cycle. Current popular software and AI product development follows adaptations of this referred to as Agile development, often using its special case predecessor called Scrum, which is characterized by short sprints, a working increment to the product, and inspect-and-adapt reviews (Beck et al., 2001; Schwaber & Sutherland, 2020).

These basic approaches provide an important context for how human-centered design can impact technology development. In fact, human factors and human-centered design requires the same kind of iteration. Gould & Lewis (1985) offered three principles: early focus on users, empirical measurement, and iterative design. Nielsen (1993) turned those principles into a usability-engineering lifecycle of analysis, design, and test, repeated until the interface is good enough. Design thinking (Brown, 2008) describes a cycle of empathize, define, ideate, prototype, and test (sometimes followed by implement). This basic loop has been standardized by ISO for human-centered design of interactive systems (International Organization for Standardization, 2019): understand the context of use, specify requirements, produce design solutions, and evaluate.

Consequently, the software development cycle and the human-centered design cycle are similar, but it is typical that human-centered design is secondary—it represents one of many technical and product constraints that need to be managed. These include security, bugs, licensing, external dependencies, business operations/sales, etc., and so human-centered design may need to fit in with the process and be opportunistic about what they can accomplish. The fast-cycling of Agile approaches also means that human-centered designers have many opportunities to have an impact, but they need to operate at a tempo such that their results are still relevant once they are complete—thus they should understand product roadmaps, redesign goals, and the like.

3.4 Human Readiness Levels

Another approach to understanding technology is the “Technological Readiness Level” framework used by NASA, the government and military funders of research and development programs to understand the maturity of a technology (Mankins, 1995). The TRL describes 9 levels of maturity, from TRL-1 and TRL-2 (describing basic research) to TRL 5-6 (Technology development and demonstration), to TRL-9 (system test and operation). On its own, this framework does not offer much direct understanding of whether the technology supports human users—it can be applied equally to rockets, naval vessels, and software tools. Importantly, there have been many products that have achieved high technological readiness, that have been fielded, but are either difficult to use, or fail to be adopted by users because of a failure of human-centered testing and design. One famous such example is the Mars Climate Orbiter, which had two subsystems that used different units (imperial vs. metric), resulting in the loss of the spacecraft. It is unclear whether a TRL assessment was conducted on the Boeing 737-Max MCAS system, but similar certifications are required by the FAA. However, its failure has been attributed to a failure to understand human factors in aviation safety (Endsley, 2019).

In response to these limitations of the TRL assessment, the Human Factors and Ergonomics Society has developed a similar system to characterize Human Readiness Levels (HRL) (See et al., 2019; Human Factors and Ergonomics Society, 2021). These map onto TRL 1-9, but characterize the extent to which a design has incorporated human-centered design and evaluation. HRL 1 to HRL-3 establish human principles, concepts, and requirements; HRL-4 through HRL-6 mature the design through part-task tests, prototypes, and high-fidelity simulation; HRL-7 through HRL-8 verify and validate human–system performance with representative users on the completed system; and HRL-9 is operational use with systematic monitoring.

The guidelines suggest that HRL should match TRL, because by TRL-6, the engineering design is often fixed, and human needs become expensive or impossible to incorporate (Human Factors and Ergonomics Society, 2021). HRL can often precede TRL, using various methods for understanding envisioned worlds (wizard-of-oz studies, part-task tests, human-centered requirements gathering, and basic human research).

Like the other lifespan approaches, the TRL and HRL model offers some understanding of technology development, adoption, and acquisition. In particular, the HRL model can help technology developers make stronger cases for including human-centered design and assessments as a critical part of the technology development program.

3.5 Threat-based stage analysis

Most of the approaches we have examined describe the timeline of development through adoption. Although the cyclic nature of several of these approaches can certainly incorporate user feedback and evaluation, it is also the case that even for fixed systems that have found a community, there are stages of deployment and use, as people learn the tool, become reliant on it, explore different use cases, and experience failures.

A framework that addresses this collaboration lifecycle (Ahmed et al., 2026) focuses on sociotechnical risks as AI is adopted and fielded. It maps human–AI work across four sequential stages—task allocation (how tasks and authority is divided among a team of humans and AI systems), interaction (communication issues, explanation, cognitive burden/overload), feedback (continuous updating to enable appropriate trust adjustments), and adoption risks (mostly secondary issues such as de-skilling, automation bias, and AI anxiety). Although these stages do not preclude continued development and improvement of an AI or automation system, the focus is on stages of activities ongoing amongst a team that is incorporating these tools into their work.

Ahmed et al. (2026) described six recurring risk clusters that are observed across many domains (including media, business, to education, healthcare, and defense): trust miscalibration, cognitive burden, accountability gaps, capability erosion, goal misalignment, and AI anxiety/technostress. Importantly, they describe a “fragmented mitigation landscape”: for these different risk clusters, there are individual mitigation frameworks, but many apply narrowly to specific domains or risks. Because these risks are socio-technical, their proposed mitigations are often central to human-centered design. For example, for the risk cluster involving cognitive burden and overload, many human-centered solutions (from automatic monitoring for fatigue, to minimizing distraction, to enabling the user control tempo) have been proposed, and it is likely that for any system, user testing is needed to determine acceptable and useful approaches.

3.6 Critical Periods

Overall, these different stage-models of technology development and adoption span from initial conception of a tool to late stages where a team must change how they work to adapt or accommodate the tool. In ideal situations, a strong human-centered analysis will precede development, as it is likely to help understand what the problems are and what the viable solutions are. But this is not always realistic. But it is also certainly the case that many products have completed the entire product development timeline and failed because they do not address the problems real users had, or were difficult to use or understand. Sometimes, conducting human-centered design or analysis can have no impact if it is done too late—after the basic system architecture is developed. It may require a complete refocus of the system to help address the issues and problems people really have, rather than the solution a piece of technology offers.

However, these two extremes are not always the case. It can often be true that resources or time does not exist to engage human-centered design prior to a proof-of-concept; it can also be the case that human-centered design can have an impact on technology that is quite mature—as many of the mitigations described by Ahmed et al. (2026) demonstrate. Furthermore, human-centered design can sometimes demonstrate effective wins at later phases, which can provide buy-in for future product development.

In response, we propose a notion inspired by research on human language development: critical periods. The basic notion is that many human-centered design and evaluation techniques are most appropriate, and can have the biggest impact, at specific stages of technology development. Some are appropriate before the system starts development; others are only available once a community of users exists, or a team has incorporated a tool in their work. We use a development-lifecycle model (Figure 3.2) that includes initial planning, development, deployment, change management, and maintenance.

In comparison to the other phase models in this chapter: waterfall, spiral, Agile, and design thinking say how planning and development are conducted (Royce, 1970; Boehm, 1988; Beck et al., 2001; Brown, 2008); Gould and Lewis's iterative principles and Nielsen's usability-engineering cycle are the human-factors versions of these cyclical design approaches (Gould & Lewis, 1985; Nielsen, 1993; International Organization for Standardization, 2019). HRLs score human-use maturity across the lifecycle timeline (1–3 planning, 4–6 development, 7–8 deployment, 9 maintenance) (See et al., 2019; Human Factors and Ergonomics Society, 2021). TAM extensions mostly focus on user adoption immediately following development, and Ahmed's risks focus on later deployment stages.

Many elements of these critical periods are logically constrained by the feasibility, cost, and information gained from applying a method at distinct points in the developmental timeline. Developing user support or assessing trustworthiness of an envisioned tool may be wasted effort, and a better use of resources at that time would be to evaluate user expectations and stakeholder desirements (HRL 1–3, not 7–8). On the other hand, many methods may do little good if applied “too late”—as when HRL lags TRL, and the system is already frozen (Human Factors and Ergonomics Society, 2021).

In the remaining sections, we will describe the stages in Figure 3.2 and some candidate methods and interventions proposed to address them. In many of these cases, the methods will be dealt with in much greater detail in subsequent chapters of the handbook, and we will offer pointers to those resources as well. Of course, these suggestions are only representative of the many different human-centered approaches that exist.

Figure 3.2. Critical periods in the technology development lifecycle, and distinct goals of human-centered AI design and practice applicable during different stages. TAM model typically describes interventions that can be done during planning and development; (Ahmed et al., 2026) describes risks that occur after deployment; HRL levels map roughly onto the timeline from left to right.

3.7 Understanding needs of users/supporters and other stakeholders

When developing technology and tools that utilize AI and machine learning, the first step should be to identify and understand the needs of users and supporters of the tool or technology. There is a large body of research on stakeholder analysis (Brugha & Varvasovszky, 2000; Schmeer, 1999; Kennon et al., 2009; Reed et al., 2009), and a growing literature on stakeholders AI systems (Hoch et al., 2023; Hoffman et al., 2023; Plass et al., 2022; Park et al., 2022; Deshpande & Sharp, 2022; Subramanian et al., 2024; Kim et al., 2024), including government guidance on responsible AI acquisition and agency reporting (Office of Management and Budget, 2024; Office of Management and Budget, 2024). These have used a variety of approaches, which we focus on next.

3.7.1 Who are the stakeholders?

A number of different stakeholders have been considered and depends on the particular technology but these have included end users, developers, researchers, regulators and legal experts, policy-makers, contracting/procurement, trainers, system evaluators, program managers, business managers, employees, AI experts, data scientists, machine operators, data protection officers, business experts, HR teams, physicians, and patients, to name a few. Notably, Hoffman et al. (2023) examined stakeholders with respect to their desirements for explanation, and found that most stakeholders played multiple roles, and so it is important to elicit feedback not based on job title but based on various goals and use patterns of any individual.

3.7.2 Methods for stakeholder analysis

There are many approaches to understanding stakeholder needs when planning a project that involves intelligent software. We discuss five that may be useful depending on the stakeholders, resources, and timeline of a given project. A useful strategy is to use a minimum-necessary rigor approach (Klein et al., 2023), striking a balance between effective information and efficient data collection that has the best chance of impacting early design.

Informed stakeholder survey.

Surveys are typically an inexpensive and fast way to get stakeholder feedback. The advantage is that many stakeholders can be reached at low cost. Such surveys may cover hypothetical scenarios, evaluate specific values and concerns, and allow more open-ended feedback. Creating representative, valid, unbiased, and informative surveys can be difficult, and is a critical skill for human factors researchers. Important issues may not be known before the survey is written, so a poorly designed survey may miss what matters to different stakeholders. Different stakeholders may also require different lines of inquiry, so multiple surveys may be needed even if each has only a few respondents. Response rates can be extremely low, and the effort needed to encourage responses may be better spent on other approaches.

Focus groups and interviews.

Focus groups and one-on-one interviews are commonly used to gather stakeholder concerns (Chapman & Shigetomi, 2018). Interview methods can allow a more intense understanding of situations, use cases, reasoning, and potential problems. Focus groups and interviews allow discussion and discovery that surveys often miss. Successful implementation often requires skilled facilitators who also know the domain (e.g., AI and the implementation field such as health care, military, transportation, etc.). The material can be costly or time-consuming to code, and without systematic process the thematic analysis can be biased by whoever extracts the themes. We discuss some of these methods in detail in Chapter 6 in the context of assessing trust in automation and AI.

Charrette for stakeholder engagement.

For especially complex projects that are likely to have large impacts on—and opposition from—communities, the charrette process (Lennertz, 2003) has become an important planning and engagement technique. These are usually extended meetings among small groups of stakeholders, often facilitated by an independent and experienced moderator, that attempt to foster ownership and reduce conflict among parties with conflicting interests.

Compared with focus groups, a charrette is more intense and focuses on participatory planning rather than feedback alone. For many AI implementation projects a charrette may not be needed. The advantage emerges when a community may (rightly or wrongly) object to the program, and a charrette can anticipate that and allow input. Any system that will rely on automation and algorithms for decisions previously made by humans, with safety or fairness impacts, may require many stakeholders to identify a viable trajectory for change. The disadvantages are experienced facilitators, extensive planning and event coordination, and a time commitment from many stakeholders; the method is designed to turn planning into a community process.

The Stakeholder Playbook.

To facilitate stakeholder analysis for AI systems, Hoffman et al. (2023) proposed the Stakeholder Playbook, with a special focus on explainability. Like a football playbook, it identifies both roles and goals (related to desirements and requirements) found in previous studies. Stakeholders often play several roles, which can lead to different and sometimes conflicting goals. The playbook can help identify which stakeholders are especially relevant for implementing AI, some of their common concerns, and a basic guideline for interviewing others. Table 3.1 summarizes major role–goal themes; Figure 3.3 shows the more detailed end-user playbook.

StakeholderGoals
JurisprudenceAnalysis of system biases, assumptions, and bounds
Contracting/procurementAnalysis of fitness for use in the operational environment
Development team leaderAnalysis of how the system integrates with other systems
DeveloperAccess to use cases that represent the operating context
System integratorExplanations at a detailed technical level to “look under the hood”
TrainerAccess to edge cases to support training users to anticipate problems
System evaluatorAccess to feedback from prospective users
Policy-makerDescriptions of system limitations and weaknesses
End-user; adopterCost–benefit analysis of tools with respect to goals
Table 3.1. Stakeholder Playbook major themes for AI systems (Hoffman et al., 2023).

Figure 3.3 shows the more detailed end-user playbook, including explanation desirements, access requirements, and cautions. A focus group can walk through a subset of those items rather than treating the playbook as a substitute for analysis.

Figure 3.3. Example avenues of investigation for end-user stakeholders (Hoffman et al., 2023). Used via CC-attribution license.
The premortem exercise.

The general stakeholder methods are useful to understand how the product will fit within its organization, market, or community of users. For smaller teams that have representative members from a number of stakeholders (designers, developers, management, users communities, etc.), a lightweight alternative is the premortem exercise (Veinott et al., 2010). It has proven successful at improving plans by identifying potential problems from the perspective of multiple stakeholders.

The premortem is designed to take around 30 minutes during initial project meetings, when stakeholders are already engaged in planning. A leader makes a statement similar to the following: “Let's assume we are several years in the future, and the project we are trying to accomplish has failed. Take a few minutes to think about why it has failed, and write the reasons down.” The prompt can be narrowed—implementation failed, or the system was implemented but did not produce the intended outcomes—but the open-ended probe is usually enough to surface unforeseen kinds of failure.

The group takes turns describing their reasons and votes when they share a similar reason, to mark importance. When that discussion is complete, a second round identifies which potential failures can be anticipated and avoided, and assigns actions to stakeholders.

The advantage is that it is lightweight and can be run without substantial training. Compared with interviews, focus groups, and charrettes, it is less intense but can still be effective. Its focus on project failure directs discussion toward a common goal of successful implementation. The short timeframe also limits how well it can identify complex problems or solicit feedback from a broad (and possibly antagonistic) set of stakeholders. The approach has been evaluated in a number of contexts, including for the hypothetical planning of AI-based systems. For example, Kannan (2026) examined a structured premortem focused on AI systems, and found that additional prompting about user expectations helped generate reasons for failure that were more likely to be related to AI, were more likely to be controllable, and were judged higher quality than unprompted pre-mortem reasons.

3.7.3 Summary of stakeholder analysis in Human-AI interaction

This section covered a number of systematic approaches to understanding stakeholders. Some are simply to help design and usability, others help characterize social impacts of a system and community opposition. For human-AI interaction, several findings have emerged from stakeholder analysis:

In the next section, we will discuss user expectations, which may come out of stakeholder analysis, or could be a more focused evaluation of existing or prospective users.

3.8 Understanding user expectations

Another goal that can be worked toward early in the development process is to understand expectations of users. This can take on various forms, but understanding the initial goals of users can help design a system that is responsive to those expectations, or develop training and marketing that will inform users about the capabilities and limitations of the system. Chapter 4 of this handbook covers a number of anthropomorphic expectations users tend to have in AI and automation: they often expect that it will behave both socially and cognitively like a human, although they can learn differently.

Zhang et al. (2021) examined human expectations of AI team-mates in multiplayer online games, and suggest expectations cover a broad range of behaviors: “instrumental skills for in-game tasks, shared understanding between humans and AI, communication capabilities, human-like behaviors and performance”. Many expectations and desires of players were not current capability of the system, and users would make suggestions about how their expectations were not met (e.g., “I think voice commands would be really good”), which could lead to changes in design or development. Thus many expectations center on ability or performance capabilities of the system.

Linja et al. (2022) reviewed user expectations on autonomous driving that appear in social media forums, and suggested they fall into several categories, including legality, safety, transparency, and anthropomorphic performance. Adding to these performance/capability and alignment (discussed further in Chapter 4), we have useful potential evaluation guidelines to understand user expectations of AI technology. These include:

  • Performance/Capability: The system is able to accomplish the task/job it is meant to under normal conditions. How fast and accurate is it compared to humans?
  • Legality: The system follows relevant laws, rules, and regulations. Does it violate fairness? Does it use information it should not? Does it violate “rules of the road” that are not explicitly codified?
  • Safety: The system does not make errors that endanger users, cost money, and violate trust. Does the system change its behavior or shut down in unknown situations? Does it include failsafes and checks to ensure its information is accurate?
  • Transparency: The system accomplishes work in a way that users understand, and when it exposes inner workings that information is correct. Does the system follow its plan? Is it easy to describe what the system is doing or predict what it will do next? Is the system inconsistent in what it says and what it does? Does it follow or ignore directions?
  • Cognitive Anthropomorphic Behavior and alignment: The system accomplishes its work in the way a human would. Does it do difficult tasks easily but fail at easy tasks? Can you tell it is an AI and not a human? Does it work in alignment with human goals?

These areas are simply ways of evaluating user expectations, and not a checklist for design. For example, we expect other drivers AND self-driving cars to follow laws and safe driving practices, but these are sometimes at odds, and most driver-assistance technologies allow the vehicle to (illegally) exceed the speed limit. A user might expect an AI to play a team game like a human would, and not “cheat” through superhuman reaction time and planning abilities, but it would not be realistic to desire an AI diagnostic system to fall prey to the same biases a human might.

These expectations can be used in several ways. We have used them to structure user interviews, to help understand specific concerns users may have. Kannan (2026) used these same categories as prompts in a pre-mortem exercise to help participants focus on how different aspects of AI might lead to failure. It may also be used to focus early development on a set of specific goals related to user expectation.

Many of these expectations (and ways of measuring them) are covered in future chapters. For example, Chapter 4 covers anthropomorphism and ways of assessing AI capability to “think”. Chapter ethics covers fairness, transparency, and safety. Finally, Chapter 5 covers models of trust/trustworthiness, and Chapter 6 many means of assessing trust, which involve each of these dimensions.

3.9 Cost-effective, efficient, and phase-appropriate evaluation

One of the main contributions made by human factors researchers is in developing assessments for evaluating aspects of the system. Later chapters cover many of the relevant assessments, approaches and instruments, including measuring trust and distrust, knowledge and understanding in a system, mental models, satisfaction in explanations, and performance in a work domain (including time, errors and costs).

3.9.1 Approaches to Evaluation

It can take substantial time and resources to fully assess the effectiveness of many AI systems. The system's learnability and the technology's usability and utility must be convincingly demonstrated by empirical evaluations of systems where humans and AI systems collaborate to complete tasks, often in comparison to the current system. Klein et al. (2023) advocated a lightweight evaluation approach for evaluating Human-AI systems called “Minimum Necessary Rigor”, with the goal of providing fast relevant feedback for designers and developers. They included ten major recommendations (see summary in Table 3.2)

CodeRequirement
P1Test individuals who will be beneficiaries of the technology.
P2Use minimal training to understand naive users.
P3Small-N studies are often sufficient
T4Tasks should be ecologically valid
T5Evaluate specific hypotheses individually rather than a large complex design
T6Conduct pilot studies to refine the methods, materials, and procedures.
T7Run a two-condition, between-participants study.
T8Run a two-condition, within-participants study.
A9Do not apply statistical analyses that are opaque or complicated.
A10Be prepared to set a high bar for determining whether or not the AI is good
A11Consider practical significance.
Table 3.2. Summaries of the requirements identified by Klein et al. (2023) for “Minimum Necessary Rigor” for Human-AI developers. P: participant-related. T: task-related. A: Analysis-related.

Evaluation extends across the development timeline and is integral to the system, but often is scheduled to occur late in the development in order to demonstrate effectiveness of the system and the value of investing in new technology (Davis, 1989; Baroni et al., 2022; Choung et al., 2023; Mustofa et al., 2025), or may only be available once a tool is in use (Human Factors and Ergonomics Society, 2021). In comparison, the MNR approach is perhaps more appropriate as a formative assessment: a lightweight early test to help steer development and identify major strengths and weaknesses. And as discussed earlier, human readiness can work ahead of technological readiness by using simulations, scenarios that users examine and evaluate, or wizard-of-oz (WOZ) approaches, where the AI or automation is fake or generated by a human. The MNR is also consistent with the design-thinking prototype/test cycle (Brown, 2008; International Organization for Standardization, 2019) that is easily integrated into cyclical product development.

3.9.2 Assessing Explainability, Transparency, and Trustworthiness

Lack of transparency and explainability has been highlighted by research as one of the largest barriers to the implementation of AI and machine learning in industry domains (Markus et al., 2021). For a human team member to properly utilize the system, the operator may need to understand the process that the AI will be utilizing as well as trust that the outcome will be favorable and usable. Explainability and trustworthiness are important for ensuring the accountability of the tool (Chamola et al., 2023). When users have a transparent view of a system's underlying principles, rationale, logic, and goals, it increases trustworthiness in the system (Linja et al., 2022). Many methods for assessing trust are described in Chapter 6.

Although these assessments may require usable prototypes or complete systems and so be most accurate late in the product lifecycle, transparency and trust have also been considered development-stage TAM/AI-TAM constructs (Baroni et al., 2022; Choung et al., 2023). Furthermore, there can be negative consequences—Ahmed et al. (2026) warned that explanation load can raise cognitive burden and reinforce automation bias. These all create a challenge, because design-time implementations of transparency and explainability may not improve (and may hurt) user performance and trust, which can only be known directly by testing users. This again suggests the importance of a lightweight evaluation process appropriate for cyclical development.

3.9.3 Assessing User Understanding of the System

Maintaining human focus within an AI system requires that users comprehend the system. A user who has an understanding of the data used as training data will have a clearer understanding of its boundary conditions (Liao et al., 2020). Increased understanding of an AI system helps users predict what outputs are valid and reduce errors if the AI produces an edge case outcome. When end users understand how the AI processes data, it also has been shown to reduce the negative emotional impact and increases acceptance of malignant outcomes (Branley-Bell et al., 2020). Chapter 7 covers assessments of mental models and system understanding in more detail. Some of the consequences of an accurate mental model include:

  • Identifying work practices and how it is used, and fits into the work system
  • Being able to mentally simulate the system to predict what it will do
  • Having other expectations about its behavior
  • Understanding its competency envelope
  • Knowing its failure points, and how to work around them or recover from them
  • Basic understanding of input, representation, algorithms, and expected results
  • Identifying costs and benefits of using it in comparison to other solutions

Although it can be helpful to engage in a detailed evaluation of user knowledge and understanding of a system among novices and new users (maybe to develop better onboarding, training, or UX), this understanding is also important to assess once the system is in place, and users have begun to employ it in their work process. At this point, it can be valuable to understand where the mismatches are between the system's intended use and user's understanding. This can be used in several ways:

  • To develop training based on expert users
  • To change the system to support how it is actually being used
  • To identify capabilities that are not being adopted and discover why
  • To invest resources in common usage patterns to make them better or more efficient

It is common to do usability assessments on novice users, which can help with adoption and onboarding. However, valuable information is also gained from more experienced users, and understanding what they know about a system can have many benefits.

3.9.4 Assessment of the Tool vs. Assessment of the System

It is important to recognize that the actual use of a tool often emerges once it becomes embedded within a system—a team of users, processes, and other tools that have interdependencies and interactivity. A specific tool is likely to be used in unanticipated ways, and also require additional changes in the work system to accommodate it. For example, LLM systems currently have replaced programming code development for many tasks, but have been adapted to the work in interesting ways such as as code, automatic unit test development, and have required new processes involving supervision and evaluation of security and mission-critical code paths. Evaluating the LLM tool in terms of benchmarks is only part of the task, which also needs to characterize the overall improvement in productivity in the context of a loss of close understanding of the codebase by developers. Thus, when evaluating a tool, it is important to understand the context it is used in and how it impacts and changes the work being done around it.

3.10 Identifying capabilities and limitations of the system and the users

When developing AI for a specific domain or task, the capabilities and limitations of both the AI and the human team members will determine the extent of utilization of the technology. As discussed in Chapter 2, Amershi et al. (2019) identified a set of 18 guidelines for human-AI interaction—the first two involve understanding specifically what the system can and cannot do, and numerous others involve guidelines for errors and other limitations of the system. Borders et al. (2024) have gone beyond this, and argued that the user's mental model includes (a) understanding of the strengths and weaknesses of the system, and (b) understanding their own strengths and weaknesses. Similarly, automation bias has been investigated in terms of users' willingness to over-rely on automation (Mosier & Skitka, 1999; Skitka et al., 2000). Mosier & Skitka (1996) argued that one (of several) explanations of automation bias was the “authority hypothesis”: a user may trust the automation's ability or advice is better than their own, and so give the automation the authority to perform autonomously. This decision process requires the user has a (possibly flawed) understanding of their own strengths and weaknesses in comparison to the automation's.

On the other hand, users are often reluctant to adopt technology because of their perceptions of their own strengths and weaknesses and those of the technology. For example, Yates et al. (2003) argued that decision aids are often ignored by users because they are perceived as providing information that is irrelevant to their decision; perhaps believing the automation helps with aspects of the decision the users are already good at and not aspects they have difficulty with. In this case, automation whose perceived strengths and weaknesses overlap with users' strengths and weaknesses may be abandoned. Understanding the strengths and weaknesses can be used in several ways. If limitations are identified and communicated effectively to end users, then those users will better be able to determine the types of tasks that the AI or automation can contribute to (Chowdhury et al., 2023). This also allows for proper task-allocation stage (Ahmed et al., 2026). Also, output quality and collaborative intention are user-focused measures of task allocation (Baroni et al., 2022).

To assess user and system strengths and weaknesses, Borders et al. (2024) proposed a technique they referred to as the “Mental Model Matrix” (see Table 3.3). This evaluation provides a structured guide that might be used to guide user interviews, or to analyze observed behavior or analyze artifacts in order to uncover users' understanding of themselves and the technology. This might be used early in the design process to help developers match the technology to the strengths of and weaknesses of users, which may impact system design. It may also be done on prototypes of the system to understand and improve interface or training. Finally, it can be done with users of mature systems to understand how the system can be improved.

CapabilitiesLimitations
System/AIGood at driving in standard conditions and in known locationsTrouble off-road or when roadway is obscured
UserExperience in driving; good at detecting and navigating unknown risks.Mental fatigue; limited attention
Table 3.3. Hypothetical Mental Model Matrix for Tesla FSD (Full self-driving) mode.

3.11 Understanding the Work Domain and Sociotechnical System

Many new applications of AI will fundamentally change how work gets done. Others will integrate into existing systems and improve or replace work previously done by humans or other less capable systems. A fundamental failing of many AI systems is that their design ignores the work domain in general, and how it not only has specific function, but changes how other work occurs surrounding that function. The critical time period for evaluating this may start early (if a work domain already exists) or late (if the AI provides capabilities that creates fundamentally new work systems), but in either case human-centered designers and human-AI interaction researchers should help focus the AI design on supporting the intended work.

The theoretical perspectives of situated cognition are central to the goals of understanding a work domain (see Clancey, 1997). Clancey (1991) argued that “...in saying that cognition is situated, we mean that reasoning processes are not merely conditional on the environment, but are inherently brought into being during an interactive process (more precisely, an interaction of different systems, neural and environmental).” In this view, it is a mistake to consider work with a system at the level of a task accomplished by a user, and instead it should be examined as how a system (involving both machines and humans) works in the context of their workplace, their organization, and their culture.

Clancey (2006) discussed a variety of approaches to understanding the work practices involved in such systems. Some of the approaches include:

  • Participatory Design. Embedding users within the design team, or design team members within the work context.
  • Contextual Design. A user-centered process incorporating ethnographic methods and field studies of users (Beyer & Holtzblatt, 1999).
  • Situated Cognition and Situated Action. Examining the cognitive processes and behaviors within the context of the work system (Clancey, 1997).
  • Ethnography. General Anthropological and sociological approaches to understanding a work system (e.g. Suchman, 1983)
  • Ethnomethodology. Approach focusing on the social construction of work systems (e.g. Clancey, 2006)
  • Cognitive Work Analysis. A structured ecological approach intended to guide the design of tools in the workplace (Rasmussen et al., 1990; Vicente, 1999)

These methods have been used in many cases in which technology is being introduced into work systems. There may be additional particular challenges when that technology is AI. Clancey et al. (1998) describe applications of this approach to Brahms, a work systems design simulation that provides guidance for integrating automation and intelligent systems into the workplace. Furthermore, Ahmed et al. (2026)'s four stages of human-AI collaboration (task allocation, interaction, feedback and adoption risks) are all situated within a work system, and characterized risks that often appear only once a tool is in the work system.

3.12 Enable Users to Learn a System

Chapter cognitive-tutorials covers tutorials and other means for helping users learn the system. User guides, documentation, tutorials, and support forums are an under-appreciated component of user-centered design. In their chapter on Documentation and User Support, Shneiderman et al. (2016) describes a broad range of user help documentation, training, and tutorials used for computer systems. Interestingly, because most computer systems (especially well-designed ones) are straightforward, most of the advice for development of user support comes from a perspective of usability, design, and readability.

However, even when user-centered design approaches are used for AI-enabled and automation-based systems, there can be unique challenges for users. Learning the system must go beyond simple guides describing the functions of different buttons, because a user must develop a mental model of the AI's capabilities—essentially a theory of mind of the AI. Mueller et al. (2009) proposed experiential training in the form of an “Experiential User Guide” as a way of allowing users to learn the more challenging aspects of automation and AI. This approach is grounded in cognitive research on learning, drawing on Cognitive Load Theory (Sweller & Chandler, 1991), ICAP (Interactive-Constructive-Active_Passive) framework (Chi & Wylie, 2014), and Bjork & Bjork (2020)'s notion of desirable difficulty. Such user guides are “cognitive” tutorials (Mueller et al., 2021) that may be necessary for helping users develop a better mental model of a system.

To support users effectively, slow and expensive classroom-based training may be impractical (Gratton, 2022), necessitating more creative approaches. While there seems to be limited literature about training strategies to support workers adjusting their skills to use AI tools, work-integrated learning (WIL; Rampersad, 2020) and on-the-job training (OJT; Zsambok et al., 1997) has been suggested to enhance workers' skills while on the job (Rampersad, 2020). By immersing users in real-world scenarios, WIL bolsters technical competence. Workers will also need to adopt and adapt to ongoing reskilling and upskilling to keep contributing their skills to the system (Zirar et al., 2023).

One additional way to support learning is to empower and enable users to support one another, which appears in Figure 3.2 as Support Collaborative Knowledge. For widely used AI systems, user forums, user groups, social media, and informal encounters with other users are an important way of learning a system.

3.13 Summary

This chapter described several frameworks for understanding the lifecycle of AI products, and many different approaches taken by human factors researchers to assess, improve, and support the development of human-centered AI and automation. We argue that different approaches have “critical periods” of maximal impact, and technologies at different stages of development and deployment can benefit from different methods and tools. The interventions appropriate during different stages of this lifespan model provide a basic roadmap for this entire handbook, and more detailed methods, strategies, and models applicable at different stages are available throughout.

Acknowledgments

This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction; Michigan Technological University; Spring 2024. The evaluation of stakeholders was adapted from work completed as part of FRA contract 693JJ6-24-C000036; Lightweight evaluation, training, and user collaboration for Human-AI Work.

Author Contributions

EM; LS: Writing and background research; STM: Conceptualization; background research, writing in all sections; general editing.

Conflicts of Interest

The authors declare no conflict of interest.

AI Usage Statement

Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.

How to Cite This Chapter

Matas, E., Sprague, L., & Mueller, S. T. (2026). Stages of technology adoption and development for human-centered AI tools and methods. In Shane T. Mueller (Ed.), A Handbook of Human-AI and Human-Automation Interaction. https://pages.mtu.edu/~shanem/human_ai/