Chapter 1
Demystifying Automation, Artificial Intelligence, and Intelligent Software Tools and Applications
Abstract. Algorithms implementing automation, artificial intelligence (AI) and other intelligent software tools are not a unitary thing. They involve a number of different concepts, purposes, algorithms, and applications. Understanding the basics of their operation can be important for improving human usability, but human factors researchers and practitioners and cognitive systems engineers interested in designing and evaluating usable AI and automation often have limited experience in such systems. The goal of this chapter is to demystify some of the common concepts in AI, automation, and intelligent software. To do this, we examine general principles of operation, representations, learning and computation algorithms, and conclude with an examination of some of the more common applications of AI and automation systems.
1.1 Introduction
When we discuss AI, Machine Learning (ML), automation, optimization, and other intelligent software tools, we often group them all together into a single general construct when asking how humans interact or understand them. But in reality, AI, ML, optimization, and automation represent dozens of different approaches applied to many different problems. Although these different approaches are well understood by researchers implementing AI or ML, they often represent black boxes to users and to cognitive systems engineers and UX/UI professionals whose job it is to think about how humans will interact with the system. Consequently, one of the goals of this chapter is to help demystify many of the algorithms and approaches used to implement automation and artificial intelligence, and give some concrete examples to help novices understand the algorithms. Although a deep understanding of particular computations may not always be necessary, having a good mental model of the process can help users understand the boundary conditions of the system: what it is good at, and where it is not appropriate. Understanding these boundaries helps users know where to trust the system and where to anticipate that it may fail. Furthermore, it helps users understand the errors these systems make. Without being able to interpret why an error is made, a naive user may attribute it to the system being bad or broken. When systematic errors can be anticipated, this helps encapsulate and contextualize the problems, provide greater understanding of the boundary conditions, and develop workarounds.
Consequently, this chapter will discuss many aspects of the algorithms underlying AI and automation, with the goal of providing some basic understanding of the landscape of this domain. It will be especially helpful for human factors researchers and practitioners who may not have a strong background in the technical side of AI and automation.
We will organize this chapter by thinking about AI and automation at several levels of abstraction, which might all be considered part of the system architecture. We will start with a few very general principles underlying many kinds of systems. Following that, we will examine some of the common knowledge and information representation approaches used across such systems (e.g., how information or data is modeled; for example as a markov network, a tree, a vector). Then, we will discuss some of the distinct inference/optimization approaches used (the way in which training data is transformed into the representation, learned from data, or parameters are optimized). Then, we will discuss different classes of computational algorithms (the computational process by which input is transformed into output; e.g., via a convolutional neural network). Finally, we will describe a number of application domains and the systems used in particular domains. It may be true that many individual systems involve a variety of different approaches along each of these dimensions, and many applications of AI use multiple combined systems to solve the problem. Later chapters take up human-centered methods (Chapter 2), lifecycle methods (Chapter 3), anthropomorphism/AGI (Chapter 4), trust (Chapters 5–6), mental models/transparency (Chapter 7), explainability (Chapter 8), ethics/fairness (Chapter 9), user support (Chapter 10), large language models and generative AI (Chapter 11), and human–machine teaming (Chapter 12). This chapter focuses on technical demystification.
Although users of such intelligent software tools may not need to know detailed aspects of the computations underlying the system, oftentimes it can make a difference for a human factors researcher. For example, understanding aspects of the architecture of the system may help identify aspects of the system that could be used to improve usability. Understanding inference algorithms may help determine whether certain adaptive systems are practical in terms of time or cost. When thinking about automated systems, a useful distinction is training versus deployment: many systems estimate parameters or generate rules during training, then process data and create output when deployed, and although the algorithms for training/creation need to be coupled with the system deployed algorithm, they are often distinct, sometimes general learning or optimization algorithms can be used across different classes of algorithms, and the choice can impact both costs and performance.
1.2 General Principles of Intelligent Software Tools
In this section, we will describe and define some of the most important high-level approaches to AI, automation, and other intelligent software systems. In many cases, these concepts are very general and have no agreed definition, but it is useful to understand how these concepts are used by various communities.
1.2.1 Automation
Automation has been defined as “The process of structuring a task, environment, or system to reduce labor and decision-making time through offloading labor and decision-making to a human-created system” (Goldberg, 2011). While often described via circular definitions (automation is often defined as simply, “to automate”), automation is largely understood as processes that reduce human labor by passing that labor and activity to a machine, algorithm, or other devised mechanism. Automation covers a wide range of technology developed over hundreds of years, and includes simple mechanical systems (a parking meter automates the job of a meter reader; a stop light automates the job of a traffic cop, an industrial machine may automate the job previously performed by a worker, an autohelm controller automates the job of keeping a boat on course).
Human factors researchers have classified how much and what kind of work is automated. Levels-of-automation scales (Sheridan & Verplank, 1978; Parasuraman et al., 2000) distinguish stages such as information acquisition, analysis, decision selection, and action implementation, each automatable to different degrees. The International Society of Automotive Engineers (SAE On-Road Automated Vehicle Standards Committee, 2014) has followed suit to create SAE J3016, a taxonomy that specifies six levels of automation–from completely human drivers to completely autonomous.
To automate something is to reduce human involvement in a task. It can be as simple as the natural proceduralization of a task to reduce decision-making labor or as complex as training an algorithm or building a machine to do the task for you. Creating an automated system generally takes time and effort but reduces time and effort later on. Additionally, not all tasks/environments/systems can or should be automated. Automation is useful within task domains that are repetitive, time-consuming, or error-prone, but may not be appropriate in situations that heavily rely on in-the-moment decision making.
Some technologies that are considered or appear to be automation may not meet our definition, because they are designed to shift the labor burden rather than reduce it. For example, grocery store self-checkouts and self-service gas station pumps are technologies that replace a job previously held by a human, but their intent is to make it possible to shift the human labor safely or economically to the buyer to reduce costs. Similarly, IKEA furniture, tools, and other housewares sold with 'assembly required' replace a factory worker and reduce shipping costs, but simply shift work from one human to another, or perhaps from a machine in a factory to a human in their living room. In contrast, an ATM automates the work of the bank teller (counting money keeping records) but does not materially shift the work burden of the customer, and might be considered true automation.
1.2.2 Artificial Intelligence
Artificial Intelligence (AI) is a blanket term for many different approaches in many different communities. Wang (2019) provided one of the most exhaustive and careful arguments understanding and defining AI. Wang argued that hundreds of forms of human intelligence have been defined (see Monett & Lewis, 2018), and that AI can be understood as an abstraction of these human intelligences. Moreover, human intelligence can be understood in general as a capacity to adapt to an environment with limited information. Wang argues that AI typically involves one of five kinds of approaches to simulate or creating human-like intelligence: Structure (simulating a brain), Behavior (simulating intelligent behavior), Capability (ability to solve problems effectively), Function (possessing of commonly-understood general cognitive functions such as memory, learning, planning, perception, etc.), and Principle (capable of fundamental principles such as rationality). Although this approach is not universally embraced (see Monett et al., 2020), many approaches to AI fit into one or another of these categories.
1.2.3 Machine Learning
Machine learning is a broad term to describe a variety of computational algorithms that are typically involved in producing behaviors or representations (classifications, labels, categories, control, generative AI systems) based on data, rather than relying on hard-coded or programmed rules (El Naqa & Murphy, 2015). An important distinction in machine learning is among supervised learning (predicting from labeled examples), unsupervised learning (identifying structure in unlabeled data), and reinforcement learning (policies via reward from interaction; see Section 4.4). Many modern systems combine these (e.g., pretraining then fine-tuning) The notion of 'learning' is quite general as well, and although may involve specific learning algorithms such as reinforcement learning or TD-learning (see Sutton & Barto, 2018), it may also involve other means of optimization (e.g., as in support vector machines, genetic algorithms, or Bayesian inference), or simply maintaining a database of past incidents (as in Case-based reasoning systems and K-nearest-neighbor classifiers). Also, machine learning systems can be contrasted with rule-based expert systems in which rules are typically identified via subject-matter experts using knowledge-elicitation approaches, although rules can also be generated via algorithms.
1.2.4 Optimization
Optimization is a fundamental concept relevant to AI, ML, and other intelligent software systems. Its origins come from calculus, which can be used to identify minima and maxima on a multivariate function. Generalizing this, many intelligent software tools involve optimizing some cost or loss function: finding a set of input values (parameters) that lead to the best outcome. Within AI, ML, and automation, optimization can be defined as the systematic and iterative process of improving the accuracy of a machine learning model or algorithm's output through testing and refinement. It can also be used to describe the process by which a system/task/environment is iteratively improved through the usage of machine learning and/or AI algorithms (Surianarayanan et al., 2023).
Optimization is used in two ways in intelligent software tools. First, problems can be framed in terms of a set of input parameters or choices (states) within a space of options, and each possible state can be associated with a cost or benefit. This use of optimization attempts to find the optimal (or at least good) parameter state that minimizes cost or maximizes benefit, with cost and benefit defined very flexibly.
The basic framework of optimization in AI and automation is that the effectiveness of behavior produced by an intelligent system can be incorporated into a cost function relevant to the particular domain and algorithm. Cost may incorporate time, efficiency, monetary costs, or other heuristic values.
For example, solving supply chain, shipping, and warehousing problems involve finding a way to minimize both costs and delivery time by determining the best way to store and ship goods to customers. Here, cost might be related to shipping and storage costs, whereas the input parameters would be the way in which a product is shipped, when it is sent and delivered, and from which source it originates. Game-playing and route-finding AI agents also frequently work from an optimization framework. For example, an AI chess player searches future game states that might result from a particular move to identify moves that minimize potential losses or maximize potential benefits.
Another way optimization is used is to describe inference processes in machine learning–the way that parameters are determined in a way that minimizes cost, error or “loss” functions. Here, the goal is to find parameter settings that optimize a system for a given set of data, so that future behavior will be efficient or effective. For a simple linear regression model, optimization is typically done via least-squares, but for more complex models other techniques are used. This optimization is often thought of as 'learning', as it allows a model to incorporate information from many cases in order to make future decisions. For example, an image classifier built with a neural network undergoes training, and typically uses a backpropagation and a learning algorithm to set weights in the network that minimize error rates in classification. Here, rather than finding the optimal solution to a particular problem, the optimization attempts to find optimal settings that will solve future problems effectively. In all such cases, there is a risk of overfitting: identifying parameters that optimize the cost for the training set, to the detriment of cases outside the training set.
1.2.5 Humans in the loop: From Cybernetics to Cognitive Systems Engineering
These different high-level constructs all need to be understood in the context of the human work they are augmenting, replacing, displacing, or improving. The scientific study of this has occurred under many labels, often considered a specialization of Human Factors Engineering and Psychology. One of the earliest terms was Wiener's (1948) concept of cybernetics, or as his book's subtitle states “Control and Communication in the Animal and the Machine.” The subfield of Cognitive Systems Engineering (CSE) (Woods & Roth, 1988) emerged from Cybernetics to deal largely with how to evaluate and design automated systems. Here, the perspective is shifted from a human operating a system to a human being part of the system. Although human-computer interaction (HCI) originally focused on 'dumb' computer systems, it is increasingly involving interaction with intelligent systems. Other subdisciplines such as human-centered AI (Shneiderman, 2022) are also popular, as well as many other approaches we will cover in other chapters (HRI: Human-Robot Interaction; coactive design; mixed-initiative interaction, social robotics, just to name a few).
It is probably true that all automation shifts human work and creates new problems while it solves others. For example, Bainbridge (1983) described “ironies of automation”, including that more automation can leave humans with harder residual tasks (such as monitoring and exception handling) that are poorly supported, and that the times when human control is needed will require skill (manual or other) that are no longer as well developed. Endsley (2023) argued that Bainbridge's warnings are still relevant in the age of artificial intelligence. Others have identified automation-induced complacency and automation bias (Wickens et al., 2015), and shifting bottlenecks to different kinds of tasks such as supervision and control (Woods, 1996), as similar specific ways that automation does not simply replace human work; it transforms it—often adding monitoring, coordination, and recovery demands.
1.3 The purpose and function of intelligent systems
Now that we are familiar with some of the broader terms, an important next step is to understand the general functional classes of systems. In many cases, from a human factors perspective, knowing the purpose and function of a system is more important than understanding the underlying logic, because it can help identify the goals of the user. There are many grey areas and overlaps among systems, and this is not comprehensive, but a brief overview includes:
- Numerical prediction or estimation. Regression analysis is a classic case of this–given some things you know about an object, can you predict a numerical value? This is how Carvana can give you a price for your car (incorporating its model, age, mileage, condition, and features), but also useful as a building block for many other applications.
- Categorical classification and detection. Given information about a case, can you identify its class? Face recognition and other machine vision algorithms are a prototypical example, but the breadth and use of classification algorithms are vast, from automated driving (detecting other cars, locations, people) to marketing, business, target detection, cancer diagnosis, etc.
- Control and regulation. Perhaps the most widely used (and simplest) intelligent systems, these sense a state while controlling inputs and attempt to manage or regulate the system for some purpose. Includes thermostats, cruise control, factory automation, autopilots.
- Rank, retrieve, and recommend. Systems that attempt to produce information based on input. Including search engines, music and movie recommendation systems, dating services, shopping applications.
- Cluster, compress, and collate. Generally 'unsupervised' systems that attempt to build a model from data that extract important regularities and similarities, including PCA and cluster analysis. Used in many classification systems to build features, cluster analysis for marketing, self-organizing maps for summarizing data patterns.
- Decide/Plan/Act. Extensions of control systems that have a representation of a problem space (often a network) and can reason about costs and benefits of different moves within the problem space, they reason to achieve a goal that is either carried out or recommended. Applications include GPS systems, logistics/supply chain management, self-driving vehicles and video game AI.
- Constraint/schedule satisfaction. Systems that attempt to achieve optimal mappings of resources based on constraints. E.g., airline crew assignment, supply-chain management, med school student matching.
- Generate/Synthesize Content. Algorithms that create new information–from predictive text entry, to algorithmic music, to image, video, and text generation.
These basic categories of functions map onto overlapping sets of algorithms, each with different strengths and weaknesses. It is important to understand that AI and automation is not a single algorithm with a single function. More details are provided about each of these in the following subsections:
1.3.1 Regression and numerical prediction models
Regression models are often not even considered AI or automation as they are fairly simple and well-studied. However, they do form important components in many other more complex systems, and there are many extensions (e.g., generalized additive models, non-linear and non-parametric regression, auto-regression) that are quite sophisticated, and all such systems are incredibly useful in many contexts. Although they are often used to make inferences about data in experimental contexts, they are also useful as predictive models and are considered machine learning models, insofar as they are 'trained' with a data set and can then be used to predict new cases. Furthermore, models with a variety of architectures far outside standard regression are useful–consider models that predict the temperature or amount of rain expected over the next 10 days, the price of a house or vehicle, or the expected value of a farmer's crop in the future when it is harvested, or the expected re-entry point and speed of a spaceship.
These models typically make predictions based on past cases. New cases outside their experience will necessarily be extrapolations that are more likely to make errors. This illustrates how, for humans using the system, its competence envelope needs to be understood, so that the user can know when to trust it, and when to not trust it.
1.3.2 Classification and classifier systems
One of the simplest classification algorithms is just a transformed regression model known as logistic regression. Here, instead of modeling a numerical value directly, you model the probability that the outcome is one of two cases (often through negation; such as 'has disease' versus 'does not have disease'). Many different approaches can be used, but the ability to place a label on input so that different courses of action can be taken is very powerful. For example, classifiers are used to identify individuals in photos, to detect copyrighted content used on youtube, to identify defects in manufacturing processes, to perform content filtering of images, text, and speech on social media, to determine diagnoses based on symptoms or radiology, and many more.
Classification typically requires large datasets of known cases that can be used to train an algorithm to distinguish new cases. Once trained, classification systems can then be distributed and used for almost any purpose. But just as with numerical prediction models, it is important to understand from a user perspective that the further the new cases are from the training, the less reliable the system is likely to be, and so its competence envelope depends on the training set.
Furthermore, many classifiers pick up on correlated features in the training set that are not easily detected by humans but are only incidental to the category. For example, most published and labeled pictures of birds are already selected by the photographer to be easy to detect, and they normally have the bird in its natural habitat. Thus, a bird that can be found in winter feeding on specific frozen berries will typically be photographed in those conditions because they are difficult to locate or identify otherwise. Consequently, a classifier is likely to include snow and berries as important features of its model, be poor at identifying the bird outside that context, and mis-identify other birds in snow as that bird.
1.3.3 Control and Regulation
Controller systems are some of the oldest and simplest automations. They require only a sensor and some type of actuator to influence the output. Even pre-mechanization involves simple controls–you add fuel to a fire when it gets too cool; you rein in a horse when you detect it is going too fast. Mechanical controllers were critical in developing the first clocks (controlling how quickly a gear rotated), but are most common in simple thermostats that detect temperature and close a circuit, often based on the expansion of two different kinds of metal. Modern controllers are much more highly advanced, including antilock brakes, automated driving systems and robotics.
A famous result in cybernetics is the good regulator theorem of Conant & Ashby (1970). Their argument, in a paper titled “Every good regulator of a system must be a model of that system", states that an effective controller of a system must embed within it the critical elements of the system it controls. This suggests an important role for “digital twins” and environment models in modern control systems, to allow forecasting and predicting the result of changes prior to making them.
A reasonable corollary of this theorem is that if a system can be modeled in simple terms, it permits a simple controller that is likely to be easy to understand. The on-off mechanical thermostat used in heating and cooling systems has a simplistic but inaccurate model of the heater. The thermostat turns the heater on when the temperature it senses is too low, and off when it reaches the desired temperature. To a first approximation, this models the basic heating system (it is either too hot or too cold), and this level of fidelity is satisfactory in many cases. However, these systems do not create a stable temperature, but creates rise-and-fall hysteresis cycles as the temperature moves to the critical trigger point and the heater turns off, then falling to a point low enough to trigger the heater again (see Figure below)
The problem that emerges is that this simple model is not sufficient for many systems. One example is consumer-level espresso machines, which typically have a thermostat controlling the temperature of the water going through the coffee grounds. As the pump is activated, the hot water in the boiler is replaced by cooler water, and eventually the temperature falls below the optimal level, at which point the thermostat triggers when more heat is needed. But this is too late, because by the time the heater starts again, the water is too cool and the result is a sour cup of coffee. Some machines are sold with or can be modified to use other controllers that model the state of the water temperature more exactly. These can cost hundreds of dollars more, but will maintain the temperature better because they detect and respond to small changes in the temperature.
These kinds of controllers are called PID (Proportional-Integral-Derivative), and are useful devices that can be used to regulate a variety of processes in industrial settings and consumer products. The PID model of the system is more sophisticated than the on-off thermostatic model. As the name implies, these systems are tuned with three separate processes: proportional control regulates how close the environment is to the set condition, integral controls change the intensity of the system's function to reach the set condition and derivative controls adjust for over-accomodation of the integral controls (Johnson, 2005). Thus, PID controllers can be good regulators–they have a more accurate model of the system they are controlling, and can better forecast future states of the system and take earlier action.
Another example is the cruise control on a vehicle. Suppose that the cruise control was a thermostat-style controller. Whenever the vehicle was travelling too slow, it would need to activate the engine full-throttle until it reached the desired speed, Then it would deactivate, coasting until it went below the specified speed. Even with a band of indifference of several MPH or KPH, it would be uncomfortable and dangerous. Instead, it controls the throttle level to maintain a set speed, and if the vehicle speed differs from the set speed, can approach it smoothly.
1.3.4 Rank/Retrieve/Recommender Systems
A number of systems work to identify, rank, and retrieve relevant content. Although the underlying systems can differ greatly, they all have a similar goal, and are critical to sensemaking in information-rich environments.
Recommender systems. Recommender systems attempt to identify recommendations about content based typically on a prior model of the user. This model could be demographic data, past purchases or preferences, or other information collected about the user that can help provide future predictions of behavior (Vultureanu-Albişi & Bădică, 2021). It can be described as a filter application that hides or displays information based on user activity and demographics. An important limitation of this type of system is that it is heavily dependent on user input (Jannach et al., 2010). In theory, a recommender system can tell you books or film you might find interesting, but it will only work well if you have provided ratings of books and films you have read or watched, or answer surveys about what types of properties of films/books you like..
A common architecture for recommender systems as collaborative filters. Here, rather than trying to build a model of the user based on the kinds of films they like, you can identify other people who liked similar content, and make recommendations about what those people also liked. This uses the power of large user databases to make informed decisions, but they can also be very blunt tools–the reason you like a specific film may be because of the director or star; another person may like the genre or writer, and so recommendations are likely to be based on reasons that don't appeal to you.
Information Retrieval and Search. Traditionally, information search and retrieval worked by creating an index of content-to-documents. As a simple example, one might look at all the keywords used for articles in a scientific journal. Over years of publication, there may be hundreds or thousands of keywords, stored efficiently in a single database or document. Then, a search engine merely needs to identify documents that match combined keywords along with rules like OR and AND to identify relevant articles. Of course, keywords are notoriously limited, and so other indices can be made (authors, words in the title, abstracts, dates, etc.), and efficient databases can easily retrieve content.
Initial web search engines worked in similar ways. They would index all web content, and create index tables mapping words to documents. Of course, some words are irrelevant (of, the, and, when, etc.), but by representing each document as the words embedded in it (a bag-of-words model), it could be easily indexed and retrieved. However, early search-engine-optimization recognized this, and many web pages would embed the entire dictionary of words on their page to increase traffic or to sabotage search engines. The important goal turned out not to be identifying pages containing content, but finding the most relevant pages–ranking them by how important they were. One strategy that led to the success of Google (see the PageRank algorithm discussed later) was to essentially use collaborative filtering–content is important if other sites link to it. This creates its own SEO problems, and modern search engines use a complex set of algorithms to provide relevant content, but they all work by attempting to rank and retrieve relevant information.
1.3.5 Cluster, compress, and collate
A variety of 'unsupervised' learning algorithms attempt to build a useful model of a data set based on its own inter-relations. Unsupervised systems take unlabelled cases and apply algorithms to map these cases onto meaningful dimensions or categories. Thus, they are akin to regression and classification models–but without the answer key.
Clustering algorithms. Clustering involves the grouping of similar objects or data using a machine-learning algorithm (Halkidi et al., 2001). Many forms of clustering models exist, each looking at the data in a specific way. Generally, clustering is a form of pattern recognition, in which the machine-learning model looks for similarities between the data points and quickly sorts them into manageable clusters of data for the researcher (Rodriguez et al., 2019). Clustering is an efficient method of evaluating data for trends and outliers, making it very useful in areas such as data science, statistics, engineering, etc. (Ezugwu et al., 2022).
Dimension Reduction and scaling. Other systems involve scaling numerical feature data into fewer dimensions that often characterize important aspects of the data. Some examples of these include eigen decomposition, principal components analysis, singular value decomposition, factor analysis, multidimensional scaling, and other self-organizing maps.
Many intelligent algorithms are centered on vector-based representations of data. That is, cases are represented by a vector of numerical values. Any data-rich domain will fit naturally into this paradigm; for example a network of sensors will produce a series of samples from each sensor that represent the state of the system at a given point in time. Similarly, data about a person in a database may be represented by the information a company has–their age, education level, zip code, etc. Data about a word may be represented by the words it appears near in documents, or other properties. In many cases, the raw vector data is extremely highly dimensional, and better inferences can be made by reducing the dimensions of the data via certain matrix arithmetic operations.
Dimensionality reduction techniques capture essential information in data while getting rid of redundant or less informative features to extract meaningful data, facilitating simplified computational tasks. Two related approaches include Principal Components Analysis (PCA) and Singular Value Decomposition (SVD).
PCA and SVD are used for dimensionality reduction in multivariate data and are a widely-used technique in data analysis and machine learning (Baruah, 2023; Chengwang, 2010). Its objective is to transform high-dimensional data into a lower-dimensional representation, capturing the most important information. The distinction between the two approaches involves primarily whether the input data are a case-by-case matrix (such as a correlation, similarity, confusion, or co-occurrence matrix) which can be reduced via PCA, or a case by context matrix (i.e., a term by document matrix), in which case SVD is used. Depending on the data set and complexity and goals, dimensions may be reduced from thousands or millions (a document x word matrix) to hundreds or even just a few. Sometimes the major dimensions (components) are interpretable, but generally this won't be true unless specific rotations or additional analyses are performed.
1.3.6 Decide/Plan/Act
This describes a very broad set of systems that involve agent-based or 'agentic' systems, decision systems, planning systems and the like. Historically, the General Problem Solver (GPS; Ernst & Newell, 1967) was perhaps one of the first comprehensive general AI systems, developed as a way of moving within a problem space, making progress toward goals by creating subgoals and using operators to move within that space. However, game-based agents are often of this nature. Holland (1999) highlighted Samuel's (1959) checker-playing robot as an emergent system that learns to play checkers based on simple pattern recognition, and more advanced systems that play Chess and Go are more complex systems that take an action based on a well-defined game state, but these often implicitly or explicitly identify series of moves–a plan–to accomplish the goal. More specific planning systems exist to generate routes within a map, and these are now embedded in all our phones and most new vehicles. Similar systems help create low-observable flight plans for stealth aircraft, identify the cheapest way to deliver a product from a warehouse to your home, and (with modern LLMs), carry out tasks such debugging computer code or making automated purchases.
1.3.7 Constraint Satisfaction
When an AI searches through the set of possible solutions to a problem, sometimes the sequence of steps used to solve the problem are important, such as GPS routing. In other cases, only the solution state is important. These types of problems are referred to as Constraint Satisfaction Problems (CSPs). Generally, we think of these problems as having variables that take on values from a particular domain and are specified with constraints that specify allowable solutions (Jones, 2008). Finding the best seating chart for a Thanksgiving dinner to minimize the chance for family drama is a CSP. Each person (the variables) can be placed in a seat at the table (the domain) near some people and far from others (the constraints). If there are enough constraints, most random solutions will violate one or more constraints, and so the goal of a CSP algorithm is to find permissible outcomes that satisfy all constraints.
A brute-force method called Generate and Test can be used to solve CSPs. It begins by generating a solution, whether randomly or guided by heuristics, and then testing to see if the constraints are satisfied. More sophisticated algorithms include backtracking and forward checking. Backtracking starts with the most constrained variable, tries adding another variable, and possibly backtracking when a constraint is found to be violated. Forward checking preemptively eliminates invalid partial solutions to reduce the search space (and correspondingly, the time to a solution).
The essence of CSPs is that the problem involves specifying the constraints, and coming up with a solution that obeys the constraints or minimizes the violations.
1.3.8 Generative AI
Generative artificial intelligence is the broader term for machine learning models that receive input (such as a prompt) from a user and create an output based on the training data the model has access to (Brynjolfsson et al., 2023). Although these systems are gaining wide usage with modern systems like ChatGPT, Shannon (1951) was perhaps one of the first to propose them in terms of a way to generate 'fake' language text using what is called a Markov model. His algorithm was simple. Take a book, open it to a page, and pick a word, and write it down. Then find that word again somewhere else in the book, and write down the word that follows it. Continue until you get a word at the end of a sentence. For example, in this chapter, I used this method to generate the following 'sentence':
A VARIETY OF IDENTIFYING AND INTERPRETING QUERIES OR OUTPUT LAYER BY OPTIMIZING SOME PURPOSE.
Notice that this is not grammatically correct or semantically meaningful at the high level, but its themes are recognizable and relevant to the current chapter. We could easily create 'fake French' text that would fool someone who doesn't speak French. If we want a better approximation, we just need to use pairs of words or larger chains (what are called n-grams). Now, we would be cutting out small pieces of language from real text, and pasting them together only when they match pairs. Or we can do it at the level of a letter. Here are some example 'sentences' made by reading a dictionary of words, and using just the letter tri-grams to generate text.
THE THEM FROW ITS DIS HATINTE WERCHAT DON ANCE ACRIVE
THE NOTHATORS MONLY SMER ING RIT IMPUT BLE ING THOSED
BUSTRARSO GAR CAUSE SED WAY KED WITS DIS URTUS THE
CALLED SON INALLIED HITY WIDE AVE TOR OFFIS VALL HEACHE
THE ITHEREGS ANDAT THE ITOUBEFFS ANE BITS TAGE USIS MEND
MINGED WIT STRATIS HIC ING CON EDISTAIR THED ANDLYSIS ALP
GUITED SENTIN THE LEGINGIS WHE THE KARD ITTLY JUSTURE BETED
SIGE CERY LOSED ROMEEDS SED SULD THEY FUNITAT WHE AND
Notice that many words are real English words, and the non-words resemble English–and are likely to fool a non-English reader. They even often appear to include orthographic rules of English. Consider MINGED, ITOUBEFFS, JUSTURE, NOTHATERS.
This is a very simple example of a generative model. If we want to approximate text better and better, we just need to use higher-order N-grams. This is simple–training involves making a matrix the size of your token set (e.g., 27 with letters + space), and dimensions the same as your desired N, and fill it all with 0s. Then, read your text corpus, and for every N-gram, find that location in the matrix and increment its value. So, for a 3-gram, if you see the word BUTTER, you would increment matrix[2,21,20] (BUT), matrix[21,20,20] (UTT), matrix[20,20,5] (TTE), matrix[20,5,18] (TER), and matrix[5,18,27] (ER_). To generate text, we simply start by finding the distribution of letters that begin words (maybe _ _ ), which includes 27 values, and sample proportionately. If we get the letter C, we then find the row of the matrix representing our current context (_C ) and sample again., which might generate an L (_CL), then an A (CLA), then an R (LAR). We can repeat this until some stopping point.
But there are several problems. First, the storage required to store N-grams goes up as a power of the number of steps. For text, 1-grams take 27 numbers, 2-grams take 27^2, 3-grams take 27^3, and n-grams take 27^N. This is very costly and inefficient, and quickly scales beyond what is possible on computers. Of course, even if you had all text ever generated, most n-grams would eventually never occur, so your matrix is mostly empty. You might be able to deal with this by clever data structures that only maintain values greater than 1 (a so-called 'sparse' matrix), but even if then, with an N of say 20, you would essentially be encoding most sentence strings uniquely. When you generated text you would really just be extracting text from your data sources directly.
Modern generative algorithms using recurrent neural networks solve the first problem, specifically by using recurrent networks that feed back into themselves so that they learn the sequence of tokens that represent language, and the second problem because their representations are inherently noisy and their learning process has the effect of generalizing. In fact, although 20-deep matrices are probably impossible with our markov approach, modern algorithms now can use at least a million tokens as the context to prompt subsequent tokens. The context includes the recent history, specific instructions, and the 'prompt' given by the user. Conceptually, this is the same as how the CLA prompt generated R. This is essentially how modern LLMs work, although there are many other details that are needed to give them the success they have.
The output of a generative AI model can vary but includes text, images, video, audio, lines of code, etc. (García-Peñalvo & Vázquez-Ingelmo, 2023). Generative AI has become a large topic of discussion in media recently, as generative AI models using Large Language Models (LLMs) such as ChatGPT (text generator), Copilot, Cursor, or Claude Code (computer code generator), and Midjourney (art generator) have risen in popularity (AbuMusab, 2023). Many find the use of generative AI models to be beneficial to users, as it may make work more efficient, but many question if the implementation of these models will impact the need for human jobs in certain fields, such as writing and media (Zewe, 2023). Both in the workforce and in education, the use of a good prompt can create a high-quality output that resembles human work, leading to many debates of whether this technology should be embedded into everyday work (AbuMusab, 2023).
1.3.9 Summary and Takehome for Human Factors Researchers
We identified 8 general classes of intelligent systems which cover many (but perhaps not all) systems that provide intelligent support for users. From the perspective of a human-centered researcher, basic literacy about the distinctions among these is important, because you may be working with researchers or developers who have solutions from one perspective or another. Understanding these also presents possibilities–if you are involved in designing a system, you may not recognize how a classifier or recommender may improve usability until you understand the roles of these systems and see examples of how they get used. It is also important to recognize that the underlying algorithms are not mutually- exclusive; modern LLM systems can achieve many of the functions listed here if prompted properly; planning systems may also use constraint-satisfaction; the same model that classifies can be re-framed as a generative model. However, it is also important to understand that some of these systems are very capable and efficient, and may provide a better solution than another more complex one. Certainly, one could embed a chatbot in a thermostat, creating an observe-and-decide cycle where it determines when to turn on a furnace, but this would be costly and may be error prone in comparison to mechanical thermostat.
1.4 Architecture: Representations and Algorithms
Many different approaches have been used to implement AI and automated software. One consideration is the underlying way in which knowledge is represented. This ultimately can impact the transparency, efficiency, and traceability of the system, and developers familiar with one kind of representation can often predict weaknesses and problems that will arise from using that architecture. Closely linked, but distinct, are the algorithms (computational procedures) used to either build the system or achieve a solution. These representations and algorithms sometimes are closely linked (back-propagation and activation in neural networks distributed representation), but sometimes interchange (a constraint-satisfaction problem solution could be achieved via several algorithms, and be represented in several distinct ways). Together, the representations and algorithms used for a particular problem can be considered the architecture of the intelligent system.
1.4.1 Common representations
Most representations used in intelligent systems come down to a few simple concepts, which map onto common data structures in computer science. These include:
- Features/vectors. These can include raw data from sensors, audio, video, etc, which closely map onto input and/or output data.
- Networks. A flexible representation that posits units connected by relationships; can be low-level distributed representations in artificial neural networks, high-level constraint networks include roadway systems, etc.
- Trees. Constrained network structure allowing only branching; algorithms can efficiently reason, search, and sort. Enables efficient O(log(N)) decision making via classification and regression.
- Symbolic type/tokens. Categorical identifiers allowing reasoning about classes of objects, including language production systems, classic expert systems, and the like. Often linked to rule-based systems.
- Rules. Often coupled with symbolic systems, but also part of decision trees; may be coupled with functions/transformations to provide stronger inference capabilities.
- Case/exemplar systems. From KNN classifiers to Case-based reasoning (CBR) systems, rather than inferring a model summarizing data, use an efficient data store to compare new cases to all existing cases.
- Probabilistic/Markov/Bayesian or Fuzzy logic. Often coupled with networks or trees, represent nodes in terms of probability distributions or partial membership in fuzzy sets..
In general, most intelligent systems use one or more of these to encode and represent the problem, data, or model it is using. Many use multiple–a vocabulary may be encoded at one level as a symbolic dictionary, but there may be underlying feature or network representations to permit similarity-based matching via an ontology.
1.4.2 Common algorithms
Just as there are common representations, there are common algorithms that should be understood, at least at a high level. Algorithms serve many different purpose, from learning to inference. A summary of many common algorithms follows, with more detail in subsequent subsections.
- Search. Search is the sine qua non intelligent systems. A solution can almost always be framed as a goal state, and the system aims to search among possible solutions for one that is optimal.
- Search: Dijkstra's Algorithm and A*. Extensively used within any well-formed problem space (including planar graph networks and road networks); used to find optimal plans, routes, and actions.
- Search: Forward/backward chaining. Used to plan a series of operations that produce a result or entail a logical requirement.
- Search: backtracking, recursive search, breadth/depth-first search. Within well-formed spaces such as trees, different strategies for efficiently finding the best or acceptable solutions.
- Inference: Case-based reasoning. Reasoning from a database of instances that represent past experience, rather than a summarized model representing typical/average situations.
- Learning: Maximum Likelihood approaches. Numerical optimization schemes that find the most likely set of parameters in a model to characterize the existing data.
- Learning/Classification: Bayesian/Monte carlo Markov-chain (MCMC) approaches. These use Bayes rule on probabilistic representations (including networks and trees) to incorporate evidence in order to identify probabilistic conclusions; results in inference about a distribution of possible parameter combinations that might produce observed data.
- Inference/classification: Fuzzy logic. ways of using fuzzy-set theory and inference to characterize uncertain knowledge and determine possibility-based set membership or classification.
- Learning: Gradient Descent/Hill-climbing. A general concept that most learning algorithms can be considered to implement; make changes in representation that reduce error or cost based on a local estimate of the slope of the error/cost landscape.
- Learning: Simulated Annealing. Addition of noise to a cost function that reduces over time in order to prevent algorithm from settling in local minima.
- Learning: Reinforcement learning. General approaches for updating expected reward for actions based on training in a model of the environment. Includes TD-learning, Q-learning, linear operator models.
- Learning: Back-propagation. Important algorithm for learning within multi-layer artificial neural networks that assigned error backward through the network.
- Optimization: Linear and dynamic programming. Provably-optimal solutions for certain kinds of well-behaved cost optimization problems, can be very efficient and the only feasible solution for many complex applications.
- Optimization: Genetic algorithms. Uses analog of evolution to explore complex feature space; iterative generations sample and exchange features which are then evaluated for fitness, to achieve overall improved cases. Useful for complex problem spaces with cost functions that are difficult to characterize.
- Unsupervised learning: Finite clustering approaches (KNN, finite mixture modeling) solves for a fixed number of clusters given similarity or feature-based representations.
- Unsupervised learning: Eigen Decomposition/SVD. Matrix algebra solutions that project data onto a smaller number of useful dimensions that can reconstruct observed data with differing levels of fidelity. Requires numerical solutions to invert a matrix, which can be efficient even on large data sets.
1.4.3 Specific Architectures in AI and intelligent systems
Together, these representations and algorithms are used across many different intelligent systems. Most complex AI systems likely use hybrid approaches combining many of the different approaches. These lists are shown to provide basic definitions and provide familiarity to a human factors researcher, but we cannot go into detail about all such representations, algorithms, and resulting architectures. However, we will review some of the more common architectures that bear understanding, as they can impact how the system is trained, its competence envelope, misconceptions users may have, and its overall costs/benefits.
1.4.3.1 Search algorithms in AI
Search algorithms have traditionally been a primary algorithm forming the cornerstone of artificial intelligence research. In general, search refers to framing a problem in terms of a state space or problem space, identifying the current state and the goal state, then using systematic algorithms and heuristics to find a path through the space (Bornstein, 2019). In addition, fitness and cost functions for either states or paths through the space are often calculated, so as to identify the optimal way to solve the problem. Several common search approaches are described in Section 4. Many AI problems can be framed as a search problem. Route planning searches through a network for an efficient route; recommender systems search through a set of products to find an option that matches ones preferences; chess AI systems search through possible futures (including likely moves of an opponent) to determine the best next move. In AI solving chess and other similar games, searching is crucial for exploring the vast number of possible moves and counter-moves in a game. Search algorithms may lack transparency, leading to uncertainty and trust issues, because solutions reflect the fitness functions used, the way a problem space is framed, and the local optima the algorithm will often produce.
1.4.3.2 Case-Based Reasoning
Case-based reasoning systems rely on a a knowledge base of previous situations and experiences to make choices or decisions based on past experience This algorithmic design is based on the concepts of scripts, memory organization packets, previous situations, and situation patterns (Schank, 1983; Schank & Abelson, 1977). Cases are contextualized experiences that include the problem, the solution, and the outcome (Watson & Marir, 1994). There are four main steps in case-based reasoning: retrieve, reuse, revise, and retain. The new problem is compared to previous similar cases and a solution is suggested based on the previously used solutions. Unless the problems are extremely similar, the previous case's solution is revised and the new solution and problem is retained for future use. Knowledge-based reasoning often fails due to the knowledge elicitation bottleneck (Hayes-Roth et al., 1983) which case-based reasoning does not experience. Case-based reasoning is particularly useful in applications where no explicit model exists and where new cases can be acquired to improve performance.
1.4.3.3 Artificial Neural Networks for inference/classification
Artificial neural networks are computational models that mirror behavior of the human brain, usually with some levels of abstraction. The human brain is made up of neurons that send signals to each other through the use of synapses and neurochemicals. An artificial neural network functions through the connection of different nodes where each node can accomplish a simple integration of information that is transformed into a simple output (Arbib, 2002). The guiding principle of artificial neural networks is that simple units and learning rules together can combine to form sophisticated and complex behavior. These networks are generally organized into multiple layers, and minimally include an input layer (taking data provided from the outside world); an output layer ( the result of the processing) and normally one or more hidden layers ( existing between the input and output) (Balcázar et al., 1997; Setiono & Leow, 2000). The connection between the nodes is known as the weight, with higher weights meaning that two nodes are more strongly connected. Individual units typically involve non-linear mappings (Vargas et al., 2017), which enable complex information integration and rule-like behaviors to emerge.
Although traditional neural networks were limited to just a few layers, modern approaches, called deep neural nets or deep learning (LeCun et al., 2015), have emerged that use many layers (e.g., 15-20), and different layers incorporate distinct processing architectures (e.g., recurrent, convolutional, etc.) The order of the layers is often defined by the programmer based on heuristics and trial-and-error. Neural networks are useful for their ability to self -learn and to find optimal solutions quickly (Wu & Feng, 2018).
1.4.3.4 Backpropagation algorithm for learning in Artificial Neural Networks
Early versions of single and multi-layer neural networks were proposed (McCulloch & Pitts, 1943; Rosenblatt, 1958) as demonstrations of the power of simple classifiers. Their capability was recognized to be limited in terms of their ability to learn complex representations such as XOR feature interactions (Minsky & Papert, 1969) unless multi-layer networks were used, but rules to properly learn weights in multi-layer networks were considered ineffective. Backpropagation was the first practical solution to this problem. Back-propagation is a scheme which starts at the final layer of a network, in which a prediction is made and compared to the ground truth. The difference between these is the prediction error, and the problem is determining which of a large number of nodes and weights are to blame for that error. In a simple perceptron node, this operates like a regression equation, and standard numerical methods can be used to choose the optimal weights. But when multiple input nodes across multiple layers make up the network, the optimization is difficult. Back-propagation works by conceiving of the error as a 'gradient descent' problem. It iteratively determines the effect each node has on the accuracy of the final layer prediction, and update weights in reverse across layers in order to reduce the prediction error. For example, if one node is off in the final layer by \(+.5\), it might be adjusted by \(-.5 * .01\) (.01 is the learning rate) to move it in the direction of the correct answer. But then, since that end node is in error, we work backward and assume the previous layer's nodes were also in error in proportion to the final layer error. Based on this error, the weights can be adjusted backward throughout the network, and with enough computational power, training examples, and a mixture of layer structures, it produces complex classification behavior. Although initial approaches to backpropagation appeared in the 1970s, (Rumelhart et al., 1986) is often credited with popularizing the approach, which was further refined by LeCun (LeCun, 1987; LeCun et al., 1989).
Although these approaches fell out of favor by the late 1990s as computational intelligence models, graphics chips that developed for mostly for gaming in the early 2000s worked by optimizing massive parallel computations on data. As algorithms and libraries were developed to optimize neural networks to perform backpropagation and computation on these computational architectures, and data sets emerged supported by internet users (including Amazon Mechanical Turk) it permitted exploring 'deep' networks with 10-20 layers of different computational and connectivity patterns, learning from large labeled data sets, and such models re-emerged by 2020 as the dominant AI architecture for image, video, and language processing. Hinton's 2024 Nobel prize was in part awarded for developing back-propagation (Rumelhart et al., 1986).
1.4.3.5 Symbolic AI Systems with forward and backward chaining
Symbols refer to tokens or identifiers that represent some concept directly. Human written language is a symbol system as groups of letters represent words which identify linguistic concepts. Early implementations of AI relied on symbolic representations in computers, typically using the LISP programming language. The benefit of this approach is that it was capable of producing complex behavior while also being relatively computationally efficient. Although symbol-based representations can be used in many ways, one typical way they are used is to represent procedures via if-then rules. For example, the knowledge embedded within a system might be represented by a series of symbolic statements:
(CAT ISA MAMMAL) (DOG ISA MAMMAL) (CANARY ISA BIRD)
The system may contain rules that match symbolic statements of knowledge to produce some outcome:
IF: ((?A ISA ?X) AND (?B ISA ?X)) THEN: (SAY ?AS AND ?BS ARE BOTH ?XS)
Then, intelligent behavior is a process of pattern matching. Here, symbolic patterns match rules that lead to actions, such as the system saying:
DOGS AND CATS ARE BOTH MAMMALS
By encoding knowledge, rules, and enabling other temporary symbols such as goals, complex behavior can be implemented. This type of inference is called forward-chaining, and matches logical inference rules like modus ponens that provide deductive proof for syllogistic logic.
Symbolic AI systems were dominant for many years but increasingly replaced by other subsymbolic representations, as compute power and big data availability increased sufficiently to support alternative approaches. However, they are still used as basic infrastructure in many applications such as game-based AI systems and case-based reasoning systems. They are also increasingly likely to be integrated with other non-symbolic representations (vector space, neural network, and probabilistic representations) to take advantage of unique advantages of different kinds of systems.
In comparison, some symbolic systems use backward-chaining, which is akin to abductive reasoning. Here, we apply the logic of modus ponens backward to identify necessary facts and rules that lead to a goal state. For example, if Diana were on a dating app, and wanted a male partner that was at least 6 feet tall with an income of at least $100,000, we might have a rule representing this goal:
Rule DATEABLE IF: (?A INCOME > 100000) AND (?A HEIGHT >72) THEN: (SAY ?A IS A DATEABLE)
A backward-chaining system would work backward to find cases that match logically. Of course, this is simple if we have a database of facts about all the available men, but we may not–we might not have income directly, but have a job title and information about incomes for different job titles
(NAVAL_CAPTAIN INCOME_IS 150000) (NAVAL_SURGEON INCOME_IS 90000) We may also have rules to help identify income of individuals Rule INCOME_PROFESSION IF: (?PROFESSION INCOME_IS ?X) (?NAME PROFESSION_IS ?PROFESSION) THEN: (SAY ?NAME INCOME_IS ?X) Now, with information about individuals, we can backtrack through rules to find conditions that satisfy the preferences: (JACK IS_A NAVAL_CAPTAIN) (STEPHEN IS_A NAVAL_SURGEON) (JACK HEIGHT>72) (STEPHEN HEIGHT > 72)
A backtracking system would iteratively search backward through the rules to identify cases that satisfy the goal, first identifying that in order to satisfy the DATEABLE rule, you need to examine the INCOME_PROFESSION rule. When that rule succeeds filling in the income of an individual through inference, it then can select JACK as the only bachelor who satisfies her constraints.
1.4.3.6 Bayesian inference and networks
Probability is a field of study that allows for the handling of uncertainty. Using probability theory, decisions can be made with only partial evidence available, which is critical for rational agents who need to make the best decision possible in their current circumstances (Russell & Norvig, 2021). One application of probability theory is the ability to discuss potential relationships between events, such as “what is the probability I have a cavity given that I have a toothache?” There is a potential relationship between having a toothache and having a cavity, and probability can be used to determine the likelihood of this relationship.
However, not all events have a relationship. For example, one could ask “what is the probability I have a cavity given that it is raining outside?” There is little to no chance of a relationship between having a dental issue and the weather in a specific location. These two events are most likely unrelated, and thus we can say they are independent (Russell & Norvig, 2021).
For events that are related, probabilities can be calculated using the information already known about the event. For example, given that the individual already has a toothache, what is the probability that the toothache is caused by a cavity compared to other potential causes? This probability would be conditional, as it is narrowed in scope to information already known and unrelated pieces of information can be discarded (Russell & Norvig, 2021).
Bayesian networks data structures that can concisely represent these conditional relationships between events, and thus a hypothetical process by which observed data are generated from a state of the world. They typically take the form of a directed graph with no cycles. Nodes in the graph will have directed paths, with the node being pointed to being a child of the node the path originates from. This directed path implies conditional probability, so the child node is conditioned on the parent and two children from the same parent node are conditionally independent. By having this representation of conditional probability, conditional probability tables can easily be generated based on the derived relationships, and interdependencies, from the network (Russell & Norvig, 2021). This network represents a probabilistic generative model of how data are created. For example, one might have data from a set of sensors intended to detect trespassers; and this might include sound, heat sensors, motion detectors, and the like. A generative model in the form of a Bayes network would represent the probabilities that each kind of sensor state might be generated by a number of different situations: the wind, a human, a mouse, a deer, etc. Finally, Bayes rule and Bayesian inference is used to compute the posterior probability of each situation (wind, human, mouse deer) given the observed data, combining both the model of how the evidence was generated optimally with information about the base rate of each case (e.g., wind occurs often, a human may be rare).
1.4.3.7 Linear programming (simplex method) and nonlinear programming
Linear programming refers to a mathematical optimization technique to help to find the optimal solution to a cost function under a set of constraints described as linear equations. In this way, linear programming helps to find the optimal solution from a set of parameters that have a linear relationship. The simplex method (Dantzig, 1982) is a widely-used algorithm for solving a linear program. It narrows down potential solutions to a simplex: the vertices of a high-dimensional convex polygon, and provides an efficient means of searching through potential optima and eliminating possible configurations that are not optimal. Linear programming is limited to linear constraints, and so other related mathematical programming techniques such as nonlinear programming are used to find the optimal solution from a set of parameters that have a nonlinear relationship (Lithmee, 2019).
1.4.3.8 Q-learning and Reinforcement Learning
Reinforcement learning (RL) represents a cross-disciplinary field at the intersection of machine learning and optimal control, addressing how intelligent agents should make decisions in dynamic environments to maximize cumulative rewards (Kaelbling et al., 1996; Sutton & Barto, 2018). Among RL algorithms, Q-learning stands out as a model-free approach for learning the value of actions in specific states, devoid of an explicit model of the environment ("model-free"). Its adaptability to problems with stochastic transitions and rewards makes it a powerful tool (Li, 2023). The versatility of Q-learning and RL extends across various disciplines, including game theory, control theory, operations research, information theory, simulation-based optimization, multi-agent systems, swarm intelligence, and statistics.
RL constitutes one of the three fundamental machine learning paradigms, alongside supervised and unsupervised learning. Unlike its counterparts, RL operates in a dynamic environment, learning from collected experiences instead of relying on a static dataset. During training, data points or experiences are accumulated through trial-and-error interactions between the environment and an agent. This distinctive feature is crucial, eliminating the need for data collection, preprocessing, and labeling before training—requirements inherent in supervised and unsupervised learning. In practical terms, this signifies that, with the right incentives, an RL model can autonomously initiate the learning of a behavior without human supervision.
While Reinforcement Learning has emerged as a pivotal aspect in the field of Machine Learning and has garnered significant attention from researchers, it encounters specific challenges outlined below: a.) Extensive Datasets: Due to the complexity of Reinforcement Learning Models, they necessitate vast datasets to enhance decision-making capabilities. b.) Dependency on Environment: Reinforcement Learning Models learn through interactions with the environment, so the agent adapts based on the current state of the environment. In scenarios where the environment is dynamic, the training of the agent becomes challenging. c.) Reward Structure Design: In real-world applications of Reinforcement Learning, there is a need to carefully analyze the problem statement and formulate an appropriate structure for when the model should be rewarded or penalized.
1.4.3.9 Dijkstra's Algorithm and A*
When a problem can be represented by a graph (for example, a set of towns connected by roads), one of the fundamental problems to solve is the minimum cost route distance between any two nodes in the graph. Among other solutions, Edsger Dijkstra's eponymous algorithm (Dijkstra, 1959) is one of the most general, and finds the solution in an iterative fashion. Starting from an initial location, set the initial distance to all nodes as “infinite.” This is not to say that the cost to visit that node is truly infinite, only that the final cost is yet unknown. From the initial location, we examine the cost, g(n), to go to every other directly connected location. The minimum value is selected from this set of costs, and that neighboring location is marked as “visited.” The same process is repeated from the newly visited location until either a specific destination location has been visited or the shortest distance between the initial location and all possible destinations has been calculated. This is considered an uninformed search, as every decision is made based on actual costs without an attempt to estimate how the current decision will affect the global view of the solution.
To use this as a simple GPS mapping algorithm, costs can be based on static factors like the distance between two cities, dynamic factors like real-time traffic information, or custom factors like how ecologically friendly the route is. A route utilizing surface roads may be shorter than the interstate highway, but a driver may choose the longer high-speed roads for a shorter trip. Conversely, a pedestrian may want the shortest distance and avoid heavy vehicular traffic.
Traditionally, Dijkstra's algorithm finds the shortest paths from the initial location to all possible destinations. The A* (pronounced “A star”) algorithm (Hart et al., 1968; Jones, 2008) is an extension of the best-first search algorithm exemplified by Dijkstra. The two main differences are that A* only calculates the shortest path between a specific pair of locations and the cost between locations is expanded to include a heuristic cost. In the case of mapping software or GPS devices, the value of the former difference should be apparent. We only want to find a route from our current location to our destination, not a path to every possible destination. The heuristic cost, h(n), allows the algorithm to estimate how a particular decision will affect the overall progress toward the goal. Consider a trip that is generally northward. The nearest neighbor may have the lowest g(n) cost, but the heuristic cost h(n) will be higher if it takes us south (and ultimately farther from our intended destination). Consequently, the algorithm will be discouraged from further exploration of that path.
While both A* and Dijkstra's algorithm make intuitive sense for pathfinding applications, any problem that can be represented as a graph can be solved optimally and correctly through either algorithm. For example, all of the possible states of a Rubik's cube or the Tower of Hanoi problem can be represented as a network, and legal moves represented as connections in that network. A* or Dijkstra's algorithm can be used to find the smallest sequence of moves to solve the problem.
1.4.3.10 Evolutionary and genetic algorithms
Evolutionary algorithms are a set of search algorithms that takes inspiration from, as implied by the name, biology and natural selection. In these algorithms, an initial population of states is generated and a fitness value is generated for each state in the population. The most fit states are then selected and recombined to produce the next generation of states. The process then repeats with this next generation of states (Russell & Norvig, 2021).
One power of evolutionary algorithms lies in how they perform the recombination process. Genetic algorithms (Holland, 1992) are one type of evolutionary algorithm and take their inspiration from human genetics. As such, the states chosen for recombination undergo a crossover process, where two states have portions of their data combined to make a new state. For example, if the 8-queens problem were being searched and the state consisted of an array of length columns where each item in the array was the row number where the queen is positioned in that column, then the first two columns might come from the first parent and the remaining six columns would come from the second parent to make a new child state based on these two parents (Russell & Norvig, 2021).
1.4.4 Summary
Together, a system's architecture involves both representation and algorithms. These are often linked closely together, but sometimes there are multiple alternatives. For example, a simple neural network might typically be trained maximum-likelihood estimation or least-squares fitting instead of backpropagation, but the resulting network might be identical. Moreover, the forward processing is distinct from the learning. When these systems are exposed to users, many of the details of the algorithms are unimportant, but there may be consequences that rear their head. For example, a rule-based system is likely to produce a deterministic result–if it gives different results with the same input, it is broken and a signal to the user that something cannot be trusted. On the other hand, a neural network is often probabilistic and for some situations will give slightly different results every time it is used. In this case, a variety of responses is a signal it is working, and if it were to give the same answer every time, that would be a signal to distrust the result.
There are many other reasons to understand aspects of the underlying algorithm. Some approaches might be infeasible because training data would be impossible to produce; others may be infeasible because the storage, computational, communication, or time requirements do not fit the system or the use case. Understanding the basic approaches can provide alternatives that may satisfy more practical needs of system development.
1.5 Looking under the hood at applications of intelligent systems
Now that we have understood the purpose, the representations, and the algorithms of intelligent machines, we can examine various application domains and identify how different applications may be constructed.
1.5.1 Path-planning for mobile robots
At present, there is an increasing application of mobile robots in a wide range of industries. For example, robotic arms have been adopted to assemble steel structures for off-site manufacturing in the construction industry (Liang et al., 2017). To control the mobile robots and run the systems smoothly, it is important to plan accurate paths through gathering and sensing information about the outside world. Inspired by the behaviors that ants discover food and birds migrate among continents, scientists have developed different algorithms to find the optimal path for robotic movement (Huang, 2023).
Previous path-planning algorithms can be categorized into four types (Zhao et al., 2020), which are respectively fuzzy logic, SLAM, reinforcement learning and intelligent optimization. The idea of fuzzy logic algorithms is based on building a database of behavioral rules that can be applied to a variety of situations. With the external feedback data from sensors taken into account, SLAM can deduce the motion model and observation model, and the robot can estimate the real-time motion state (Huang, 2023). Reinforcement learning enables robots to choose the step with maximum reward at each node in a map, so the optimal path can be found eventually. Through many iterations, intelligent optimization algorithms, such as the ant colony algorithm, endow robots with excellent optimization capabilities.
With the increasing research and continued focus on path-planning for mobile robotics, A*, Dijkstra, rapidly-exploring random tree (RRT) and other algorithms have appeared in recent years (Huang, 2023). However, the common problem of these algorithms is how to get a shorter and smoother path and accelerate the convergence speed. In addition to improving the algorithm itself, combining different algorithms can lead to better path solutions, which is considered an effective solution in the future.
1.5.2 Gesture recognition
Gesture recognition is the process of identifying and interpreting meaningful expressions of motion exhibited by a human, encompassing movements of the hands, arms, face, head, and/or body (Mitra & Acharya, 2007). This interdisciplinary field combines elements of computer vision, machine learning, and signal processing to enable machines to comprehend and respond to human gestures, fostering natural and intuitive human-computer interaction. Therefore, gesture recognition has a large range of applications in various industries, such as virtual worlds, robotics, intelligent surveillance, and sign language translation (Chang et al., 2023).
Liu & Wang (2018) introduced a comprehensive model tailored for gesture recognition in human-robot collaboration. This model includes five crucial steps: sensor data collection, gesture identification, gesture tracking, gesture classification and gesture mapping. To start a gesture recognition process, raw gesture data needs to be collected by image based (e.g., marker, single camera, stereo camera, and depth sensor) and non-image based (e.g., glove, band, non-wearable) approaches (Liu & Wang, 2018). Second, gesture recognition refers to locating gestures from static raw data captured by sensors. Next, the located gesture is then tracked during the gesture movement. In the fourth step, tracked gesture movement is classified according to pre-defined gesture types based on classification algorithms (e.g., Hidden Markov Model, Support Vector Machine, Deep Learning). Finally, gesture mapping translates recognition results into machine or human understandable feedback.
Gesture recognition technology faces challenges that researchers are actively addressing. One obstacle involves the accurate detection and recognition of human actions, hindered by factors like the variability in human body parts and the surrounding environmental conditions. Another constraint arises from the requirement for close operation and critical point sensitivity in gesture control systems utilizing sensor boards. Moreover, vision-centered human activity analysis through computer vision encounters limitations in detecting actions behind walls or in dimly lit places. Additionally, the intricate degrees of freedom and variability in hand gestures present a recognition challenge for computer vision-based systems. These limitations underscore the imperative for ongoing research and development in the field of gesture recognition technology.
1.5.3 LSA to model linguistic data
One application of SVD is Latent Semantic Analysis (LSA),a natural language processing technique that involves generating a set of concepts associated with documents and terms to analyze the relationships between a set of documents and the terms they contain. The purpose is to reveal the latent semantic structure within a collection of documents, often accomplished by adopting Singular Value Decomposition (SVD) to reduce dimension (Landauer & Dumais, 1997).
LSA uses SVD to decompose a word (term) by document matrix to extract a term by dimension matrix, generally using a few hundred latent dimensions to describe all words. PCA is used to decompose a term by term or variable by variable matrix into a similar variable by latent dimension matrix. The resulting simplified representation has been demonstrated to capture many aspects of semantic information: words that are considered similar by human judges are generally nearby in the resulting high-dimensional space. Because of this, content and sentiment analysis can be performed, generally by applying regression or classification algorithms against a spatial representation for a word, sentence, or document.
1.5.4 Information retrieval and the PageRank algorithm
Google PageRank is a patented algorithm used by Google Search to rank web pages in its search results. It was developed by Google's Larry Page and Sergey Brin, and named after Page (Wikipedia, 2024). The basic idea behind PageRank is that the importance of a webpage is determined by the number and quality of links pointing to it.. This concept extends beyond search engines, being employed in academic citation systems and some social networks. While PageRank provides valuable insights, understanding its limitations is essential. Vulnerabilities to manipulation, a bias towards established sites, and static analysis may impact the accuracy of rankings. Acknowledging these limitations is crucial for users to interpret search results critically and for developers designing systems relying on PageRank principles. Awareness of these challenges helps maintain trust by recognizing that rankings may not always reflect the most relevant or dynamic information.
1.5.5 Predictive text entry
Personal computing devices of all sorts now offer predictive text entry. As the user types, suggestions to complete a partially-typed word or phrase are presented and potentially accepted. By looking at a corpus of any language, we can calculate the frequency of individual letters, digrams (two-letter combinations), trigrams (three-letter combinations), or generic N-grams. Speakers of any language have a reasonably strong intuitive understanding of these probabilities (Shannon, 1951). Whether the approach is to maximize the likelihood or to minimize the entropy (the level of surprise or disorganization), computing systems can be programmed with similar predictive capabilities. Indeed, the original iPhone autocorrect was manually programmed with N-grams to overcome the limitations of typing on the small screen (Kocienda, 2018). A larger corpus of words should lead to more accurate predictions. On-going updates to the corpus specific to a user will lead to more accurate predictions. However, the nuances of human communication, the context-specific nature of certain forms of communication (e.g., a work setting versus a casual conversation), and the ongoing changes to a living language contribute to inaccurate predictions.
1.5.6 Grammarly tone/sentiment analysis
Sentiment analysis is a form of classification problem that seeks to classify a piece of text as carrying an overall tone or sentiment (Medhat et al., 2014). There are many ways that artificial intelligence can accomplish this task, from more statistical methods that rely on tools and methods such as pre-classified lexicons, n-gram models, and naive Bayes classifiers to machine learning approached with recurrent neural networks and large language models (Medhat et al., 2014; Russell & Norvig, 2021).
Grammarly's tone indicator is one example of sentiment analysis application. In a post on the company's blog (Calonia, 2020), tone is defined as “...the author's attitude about a subject or topic to their reader.” The author then goes on to provide examples of different tones, noting that choosing the correct tone can help ensure that one is making grammar, stylistic, and vocabulary choices that align with what one is trying to communicate.
To assist with this, Grammarly released a tone detector in their system in 2019. Introducing the tone detector in another blog post (Grammarly, 2019), Grammarly described their detector as a mixture of both a rules-based and machine-learning approach that analyzes vocabulary selection, phrasing, punctuation, and style choices in things like capitalization. They end the post with the following rationale for creating their tone detector:
It's true that you can't control how someone reacts to your writing, but by making thoughtful decisions about the way you deliver your message, you can increase the likelihood that your recipient will focus not on unintended effects of your tone but on the information you're sharing.
1.5.7 SIRI and other assistants
Siri and other voice assistants all exist to aid the user in some manner. In general, these devices provide information, play specific media requests, or complete certain actions or tasks (Hoy, 2018). These devices listen for specific keywords, often their own name, to understand that they are about to listen for a prompt or request. In the case of Siri, a deep neural network is used to convert acoustic values of a person's voice into a probability distribution of potential requests. Once this is done, a confidence score is created. If the confidence score meets a certain threshold, then the device responds to you. Since voice assistants are connected to the internet, they can respond to a larger number of prompts. In addition, as the amount of text on the internet has increased, natural language processing has helped to create better responses. Natural language processing also allows for a number of different prompts to get the same results, as opposed to previous technologies where a specific phrase had to be used.
Voice assistants have some limitations of security and privacy. In general, anyone who can access a voice assistant can request personal information about a user from it. Researchers have shown that voice assistants will respond to commands given at ultrasonic frequencies that would be inaudible to humans (Zhang et al., 2017). By nature of the device, voice assistants are listening at all times. Tech companies like Google and Apple claim they are not collecting data at all times, but Google's voice assistant has been known to record at all times and save those recordings (Stegner, 2023).
1.5.8 Smart home automation
A smart home automation integrates household appliances and systems through a network to achieve monitoring and management functions. The central management system enhances convenience and efficiency of the machines and systems. When combined with artificial intelligence (AI), the smart home automation collects user data, automates settings for home appliances such as temperature and lighting systems, offers personalized setup, and takes proactive actions in response to detected abnormal actions (Dasgupta et al., 2022).
1.5.9 Facial Recognition
Facial recognition is a form of technology capable of evaluating human faces and informing the user of who the individual is (Lewis & Crumpler, 2021). There are many variations of facial recognition, from smaller uses such as Face ID on a cellphone to evaluating an image or video clip from a crime scene of an unknown individual against a large database to determine if an identification can be made (Adjabi et al., 2020). Facial recognition works by evaluating small features of the face against another image to determine if similarities exist (Qinjun et al., 2023). This technology is rapidly changing and evolving, however, there are many issues and limitations with the use of facial recognition technology. Issues such as accuracy, discrimination, legality, privacy, and data security have been heavily debated in recent years as facial recognition has become more prominent in various domains such as healthcare, retail, government, and law enforcement (Qinjun et al., 2023). While this technology may have challenges, some see its usage as an aid in human decision-making abilities to help error reduction in identification (Lewis & Crumpler, 2021).
1.5.10 Rule-based and expert systems
Rule-based and expert systems are a type of AI that uses rules and logic to make decisions, mimicking human experts in specific domains like medical diagnosis or stock trading. The underlying algorithms involve ways of combining data about a current situation with a complex set of context-dependent rules to determine an outcome or recommendation. These emerged in the 1970s and 1980s as one of the first bubbles in AI systems. Developers of such systems needed a detailed understanding of the domain and the rules that govern a rule-based system, as end users typically don't need to comprehend its inner workings. This often meant engaging with experts within the domain to elicit their knowledge, identify the rules that described their decisions, and carefully delineate many edge cases, and eliciting and creating such rules came to be recognized as the greatest challenge for creating and maintaining such systems.
Some examples of rule-based systems include MYCIN (Shortliffe, 1976) (a medical diagnosis system), XCON (a configuration management system), and CLIPS (a tool for building expert systems) (see Berka, 2015). Rule-based systems can be limited in their ability to handle unexpected situations that don't match the predefined rules. Errors may occur if the rules are incomplete or incorrectly specified. Understanding rule-based systems' limitations can enhance users' trust, allowing them to comprehend their capabilities and limitations rather than blindly trusting their accuracy or completeness (Wanner et al., 2022).
1.5.11 Adaptive cruise control
Cruise control is a commonly used car feature that can maintain a set speed. Traditionally, cruise controllers have used simple PID controllers, adjusting the throttle to smoothly maintain a set speed. Adaptive cruise control adds additional intelligence and incorporates external factors into the maintenance of the vehicle's speed, including sensors in the front of the vehicle to measure distances to objects in front of the vehicle and use of braking for added safety. A set speed can be superseded by the car's speed in front of the operating vehicle and includes a set distance between the two cars to maintain comfortable spacing. If the front vehicle moves faster than the set speed, then the vehicle will only move as fast as the set speed. Adaptive cruise control can be considered another step toward fully autonomous vehicles as the car can regulate to external factors beyond a desired speed limit set by the operator.
1.6 Summary and Discussion
The goal of this chapter is to demystify many of the terms, algorithms, concepts, and constructs used in AI, Machine Learning, and other related domains. Basic understanding of these concepts is important for a professional or researcher in human factors or cognitive systems engineering who wants to interact with engineers working with such systems, and who want to improve human use of these systems. The particulars of these systems need not always be a black box, a understanding the basic strengths, weaknesses, and constraints of these systems can be critical for helping users develop justified trust and mistrust of the system, by understanding the competence envelope of the system, its boundary conditions for proper use, and situations in which it is not appropriate.
Acknowledgments
This chapter was developed as part of a class project for the MTU ACSHF graduate course HF 5430: Human-AI Interaction during Spring 2024.
Author Contributions
Contributors to the chapter are listed in alphabetical order.
CC: 2.6, 5.2, 6.4:; BF: 4.2, 5.3, 6.5; EM: 2.5, 3.4, 5.6, 6.3; KU: 3.1, 5.2, 6.6; ;LS: 2.1, 2.4, 5.1, 5.5; JW: 3.2, 3.3, 4.3, 6.7 ; SW: 5.3, 5.7, 6.8; YJW: 4.4, 6.1, 6.2; BW: 5.1, 5.4, 6.9. STM: Conceptualization, Review, Editing, and additional writing of all sections.
Conflicts of Interest
The authors declare no conflicts of interest.
AI Usage Statement
Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.
How to Cite This Chapter
Cischke, C., Frisch, B., Matas, E., Mueller, S. T., Sprague, L., Ulinski, K., Walker, S., Wang, J., Wang, Y., & Woolman, B. (2026). Demystifying automation, artificial intelligence, and intelligent software tools and applications. In Shane T. Mueller (Ed.), A Handbook of Human-AI and Human-Automation Interaction. https://pages.mtu.edu/~shanem/human_ai/