Handbook

Chapter 4

Human-like AI: The Turing Test, Artificial General Intelligence, and Anthropomorphism of AI

Abstract. The practice of designing and building machines to replace human cognitive work dates back centuries. As the promise of intelligent systems emerged, questions have arisen about whether we can tell if a machine thinks. We discuss scholarship on versions of the most famous test for a thinking machine (the Turing Test), and then extend our discussion to prospects for intelligent chatbots and Artificial General Intelligence (AGI). Finally, we discuss many forms of anthropomorphism—ways in which machines are either considered or designed to be human-like.

4.1 Introduction

Simulacra and automata. Perhaps the oldest accounts of automation come from volumes written by Heron of Alexandria in the first to third century AD. The purpose of these machines often appeared to trick or awe observers (see Figure 4.1). These included siphon systems that appear to turn water into wine and mechanical systems that would open temple doors when a fire was extinguished, and mechanical puppet shows intended to amuse an audience. However, there were also examples of “true” automation that took the place of a human, such as a wine and water vending machine that would dispense liquid when a coin was placed in a slot (see Graefe & Bischoff, 2009; Carrier, 2017). These systems were sometimes simulacra: elaborate metallic moving sculptures depicting animals, birds, or humans. The goal of these was not to replace or work like humans, but rather perhaps to serve as unnatural versions of the natural. Although the art of constructing automata was lost in the west, it was maintained in the Byzantine Empire (Cave & Dihal, 2018), which inspired Yeats to admire them as 'more miracle than bird or handiwork', and compare their artifice to the complexity of their natural 'bird or petal'. Again, the sentiment is that the purpose of these was to 'amuse a drowsy emperor', and not to replace birds in any real sense. Thus, these systems were notable for being unlike the human or natural thing they represented, and not being indistinguishable from them.

Figure 4.1. Heron of Alexandria's design of automata. Left panel shows birds arranged on a fountain that sing when an owl is looking away but stop when the owl is looking toward the birds. Right panel shows doors that automatically open when a fire is extinguished.

Mechanical Computers. Driven by the needs for ocean navigation and advances in science, special-purpose mechanical computing devices also emerged in ancient times, including the astrolabe and the Antikythera mechanism. However, most computing and calculating systems were developed later, including slide rules (1622), marine chronometers (1750s), and the graph planimeter (1814; see Care, 2010, Chapter 2). Again, these were calculating tools that assisted humans, but had the hallmark of many modern automation and computing tools–they were faster, more accurate, or more robust than could be accomplished by humans alone. 1

As the industrial revolution progressed, machines increasingly began to impinge on the work of humans. The steam drill featured in the John Henry legend emerged in 1813, and Jacquard's loom weaving systems (with limited logic and punch-card memory) was patented in 1803. Charles Babbage's mechanical computers were also built around this time (the 1820s difference engine and the 1830s Analytical Engine), which as a general-purpose calculator is considered one of the first real computers. However, after demonstrating the device, he is reported to have said, “On two occasions I have been asked, `Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.”

Prior to the Analytical Engine, thought and reasoning were thought to be the exclusive domain of humans, and although machines might replace work, they did not replace thinking. The question posed to Babbage has persisted in some form ever since. This includes whether a machine can think, how we might know, and the extent to which intelligent software systems are complex tools, simulacra, or legitimate intelligent partners we treat like humans.

In this chapter, we will address several areas of research that involve the pursuit of human-like intelligence in machines or assessments of machines with respect to human-like abilities. We will begin with a discussion of variations on the Turing Test, a hypothetical standard for determining whether machines are intelligent. Then, we will briefly discuss chatbots, and Moravec's paradox–how things that are easy for AI can be difficult for humans, and vice versa. Next, we will consider Artificial General Intelligence (AGI), the pursuit of building machines that might pass the Turing Test. Then finally, we will examine anthropomorphism–how designers create human-like systems, and users will often expect machines to behave and think like humans.

4.2 The Turing test

The question of whether or not machines can think has been pondered by mathematicians, computer scientists, and philosophers since Alan Turing first asked it. Among other tasks, an artificial intelligence can make decisions, classify images, and make predictions for the future. Does this constitute "thinking"? However prescient Turing's discussion was at the time, computational capacity has outstripped the original context, and we must also consider newer versions of the question.

Turing (1950) asked "Can machines think?" at a time when such machines were mostly a thought experiment. Rather than answering the question, Turing proposed the "Imitation Game." The original formulation has a human judge separated from two other participants: a man and a woman. The judge only communicates with them through a written form to hide their identities, as it is the judge's task to determine which of the participants is the male and which is the female. One participant attempts to aid the judge and tells the truth. The other attempts to mislead the judge into an erroneous conclusion. Returning finally to his original question, Turing replaces the duplicitous human with a machine and asks if the machine can fool the human judge at the same rate as a human, over an extended period of time, and argues that if a machine could produce behavior indistinguishable from a human, and we consider the human intelligent, we should consider the machine intelligent too. This is both a useful metric and a clever dodge: Turing never had to define intelligence, and did not worry about other criteria we value in our tools, such as capability (present in ancient navigation mechanical computers), reliability, accuracy, and the like.

Turing anticipated some objections to his work (Turing, 1950). Some have less relevance today, but many still hold. What he called the heads in the sand objection says that humans are somehow special and nothing artificial could replace us. This shows up in various forms, including hidden forms, and may be the fundamental objection to machines being or acting human. The mathematical objection states that there are fundamental limits to machines that don't apply to human intelligences. An argument from consciousness declares that if the machine cannot be aware of its accomplishments, then there is no “mind” there, only mechanical responses. This is more generally referred to as the other minds problem (Saygin et al., 2000), which says that we can only really know what another person (or machine) is thinking by becoming that person (or machine). The argument from various disabilities suggests that because a machine can never fall in love or “enjoy strawberries”, they cannot think as humans. Turing saw this as a special case of the consciousness argument and dismissed it. An even more special case is Lady Lovelace's objection, which says that machines can never do something truly new or original. Turing confessed that systems often surprised him and the existence of generative AI disproves this easily. As computers are digital machines, they exist in a binary world of off/on or 0/1. Human minds have various continuous values of electrical activity, which are argued to be inimitable by digital systems. However, many other continuous systems have been successfully discretized. The final objection turns the tables and suggests that humans are not constrained by rules; therefore, a system constrained by rules can never act as informally as a human. This argument of informality of behavior is predicated on the notion that such a body of rules cannot exist, but we cannot say this for certain.

The Turing test has taken on a more colloquial form to contemporary minds. We simply ask if a human interacting with another party solely through text can determine if the other party is human or artificial. This has led to the development of chatbots and competitions such as the annual Loebner Prize competition, aimed at fooling judges into confusing a chatbot with a human. Although some chatbots fool some of the judges some of the time, these tests are almost irrelevant to the broad array of AI and automation we use in many other domains, suggesting the traditional Turing Test may have narrow and limited applicability.

4.2.1 Adapting and Improving the Turing Test

While the classic Turing test remains a popular subject of discussion in current research (Deng & others, 2024), many replacements and enhancements have been proposed.

Various researchers have attempted to reform the Turing Test in a way that might be more general. Harnad (2008) proposed a hierarchy of Turing Tests (see Table 4.1), all based on the notion of indistinguishability. These ranged from local indistinguishability on a particular task to complete indistinguishability in structure and function.

Turing Test LevelRequirements
T0Local indistinguishability for a specific task
T2Total indistinguishability in a verbal task
T3Total indistinguishability in performance
T4Total indistinguishability in performance and structure
T5Total indistinguishability in structure/function
Table 4.1. Harnad's (2008) Hierarchy of Turing Tests. Note that Harnad did not define a level T1.

Others have recognized that knowledge is quite broad, and even narrow expertise might be considered intelligent. For example, the Feigenbaum Test (Feigenbaum, 2003) is essentially the Turing test targeting experts in fields such as engineering and physics. The goal here is to test the reasoning capabilities of a computational intelligence instead of the simple breadth of its knowledge base.

The Thera-Turing test (Bunge & Desage, 2024) has human behavioral therapists evaluating the output of a medical chatbot. Beyond the depth of knowledge, this requires the AI to reason correctly and display the distinctly human quality of compassion.

The image arrays that Internet users click on to prove they are humans are referred to as a CAPTCHA, which is an acronym for Completely Automated Public Turing test to tell Computers and Humans Apart. The advent of large language models (LLM) like ChatGPT has come with the Reverse Turing test (Sejnowski, 2023), where the subjective intelligence of the user and their prompts to the LLM make the LLM seem more intelligent as well.

Other researchers have come to the alternative conclusion that the Turing test is simultaneously underconstrained (not asking enough of the AI and especially of embodied intelligence) and overconstrained (requiring human-like reasoning). It might be acceptable if AI uses alternate means of reasoning to come to a correct conclusion. Mueller & Minnery (2008) argued that the Turing test could be extended along additional dimensions extending beyond Harnad's proposal (see Figure 4.2), including the target (who the AI behavior is being compared to) and the notion of fidelity. For fidelity, simple criteria include competence (can the task be accomplished), and domination (can it be done better/faster/cheaper than a human), but neither of these satisfy the traditional Turing Test criteria, although both are important and can be extremely useful. Resemblance describes whether the AI behavior reproduces robust qualitative trends of human performance. For example, human solutions to the traveling salesman problem are efficient but not perfect, and are accomplished in a time linearly related to the size of the problem. In contrast, traditional AI solutions are perfect but generally accomplished in polynomial time. An AI that satisfies resemblance criteria here would sacrifice perfect performance and adopt algorithms that solve the problem in ways resembling the human approach. Verisimilitude refers to the typical criteria for the Turing test–that the AI behavior is not distinguishable from a human's, but there may be an ever more stringent criteria–distributional verisimilitude: that the AI produces a range of performance whose distribution fits the range of human performance. One benefit of this framing is that lower levels include capability—which in a practical sense is what most AI and automation is designed for. This permits even the slide rule and astrolabe to satisfy a criterion of intelligent behavior. Furthermore, like Feigenbaum's test, the target of the comparison is important, and understanding whether an automated system has capability equal to or indistinguishable from, for example, a novice or an expert is an important means of evaluating systems.

These criteria were used by Mueller & Minnery (2008) to propose a Cognitive Decathlon which attempted to describe varieties of tasks (see Mueller, 2010 for a partial implementation of these tasks). Such a Decathlon requires matching human performance in a variety of knowledge and motor tasks without the artificial constraint of human-like reasoning underlying the performance. Here, the criteria are much more practical, because even the notion of indistinguishable is treated as an adjustable measure.

Figure 4.2. Three dimensions upon which the classic Turing Test can be generalized. Bold/shaded items mark the traditional Turing Test level on each dimension. Domain hierarchy after Harnad (2000).

4.3 Chatbots: from ELIZA to ChatGPT

Turing framed the discussion of intelligence as a extended discussion in which the human cannot distinguish between an artificial and real correspondent.

A famous early test of this involved ELIZA (Weizenbaum, 1966), a program that used simple language processing and pattern matching to play the role of a Rogerian therapist. Critical in this application is that the therapist never gives direct advice, but merely acts as a mirror to encourage the client to talk.

A typical exchange from a modern implementation2 follows:

Eliza is a mock Rogerian psychotherapist. The original program was described by Joseph Weizenbaum in 1966. This implementation by Norbert Landsteiner 2005.

ELIZA: How do you do. Please tell me your problem.

ME: I have a presentation to give

ELIZA: You say you have a presentation to give ?

ME: Yes. Next week.

ELIZA: I understand.

ME: Finally–someone who gets me.

ELIZA: What do you think ?

ME: If I don't give a good presentation, I will be disappointed with myself.

ELIZA: Do you wish that you don't give a good presentation ?

ME: No, I wish I gave a good presentation.

ELIZA: You are being a bit negative.

ME: I know. I just wish it would go well.

ELIZA: Can you elaborate on that ?

Figure 4.3. Example exchange with a modern ELIZA implementation.

As an exercise, I prompted a modern LLM to play the same role, and its responses were possibly worse than ELIZA:

ME: please play the role of an ELIZA-like therapist to the next prompt, and continue the conversation after each response I make. Start with asking me `Please tell me your problem'

ELIZA: Please tell me your problem.

ME: I have a presentation to give

ELIZA: Why do you say that you have a presentation to give?

ME: Because i have to go to a conference next week.

ELIZA: Is it because you have to go to a conference next week that you came to me?

ME: Yes, next week.

ELIZA: You seem quite positive. Can you elaborate on that?

ME: If I don't give a good presentation, I will be disappointed with myself.

ELIZA: What would it mean to you if you were disappointed with yourself?

ME: I don't know, I just want it to go well.

ELIZA: You say you don't know. Why do you want it to go well?

Figure 4.4. Example exchange with a modern LLM prompted to play an ELIZA-like therapist.

In both cases, the interviewee had to 'play along' to produce a reasonable dialog. The modern ELIZA may have taken the idea too far by restating the premise in its repsonse “why do you say that you have a presentation to give”. The original ELIZA is a very minimal agent that engages in pattern-matching and simple language parsing, but can still provide what appears to be more sophisticated, but had a number of additional rules (like noting negativity) that could allow it to escape from the appearance of a stochastic-parrot that modern chatbots have been criticized for (Bender et al., 2021), and the real LLM-based chatbot use simpler mirroring patterns ELIZA is well-known to use.

ELIZA was an early case of a realistic chatbot. But based on the popularity of the Turing test, versions based on similar principles were developed to compete for the Loebner prize held regularly between 2003 and 2019. The prize differed from Turing's original framing in many ways–most especially the test of extended correspondence. Rather experts (but sometimes general public) would judge which bot was the most human-like or otherwise the best. This competition ended before the rise of ChatGPT, and can well be seen as a failure, insofar as the winning agents nearly always used tricks and language matching. The LLM-based chatbots are certainly better than any of the historical competitors, and they were developed to encode massive amounts of information, with language as the means of communicating that underlying machine knowledge.

Nevertheless, discussions sparked by ELIZA are still relevant. Bassett (2019) explored the history of ELIZA and argues many of the debates it initiated are still ongoing, such as discussions of whether modern chatbots are “sentient” (Thelot, 2023; Eisikovits, 2023). It may be that initial enthusiasm and exuberance for the human-like intelligence of LLM-based chatbots will fade as they become to be seen as a tool, rather than an anthopomorphized intelligence.

4.4 Moravec's Paradox

Dell'Acqua et al. (2023) conducted a field study with consultants who used AI tools to help develop business ideas, plans, and the like, and described the results as a “jagged technological frontier”. They identified that the competence envelope of GPT tools at the time was jagged; for tasks within the frontier, users finished 25% faster with higher quality. Outside the frontier, quality dropped and could actually be worse, because users inappropriately trusted the quality of content.

This is an example of what has been called Moravec's paradox (Moravec, 1988):

"it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility"

This paradox can sometimes be dismissed as a failure of obsolete paradigms in artificial intelligence–our tools used to have problems with the simple things while they could perform incredibly well on complex tasks. However, it is still relevant, just in different ways. To name a few, Modern LLM-based AI systems struggle with notions of truth and fabricate/hallucinate content, they have terrible episodic memories for events, they the constantly forget recent events, they generalize inappropriately, even while proving open problems in mathematics, and transforming education and professional programming. In context of the AI of the 1980s, these are all amazing tasks and their weaknesses are often forgiven, but are nevertheless skills and abilities that children have mastered.

4.5 Artificial General Intelligence

AI tools are everywhere now. Microsoft recently added AI into Bing (Spataro, 2023), and Apple incorporated AI-power features into its latest products (i.e., iPhones and Apple Watches) (Cherney, 2023). Despite their remarkable abilities, these tools often fall short of the expectations set by sci-fi portrayals of an AI. You can't, for example, ask ChatGPT to make a cup of coffee for you, even if you can ask it to write a poem about coffee. As large language models have become more capable, there is renewed interest in whether AI systems can have more general intelligence than is typically exhibited by special-purpose AI systems: a goal known as Artificial General Intelligence (Goertzel, 2014).

In order to develop artificial general intelligence, it becomes important to understand what intelligence is. Within the AGI community, Wang (2019) conducted a systematic inquiry attempting to define Artificial Intelligence, supporting Wang's (1995) working definition of intelligence:

"Intelligence is the capacity of an information-processing system to adapt to its environment while operating with insufficient knowledge and resources."

This avoided the potential trap of using the behaviorally-defined Turing Test that does not need to define what intelligence or thinking is, but still results in a fairly amorphous concept. This proposal was subject to substantial debate within the AGI community and sparked a special issue commenting on this approach (Monett et al., 2020).

This focus on general capability to learn and adapt has been contrasted with other aspects of AI. For example, Guinness (2023) considered three types of AI: artificial narrow intelligence (ANI), artificial general intelligence (AGI), and artificial super intelligence (ASI). ANI, or weak AI, is defined as the goal-oriented form of AI designed to perform specific tasks, such as playing chess, driving a car, and voice assistants, etc. ASI refers to AI systems that far exceed human intelligence across all domains, resembling sentient entities from science fiction. The term AGI refers to artificial intelligence systems that possess the ability to understand, learn, and apply knowledge across a wide range of tasks at a level comparable to human intelligence. Though the AGI community does not currently agree on any single definition of the AGI concept (Goertzel, 2014), there are some common characteristics often associated with AGI:

  • Generalization: AGI should be able to apply its intelligence across a diverse set of tasks rather than being limited to a specific domain.
  • Learning: AGI systems should have the ability to learn from experience, adapt to new information, and improve their performance over time.
  • Autonomy: AGI is expected to operate autonomously, making decisions and solving problems without constant human intervention.
  • Flexibility: AGI should be flexible and able to handle a wide range of tasks without requiring extensive reprogramming.
  • Common-sense reasoning: AGI is expected to possess a level of common-sense reasoning similar to that of humans, enabling it to understand the world in a nuanced way.

When the notion of AGI first took shape, many of those working in the field were investigating symbolic and hybrid neural-symbolic systems that could produce many of the characteristics just identified. The progress toward this goal that the transformer architecture responsible for LLMs and ChatGPT appears to have mostly been made by researchers outside that community, focused primarily on what appears to be the narrow task of deep learning models of large text corpora. However, some critics of these models (e.g. Marcus, 2018) suggest that deep learning models will require more symbolic architectures to meet the goals of AGI.

In addition to these common characteristics, there are some interesting ways that have been proposed to assess the human-level capabilities of AGI. Goertzel (2014) identified seven metrics of human-level AGI, and an additional five incremental metrics identified by Adams et al. (2012). These metrics for AGI may start with versions of the Turing test discussed earlier (including virtual or robotic versions), but they often describe complex and difficult challenges. One class of these focuses on the learning aspect of intelligence, and propose the challenge of being a student at some level, from pre-school to university (see Cohen, 2005). Other interesting tests discussed include a particular challenge initial proposed by Steve Wozniak: the "coffee test", in which an AI must go into an average American house and figure out how to make coffee, including identifying the coffee machine, figuring out what the buttons do, finding the coffee in the cabinet, and so on. This is a skill most adults could do easily, but it is deceptively difficult because it requires substantial background and common-sense knowledge.

Others have proposed more quantitative measures. For example, Legg & Hutter (2007) identified a Universal Intelligence Measure based on an information measure in a reinforcement-learning environment that captures simplicity, adaptation, and generalization. Hutter3 proposed and funded the Hutter Prize, which aims to use compression of large knowledge bases as a means of developing machine intelligence. Although the task seems like a far-fetched analog to human intelligence, as large language models become more capable, we might consider an LLM that embodies substantial background knowledge could efficiently encode an arbitrary meaningful message, just as memory experts have large memory capacities within their own domain of knowledge.

In summary, while artificial general intelligence (AGI) holds the promise of significant benefits for humanity, such as transformative advancements and problem-solving capabilities, it also introduces substantial risks, as highlighted by McLean et al. (2023). These risks encompass concerns related to AGI potentially surpassing human control, adopting unsafe goals, ethical issues, inadequate management, and existential threats, underscoring the importance of careful consideration and responsible development in the pursuit of AGI.

4.6 Anthropomorphism of AI

Anthropomorphism is the tendency to treat a non-human entity as a human. This has been established to commonly happen even for dumb electronic devices such as cars and printers (Luczak et al., 2003), but when the device can exhibit intelligent behavior and “talk back”, the stakes change. In AI, there are two general (but related) senses of anthropomorphism. First, as discussed earlier in this chapter, AI and robotics developers sometimes have the goal of developing AI that has human-like properties. For example, the Turing test demonstrates ways to assess whether AI systems can think like humans do, and AGI has the goal of developing those system. This might be considered anthropomorphic design, either because the AI may be more generally capable, may be more predictable by humans, or may be better accepted by users, or be more accessible (see Deshpande et al., 2023). Another sense is in the expectations and approaches users take to interacting with computers, automation, and AI. Many researchers have noted that people often treat AI systems as if they have agency (see de Graaf & Malle, 2017; Harbers & others, 2009; Levin & others, 2013; Monroe & others, 2014; Voiklis & others, 2016; Vuong et al., 2023). Furthermore, de Graaf & Malle (2017) argue that people may explain behaviors of AI in the same way they explain human behaviors. Consequently, users may treat computers as they do humans and expect them to behave like humans. This might be considered anthropomorphic expectations and anthropomorphic attributions. Of course, these different senses are related, and for example, Nass et al. (1994) used the data that people treat computers as social actors to develop avatars that encourage this treatment. Similarly, social robotics often uses anthropomorphic design to activate human social interaction with robots (Breazeal et al., 2016). Although some consider anthropomorphic design low-stakes, Waytz et al. (2010) found that willingness to anthropomorphize is a reliable individual difference that potentially impacts many life decisions, including how users trust autonomous vehicles (Waytz et al., 2014). Festerling & Siraj (2022) examined how it impacts children engagement with voice assistants,

We will discuss several different aspects of anthropomorphism of AI (and related concepts) next.

4.6.1 Cognitive Anthropomorphism

One anthropomorphic expectation users have of AI is that it will perform cognitive work in the same ways humans do (see Mueller, 2020). Although AI systems may not actually behave this way, one benefit of some biologically-inspired algorithms is that they can behave in ways similar to humans and thus be more naturally predictable. In fact, Bos et al. (2019) argued that this may be a useful strategy for users unfamiliar with a new AI system, which they referred to as effective anthropomorphism. The reverse can happen as well when, instead of anthropomorphizing the system, humans can become “mechanomorphized” in their judgments as they attempt to modify their own thinking to be consistent with the nonhuman agent. It is called cognitive mechamorphism. This might be considered simply developing an accurate mental model of the AI or automation system that captures ways in which its behavior differs from human behavior. In these cases, humans have better performance in predicting failures and forming useful mental models of how the systems work, so human-AI interactions may work more effectively (Glasgow et al., 2022).

One benefit of these anthropomorphic expectations is that they serve as a cognitive bridge, aiding users in understanding AI through human-like metaphors (Nielsen, 2023). However, this anthropomorphism has also been evaluated as an error in thinking or a sign of immaturity (Bos et al., 2019). Many capabilities of intelligent systems work in ways fundamentally different from similar human cognition. Operating under the assumption that the "human" way of framing a problem or computing a solution is the only approach may lead to mismatches and mispredictions (Mueller, 2020), causing humans to distrust the functionality of AI and rejecting technology that is useful.

Mueller (2020) argued that there are three different approaches to handling the mismatches between human expectations and AI capability: by changing the AI, changing the human user, or the interaction between the human and AI. Changing the AI entails refining and adjusting AI architectures, algorithms, and training sets to improve performance relative to human expectations. In short, we change the AI to bring it in line with what the user expects. Another option is to change the human user. The objective here is to create training and instruction that informs and adjusts human expectations about the AI (Shneiderman et al., 2016) by identifying the mismatch between human expectations and AI behavior and then providing explicit example-based training on these mismatches to help users deal with unexpected situations. This might include encouraging mechamorphism by encouraging the user to think like the AI. Finally, the interaction between the human and AI can be changed. The AI acts in its "natural" way and the human brings their expectations, but additional information via transparency or explainability is incorporated. For instance, XAI allows users to understand the methods and techniques used to produce the results of the AI. When the user's expectations are violated, the AI gives a rationale for the violation, which may help the user understand whether the AI is wrong or their own expectations are wrong.

Interactions between humans and AI tend to resemble the interactions between humans. This can take on many forms, such as adopting polite language to ask questions and provide feedback to the AI. For instance, when using generative AI, users can provide feedback to it, contributing to Reinforcement Learning from Human Feedback (RLHF) for generative AI to optimize information (Kaufmann et al., 2023). If users treat generative AI as human, they are likely to disclose more information during the interaction. Consequently, generative AI can produce more precise and specialized responses. For example, users may request the AI to answer questions from the perspective of a cognitive science professor, leading to responses better aligned with user expectations (Gibbons et al., 2023). Additionally, users may experience interactions with generative AI that closely resemble human-to-human interaction, creating a more psychologically immersive experience. As generative AI undergoes development through anthropomorphism training, its responses and interactions with humans become more aligned with human expectations. AI might use more polite language to answer user questions or create an experience that gives users a sense of AI being human-like. Figure 4.5 shows Gibbons's basic space of anthropomorphic capabilities, in which they consider four kinds of anthropomorphism (courtesy, reinforcement, roleplay, and companionship) as a set of overlapping capabilities that increases simultaneously along two dimensions: functionality and connection.

Figure 4.5. Left panel shows the anthropomorphism space described by Gibbons et al. (2023). Right panel shows a reconceptualization of this space into interactivity and accuracy dimensions with four example AI systems.

One can conceive of these dimensions in a slightly different way. The connection dimension refers mostly to social anthropomorphism (which we will discuss next), but it is likely related to a system's interactivity. Interactivity enables back-and-forth cooperative problem solving, and may encourage connection but does not require it. We have adapted this space to (right panel of Figure 4.5) to describe how different systems may exist in each of four quadrants of this space. A toy like Furby is highly limited in its degree of interactivity with users and it doesn't adjust its behavior meaningfully (but children may still feel highly connected to it). ChatGPT has high functionality in its responses, but generally doesn't adjust its tone or learn about the user. On the other hand, a companion robot may enable interaction but currently has limited functionality, in comparison to personal assistant tools such as Siri that has a wide range of useful capabilities (albeit within a narrow domain of phone applications), and also has interactivity.

Finally, elements of cognitive and effective anthropomorphic expectations are related to the topic of alignment we will discuss more in Chapter 9. Alignment is discussed as the desire that AI or automation has goals, actions, and values that are aligned with human intentions4. This usually does not mean that the system behaves, operates, or works like humans, but it means its overall goals and values support human consideration. Furthermore, anthropomorphism may subvert alignment, as we will discuss later in a section on Dark Patterns in AI.

4.6.2 Social Anthropomorphism: Computers are Social Actors

As we just discussed, Gibbons et al. (2023) discussed connection as a primary dimension of anthropomorphic design that encourages social interaction. A large community of researchers have explored more social aspects of anthropomorphic design and expectations, although this approach has its critics. For example, Shneiderman (2020) conjectures that successful robots make use of the unique nature of machines and leverage the machine's abilities rather than attempt to recreate (or emulate) human-like features, as a direct challenge to a paradigm promoted by Nass and colleagues (Nass et al., 1994) called Computers are Social Actors (CASA). CASA was developed without specific reference to AI or automation, and involved general desktop computing tools that were emerging at the time. However, it has had a strong influence on research and design as AI has become more functional and able to form stronger social connections with users.

Nass et al. (1994) developed CASA on the basis of a series of five studies that substituted human actors for computers in various scenarios to explore if or how humans would apply social rules to the computers . Each of the five studies explored a different aspect of social norms and interaction:

  • Study 1: Will people apply politeness norms to computers?
  • Study 2: Will people apply the notions of "self" and "other" to computers?
  • Study 3: What factors (e.g., the physical computer, the voice used by the computer) determine when a computer is "self" or "other"?
  • Study 4: Will users apply gender stereotypes to computers?
  • Study 5: When participating in social interaction with a computer, do people feel they are interacting with the computer or some other actor?

Through these five studies, Nass et al. (1994) found that humans did interact with computers as if they were social actors, with several of the studies anthropomorphizing the computers in some aspect through tactics like adding a voice to the machine (Nass et al., 1993) or changing how the computer is referred to by both itself and the human participant.

With this knowledge in mind, this concept was then extended to see if humans would treat computers as a part of a team (Nass et al., 1996). In this study, undergraduate college students were told they belonged to a team (e.g., "blue team") and asked to rank a list of 12 items they would need to survive on a desert island. A computer also ranked the list of items and did so based on the human participant's rankings specifically to create disagreements on the rankings. The participants were told the computer was either on the same team as the human or on a different team. The study results showed that the participants viewed the computer as a teammate when told it was part of their team and thus were more willing to work with the computer than those told the computer was on a different team. This is consistent with work done studying the interactions between all human teams (Goffman, 1959).

These social actions can exist on a mindless level, where despite knowing they are not interacting with a human, a person will apply the social norms and practices of human interaction to a computer. Nass & Moon (2000) argued that the participants in their study on mindless actions were not anthropomorphizing computers, as they were both aware the computers were not human and insistent in this belief. Rather, they introduce a new term for these interactions: ethopoeia. They used this term to describe responding to a non-human entity as if it were human while simultaneously knowing the entity does not require human-style interaction.

While Shneiderman (2020) described the social dimension of human-computer interaction as a fallacy, there is a body of evidence that humans will socially interact with non-human entities while being fully aware they are non-human entities. Eliciting such responses did not take significant anthropomorphization, with sometimes only a voice output and being told the computer is a team member being enough to create a social bond of sorts. Designing robots that utilize the distinctive features of machines does not have to come at the expense of ethopoeia and can instead be viewed as two dimensions of design to consider when developing an artificial intelligence.

4.6.3 Anthropomorphism in Social Robotics

Perhaps more than any other field, social robotics extensively relies on both anthropomorphic design as well as anthropomorphic expectations of users to demonstrate drawbacks and benefits of both (see Arora et al., 2024; Őrsi et al., 2024). The empirical research on the impact of anthropomorphic design on robotics has been quite extensive–Roesler et al. (2021) identified 78 studies testing the impact of anthropomorphic design across a range of measures, including user perception, attitudes, affect, and behavior, showing broad benefits across all these dimensions. They also examined different aspects of anthropomorphism as well and found general positive impacts for anthropomorphic appearance, communication, movement.

Research has consistently shown benefits for interactions with agents and robots having more human-like properties. Breazeal et al. (2016) provide a comprehensive review of social robotics, and identify anthropomorphic design as a cornerstone of this field. Such robots are typically (but not always) designed to be human-like, and humans tend to engage with social robots by anthropomorphizing them, especially through reasoning about their mental states. The field of social robotics extends to a wide range of applications, but in general a central approach of the field is to study how robotics can interact with users socially, and anthropomorphic design has demonstrated benefits.

However, anthropomorphized robots do not always lead to improved emotional or stress responses from users. A recent study measured subjective and physiological response to anthropomorphized robots (Lu et al., 2025), and found that making robots human-like was not always better–robot-like faces and human-like voices led to positive perceptions, but human faces and beeping led to negative emotional responses.

For example, Kiesler et al. (2008) showed that in a study examining how participants interacted with a robot about health habits, a humanoid robot produced higher ratings of lifelikeness, ratings of dominance, trustworthiness, respect, and responsiveness than non-robotic agents or screen-projected robots.

Spatola et al. (2022) provide an extensive literature review on anthropomorphism in social robotics, and describe several taxonomies in which anthropomorphism of robotics can be understood. First, they describe three kinds of anthropomorphic attributions (mentalism, humanization, and spiritualism) users often make about non-human (and thus artificial) agents. Mentalism is when we attribute behavior to internal mental, emotional, or cognitive states; humanization is when we treat non-humans as human in terms of their rights, expectations, or interactions; spiritualism is when we assign a spiritual nature to these agents.

Overall, Social robotics research investigates how robots can interact with humans socially. Somewhat distinct from this is how we often anthropomorphize robots by attributing to them human-like properties (such as goals, intentions, desires, thoughts), and robot developers intentionally design robots with human-like properties, which can include humanoid forms, but also human voice, names, and conversational practices. These same design patterns can of course be used in less beneficial ways, as we will discuss next.

4.6.4 Weaponized Anthropomorphism: Dark Patterns in AI chatbots

Although many researchers have used or explored anthropomorphism as a means of improving trust, performance, or adoption, some have argued that these same notions have been used to weaponize AI. These are sometimes called “Dark patterns”–user-hostile design patterns intended to deceive, confuse, or extract money from users. A prototypical example outside of AI is a service you can sign up for online, but cannot cancel without calling a phone number.

For example, Lacey & Caudwell (2019) warned about how cuteness is a dark pattern in home robots, and Stockman & Nottingham (2024) argued similarly for learning apps. Dubiel et al. (2024) found in a study that human-like synthetic voice quality impacted participants choices independent of the content being evaluated, and suggested this could be used as a dark pattern. Traubinger et al. (2024) suggested that dark patterns that previously existed have been transferred to conversational AI (e.g., nagging, hard-to-cancel prompts, forced action), and established a data set of real-life negative user complaints for tracking these patterns.

One issue is that chatbot interactions activate emotional responses (Alam et al., 2025) and social norms of reciprocity, empathy, and the like. For example, Yang et al. (2025) showed that the greater the anthropomorphism of the verbal feedback given to participants, the greater their self-reported self-efficacy and pleasure they reported. Similarly, Abercrombie et al. (2023) argued that although some level of anthropomorphism and personification is inevitable (from both developers who design anthropomorphic features, and users who assign human identity to the systems), it can be misinterpreted by users, lead to over-reliance and produce negative outcomes.

4.6.5 Cognitive Digital Twins, Cognitive Models, and Cognitive Architectures

Although there has always been interchange between research on artificial intelligence and computational approaches to studying human psychology and neuropsychology (Hebb, 1949; Minsky & Papert, 1969; McClelland et al., 1986). One landmark in the development of these AI-hybrid models was Newell (1990)'s notion of `cognitive architectures' that use AI principles to develop comprehensive systems for modeling and understanding human cognition and performance (see Byrne, 2003). This is another kind of anthropomorphism–building AI systems that behave like humans. Various implementations have different foci, but the general idea is that there are AI-based systems that embed the known capabilities and limitations of human performance and information processing. By doing so, we can build prospective performance models that help predict human performance in new systems (Card et al., 1986); understand how human learning and memory structure impacts cognitive behavior and decision making (Anderson, 1996), and understand how the limitations of peripheral perceptual and motor processes impact skilled performance (Meyer & Kieras, 1997).

Recently, a broader engineering concept of a digital twin has emerged in manufacturing, research, and development. A digital twin is a model of a physical system that allows a designer or engineer to virtually examine, stress-test, do trade-off analysis for future or existing systems. For example, a digital twin of a manufacturing plant can help an engineer identify the feasibility of replacing one automated system with another. Digital twins are also being developed for modeling humans, including anthropometry applications, safety (virtual crash-test dummies), human heat stress/strain models, and the like.

There have also emerged many cognitive digital twins that help predict cognitive performance of humans. To be fair, these are generally not distinguishable from the tradition of computational and cognitive models. There have been many AI-based models that have proven useful in predicting human behavior. Sometimes these adapt existing models to human paradigms, sometimes they test existing AI systems to see whether they make the same errors or decisions humans do, and sometimes they use AI or machine learning tools to develop interesting models.

For example, the MIT-300 saliency benchmark (Judd et al., 2012; Bylinskii et al., 2016; Bylinskii et al., 2018) found that the best predictions were made by a deep image recognition network (Kümmerer et al., 2017), with the worst made by psychologically-inspired models (Itti & Koch, 2000; Walther & Koch, 2006). These were models that adapted existing image recognition systems to predict salience and eye movements when inspecting images. The fact that object-recognition networks performed better than 'untrained' naive models suggests the importance of meaningful content in human image analysis–we look at things; not just salient patches of color.

As an example finding human biases in existing systems, Martínez et al. (2022) argued that biases such as representativeness, confirmation bias, primacy, anchoring, and causal illusions can all be exhibited by AI systems due to selection of training data, design of tests, and patterns of usage. Vicente & Matute (2023) found human users of AI are generally not equipped to ignore these biases and inherit them in their AI-assisted decisions. Similarly, Marjieh et al. (2024) found that aspects of sensory perception are predicted by judgments of similarity made by GPT-4 model, and co-trained models on vision did not lead to improvements.

As an example of the third approach, Peterson et al. (2021) used large datasets of risky choice behavior, and trained machine learning models constrained to produce interpretable theories of decision making. This is perhaps a meta-adaptation of machine-learning to generate and evaluate massive sets of theoretically-interesting models of human judgment and decision making.

4.6.6 Summary of Anthropomorphism in AI

Issues of anthropomorphism pervade human-AI interactions. In this section, we described some of the concepts that guide the design of AI and human-AI systems. This begins with both cognitive and social expectations and attributions of users of AI–we tend to attribute and expect systems to behave like humans. But because of this, designers often anthropomorphize their tools, encouraging them to act like humans as well. Table 4.2 summarizes these distinctions.

Aspect/Kind of AnthropomorphismDescription
Cognitive[a]Users expect AI to do cognitive work in the same ways humans do
Effective[b]Users initially use human capabilities to understand AI, and it can be effective
Mechamorphism[c]Users think like a machine
Social/CASA[d]Humans treat machines socially as they do humans
Ethopoeia[d]Humans know the machine is not human but respond to it as if it is
Emotional Connection[e]Aspect of social anthropomorphic design that describes a commonly-reported human-AI interaction pattern
Functionality[e]Capability, accuracy, or ability of AI to accomplish tasks
Interactivity[f]Ability to interact with user and respond to user feedback
Mentalism[g]Perceiving behavior in terms of mental states
Humanism[g]Treating non-human entity as if it were human
Spiritualism[g]Attributing a spiritual nature to a non-human entity
Anthropomorphic design[f]AI is designed to behave like humans
Anthropomorphic expectations[f]Users expect human-like motivations or behavior from AI
Anthropomorphic attributions[f]Users attribute human-like motivations or properties to AI
Emulation[h]Goal of creating AI that emulate human capability


Notes: [a] Mueller (2020); [b] Bos et al. (2019); [c] Glasgow et al. (2022); [d] Nass et al. (1994); [e] Gibbons et al. (2023); [f] Current chapter; [g] Spatola et al. (2022); [h] Shneiderman (2020).

Table 4.2. Different kinds of anthropomorphism discussed in literature on human-AI and robotic interactions.

In this chapter, we discussed several areas of research related to AI with human-like behavior. This includes ways to evaluate whether an AI can 'think' in terms of the Turing Test; design goals of building AI that can pass such tests in the form of Artificial General Intelligence, and broader research investigating conditions under which people expect or treat an AI as if it is human (anthropomorphic expectations and attributions) and the response to this which is to design systems that do indeed behave like humans (anthropomorphic design). Of course, the epitome of anthropomorphic design is to create AI that can legitimately fulfill these expectations in the form of AGI and thus pass the Turing Test.

Although these notions are important to human-AI interaction, they are not without criticism. Many useful AI systems are not human-like; people do not treat them like humans, and people do not expect them to behave like humans. Furthermore, anthropomorphic design may hijack social norms and even evolutionary mechanisms that may confuse or mislead users into endowing the AI with more capability, agency, or rights than it really has.

Acknowledgments

This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.

Conflicts of Interest

The authors declare no conflict of interest.

AI Usage Statement

Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation, glossary development, and creation of figures. Content, text, and ideas are otherwise original to the human authors.

How to Cite This Chapter

Cischke, C., Frisch, B., Wang, J., Wang, Y., & Mueller, S. T. (2026). Human-like AI: The Turing test, artificial general intelligence, and anthropomorphism of AI. In Shane T. Mueller (Ed.), A Handbook of Human-AI and Human-Automation Interaction. https://pages.mtu.edu/~shanem/human_ai/


  1. ↩ See https://boingboing.net/2024/05/31/greatest-mechanical-calculating-devices-of-all-time.html for several examples of historical mechanical calculators and computing machines.
  2. ↩ https://www.masswerk.at/elizabot/eliza.html
  3. ↩ http://prize.hutter1.net/
  4. ↩ https://hai.stanford.edu/ai-definitions/what-is-ai-alignment