Chapter 9
Safe, Aligned, Fair, and Ethical Considerations for Artificial Intelligence and Automation
Abstract. [Abstract pending in source draft.]
*****Notes: A Similar adaptive management approach called AI Alignment has been proposed (Wixom et al., 2020) with the goal of achieving consistency across diverse stakeholder needs. These principles advocate for cross-disciplinary collaboration and require rigor to justify the investment to develop, deploy, and assess the human-AI system. May belong in chapter 13.
9.1 Introduction
9.2 Alignment Problem in AI
What happens when artificial intelligence no longer acts in ways that benefit humans or are in accordance with human values? Popular culture has addressed this question in various ways. AI misbehavior has been displayed in many popular works of fiction, like 2001: A Space Odyssey, The Matrix series, the Terminator series, and Ex Machina. In all of those works, the AI ceases to align its behavior and actions with human values. Thus, we define AI alignment as the state in which an artificial intelligence system follows human intent and does what humans want it to do, even in situations containing considerable uncertainty. The problem of AI alignment is creating and maintaining alignment in developed systems, which AI developers and researchers consider to be a fundamental problem in the field. In one survey, 5
An early (and fictional) version of alignment comes from the author Isaac Asimov, who produced the Three Laws of Robotics:
- A robot may not injure a human being or, through inaction, allow a human being to come to harm.
- A robot must obey orders given it by human beings except where such orders would conflict with the First Law.
- A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
While a fine starting place, a more modern set of guidelines for alignment rests on four principles: Robustness, Interpretability, Controllability, and Ethicality (Ji, et al., 2024). The so-called RICE principles break down into the resilience of AI systems operating across diverse or adversarial scenarios, the ability of humans to understand the inner reasoning of the AI system, the ability of humans to control (and even shut down) the system, and an unwavering commitment from the system to uphold human norms and values.
The key difficulty is communicating those values to the AI. One AI vendor, Anthropic, adheres to a constitutional approach to alignment. This approach differs little from Asimov's Three Laws except that reinforcement learning refines the AI's understanding of the constitutional principles originally specified (Bai et al., 2022). Specific principles are better, but some experiments show that a single, simple principle of "do what's best for humanity" can be generalized by simple AI models (Kundu et al., 2023).
A structured approach to this starts by decomposing the problem into three subdomains of alignment: specification alignment, process alignment, and evaluation alignment (Terry et al., 2023). In the specification alignment stage, the user tells the AI what the task is, both generally and in specifics. For example, we may ask the AI to generate code to implement a search algorithm (the general task), but that it should be written in the Python language in a manner simple enough for a novice programmer to understand (the specifics). In the process alignment stage, we constrain how the AI performs the task. In our Python example, we may specify that all code should be properly attributed to a source or that an intermediate stage describing the algorithm should be produced to allow the user to tweak the algorithm before the final code generation. Finally, in the evaluation alignment stage, we validate the output for correctness. If our algorithm doesn't properly execute, then the AI is unaligned.
As with other AI issues, the solution may not come from a technical domain. Alignment may be an economic or sociological problem. Humans have had incomplete or misspecified contracts between themselves with well-studied results. Even with a constitutional alignment, the contract between the human and AI may be similarly incomplete because the designer of the AI failed to think of all possible contingencies, insufficient effort was put into the principles, or some principles are difficult (or impossible) to contract. To this end, all reward functions used by AI models are misspecified or incomplete because there are implied portions that human contracts do not need (Hadfield-Menell and Hadfield, 2019).
Many of the implied rules may be considered "silly" (Hadfield-Menell et al., 2019). Alignment with human values requires knowing when silly rules need to be applied and when high-priority rules take precedence. For example, we may tell an autonomous vehicle that it shouldn't hit small animals crossing the road, but this rule should be violated if the alternative endangers human life. The balance between the important and silly rules is critical, as is understanding the relative frequency of needing to apply them.
Perhaps the approach may be to teach the AI what humans value (Butlin, 2019). Humans themselves operate under reinforcement learning – performing some action X yielded positive results, both intrinsically and extrinsically. Therefore, we will do that action more often. Conversely, some other action Y yielded negative results and will be performed less often. Since these results include relational and cultural outcomes (not only did I arrive at my intended destination, I did it without endangering other drivers and pedestrians), an AI that understands human values may be able to fill in the gaps in a contract. However, our poor introspection into our own values, the difficulty in quantifying some types of intrinsic motivations, or the indirect link between certain motivations and their outcomes may make this approach difficult. Moreover, some believe that a universal set of "human values" does not exist (Turchin, 2019), and using them to define AI safety is futile.
Ultimately, alignment is not a "set it and forget it" task. The initial alignment, whether constitutional or value-based, could be considered "forward alignment" (Ji et al. 2024), whereas the ongoing alignment refinement after the system's deployment is a backward alignment. Both are necessary to ensure AI continues to be a long-term benefit to human society.
P(Doom)
Discussion of how there is a discussion of p(doom), the probability that AI destroys humanity.
9.3 Framework for Harms
In 2019, Jobin and colleagues established a list of eleven overarching ethical principles most relevant to human/ai interaction: transparency, justice and fairness, non-maleficence, responsibility, privacy, beneficence, freedom and autonomy, trust, sustainability, dignity, and solidarity. These principles may be used to establish best practices for harm reduction when interacting with ai and machine learning systems. While this is not a comprehensive list of principles that ai developers should use when designing and implementing these tools they do create a solid framework for establishing fair and ethical ai tools.
*Transparency and Trust.* Based on previous research in the field of human and ai interaction, transparency is a topic that arises frequently. Jobin and colleagues highlighted 73 sources in their content analysis that found transparency to be the most important ethical principle to consider when developing ai tools. There was quite a bit of interpretation over what transparency means in terms of ai tools but most agree that it can be framed as an effort to increase explainability, interpretability, or other acts of communication or disclosure. Many researchers also emphasized a need for transparency as it relates to trust in a partner for human-machine teaming (Endsley, 2023; Morley et al., 2020). Transparency in terms of ai tools can be many different things such as disclosing of training data, labeling of algorithm code, or releasing results of testing and evaluation (Peckham, 2024). Those developing AI tools and systems can achieve greater transparency through the disclosure of information that includes but is not limited to boundary conditions, societal impact, natural language explanations, investors and stakeholders, and source code (Nikolinakos, 2023).
*Justice and Fairness.* The concept of justice and fairness are difficult to define but in terms of human-AI interactions, they relate to distribution and access to tools, accessibility, equity, and non-discrimination (Peckham, 2024). AI tools have a history of bias (i.e. racial bias in policing algorithms, facial recognition, and finance tools, gender bias in resume scanning algorithms) thus justice and fairness should be implemented to reduce discrimination and increase diversity, inclusion, and equality in AI-human collaboration . AI tools should evaluate all human collaborators equally regardless of race, gender, or other protected characteristics (Davidson & Ravi, 2020). In their research into anomaly detection, Zhang and Davison (2021) found that the model could remove bias toward and individual but not a group as well as the reverse. Whether AI tools should be trained to be fair and just to the individual or the group determines on the use case for the tool. When designing a personalized chat-bot it is likely to be more beneficial to the end user if it is fair on an individual basis while policing algorithms should be focused on removing bias against groups due to a history of unfair police practices towards specific demographics in the US (Beattie et al., 2022: Fountain, 2022).
*Non-maleficence*. When considering human-AI interactions, Textor and colleagues (2022) found that violations of non-maleficence (do no harm) were considered to be more egregious violations than other ethical violations like prospect of success and proportionality. When designing AI tools for human teaming, developers should avoid causing harm to individuals, communities, or society as a whole depending on the use case for the tool. This concept is closely related to justice and fairness as non-maleficence emphasizes removing discriminatory bias and the reduction of physical or social harms. This can be implemented through harm-prevention strategies that focus on technical measures and governance strategies at the level of AI research, design, development, or deployment (Jobin et al., 2019).
*Responsibility.* Jobin and colleagues highlighted 60 sources that consider responsibility to be an important ethical guideline to consider when designing Human-AI interaction tools and environments. In spite of the numerous calls for responsible AI, researchers are hesitant to define what that means. Some sources define it as AI tools operating with "integrity" and establishing legal liability initially (Matsuo, 2017). Others frame it in terms of the broader societal impact (Siqueira de Cerqueira et al., 2021). With all of this to consider, the key aspects of responsible AI in human-AI interaction should be transparency, fairness, and accountability to the end user (Jobin et al., 2019).
*Privacy*. Privacy is often at the heart of Ethical AI as it can be considers both a value to uphold and a protected right. It is most frequently defined in the terms of data security and protection. Achieving privacy in human-AI interactions can be done through technical solutions, access control, regulatory processes, more research and awareness, and through the creation and adaptation of legal frameworks to support these tools (Jobin et al., 2019). Differential privacy, a model which guarantees that an individual's data has little impact on output, is one technical solution that has been implemented into AI tools (Zhu et al., 2020).
*Beneficience*. This ethical concept of promoting good (beneficence) is mentioned frequently in AI-human collaboration research, it was very rarely defined when Jobin and colleagues performed their content analysis. AI research in the private sector is often focused on the benefit or good that the customer while some research calls for the benefit that AI bring to be shared among humanity, society, and the environment (Jobin et al., 2019; Siqueira de Cerqueira et al., 2021). There is an expectation that AI tools will be aligned with human values but this expectation fails to take into account that human values differ from culture to culture and change over time . This issue could be mitigated by developing AI tools based on the society and cultures that will be utilizing them.
*Freedom and Autonomy.* Autonomy is considered to be the most fundamental human right to self-determination through democratic means, right to establish and develop relationships with other human beings, the freedom to withdraw consent, and the freedom to flourish (Jobin et al., 2019). When considering how AI-human interactions can safeguard human freedom and autonomy, many researchers promote it through transparency and predictable AI, increasing the general public's knowledge about AI, giving notice and consent, or refraining from collecting personal data without informed consent (Jobin et al., 2019).
*Sustainability*. When considering how sustainability intersects with human-AI interactions, one can consider both the enviornmental impact of using the AI tool and how the AI tool contributes to sustainability practices (Wynshberghe, 2021). For AI tools to be considered "Sustainable AI" they must accomplish both. This indicates that not all AI tools can or even should strive for the Sustainable AI label as not all tools should be used to contribute to outcomes that improve sustainability. Some AI tools are run off of equipment that uses renewable resources as well as carbon off-set resources (Wu et al., 2022). However, the high energy demands of these models is still a considerable concern (Crawford, 2021).
*Dignity*. Human dignity is another foundational human right that states that all people hold a special value that is inherent and fundamentally and solely tied to their humanity and not relative to any other factor. This suggests that humanity is worth preserving the existence of for its own sake. With AI tools that interface with human beings researchers call for the promoting of humanity. AI should not diminish or destroy human dignity but preserve and protect it (Jobin et al., 2019). This could be accomplished through new legislation and governance initiatives or through technical guidelines (Van Est et al., 2017).
*Solidarity*. When discussing solidarity as it relates to human-AI interactions, most authors reference this in terms of the labor market and how these tools are likely to impact unemployment rates and the reduction of unskilled labor jobs. This underlines the need for distributing the benefits of AI in order to respect persons or groups who may be vulnerable (Jobin et al., 2019). This could be in terms of upskilling current employees and training employees to be the managers of the AI tools (Sofia et al., 2023).
9.4 Bias in algorithms
In artificial intelligence, bias refers to the occurrence of skewed information or output from an algorithm (Schwartz et al., 2022). Bias in AI can lead to incorrect outputs, impacting the accuracy of an AI algorithm. Schwartz et al. (2022) identified three categories of bias in artificial intelligence:
- Human bias: Biases in human thought that impact how a human perceives information, (e.g. output from an AI system), that they then use to make a decision, inform others, etc.
Many examples of AI bias have been present in society and media. Mozafari et al. (2020) found that an AI system misclassified posts from people of color as hate speech due to differences in regional dialects. This led to the AI system banning or removing a higher number of people of color from the platform, indicating a bias present in the AI. In another study, Thomas & Thomson (2023) evaluated a generative AI application's ability to generate an image of a journalist. They identified many biases in the AI output, including ageism, sexism, classism, and racial bias. Despite changing the user input to try to mitigate some biases, the generative AI program still produced a biased output (Thomas & Thomson, 2023).
Due to the prevalence of bias in AI algorithms, recent research has been focused on ways to mitigate or reduce bias. Roselli et al. (2019) described many ways to mitigate biases in AI, including vetting the training data before implementing it in the AI system, monitoring the production data and output, and changing the input and training data appropriately based on biases present in the output. Utilizing explainability in bias detection is important for making the decision-making process of the AI outcome explainable to the human user and AI developers (Roselli et al., 2019; Mosqueria-Rey et al., 2023). Human-in-the-loop systems also can help in bias detection by having a human interact with the AI output "behind the curtain" to mitigate bias (Mosqueria-Rey et al., 2023).
Bias detection and mitigation are important, but issues can arise where bias mitigation goes too far. Oversteps in bias mitigation can increase other biases that were not originally present in the data before the data was manipulated to mitigate one bias (Safarlou et al., 2023). For example, if a dataset is over-manipulated to mitigate racial bias found in the data, a gender bias may then emerge in the dataset that was not present prior to the manipulation of training data. Understanding how to mitigate bias in an appropriate way is important to make the AI system output as accurate as possible (Safarlou et al., 2023)
9.5 Governance, licensure
9.6 Designing approaches for trustworthy AI
9.7 Conclusions
This section is not mandatory, but can be added to the manuscript if the discussion is unusually long or complex.
Acknowledgments
This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.
Conflicts of Interest
The authors declare no conflict of interest.
AI Usage Statement
Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation. Content and ideas are otherwise original to the human author contributors.