Chapter 8
Explainable and Interpretable AI: Human-Centered Aspects
Abstract. [Abstract pending in source draft.]
8.1 Introduction
8.2 What Can Explanations Provide
2.1 Improving Transparency
Explainable AI works to show what is happening when an AI system does something. Creating human-understandable explanations can improve transparency for the users (Das & Rad, 2020). This allows for the user to use the AI more productively and prevent bad decisions and conclusions. When the user understands what is going on within the AI system, they are better prepared to use the AI. An explanation of what transparency is and why it matters can be found in Chapter 7.
8.2.1 Improving Trust
Trust in the system is important to the user and their understanding of how and when to use the system. Explainability can help to form this trusting relationship (Shin, 2021). When the user understands what the system is doing or how it came to a specific conclusion, it is easier to trust. The user has something that allows them to form trust with this model beyond trial and error or pure faith alone. The presentation of the explanation can also help to form a trusting relationship. When the presentation is based on the user's knowledge, the user may trust the information more. They are able to form a connection with it since it provides information that they can understand. For more information on trust see Chapters 5 and 6.
8.2.2 Improving Understanding of Shortcomings
Understanding what an AI system can and can't do is an important feature of use. Explainable AI can help users understand what the system can't do (Adadi & Berrada, 2018). This allows the user to have enhanced control when something has gone wrong or not worked as expected. The explanation can help to describe why the system may have been unable to do what the user expected. This allows the user to redefine the way they view the system and how they interact with it. Additionally, the model could better improve because you are interacting with it in a way that it can use and understand (Adadi & Berrada, 2018). The system is better able to fit into the expectations of the user since the user is informed about the abilities the system possesses.
8.2.3 Limitations
O'Hara (2020) argues that no AI system can be truly and entirely explainable because the system makes connections that the user is unable to understand. Due to this, O'Hara states that some portion of the system is based solely on trust and that fact can't be changed. Even when you try to create better explanations, some portion of the system will stay obscured to the user. AI doesn't, and shouldn't, make the decisions, but rather be used as a decision aid for the human user. There are some AI systems that do not need human interaction to work, but O'Hara argues that the user is still making the decision to not interact with it.
Another limitation that explanations have is the specificity of the language that they use within the explanations (Yang et al., 2023). Some explanations are too generic, not providing enough information for the user to truly understand the system to the extent that they wish. On the other hand, some explanations use jargon that is too domain specific or too high level that everyday users may not understand. According to [citation] it is extremely difficult and possibly impossible to create an explanation that would work for all users. Explainable AI needs to be based around the specific audience that will be using the AI system. [citation] further explains that most metrics for explanations do not include the subjective experience of using the explanation is rarely included. Rather, explanations are judged by how accurate the explanation is, which won't help to understand how well the user understands the material.
Yang et al. (2023) also explains another limitation within the long term use of explanations. As the AI system is updated and changed, the explanation must also be updated and changed. If the explanation is too stagnant or can't be updated, the explanation will become completely obsolete. In the worst cases, an outdated explanation may confuse the user about how it works or give them factually incorrect information. For the users that understand that this is wrong, it will break down user trust and make the explanation decrease value rather than add it.
3. Who are these for? (BW) When to Use Explainable AI
8.3 White Boxes Versus Black Boxes (Interpretable Models versus PostHoc Explanations)
Interpretable AI, or White Box AI, generally refers to systems where the decision-making process is both transparent and explainable. White Box models are considered to be intrinsically interpretable and one can explain how they make decisions, what the influencing variables are, and how they produce predictions. White-box systems can be examined using interpretable analysis as opposed to post-hoc analysis. Interpretable analysis examines the model transparency, weighs feature importance (SHAP), and plots partial dependence (Ma et al., 2023). It highlights the contribution of each input variable to the output, increasing explainability. Linear regression and decision trees are common examples of white-box models as they provide less predictive capacity but are considered highly interpretable. White-box systems can often be interpreted by laypersons as the output may be in the form of visualizations, natural language explanations, or simplified rule sets (Hendricks et al., 2016; Liu et al., 2020). In addition to having less predictive capacity, highly explainable and interpretable systems are often less accurate. These types of models are the counterpoint to black-box models which prioritize accuracy over explainability and transparency. With both white-box systems and black-box systems, however, the performance of the system depends on the application domain and input data (Loyola-Gonzalez, 2019).
Black-box models such as neural networks, deep-learning models, or complex ensembles produce highly accurate output but users are unable to understand or explain how the model produced the output. When experts in practical applications such as medical professionals, financial experts, and the military interact with these models and may not be able to interpret the output, accurate or not. Support vector machine (SVM) models are utilized widely in author profiling, but the mathematical operations that support them are difficult to understand for both machine learning experts as well those extracting the demographic information (Loyola-Gonzalex, 2019). Since these and other black-box models are difficult to interpret and explain interpretable analysis cannot be utilized. Using post-hoc analysis can explain the behavior of these models. Behavior and sensitivity analysis can be used to examine the responses to various inputs even when the changes to input data are insignificant (Borgonovo & Pilschke, 2016). Surrogate models can also be examined to approximate the behavior of the black-box model (Kuttichira et al., 2019). These aspects of post-hoc analysis of black-box models can improve explainability and interpretability but have limitations in fully explaining the decision boundaries or correctly capturing the interactions between features. It is important also to remember that any explanation generated by a post-hoc analysis should be considered an approximation rather than an explicit representation of the model.
8.4 Approaches to Explanations
Shane:
- Local and Global explanations
- Examples
- Comparisons
- contrasts and counterfactuals
- Highlighting
- Self-explanation scorecard
- PDP/ALE
8.5 Examples
8.5.1 Accumulated Local Effects
One example of an automated explanation providing global explanations based on features is the Accumulated Local Effects (ALE) plot. Consider a machine learning model that attempts to predict if a patient may develop diabetes. The model will be trained on certain features (e.g., blood glucose levels) and patient outcomes (a diabetes diagnosis or not). If the modeler wants to understand how each individual feature contributes to the prediction, ALE examines the importance of each feature in an algorithmic fashion and plots them. Since this does not give an individual prediction or outcome, it is a global explanation of the model. Another type of analysis in this space is a Partial Dependence Plot (PDP). The main problem with PDP occurs when features are highly correlated. For example, people may tend to buy more ice cream when the outside temperature is warmer. If a beach is trying to ascertain what times there are more lifeguard rescues, it may be hard to discern the influence between these two features and erroneously conclude that people who bought ice cream on a given day are more likely to require rescuing. ALE plots, due to the method of computation, are more able to discriminate between these features. Largely speaking, ALE plots are interpretable by any layperson, provided that they understand the real-world details of the feature shown in the plot. Shown in Figure X is a single ALE plot where a simple neural network was built from an open-source diabetes dataset.
The plot shows that as insulin levels increase (along the horizontal axis), the prevalence of a diabetes diagnosis increases (along the vertical axis). Since we know that a feature of diabetes is elevated insulin production due to insulin-resistant cells, this makes sense. An individual can find where they fit on the plot with relative ease, but their local information is not included by default. The hash marks at the bottom of the plot indicate the amount of data that led to the prediction. In Figure X, there are many data points for insulin levels less than about 300. There are very few past 500. The model should be considered more reliable for the ranges on which it was trained. ALE plots comparing all of the features can easily be generated to investigate the individual feature effects, as shown in Figure Y.
8.5.2 Explainable Neural Networks*
Apart from data scientists and programmers, neural networks (NNs) are not informative for the typical end-user. Explainable network nodes can provide some insight for the data scientists to make decisions about which nodes are important and how many layers would be appropriate given the problem space. These nodes can be informative to those that understand how the NNs function and be almost zero help to the end user trying to make sense of the output.
Image classification and NNs might be beneficial for some users who specialize in identifying abnormalities (ex. radiologists, doctors, quality control technician). These explainable systems might provide extra evidence and draw attention to regions where an issue occurs and provide some form of reasoning. For example, an explainable neural network could provide some reasoning to a radiologist where it detected cancerous cells based on classifications throughout the layers of the network. One such NN was able to identify regions of damage due to Covid-19 in patient lungs through x-ray image classification (Angelov & Soares, 2020). If NNs can provide reasoning and steps to identify problems, then they might provide some valuable information to specialized end-users who can read the data output.
Alternative to direct output, a post-analysis explanation of NNs might benefit some users. If the nodes of a network can be explained as a binary decision, then the network might be explainable for end-users. NNs can become more explainable to describe what happens at the node- and layer- levels with language end-users can understand. Blazek and Lin (2021) used this approach to make an explainable neural network that classifies complex shapes. Work with explainable NNs is relatively new and examples are limited for systems that would help end-users. Currently, neural networks are explainable almost exclusively to data scientists familiar with these systems.
8.5.3 6.X Local Interpretable Model-Agnostic Explanation (LIME) ?
A local surrogate model is a type of model that works with a different, black-box model and is trained to approximate the predictions of that black-box model (Molnar, 2023). This surrogate model is an interpretable model, such as a decision tree, and can be used to make good local predictions analogous to those of the black-box model to help shed insight into what the black-box model is doing. The idea of local predictions means that the model will make good predictions when working on a task that is similar to the one the black-box model is working (Ribeiro et al., 2016). If you take an interpretable model made by LIME and try to apply it to a problem outside of the target domain, there is not a guarantee that it will generate as good of predictions as those made in the target domain.
LIME models can be used on various forms of data, including tabular, text, and image data (Molnar, 2023). For tabular data, a LIME model could be used to examine the different features used by a model and generate a visualization of the impact they have on a certain predictor, such as what features will have have an impact on the amount of bikes rented on a given day. For text data, LIME will perform a similar task to determine which words are impacting the prediction. Let's say we want to predict whether a comment on social media is spam or genuine. The text of the comment can be taken and variations can be made, with different variations having only a subset of the original words. Each of these variations can then be used in the prediction, with the impact of each word on the prediction becoming measurable across the variations to provide the most likely words that lead to a spam classification. For image data, variations are also generated, but instead of individual pixels being masked spans of pixels with similar coloring are replaced with a different color to measure the impact that region has on the prediction. This might not be able to explain why those spans of pixels are impacting the prediction, but LIME can at least be used for identification purposes with image data.
8.5.4 Shapley Values and the SHAP Algorithm
Shapley values come from game theory and have been used in machine learning to measure feature importance (Molnar, 2023). In the machine learning domain, the game to perform is the prediction of the model and the features are the players of the game, with each feature having an impact on the outcome of the game. This impact can be calculated by taking one data point and calculating the prediction for this data point. Another data point is then taken at random and each feature of the original data point is swapped for the value from the randomly selected data point and the prediction is run again. The difference in the predicted value is measured for each feature, and these differences are averaged to calculate the Shapley value (Molnar, 2023).
The Shapley Additive Explanation (SHAP) algorithm is an algorithm that generates values that attribute to each feature the change in the expected model prediction when conditioning on that feature (Lundberg & Lee, 2017), or, more plainly, their Shapley value. This algorithm will generate a visualization of the importance of each feature, such as a bar graph, with a more important feature being shown with a larger bar in the bar graph. Figure A shows a visualization generated on a California housing price dataset, showing that the location of the house had the biggest impact on the house price. This feature importance is represented using a linear model, which means that the SHAP algorithm can also be used with LIME to make a connection between the generated Shapley values and the explainability of those values using LIME (Molnar, 2023).
6.X. GAMs (EM)
8.5.5 Example-based explanation
Example-based explanation methods select particular examples of the dataset to explain the behavior of machine learning models rather than creating summaries of features (such as feature importance or partial dependence). Representing an instance of data in a human-understandable way is crucial for example-based explanations (Molnar, 2020). Because an instance may consist of hundreds or thousands of (less structured) features, listing all feature values to describe an instance is usually impractical. When there are only a few features or we can effectively summarize the instance effectively, example-based explanations work well.
Example-based explanations help in building mental models of machine learning models and the data the machine learning model has been trained on. For example, a doctor was seeing a patient who was coughing and had a fever. The symptoms reminded her of previous similar examples. Thus, she inferred this patient may have the same disease. She did a blood sample to test for him. The blueprint for example-based explanations is: if thing B is similar to thing A and A leads to Y, then we predict that B will also lead to Y. Implicitly, some machine learning methods are based on examples.
Figure placeholder
Save image as: figures/fig-ch08-01-example-based-explanations-for-miscl.png
Export from Google Drive as PNG or SVG, place the file in figures/, then rebuild.
Sometimes AI makes classification errors, like misclassifying bird images as airplanes (Sigler, 2022). Using example-based explanation to retrieve other images from the training data that are most similar to the misclassified bird image in the latent space. Analyzing these similar images reveals that both misclassified bird images and similar images were dark silhouettes. To understand the reason, the study expanded the similar example search and showed the 20 nearest neighbors. Among them, 15 were airplane images, and 5 were bird images. The results indicated a lack of bird images with dark silhouettes in the AI training data because only one of the 5 bird images had a dark silhouette. Therefore, improving the model can be achieved by collecting more data with bird silhouette images.
8.5.6 Counterfactual Explanations
Counterfactual explanations are ones that create potential scenarios based on a small change that produces a different outcome. They produce a "what if?" explanation to show what the outcome of a scenario could be if one aspect, mainly the input, of the scenario were different (Verma et al., 2020).
A very common example of a counterfactual explanation is when applying for and being denied a loan (Verma et al., 2020). If someone is denied a loan, they want to know why, especially what factors are contributing most to the decision. If their credit score was higher, they had a different zip code, had higher income, would their outcome be different? A counterfactual explanation would provide them with information about what aspects of their information are affecting the decision, such as "Your credit score is too low for approval. The average credit score for loan approval is X. If you improve your credit score to X, you can get the loan". This explains why a person may not have gotten approval, shows which factor is causing the issue, and what the range is to be approved.
Counterfactual explanations are typically used to explain the outcome to the user, such as in the loan example. However, there is the possibility of using counterfactual explanations to predict human behavior by examining all possible changes. According to the MIT Technology Review, Spotify has begun to use counterfactual analysis and explanations to predict user listening behavior to provide music recommendations (Heaven, 2023). Instead of providing an explanation to the user about what change in the input could lead to a different outcome, Spotify is providing developers and algorithms with explanations of how user listening behavior could change to narrow down personal recommendations on the platform (Heaven, 2023).
8.5.7 Explainability/Interpretability for LLMs
Large language models (LLMs) refer to transformer-based neural language models that contain tens to hundreds of billions of parameters, and which are pre-trained on massive text data (Singh et al. 2024), e.g., PaLM24, LLaMA12, and GPT-4. While they have shown impressive capabilities in understanding and generating human-like texts, LLMs are notoriously complex "black-box" systems. This is because their inner working mechanisms are opaque, and the high complexity makes model interpretation much more challenging.
Why is it important to improve the explanation for LLMs? By elucidating the decision-making process, stakeholders can verify that the model operates within the bounds of privacy, fairness, and non-discrimination principles, thus mitigating potential legal and ethical risks (Singh et al. 2024).
- Enhancing trust and accountability: Clear explanations of LLM outputs improve user understanding and trust. This transparency fosters accountability by allowing stakeholders to assess the model's ethical alignment.
- Error diagnosis and improvement: Refined explanations help identify model weaknesses, biases, and errors, enabling targeted improvements to enhance accuracy and reliability.
- User understanding and acceptance: Clear explanations empower users to grasp the model's behavior, fostering acceptance and effective utilization.
- Safety and robustness: Prioritizing explanation enhances model resilience by enabling safeguards against harmful outputs and adversarial attacks.
- Legal and ethical compliance: Interpretable outputs aid in assessing model adherence to regulations and ethical guidelines, mitigating potential risks.
Understanding Large Language Models (LLMs) presents significant opportunities through the integration of natural-language interfaces and interactive explanations. Leveraging natural language, a medium familiar to humans, offers a promising avenue for elucidating complex patterns inherent in LLM outputs. This approach serves as a conduit for bridging understanding across diverse modalities, including DNA sequences, chemical compounds, and images. Additionally, interactive explanations provide a dynamic means for users to engage with LLMs, enabling tailored insights to meet individual needs. Through dialogue and analysis, users can delve deeper into explanations, fostering a collaborative exploration of complex concepts and enhancing comprehension.
However, these opportunities are countered by formidable challenges that necessitate careful consideration. Chief among these challenges is the phenomenon of hallucination, wherein LLMs may generate incorrect or baseless explanations, leading to potential confusion or misinformation. Addressing hallucination is paramount to ensuring the reliability and trustworthiness of LLM interpretations. Furthermore, the sheer immensity and opaqueness of LLMs, characterized by models containing tens or hundreds of billions of parameters, present significant hurdles to interpretation efforts. Understanding the inner workings of these complex systems requires innovative approaches to distill essential insights and present them in a comprehensible format, thereby facilitating understanding and trust in LLM outputs.
Zhao et al., (2024) categorize techniques based on the training paradigms of LLMs: traditional fine-tuning-based paradigm and prompting-based paradigm. The Traditional Fine-Tuning Paradigm involves initially pre-training a language model on unlabeled text data, followed by fine-tuning it on labeled data from a specific domain. During fine-tuning, additional fully connected layers are appended above the final encoder layer, facilitating adaptation to various tasks. This approach has demonstrated success primarily with medium-sized models like BERT or RoBERTa, typically containing up to one billion parameters. Explanations within this paradigm predominantly revolve around comprehending how self-supervised pre-training enables models to grasp language fundamentals and how subsequent fine-tuning equips them to effectively address specific tasks.
Conversely, the Prompting Paradigm employs prompts, such as natural language sentences with blanks, for zero-shot or few-shot learning without requiring additional training data. Models within this paradigm are categorized into Base and Assistant Models. As LLMs increase in size and training data, they exhibit notable advancements, including the ability for few-shot learning through prompting. Explanations concerning base models center on understanding how these models utilize their pre-trained knowledge in response to prompts, although they often encounter difficulties in following user instructions and may generate biased or toxic content. To mitigate these shortcomings, base models undergo supervised fine-tuning to attain human-level abilities, such as engaging in open-domain dialogue. Explanations in this context concentrate on elucidating how models acquire interactive behaviors from conversations, with the objective of aligning responses with human feedback and preferences.
8.6 Conclusions
This section is not mandatory, but can be added to the manuscript if the discussion is unusually long or complex.
Acknowledgments
This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.
Conflicts of Interest
The authors declare no conflict of interest.
AI Usage Statement
Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation. Content and ideas are otherwise original to the human author contributors.