Handbook

Chapter 10

Cognitive Tutorials and User Documentation

Abstract. The abstract should be a total of about 200 words maximum. The abstract should be a single paragraph and should follow the style of structured abstracts, but without headings: 1) Background: Place the question addressed in a broad

10.1 Importance of Documentation and User Support

Learning something new is not an easy task, especially when it comes to new technology. Challenging menus, instructions, and displays make learning new technology even more difficult, which is why documentation and user support are critical elements of implementing new technology (Schneiderman, 2016). User guides are assistive instructions that are provided to users of a product or system to help them navigate any challenges or learning curves associated with the novelty of the product to avoid frustration (Schneiderman, 2016). There are many formats of user guides and assistive manuals that can be created for the user's benefit and ease.

  • Paper Documentation: Paper documentation is the traditional form of reference material for new products and systems (Schneiderman, 2016). This type of documentation is typically in the form of a physical manual or reference guide that is given to the user when a product is purchased. Paper documentation is useful, as it is accessible to everyone who has purchased the product or system, however, it is not the easiest form of reference to quickly navigate if looking for help in one specific area (van der Meij, 1996).
  • Online Help: Online help is another commonly used form of user support guides and documentation. Instead of a physical manual given to each customer, the manual is online for all customers to access (Schneiderman, 2016). Examples of online help include help menus and frequently asked question (FAQ) pages. While this form of user support is easy to navigate, it may not contain all the necessary information the user may need if dealing with an uncommon issue (Schneiderman, 2016).
  • Online Tutorials: Online tutorials provide a walkthrough of how to use a new system or product. These range from online video tutorials to interactive training tutorials where users click through a tutorial within the system explaining key functions. These tutorials are beneficial to users who want a brief explanation of how to use a system or product, but may not always be useful for users who need to troubleshoot an issue that is occurring or need more assistance than a brief tutorial (Schneiderman, 2016).
  • Online Communities: Online communities and forums are a way to provide user support through other users (Schneiderman, 2016). These online communities sometimes exist for one specific product, but may also be a general question and answer website that can be used for a variety of subjects. Users are able to post questions or issues they are having with the product or system and other users can reply with assistance or troubleshooting information. This format of user support is beneficial to users experiencing uncommon problems that may not be addressed in other forms of documentation by connecting them with people who have experienced the same or similar issues before. However, it is not always easy to tell if the troubleshooting information provided in the online community is from a reliable source or is accurate, which could lead to further issues or damage to the product or system (Schneiderman, 2016).
  • Virtual Assistance: Virtual assistance is a newer form of user support that allows the user to chat with an AI assistant to troubleshoot or learn to use a system. The AI assistants are good for asking simple questions or navigating a form of documentation such as an online help menu, but may not have the ability to answer questions regarding specific issues the user is having with the product or system (Zaidi et al., 2021).

Once the format of the documentation and user guide has been selected and created, it is important to make sure that all users are capable of using the help guides. Ensuring that the documentation is easily understandable, meaning it is able to be used by users of different reading comprehension levels, and is appropriate for all users (i.e. readable font and sizing, appropriate use of color, etc.) is critical in the formation of user documentation and support (Schneiderman, 2016). It is also important to make sure the documentation format is logical, meaning that the path the user takes to get the information makes sense for the product or system (Schneiderman, 2016). Another necessary aspect of user support documentation is to keep the support guide updated along with updates in the product or system, especially when it comes to technology. Technology is constantly advancing, so it is important to make sure support guides are being updated appropriately (Schneiderman, 2016).

Tutorial and User Guide Best Practices (BW)

Tutorials are meant to inform users of how they can interact with a system/program. These guides not only explain the necessary parts of a program but detail the different commands the system responds to and what the system does with those inputs. Good tutorials tend to use a variety of descriptions to make sure that users get the necessary information out of them. Some ways that tutorials inform users are: example situations, controller diagrams, descriptions of various components or full worked examples. While there are many ways that tutorials can be made, a poorly designed tutorial will confuse users and make learning a new system harder than it needs to be.

One way that tutorials fail to inform users is by not optimizing the amount of information that is necessary for users to understand the system. If a guide is too short, it may not fully encapsulate all the essential tools that a user should understand before starting the program. Providing users with an appropriate amount of information likely means limiting prose and includes multiple forms of communication (ie. bullet points, figures, tables). By using multiple types of media the user isn't overloaded while dealing with a tutorial (Ibili & Billinghurst, 2019). This means when one form of information is overused, lengthy text as an example, users might get bored or have issues remembering all of the information from the tutorial. Finding an appropriate length for tutorials depends on context, something like a controller for a video game might not require a long tutorial, but flight controllers may need something far longer. Experienced user guides might require more complicated tutorials to effectively support users navigating AI (Mueller, Klein & Burns, 2009). Some ways these guides interact with users is outlined in chapter 2.

An effective way that tutorials help users is through showing how something works, rather than explaining it. When possible, examples and figures can help limit text and give visuals to provide a quick reference guide to understand what should happen while interacting with the system. When a system contains many components, well organized and limited information might be more effective to inform users (Klein, Mueller & Rasmussen, 2010). User guides that contain a lot of clutter would make retaining the information in them difficult and would likely confuse users.

Strong tutorials provide users with information that inform about the essential parts of a program while not overloading the user with information. Finding the right amount of information to display is highly dependent on the complexity of the program. Usually, shorter guides with lots of examples and figures are most effective at teaching users quickly. Quick information to understand the important components of a program can be incredibly useful before actually experimenting with it; see later in this chapter for example model cards.

10.2 Training Humans for Trained Machines: Updating the Human Mental Models of AI

Shane, let me know if you don't think I need to write this section.

10.3 Domain Applications for Model Cards

Model cards are typically used in the domain of artificial intelligence and machine learning, particularly in the context of deploying AI models. Here are some common domain applications for model cards:

  • Data Science and Machine Learning: Model cards are used by data scientists and machine learning engineers to document and communicate essential information about AI models they develop. This includes details about the dataset used for training, model architecture, hyperparameters, evaluation metrics, and potential biases.
  • AI Ethics and Responsible AI: Model cards play a crucial role in promoting transparency, fairness, and accountability in AI systems. They help stakeholders understand the ethical implications and potential biases associated with deploying a particular model in real-world applications.
  • Regulatory Compliance: In domains where regulatory compliance is essential, such as healthcare, finance, and autonomous vehicles, model cards can serve as documentation to demonstrate compliance with regulations and standards related to AI transparency and fairness.
  • Product Development and Deployment: Tech companies and organizations integrate model cards into their product development and deployment processes to provide consumers, clients, and regulators with comprehensive information about AI models embedded in their products or services.
  • Academic Research: Researchers use model cards to document their AI models and experiments, facilitating reproducibility and enabling other researchers to build upon their work. Model cards also help in peer review by providing reviewers with detailed information about the model and its associated datasets.
  • Education and Training: Model cards can be valuable educational resources for teaching concepts related to AI ethics, fairness, transparency, and responsible AI development. They help students understand the multifaceted aspects of deploying AI models in real-world scenarios.

Overall, model cards serve as a standardized way to communicate essential information about AI models, promoting transparency, accountability, and trust in AI systems across various domains and applications.

10.4 Worked Examples of Rule Cards

Mitchell et al. (2019) introduced the concept of Model Cards as short references for users of trained machine learning models. Indeed, the Model Card defines who the intended users for a particular model are, as seen by the developers of the model. Figure X shows a hypothetical Model Card for the Copilot programming assistant, though created by an end user.

Figure placeholder

Save image as: figures/fig-ch10-01-a-sample-model-card-for.png

Export from Google Drive as PNG or SVG, place the file in figures/, then rebuild.

Figure 10.1. A Sample Model Card for GitHub Copilot.

This Model Card shows how some users may misunderstand the difference between an AI tool and the model that powers it. Copilot is the marketed tool, as well as its integration with various integrated development environments (IDEs), but the trained model powering Copilot is the OpenAI Codex transformer model. One difficulty in filling out this card was the Factors section, which is admirably aimed at describing how the model's training may affect its performance in intersectional cultural contexts. When the model is trained on computer code and is primarily for the generation of computer code, many of those intersectional factors are less relevant. Simultaneously, the Model Card concisely informs novice users and programming students that the model is not a substitute for learning to program or solving a problem without further audits of the code. Some aspects of the canonical Model Card are difficult to include (e.g., false positive rates) if the card is created by an outside source.

QUESTION TO SHANE AND BRANDON: Do we need another Model Card example here?

10.4.1 Model card for GPT4-V

The model card provides an overview of the GPT-4V (ision) model developed by OpenAI. GPT-4V is an extension of the GPT-4 language model that incorporates vision capabilities, allowing it to process and interpret visual inputs such as photographs, screenshots, and documents. It was launched by OpenAI in 2023 and is based on the GPT-4 model.

The development process for GPT-4V involved two key steps. First, a transformer-based model was pre-trained on a large dataset of text and image data from the internet as well as licensed sources. This pre-training stage focused on training the model to predict the next word in a document. Subsequently, the model underwent fine-tuning using reinforcement learning from human feedback, a technique known as RLHF (Reinforcement Learning from Human Feedback).

One of the key capabilities of GPT-4V is its ability to handle various types of visual inputs, including photographs, screenshots, and documents. It can perform tasks such as object detection and analysis within images, identifying and interpreting objects present in the visual data. Additionally, GPT-4V excels at data analysis, allowing it to interpret and analyze graphs, charts, and other data visualizations. Another notable feature is its text deciphering ability, enabling it to read and interpret handwritten notes and text within images.

While GPT-4V offers powerful vision capabilities, the model card also highlights several key risk areas that users should be aware of. These include potential issues with scientific proficiency, such as image-text combination problems, hallucinations, and factual errors. Medical advice provided by the model may be vulnerable to inaccuracies and misinterpretations of imaging data. The model may also exhibit stereotyping and ungrounded inferences, leading to biases, open-ended questions, and unintended topics.

Furthermore, the model card cautions about disinformation risks associated with image-text pairing, tailored content, and potential disinformation. It also warns about the possibility of hateful content, such as content related to the Templar Cross or hate groups, as well as praise for such content. Finally, the model card highlights visual vulnerabilities, including potential issues with open-source code, corporate styles, and image ordering.

Overall, the model card serves as a comprehensive overview of the GPT-4V model, outlining its vision capabilities, development process, key features, and potential risk areas, providing users with valuable information to consider when utilizing or deploying the model.

Explainability fact sheets:

https://dl.acm.org/doi/10.1145/3351095.3372870

10.5 Collaborative Explanation & SQA

Social Q&A (SQA) refers to the process of people asking questions, answering questions, and rating the content of other answers (Gazan, 2011). Two examples of this kind of system are Stack Exchange and Stack Overflow. These systems depend on the collective knowledge of the user base to answer specific questions of the user base. The ranking ability that these websites have allow for users and the community to decide what the best or most useful answers are. The community can use the ranking system to control the information that is seen and to attempt to moderate answers to some extent. Wang et al. (2013) argue that SQA sites are a method of crowdsourcing information for specific information and use the knowledge and experiences gained by others. According to Harper et al. (2010) there are 6 major types of questions asked on SQA sites. These types are: advice; identification; (dis)approval; quality; perspective; and factual. All of these question types could be used in SQA for AI systems.

Bao (2018) suggests that SQA is dependent on social cognitive theory. Social cognitive theory refers to the idea that personal cognition in a social environment can shape and control their behavior and a dynamic and reciprocal interaction of personal cognition, environmental elements, and human behavior is created (Bandura, 1986). In this sense, all people within the system contribute to and shape the environment with each question or answer they provide. The community is able to create meaningful answers partly because they all agreed to. All members of the community understand that they will not be judged for their personal actions, but rather through the answers that they provide (Bao, 2018). The community informs the way that each individual is expected to interact with the questions and what methods of communication are appropriate.

When studying the effectiveness of using a SQA system for explainable AI, Mueller et al. (2021) showed that SQA can be used well for explaining AI. Users were given two different methods of searching for answers, using the SQA interface or searching a database of information. Users were able to find information that they desired faster than if they just searched the database. Users rated the system as more satisfying, sufficient, complete, and trustworthy as compared to an example base explanation. This shows that SQA can be a useful and productive method of explanation that users are already familiar with. The use of this method relies on everyone involved, removing some of the burden from the AI system's developers. In addition, the explanations may be more sensitive to what a user wants to know since other users provide the questions and answers. The questions can also be more generalized about AI, rather than specific to the particular system that is currently being used.

Acknowledgments

This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.

Conflicts of Interest

The authors declare no conflict of interest.

AI Usage Statement

Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation. Content and ideas are otherwise original to the human author contributors.