Chapter 11
Large Language Models and Prompt Creation
Abstract. The abstract should be a total of about 200 words maximum. The abstract should be a single paragraph and should follow the style of structured abstracts, but without headings: 1) Background: Place the question addressed in a broad
11.1 Introduction
Maybe link Ch. 2
11.2 An overview of the Transformer Architecture and Generative Pretrained Transformers (GPTs)
Modern Large-Language Models (LLMs) often take the form of a Transformer. This type of architecture is used by neural networks and is well-suited for natural language processing (Transformer, 2017). A Transformer model builds on the concept of Sequence-to-Sequence learning (Sutskever et al., 2014), where both the inputs and outputs of a model can be an entire sequence of words instead of a single input or output. Combining sequence-to-sequence learning with a Recurrent Neural Network (RNN) allowed for machine translation to be performed. An RNN would use an encoder to process the entire input sequence, generate a series of vectors (the hidden layers within the encoder), and pass them to a decoder. The decoder can then use these vectors to produce an output sequence one word at a time, with the vectors providing a context for how the input sequence words line up with the output sequence words (Alammar, 2018a). This is a mechanism called attention.
A Transformer model also uses an encoder-decoder architecture as well as attention (Vaswani et al., 2023). However, unlike the attention used in an RNN, a Transformer uses self-attention. The encoder of a Transformer will still generate vectors, in this case by using a feed-forward neural network, but the encoder also has a self-attention layer preceding the feed-forward network. This self-attention layer creates a representation of the entire input sequence that, in part, contains the connection of each word in the sequence to all the other words in the sequence (Alammar, 2018b). This allows for connections like those between the "it" and "rock" in the sentence "The person couldn't move the rock because it was heavy" to be embedded into the vector representation. The decoder can then use this relationship embedding when generating one word at a time to maintain relationships between connected words (Transformer, 2017).
Generative Pretrained Transformers (GPTs) are Transformer models that combine both unsupervised and supervised learning to teach the model broader natural language understanding (Radford et al., n.d.). GPTs undergo unsupervised training on very large amounts of data, such as GPT-1 by OpenAI being trained on the BookCorpus[^1] dataset that contains over 7,000 books, to give the model a broad understanding of natural language as a whole. Such general understanding can even lead to the model performing tasks it has not specifically been trained to perform (OpenAI, 2018). However, training the model further on a specific task can lead to increased performance on that task beyond that of general understanding alone. This further training is done through a supervised training process called fine-tuning, where the model is given labeled examples of a specific task to learn how to perform that task in a greater capacity than only a general understanding would allow (OpenAI, 2018; Radford et al., n.d.).
11.3 Limitations of GPT models
11.4 Prompt Engineering
Prompt engineering involves systematically generating and refining prompts to produce targeted responses from Large Language Models (LLMs) (Aljanabi et al., 2023). Key principles of effective prompt engineering include providing clarity, specificity, and context so that LLMs understand the instructions and then produce the intended output. Incorporating illustrative examples enhances clarity, especially in creative tasks like poetry or scriptwriting (Aljanabi et al., 2023). Prompt engineering is iterative, requiring ongoing refinement to optimize communication between humans and machines.
Whether we are aware of it or not, humans use many types of prompts when interacting with LLMs, from zero-shot to few-shot to chain-of-thought prompts, to name a few. Each of these types ask for and retrieve different information. Zero-shot prompting is simply asking an AI model to provide information or perform a task without the operator (?) first providing instructions or examples (Prompt Engineering Guide). Few-shot prompting is perhaps the most common approach, which includes prompting the model with several input–output exemplars demonstrating the task (Wei et al., 2023). Chain-of-thought prompting can be successful in completing more complex tasks where reasoning is required. Chain-of-thought uses several intermediate reasoning steps to complete the task.
In everyday life, few-prompt strategies are sufficient, but to bring even more accuracy and a better fit to projects, using a template like the CO-STAR method substantially improves the quality of LLM responses. CO-STAR is a template developed by GovTech Singapore's Data Science & AI team for setting up prompts that considers (C ) Context, (O) Objective, (S) Style, (T) Tone, (A) Audience, and (R ) Response (Teo, 2023) when delivering the output. This structured approach ensures that prompts are clear, specific, and aligned with the intended goals, facilitating better communication with LLMs. Moreover, the standardized nature of CO-STAR templates promotes consistency across different tasks and enables easier evaluation of LLM responses (Teo, 2023). Additionally, CO-STAR empowers prompt creators to provide comprehensive instructions and guidelines to LLMs, leading to more efficient collaboration and task execution. Overall, incorporating templates like CO-STAR into prompt engineering processes can significantly enhance the quality and efficiency of interactions with LLMs, ultimately improving outcomes.
11.5 LLMs as coding assistants
Acknowledgments
This chapter was originally developed as part of a course project for HF 5430 Human-AI Interaction, Michigan Technological University, Spring 2024.
Conflicts of Interest
The authors declare no conflict of interest.
AI Usage Statement
Generative AI models were used for additional research, to identify missing concepts, to support better organization, and for editorial tasks such as formatting, evaluating grammar/clarity and citation collation. Content and ideas are otherwise original to the human author contributors.