Understanding AI
It is easy to mistake a good answer for a good understanding. A language model can explain an unfamiliar idea beautifully and then cite a paper that does not exist. To use one well, it helps to know what is happening between the question and the reply.
This is a guide to text-generating language models, not every form of AI. I’ll follow the path from text and training to a finished answer, then look at what makes that answer worth trusting.
Start with the right picture
Ordinary software follows rules written by people. Machine learning also runs as software, but many of its useful rules are learned from examples. In a language model, a large number of numerical settings, called parameters, shape which text is likely to come next. The developer chooses the architecture and training process; the examples shape those settings.
This does not mean the model is a person with an internal library of facts. Nor does it mean its output is random. Given an input, it computes scores for possible continuations. The application then chooses a continuation according to its generation settings. Even a repeatable answer can be false.
The model sees pieces, not words
Before text reaches the network, a tokenizer splits it into tokens. A token may be a word, part of a word, punctuation or a single character. The split depends on the model and its encoding. Each piece becomes a number, and that number is mapped to a learned vector: a list of values the network can work with.1
This is why counting words is a poor way to estimate how much room a prompt will need. The same sentence can use different numbers of tokens in different models. A misspelling, a rare name or a code fragment may also be divided in surprising ways.
01 / FROM TEXT TO TOKENS
The shipment arrived.
Each token receives an ID, then a learned numerical representation.
Where the ability comes from
A neural network is a stack of mathematical operations. Early layers turn token vectors into new representations; later layers transform them again. A layer combines its inputs using learned weights and a non-linear operation. The result is passed on. The word deep simply means there are many such layers.
That sounds mechanical because it is. The surprising part is what the system can learn when these operations are repeated at scale. A parameter is not a labelled entry for “bicycle” or “good answer.” Useful behaviour comes from the combined effect of many parameters across many layers.
02 / A NETWORK MAKES A PREDICTION
For many text models, the first large training stage is pretraining. The model sees a sequence and tries to predict the next token. It is shown the actual next token and is scored on how much probability it gave that token. This objective has produced systems that can work across many language tasks, even without a separate set of examples for each task.2
The score is called a loss. A calculation called backpropagation works backward through the network to estimate how each parameter contributed to that loss. An optimizer uses those estimates to adjust the parameters. Predict, measure, adjust; then repeat over many batches of examples.3
03 / ONE TRAINING STEP
The next example starts the cycle again.
Training material needs selection and cleaning. The model can learn useful patterns from it, but it can also absorb mistakes, bias and material it may reproduce later. Its parameters are not a searchable archive with dependable page references. A fluent answer does not tell you which example, if any, supports a particular sentence.
Why earlier text changes later text
Many large language models use a transformer. Its attention mechanism lets each token’s representation draw information from other tokens available in the sequence. Position also matters: “dog bites man” and “man bites dog” contain the same words in a different order. The original transformer paper introduced an architecture built around attention rather than reading the sequence through a recurrent loop.4
Take “The red bicycle leaned against a wall. It needed a new tire.” The earlier word “bicycle” helps make sense of “It.” Real models use many attention patterns across many layers; we cannot read off a model’s full reasoning from one simple arrow. Attention helps use context. It does not check whether the sentence is true.
04 / CONTEXT CHANGES A TOKEN
The red bicycle leaned against a wall.
It needed a new tire.
Earlier words can help interpret a later one.
The prompt, selected conversation history and any supplied documents must fit within a model’s context window, measured in tokens. The application decides what to include when a conversation becomes too long. Anything omitted is unavailable in that request, though the model’s previously learned parameters still affect the answer.1
Why a chatbot follows instructions
Pretraining gives a model a powerful completion habit. It does not guarantee that it will answer a question clearly, admit uncertainty or follow a useful format. A later stage, often called post-training, tries to improve that behaviour.
One documented approach trains on examples of instructions and helpful responses, then uses human comparisons of candidate answers to steer further training. InstructGPT used reinforcement learning from human feedback, or RLHF, for that second part.5 Other methods, including direct preference optimisation, learn from preferred and rejected answers without the same reinforcement-learning loop.6 A particular product’s recipe may differ or may not be public.
05 / A PREFERENCE EXAMPLE
“Are these figures confirmed?”
Yes, every figure is correct.
I can summarise them, but I have not checked the source.
A reviewer’s preference supplies a training signal, not proof that every later answer is honest.
What happens after you press Send
Using a trained model to answer a new request is called inference. For text, it computes scores for possible next tokens, selects one, appends it to the sequence and repeats. The text you see is built piece by piece. Your prompt changes the context for that process; a normal conversation does not update the model’s lasting parameters.
Some systems let the application vary temperature, a setting that changes how strongly generation favours high-scoring tokens. Lower values usually make the selection more concentrated; higher values allow more variation. Availability and behaviour depend on the model. Neither setting checks facts, and the highest-scoring next token is not necessarily part of a correct answer.7
06 / CHOOSING A NEXT TOKEN
“The boat entered the …”
These bars illustrate a ranking. They are not measured outputs or probabilities from a model.
Some models also spend additional computation on intermediate steps before a final response. That can help on difficult tasks, but a polished explanation of the steps is still an output to examine, not a certificate that the result is right.8
Give it evidence when evidence matters
A model’s parameters do not give it a live view of today’s documents. An application can retrieve relevant passages or call a tool, then place the results into the context for the next part of the answer. Retrieval-augmented generation is one researched way to combine a model with an external store of passages.9
Suppose someone asks when a new policy begins. The application finds the policy document and passes a passage to the model. A useful answer should state the date and point to that passage. It should also distinguish a missing date from a known one. A retrieved page may be outdated or wrong, and a citation is only useful if it supports the exact claim beside it.
07 / FROM QUESTION TO CHECKED ANSWER
Check that the passage is current and really refers to the policy in question.
Without such checks, a model can give a confident false statement, invent a reference or quietly ignore a constraint. These are often grouped under hallucinations, although they have different causes. Some training and evaluation practices can even reward a plausible guess over an honest “I don’t know.”10
Smaller models make different trade-offs
Not every job needs the largest model. Distillation trains a smaller student model using information from a larger teacher model, often its outputs or softened predictions. The aim is to keep useful behaviour while reducing the cost of running it. A smaller model can be faster or cheaper, but quality must be tested on the actual task; it will not preserve every ability of the teacher.11
Decide whether it works
A benchmark score tells you something about a model. It does not tell you whether your whole workflow has improved. For a real application, build a set of representative cases before tuning prompts or choosing a model. Include ordinary cases, missing information, conflicting sources and requests that should be declined. Decide what a good answer looks like for each one.12
08 / TEST THE TASK, NOT THE DEMO
Then measure more than the first draft: time spent checking, corrections, quality, failures and the cost of a wrong answer. Repeat the same cases when the model, prompt, tools or documents change. For consequential decisions, arrange qualified human review and reliable sources.
The model is most useful when its strengths and its limits are both visible. Treat training as the source of a capability, context as the information supplied for this request, and evaluation as the way to find out whether the complete system helps.
Resources
- OpenAI, Understanding and counting tokens
- Tom Brown and colleagues, Language Models are Few-Shot Learners
- PyTorch, Optimizing model parameters
- Ashish Vaswani and colleagues, Attention Is All You Need
- Long Ouyang and colleagues, Training language models to follow instructions with human feedback
- Rafael Rafailov and colleagues, Direct Preference Optimization
- OpenAI API, Sampling temperature
- OpenAI, Learning to reason with LLMs
- Patrick Lewis and colleagues, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- OpenAI, Why language models hallucinate
- Geoffrey Hinton and colleagues, Distilling the Knowledge in a Neural Network
- OpenAI, Evaluation best practices