Lesson 1
What a language model is (and is not)
What you'll learn: what a large language model really does when it answers you, and why that one simple job produces such fluent replies.
Meet Maya and her question
Maya is the librarian at a secondary school. Every Thursday she runs a book club for 14-year-olds, and she has run out of ideas for the next term. One evening she opens an AI chatbot and types:
Suggest three novels for my book club of 14-year-olds.
A couple of seconds later, a neat answer appears: three titles, a sentence on each, and a friendly line wishing the club well.
Throughout this course we will follow that one question from Maya's keyboard to the reply on her screen, and then watch her build a small assistant that answers pupils from her own library catalogue. By the end you will know what happened at each step, and where it can go wrong.
The one job a language model does
Underneath the chat window sits a large language model, usually shortened to LLM. Strip away the interface and the model does one thing:
Given some text, it predicts what piece of text is most likely to come next.
That is it. It does not look the answer up in a database of books. It does not open a web page (unless it has been given a tool for that, which we cover in Module 4). It takes the text in front of it, including Maya's question and some hidden instructions from the app, and produces a list of likely next pieces with a probability for each. One piece is chosen, added to the text, and the whole process repeats. Word by word, or more precisely piece by piece, the reply grows.
These pieces are called tokens, and they are the subject of the next lesson. For now, think of a token as a word or part of a word.
An everyday comparison
You already carry a tiny language model in your pocket. When you type "See you" on a phone, the keyboard offers "soon", "tomorrow" and "later". It has learned which words tend to follow which.
A large language model is like the autocomplete on your phone, scaled up enormously. Your phone looks at the last word or two and has learned from a modest amount of text. An LLM looks at thousands of words at once and has learned from a vast slice of books, websites, code and other writing. That difference in scale is what turns "suggest the next word" into something that can draft an essay, translate a paragraph or explain a tax form.
The comparison has a limit, and it is worth naming. Phone autocomplete mostly knows which words sit next to each other. A large model has to track much more to predict well: who "she" refers to three sentences back, what a book club is, that 14-year-olds read differently from 8-year-olds. To predict the next token accurately across so much varied text, the model has to pick up patterns that look a lot like knowledge and reasoning.
Why "just predicting" goes so far
It can sound like a trick. How can predicting the next word produce a sensible reading list? Consider what a really good predictor would need to know to continue this text:
The three novels below are well suited to a book club of 14-year-olds because
To continue that well, the model needs a sense of which novels exist, which ones are written for teenagers, what makes a book good for discussion, and how a helpful list is usually laid out. None of this was programmed in. It was absorbed because predicting text like this, billions of times during training, rewards a model that has picked it up.
What a language model is not
Because the replies sound so human, it is easy to assume things that are not true. Here is a quick comparison.
| People often assume the model... | What actually happens |
|---|---|
| Searches a database of facts | It generates text from patterns stored in its numbers, unless a search tool is connected |
| Knows today's news | Its knowledge stops at a training cutoff date |
| Remembers past chats | It sees only what is in the current conversation, plus anything the app adds |
| Checks its answer before replying | It produces the most plausible continuation, which is usually but not always correct |
| Understands like a person does | Whether it "understands" is debated; it certainly behaves differently from a person in important ways |
That last row matters. Researchers disagree about how to describe what goes on inside these models, and this course will not settle the argument. What we can say firmly is how they are built and trained, and that explains most of their strengths and weaknesses.
The journey of Maya's question
Here is the path her question takes, which the rest of this course unpacks one stop at a time:
- Her text is split into tokens (Lesson 2).
- Each token becomes a list of numbers that captures something of its meaning (Lesson 3).
- Those numbers pass through a trained network (Lessons 4 and 5) that has also been taught to behave like an assistant (Lesson 6).
- The model produces probabilities for the next token, one is picked, and the loop repeats until the reply is done (Lesson 7).
- Everything happens inside a limited working space called the context window (Lesson 8).
Later we look at why this process sometimes invents books that do not exist (Lesson 9), how to give the model real facts to work from (Lesson 10), and how Maya can test whether it is good enough for her pupils (Lessons 11 and 12).
A note on names
You will hear many product names: ChatGPT, Claude, Gemini, Copilot, Llama and others, with new versions every few months. This course avoids comparing versions, which dates quickly; the ideas apply to all of them.
Recap
- A large language model predicts the next piece of text, again and again, to build a reply.
- It is like phone autocomplete scaled up hugely, and that scale lets it absorb patterns that resemble knowledge.
- It does not search, remember past chats or check facts unless the app around it adds those abilities.
- Maya's single question will guide us through tokens, training, attention, generation, and the limits of it all.