Sample course · Beginner · 12 lessons

How large language models actually work

From tokens to transformers, in plain language, with one librarian's question as the guide

A plain-language tour of what happens inside a large language model: how text becomes tokens and numbers, how training on next-token prediction works, what attention does, and how a raw model becomes a helpful assistant. You will finish able to explain why chatbots sound confident when they are wrong, what a context window limits, how tools and retrieval add fresh facts, and how to test whether a model is good enough for your own task.

What you'll learn

  • Explain next-token prediction and why it is enough to produce fluent answers
  • Describe how text is split into tokens and turned into embeddings
  • Give an intuitive account of training, transformers and attention
  • Explain how fine-tuning and human feedback turn a raw model into an assistant
  • Predict how sampling settings and the context window shape a model's replies
  • Recognise why hallucinations happen and how retrieval and tools reduce them
  • Design a simple test set to judge whether a model is fit for a real task

Who it's for

  • People who use chatbots at work or home and want to know what is going on underneath
  • Non-programmers deciding whether an AI tool can be trusted with a task
  • Students and curious readers who want the concepts without the maths

Syllabus

  1. 1.From words to numbers

    What a language model actually does, and how it turns your text into pieces and then into numbers it can work with.

    1. What a language model is (and is not)
    2. Tokens: how text is chopped into pieces· checkpoint
    3. Embeddings: meaning as a place on a map
  2. 2.How a model learns

    Pretraining on huge amounts of text, the transformer and its attention mechanism, and the extra training that turns a text predictor into a helpful assistant.

    1. Training: learning by guessing the next token· checkpoint
    2. Transformers and attention, intuitively
    3. From text predictor to helpful assistant· checkpoint
  3. 3.What happens when you hit send

    How a reply is produced token by token, what the context window lets a model see, and why fluent answers can still be false.

    1. Writing a reply one token at a time
    2. Context windows: what the model can see right now· checkpoint
    3. Why models make things up
  4. 4.Making models useful, and knowing their limits

    Retrieval and tools that give a model fresh, checkable facts, how to evaluate a model for your own task, and the limits worth keeping in mind.

    1. Tools and retrieval: an open-book model· checkpoint
    2. Testing a model on your own task
    3. Limits, costs and how to think about what comes next· checkpoint

Lesson 1

What a language model is (and is not)

What you'll learn: what a large language model really does when it answers you, and why that one simple job produces such fluent replies.

Meet Maya and her question

Maya is the librarian at a secondary school. Every Thursday she runs a book club for 14-year-olds, and she has run out of ideas for the next term. One evening she opens an AI chatbot and types:

Suggest three novels for my book club of 14-year-olds.

A couple of seconds later, a neat answer appears: three titles, a sentence on each, and a friendly line wishing the club well.

Throughout this course we will follow that one question from Maya's keyboard to the reply on her screen, and then watch her build a small assistant that answers pupils from her own library catalogue. By the end you will know what happened at each step, and where it can go wrong.

The one job a language model does

Underneath the chat window sits a large language model, usually shortened to LLM. Strip away the interface and the model does one thing:

Given some text, it predicts what piece of text is most likely to come next.

That is it. It does not look the answer up in a database of books. It does not open a web page (unless it has been given a tool for that, which we cover in Module 4). It takes the text in front of it, including Maya's question and some hidden instructions from the app, and produces a list of likely next pieces with a probability for each. One piece is chosen, added to the text, and the whole process repeats. Word by word, or more precisely piece by piece, the reply grows.

These pieces are called tokens, and they are the subject of the next lesson. For now, think of a token as a word or part of a word.

An everyday comparison

You already carry a tiny language model in your pocket. When you type "See you" on a phone, the keyboard offers "soon", "tomorrow" and "later". It has learned which words tend to follow which.

A large language model is like the autocomplete on your phone, scaled up enormously. Your phone looks at the last word or two and has learned from a modest amount of text. An LLM looks at thousands of words at once and has learned from a vast slice of books, websites, code and other writing. That difference in scale is what turns "suggest the next word" into something that can draft an essay, translate a paragraph or explain a tax form.

The comparison has a limit, and it is worth naming. Phone autocomplete mostly knows which words sit next to each other. A large model has to track much more to predict well: who "she" refers to three sentences back, what a book club is, that 14-year-olds read differently from 8-year-olds. To predict the next token accurately across so much varied text, the model has to pick up patterns that look a lot like knowledge and reasoning.

Why "just predicting" goes so far

It can sound like a trick. How can predicting the next word produce a sensible reading list? Consider what a really good predictor would need to know to continue this text:

The three novels below are well suited to a book club of 14-year-olds because

To continue that well, the model needs a sense of which novels exist, which ones are written for teenagers, what makes a book good for discussion, and how a helpful list is usually laid out. None of this was programmed in. It was absorbed because predicting text like this, billions of times during training, rewards a model that has picked it up.

What a language model is not

Because the replies sound so human, it is easy to assume things that are not true. Here is a quick comparison.

People often assume the model...What actually happens
Searches a database of factsIt generates text from patterns stored in its numbers, unless a search tool is connected
Knows today's newsIts knowledge stops at a training cutoff date
Remembers past chatsIt sees only what is in the current conversation, plus anything the app adds
Checks its answer before replyingIt produces the most plausible continuation, which is usually but not always correct
Understands like a person doesWhether it "understands" is debated; it certainly behaves differently from a person in important ways

That last row matters. Researchers disagree about how to describe what goes on inside these models, and this course will not settle the argument. What we can say firmly is how they are built and trained, and that explains most of their strengths and weaknesses.

The journey of Maya's question

Here is the path her question takes, which the rest of this course unpacks one stop at a time:

  1. Her text is split into tokens (Lesson 2).
  2. Each token becomes a list of numbers that captures something of its meaning (Lesson 3).
  3. Those numbers pass through a trained network (Lessons 4 and 5) that has also been taught to behave like an assistant (Lesson 6).
  4. The model produces probabilities for the next token, one is picked, and the loop repeats until the reply is done (Lesson 7).
  5. Everything happens inside a limited working space called the context window (Lesson 8).

Later we look at why this process sometimes invents books that do not exist (Lesson 9), how to give the model real facts to work from (Lesson 10), and how Maya can test whether it is good enough for her pupils (Lessons 11 and 12).

A note on names

You will hear many product names: ChatGPT, Claude, Gemini, Copilot, Llama and others, with new versions every few months. This course avoids comparing versions, which dates quickly; the ideas apply to all of them.

Recap

  • A large language model predicts the next piece of text, again and again, to build a reply.
  • It is like phone autocomplete scaled up hugely, and that scale lets it absorb patterns that resemble knowledge.
  • It does not search, remember past chats or check facts unless the app around it adds those abilities.
  • Maya's single question will guide us through tokens, training, attention, generation, and the limits of it all.

Lesson 2

Tokens: how text is chopped into pieces

What you'll learn: how a model splits text into tokens, why it does it that way, and how tokens explain prices, limits and a few odd mistakes.

The first stop: Maya's text is chopped up

When Maya presses send, her question does not travel to the model as a sentence. A program called a tokenizer first cuts it into pieces called tokens, and replaces each piece with a number. The model only ever sees those numbers.

Each model family has its own tokenizer, so the exact split varies. One plausible split of her question looks like this:

Piece of textWhat it is
Suggesta whole word, one token
threea word with its leading space, one token
novelsone token
for, my, book, club, offive common words, one token each
14a number, often one or two tokens
-year, -old, sa compound broken into parts
.punctuation, one token

That is about 14 tokens for a 10-word question. Notice two things. Spaces usually travel with the word that follows them, and a less common form like "14-year-olds" gets split into several parts.

Why not just use whole words, or single letters?

The designers of a tokenizer face a trade-off.

If every letter were a token, the vocabulary would be tiny, but every sentence would become a very long sequence. The model would spend enormous effort just assembling letters into words before it could think about meaning.

If every whole word were a token, sequences would be short, but the vocabulary would need millions of entries to cover names, typos, technical terms, other languages and words invented last week. Any word not in the list would be impossible to represent.

Tokens sit in between. Think of a box of LEGO bricks. A good set has some large specialised pieces for shapes you build all the time, plus small basic bricks you can combine into anything else. Common words like "the" or "book" get their own large piece. Rare words like "Thessaloniki" or "unputdownable" get built from several smaller pieces. Nothing is ever impossible to build, because in the worst case you can fall back to the smallest bricks, individual characters or even raw bytes.

The most common method for choosing these pieces is called byte-pair encoding. It starts with single characters, then repeatedly merges the pairs that appear together most often in a large sample of text, until it reaches a target vocabulary size. Modern vocabularies typically hold somewhere between tens of thousands and a couple of hundred thousand tokens.

A useful rule of thumb

For ordinary English prose, one token averages about three quarters of a word, or roughly four characters. So:

  1. 100 English words is about 130 tokens.
  2. A 1,500-word article is about 2,000 tokens.
  3. A 300-page novel of roughly 90,000 words is about 120,000 tokens.

These are averages, not laws. Text full of code, numbers, unusual names or formatting takes more tokens per word. Many other languages also take more tokens per word than English, because tokenizers were often built from samples dominated by English text. The gap has narrowed in newer tokenizers but has not disappeared, and it varies by language.

Why tokens matter to you

Tokens are not just an internal detail. They shape three things you will notice as a user.

Cost

When you use a model through a paid service or an API, the bill is usually counted in tokens, both the ones you send and the ones the model writes back. Prices are often quoted per million tokens, and output tokens commonly cost more than input tokens. A prompt that pastes in a whole report costs more than a short question.

Limits

Every model has a maximum number of tokens it can handle at once, called its context window (Lesson 8 is all about it). When an app says a document is "too long", it means too many tokens, not too many pages.

Strange blind spots

Because the model sees token numbers rather than letters, some tasks that are trivial for you are awkward for it. Asking "How many r's are in strawberry?" can trip models up, because "strawberry" may arrive as one or two tokens, and the individual letters are not directly in view. The same goes for rhyming, spelling backwards or counting characters. Newer models often handle these better, partly through extra training and partly by writing the word out letter by letter first, but the root cause is the tokenizer.

Arithmetic with long numbers can suffer for a similar reason. A number like 48,213 might be split as "48", ",", "213", and the split may differ from one number to the next, which makes column-by-column addition harder than it looks.

Back to Maya

Maya's question is tiny, about 14 tokens. But the app adds its own hidden instructions before her text (things like "You are a helpful assistant"), so the model may actually receive a few hundred tokens or more before it starts writing. Her reply of about 120 words will be around 160 tokens.

Next term, when Maya wants the model to work from her library catalogue of 8,000 books, tokens will suddenly matter a lot. At 20 words per catalogue entry, that is 160,000 words, or more than 200,000 tokens. Whether that fits, and what it costs, comes down to the arithmetic you just learned.

Recap

  • A tokenizer cuts text into tokens and turns each into a number; the model only sees the numbers.
  • Tokens are a middle ground between letters and whole words, built like a box of LEGO bricks with common pieces and small fallback pieces.
  • In English, one token is about three quarters of a word; other languages, code and numbers often use more.
  • Tokens determine cost, length limits, and some odd weaknesses such as counting letters.

This lesson ends with a 2-question checkpoint, graded in the app.

10 more lessons in this course

Start it in Akadyo to read on, take the checkpoints and keep your place, with a tutor beside every lesson.