How AI Writes Text
ChatGPT-style models don't 'understand' language the way you do — they predict the most likely next word, over and over, at incredible speed. Here's the pipeline behind it.
A Large Language Model (LLM) generates text one small piece at a time. Given everything written so far, it calculates a probability for every possible next 'token' (a word or word-fragment), picks one, appends it, and repeats — hundreds of times per response.
The text-generation pipeline
Your Prompt
"Write a haiku about the ocean"
Tokenization
Text is split into sub-word tokens
Next-Token Prediction
Model scores every possible next token
Sampling
One token is chosen, sometimes with randomness
Output Text
Tokens are decoded back into words
What happens in each generation step
- 1
Tokenize
Your text is broken into tokens the model was trained to recognize.
- 2
Predict
The model outputs a probability distribution over its entire vocabulary for 'what comes next'.
- 3
Sample
A token is selected — often the highest-probability one, sometimes a slightly less likely one for variety.
- 4
Repeat
The new token is appended to the input, and the whole process runs again for the next token.
Definition
This is why LLMs can be trained on trillions of words yet still 'hallucinate' — they're not looking up facts, they're generating the statistically most plausible continuation of the text so far.
Key takeaways
- LLMs generate text one token at a time by predicting the most likely next token, repeatedly.
- The pipeline is: tokenize → predict probabilities → sample a token → repeat.
- Because generation is probabilistic, not fact-lookup, models can produce fluent but incorrect statements — always verify important claims.
Check your understanding
0/2 answered1.What does an LLM actually predict at each generation step?
2.LLMs generate entire responses in a single step rather than token by token.
Lesson summary
AI writes text by repeatedly predicting and sampling the next most likely token — a statistical process, not a lookup, which is exactly why fluent AI text can still be factually wrong.
AI-generated notes