Transformers
The 2017 paper 'Attention Is All You Need' introduced the architecture behind every modern LLM. Here's what makes transformers different — and why they took over.
Inside a transformer
Token Embeddings
Words become numeric vectors
Positional Encoding
Word order is injected into the vectors
Self-Attention
Every token weighs relevance of every other token
Feed-Forward Layers
Each token's representation is refined
Output Probabilities
Predicted next-token distribution
Self-attention, in plain terms
For every word in a sentence, self-attention asks: 'which other words in this sentence matter most for understanding me?' In "The trophy didn't fit in the suitcase because it was too big", attention is what lets the model figure out that 'it' refers to the trophy, not the suitcase.
Transformers vs. older RNN architectures
| RNNs (older) | Transformers | |
|---|---|---|
| Processing order | Word by word, in sequence | All words processed in parallel |
| Long-range context | Struggles with long sentences | Attention captures long-range relationships well |
| Training speed | Slow — sequential by design | Fast — highly parallelizable on GPUs |
Key takeaways
- Transformers process every word in parallel, unlike older RNNs which worked word-by-word.
- Self-attention lets each word weigh the relevance of every other word, capturing long-range context.
- Parallelization is why transformers could be trained on internet-scale data — this is the architecture behind GPT, Claude, and Gemini.
Check your understanding
0/2 answered1.What is 'self-attention' primarily responsible for?
2.A key advantage of transformers over RNNs is that they can process all words in a sequence in parallel.
Lesson summary
Transformers use self-attention to let every word consider every other word in parallel — the architectural breakthrough behind every modern LLM.
AI-generated notes