Hands-on Practice
student workspace
0 XP
PS
AI Foundation Program/How AI Works
intermediate8 min read

Transformers

The 2017 paper 'Attention Is All You Need' introduced the architecture behind every modern LLM. Here's what makes transformers different — and why they took over.

Inside a transformer

Self-attention, in plain terms

For every word in a sentence, self-attention asks: 'which other words in this sentence matter most for understanding me?' In "The trophy didn't fit in the suitcase because it was too big", attention is what lets the model figure out that 'it' refers to the trophy, not the suitcase.

Transformers vs. older RNN architectures

RNNs (older)Transformers
Processing orderWord by word, in sequenceAll words processed in parallel
Long-range contextStruggles with long sentencesAttention captures long-range relationships well
Training speedSlow — sequential by designFast — highly parallelizable on GPUs

Key takeaways

  • Transformers process every word in parallel, unlike older RNNs which worked word-by-word.
  • Self-attention lets each word weigh the relevance of every other word, capturing long-range context.
  • Parallelization is why transformers could be trained on internet-scale data — this is the architecture behind GPT, Claude, and Gemini.

Check your understanding

0/2 answered

1.What is 'self-attention' primarily responsible for?

2.A key advantage of transformers over RNNs is that they can process all words in a sequence in parallel.

Lesson summary

Transformers use self-attention to let every word consider every other word in parallel — the architectural breakthrough behind every modern LLM.

AI-generated notes