Attention and Transformers from Scratch
From RNN Seq2Seq to Bahdanau and Luong attention to a full Transformer — built, trained and explained in PyTorch
What you'll learn
- ✓Explain the fixed-context-vector problem of RNN Seq2Seq models and measure it yourself
- ✓Derive additive (Bahdanau) and multiplicative (Luong) attention and implement both in PyTorch
- ✓Read 'Attention Is All You Need' and map each component to code
- ✓Implement positional encoding, masking, multi-head attention, Add & Norm and feed-forward blocks
- ✓Assemble, train and use a complete Transformer for translation and inspect its cross-attention
About this course
Attention is the idea behind every modern language model, and it was invented to solve one concrete problem: a sequence-to-sequence model squeezes a whole sentence into one fixed vector, and long sentences fall apart. This course starts from that problem and builds all the way up to a working Transformer.
We begin with what language models can do and the roadmap of an NLP engineer. Then we implement an RNN Seq2Seq model in PyTorch and measure its limits: the fixed context vector, what a wider hidden state does and does not fix. We derive the attention mechanism in depth — the core idea, the alignment weights, Bahdanau's additive attention — and implement Bahdanau and Luong attention with alignment visualisations. Next we read "Attention Is All You Need" section by section: self-attention, multi-head attention, positional encoding, feed-forward blocks, encoder and decoder. Finally we assemble a Transformer from parts — positional encoding, masking, multi-head attention, Add & Norm, encoder and decoder layers — train it, translate with it, and inspect cross-attention.
Theory sessions use slides; implementation sessions show the notebook on screen while the narration explains every cell. Small models that train in minutes, full understanding of what each line does. Resources also include a bonus notebook that trains Bahdanau, Luong and scaled dot-product attention on the same corpus, so you can compare the three scoring rules and their alignment heatmaps side by side.
Curriculum8 sections · 28 lectures · 3h 24m
Section 1. Course Introduction
Free previewSection 2. Limits of RNN Seq2Seq
- 🔒Opening questions and the Seq2Seq idea5:54
- 🔒How a Seq2Seq model is implemented8:30
- 🔒The problem with a fixed context vector7:30
- 🔒The arrival of attention7:48
Section 3. Building a Seq2Seq Model
- 🔒Data encoder and decoder in PyTorch9:18
- 🔒Training setup and the first run7:12
- 🔒Does a wider hidden state help Sweeps5:18
Section 4. The Attention Mechanism in Depth
- 🔒The core idea of attention6:12
- 🔒Bahdanau attention in detail9:18
- 🔒Implementing Bahdanau attention5:48
Section 5. Implementing Attention
- 🔒Bahdanau attention additive scoring training and alignment12:54
- 🔒Luong attention multiplicative scoring9:24
Section 6. Transformer Architecture
- 🔒Self attention8:24
- 🔒Multi head attention and positional encoding9:54
- 🔒Reading Attention Is All You Need9:48
- 🔒Feed forward encoder decoder and assembly5:54
Section 7. Assembling the Transformer
- 🔒Positional encoding8:30
- 🔒Masking5:12
- 🔒From one head to many5:12
- 🔒Multi head attention and what the loop costs6:30
- 🔒Add and Norm feed forward encoder and decoder layers8:30
- 🔒Final assembly and training7:24
- 🔒Translating cross attention and what we built8:00
Section 8. Wrap-up
- 🔒What we learned6:06
- 🔒Where to go next5:42
Requirements
- · A computer with Python 3.10+ (a free Google Colab account is enough for most sessions)
- · Comfort reading Python code; you do not need to be an expert
- · Basic PyTorch and an idea of what an RNN does
Who this is for
- · Developers who use Transformers and want to understand the mechanism
- · Students starting NLP or preparing for deep-learning interviews
- · Engineers who want to build a Transformer from parts rather than import one
Read alongside the course
- Building Seq2Seq from Scratch: How the First Neural Architecture Solved Variable-Length I/O
- Why Your Translation Model Fails on Long Sentences: Context Vector Bottleneck Explained
- Bahdanau vs Luong Attention: Which One Should You Actually Use? (Spoiler: Luong)
- LLM Inference Optimization Part 1 — Attention Mechanism Deep Dive
- Karpathy's microgpt.py Dissected: Understanding GPT's Essence in 150 Lines
The voice-over in this course is synthesized with a text-to-speech model from scripts written and reviewed by the instructor, and the on-screen material (notebooks, code, slides) is the instructor's own work.