IntermediateEnglish28 lectures · 3h 24m

Attention and Transformers from Scratch

From RNN Seq2Seq to Bahdanau and Luong attention to a full Transformer — built, trained and explained in PyTorch

Watch 3 lectures free

What you'll learn

  • Explain the fixed-context-vector problem of RNN Seq2Seq models and measure it yourself
  • Derive additive (Bahdanau) and multiplicative (Luong) attention and implement both in PyTorch
  • Read 'Attention Is All You Need' and map each component to code
  • Implement positional encoding, masking, multi-head attention, Add & Norm and feed-forward blocks
  • Assemble, train and use a complete Transformer for translation and inspect its cross-attention

About this course

Attention is the idea behind every modern language model, and it was invented to solve one concrete problem: a sequence-to-sequence model squeezes a whole sentence into one fixed vector, and long sentences fall apart. This course starts from that problem and builds all the way up to a working Transformer.

We begin with what language models can do and the roadmap of an NLP engineer. Then we implement an RNN Seq2Seq model in PyTorch and measure its limits: the fixed context vector, what a wider hidden state does and does not fix. We derive the attention mechanism in depth — the core idea, the alignment weights, Bahdanau's additive attention — and implement Bahdanau and Luong attention with alignment visualisations. Next we read "Attention Is All You Need" section by section: self-attention, multi-head attention, positional encoding, feed-forward blocks, encoder and decoder. Finally we assemble a Transformer from parts — positional encoding, masking, multi-head attention, Add & Norm, encoder and decoder layers — train it, translate with it, and inspect cross-attention.

Theory sessions use slides; implementation sessions show the notebook on screen while the narration explains every cell. Small models that train in minutes, full understanding of what each line does. Resources also include a bonus notebook that trains Bahdanau, Luong and scaled dot-product attention on the same corpus, so you can compare the three scoring rules and their alignment heatmaps side by side.

Curriculum8 sections · 28 lectures · 3h 24m

Section 2. Limits of RNN Seq2Seq

  • 🔒Opening questions and the Seq2Seq idea5:54
  • 🔒How a Seq2Seq model is implemented8:30
  • 🔒The problem with a fixed context vector7:30
  • 🔒The arrival of attention7:48

Section 3. Building a Seq2Seq Model

  • 🔒Data encoder and decoder in PyTorch9:18
  • 🔒Training setup and the first run7:12
  • 🔒Does a wider hidden state help Sweeps5:18

Section 4. The Attention Mechanism in Depth

  • 🔒The core idea of attention6:12
  • 🔒Bahdanau attention in detail9:18
  • 🔒Implementing Bahdanau attention5:48

Section 5. Implementing Attention

  • 🔒Bahdanau attention additive scoring training and alignment12:54
  • 🔒Luong attention multiplicative scoring9:24

Section 6. Transformer Architecture

  • 🔒Self attention8:24
  • 🔒Multi head attention and positional encoding9:54
  • 🔒Reading Attention Is All You Need9:48
  • 🔒Feed forward encoder decoder and assembly5:54

Section 7. Assembling the Transformer

  • 🔒Positional encoding8:30
  • 🔒Masking5:12
  • 🔒From one head to many5:12
  • 🔒Multi head attention and what the loop costs6:30
  • 🔒Add and Norm feed forward encoder and decoder layers8:30
  • 🔒Final assembly and training7:24
  • 🔒Translating cross attention and what we built8:00

Section 8. Wrap-up

  • 🔒What we learned6:06
  • 🔒Where to go next5:42

Requirements

  • · A computer with Python 3.10+ (a free Google Colab account is enough for most sessions)
  • · Comfort reading Python code; you do not need to be an expert
  • · Basic PyTorch and an idea of what an RNN does

Who this is for

  • · Developers who use Transformers and want to understand the mechanism
  • · Students starting NLP or preparing for deep-learning interviews
  • · Engineers who want to build a Transformer from parts rather than import one

Read alongside the course

The voice-over in this course is synthesized with a text-to-speech model from scripts written and reviewed by the instructor, and the on-screen material (notebooks, code, slides) is the instructor's own work.

Attention and Transformers from Scratch | SOTAAZ Blog