Module 8 – Attention Mechanism
What you'll learn
Builds on Modules 1–7
≈6 h · 2 h video · 2 h reading · 2 h coding
- Implement QKV attention; visualize attention weights
- Build a toy seq2seq (copy/reverse) with attention; interpret alignments
- Run a small ablation on attention variants and report findings
Lecture 8 – Attention Mechanism (Part 1)
Lecture 8.2 – Attention Mechanism (Part 2)
Lecture 8.3 – Implementing Attention in seq2seq Decoder
📚 Resources & Live Coding
Required reading · course book
Research lens · read after the course book
These papers use the regression lens as architecture-design language; begin with the named sections after completing the book derivation.
Homework 4: Add cross-attention to your prior GRU-based seq2seq model so that the decoder can attend over all encoder hidden states at each decoding step (same translation task). Report and compare accuracy/bleu vs your previous best.
Seq2Seq Cross-Attention Diagram preview: