Attention: Learning the Similarity
Part III leaves us with a whole trajectory of source states and one question: which of them matters now? We begin with a fixed similarity-weighted average, the classical version of a query reading from a memory bank. Then the comparison itself becomes learnable through query, key, and value maps. Self-attention shares that rule across every pair of positions; positional representations restore order, and masking declares what each position may see. The test-time-regression interlude holds the prediction problem fixed and compares what different memory solvers must retain; pretraining and patch tokens then carry the representation across tasks and images. Scaling closes the Part with a systems bill: a training-optimal backbone may still be too expensive to copy, update, or store for every task.