The best short explanation of transformer mechanics in any medium: embeddings as directions in space, attention as tokens updating each other's meaning, softmax at the end. You will not be able to implement one afterward, and you will understand every subsequent conversation better.
What to do Watch it once at normal speed, on the page, without taking notes. Then rewatch only the attention section and say out loud what a single attention head does to one token. Replay that section if the explanation still feels slippery. That sentence makes the rest of the map easier to follow.
Also in How it works.
27 min YouTube · free
