A transformer block is six tensors and a bus: read it like a pipeline
Originally published at cchinchilla.dev. Part 5 of From code to weights, a 12-part series on ML fundamentals for engineers. Part 4 ended on one stage. A decoder is a stack of blocks, twelve in GPT-2 small, and attention is one of two stages in each. This post reads a whole block the way you'd read a pipeline: what flows in, what each stage writes, and where the parameters and the cost end up.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in