ISOM-R2: Streaming 1,055,402 Tokens on 3.24 GB Peak VRAM
Standard Transformer attention has a memory problem at scale. For a 1.05 million-token context with 28 layers, 2 KV heads, and a head dimension of 128, the FP16 Key-Value cache alone requires: 2 × 28 × 2 × 128 × 1,055,402 × 2 bytes ≈ 28.2 GiB That is before loading a single model weight. On a 40 GB A100, the KV cache would consume over 70% of total available memory. On consumer GPUs, it simply crashes. ISOM-R2 (Isometric State Space / Virtual SVD) is a new-gen recurrent memory architecture built to solve this directly.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in