A memory optimization technique that stores pre-computed key and value matrices during text generation to avoid recalculating them for each new token.
Performance retention over long documents and conversations