Specialized components in transformer models that store and process key-value pairs to help the model focus on relevant parts of the input when generating each output token.
Performance retention over long documents and conversations
Multi-step reasoning, logic puzzles, mathematical problem-solving