An LLM design combining softmax attention layers with linear-attention layers to balance expressiveness and efficiency.