A computational approach where the time and memory required to process a sequence grows proportionally to its length, rather than exponentially like attention mechanisms.
Performance retention over long documents and conversations