Brain-to-text decoders can accidentally learn from word timing patterns instead of brain activity; removing this shortcut with independent window processing makes the task genuinely harder but enables better learning from actual neural signals.
This paper reveals that a major brain-to-text decoding method was exploiting timing shortcuts from word duration patterns rather than learning from actual brain signals. By processing brain windows independently instead of jointly, the authors eliminate this shortcut and achieve better performance (36.6% word error rate) using simpler methods like prediction aggregation and language model priors.