Linear STFT outperforms the more complex HCQT for vocal ensemble pitch estimation while reducing computational cost—simpler input representations can be more effective when properly designed.
This paper challenges the conventional use of harmonic constant-Q transform (HCQT) for multi-pitch estimation in vocal ensembles by showing that simpler linear STFT representations actually perform better while being much faster to compute.