Token-level tolerance to transcription ambiguity in ASR training reduces word error rates by ~9.5% on average by letting models skip disputed individual characters while keeping supervision for the rest of the word.
This paper addresses a real problem in speech recognition: reference transcripts often contain ambiguous pronunciations or spellings that the audio doesn't uniquely determine.