Audio language models don't reliably represent the same phonetic features in the same direction across speech and text—suggesting these models may process the two modalities quite differently despite using a shared decoder.
This paper investigates whether audio language models represent phonetic features consistently across speech and text inputs. Using minimal pairs of phonemes that differ in single features, researchers measured whether the same distinctive features (like voicing) are encoded in the same direction in both modalities across 6 models, 7 features, and 15 languages.