Audio-language models can learn complex reasoning by generating their own reasoning trajectories while receiving token-level guidance from text teachers, without requiring massive amounts of audio reasoning data.
This paper presents X³-OPD, a method for teaching audio-language models to reason better by learning from text-based models.