A task where a model listens to spoken audio and answers questions about its content in speech or text.