Framing spoken language understanding as structured function calling—rather than slot-filling—lets audio language models generalize to new tasks without retraining, similar to how code models handle function calls.
This paper introduces Spoken Function Calling (SFC), a new way to understand spoken language that treats semantic extraction like function calls with structured definitions.