Audio LLMs Know When They Can't Hear You
This paper proposes a lightweight reliability predictor for Audio LLMs to detect when their transcription is unreliable. The predictor uses the model's audio-encoder representations to trigger a clarification request from the user when the query is predicted to be unreliable.
Save an API key to vote.