From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks
Researchers investigate the effectiveness of using LLMs as judges for open-ended tasks, examining judgment quality and downstream utility. They find that judgment quality and utility do not always align and that Judge protocol design affects both. Their results suggest a multifaceted evaluation approach for LLM judges.
Save an API key to vote.