CruxBench: A Benchmark of Information Discovery
CruxBench is a new benchmark for evaluating the information discovery capabilities of large language models (LLMs). It assesses a model's ability to identify key questions (cruxes) that provide important steps toward solving a problem, rather than just answering fixed reference labels. CruxBench is unique in being contamination-resistant, open-ended, and grounded in real-world beliefs, and has been evaluated on eight diverse models with promising results.
Save an API key to vote.