EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
EurekaBench is a benchmark for measuring AI agents' ability to make scientific discoveries. It tests agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. The benchmark evaluates three axes of scientific discovery: following scientific constraints, predictive accuracy, and deriving scientific insights. Current AI agents perform well in predictive accuracy but struggle to derive meaningful scientific insights.
Save an API key to vote.