CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
CompMat-Bench is a benchmark for evaluating AI agents on computational materials science tasks. It reproduces research steps and assesses agents on preparing inputs and analyzing outputs for expensive simulations, supporting four evaluation conditions. The benchmark demonstrates the ability of agents based on three LLMs to complete individual materials research steps with pass rates of 66.0-90.4%.
Save an API key to vote.