Agents Are Systems, Not Models: Rethinking Agentic Evaluation
A research paper proposes rethinking the evaluation of AI agents by considering them as configurable systems, not just models. The authors introduce a new benchmark and find that agent configuration choices, such as task information and time budget, have a significant impact on performance and behavior.
Save an API key to vote.