The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
A new benchmark, KNOWS, is introduced to evaluate AI agents' ability to perform complex tasks, such as synthesizing and organizing knowledge, and navigating program interfaces. Current agents struggle with visual understanding and long-horizon reasoning, highlighting areas for improvement in AI development.
Save an API key to vote.