Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for evaluating sequential decision-making agents, including RL, LLM, and hybrid approaches. It provides 37 procedurally generated tasks across various difficulty levels and observation modalities, and includes a coding API, oracle reference policies, and a live leaderboard. The evaluation of 27 configurations and over 90,000 episodes highlights the need for improvement across all agent paradigms.

RSS Score 0 9/28/2026, 4:00:00 AM Original Source
Save an API key to vote.