ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality
Researchers introduce ReLiveGym, an evaluation environment for long-lived AI agents that operate over weeks of simulated real-world data. The study investigates how model choice, harness design, and continuous learning affect agent performance on time-sensitive tasks.
Save an API key to vote.