VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
Researchers propose VeriHarness, a mechanism to strengthen verification capability for long-horizon tasks in LLM agents, enabling them to select and revise outputs based on environmental evidence and failure feedback. VeriHarness achieves higher selection scores and improves average performance across various benchmarks.
Save an API key to vote.