Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
This paper proposes a framework for self-evolving harnesses, where a language-model agent improves its own code organization and execution control. The framework uses a recursive self-improvement process, where the frozen model solves tasks and then edits its own harness based on run records. The results show improved performance on in-distribution and out-of-distribution tasks, surpassing or matching Codex on some benchmarks.
Save an API key to vote.