Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Researchers compare the performance of 66 configurations of language models paired with different harnesses, finding that the best model-harness pairing can vary depending on the task, and that a higher cost does not always buy a higher score. The study provides a framework for evaluating the fit between models and harnesses, and releases the evaluation code and harness adapters for reproducibility.
Save an API key to vote.