Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?
A study on the effectiveness of using off-the-shelf small language models (SLMs) in agent harnesses, finding an eligibility gap for microtasks involving large language models. The study proposes a benchmark and analysis framework to assess SLMs and suggests using a baseline that meets a context-informed (CI-backed) threshold.
Save an API key to vote.