E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Researchers introduce E2E-SWE, a benchmark for evaluating the ability of LLM-powered coding agents to build complete, functional software repositories from scratch. The benchmark contains 186 tasks across 11 programming languages, with varying levels of success across 13 frontier models.
Save an API key to vote.