E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Researchers introduce E2E-SWE, a benchmark for evaluating the ability of LLM-powered coding agents to build complete, functional software repositories from scratch. The benchmark contains 186 tasks across 11 programming languages, with varying levels of success across 13 frontier models.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.