AI product Open source
SWE-bench is an execution-based benchmark for evaluating large language models on real-world software issues collected from GitHub. Given a repository codebase and an issue, a model must generate a patch intended to resolve the problem; the patch is evaluated by running repository tests in Docker-based environments. The project provides datasets and an evaluation command-line tool, with dataset variants including full, verified, multimodal, multilingual, and locally supplied datasets. It is distributed as open-source code through its GitHub repository and can also be loaded from Hugging Face.
2 uses taken from transcripts — each links to the moment in the video.
Serves as the original coding benchmark that Senior SWE-bench extends toward senior-engineer-level work.
Execution-based coding benchmark that evaluates whether a model can modify a repository and solve a bug ticket using tests.
2 in the library.