AI product Open source

SWE-bench

SWE-bench is an execution-based benchmark for evaluating large language models on real-world software issues collected from GitHub. Given a repository codebase and an issue, a model must generate a patch intended to resolve the problem; the patch is evaluated by running repository tests in Docker-based environments. The project provides datasets and an evaluation command-line tool, with dataset variants including full, verified, multimodal, multilingual, and locally supplied datasets. It is distributed as open-source code through its GitHub repository and can also be loaded from Hugging Face.

View repository Visit site Mentioned in 2 videos ↓

What SWE-bench is used for

2 uses taken from transcripts — each links to the moment in the video.

  • Serves as the original coding benchmark that Senior SWE-bench extends toward senior-engineer-level work.

  • Execution-based coding benchmark that evaluates whether a model can modify a repository and solve a bug ticket using tests.

Videos mentioning SWE-bench

2 in the library.