门户首页
教学网站 · 领域基础

SWE-bench 数据长什么样:Schema 详解与实例

SWE-bench(SWE = Software Engineering 软件工程 + bench = benchmark 基准测试,用真实 GitHub issue 评测代码 agent 的软件工程基准)数据集是 jsonl 格式,每行一条 instance。 这一页把每条 instance 的完整字段(instance_id / repo / base_commit / patch / test_patch / problem_statement / fail_to_pass / pass_to_pass / version)逐个讲清——含类型、用途、典型值;并配真实实例(django__django-10914)让你一眼看到结构。 最后讲本仓零依赖加载器怎么用,读完你就能自己加载、过滤、统计 SWE-bench 数据。

生成时间:2026-09-02 · 版本 v0.2 · 生成 Agent:MiniMax Code (LLM: MiniMax-M3) · 载体:agentsoft-research-platform teaching-web-platform

概览

jsonl
数据格式
9
核心字段
2,294
原版实例数
12
Python 仓库
项目说明
本卡定位SWE-bench 数据 schema 详解 · 1.0 领域基础 第 2 张
前置swe-bench-intro.html(背景 + F2P/P2P 范式)
读完会什么能用本仓零依赖加载器加载 SWE-bench 数据;理解每个字段含义;能写出过滤/统计脚本
姊妹swe-bench-evaluation.html(数据怎么被 S1-S7 流水线消费)
权威源huggingface.co/datasets/SWE-bench/SWE-bench · github.com/SWE-bench/SWE-bench

§ 1 数据格式:jsonl(每行一条 instance)

SWE-bench 原始数据集是 jsonl(JSON Lines,每行一个 JSON(JavaScript Object Notation,人类可读的文本数据格式)对象)格式——每行一个完整 JSON 对象,对应一条 instance。

为什么用 jsonl 而不是单个 JSON 数组
  • 流式读取:jsonl 可以逐行解析,2,294 条不用一次全部加载到内存
  • 容易切分:用 head -100 / wc -l 等 shell 工具就能做粗筛
  • 容错强:某行 JSON 解析失败不影响其他行
  • git 友好:新增 1 行就是 1 个 diff,PR(Pull Request,GitHub 上的代码合并请求)review 容易
文件结构
# swe-bench.jsonl 文件结构(每行 1 个 JSON)
{"instance_id": "django__django-10914", "repo": "django/django", ...}
{"instance_id": "django__django-11095", "repo": "django/django", ...}
{"instance_id": "matplotlib__matplotlib-13987", "repo": "matplotlib/matplotlib", ...}
...
(总共 2,294 行 / 文件约 380 MB)
文件大小注意:原版 swe-bench.jsonl 约 380 MB,超过 GitHub 单文件 100 MB 限制,所以本仓用 .gitignore 排除——首次使用需手动从 HuggingFace 拉取(见 §6)。

§ 2 完整 Schema(9 个核心字段)

每条 instance 必含 9 个核心字段。下面用 TypeScript 风格表达 schema(? = 可选),并对每个字段详解。

完整 TypeScript 类型
type SWEBenchInstance = {
  // —— 身份与定位 ——
  instance_id: string,                    // "django__django-10914"
  repo: string,                          // "django/django"
  base_commit: string,                    // "abc123def..."(git commit hash)
  version: string,                       // "3.10" / "3.11" 等(Python 版本)

  // —— 题面 ——
  problem_statement: string,              // issue 文本(含 title + body)

  // —— 答案(ground truth,agent 看不到)——
  patch: string,                          // fix PR 的 git diff
  test_patch: string,                     // 配套测试改动

  // —— 评测契约(F2P / P2P)——
  fail_to_pass: string[],                 // F2P 测试 ID 列表
  pass_to_pass: string[]                  // P2P 测试 ID 列表
}
字段详解

instance_id · 唯一标识

维度值
类型string
必填是
格式<repo>__<issue-number>(双下划线分隔)
例子"django__django-10914"、"matplotlib__matplotlib-13987"
作用作为 主键——预测文件、结果文件、数据库表都靠它关联

注意:django__django 是"仓库名 / 仓库名"格式的冗余(因为 GitHub owner 和 repo name 可能相同也可能不同,如 pallets__flask 的 owner 是 pallets、repo 是 flask)。这个命名沿用 GitHub Issue URL 的格式。

repo · GitHub 仓库标识

维度值
类型string
必填是
格式<owner>/<repo>
例子"django/django"、"matplotlib/matplotlib"、"pallets/flask"
作用对应 GitHub github.com/<owner>/<repo>,用于拉代码、跑测试

base_commit · 仓库起点 commit

维度值
类型string(git SHA-1 hash)
必填是
格式40 位 hex(如 "a4e35c2f8b9d12e7c...")
作用agent 从这个 commit 状态开始改代码;评测时从这拉 base 状态

version · Python 版本

维度值
类型string
必填是
格式"3.10" / "3.11" 等
作用对应 answer_evaluator/harness/constants.py 的 MAP_REPO_VERSION_TO_SPECS 字典的 key——缺了会 KeyError
重要:version 字段必须与本仓 MAP_REPO_VERSION_TO_SPECS 字典对齐,否则 make_test_spec() 抛 KeyError。详见 swe-bench-evaluation.html。

problem_statement · issue 文本

维度值
类型string(Markdown)
必填是
典型长度~200 词(含 title + body)
内容issue 的 title + body(markdown)
作用唯一给 agent 看的题面——agent 只看到这个,不知道 patch / test_patch

patch · fix PR 的代码改动(ground truth)

维度值
类型string(unified diff 格式)
必填是
典型规模~50 行(含上下文)
格式diff --git a/file.py ... @@ -10,3 +10,4 @@ ...
作用"正确答案"——评测时不直接用,只在分析/对比时用

test_patch · 配套测试改动

维度值
类型string(unified diff)
必填是
内容定义 F2P 测试(验证 bug 修了)+ 修测试自身(让 F2P 在 pre-fix 状态 fail)
作用评测时必须先 apply——让 F2P 测试"在 base_commit 状态 fail、在 model_patch 后 pass"
test_patch 的关键作用:没有 test_patch,F2P 测试在 base_commit 状态可能是 pass 的——那 agent 啥也不用改,F2P 就过了。test_patch 引入"验证 bug 的新测试",确保评测的起点是真有 bug的。

fail_to_pass · F2P 测试 ID 列表

维度值
类型string[](pytest test node ID)
必填是
例子["django.test.utils_tests.test_functional.TestDatetimeWarnings.test_naive_datetime_warning"]
作用该列表中的测试必须 (pre: fail) ∧ (post: pass) 才能算 resolved

pass_to_pass · P2P 测试 ID 列表

维度值
类型string[]
必填是
例子["django.db.models.fields.test_booleanfield.test_boolean_field_save_load"]
作用该列表中的测试必须 (pre: pass) ∧ (post: pass)——防回归

§ 3 真实实例展示(django__django-10914)

从 Lite 子集挑一条真实 instance,让你看到 9 个字段实际长什么样。为节省篇幅,patch / test_patch / problem_statement 截断显示。

完整 jsonl 一行(截断版)
{
  "instance_id": "django__django-10914",
  "repo": "django/django",
  "base_commit": "d8b8f8f8f8f8f8f8f8f8f8f8f8f8f8f8f8f8f8f8",
  "version": "3.10",
  "problem_statement": "FILE_UPLOAD_PERMISSION default to 0o644 is unsafe on multi-user systems...\n\n## Description\n\nWhen using `FileSystemStorage`, the default `FILE_UPLOAD_PERMISSION` is `0o644`, which means other users on the same system can read uploaded files...\n\n## Steps to Reproduce\n\n1. Set up a Django project with `FileSystemStorage`\n2. Upload a file via a form\n3. Check the file permissions: `ls -l /var/www/media/upload.txt`\n4. Observe: `-rw-r--r--` (world-readable)\n\n## Expected Behavior\n\nThe default should be `0o644` only if safe, but at minimum should not be world-writable...",

  "patch": "diff --git a/django/core/files/storage.py b/django/core/files/storage.py\nindex abc..def 100644\n--- a/django/core/files/storage.py\n+++ b/django/core/files/storage.py\n@@ -290,7 +290,7 @@ class FileSystemStorage(...)  :\n     def _save(self, name, content):\n         ...\n-        os.chmod(full_path, self.file_permissions_mode)\n+        os.chmod(full_path, self.file_permissions_mode & 0o777)\n         ...",

  "test_patch": "diff --git a/tests/file_storage/tests.py b/tests/file_storage/tests.py\nindex 123..456 100644\n--- a/tests/file_storage/tests.py\n+++ b/tests/file_storage/tests.py\n@@ -50,6 +50,15 @@ class FileStorageTests(SimpleTestCase):\n     def test_file_upload_permissions(self):\n         ...\n+    def test_file_upload_permissions_not_world_writable(self):\n+        # 测试:默认权限不应是 world-writable\n+        ...\n",

  "fail_to_pass": [
    "tests.file_storage.tests.FileStorageTests.test_file_upload_permissions_not_world_writable"
  ],
  "pass_to_pass": [
    "tests.file_storage.tests.FileStorageTests.test_file_upload_permissions",
    "tests.file_storage.tests.FileStorageTests.test_save_overwrite",
    "tests.file_storage.tests.FileStorageTests.test_delete",
    "... 约 200 个其他测试 ..."
  ]
}
这 9 个字段之间的关系(一张图看懂)
                    problem_statement (题面)
                            ↓
                ┌─────────── agent ───────────┐
                │   读 base_commit 状态 + 题面  │
                │   写出 model_patch             │
                └─────────────────────────────┘
                            ↓
                       model_patch
                            ↓
              ┌─────────── 评测 ─────────────┐
              │  apply test_patch (定义 F2P)   │
              │  apply model_patch (agent 输出)│
              │  跑 fail_to_pass + pass_to_pass│
              └─────────────────────────────┘
                            ↓
                  ┌────── 判定 ──────┐
                  │  F2P 全过 + P2P 全过?  │
                  │    → resolved ✓    │
                  └─────────────────────┘

ground truth:  patch (agent 看不到,评测用不到,仅作分析)
flowchart LR subgraph Q["题面区(agent 可见)"] IID["instance_id"] PS["problem_statement
题面:问题描述"] REPO["repo"] BC["base_commit
评测起点"] end subgraph A["答案区(agent 不可见)"] PA["patch
ground truth 修复"] TP["test_patch
定义 F2P 的测试改动"] end subgraph C["评测契约区"] F2P["FAIL_TO_PASS
pre fail → post pass"] P2P["PASS_TO_PASS
pre pass → post pass"] VER["version
对齐环境 specs"] end AGENT["Agent
读题面 + base_commit 代码"] MP["model_patch
agent 输出的修复"] EVAL["评测
apply test_patch + model_patch
跑 F2P + P2P"] JUDGE{"判定
F2P 全过 ∧ P2P 全过?"} OK["resolved ✓"] PS --> AGENT BC --> AGENT AGENT --> MP MP --> EVAL TP --> EVAL F2P --> EVAL P2P --> EVAL EVAL --> JUDGE JUDGE -->|是| OK
图 1 · 9 字段三分区与评测流向:题面区喂给 agent,契约区约束评测,答案区仅作 ground truth

§ 4 fail_to_pass / pass_to_pass 字段详解

这两个字段是 F2P/P2P 评测的实际清单——告诉评测 harness 要跑哪些测试、怎么判定。

字段格式
  • 每个条目是 pytest 的 test node ID(完整路径)
  • 例子:"django.test.utils_tests.test_functional.TestDatetimeWarnings.test_naive_datetime_warning"
  • 格式:tests/<file>.py::<ClassName>::<test_method> 或 tests/<file>.py::<test_function>
F2P / P2P 怎么用(评测时)
# 1. 在 base_commit 状态跑测试(不 apply model_patch)
pre_results = pytest.run(test_ids = fail_to_pass + pass_to_pass)

# 2. apply model_patch + test_patch,跑同一批测试
post_results = pytest.run(test_ids = fail_to_pass + pass_to_pass)

# 3. 判定
for test_id in fail_to_pass:
    assert pre_results[test_id] == "FAIL"  # 起点 fail
    assert post_results[test_id] == "PASS"  # 修复后 pass

for test_id in pass_to_pass:
    assert pre_results[test_id] == "PASS"  # 起点 pass
    assert post_results[test_id] == "PASS"  # 修复后仍 pass

# 4. 全过 = resolved
resolved = all_f2p_passed and all_p2p_passed
F2P / P2P 测试规模典型值
字段典型规模说明
fail_to_pass1-5 个专门验证 bug 修了——很少,因为只需 1-2 个测试就能覆盖 bug
pass_to_pass200-500 个整个仓库的回归测试集——很多,因为要确保没破坏其他功能
F2P / P2P 不对称的设计意图:F2P 测试少而精(专门盯 bug),P2P 测试多而广(覆盖整个仓库)。这种设计让评测敏感但不冗余:模型得真修 bug 才过 F2P;改坏其他地方就挂 P2P。评测时间 ≈ 跑 P2P 测试时间,所以 Lite 子集 300 题需要几小时跑完(每题 ~5 分钟跑 P2P)。

§ 5 字段速查表(按字母序)

字段类型必填一句话
base_commitstring★40 位 hex git SHA-1,agent 改代码的起点
fail_to_passstring[]★F2P 测试 ID 列表(pytest node ID)
instance_idstring★主键,格式 <repo>__<issue-number>
pass_to_passstring[]★P2P 测试 ID 列表(防回归)
patchstring★fix PR 的 git diff(ground truth)
problem_statementstring★issue 文本(题面)
repostring★<owner>/<repo> 格式
test_patchstring★配套测试改动(定义 F2P + 修复 pre-fix 状态)
versionstring★Python 版本,对应 MAP_REPO_VERSION_TO_SPECS
所有 9 个字段都是必填的。SWE-bench 协议是严格 schema——任何字段缺失或类型错都会导致加载失败。本仓加载器会做严格校验,详见 §6。

§ 6 怎么加载数据(本仓零依赖加载器)

本仓 swe_bench_warehouse/api/_internal 提供纯标准库加载器(不依赖 datasets / pandas),教学友好——你只需会 json + pathlib 就能用。

3 个常用 API(Application Programming Interface,应用程序接口——这里指加载器暴露的函数入口)
API用途典型用法
load(name)加载整个数据集load("swe-bench") → list[Instance](2,294 条)
get(name, id)按 ID 查 1 条get("swe-bench", "django__django-10914")
filter_by_repo(name, repo)按 repo 过滤filter_by_repo("swe-bench", "django/django")
filter_by_language(name, lang)按语言过滤(Pro 变体用)filter_by_language("swe-bench-pro", "go")
stats(name)统计数据集元信息stats("swe-bench") → 总数 / 仓库数 / 语言数
search(name, kw)关键字搜索 issuesearch("swe-bench", "memory leak")
完整使用示例(Python)
import sys
sys.path.insert(0, ".")  # 把项目根目录加进 path
from swe_bench_warehouse.api._internal import load, get, filter_by_repo, stats

# 1. 加载整个数据集
instances = load("swe-bench")                # list[Instance] · 2,294 条
print(f"总数: {len(instances)}")
print(f"前 1 条 instance_id: {instances[0].instance_id}")

# 2. 按 ID 查
inst = get("swe-bench", "django__django-10914")
print(f"问题: {inst.problem_statement[:100]}...")
print(f"F2P 测试数: {len(inst.fail_to_pass)}")
print(f"P2P 测试数: {len(inst.pass_to_pass)}")

# 3. 按 repo 过滤(django 全部题)
django_all = filter_by_repo("swe-bench", "django/django")
print(f"django 全部题: {len(django_all)}")

# 4. 统计元信息
s = stats("swe-bench")
print(f"总数: {s.total}, 仓库数: {s.repos}, 语言数: {s.languages}")

# 5. 关键字搜索
hits = search("swe-bench", "memory leak")    # 在 problem_statement 里搜
print(f"含 'memory leak' 的题: {len(hits)}")
Instance 对象属性访问
属性对应字段类型
inst.instance_idinstance_idstr
inst.reporepostr
inst.base_commitbase_commitstr
inst.versionversionstr
inst.problem_statementproblem_statementstr
inst.patchpatchstr
inst.test_patchtest_patchstr
inst.fail_to_passfail_to_passList[str]
inst.pass_to_passpass_to_passList[str]
加载器设计哲学:本仓加载器纯标准库(json + pathlib + functools + collections)——不引入 datasets / pandas 等重依赖。教学友好:你只需会标准库即可浏览数据,不用先装环境。这是本仓"零依赖"原则在数据层的体现。

§ 7 本仓数据集清单(swe_bench_warehouse/)

本仓 swe_bench_warehouse/raw/ 收录了 2 个 SWE-bench 变体(其他变体未收录)。

已收录数据集
文件名来源任务数是否在 git
swe-bench.jsonlHuggingFace SWE-bench/SWE-bench2,294❌(约 380 MB,超 100 MB 限制)
swe-bench-lite.jsonlHuggingFace SWE-bench/SWE-bench_Lite300❌(同上原因)
拉取流程(首次使用)
# 方法 1:用 huggingface CLI(推荐)
pip install datasets
huggingface-cli download SWE-bench/SWE-bench_Lite \
  --repo-type dataset \
  --local-dir ./swe_bench_warehouse/raw/

# 方法 2:用本仓构建脚本
python -m database.build_swe_bench --lite

# 拉完后文件位置
ls -la swe_bench_warehouse/raw/
# swe-bench.jsonl         (380 MB, 2,294 条)
# swe-bench-lite.jsonl    (50 MB, 300 条)
为什么 Lite 是项目基线:Lite 是从 12 仓库各抽 25 道题的"代表性"子集(人工筛选过),~50 MB 跑得动、评测快(1-2 小时)、信噪比高(去掉了一些 issue 描述不清 / 测试不准的题)。本仓所有 demo 和开发都用 Lite,原版 2,294 题是正式评测时用。

引用与配套资料

来源链接 / 路径说明
数据集(HuggingFace)huggingface.co/datasets/SWE-bench/SWE-bench原版 2,294 题
数据集(Lite)huggingface.co/datasets/SWE-bench/SWE-bench_LiteLite 300 题(项目基线)
GitHub 仓库github.com/SWE-bench/SWE-bench数据集元信息 + 评测工具
本仓加载器源码swe_bench_warehouse/api/_internal零依赖加载器(标准库)
本仓数据集入仓swe_bench_warehouse/raw/jsonl 文件位置(.gitignore 排除)
本仓姊妹讲义swe-bench-intro.html背景与任务定义
本仓姊妹讲义swe-bench-evaluation.html评测流程与 S1-S7 映射
本仓 wikiwiki/50-swe-bench-paper.md论文 wiki 索引
本仓 wikiwiki/04-swe-bench-lite-tasks.mdLite 子集逐条说明