专题 · UML 建模 · 第 5 层
用 UML 拆解两个 Agent 的骨架
把 mini-swe-agent 与 SWE-agent 的源码反向拆成七张 UML 图——看清同样是「Agent 主循环」,在两个项目里为何会长成完全不同的样子;最后用两张行为图,直击「结构相似、行为差异」的本质。
生成时间:2026-09-18 · 版本 v0.2 · 生成 Agent:WorkBuddy · 载体:agentsoft-research-platform teaching-web-platform
概览
| 项目 | 说明 |
| 本页定位 | 给「架构演进调研」补可视化层。姊妹专题 02 用文字与表格讲清了两个项目的演进脉络,本页回答它没能画出来的部分:这些概念在源码里具体是怎么组织的 |
| 标本版本 | mini-swe-agent 04d809c(包版本 2.4.6)· SWE-agent 3ea751c。全部结论锚定到具体 commit,禁止用 main 分支 |
| 前置知识 | Agent harness 循环(模型 / 工具 / 观察三步走)、Agent 最小六部分。没读过先补这两讲 |
| 读完会什么 | 能画出两个项目的包依赖;能说出五种 Python 抽象机制分别对应哪种 UML 关系;能解释「为什么终止判据一个靠消息、一个靠返回值」——这是两个项目最本质的一处分岔;还能说出两者「结构相似、行为差异」的真相——SWE-agent 真正值钱的可复用经验是 分级错误兜底(requery / autosubmit)与 hook 扩展点设计 |
| 不讲什么 | 不讲 ACI 接口设计原则(见 SWE-agent 论文讲义)、不讲运行命令与参数调优、不讲性能评测数据 |
第一篇 · 包图
第一篇 · 代码怎么组织(第 1 图)
先看森林再看树:两个项目的包各有哪些,依赖箭头指向哪里。
图 1 · 包依赖对照
看清源码目录的分层与依赖方向
两侧的分层其实高度相似:都有一个「控制中枢」包、一个模型包、一个环境包、一个配置区。差别在于
SWE-agent 多出了 tools 与 types 两个包——前者承载工具系统,后者把跨模块传递的数据结构单独收拢,避免循环依赖。
mini-swe-agent 没有这两个包,因为它的工具只有一个 bash,数据结构就是普通的消息字典。
%% UML 包图 · mini-swe-agent vs SWE-agent 源码组织对照
%% 标本: mini 04d809c / SWE-agent 3ea751c
%% ⚠️ mermaid 无原生 Package Diagram 语法,本图以 flowchart + subgraph 近似表达
flowchart TB
subgraph PKG_MINI["「mini-swe-agent」源码组织 · commit 04d809c"]
direction TB
ID_mini_root["minisweagent/__init__
Model·Environment·Agent Protocol"]
ID_mini_run["run/
CLI 入口 (typer)"]
ID_mini_agents["agents/
DefaultAgent·InteractiveAgent"]
ID_mini_models["models/
Litellm·OpenRouter·Portkey·Requesty"]
ID_mini_envs["environments/
local·docker·singularity"]
ID_mini_cfg["config/
default.yaml·mini.yaml"]
ID_mini_run --> ID_mini_root
ID_mini_run --> ID_mini_agents
ID_mini_run --> ID_mini_models
ID_mini_run --> ID_mini_envs
ID_mini_agents --> ID_mini_models
ID_mini_agents --> ID_mini_envs
ID_mini_models --> ID_mini_root
ID_mini_envs --> ID_mini_root
ID_mini_cfg -.->|"渲染/装配"| ID_mini_run
end
subgraph PKG_SWE["「SWE-agent」源码组织 · commit 3ea751c"]
direction TB
ID_swe_types["types.py
StepOutput·TrajectoryStep"]
ID_swe_run["run/
RunSingle·RunBatch·hooks"]
ID_swe_agent["agent/
AbstractAgent·RetryAgent·DefaultAgent"]
ID_swe_env["environment/
SWEEnv·repo·hooks"]
ID_swe_tools["tools/
ToolHandler·parsing·commands"]
ID_swe_utils["utils/
log·patch 工具"]
ID_swe_run --> ID_swe_agent
ID_swe_run --> ID_swe_env
ID_swe_agent --> ID_swe_tools
ID_swe_agent --> ID_swe_types
ID_swe_tools --> ID_swe_types
ID_swe_env --> ID_swe_types
ID_swe_agent --> ID_swe_utils
ID_swe_env --> ID_swe_utils
end
ID_NOTE_EXC["横切关注点(两侧共用,不入包图主干)
mini: exceptions.py(中断型异常树)
SWE: run/hooks/ · environment/hooks/(生命周期钩子)"]
classDef cfg fill:#f5f7fb,stroke:#4f5dd5,stroke-width:1px
classDef root fill:#eef1f8,stroke:#185fa5,stroke-width:1px
classDef note fill:#fffbe6,stroke:#ba7517,stroke-width:1px,stroke-dasharray:4 3
class ID_mini_cfg,ID_swe_utils cfg
class ID_mini_root,ID_swe_types root
class ID_NOTE_EXC note
图 1 · 两个项目的包级依赖对照(以 flowchart 近似表达,mermaid 无原生包图语法 · mini 04d809c / SWE-agent 3ea751c)
核心知识点
- 两侧都遵循「上层依赖下层」:run 依赖 agents,agents 同时依赖 models 与 environments
- agents 是唯一同时依赖模型与环境的层——它就是主循环所在地,这个位置由依赖格局决定,不是命名巧合
- 配置层到业务层是虚线(渲染 / 装配关系),不是实线调用:YAML 不参与编译期依赖
- SWE-agent 的 types.py 是横切的数据契约层,被 agent / tools / environment 三方共同依赖
课堂实训
- 打开两个仓库,对照本图数一数每个包的 Python 文件数量,验证「谁最重」
- 尝试给 mini-swe-agent 加一个 tools 包,思考需要改动哪几条依赖箭头
思考与讨论
- SWE-agent 为什么需要单独的 types 包?(提示:读它的文件头注释,那里写了原因)
- 如果两个包互相 import 怎么办?本图是否出现了这种情况?
证据:agents/default.py:14 导入 Model 与 Environment
证据:SWE-agent/types.py:1 文件头注释说明跨模块原因
第二篇 · 类图
第二篇 · 契约长什么样(第 2 图)
同一个「Agent / Model / Environment」 triad,四个里面没有一个是同一个 Python 机制。
图 2 · 契约层对照
用一张图看全五种抽象机制
这是整套图里信息密度最高的一张,也是最有教学价值的一张。它会告诉你一个容易忽略的事实:
「契约」在 Python 里有五种表达强度,从最严格的 Protocol,到最松散的「压根没有契约」。
两个项目恰好把这五种都用上了。
%% UML 类图 · mini-swe-agent vs SWE-agent 契约层对照
%% 标本: mini 04d809c / SWE-agent 3ea751c
%% mermaid classDiagram 原生语法,支持 <<annotation>> 与五种关系
classDiagram
%% ============ mini-swe-agent 侧 ============
class mini_Agent {
<<interface>>
+config
+run(task) dict
+save(path) dict
}
class mini_Model {
<<interface>>
+config
+query(messages) dict
+format_message() dict
+format_observation_messages() list
+get_template_vars() dict
+serialize() dict
}
class mini_Environment {
<<interface>>
+config
+execute(action, cwd) dict
+get_template_vars() dict
+serialize() dict
}
class mini_DefaultAgent {
+messages list
+cost float
+n_calls int
+run(task) dict
+step() list
+query() dict
+execute_actions(message) list
+add_messages() list
}
class mini_LitellmModel {
+abort_exceptions list
+query(messages) dict
+_parse_actions(response) list
+_calculate_cost(response) dict
}
class mini_LocalEnvironment {
+execute(action, cwd) dict
-_check_finished(output) void
}
class mini_AgentConfig {
<<config extends BaseModel>>
+step_limit int
+cost_limit float
+wall_time_limit_seconds int
+max_consecutive_format_errors int
}
class mini_InterruptAgentFlow {
+messages tuple
}
class mini_Submitted {
}
class mini_FormatError {
}
mini_Agent <|.. mini_DefaultAgent : duck-typed impl
mini_Model <|.. mini_LitellmModel : duck-typed impl
mini_Environment <|.. mini_LocalEnvironment : duck-typed impl
mini_DefaultAgent *-- mini_LitellmModel : owns
mini_DefaultAgent *-- mini_LocalEnvironment : owns
mini_DefaultAgent --> mini_AgentConfig : configures
mini_DefaultAgent ..> mini_InterruptAgentFlow : catches
mini_InterruptAgentFlow <|-- mini_Submitted
mini_InterruptAgentFlow <|-- mini_FormatError
%% ============ SWE-agent 侧 ============
class swe_AbstractAgent {
<<semantic abstract>>
+step() StepOutput
+run() AgentRunResult
+from_config() Self
+add_hook() None
}
class swe_DefaultAgent {
+messages list
+trajectory Trajectory
+setup() None
+step() StepOutput
+run(env, problem) AgentRunResult
+forward(history) StepOutput
+forward_with_handling(history) StepOutput
+handle_action(step) StepOutput
}
class swe_AbstractModel {
<<abstract>>
+query(messages) dict
+stats InstanceStats
}
class swe_LiteLLMModel {
+query(messages) dict
}
class swe_SWEEnv {
<<no abstraction>>
+config EnvironmentConfig
+communicate(input) dict
+step(action) dict
}
class swe_StepOutput {
<<config>>
+done bool
+exit_status str
+submission str
+thought str
+action str
+output str
}
class swe_AbstractParseFunction {
<<abstract>>
+__call__(model_response) list
}
class swe_ToolHandler {
+config ToolConfig
+get_action(step) str
+should_block_action() bool
}
class swe_XmlThoughtActionParser {
<<parser impl>>
+__call__(model_response) list
}
class swe_EnvHook {
<<semantic abstract>>
+on_step_start() None
+on_step_done() None
}
class swe_RunHook {
<<semantic abstract>>
+on_instance_start() None
+on_instance_completed() None
}
swe_AbstractAgent <|-- swe_DefaultAgent
swe_AbstractModel <|-- swe_LiteLLMModel
swe_AbstractParseFunction <|-- swe_XmlThoughtActionParser
swe_DefaultAgent --> swe_AbstractModel : uses
swe_DefaultAgent --> swe_SWEEnv : uses
swe_DefaultAgent --> swe_ToolHandler : uses
swe_DefaultAgent --> swe_StepOutput : returns
swe_DefaultAgent ..> swe_RunHook : notifies
swe_SWEEnv ..> swe_EnvHook : notifies
swe_ToolHandler --> swe_AbstractParseFunction : delegates
图 2 · 契约层对照:五档抽象语义在同一张图里并列(mermaid classDiagram 原生语法)
核心知识点
- 第一档 «interface»:class X(Protocol) —— mini 的三大契约全部如此,虚线空心三角表示实现
- 第二档 «abstract»:ABC + @abstractmethod —— SWE-agent 的 AbstractModel 是真正的语言级抽象
- 第三档 «semantic_abstract»:类名叫 Abstract、方法体是省略号,但没有 ABC —— SWE-agent 的 AbstractAgent 正是如此,靠的是约定不是机制
- 第五档 «no_abstraction»:SWEEnv 从头到尾就一个类,没有任何抽象层。这也是答案,不是缺口
- mini 的 DefaultAgent 不继承任何东西却能赋值给 Agent 类型——这就是结构化子类型(鸭子类型),图上用虚线空心三角
课堂实训
- 用 inspect.getmro() 和 __abstractmethods__ 分别检查这五个类,亲手确认表格里的判定
- 给 AbstractAgent 加上 ABC 与 abstractmethod,看看 pytest 会不会报错——体会「约定」与「机制」的差别
思考与讨论
- 为什么 SWE-agent 的 Environment 不需要抽象层?如果明天要加第二种环境,会发生什么?
- 第三档「语义抽象」算不算设计缺陷?如果让你重构,你会补 ABC 还是删掉类名里的 Abstract?
证据:src/minisweagent/__init__.py:43,61,73
证据:sweagent/agent/agents.py:224 无 ABC 方法体为省略号
SWE-agent (NeurIPS 2024)
第三篇 · 顺序图
第三篇 · 主循环怎么走(第 3–4 图)
把时间维度拉开:一次任务里,谁在什么时候调用了谁。
图 3 · mini-swe-agent 主循环
看消息驱动的循环长什么样
mini 的主循环只有 37 行。注意一个关键现象:历史缓冲区(messages)被画成独立生命线,
所有写入都收敛到 add_messages 一处。另一个要点是三方分工——
提交由环境抛出、格式错误由模型抛出、预算超限由 Agent 抛出,谁最清楚就谁负责。
%% UML 顺序图 · mini-swe-agent 主循环(标本 commit 04d809c)
%% ⚠️ sequenceDiagram 不支持 subgraph,故双项目拆为两个文件并列对照
%% 证据锚点全部来自 DefaultAgent.run / step / query / execute_actions
sequenceDiagram
autonumber
participant U as 用户 CLI
participant A as DefaultAgent
participant H as MessageHistory
participant M as LitellmModel
participant E as LocalEnvironment
participant T as 轨迹文件
U->>A: run task
Note over A,H: run 清空历史 L91
A->>H: 追加 system 与 instance 两条种子消息 L92-95
loop while True L96 终止条件是末条消息 role 为 exit L122
A->>A: step 等于 execute_actions 包 query L126-128
alt 预算已超限 L132 或 L140
A-->>A: 抛 LimitsExceeded 或 TimeExceeded 携带 exit 消息
else 预算正常
A->>M: query messages L149
opt API 可重试失败
M->>M: retry 循环重试直到 abort_exceptions L82
end
M-->>A: 返回 message 含 actions 与 cost
A->>H: add_messages message L151
Note over A: 解析失败抛 FormatError actions_toolcall.py L41
end
A->>E: execute action L156
Note over E: execute 内调用 _check_finished local.py L42
alt 首行为 COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT local.py L48
E-->>A: 抛 Submitted 携带 submission local.py L50
else 正常命令输出
E-->>A: 返回 output 与 returncode 与 exception_info
end
A->>H: add_messages observation L157
alt 捕获 FormatError 且未连续超限 L100-114
A->>H: 追加错误消息后继续循环
else 连续超限 L104
A->>H: 追加 exit 消息 RepeatedFormatError
end
A->>T: save 落盘 finally L120-121 每步都写
end
A-->>U: 返回末条消息的 extra 即 exit_status 与 submission
图 3 · mini-swe-agent 一次完整任务的时序(sequenceDiagram 原生 · loop / alt / opt 三层控制流)
核心知识点
- 循环框架是 while True,退出靠末条消息 role == "exit"
- alt 表达「必选其一」的出口(预算 / 提交 / 正常),opt 表达「某些实现才有」的可选路径(API 重试),两者职责不同
- 每一步都在 finally 里落盘轨迹—— crash 也不丢数据
- 观测消息是追加进 messages,而不是停在 Agent 内部
课堂实训
- 用 30 行伪代码把这张图复述一遍,重点说出三个异常各自来自哪一层
- 在 add_messages 上加一行 print,观察一轮循环里消息条数的增长曲线
思考与讨论
- 为什么「提交」要让 Environment 来判断,而不是 Agent?(提示:谁先看到 shell 的输出)
- 把 exit 信号混在消息流里,有什么代价?如果你来设计,会把控制信号放在哪里?
证据:agents/default.py:96-124 run 主循环
证据:environments/local.py:48 提交标记检测
图 4 · SWE-agent 主循环
看返回值驱动的循环长什么样
同样是「主循环」,这里的结构明显更重:step 之下还有 forward_with_handling,
其内部再套一层重查询循环。工具执行要经过 ToolHandler 放行,
每一步前后还有 RunHook 钩子。请特别注意与图 3 的对照:终止判据是 StepOutput.done 这个布尔字段,不再依赖消息里的某个标记。
%% UML 顺序图 · SWE-agent 主循环(标本 commit 3ea751c)
%% ⚠️ sequenceDiagram 不支持 subgraph,故双项目拆为两个文件并列对照
%% 证据锚点全部来自 agents.py 的 run / step / forward_with_handling / forward / handle_action
sequenceDiagram
autonumber
participant U as RunSingle
participant A as DefaultAgent
participant H as messages 历史
participant P as ToolHandler 与 parser
participant M as LiteLLMModel
participant E as SWEEnv
participant K as RunHook 钩子
U->>A: run env 与 problem_statement L1265
A->>A: setup 装工具并写入模板消息 L1279
A->>K: on_run_start L1282
loop while not step_output.done L1284
A->>K: on_step_start L1248
A->>A: step L1235 转 forward_with_handling self.messages L1252
loop 内层重查询循环 n_format_fails 小于 max_requeries L1107
A->>M: forward 调用 query L1006
P->>P: parse_actions 生成 actions
A->>P: ToolHandler 判定动作是否放行 L936
alt FormatError 或 BlockedAction 或 BashSyntax L1120-1141
Note over A: requery 不退出循环,次数累加
else 触发退出类异常 L1152-1210
A-->>A: attempt_autosubmission_after_error 兜底提交 patch
else 返回正常 StepOutput
A->>E: handle_action 执行 L936
E-->>A: observation 含 returncode
end
end
A->>H: add_step_to_history step_output L1253
A->>A: save_trajectory L1286 每步落盘
A->>K: on_step_done L1262
end
A->>K: on_run_done L1287
A-->>U: 返回 AgentRunResult 含 info 与 trajectory
图 4 · SWE-agent 一次完整任务的时序(双层循环:外层 step + 内层 requery)
核心知识点
- 外层循环条件:while not step_output.done(结构化返回值驱动)
- 内层重查询循环:while n_format_fails < max_requeries,格式错在这层消化掉,不打扰外层——这是 mini 没有的一层
- 单步返回值是 StepOutput(Pydantic 模型),而不是消息字典
- RunHook 在 run 开始/结束、step 开始/结束四个时点被通知,这是扩展能力的接入方式
课堂实训
- 把图 3 与图 4 叠在一起看:标出两侧「同一件事」各自用了多少层调用
- 故意让模型输出一段格式错误的 YAML,观察内层重查询循环被触发几次
思考与讨论
- 「返回值驱动」比「消息流驱动」好在哪里?代价是什么?
- 内层重查询循环存在的前提是什么?(提示:弱模型才会频繁产出错误格式)
证据:agents.py:1284 外层循环条件
证据:agents.py:1107 内层重查询循环
第四篇 · 状态机图
第四篇 · 什么时候停、能不能重试(第 5–6 图)
终态与可恢复态的分界,是可靠性设计最诚实的一面。
图 5 · mini-swe-agent 状态机
五个终态全部收敛到同一个 except
mini 的状态机有一个精巧之处:所有中断型异常都继承同一个基类 InterruptAgentFlow,
于是主循环只需一处 except InterruptAgentFlow 就能统一收集,再把携带的消息追加进历史,
最后由单一判据检出。「收敛」是这张图的关键词。
%% UML 状态机图 · mini-swe-agent(标本 commit 04d809c)
%% 迁移标签统一 [event]/action 三段式;终态出边仅指向 [*]
%% 证据:agents/default.py run L96-124 / query L130-152 / execute_actions L154-157;exceptions.py
%% 关键事实:所有中断型异常均由同一处 except InterruptAgentFlow(L115) 收敛成 role=exit 消息
stateDiagram-v2
[*] --> Initializing : [agent.run() L88]/messages 清空
Initializing --> Thinking : [写入种子消息 L92-95]/messages += system 与 instance
Thinking --> Acting : [query() 成功 L149]/n_calls++ 且 cost 累加
Thinking --> Terminating : [check_limits() L132]/raise LimitsExceeded
Thinking --> Terminating : [check_wall_time() L140]/raise TimeExceeded
Thinking --> FormatError : [_parse_actions 失败 actions_toolcall L41]/raise FormatError
Thinking --> Failed : [未捕获异常 L117]/追加 exit 消息后 raise 逃逸
FormatError --> Thinking : [单次失败 L113-114]/cost 补记后继续循环
FormatError --> Terminating : [连续超限 L104-112]/追加 RepeatedFormatError 的 exit 消息
Acting --> Observing : [execute_actions L156-157]/messages += observation
Acting --> Terminating : [_check_finished local.py L48-50]/raise Submitted
Observing --> Thinking : [末条消息 role 不是 exit L122]/进入下一轮 while True
Observing --> Failed : [未捕获异常 L117]/raise 逃逸
Terminating --> Submitted : [exit_status 为 Submitted L122]/break 返回 submission
Terminating --> LimitsExceeded : [exit_status 为 LimitsExceeded]/break 返回空 submission
Terminating --> TimeExceeded : [exit_status 为 TimeExceeded]/break 返回空 submission
Terminating --> RepeatedFormatError : [exit_status 为 RepeatedFormatError]/break 返回空 submission
Submitted --> [*] : [done]/正常结束
LimitsExceeded --> [*] : [done]/预算耗尽
TimeExceeded --> [*] : [done]/超时
RepeatedFormatError --> [*] : [done]/格式连续失败
Failed --> [*] : [done]/异常向上传播
note right of Terminating
收敛点:except InterruptAgentFlow L115-116
统一把异常携带的消息追加进历史,
再由 L122 单一判据检出 role == exit
end note
note left of Observing
落盘时机:finally 子句每步调用 save L120-121
end note
图 5 · mini-swe-agent 状态迁移(迁移标签为 UML 标准的 [event]/action 三段式)
核心知识点
- Terminating 是收敛态:四类终态异常在此汇合,统一转成 exit 消息
- 可恢复态只有 FormatError 一个:单次失败回到 Thinking,连续超限才进 Terminating
- 五个终态:Submitted / LimitsExceeded / TimeExceeded / RepeatedFormatError / Failed,出边一律指向终态
- 迁移标签统一写成 [事件]/动作,且没有用 Submitted 同时当终态名和事件名——这是 UML 最容易出错的地方
课堂实训
- 把 max_consecutive_format_errors 设为 1,观察 Agent 的行为变化
- 给每个终态写一句「出现时我应该怎么排查」,形成一张排错表
思考与讨论
- 为什么 Failed 要 raise 逃逸而不是变成 exit 消息?(提示:区分「Agent 判断无法继续」与「程序本身出错」)
- 如果这个循环连续 N 次 FormatError 仍未收敛,状态机应该怎么改?
证据:exceptions.py:1-27 异常继承树
证据:agents/default.py:115 统一收敛点
图 6 · SWE-agent 状态机
九类退出先兜底再结束
这张图最值得注意的构件是 Autosubmitting:SWE-agent 在绝大多数异常退出前,
会先尝试从环境里把改动导出成 patch 兜底提交。这在 mini 里完全不存在。
它不是多余的好心,而是为「强模型会犯的错」买的保险——这恰恰是复杂度迁移最典型的样例。
%% UML 状态机图 · SWE-agent(标本 commit 3ea751c)
%% 迁移标签统一 [event]/action 三段式;终态出边仅指向 [*]
%% 证据:agents.py run L1265-1294 / step L1235 / forward_with_handling L1062-1218;types.py StepOutput
%% 关键事实:终止靠结构化返回对象 StepOutput.done,而非消息流;九类异常先 autosubmit 兜底
stateDiagram-v2
[*] --> Initializing : [RunSingle.from_config]/构造 RunSingle
Initializing --> Setup : [agent.run() L1279]/setup 装工具并写模板消息
Setup --> Thinking : [on_run_start L1282]/进入 while not done
Thinking --> Acting : [forward L1006 查询成功]/StepOutput 含 action
Thinking --> Requery : [forward_with_handling L1107]/n_format_fails++ 后重试
Requery --> Thinking : [_BlockedAction L1125 或 BashSyntax L1135]/history 换错误消息后重试
Requery --> Thinking : [RetryWithOutput L1142]/追加观测后重试
Requery --> Autosubmitting : [n_format_fails 达到 max_requeries L1211]/exit_format
Acting --> Observing : [handle_action L936 执行成功]/observation 写入 StepOutput
Acting --> Autosubmitting : [_ExitForfeit L1154]/exit_forfeit
Acting --> Autosubmitting : [CommandTimeoutError L1168]/exit_command_timeout
Observing --> Thinking : [step_output.done 为 False L1284]/add_step_to_history 后下一步
Observing --> Done : [submission 已产生 L1255]/done 置为 True
Thinking --> Autosubmitting : [ContextWindowExceeded L1175]/exit_context
Thinking --> Autosubmitting : [CostLimitExceeded L1182]/exit_cost
Thinking --> Autosubmitting : [TotalExecutionTimeExceeded L1161]/exit_total_execution_time
Thinking --> Failed : [TotalCostLimitExceeded L1180]/唯一直接 raise 不兜底
Autosubmitting --> Done : [attempt_autosubmission_after_error L823]/从环境提取 patch 兜底提交
Done --> [*] : [while 条件不再成立]/返回 AgentRunResult
Failed --> [*] : [done]/异常向上传播
note right of Autosubmitting
兜底提交是 SWE-agent 独有机制:
出错时仍尝试从环境导出 patch 提交,
mini-swe-agent 没有对应设计
end note
note left of Observing
落盘时机:save_trajectory L1286 每步调用
同时 RunHook.on_step_done L1262 触发
end note
图 6 · SWE-agent 状态迁移(注意 Autosubmitting 兜底状态是 mini 所没有的)
核心知识点
- exit_status 多达十余种(exit_cost / exit_context / exit_environment_error …),粒度远细于 mini
- Requery 是可恢复态,三种异常在此重试而不上抛
- 唯一绕过兜底的是 TotalCostLimitExceededError——总预算超了不能继续花钱,直接 raise
- 每次 step 结束都要经过 add_step_to_history 与落盘
课堂实训
- 把图 5 与图 6 的终态数量、可恢复态数量做成对照表,讨论「多出来的这些带来了什么」
- 找到 attempt_autosubmission_after_error,读一遍它到底从环境里取了什么
思考与讨论
- 兜底提交是好的设计吗?它会不会掩盖真正的失败?
- 如果模型已经足够强,这九类兜底里有哪些可以删掉?(对应姊妹专题 02 讲的「复杂度迁移」)
证据:agents.py:1076-1086 兜底提交实现
证据:agents.py:1180 唯一不兜底的分支
第五篇 · 组件图
第五篇 · 运行时怎么装配(第 7 图)
包图看「源码怎么摆」,组件图看「跑起来以后谁被谁拿起来了」。
图 7 · 运行时装配对照
两种装配哲学:数据驱动 vs 代码驱动
这张图回答一个实践问题:我要换一个 Model 实现,需要改多少地方?
mini 的做法是字符串映射表加 importlib,加一行字典即可;
SWE-agent 用的是显式 if-elif 分支,得改控制流代码。前者是数据驱动,后者是代码驱动。
%% UML 组件图 · mini-swe-agent vs SWE-agent 运行时装配对照
%% 标本: mini 04d809c / SWE-agent 3ea751c
%% ⚠️ mermaid 无原生 Component Diagram 语法,本图以 flowchart + subgraph 近似表达
%% 与包图的分工:包图画「源码怎么组织」,组件图画「运行时怎么装配」
flowchart TB
subgraph RUN_MINI["「mini-swe-agent」运行时装配"]
direction TB
ID_m_cli["『CLI 构件』 typer app
run/mini.py main"]
ID_m_cfg["「config 构件」
YAML + Jinja2 模板"]
ID_m_fac["『装配构件』 工厂三件套
get_model / get_environment / get_agent"]
ID_m_agent["『控制构件』 DefaultAgent"]
ID_m_model["『模型构件』 LitellmModel"]
ID_m_env["『环境构件』 LocalEnvironment"]
ID_m_hist["『存储构件』 messages + 轨迹 JSON"]
ID_m_cfg -.->|«configures»| ID_m_cli
ID_m_cli -->|"读配置, 调工厂"| ID_m_fac
ID_m_cfg -.->|«configures»| ID_m_fac
ID_m_fac -->|"注入"| ID_m_agent
ID_m_fac -->|"构造"| ID_m_model
ID_m_fac -->|"构造"| ID_m_env
ID_m_agent ==>|"持有并调用"| ID_m_model
ID_m_agent ==>|"持有并调用"| ID_m_env
ID_m_agent -->|"每步写"| ID_m_hist
end
subgraph RUN_SWE["「SWE-agent」运行时装配"]
direction TB
ID_s_cli["『CLI 构件』 RunSingle / RunBatch"]
ID_s_cfg["「config 构件」
YAML + Jinja2 模板"]
ID_s_fac["『装配构件』 get_agent_from_config
显式 if-elif 分支"]
ID_s_agent["『控制构件』 DefaultAgent"]
ID_s_model["『模型构件』 LiteLLMModel"]
ID_s_tools["『工具构件』 ToolHandler
含 parser 家族"]
ID_s_env["『环境构件』 SWEEnv
含 swerex 容器会话"]
ID_s_hook["『扩展构件』 RunHook / EnvHook"]
ID_s_out["『存储构件』 StepOutput + Trajectory"]
ID_s_cfg -.->|«configures»| ID_s_cli
ID_s_cli -->|"读配置"| ID_s_fac
ID_s_cfg -.->|«configures»| ID_s_fac
ID_s_fac -->|"注入"| ID_s_agent
ID_s_agent ==>|"调用"| ID_s_model
ID_s_agent ==>|"调用"| ID_s_tools
ID_s_tools -->|"解析后交给"| ID_s_env
ID_s_agent ==>|"调用"| ID_s_env
ID_s_agent -.->|"通知"| ID_s_hook
ID_s_env -.->|"通知"| ID_s_hook
ID_s_agent -->|"每步写"| ID_s_out
end
ID_NOTE_DIFF["装配方式差异
mini: 字符串映射表 + importlib 动态导入(数据驱动, 加一行即可扩展)
SWE: if-elif 显式分支( agents.py L242-254, 加一种要改代码)"]
classDef cfg fill:#f5f7fb,stroke:#4f5dd5,stroke-width:1px
classDef store fill:#eef1f8,stroke:#185fa5,stroke-width:1px
classDef note fill:#fffbe6,stroke:#ba7517,stroke-width:1px,stroke-dasharray:4 3
class ID_m_cfg,ID_s_cfg cfg
class ID_m_hist,ID_s_out store
class ID_NOTE_DIFF note
图 7 · 运行时装配对照(以 flowchart 近似表达 · 配置构件用虚线标注 «configures» 关系)
核心知识点
- 配置构件(«config»)与业务构件之间是虚线的装配关系,不是实线调用
- Agent 对被替代构件的粗实心箭头表示持有并调用——生命周期绑定
- SWE-agent 比 mini 多出「工具构件」与「扩展构件」两类运行期协作者
- 存储形态不同:mini 写消息字典 + 轨迹 JSON,SWE 写结构化 StepOutput + Trajectory
课堂实训
- 给 mini-swe-agent 加一种新的 Environment 实现,记录你一共改了几处文件
- 对照 agents/__init__.py:8 的映射表与 SWE-agent agents.py:242 的 if-elif,说出各自的扩展成本
思考与讨论
- 数据驱动装配一定更好吗?它在可发现性(能不能一眼看出有哪些实现)上有什么代价?
- 如果让你为本项目的某个模块做同样的装配改造,你会选哪种?为什么?
证据:src/minisweagent/agents/__init__.py:8-19 字符串映射表
证据:sweagent/agent/agents.py:242-254 显式分支
第六篇 · 泳道活动图
第六篇 · 错误处理行为差异(第 8 图)
前面五张图都在回答「结构」,这一张回答「行为」——两个项目的主循环几乎相同,真正拉开差距的是错误处理兜底策略。
图 8 · 错误处理取舍 vs 结构化对照
为什么 Mini 190 行、SWE 数千行——差全在错误兜底
前面五张图可能给你一种错觉:两个项目结构挺像。这句话只说对了一半。主循环确实几乎一样(query → parse → execute → observe → repeat 一条直线),
但把它放到「错误处理」维度看,差异立刻被放大。这张活动图用三条颜色的分支标出两种处理哲学:
Mini 只有一层 requery(蓝色),连错三次就直终;SWE 额外多了一层 autosubmit 兜底(橙色),九类异常都会先尝试把改动导成 patch 提交再退出。这正是「迷你 vs 生产级」的本质差。
%% UML 泳道活动图 · mini-swe-agent vs SWE-agent 错误处理与重试行为对照
%% 标本: mini 04d809c / SWE-agent 3ea751c
%% ⚠️ mermaid 无原生 Activity Diagram / swimlane,以 flowchart 上下双泳道近似
%% 三色约定:主链(灰)正常流转 │ 蓝(requery) 可恢复重试 │ 橙(autosubmit) SWE 独有兜底 │ 红(terminate) 直接终止
flowchart LR
classDef requery fill:#e1ecff,stroke:#1f6fd6,stroke-width:1.5px,color:#0b3f8f
classDef autosub fill:#ffe4cc,stroke:#d97b1f,stroke-width:1.5px,color:#8a4b08
classDef term fill:#ffd9d9,stroke:#d6453d,stroke-width:1.5px,color:#8a1210
classDef main fill:#f5f7fa,stroke:#5b6572,stroke-width:1px
subgraph mini_sw["「mini-swe-agent」错误处理 · default.py run L96-124"]
direction TB
m1["run(task) L88
注入两条种子消息"]
m2{{"while True L96"}}
m3["step = query + execute_actions L126"]
mQ{{"query() L130 查预算"}}
m4["model.query L149 + add_messages"]
m5{{"parse 成功?"}}
mE["env.execute L156
_check_finished 判提交"]
mO["format_observation 追加 L157"]
mN["role == exit ? L122"]
mEnd(("返回 exit_status"))
mT1["LimitsExceeded L133"]:::term
mT2["TimeExceeded L141"]:::term
mT3["RepeatedFormatError L109"]:::term
mF["FormatError 累计 L100"]:::requery
mX["无 autosubmit 兜底
连续错 = 直接结束"]:::autosub
m1 --> m2
m2 -->|每次| m3
m3 --> mQ
mQ -->|超限| mT1 & mT2
mQ -->|正常| m4
m4 --> m5
m5 -->|成功| mE
m5 -->|失败| mF
mE --> mO --> mN
mN -->|exit| mEnd
mN -->|非 exit| m2
mF -->|单次| m2
mF -->|超3次| mT3
mX -.-> mT3
mT1 & mT2 & mT3 --> mEnd
end
subgraph swe_sw["「SWE-agent」错误处理 · agents.py forward_with_handling L1062-1218"]
direction LR
s1["run(env,problem) L1265"]
s2{{"while not done L1284"}}
s3["step() L1235"]
s4{{"n_format_fails < max_requeries L1107"}}
s5["forward→query L1006"]
s6{{"parse 结果?"}}
s7["ToolHandler 放行 L936"]
s8["env.execute L936"]
sE["add_step_to_history L1253
save_trajectory L1286"]
sEnd(("返回 AgentRunResult"))
sR1["requery 分支 L1120-1140"]:::requery
sA1["_ExitForfeit L1154"]:::autosub
sA2["TotalExecTime L1161"]:::autosub
sA3["CommandTimeout L1168
ContextWindow L1175"]:::autosub
sA4["CostLimit L1182
Retry/Swerex/RuntimeError L1187-1207"]:::autosub
sA0{{"异常 → 统一 autosubmit L1152"}}:::autosub
sA5["exit_format 兜底 L1215"]:::autosub
sAS["attempt_autosubmission_after_error
导 patch 提交 L823-868"]:::autosub
sT["TotalCostLimit
唯一直接 raise L1180"]:::term
s1 --> s2
s2 -->|每次| s3
s3 --> s4
s4 -->|正常| s5
s5 --> s6
s6 -->|Format/Blocked/BashSyntax L1120-1140| sR1
s6 -->|正常| s7
s7 --> s8 --> sE -->|done=False| s2
sE -->|done=True| sEnd
sR1 -->|未超限| s4
sR1 -->|达到上限 L1211| sA5
s6 -->|退出类异常| sA0
s8 -->|CommandTimeout L1168| sA0
sA0 --> sA1 & sA2 & sA3 & sA4
sA0 --> sAS
sA1 & sA2 & sA3 & sA4 --> sA5
sAS --> sEnd
s6 -->|总成本超限| sT
sT --> sEnd
end
图 8 · 错误处理活动图:主循环几乎一致(灰线),差异集中在蓝色 requery 重试层与橙色 autosubmit 兜底层(以 flowchart 双泳道近似 · mermaid 无原生活动图语法)
核心知识点
- 横轴是「结构相似」:两者的主链都是 query → parse → execute → observe → repeat,这条贯穿两泳道的灰线正是「核心循环几乎一样」的图证
- 纵轴是「行为差异」:Mini 只有一层 requery(蓝),FormatError 单次回到循环、连续 3 次直终;SWE 在此之上多出 autosubmit(橙)——九类异常统一走 attempt_autosubmission_after_error(agents.py:823)导 patch 兜底提交
- 红色分支是唯一不兜底处:TotalCostLimitExceededError(agents.py:1180)总预算超了直接 raise,因为「继续花钱」本身没有意义
- Mini 图的橙色框是一个「缺位」标注——它提醒学生:mini 不是「也有 autosubmit」,而是「刻意没有」。这是复杂度迁移的直观呈现
课堂实训
- 给模型喂一段会连续产出错误格式的 system prompt,分别数两个项目在终止前各自重试了多少次
- 在 SWE-agent 里把 max_requeries 调成 0,观察原本会被 requery 吸收的格式错如何立刻触发 autosubmit 兜底
思考与讨论
- 为什么 mini-swe-agent 不实现 autosubmit?「减少代码」是否等价于「减少能力」?
- autosubmit 把失败伪装成成功了吗?它能提高 SWE-bench 得分,但会不会提高「看似解决实则不然」的比例?
证据:agents.py:823 兜底提交
证据:agents.py:1180 唯一不兜底
第七篇 · Hook 织入时序图
第七篇 · 扩展点设计(第 9 图)
「怎么在不改主循环的前提下加能力?」SWE-agent 的答案是 hook——把主循环走一圈,标注 13 个可插入回调的点。
图 9 · SWE-agent 扩展点织入
13 个回调点,就是「不用改主逻辑就能加能力」的设计
mini-swe-agent 要加能力靠继承 override(改子类);SWE-agent 则把观察点抽成 AbstractAgentHook,
通过 add_hook(agents.py:519)注册进 CombinedAgentHook,在主循环的 setup / run / step / query 各阶段织入 13 个回调。
这张图的价值在于:它让你看清「哪里可以挂钩子」——这直接决定了这个框架的扩展边界。
%% UML Hook 织入时序图 · SWE-agent 扩展点(标本 commit 3ea751c)
%% 一帧 step 内密集织入多回调;13 个回调定义见 hooks/abstract.py
%% 教学落点:新增能力只需加一个 AbstractAgentHook 子类经 add_hook 注入,不改主循环
sequenceDiagram
autonumber
participant R as RunSingle
participant A as DefaultAgent
(agents.py)
participant H as CombinedAgentHook
(hooks/abstract.py)
participant M as AbstractModel
participant E as SWEEnv
Note over R,A: === 装配 ===
A->>H: add_hook(hook) 注册自定义钩子 L519
A->>A: on_init L287 构造期钩子(慎用)
Note over A,E: === setup 阶段(一次) ===
A->>H: on_tools_installation_started L592
A->>H: on_setup_attempt L594
A->>H: on_setup_done L606
Note over A,H: === run 外层 ===
A->>H: on_run_start L1282
loop while not step_output.done L1284
A->>H: on_step_start L1248
rect rgb(235, 245, 255)
Note over A,M: requery 内层循环 while n_format_fails < max_requeries L1107
A->>H: on_query_message_added L558 (每加一条消息)
A->>H: on_model_query L1029 (带完整 history 查询)
A->>M: query(history)
M-->>A: model response
A->>H: on_actions_generated L1051 (动作已生成未执行)
end
A->>H: on_action_started L959
A->>E: handle_action→execute L936
E-->>A: observation
A->>H: on_action_executed L992
A->>A: add_step_to_history + save_trajectory L1253/1286
A->>H: on_step_done L1262
end
A->>H: on_run_done L1287 (整个 run 结束)
R-->>A: 返回 AgentRunResult
图 9 · hook 织入时序:13 个回调点在 setup / run / step / query 各阶段的织入时机(sequenceDiagram 原生语法)
核心知识点
- 三类阶段对应三种扩展意图:setup 钩子管「准备」,step/query/action 钩子管「执行中观测」,run/step_done 钩子管「收尾记录」
- 最密集处是 requery 内层循环:on_query_message_added(每条消息)、on_model_query(每次真正查询)、on_actions_generated(每次动作产出)
- 对比图 8 的橙框:Mini 没有 hook,相当于把「观测能力」硬编码进主循环;SWE 把观测抽成回调,新增能力(轨迹记录 / 风险检查)只需加一个 hook
- 这个「不改主循环、只加观察点」的思路不是 SWE-agent 独有——你在本项目的 Agent 适配层也会看到类似的扩展方式,可自行对照 experiment_modules/solving/agent_runtime/ 印证
课堂实训
- 写一个 10 行的 hook 子类,override on_action_started 打印每次动作,运行后确认它被调用
- 数一遍:为「记录每一步耗时」这个能力,Mini 需要改主循环几处,SWE 只需 new 一个 hook?
思考与讨论
- hook 回调点太多会不会过度设计?哪些阶段你会合并?
- 如果 mini-swe-agent 想加 hook,哪个点最值得先加?(提示:对照你上一讲学到的上下文压缩需求)
证据:agents.py:519 add_hook 注入
证据:hooks/abstract.py:11-53 13 回调定义
已知局限(诚实交代):
① 本页的包图、组件图与泳道活动图是用 flowchart 近似表达的——mermaid 没有原生 UML 包图 / 组件图 / 活动图(swimlane)语法,已在 MODEL_MANIFEST 说明;
② AbstractAgent / EnvHook / RunHook 的依据是「命名 + 空方法体」而非 ABC 机制,属 medium 置信度;
③ 本次为纯静态建模,未运行被测代码,运行时的动态覆盖(monkey patch 等)不体现在图中。