门户首页
专题 · UML 建模 · 第 5 层

用 UML 拆解两个 Agent 的骨架

把 mini-swe-agent 与 SWE-agent 的源码反向拆成七张 UML 图——看清同样是「Agent 主循环」,在两个项目里为何会长成完全不同的样子;最后用两张行为图,直击「结构相似、行为差异」的本质。

生成时间:2026-09-18 · 版本 v0.2 · 生成 Agent:WorkBuddy · 载体:agentsoft-research-platform teaching-web-platform

概览

7
产出图数
2
对照项目
112
源码锚点
100%
证据覆盖率
项目说明
本页定位给「架构演进调研」补可视化层。姊妹专题 02 用文字与表格讲清了两个项目的演进脉络,本页回答它没能画出来的部分:这些概念在源码里具体是怎么组织的
标本版本mini-swe-agent 04d809c(包版本 2.4.6)· SWE-agent 3ea751c。全部结论锚定到具体 commit,禁止用 main 分支
前置知识Agent harness 循环(模型 / 工具 / 观察三步走)、Agent 最小六部分。没读过先补这两讲
读完会什么能画出两个项目的包依赖;能说出五种 Python 抽象机制分别对应哪种 UML 关系;能解释「为什么终止判据一个靠消息、一个靠返回值」——这是两个项目最本质的一处分岔;还能说出两者「结构相似、行为差异」的真相——SWE-agent 真正值钱的可复用经验是 分级错误兜底(requery / autosubmit)与 hook 扩展点设计
不讲什么不讲 ACI 接口设计原则(见 SWE-agent 论文讲义)、不讲运行命令与参数调优、不讲性能评测数据

第一篇 · 包图

第一篇 · 代码怎么组织(第 1 图)
先看森林再看树:两个项目的包各有哪些,依赖箭头指向哪里。
图 1 · 包依赖对照

看清源码目录的分层与依赖方向

两侧的分层其实高度相似:都有一个「控制中枢」包、一个模型包、一个环境包、一个配置区。差别在于 SWE-agent 多出了 tools 与 types 两个包——前者承载工具系统,后者把跨模块传递的数据结构单独收拢,避免循环依赖。 mini-swe-agent 没有这两个包,因为它的工具只有一个 bash,数据结构就是普通的消息字典。

%% UML 包图 · mini-swe-agent vs SWE-agent 源码组织对照 %% 标本: mini 04d809c / SWE-agent 3ea751c %% ⚠️ mermaid 无原生 Package Diagram 语法,本图以 flowchart + subgraph 近似表达 flowchart TB subgraph PKG_MINI["「mini-swe-agent」源码组织 · commit 04d809c"] direction TB ID_mini_root["minisweagent/__init__
Model·Environment·Agent Protocol"] ID_mini_run["run/
CLI 入口 (typer)"] ID_mini_agents["agents/
DefaultAgent·InteractiveAgent"] ID_mini_models["models/
Litellm·OpenRouter·Portkey·Requesty"] ID_mini_envs["environments/
local·docker·singularity"] ID_mini_cfg["config/
default.yaml·mini.yaml"] ID_mini_run --> ID_mini_root ID_mini_run --> ID_mini_agents ID_mini_run --> ID_mini_models ID_mini_run --> ID_mini_envs ID_mini_agents --> ID_mini_models ID_mini_agents --> ID_mini_envs ID_mini_models --> ID_mini_root ID_mini_envs --> ID_mini_root ID_mini_cfg -.->|"渲染/装配"| ID_mini_run end subgraph PKG_SWE["「SWE-agent」源码组织 · commit 3ea751c"] direction TB ID_swe_types["types.py
StepOutput·TrajectoryStep"] ID_swe_run["run/
RunSingle·RunBatch·hooks"] ID_swe_agent["agent/
AbstractAgent·RetryAgent·DefaultAgent"] ID_swe_env["environment/
SWEEnv·repo·hooks"] ID_swe_tools["tools/
ToolHandler·parsing·commands"] ID_swe_utils["utils/
log·patch 工具"] ID_swe_run --> ID_swe_agent ID_swe_run --> ID_swe_env ID_swe_agent --> ID_swe_tools ID_swe_agent --> ID_swe_types ID_swe_tools --> ID_swe_types ID_swe_env --> ID_swe_types ID_swe_agent --> ID_swe_utils ID_swe_env --> ID_swe_utils end ID_NOTE_EXC["横切关注点(两侧共用,不入包图主干)
mini: exceptions.py(中断型异常树)
SWE: run/hooks/ · environment/hooks/(生命周期钩子)"] classDef cfg fill:#f5f7fb,stroke:#4f5dd5,stroke-width:1px classDef root fill:#eef1f8,stroke:#185fa5,stroke-width:1px classDef note fill:#fffbe6,stroke:#ba7517,stroke-width:1px,stroke-dasharray:4 3 class ID_mini_cfg,ID_swe_utils cfg class ID_mini_root,ID_swe_types root class ID_NOTE_EXC note
图 1 · 两个项目的包级依赖对照(以 flowchart 近似表达,mermaid 无原生包图语法 · mini 04d809c / SWE-agent 3ea751c)
核心知识点
  • 两侧都遵循「上层依赖下层」:run 依赖 agents,agents 同时依赖 models 与 environments
  • agents 是唯一同时依赖模型与环境的层——它就是主循环所在地,这个位置由依赖格局决定,不是命名巧合
  • 配置层到业务层是虚线(渲染 / 装配关系),不是实线调用:YAML 不参与编译期依赖
  • SWE-agent 的 types.py 是横切的数据契约层,被 agent / tools / environment 三方共同依赖
课堂实训
  • 打开两个仓库,对照本图数一数每个包的 Python 文件数量,验证「谁最重」
  • 尝试给 mini-swe-agent 加一个 tools 包,思考需要改动哪几条依赖箭头
思考与讨论
  • SWE-agent 为什么需要单独的 types 包?(提示:读它的文件头注释,那里写了原因)
  • 如果两个包互相 import 怎么办?本图是否出现了这种情况?
证据:agents/default.py:14 导入 Model 与 Environment 证据:SWE-agent/types.py:1 文件头注释说明跨模块原因

第二篇 · 类图

第二篇 · 契约长什么样(第 2 图)
同一个「Agent / Model / Environment」 triad,四个里面没有一个是同一个 Python 机制。
图 2 · 契约层对照

用一张图看全五种抽象机制

这是整套图里信息密度最高的一张,也是最有教学价值的一张。它会告诉你一个容易忽略的事实: 「契约」在 Python 里有五种表达强度,从最严格的 Protocol,到最松散的「压根没有契约」。 两个项目恰好把这五种都用上了。

%% UML 类图 · mini-swe-agent vs SWE-agent 契约层对照 %% 标本: mini 04d809c / SWE-agent 3ea751c %% mermaid classDiagram 原生语法,支持 <<annotation>> 与五种关系 classDiagram %% ============ mini-swe-agent 侧 ============ class mini_Agent { <<interface>> +config +run(task) dict +save(path) dict } class mini_Model { <<interface>> +config +query(messages) dict +format_message() dict +format_observation_messages() list +get_template_vars() dict +serialize() dict } class mini_Environment { <<interface>> +config +execute(action, cwd) dict +get_template_vars() dict +serialize() dict } class mini_DefaultAgent { +messages list +cost float +n_calls int +run(task) dict +step() list +query() dict +execute_actions(message) list +add_messages() list } class mini_LitellmModel { +abort_exceptions list +query(messages) dict +_parse_actions(response) list +_calculate_cost(response) dict } class mini_LocalEnvironment { +execute(action, cwd) dict -_check_finished(output) void } class mini_AgentConfig { <<config extends BaseModel>> +step_limit int +cost_limit float +wall_time_limit_seconds int +max_consecutive_format_errors int } class mini_InterruptAgentFlow { +messages tuple } class mini_Submitted { } class mini_FormatError { } mini_Agent <|.. mini_DefaultAgent : duck-typed impl mini_Model <|.. mini_LitellmModel : duck-typed impl mini_Environment <|.. mini_LocalEnvironment : duck-typed impl mini_DefaultAgent *-- mini_LitellmModel : owns mini_DefaultAgent *-- mini_LocalEnvironment : owns mini_DefaultAgent --> mini_AgentConfig : configures mini_DefaultAgent ..> mini_InterruptAgentFlow : catches mini_InterruptAgentFlow <|-- mini_Submitted mini_InterruptAgentFlow <|-- mini_FormatError %% ============ SWE-agent 侧 ============ class swe_AbstractAgent { <<semantic abstract>> +step() StepOutput +run() AgentRunResult +from_config() Self +add_hook() None } class swe_DefaultAgent { +messages list +trajectory Trajectory +setup() None +step() StepOutput +run(env, problem) AgentRunResult +forward(history) StepOutput +forward_with_handling(history) StepOutput +handle_action(step) StepOutput } class swe_AbstractModel { <<abstract>> +query(messages) dict +stats InstanceStats } class swe_LiteLLMModel { +query(messages) dict } class swe_SWEEnv { <<no abstraction>> +config EnvironmentConfig +communicate(input) dict +step(action) dict } class swe_StepOutput { <<config>> +done bool +exit_status str +submission str +thought str +action str +output str } class swe_AbstractParseFunction { <<abstract>> +__call__(model_response) list } class swe_ToolHandler { +config ToolConfig +get_action(step) str +should_block_action() bool } class swe_XmlThoughtActionParser { <<parser impl>> +__call__(model_response) list } class swe_EnvHook { <<semantic abstract>> +on_step_start() None +on_step_done() None } class swe_RunHook { <<semantic abstract>> +on_instance_start() None +on_instance_completed() None } swe_AbstractAgent <|-- swe_DefaultAgent swe_AbstractModel <|-- swe_LiteLLMModel swe_AbstractParseFunction <|-- swe_XmlThoughtActionParser swe_DefaultAgent --> swe_AbstractModel : uses swe_DefaultAgent --> swe_SWEEnv : uses swe_DefaultAgent --> swe_ToolHandler : uses swe_DefaultAgent --> swe_StepOutput : returns swe_DefaultAgent ..> swe_RunHook : notifies swe_SWEEnv ..> swe_EnvHook : notifies swe_ToolHandler --> swe_AbstractParseFunction : delegates
图 2 · 契约层对照:五档抽象语义在同一张图里并列(mermaid classDiagram 原生语法)
核心知识点
  • 第一档 «interface»:class X(Protocol) —— mini 的三大契约全部如此,虚线空心三角表示实现
  • 第二档 «abstract»:ABC + @abstractmethod —— SWE-agent 的 AbstractModel 是真正的语言级抽象
  • 第三档 «semantic_abstract»:类名叫 Abstract、方法体是省略号,但没有 ABC —— SWE-agent 的 AbstractAgent 正是如此,靠的是约定不是机制
  • 第五档 «no_abstraction»:SWEEnv 从头到尾就一个类,没有任何抽象层。这也是答案,不是缺口
  • mini 的 DefaultAgent 不继承任何东西却能赋值给 Agent 类型——这就是结构化子类型(鸭子类型),图上用虚线空心三角
课堂实训
  • 用 inspect.getmro() 和 __abstractmethods__ 分别检查这五个类,亲手确认表格里的判定
  • 给 AbstractAgent 加上 ABC 与 abstractmethod,看看 pytest 会不会报错——体会「约定」与「机制」的差别
思考与讨论
  • 为什么 SWE-agent 的 Environment 不需要抽象层?如果明天要加第二种环境,会发生什么?
  • 第三档「语义抽象」算不算设计缺陷?如果让你重构,你会补 ABC 还是删掉类名里的 Abstract?
证据:src/minisweagent/__init__.py:43,61,73 证据:sweagent/agent/agents.py:224 无 ABC 方法体为省略号 SWE-agent (NeurIPS 2024)

第三篇 · 顺序图

第三篇 · 主循环怎么走(第 3–4 图)
把时间维度拉开:一次任务里,谁在什么时候调用了谁。
图 3 · mini-swe-agent 主循环

看消息驱动的循环长什么样

mini 的主循环只有 37 行。注意一个关键现象:历史缓冲区(messages)被画成独立生命线, 所有写入都收敛到 add_messages 一处。另一个要点是三方分工—— 提交由环境抛出、格式错误由模型抛出、预算超限由 Agent 抛出,谁最清楚就谁负责。

%% UML 顺序图 · mini-swe-agent 主循环(标本 commit 04d809c) %% ⚠️ sequenceDiagram 不支持 subgraph,故双项目拆为两个文件并列对照 %% 证据锚点全部来自 DefaultAgent.run / step / query / execute_actions sequenceDiagram autonumber participant U as 用户 CLI participant A as DefaultAgent participant H as MessageHistory participant M as LitellmModel participant E as LocalEnvironment participant T as 轨迹文件 U->>A: run task Note over A,H: run 清空历史 L91 A->>H: 追加 system 与 instance 两条种子消息 L92-95 loop while True L96 终止条件是末条消息 role 为 exit L122 A->>A: step 等于 execute_actions 包 query L126-128 alt 预算已超限 L132 或 L140 A-->>A: 抛 LimitsExceeded 或 TimeExceeded 携带 exit 消息 else 预算正常 A->>M: query messages L149 opt API 可重试失败 M->>M: retry 循环重试直到 abort_exceptions L82 end M-->>A: 返回 message 含 actions 与 cost A->>H: add_messages message L151 Note over A: 解析失败抛 FormatError actions_toolcall.py L41 end A->>E: execute action L156 Note over E: execute 内调用 _check_finished local.py L42 alt 首行为 COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT local.py L48 E-->>A: 抛 Submitted 携带 submission local.py L50 else 正常命令输出 E-->>A: 返回 output 与 returncode 与 exception_info end A->>H: add_messages observation L157 alt 捕获 FormatError 且未连续超限 L100-114 A->>H: 追加错误消息后继续循环 else 连续超限 L104 A->>H: 追加 exit 消息 RepeatedFormatError end A->>T: save 落盘 finally L120-121 每步都写 end A-->>U: 返回末条消息的 extra 即 exit_status 与 submission
图 3 · mini-swe-agent 一次完整任务的时序(sequenceDiagram 原生 · loop / alt / opt 三层控制流)
核心知识点
  • 循环框架是 while True,退出靠末条消息 role == "exit"
  • alt 表达「必选其一」的出口(预算 / 提交 / 正常),opt 表达「某些实现才有」的可选路径(API 重试),两者职责不同
  • 每一步都在 finally 里落盘轨迹—— crash 也不丢数据
  • 观测消息是追加进 messages,而不是停在 Agent 内部
课堂实训
  • 用 30 行伪代码把这张图复述一遍,重点说出三个异常各自来自哪一层
  • 在 add_messages 上加一行 print,观察一轮循环里消息条数的增长曲线
思考与讨论
  • 为什么「提交」要让 Environment 来判断,而不是 Agent?(提示:谁先看到 shell 的输出)
  • 把 exit 信号混在消息流里,有什么代价?如果你来设计,会把控制信号放在哪里?
证据:agents/default.py:96-124 run 主循环 证据:environments/local.py:48 提交标记检测
图 4 · SWE-agent 主循环

看返回值驱动的循环长什么样

同样是「主循环」,这里的结构明显更重:step 之下还有 forward_with_handling, 其内部再套一层重查询循环。工具执行要经过 ToolHandler 放行, 每一步前后还有 RunHook 钩子。请特别注意与图 3 的对照:终止判据是 StepOutput.done 这个布尔字段,不再依赖消息里的某个标记。

%% UML 顺序图 · SWE-agent 主循环(标本 commit 3ea751c) %% ⚠️ sequenceDiagram 不支持 subgraph,故双项目拆为两个文件并列对照 %% 证据锚点全部来自 agents.py 的 run / step / forward_with_handling / forward / handle_action sequenceDiagram autonumber participant U as RunSingle participant A as DefaultAgent participant H as messages 历史 participant P as ToolHandler 与 parser participant M as LiteLLMModel participant E as SWEEnv participant K as RunHook 钩子 U->>A: run env 与 problem_statement L1265 A->>A: setup 装工具并写入模板消息 L1279 A->>K: on_run_start L1282 loop while not step_output.done L1284 A->>K: on_step_start L1248 A->>A: step L1235 转 forward_with_handling self.messages L1252 loop 内层重查询循环 n_format_fails 小于 max_requeries L1107 A->>M: forward 调用 query L1006 P->>P: parse_actions 生成 actions A->>P: ToolHandler 判定动作是否放行 L936 alt FormatError 或 BlockedAction 或 BashSyntax L1120-1141 Note over A: requery 不退出循环,次数累加 else 触发退出类异常 L1152-1210 A-->>A: attempt_autosubmission_after_error 兜底提交 patch else 返回正常 StepOutput A->>E: handle_action 执行 L936 E-->>A: observation 含 returncode end end A->>H: add_step_to_history step_output L1253 A->>A: save_trajectory L1286 每步落盘 A->>K: on_step_done L1262 end A->>K: on_run_done L1287 A-->>U: 返回 AgentRunResult 含 info 与 trajectory
图 4 · SWE-agent 一次完整任务的时序(双层循环:外层 step + 内层 requery)
核心知识点
  • 外层循环条件:while not step_output.done(结构化返回值驱动)
  • 内层重查询循环:while n_format_fails < max_requeries,格式错在这层消化掉,不打扰外层——这是 mini 没有的一层
  • 单步返回值是 StepOutput(Pydantic 模型),而不是消息字典
  • RunHook 在 run 开始/结束、step 开始/结束四个时点被通知,这是扩展能力的接入方式
课堂实训
  • 把图 3 与图 4 叠在一起看:标出两侧「同一件事」各自用了多少层调用
  • 故意让模型输出一段格式错误的 YAML,观察内层重查询循环被触发几次
思考与讨论
  • 「返回值驱动」比「消息流驱动」好在哪里?代价是什么?
  • 内层重查询循环存在的前提是什么?(提示:弱模型才会频繁产出错误格式)
证据:agents.py:1284 外层循环条件 证据:agents.py:1107 内层重查询循环

第四篇 · 状态机图

第四篇 · 什么时候停、能不能重试(第 5–6 图)
终态与可恢复态的分界,是可靠性设计最诚实的一面。
图 5 · mini-swe-agent 状态机

五个终态全部收敛到同一个 except

mini 的状态机有一个精巧之处:所有中断型异常都继承同一个基类 InterruptAgentFlow, 于是主循环只需一处 except InterruptAgentFlow 就能统一收集,再把携带的消息追加进历史, 最后由单一判据检出。「收敛」是这张图的关键词。

%% UML 状态机图 · mini-swe-agent(标本 commit 04d809c) %% 迁移标签统一 [event]/action 三段式;终态出边仅指向 [*] %% 证据:agents/default.py run L96-124 / query L130-152 / execute_actions L154-157;exceptions.py %% 关键事实:所有中断型异常均由同一处 except InterruptAgentFlow(L115) 收敛成 role=exit 消息 stateDiagram-v2 [*] --> Initializing : [agent.run() L88]/messages 清空 Initializing --> Thinking : [写入种子消息 L92-95]/messages += system 与 instance Thinking --> Acting : [query() 成功 L149]/n_calls++ 且 cost 累加 Thinking --> Terminating : [check_limits() L132]/raise LimitsExceeded Thinking --> Terminating : [check_wall_time() L140]/raise TimeExceeded Thinking --> FormatError : [_parse_actions 失败 actions_toolcall L41]/raise FormatError Thinking --> Failed : [未捕获异常 L117]/追加 exit 消息后 raise 逃逸 FormatError --> Thinking : [单次失败 L113-114]/cost 补记后继续循环 FormatError --> Terminating : [连续超限 L104-112]/追加 RepeatedFormatError 的 exit 消息 Acting --> Observing : [execute_actions L156-157]/messages += observation Acting --> Terminating : [_check_finished local.py L48-50]/raise Submitted Observing --> Thinking : [末条消息 role 不是 exit L122]/进入下一轮 while True Observing --> Failed : [未捕获异常 L117]/raise 逃逸 Terminating --> Submitted : [exit_status 为 Submitted L122]/break 返回 submission Terminating --> LimitsExceeded : [exit_status 为 LimitsExceeded]/break 返回空 submission Terminating --> TimeExceeded : [exit_status 为 TimeExceeded]/break 返回空 submission Terminating --> RepeatedFormatError : [exit_status 为 RepeatedFormatError]/break 返回空 submission Submitted --> [*] : [done]/正常结束 LimitsExceeded --> [*] : [done]/预算耗尽 TimeExceeded --> [*] : [done]/超时 RepeatedFormatError --> [*] : [done]/格式连续失败 Failed --> [*] : [done]/异常向上传播 note right of Terminating 收敛点:except InterruptAgentFlow L115-116 统一把异常携带的消息追加进历史, 再由 L122 单一判据检出 role == exit end note note left of Observing 落盘时机:finally 子句每步调用 save L120-121 end note
图 5 · mini-swe-agent 状态迁移(迁移标签为 UML 标准的 [event]/action 三段式)
核心知识点
  • Terminating 是收敛态:四类终态异常在此汇合,统一转成 exit 消息
  • 可恢复态只有 FormatError 一个:单次失败回到 Thinking,连续超限才进 Terminating
  • 五个终态:Submitted / LimitsExceeded / TimeExceeded / RepeatedFormatError / Failed,出边一律指向终态
  • 迁移标签统一写成 [事件]/动作,且没有用 Submitted 同时当终态名和事件名——这是 UML 最容易出错的地方
课堂实训
  • 把 max_consecutive_format_errors 设为 1,观察 Agent 的行为变化
  • 给每个终态写一句「出现时我应该怎么排查」,形成一张排错表
思考与讨论
  • 为什么 Failed 要 raise 逃逸而不是变成 exit 消息?(提示:区分「Agent 判断无法继续」与「程序本身出错」)
  • 如果这个循环连续 N 次 FormatError 仍未收敛,状态机应该怎么改?
证据:exceptions.py:1-27 异常继承树 证据:agents/default.py:115 统一收敛点
图 6 · SWE-agent 状态机

九类退出先兜底再结束

这张图最值得注意的构件是 Autosubmitting:SWE-agent 在绝大多数异常退出前, 会先尝试从环境里把改动导出成 patch 兜底提交。这在 mini 里完全不存在。 它不是多余的好心,而是为「强模型会犯的错」买的保险——这恰恰是复杂度迁移最典型的样例。

%% UML 状态机图 · SWE-agent(标本 commit 3ea751c) %% 迁移标签统一 [event]/action 三段式;终态出边仅指向 [*] %% 证据:agents.py run L1265-1294 / step L1235 / forward_with_handling L1062-1218;types.py StepOutput %% 关键事实:终止靠结构化返回对象 StepOutput.done,而非消息流;九类异常先 autosubmit 兜底 stateDiagram-v2 [*] --> Initializing : [RunSingle.from_config]/构造 RunSingle Initializing --> Setup : [agent.run() L1279]/setup 装工具并写模板消息 Setup --> Thinking : [on_run_start L1282]/进入 while not done Thinking --> Acting : [forward L1006 查询成功]/StepOutput 含 action Thinking --> Requery : [forward_with_handling L1107]/n_format_fails++ 后重试 Requery --> Thinking : [_BlockedAction L1125 或 BashSyntax L1135]/history 换错误消息后重试 Requery --> Thinking : [RetryWithOutput L1142]/追加观测后重试 Requery --> Autosubmitting : [n_format_fails 达到 max_requeries L1211]/exit_format Acting --> Observing : [handle_action L936 执行成功]/observation 写入 StepOutput Acting --> Autosubmitting : [_ExitForfeit L1154]/exit_forfeit Acting --> Autosubmitting : [CommandTimeoutError L1168]/exit_command_timeout Observing --> Thinking : [step_output.done 为 False L1284]/add_step_to_history 后下一步 Observing --> Done : [submission 已产生 L1255]/done 置为 True Thinking --> Autosubmitting : [ContextWindowExceeded L1175]/exit_context Thinking --> Autosubmitting : [CostLimitExceeded L1182]/exit_cost Thinking --> Autosubmitting : [TotalExecutionTimeExceeded L1161]/exit_total_execution_time Thinking --> Failed : [TotalCostLimitExceeded L1180]/唯一直接 raise 不兜底 Autosubmitting --> Done : [attempt_autosubmission_after_error L823]/从环境提取 patch 兜底提交 Done --> [*] : [while 条件不再成立]/返回 AgentRunResult Failed --> [*] : [done]/异常向上传播 note right of Autosubmitting 兜底提交是 SWE-agent 独有机制: 出错时仍尝试从环境导出 patch 提交, mini-swe-agent 没有对应设计 end note note left of Observing 落盘时机:save_trajectory L1286 每步调用 同时 RunHook.on_step_done L1262 触发 end note
图 6 · SWE-agent 状态迁移(注意 Autosubmitting 兜底状态是 mini 所没有的)
核心知识点
  • exit_status 多达十余种(exit_cost / exit_context / exit_environment_error …),粒度远细于 mini
  • Requery 是可恢复态,三种异常在此重试而不上抛
  • 唯一绕过兜底的是 TotalCostLimitExceededError——总预算超了不能继续花钱,直接 raise
  • 每次 step 结束都要经过 add_step_to_history 与落盘
课堂实训
  • 把图 5 与图 6 的终态数量、可恢复态数量做成对照表,讨论「多出来的这些带来了什么」
  • 找到 attempt_autosubmission_after_error,读一遍它到底从环境里取了什么
思考与讨论
  • 兜底提交是好的设计吗?它会不会掩盖真正的失败?
  • 如果模型已经足够强,这九类兜底里有哪些可以删掉?(对应姊妹专题 02 讲的「复杂度迁移」)
证据:agents.py:1076-1086 兜底提交实现 证据:agents.py:1180 唯一不兜底的分支

第五篇 · 组件图

第五篇 · 运行时怎么装配(第 7 图)
包图看「源码怎么摆」,组件图看「跑起来以后谁被谁拿起来了」。
图 7 · 运行时装配对照

两种装配哲学:数据驱动 vs 代码驱动

这张图回答一个实践问题:我要换一个 Model 实现,需要改多少地方? mini 的做法是字符串映射表加 importlib,加一行字典即可; SWE-agent 用的是显式 if-elif 分支,得改控制流代码。前者是数据驱动,后者是代码驱动。

%% UML 组件图 · mini-swe-agent vs SWE-agent 运行时装配对照 %% 标本: mini 04d809c / SWE-agent 3ea751c %% ⚠️ mermaid 无原生 Component Diagram 语法,本图以 flowchart + subgraph 近似表达 %% 与包图的分工:包图画「源码怎么组织」,组件图画「运行时怎么装配」 flowchart TB subgraph RUN_MINI["「mini-swe-agent」运行时装配"] direction TB ID_m_cli["『CLI 构件』 typer app
run/mini.py main"] ID_m_cfg["「config 构件」
YAML + Jinja2 模板"] ID_m_fac["『装配构件』 工厂三件套
get_model / get_environment / get_agent"] ID_m_agent["『控制构件』 DefaultAgent"] ID_m_model["『模型构件』 LitellmModel"] ID_m_env["『环境构件』 LocalEnvironment"] ID_m_hist["『存储构件』 messages + 轨迹 JSON"] ID_m_cfg -.->|«configures»| ID_m_cli ID_m_cli -->|"读配置, 调工厂"| ID_m_fac ID_m_cfg -.->|«configures»| ID_m_fac ID_m_fac -->|"注入"| ID_m_agent ID_m_fac -->|"构造"| ID_m_model ID_m_fac -->|"构造"| ID_m_env ID_m_agent ==>|"持有并调用"| ID_m_model ID_m_agent ==>|"持有并调用"| ID_m_env ID_m_agent -->|"每步写"| ID_m_hist end subgraph RUN_SWE["「SWE-agent」运行时装配"] direction TB ID_s_cli["『CLI 构件』 RunSingle / RunBatch"] ID_s_cfg["「config 构件」
YAML + Jinja2 模板"] ID_s_fac["『装配构件』 get_agent_from_config
显式 if-elif 分支"] ID_s_agent["『控制构件』 DefaultAgent"] ID_s_model["『模型构件』 LiteLLMModel"] ID_s_tools["『工具构件』 ToolHandler
含 parser 家族"] ID_s_env["『环境构件』 SWEEnv
含 swerex 容器会话"] ID_s_hook["『扩展构件』 RunHook / EnvHook"] ID_s_out["『存储构件』 StepOutput + Trajectory"] ID_s_cfg -.->|«configures»| ID_s_cli ID_s_cli -->|"读配置"| ID_s_fac ID_s_cfg -.->|«configures»| ID_s_fac ID_s_fac -->|"注入"| ID_s_agent ID_s_agent ==>|"调用"| ID_s_model ID_s_agent ==>|"调用"| ID_s_tools ID_s_tools -->|"解析后交给"| ID_s_env ID_s_agent ==>|"调用"| ID_s_env ID_s_agent -.->|"通知"| ID_s_hook ID_s_env -.->|"通知"| ID_s_hook ID_s_agent -->|"每步写"| ID_s_out end ID_NOTE_DIFF["装配方式差异
mini: 字符串映射表 + importlib 动态导入(数据驱动, 加一行即可扩展)
SWE: if-elif 显式分支( agents.py L242-254, 加一种要改代码)"] classDef cfg fill:#f5f7fb,stroke:#4f5dd5,stroke-width:1px classDef store fill:#eef1f8,stroke:#185fa5,stroke-width:1px classDef note fill:#fffbe6,stroke:#ba7517,stroke-width:1px,stroke-dasharray:4 3 class ID_m_cfg,ID_s_cfg cfg class ID_m_hist,ID_s_out store class ID_NOTE_DIFF note
图 7 · 运行时装配对照(以 flowchart 近似表达 · 配置构件用虚线标注 «configures» 关系)
核心知识点
  • 配置构件(«config»)与业务构件之间是虚线的装配关系,不是实线调用
  • Agent 对被替代构件的粗实心箭头表示持有并调用——生命周期绑定
  • SWE-agent 比 mini 多出「工具构件」与「扩展构件」两类运行期协作者
  • 存储形态不同:mini 写消息字典 + 轨迹 JSON,SWE 写结构化 StepOutput + Trajectory
课堂实训
  • 给 mini-swe-agent 加一种新的 Environment 实现,记录你一共改了几处文件
  • 对照 agents/__init__.py:8 的映射表与 SWE-agent agents.py:242 的 if-elif,说出各自的扩展成本
思考与讨论
  • 数据驱动装配一定更好吗?它在可发现性(能不能一眼看出有哪些实现)上有什么代价?
  • 如果让你为本项目的某个模块做同样的装配改造,你会选哪种?为什么?
证据:src/minisweagent/agents/__init__.py:8-19 字符串映射表 证据:sweagent/agent/agents.py:242-254 显式分支

第六篇 · 泳道活动图

第六篇 · 错误处理行为差异(第 8 图)
前面五张图都在回答「结构」,这一张回答「行为」——两个项目的主循环几乎相同,真正拉开差距的是错误处理兜底策略。
图 8 · 错误处理取舍 vs 结构化对照

为什么 Mini 190 行、SWE 数千行——差全在错误兜底

前面五张图可能给你一种错觉:两个项目结构挺像。这句话只说对了一半。主循环确实几乎一样(query → parse → execute → observe → repeat 一条直线), 但把它放到「错误处理」维度看,差异立刻被放大。这张活动图用三条颜色的分支标出两种处理哲学: Mini 只有一层 requery(蓝色),连错三次就直终;SWE 额外多了一层 autosubmit 兜底(橙色),九类异常都会先尝试把改动导成 patch 提交再退出。这正是「迷你 vs 生产级」的本质差。

%% UML 泳道活动图 · mini-swe-agent vs SWE-agent 错误处理与重试行为对照 %% 标本: mini 04d809c / SWE-agent 3ea751c %% ⚠️ mermaid 无原生 Activity Diagram / swimlane,以 flowchart 上下双泳道近似 %% 三色约定:主链(灰)正常流转 │ 蓝(requery) 可恢复重试 │ 橙(autosubmit) SWE 独有兜底 │ 红(terminate) 直接终止 flowchart LR classDef requery fill:#e1ecff,stroke:#1f6fd6,stroke-width:1.5px,color:#0b3f8f classDef autosub fill:#ffe4cc,stroke:#d97b1f,stroke-width:1.5px,color:#8a4b08 classDef term fill:#ffd9d9,stroke:#d6453d,stroke-width:1.5px,color:#8a1210 classDef main fill:#f5f7fa,stroke:#5b6572,stroke-width:1px subgraph mini_sw["「mini-swe-agent」错误处理 · default.py run L96-124"] direction TB m1["run(task) L88
注入两条种子消息"] m2{{"while True L96"}} m3["step = query + execute_actions L126"] mQ{{"query() L130 查预算"}} m4["model.query L149 + add_messages"] m5{{"parse 成功?"}} mE["env.execute L156
_check_finished 判提交"] mO["format_observation 追加 L157"] mN["role == exit ? L122"] mEnd(("返回 exit_status")) mT1["LimitsExceeded L133"]:::term mT2["TimeExceeded L141"]:::term mT3["RepeatedFormatError L109"]:::term mF["FormatError 累计 L100"]:::requery mX["无 autosubmit 兜底
连续错 = 直接结束"]:::autosub m1 --> m2 m2 -->|每次| m3 m3 --> mQ mQ -->|超限| mT1 & mT2 mQ -->|正常| m4 m4 --> m5 m5 -->|成功| mE m5 -->|失败| mF mE --> mO --> mN mN -->|exit| mEnd mN -->|非 exit| m2 mF -->|单次| m2 mF -->|超3次| mT3 mX -.-> mT3 mT1 & mT2 & mT3 --> mEnd end subgraph swe_sw["「SWE-agent」错误处理 · agents.py forward_with_handling L1062-1218"] direction LR s1["run(env,problem) L1265"] s2{{"while not done L1284"}} s3["step() L1235"] s4{{"n_format_fails < max_requeries L1107"}} s5["forward→query L1006"] s6{{"parse 结果?"}} s7["ToolHandler 放行 L936"] s8["env.execute L936"] sE["add_step_to_history L1253
save_trajectory L1286"] sEnd(("返回 AgentRunResult")) sR1["requery 分支 L1120-1140"]:::requery sA1["_ExitForfeit L1154"]:::autosub sA2["TotalExecTime L1161"]:::autosub sA3["CommandTimeout L1168
ContextWindow L1175"]:::autosub sA4["CostLimit L1182
Retry/Swerex/RuntimeError L1187-1207"]:::autosub sA0{{"异常 → 统一 autosubmit L1152"}}:::autosub sA5["exit_format 兜底 L1215"]:::autosub sAS["attempt_autosubmission_after_error
导 patch 提交 L823-868"]:::autosub sT["TotalCostLimit
唯一直接 raise L1180"]:::term s1 --> s2 s2 -->|每次| s3 s3 --> s4 s4 -->|正常| s5 s5 --> s6 s6 -->|Format/Blocked/BashSyntax L1120-1140| sR1 s6 -->|正常| s7 s7 --> s8 --> sE -->|done=False| s2 sE -->|done=True| sEnd sR1 -->|未超限| s4 sR1 -->|达到上限 L1211| sA5 s6 -->|退出类异常| sA0 s8 -->|CommandTimeout L1168| sA0 sA0 --> sA1 & sA2 & sA3 & sA4 sA0 --> sAS sA1 & sA2 & sA3 & sA4 --> sA5 sAS --> sEnd s6 -->|总成本超限| sT sT --> sEnd end
图 8 · 错误处理活动图:主循环几乎一致(灰线),差异集中在蓝色 requery 重试层与橙色 autosubmit 兜底层(以 flowchart 双泳道近似 · mermaid 无原生活动图语法)
核心知识点
  • 横轴是「结构相似」:两者的主链都是 query → parse → execute → observe → repeat,这条贯穿两泳道的灰线正是「核心循环几乎一样」的图证
  • 纵轴是「行为差异」:Mini 只有一层 requery(蓝),FormatError 单次回到循环、连续 3 次直终;SWE 在此之上多出 autosubmit(橙)——九类异常统一走 attempt_autosubmission_after_error(agents.py:823)导 patch 兜底提交
  • 红色分支是唯一不兜底处:TotalCostLimitExceededError(agents.py:1180)总预算超了直接 raise,因为「继续花钱」本身没有意义
  • Mini 图的橙色框是一个「缺位」标注——它提醒学生:mini 不是「也有 autosubmit」,而是「刻意没有」。这是复杂度迁移的直观呈现
课堂实训
  • 给模型喂一段会连续产出错误格式的 system prompt,分别数两个项目在终止前各自重试了多少次
  • 在 SWE-agent 里把 max_requeries 调成 0,观察原本会被 requery 吸收的格式错如何立刻触发 autosubmit 兜底
思考与讨论
  • 为什么 mini-swe-agent 不实现 autosubmit?「减少代码」是否等价于「减少能力」?
  • autosubmit 把失败伪装成成功了吗?它能提高 SWE-bench 得分,但会不会提高「看似解决实则不然」的比例?
证据:agents.py:823 兜底提交 证据:agents.py:1180 唯一不兜底

第七篇 · Hook 织入时序图

第七篇 · 扩展点设计(第 9 图)
「怎么在不改主循环的前提下加能力?」SWE-agent 的答案是 hook——把主循环走一圈,标注 13 个可插入回调的点。
图 9 · SWE-agent 扩展点织入

13 个回调点,就是「不用改主逻辑就能加能力」的设计

mini-swe-agent 要加能力靠继承 override(改子类);SWE-agent 则把观察点抽成 AbstractAgentHook, 通过 add_hook(agents.py:519)注册进 CombinedAgentHook,在主循环的 setup / run / step / query 各阶段织入 13 个回调。 这张图的价值在于:它让你看清「哪里可以挂钩子」——这直接决定了这个框架的扩展边界。

%% UML Hook 织入时序图 · SWE-agent 扩展点(标本 commit 3ea751c) %% 一帧 step 内密集织入多回调;13 个回调定义见 hooks/abstract.py %% 教学落点:新增能力只需加一个 AbstractAgentHook 子类经 add_hook 注入,不改主循环 sequenceDiagram autonumber participant R as RunSingle participant A as DefaultAgent
(agents.py) participant H as CombinedAgentHook
(hooks/abstract.py) participant M as AbstractModel participant E as SWEEnv Note over R,A: === 装配 === A->>H: add_hook(hook) 注册自定义钩子 L519 A->>A: on_init L287 构造期钩子(慎用) Note over A,E: === setup 阶段(一次) === A->>H: on_tools_installation_started L592 A->>H: on_setup_attempt L594 A->>H: on_setup_done L606 Note over A,H: === run 外层 === A->>H: on_run_start L1282 loop while not step_output.done L1284 A->>H: on_step_start L1248 rect rgb(235, 245, 255) Note over A,M: requery 内层循环 while n_format_fails < max_requeries L1107 A->>H: on_query_message_added L558 (每加一条消息) A->>H: on_model_query L1029 (带完整 history 查询) A->>M: query(history) M-->>A: model response A->>H: on_actions_generated L1051 (动作已生成未执行) end A->>H: on_action_started L959 A->>E: handle_action→execute L936 E-->>A: observation A->>H: on_action_executed L992 A->>A: add_step_to_history + save_trajectory L1253/1286 A->>H: on_step_done L1262 end A->>H: on_run_done L1287 (整个 run 结束) R-->>A: 返回 AgentRunResult
图 9 · hook 织入时序:13 个回调点在 setup / run / step / query 各阶段的织入时机(sequenceDiagram 原生语法)
核心知识点
  • 三类阶段对应三种扩展意图:setup 钩子管「准备」,step/query/action 钩子管「执行中观测」,run/step_done 钩子管「收尾记录」
  • 最密集处是 requery 内层循环:on_query_message_added(每条消息)、on_model_query(每次真正查询)、on_actions_generated(每次动作产出)
  • 对比图 8 的橙框:Mini 没有 hook,相当于把「观测能力」硬编码进主循环;SWE 把观测抽成回调,新增能力(轨迹记录 / 风险检查)只需加一个 hook
  • 这个「不改主循环、只加观察点」的思路不是 SWE-agent 独有——你在本项目的 Agent 适配层也会看到类似的扩展方式,可自行对照 experiment_modules/solving/agent_runtime/ 印证
课堂实训
  • 写一个 10 行的 hook 子类,override on_action_started 打印每次动作,运行后确认它被调用
  • 数一遍:为「记录每一步耗时」这个能力,Mini 需要改主循环几处,SWE 只需 new 一个 hook?
思考与讨论
  • hook 回调点太多会不会过度设计?哪些阶段你会合并?
  • 如果 mini-swe-agent 想加 hook,哪个点最值得先加?(提示:对照你上一讲学到的上下文压缩需求)
证据:agents.py:519 add_hook 注入 证据:hooks/abstract.py:11-53 13 回调定义
已知局限(诚实交代): ① 本页的包图、组件图与泳道活动图是用 flowchart 近似表达的——mermaid 没有原生 UML 包图 / 组件图 / 活动图(swimlane)语法,已在 MODEL_MANIFEST 说明; ② AbstractAgent / EnvHook / RunHook 的依据是「命名 + 空方法体」而非 ABC 机制,属 medium 置信度; ③ 本次为纯静态建模,未运行被测代码,运行时的动态覆盖(monkey patch 等)不体现在图中。