门户首页
API Reference · LLM 协议参考

三大 LLM 对话协议 Schema 详解

OpenAI Chat Completions · OpenAI Responses · Anthropic Messages 三套主流协议的完整 schema 定义、字段语义、往返示例、翻译对照表。讲"结构是什么"。 先对齐三个名词:LLM = Large Language Model(大语言模型);API = Application Programming Interface(应用程序接口——本页指调用大模型的 HTTP 接口);Schema = 模式定义(约定请求/响应里有哪些字段、什么类型、什么语义)。

生成时间:2026-09-02 · 版本 v0.2 · 生成 Agent:MiniMax Code · 载体:agentsoft-research-platform docs · 2026-09-06 修订:4 处「课堂实训」升级为任务型(lab-task 组件试点,命令均经实跑验证)+ 新增第六篇自测

概览:三协议核心维度

3
主流协议
18+
核心字段族
3
工具往返协议
3
流式形态
协议端点核心语义设计阶段
OpenAI Chat CompletionsPOST /v1/chat/completions扁平 message + 挂载 tool_calls2023 至今(事实标准)
OpenAI ResponsesPOST /v1/responses判别联合 item + 服务端状态2025.03 起(Agent 原语)
Anthropic MessagesPOST /v1/messages判别联合 content block + max_tokens 必填2023 至今(Claude 生态)

三协议工具调用往返 · 并排时序对比

开场总览 · 同一场景,三套协议
用户说"跑一下测试" → 模型调 run_tests 工具 → 回填结果 → 总结。四条泳道并排,一眼看清三家协议在"请求 / 工具调用 / 回填 / 总结"四步上的结构差异。
sequenceDiagram participant C as Client(你的程序) participant CC as Chat Completions participant RP as Responses participant MS as Messages Note over C,CC: ── Chat:messages 数组全量回放 ── C->>CC: messages(user: 跑一下测试) + tools CC-->>C: assistant.tool_calls(id=call_abc,arguments 是 JSON 字符串),finish_reason=tool_calls C->>CC: 全量回放 + 独立 role:tool 消息回填(tool_call_id=call_abc 严格配对,42 passed) CC-->>C: 总结:测试全部通过,finish_reason=stop Note over C,RP: ── Responses:input 可传字符串,output 是 item 数组 ── C->>RP: input 直接传字符串「跑一下测试」+ 扁平工具定义 RP-->>C: output item 数组含 function_call(call_id=fc_abc) C->>RP: 回填 function_call_output item(call_id 配对);可改用 previous_response_id 续接、不重传 history RP-->>C: 总结 message item,status=completed Note over C,MS: ── Messages:max_tokens 必填 + system 顶层 ── C->>MS: max_tokens 必填 + 顶层 system + messages(user: 跑一下测试) MS-->>C: content 含 tool_use block(id=toolu_xyz,input 已是对象非 JSON 字符串),stop_reason=tool_use C->>MS: assistant 原样回放 + user 消息内嵌 tool_result block(tool_use_id 配对) MS-->>C: 总结 text block,stop_reason=end_turn
图 1 · 三协议工具调用往返并排时序——同一"跑一下测试"场景下,Chat 靠 messages 全量回放 + role:tool 回填,Responses 用判别 item + previous_response_id 续接,Messages 以 tool_result 嵌进 user 消息且 input 已是对象

第一篇 · OpenAI Chat Completions Schema

第一篇 · Chat Completions(事实标准)
回答"全生态兼容的普通话长什么样、消息怎么排、响应怎么拆"。
卡片 1 · 端点 + 消息四元组

扁平 message + role 四元组

对话载体是 messages: MessageParam[],按 role 字段判别身份。 除 model 和 messages 外全部字段可选——这是"宽松 schema"的典型形态。

核心知识点
  • system / developer:全局指令,developer 是 2025 年新增的更细粒度角色
  • user:用户输入,content 支持 string 或多模态部件数组
  • assistant:模型历史回复;发起工具调用时 content 可为 null,tool_calls 挂在同层
  • tool:工具结果,必带 tool_call_id 与 assistant 的 tool_calls[].id 严格配对
思考与讨论
  • 为什么 role 配字段的约束(如 tool 必带 tool_call_id)不写进类型定义?服务端运行时校验兜底,是事故的结构根源
spec §1.3 消息类型 paper: SWE-bench (ICLR 2024 · SWE = Software Engineering 软件工程 + bench = benchmark 基准测试;ICLR = International Conference on Learning Representations,机器学习顶会)
卡片 2 · 响应 + 状态机

finish_reason 四态机

响应正文在 choices[0].message,循环分支的唯一依据是 finish_reason。 手写 Agent 的主循环只认这个字段做分支,四种取值必须全覆盖。

核心知识点
  • stop:自然结束 → 收尾、提交答案、退出循环
  • tool_calls:模型要求调工具 → 执行工具 → role:"tool" 回填 → 再次请求(核心循环)
  • length:撞 max_tokens 被截断 → 扩预算或续写,不要盲目重试(同样输入大概率同样截断)
  • content_filter:触发安全策略 → 降级 / 改写 / 上报
课堂实训
  • 用 if response.choices[0].message.content: 判循环结束? 反模式:模型完全可能返回空 content + 有效 tool_calls,这个判断会把工具循环掐死
动手做 · 在本仓找 finish_reason 的真实用法
cd agentsoft-research-platform   # 本仓根目录
grep -rn "finish_reason" experiment_modules/solving/adapters --include="*.py" | grep -v tests

预期:3 处命中(mimo/_mimo_json.py:63、mimo/agent.py:143、opencode/_opencode_json.py:71)。打开 mimo/_mimo_json.py 读 63 行前后,回答:本仓 adapter 把 finish_reason 存到哪里、为什么只记录不用它做循环分支?(提示:循环终止由 CLI 封装层处理;adapter 只把原因沉淀进轨迹 extras,供 S6 评分 / S7 过程分析回溯)

spec §1.5 响应对象 spec §3.2 finish_reason 状态机
卡片 3 · 流式 chunk 拼接

delta + index 聚合:四条规则

stream: true 时服务端返回 SSE(Server-Sent Events,服务器单向推送的流式传输)流。增量在 delta 而非 message,把每个 chunk 的 delta.content 顺序拼接才是完整文本。

核心知识点
  • tool_calls 的 arguments 同样分片到达 → 按 delta.tool_calls[].index 聚合,先拼 name 再拼 arguments
  • arguments 拼接完成后再一次性 json.loads,中途解析必然失败
  • [DONE] 是终止哨兵,收到即关闭连接
  • finish_reason 出现在最后一个 chunk的 delta 旁
spec §1.6 流式 chunk spec §5 SSE 解析
卡片 4 · Schema 全景

完整 Schema + 实战核心子集 + 工具调用往返

前 3 张卡片拆开了 Chat Completions 的三个剖面(消息、结束态、流式)。这一张把"完整 schema"贴回原样——按 OpenAI 官方 API 参考浓缩,包含请求体、响应体、工具定义三个对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环 4 段 JSON(JavaScript Object Notation,人类可读的文本数据格式)把全流程跑通。

完整请求体(ChatCompletionRequest)
type ChatCompletionRequest = {
  // —— 必填:身份与上下文 ——
  model: string,                          // "gpt-4o" / "gpt-5" / "deepseek-chat"...
  messages: MessageParam[],               // 扁平、按时间序的完整历史(无状态全量回放)

  // —— 采样与长度 ——
  temperature?: number,                   // 默认 1.0;推理模型忽略
  top_p?: number,                         // 默认 1.0(与 temperature 二选一)
  max_completion_tokens?: number,         // 新字段(推理模型时代)
  max_tokens?: number,                    // ⚠ deprecated,但兼容端点普遍只认它
  n?: number, seed?: number, stop?: string | string[],

  // —— 工具调用(Agent 核心)——
  tools?: ToolDefinition[],               // 工具定义(见卡片 1 消息四元组的 assistant 段;TypeScript 定义见 1.4)
  tool_choice?: "none" | "auto" | "required"
              | { type: "function", function: { name: string } },
  parallel_tool_calls?: boolean,

  // —— 结构化输出 ——
  response_format?: { type: "text" }
                  | { type: "json_object" }                       // JSON 模式
                  | { type: "json_schema", json_schema: {         // 强约束 schema
                        name: string, schema: object, strict?: boolean } },

  // —— 流式 ——
  stream?: boolean,
  stream_options?: { include_usage?: boolean },

  // —— 元信息 & 惩罚 & 微调 ——
  user?: string, store?: boolean, metadata?: object,
  service_tier?: "auto" | "default" | "flex",
  reasoning_effort?: "low" | "medium" | "high",   // 推理模型专属(方言级)
  presence_penalty?: number, frequency_penalty?: number,
  logit_bias?: Record<string, number>,
  logprobs?: boolean, top_logprobs?: number,
};
完整响应体(ChatCompletion)
type ChatCompletion = {
  id: string,                            // "chatcmpl-...",每次请求唯一
  object: "chat.completion",
  created: number,                       // unix 秒
  model: string,                         // 实际使用的快照版(与请求可能不同)
  choices: [{
    index: number,                       // n>1 时区分候选
    message: {
      role: "assistant",
      content: string | null,            // 工具调用时可为 null!
      tool_calls?: ToolCall[],
      refusal?: string | null,
      annotations?: Annotation[]
    },
    finish_reason: "stop" | "length" | "tool_calls"
                 | "content_filter" | "function_call",   // 末项 legacy
    logprobs?: object
  }],
  usage: {
    prompt_tokens: number, completion_tokens: number, total_tokens: number,
    prompt_tokens_details?: { cached_tokens?: number, audio_tokens?: number },
    completion_tokens_details?: { reasoning_tokens?: number }
  },
  system_fingerprint?: string
};

type ToolCall = {
  id: string,                            // "call_...",回填配对用
  type: "function",
  function: { name: string, arguments: string }   // ⚠ arguments 是 JSON 字符串
};
实战核心子集:上面 ~30 个字段,手写 Agent 80% 时间只摸这 10 个——model / messages / temperature / stream / tools / tool_choice / max_tokens(或 max_completion_tokens)/ response_format / stop。其余字段(logit_bias / parallel_tool_calls / service_tier / logprobs 等)属于"知道在哪、想用时翻"。
核心字段速查表(实战 10 + 推荐 2)
字段位置含义必用?
model请求顶层模型 ID(gpt-4o、deepseek-chat)★ 必填
messages请求顶层对话历史,全量回放(无状态协议)★ 必填
temperature请求顶层采样温度,0 = 贪心,1 = 默认;推理模型忽略★ 常用
stream请求顶层是否流式(SSE);前端打字机场景必用★ 常用
tools请求顶层工具定义(声明"能做什么",不是"要做什么")★ Agent 必用
tool_choice请求顶层工具选择:auto / none / required / 强制某函数★ Agent 必用
max_tokens / max_completion_tokens请求顶层最大生成 token 数;旧/新两套并存,跨厂商要按端点切换★ 常用
response_format请求顶层结构化输出:json_object / json_schema★ 解析友好场景
stop请求顶层停止序列(最多 4 个)★ 常用
user请求顶层终端用户标识(OpenAI 用于滥用检测)○ 推荐
choices[0].finish_reason响应循环分支唯一依据:stop / tool_calls / length / content_filter★ 必读
usage响应token 计数(成本/限速面板用)○ 推荐
完整往返示例:工具调用最小环(4 段 JSON)
// ── ① 请求:用户问"跑一下测试",声明一个 run_tests 工具 ──
{
  "model": "gpt-4o",
  "messages": [{"role": "user", "content": "跑一下测试"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "run_tests",
      "parameters": {
        "type": "object",
        "properties": {"suite": {"type": "string"}},
        "required": ["suite"]
      }
    }
  }]
}
// ── ② 响应:模型要求调工具(finish_reason = "tool_calls")──
{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": null,                          // 注意:工具调用时 content 为 null
      "tool_calls": [{
        "id": "call_abc",
        "type": "function",
        "function": {"name": "run_tests", "arguments": "{\"suite\":\"all\"}"}   // arguments 是 JSON 字符串
      }]
    },
    "finish_reason": "tool_calls"
  }],
  "usage": {"prompt_tokens": 88, "completion_tokens": 12, "total_tokens": 100}
}
// ── ③ 回填请求:append 两条消息(assistant 原样 + tool 结果)──
{
  "model": "gpt-4o",
  "messages": [
    {"role": "user", "content": "跑一下测试"},
    {"role": "assistant", "content": null, "tool_calls": [
      {"id": "call_abc", "type": "function",
       "function": {"name": "run_tests", "arguments": "{\"suite\":\"all\"}"}}
    ]},
    {"role": "tool", "tool_call_id": "call_abc", "content": "42 passed"}    // tool_call_id 严格配对 call_abc
  ],
  "tools": ["... 同上 ..."]
}
// ── ④ 最终响应:模型总结(finish_reason = "stop")──
{
  "choices": [{
    "message": {"role": "assistant", "content": "测试全部通过。"},
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 142, "completion_tokens": 8, "total_tokens": 150}
}
课堂实训
  • 数一数上面"完整响应体"里 id 字段出现的位置——它每次请求都不同。所以 id 不能当幂等键,要幂等得自己用 messages 的 hash
  • 把第 ③ 段里的 tool_call_id 改成 call_xyz 再发,会发生什么?(提示:服务端会 400,因为配不上 assistant 消息里的 tool_calls[0].id)
  • 第 ① 段没传 temperature,模型实际用了什么值?再查官方文档确认——这就是"宽松 schema"的代价:默认值散落各处
动手做 · 亲手造一个幂等键(纯标准库,10 秒)
python3 -c "
import json, hashlib
msgs = [{'role':'user','content':'跑一下测试'}]
key = hashlib.sha256(json.dumps(msgs, ensure_ascii=False, sort_keys=True).encode()).hexdigest()[:16]
print('幂等键:', key)"

预期输出:幂等键: 62de41c91b2f7e9a(已实跑验证)。重复跑结果不变;把 content 改一个字再跑,键完全改变——这就是「服务端 id 每次不同、幂等键只能自己用 messages 的 hash 造」的含义。

spec §1.2-1.5 请求/响应/工具 ref: OpenAI Chat Completions API 📎 完整 schema 附录

第二篇 · OpenAI Responses API Schema

第二篇 · Responses(Agent 原语)
回答"判别联合 item 怎么用、为什么 GPT-5 之后 Chat 进入维护模式"。
卡片 1 · Item 判别联合

工具调用是一等 item,不再挂在 message 上

与 Chat 的"扁平 message + 可选 tool_calls"不同,Responses 把所有参与者建模为判别联合 item——message / function_call / function_call_output / reasoning 都是顶层数组元素。

核心知识点
  • function_call item:call_id(命名从 Chat 的 id 改)+ name + arguments(仍是 JSON 字符串)
  • function_call_output item:call_id + output(无 role:"tool" 概念)
  • reasoning item:推理链可跨轮保留,encrypted_content 形态回传服务端续接推理
  • item_reference item:引用既有 item,免重传
思考与讨论
  • 为什么 Responses 抛弃 role:"tool"?工具调用升级为一等公民,消息仅用于文本对话,结构更对称
spec §2.3 Item 联合 paper: SWE-agent (arXiv 2405.15793)
卡片 2 · 服务端状态 + 推理持久化

previous_response_id:服务端替客户端维护历史

Responses 的关键演进:Chat 的"无状态全量回放"在长程 Agent 里代价线性膨胀,Responses 让你只发增量——服务端用 previous_response_id 替客户端维护对话状态。

核心知识点
  • previous_response_id:服务端续接上一轮,只发增量(也可继续手动传 items,保持客户端全控)
  • store: true(默认):服务端存储响应;可关掉做无状态
  • reasoning.effort:"minimal" | "low" | "medium" | "high",推理模型专属
  • 推理持久化实打实降本:Chat 协议下每个 tool loop 轮推理链都会丢失重算,Responses 可加密续接
spec §2.2 请求对象 paper: SWE-bench (ICLR 2024)
卡片 3 · Schema 全景

完整 Schema + 实战核心子集 + 工具调用往返

前 2 张卡片拆开了 Responses 的两个剖面(Item 判别联合、服务端状态)。这一张把"完整 schema"贴回原样——按 OpenAI 官方 API 参考浓缩,包含请求体、判别联合 item、响应体、流式事件四组对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环跑通——展示 Responses 的"无 role:tool、无 messages、判别 item"三大差异如何落地。

完整请求体(ResponseRequest)
type ResponseRequest = {
  model: string,                          // "gpt-5" / "o3" ...
  input: string | ResponseItem[],         // ★ 直接传字符串也合法(极简入口)

  instructions?: string,                  // ★ system 的替代(顶层,不进 input)
  max_output_tokens?: number,
  temperature?: number, top_p?: number,

  // —— 状态管理(与 Chat 的最大区别)——
  previous_response_id?: string,          // ★ 服务端续接上一轮,只发增量
  store?: boolean,                        // 服务端存储响应(默认 true)
  conversation?: string,                  // 会话对象(更高级的状态原语)

  // —— 工具(扁平化 + 内置托管工具)——
  tools?: (FunctionTool | WebSearchTool | FileSearchTool
         | CodeInterpreterTool | McpTool)[],
  tool_choice?: "auto" | "none" | "required",
  parallel_tool_calls?: boolean,

  // —— 推理控制(GPT-5 推理模型专属)——
  reasoning?: {
    effort?: "minimal" | "low" | "medium" | "high",
    summary?: "auto" | "concise" | "detailed"
  },

  // —— 输出格式 ——
  text?: { format: { type: "text" } | { type: "json_object" }
                | { type: "json_schema", name, schema, strict? } },
  stream?: boolean,
  metadata?: object
}

type FunctionTool = { type: "function", name: string,         // ★ 扁平,无 function 包装层
                      description?: string,
                      parameters?: JSONSchema,
                      strict?: boolean }
Input/Output Item 判别联合(与 Chat 的本质差异)
type ResponseItem =
  | { type: "message",                     // 文本消息
      role: "user" | "assistant" | "system" | "developer",
      content: string | ContentPart[] }

  | { type: "function_call",               // ★ 工具调用是一等 item
      call_id: string,                     // ★ 命名从 id → call_id
      name: string,
      arguments: string }                 // JSON 字符串

  | { type: "function_call_output",        // ★ 工具结果也是一等 item(无 role 概念)
      call_id: string,
      output: string }

  | { type: "reasoning",                   // ★ 推理链 item:跨轮保留思维链
      summary?: SummaryPart[],
      encrypted_content?: string }         // 加密形态,回传服务端续接推理

  | { type: "item_reference", id: string } // 引用既有 item,免重传
完整响应体(Response)
type Response = {
  id: string,                             // "resp_..."
  object: "response",
  model: string,
  status: "completed" | "failed" | "in_progress" | "cancelled",
  output: ResponseItem[],                 // ★ 结构与 input 同构(含 reasoning item)
  usage: {
    input_tokens: number,                 // ★ 命名从 prompt_tokens → input_tokens
    input_tokens_details?: { cached_tokens: number },
    output_tokens: number,
    output_tokens_details?: { reasoning_tokens: number },
    total_tokens: number
  },
  previous_response_id?: string | null,
  error?: { code: string, message: string } | null
}
实战核心子集:上面 ~25 个字段,手写 Agent 80% 时间只摸这 10 个——model / input / instructions / previous_response_id / tools / tool_choice / reasoning.effort / max_output_tokens / stream / text.format。其余字段(conversation / parallel_tool_calls / 内置托管工具的 type 变体)属于"知道在哪、想用时翻"。
核心字段速查表(实战 10 + 推荐 2)
字段位置含义必用?
model请求顶层模型 ID(gpt-5 / o3)★ 必填
input请求顶层对话载体:string 或 Item[](判别联合)★ 必填
instructions请求顶层系统指令(Chat 的 role:system 在此顶层)★ 常用
previous_response_id请求顶层服务端续接历史,只发增量★ Agent 必用
tools请求顶层工具定义;FunctionTool 是扁平结构(无 function 包装层)★ Agent 必用
tool_choice请求顶层auto / none / required(无"强制某函数"形式)★ 常用
reasoning.effort请求顶层 reasoning推理强度:minimal/low/medium/high,GPT-5 推理模型专属★ 推理场景
max_output_tokens请求顶层最大生成 token;output_ 前缀(区别于 Chat 的 completion_)★ 常用
stream请求顶层是否流式(类型化事件流,非同构 chunk)★ 常用
text.format请求顶层 text结构化输出:json_object / json_schema★ 解析友好场景
store请求顶层是否服务端存储响应(默认 true)○ 隐私场景
output[] + status响应循环分支:completed 收尾 / failed 兜底 / in_progress 中转★ 必读
usage响应token 计数:input_tokens / output_tokens(含 reasoning_tokens)○ 推荐
完整往返示例:工具调用最小环(3 段 JSON + 1 个续接)
// ── ① 请求:直接传字符串 input(极简入口)+ 扁平工具定义 ──
{
  "model": "gpt-5",
  "input": "跑一下测试",
  "tools": [{"type": "function", "name": "run_tests",            // 扁平!无 function 包装层
    "parameters": {"type": "object",
      "properties": {"suite": {"type": "string"}}}}]
}
// ── ② 响应:output 是 item 数组(含 reasoning + function_call 两种 item)──
{
  "id": "resp_001",
  "status": "completed",
  "output": [
    {"type": "reasoning", "summary": []},                       // 推理链 item
    {"type": "function_call", "call_id": "fc_abc",              // call_id 命名
     "name": "run_tests", "arguments": "{\"suite\":\"all\"}"}
  ],
  "usage": {"input_tokens": 80, "output_tokens": 40, "total_tokens": 120}
}
// ── ③ 回填:两条 item(function_call 原样 + function_call_output)──
// 注意:与 Chat 相比没有 role:"tool"、没有 messages
{
  "model": "gpt-5",
  "input": [
    {"type": "function_call", "call_id": "fc_abc",
     "name": "run_tests", "arguments": "{\"suite\":\"all\"}"},
    {"type": "function_call_output", "call_id": "fc_abc",
     "output": "42 passed"}
  ]
}
// ── ④ 进阶:previous_response_id 续接(不用重传 history)──
{
  "model": "gpt-5",
  "previous_response_id": "resp_001",         // 服务端已有 resp_001 的状态
  "input": "那覆盖率呢?"                       // 只发新增输入
}
课堂实训
  • 第 ① 段的 input 是字符串"跑一下测试"——Responses 允许这种极简入口。换成 Chat 协议就必须包成 messages: [{role:"user", content:...}],对比一下两边的"最短可工作请求"长度
  • 把第 ③ 段 function_call item 的 call_id 改成 fc_xyz 再发,会发生什么?服务端会 400——这是 Chat 的 tool_call_id 配对的 Responses 等价坑
  • 第 ④ 段没传 input 数组,只传了字符串"那覆盖率呢?"——服务端会拿这个字符串 + 上一轮的 resp_001 上下文续接。这是 Responses 相对 Chat 的最大成本优势(长程 Agent 不用每轮都重传全量 history)
动手做 · 量一量「最短可工作请求」的长度差
python3 -c "
import json
chat = json.dumps({'model':'gpt-5','messages':[{'role':'user','content':'跑一下测试'}]}, ensure_ascii=False)
resp = json.dumps({'model':'gpt-5','input':'跑一下测试'}, ensure_ascii=False)
print(f'Chat {len(chat)} 字符 vs Responses {len(resp)} 字符')"

预期输出:Chat 70 字符 vs Responses 36 字符(已实跑验证)——字符串 input 入口几乎省一半。进阶:给两边各补一轮工具回填消息再测,观察差距如何随轮次累计(长程 Agent 的成本账)。

spec §2.2-2.5 请求/Item/响应/流式 ref: OpenAI Responses API 📎 完整 schema 附录

第三篇 · Anthropic Messages Schema

第三篇 · Messages(Claude 生态)
回答"Claude 的语言长什么样、为什么 max_tokens 必填、块边界怎么用流式事件表达"。
卡片 1 · 端点 + content block

顶层 system + content 必为 block 数组 + max_tokens 必填

Messages 与 Chat 的三个结构差异:system 顶层不进 messages、content 必为数组(text + tool_use 可同轮并存)、max_tokens 必填。三个差异一起决定了协议严格度上限。

核心知识点
  • system 顶层参数:可传 string 或 {type:"text", text, cache_control?} 数组
  • max_tokens 必填:强制显式声明资源上限(OpenAI 侧全可选)
  • cache_control: {type:"ephemeral"}:显式缓存断点,大 system / 工具列表的省钱利器
  • 输入侧 block 联合:text / image / tool_result / document,tool_result 嵌在 user 消息里(结构性不同)
spec §3.2-3.4 请求与消息
卡片 2 · 流式 + thinking 块

content_block_start/stop:流式块边界靠事件

Messages 的流式是显式状态机:每个 SSE 事件带 type 字段,content_block_start / stop 明确告诉解析器"这个 tool_use / thinking 到哪结束"——无需靠 index 拼凑。

核心知识点
  • 事件族:message_start / content_block_start / content_block_delta / content_block_stop / message_delta / message_stop / ping
  • delta 类型:text_delta / input_json_delta(工具入参分片) / thinking_delta(思维链) / signature_delta
  • thinking block:思维链内容 + 防篡改 signature,回填多轮时必须原样保留
  • ping 心跳:协议内建连接活性判断,不用自己造
spec §3.6 流式事件 spec §3.3 输出侧 block
卡片 3 · Schema 全景

完整 Schema + 实战核心子集 + 工具调用往返

前 2 张卡片拆开了 Messages 的两个剖面(content block 判别联合、流式事件状态机)。这一张把"完整 schema"贴回原样——按 Anthropic 官方 API 参考浓缩,包含请求体、消息与内容块、响应体、流式事件四组对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环跑通——展示 Messages 与 Chat/Responses 的三个核心差异(system 顶层 / max_tokens 必填 / tool_result 嵌在 user 里)。

端点与认证(注意 anthropic-version 头必填)
POST /v1/messages
x-api-key: <api-key>                    // ★ 不用 Bearer,用自定义 x-api-key 头
anthropic-version: 2023-06-01            // ★ 显式版本头(必填)
content-type: application/json
完整请求体(MessagesRequest)
type MessagesRequest = {
  model: string,                          // "claude-sonnet-4-5" / "claude-opus-4-1" ...
  messages: MessageParam[],               // 必填
  max_tokens: number,                     // ★ 必填!强制显式声明资源上限

  // —— system 是顶层参数(不进 messages)——
  system?: string
        | { type: "text", text: string, cache_control?: CacheControl }[],

  // —— 采样 ——
  temperature?: number,                   // 0~1
  top_p?: number, top_k?: number,
  stop_sequences?: string[],

  // —— 工具(注意:tool_choice 是结构化对象,不是字符串枚举)——
  tools?: ToolDefinition[],
  tool_choice?: { type: "auto" }                        // 模型自决
              | { type: "any" }                         // 必须调某个工具
              | { type: "tool", name: string },         // 强制调指定工具

  // —— 其他 ——
  stream?: boolean,
  metadata?: { user_id?: string }
}

type CacheControl = { type: "ephemeral" }  // ★ 显式缓存断点
消息与内容块(判别联合)
type MessageParam = {
  role: "user" | "assistant",             // ★ 只有两种 role(无 system / tool / developer)
  content: string | ContentBlock[]        // 空字符串不允许;块可为空数组占位
}

// —— 输入侧块(请求中可出现)——
type ContentBlock =
  | { type: "text", text: string, cache_control?: CacheControl }

  | { type: "image",
      source: { type: "base64", media_type: "image/jpeg"|"image/png"|"image/gif"|"image/webp",
                data: string }
             | { type: "url", url: string } }

  | { type: "tool_result",                // ★ 工具结果:嵌在 user 消息里
      tool_use_id: string,                // 与 ToolUseBlock.id 配对
      content?: string | ContentBlock[],
      is_error?: boolean,                 // ★ schema 原生错误标记(区别于 Chat)
      cache_control?: CacheControl }

  | { type: "document",                   // PDF 支持
      source: {...} }

// —— 输出侧块(响应中出现;thinking 块回填时必须原样保留)——
type OutputBlock =
  | { type: "text", text: string }
  | { type: "tool_use",
      id: string,                         // "toolu_..."
      name: string,
      input: object }                     // ★ 已是对象,不是 JSON 字符串!
  | { type: "thinking",
      thinking: string,                   // 思维链内容
      signature: string }                 // 防篡改签名(回填必须原样保留)
  | { type: "redacted_thinking", data: string }   // 加密思维链
完整响应体(Message)
type Message = {
  id: string,                             // "msg_..."
  type: "message",
  role: "assistant",
  model: string,
  content: OutputBlock[],                 // ★ 必为数组(text + tool_use 可同轮并存)
  stop_reason: "end_turn" | "max_tokens" | "stop_sequence"
             | "tool_use" | "refusal" | "pause_turn",
  stop_sequence: string | null,
  usage: {
    input_tokens: number,
    output_tokens: number,
    cache_creation_input_tokens?: number, // ★ 缓存计费一级公民
    cache_read_input_tokens?: number
  }
}
实战核心子集:上面 ~20 个字段,手写 Agent 80% 时间只摸这 10 个——model / messages / max_tokens / system / tools / tool_choice / temperature / stream / cache_control / stop_sequences。其余字段(top_k / top_p / metadata)属于"知道在哪、想用时翻"。注意 max_tokens 是必填——OpenAI 侧全可选,跨厂商代码要兜底默认值。
核心字段速查表(实战 10 + 推荐 2)
字段位置含义必用?
model请求顶层模型 ID(claude-sonnet-4-5)★ 必填
messages请求顶层对话历史;只有 user/assistant 两种 role★ 必填
max_tokens请求顶层最大生成 token;必填(Anthropic 严格度分水岭)★ 必填
system请求顶层系统指令(不进 messages,独立顶层参数)★ 常用
tools请求顶层工具定义;name/input_schema 扁平(无包装层)★ Agent 必用
tool_choice请求顶层结构化对象:{type:"auto"} / "any" / "tool",name★ 常用
temperature请求顶层0~1(OpenAI 0~2)——跨厂商代码要钳制范围★ 常用
stream请求顶层是否流式;触发显式状态机事件流(content_block_start/stop + ping)★ 常用
cache_control块级 / system 数组元素 / 工具定义显式缓存断点;大 system / 工具列表的省钱利器★ 长 system 必用
stop_sequences请求顶层停止序列(最多 4 个,命名与 OpenAI stop 不同)★ 常用
x-api-key + anthropic-versionHTTP 头认证头 + 版本头(必填,区别于 OpenAI 的 Authorization: Bearer)★ 必填
stop_reason响应循环分支:end_turn / tool_use / max_tokens / refusal★ 必读
usage.cache_*_input_tokens响应缓存命中/创建的 token 数(成本面板必看)○ 推荐
完整往返示例:工具调用最小环(3 段 JSON)
// ── ① 请求:max_tokens 必填 + system 顶层 + 扁平工具定义 ──
{
  "model": "claude-sonnet-4-5",
  "max_tokens": 1024,                                       // ★ 必填
  "system": "你是测试助手",                                  // 顶层,不进 messages
  "messages": [{"role": "user", "content": "跑一下测试"}],
  "tools": [{"name": "run_tests", "description": "运行测试",   // ★ 扁平!无 type/function 包装
    "input_schema": {"type": "object",
      "properties": {"suite": {"type": "string"}}}}]
}
// ── ② 响应:stop_reason = "tool_use";input 已是对象(不是 JSON 字符串)──
{
  "id": "msg_001",
  "type": "message",
  "role": "assistant",
  "content": [
    {"type": "text", "text": "我先运行测试"},
    {"type": "tool_use", "id": "toolu_xyz",
     "name": "run_tests", "input": {"suite": "all"}}            // ★ input 是对象!
  ],
  "stop_reason": "tool_use",
  "usage": {"input_tokens": 95, "output_tokens": 15}
}
// ── ③ 回填:assistant 原样回放 + user 内嵌 tool_result ──
// 注意:tool_result 嵌在 user 消息里(不是独立 role),用 tool_use_id 配对
{
  "messages": [
    {"role": "user", "content": "跑一下测试"},
    {"role": "assistant", "content": [
      {"type": "text", "text": "我先运行测试"},
      {"type": "tool_use", "id": "toolu_xyz",
       "name": "run_tests", "input": {"suite": "all"}}
    ]},
    {"role": "user", "content": [
      {"type": "tool_result", "tool_use_id": "toolu_xyz",     // tool_use_id 严格配对
       "content": "42 passed"}
    ]}
  ]
}
课堂实训
  • 把第 ① 段里的 max_tokens 删掉再发,会发生什么?(提示:Anthropic 侧 400 "max_tokens: Field required",对比 Chat 的可选)
  • 第 ② 段 tool_use.input 已经是 {"suite": "all"} 对象——Chat/Responses 是 JSON 字符串 "{\"suite\":\"all\"}"。改协议时记得 json.loads 一下
  • 第 ③ 段把 tool_use_id 改成 toolu_abc 再发会怎样?400——配对错误。Chat 报 tool_call_id、Responses 报 call_id、Messages 报 tool_use_id,三家命名不同但语义一致
  • 如果模型在第 ② 段同时输出了 thinking 块(思维链),第 ③ 段回填时必须原样保留(含 signature)——否则服务端会判为篡改、拒绝续接
动手做 · 亲手触发一次 400(本地模拟必填校验)
python3 -c "
req = {'model':'claude-opus-4-8','messages':[{'role':'user','content':'hi'}]}
required = ['model','max_tokens','messages']
missing = [k for k in required if k not in req]
print('400:', f'{missing[0]}: Field required' if missing else 'OK')"

预期输出:400: max_tokens: Field required(已实跑验证)——与 Anthropic 真实报错文案一致。给 req 补上 'max_tokens': 1024 再跑,输出 OK;再把 required 换成 OpenAI 的 ['model','messages'],体会「严格度分水岭」。

spec §3.2-3.6 请求/消息/响应/流式 ref: Anthropic Messages API 📎 完整 schema 附录

第四篇 · 三协议类型映射总表("翻译词典")

第四篇 · 协议翻译表
写跨厂商客户端 / 协议转换层时按此表对齐——配对键错一个就 400。
翻译第一原则:工具往返的配对键是三家共同的可靠性锚点——id(Chat) / call_id(Responses) / tool_use_id(Messages tool_result)。配错就 400,转换层必须精确映射。
概念Chat CompletionsResponses APIMessages API
对话载体messages[]input: string | Item[]messages[]
系统指令messages[0].role="system"顶层 instructions顶层 system
消息角色system / developer / user / assistant / tool仅 message item 内 user / assistant / system / developer仅 user / assistant
内容建模扁平 message + 可选字段判别联合 item判别联合 block
工具定义{type:"function", function:{name, parameters}} 两层{type:"function", name, parameters} 扁平{name, input_schema} 无包装
模型发起调用message.tool_calls[].iditem{type:"function_call"}.call_idblock{type:"tool_use"}.id
调用参数arguments(JSON 字符串)arguments(JSON 字符串)input(对象)
结果回填独立 {role:"tool", tool_call_id}{type:"function_call_output", call_id}user 内 tool_result block
错误标记无(content 文本自约定)无(output 文本自约定)is_error: true
长度上限max_completion_tokens(可选)max_output_tokens(可选)max_tokens(必填)
结束信号finish_reason: stop/length/tool_callsstatus: completed/failedstop_reason: end_turn/max_tokens/tool_use
流式形态同构 chat.completion.chunk 流类型化 response.* 事件类型化 message/block/delta 事件 + ping
流式块边界无(按 index 聚合自拼)output_item.added/donecontent_block_start/stop
缓存控制自动(usage 明细报 cached)自动cache_control 显式断点
推理链暴露不暴露(仅 reasoning_tokens 计数)reasoning item(可加密续接)thinking block(带签名回放)
状态管理无状态previous_response_id无状态
token 计数字段prompt / completion_tokensinput / output_tokensinput / output_tokens + cache 两项
版本管理无版本头无版本头anthropic-version 头

第五篇 · 设计规律与权威来源

第五篇 · 设计规律 + 官方文档
回答"为什么值得对照学三家协议 + 落地实现前去哪里查权威 schema"。
卡片 1 · 设计规律

五条规律:从三家协议收敛方向看未来

三家后发 schema 都收敛到判别联合结构,工具往返配对键是共同的可靠性锚点——理解这五条对自研 Agent 协议 / 轨迹存储格式有直接指导。

核心知识点
  • 判别联合是终局:Chat 扁平 message → Responses item、Messages block,两个阵营后发 schema 都收敛到 {type: ...} 自判别结构。自研直接选判别联合
  • 工具往返配对键是可靠性锚点:三家共同的 id / call_id / tool_use_id 配对错就 400
  • "字符串化 JSON" 两种立场:OpenAI 系把工具参数定为 JSON 字符串(流式可追加分片);Anthropic 定为结构化对象(类型系统友好)
  • 严格度分水岭:max_tokens 必填(Anthropic)vs 全可选(OpenAI)——做兼容层时给 Anthropic 侧补默认上限
  • 流式协议两代形态:同构 chunk 流(Chat,靠字段缺失表义)vs 类型化事件流(Responses / Messages,显式状态机 + 块边界 + 心跳)。新协议设计一律选后者
思考与讨论
  • 如果让你自研 Agent 轨迹存储格式,你会用判别联合 item 模型(仿 Responses)还是扁平 message(仿 Chat)?为什么?
spec §6 设计规律
卡片 2 · 权威来源

官方文档索引:落地实现前必查

三家 schema 均在快速演进(尤其 Responses 与 thinking 相关结构),本文档为教学浓缩版,落地前以官方文档为准。

核心知识点
  • OpenAI:platform.openai.com/docs/api-reference/chat / /responses / /guides/migrate-to-responses
  • Anthropic:docs.anthropic.com/en/api/messages + platform.claude.com/docs 双域名同源;code.claude.com/docs 是 Claude Code 工具文档勿混
  • MCP(工具生态层):modelcontextprotocol.io,Agent 怎么以标准化方式连接成百上千个外部工具
  • OpenAI 部分参考页需登录开发者账号查看完整内容
spec §5 官方文档索引 paper: SWE-bench (ICLR 2024)

第六篇 · 自测(动手做完,来对答案)

自测 · 3 题(点击展开答案)
覆盖本页三个最易踩的坑:循环终止判据、配对键命名、续接成本。答不上来,就回对应篇的「课堂实训」把动手做跑一遍。
Q1 · 模型返回「空 content + 有效 tool_calls」时,Agent 主循环应该退出吗?
不应该。此时 finish_reason = "tool_calls"——模型在要求执行工具:执行 → 以 role:"tool" 回填 → 再次请求。用 if message.content: 判结束会把工具循环掐死(第一篇卡片 2 的反模式);分支依据只能是 finish_reason。
Q2 · 三家协议的工具往返配对键分别叫什么?配错会怎样?
Chat = tool_call_id,Responses = call_id,Messages = tool_use_id(嵌在 user 消息的 tool_result 块里)。语义一致、命名不同;配错服务端直接 400——写协议转换层时这是第一对齐项(第四篇翻译表)。
Q3 · Responses 的 previous_response_id 相对 Chat 每轮重传全量 history,省的是什么?
省每轮重传的 prompt tokens——服务端已保存上一轮响应状态,客户端只发增量输入。轮次越多、history 越长,成本差越大,这是长程 Agent 场景 Responses 相对 Chat 的最大成本优势(第二篇「动手做」可亲手量出入口长度差)。