三大 LLM 对话协议 Schema 详解
OpenAI Chat Completions · OpenAI Responses · Anthropic Messages 三套主流协议的完整 schema 定义、字段语义、往返示例、翻译对照表。讲"结构是什么"。 先对齐三个名词:LLM = Large Language Model(大语言模型);API = Application Programming Interface(应用程序接口——本页指调用大模型的 HTTP 接口);Schema = 模式定义(约定请求/响应里有哪些字段、什么类型、什么语义)。
概览:三协议核心维度
| 协议 | 端点 | 核心语义 | 设计阶段 |
|---|---|---|---|
| OpenAI Chat Completions | POST /v1/chat/completions | 扁平 message + 挂载 tool_calls | 2023 至今(事实标准) |
| OpenAI Responses | POST /v1/responses | 判别联合 item + 服务端状态 | 2025.03 起(Agent 原语) |
| Anthropic Messages | POST /v1/messages | 判别联合 content block + max_tokens 必填 | 2023 至今(Claude 生态) |
三协议工具调用往返 · 并排时序对比
第一篇 · OpenAI Chat Completions Schema
扁平 message + role 四元组
对话载体是 messages: MessageParam[],按 role 字段判别身份。 除 model 和 messages 外全部字段可选——这是"宽松 schema"的典型形态。
- system / developer:全局指令,developer 是 2025 年新增的更细粒度角色
- user:用户输入,content 支持 string 或多模态部件数组
- assistant:模型历史回复;发起工具调用时 content 可为 null,tool_calls 挂在同层
- tool:工具结果,必带 tool_call_id 与 assistant 的 tool_calls[].id 严格配对
- 为什么 role 配字段的约束(如 tool 必带 tool_call_id)不写进类型定义?服务端运行时校验兜底,是事故的结构根源
finish_reason 四态机
响应正文在 choices[0].message,循环分支的唯一依据是 finish_reason。 手写 Agent 的主循环只认这个字段做分支,四种取值必须全覆盖。
- stop:自然结束 → 收尾、提交答案、退出循环
- tool_calls:模型要求调工具 → 执行工具 → role:"tool" 回填 → 再次请求(核心循环)
- length:撞 max_tokens 被截断 → 扩预算或续写,不要盲目重试(同样输入大概率同样截断)
- content_filter:触发安全策略 → 降级 / 改写 / 上报
- 用 if response.choices[0].message.content: 判循环结束? 反模式:模型完全可能返回空 content + 有效 tool_calls,这个判断会把工具循环掐死
cd agentsoft-research-platform # 本仓根目录 grep -rn "finish_reason" experiment_modules/solving/adapters --include="*.py" | grep -v tests
预期:3 处命中(mimo/_mimo_json.py:63、mimo/agent.py:143、opencode/_opencode_json.py:71)。打开 mimo/_mimo_json.py 读 63 行前后,回答:本仓 adapter 把 finish_reason 存到哪里、为什么只记录不用它做循环分支?(提示:循环终止由 CLI 封装层处理;adapter 只把原因沉淀进轨迹 extras,供 S6 评分 / S7 过程分析回溯)
delta + index 聚合:四条规则
stream: true 时服务端返回 SSE(Server-Sent Events,服务器单向推送的流式传输)流。增量在 delta 而非 message,把每个 chunk 的 delta.content 顺序拼接才是完整文本。
- tool_calls 的 arguments 同样分片到达 → 按 delta.tool_calls[].index 聚合,先拼 name 再拼 arguments
- arguments 拼接完成后再一次性 json.loads,中途解析必然失败
- [DONE] 是终止哨兵,收到即关闭连接
- finish_reason 出现在最后一个 chunk的 delta 旁
完整 Schema + 实战核心子集 + 工具调用往返
前 3 张卡片拆开了 Chat Completions 的三个剖面(消息、结束态、流式)。这一张把"完整 schema"贴回原样——按 OpenAI 官方 API 参考浓缩,包含请求体、响应体、工具定义三个对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环 4 段 JSON(JavaScript Object Notation,人类可读的文本数据格式)把全流程跑通。
type ChatCompletionRequest = {
// —— 必填:身份与上下文 ——
model: string, // "gpt-4o" / "gpt-5" / "deepseek-chat"...
messages: MessageParam[], // 扁平、按时间序的完整历史(无状态全量回放)
// —— 采样与长度 ——
temperature?: number, // 默认 1.0;推理模型忽略
top_p?: number, // 默认 1.0(与 temperature 二选一)
max_completion_tokens?: number, // 新字段(推理模型时代)
max_tokens?: number, // ⚠ deprecated,但兼容端点普遍只认它
n?: number, seed?: number, stop?: string | string[],
// —— 工具调用(Agent 核心)——
tools?: ToolDefinition[], // 工具定义(见卡片 1 消息四元组的 assistant 段;TypeScript 定义见 1.4)
tool_choice?: "none" | "auto" | "required"
| { type: "function", function: { name: string } },
parallel_tool_calls?: boolean,
// —— 结构化输出 ——
response_format?: { type: "text" }
| { type: "json_object" } // JSON 模式
| { type: "json_schema", json_schema: { // 强约束 schema
name: string, schema: object, strict?: boolean } },
// —— 流式 ——
stream?: boolean,
stream_options?: { include_usage?: boolean },
// —— 元信息 & 惩罚 & 微调 ——
user?: string, store?: boolean, metadata?: object,
service_tier?: "auto" | "default" | "flex",
reasoning_effort?: "low" | "medium" | "high", // 推理模型专属(方言级)
presence_penalty?: number, frequency_penalty?: number,
logit_bias?: Record<string, number>,
logprobs?: boolean, top_logprobs?: number,
};
type ChatCompletion = {
id: string, // "chatcmpl-...",每次请求唯一
object: "chat.completion",
created: number, // unix 秒
model: string, // 实际使用的快照版(与请求可能不同)
choices: [{
index: number, // n>1 时区分候选
message: {
role: "assistant",
content: string | null, // 工具调用时可为 null!
tool_calls?: ToolCall[],
refusal?: string | null,
annotations?: Annotation[]
},
finish_reason: "stop" | "length" | "tool_calls"
| "content_filter" | "function_call", // 末项 legacy
logprobs?: object
}],
usage: {
prompt_tokens: number, completion_tokens: number, total_tokens: number,
prompt_tokens_details?: { cached_tokens?: number, audio_tokens?: number },
completion_tokens_details?: { reasoning_tokens?: number }
},
system_fingerprint?: string
};
type ToolCall = {
id: string, // "call_...",回填配对用
type: "function",
function: { name: string, arguments: string } // ⚠ arguments 是 JSON 字符串
};
| 字段 | 位置 | 含义 | 必用? |
|---|---|---|---|
| model | 请求顶层 | 模型 ID(gpt-4o、deepseek-chat) | ★ 必填 |
| messages | 请求顶层 | 对话历史,全量回放(无状态协议) | ★ 必填 |
| temperature | 请求顶层 | 采样温度,0 = 贪心,1 = 默认;推理模型忽略 | ★ 常用 |
| stream | 请求顶层 | 是否流式(SSE);前端打字机场景必用 | ★ 常用 |
| tools | 请求顶层 | 工具定义(声明"能做什么",不是"要做什么") | ★ Agent 必用 |
| tool_choice | 请求顶层 | 工具选择:auto / none / required / 强制某函数 | ★ Agent 必用 |
| max_tokens / max_completion_tokens | 请求顶层 | 最大生成 token 数;旧/新两套并存,跨厂商要按端点切换 | ★ 常用 |
| response_format | 请求顶层 | 结构化输出:json_object / json_schema | ★ 解析友好场景 |
| stop | 请求顶层 | 停止序列(最多 4 个) | ★ 常用 |
| user | 请求顶层 | 终端用户标识(OpenAI 用于滥用检测) | ○ 推荐 |
| choices[0].finish_reason | 响应 | 循环分支唯一依据:stop / tool_calls / length / content_filter | ★ 必读 |
| usage | 响应 | token 计数(成本/限速面板用) | ○ 推荐 |
// ── ① 请求:用户问"跑一下测试",声明一个 run_tests 工具 ──
{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "跑一下测试"}],
"tools": [{
"type": "function",
"function": {
"name": "run_tests",
"parameters": {
"type": "object",
"properties": {"suite": {"type": "string"}},
"required": ["suite"]
}
}
}]
}
// ── ② 响应:模型要求调工具(finish_reason = "tool_calls")──
{
"choices": [{
"message": {
"role": "assistant",
"content": null, // 注意:工具调用时 content 为 null
"tool_calls": [{
"id": "call_abc",
"type": "function",
"function": {"name": "run_tests", "arguments": "{\"suite\":\"all\"}"} // arguments 是 JSON 字符串
}]
},
"finish_reason": "tool_calls"
}],
"usage": {"prompt_tokens": 88, "completion_tokens": 12, "total_tokens": 100}
}
// ── ③ 回填请求:append 两条消息(assistant 原样 + tool 结果)──
{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "跑一下测试"},
{"role": "assistant", "content": null, "tool_calls": [
{"id": "call_abc", "type": "function",
"function": {"name": "run_tests", "arguments": "{\"suite\":\"all\"}"}}
]},
{"role": "tool", "tool_call_id": "call_abc", "content": "42 passed"} // tool_call_id 严格配对 call_abc
],
"tools": ["... 同上 ..."]
}
// ── ④ 最终响应:模型总结(finish_reason = "stop")──
{
"choices": [{
"message": {"role": "assistant", "content": "测试全部通过。"},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 142, "completion_tokens": 8, "total_tokens": 150}
}
- 数一数上面"完整响应体"里 id 字段出现的位置——它每次请求都不同。所以 id 不能当幂等键,要幂等得自己用 messages 的 hash
- 把第 ③ 段里的 tool_call_id 改成 call_xyz 再发,会发生什么?(提示:服务端会 400,因为配不上 assistant 消息里的 tool_calls[0].id)
- 第 ① 段没传 temperature,模型实际用了什么值?再查官方文档确认——这就是"宽松 schema"的代价:默认值散落各处
python3 -c "
import json, hashlib
msgs = [{'role':'user','content':'跑一下测试'}]
key = hashlib.sha256(json.dumps(msgs, ensure_ascii=False, sort_keys=True).encode()).hexdigest()[:16]
print('幂等键:', key)"
预期输出:幂等键: 62de41c91b2f7e9a(已实跑验证)。重复跑结果不变;把 content 改一个字再跑,键完全改变——这就是「服务端 id 每次不同、幂等键只能自己用 messages 的 hash 造」的含义。
第二篇 · OpenAI Responses API Schema
工具调用是一等 item,不再挂在 message 上
与 Chat 的"扁平 message + 可选 tool_calls"不同,Responses 把所有参与者建模为判别联合 item——message / function_call / function_call_output / reasoning 都是顶层数组元素。
- function_call item:call_id(命名从 Chat 的 id 改)+ name + arguments(仍是 JSON 字符串)
- function_call_output item:call_id + output(无 role:"tool" 概念)
- reasoning item:推理链可跨轮保留,encrypted_content 形态回传服务端续接推理
- item_reference item:引用既有 item,免重传
- 为什么 Responses 抛弃 role:"tool"?工具调用升级为一等公民,消息仅用于文本对话,结构更对称
previous_response_id:服务端替客户端维护历史
Responses 的关键演进:Chat 的"无状态全量回放"在长程 Agent 里代价线性膨胀,Responses 让你只发增量——服务端用 previous_response_id 替客户端维护对话状态。
- previous_response_id:服务端续接上一轮,只发增量(也可继续手动传 items,保持客户端全控)
- store: true(默认):服务端存储响应;可关掉做无状态
- reasoning.effort:"minimal" | "low" | "medium" | "high",推理模型专属
- 推理持久化实打实降本:Chat 协议下每个 tool loop 轮推理链都会丢失重算,Responses 可加密续接
完整 Schema + 实战核心子集 + 工具调用往返
前 2 张卡片拆开了 Responses 的两个剖面(Item 判别联合、服务端状态)。这一张把"完整 schema"贴回原样——按 OpenAI 官方 API 参考浓缩,包含请求体、判别联合 item、响应体、流式事件四组对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环跑通——展示 Responses 的"无 role:tool、无 messages、判别 item"三大差异如何落地。
type ResponseRequest = {
model: string, // "gpt-5" / "o3" ...
input: string | ResponseItem[], // ★ 直接传字符串也合法(极简入口)
instructions?: string, // ★ system 的替代(顶层,不进 input)
max_output_tokens?: number,
temperature?: number, top_p?: number,
// —— 状态管理(与 Chat 的最大区别)——
previous_response_id?: string, // ★ 服务端续接上一轮,只发增量
store?: boolean, // 服务端存储响应(默认 true)
conversation?: string, // 会话对象(更高级的状态原语)
// —— 工具(扁平化 + 内置托管工具)——
tools?: (FunctionTool | WebSearchTool | FileSearchTool
| CodeInterpreterTool | McpTool)[],
tool_choice?: "auto" | "none" | "required",
parallel_tool_calls?: boolean,
// —— 推理控制(GPT-5 推理模型专属)——
reasoning?: {
effort?: "minimal" | "low" | "medium" | "high",
summary?: "auto" | "concise" | "detailed"
},
// —— 输出格式 ——
text?: { format: { type: "text" } | { type: "json_object" }
| { type: "json_schema", name, schema, strict? } },
stream?: boolean,
metadata?: object
}
type FunctionTool = { type: "function", name: string, // ★ 扁平,无 function 包装层
description?: string,
parameters?: JSONSchema,
strict?: boolean }
type ResponseItem =
| { type: "message", // 文本消息
role: "user" | "assistant" | "system" | "developer",
content: string | ContentPart[] }
| { type: "function_call", // ★ 工具调用是一等 item
call_id: string, // ★ 命名从 id → call_id
name: string,
arguments: string } // JSON 字符串
| { type: "function_call_output", // ★ 工具结果也是一等 item(无 role 概念)
call_id: string,
output: string }
| { type: "reasoning", // ★ 推理链 item:跨轮保留思维链
summary?: SummaryPart[],
encrypted_content?: string } // 加密形态,回传服务端续接推理
| { type: "item_reference", id: string } // 引用既有 item,免重传
type Response = {
id: string, // "resp_..."
object: "response",
model: string,
status: "completed" | "failed" | "in_progress" | "cancelled",
output: ResponseItem[], // ★ 结构与 input 同构(含 reasoning item)
usage: {
input_tokens: number, // ★ 命名从 prompt_tokens → input_tokens
input_tokens_details?: { cached_tokens: number },
output_tokens: number,
output_tokens_details?: { reasoning_tokens: number },
total_tokens: number
},
previous_response_id?: string | null,
error?: { code: string, message: string } | null
}
| 字段 | 位置 | 含义 | 必用? |
|---|---|---|---|
| model | 请求顶层 | 模型 ID(gpt-5 / o3) | ★ 必填 |
| input | 请求顶层 | 对话载体:string 或 Item[](判别联合) | ★ 必填 |
| instructions | 请求顶层 | 系统指令(Chat 的 role:system 在此顶层) | ★ 常用 |
| previous_response_id | 请求顶层 | 服务端续接历史,只发增量 | ★ Agent 必用 |
| tools | 请求顶层 | 工具定义;FunctionTool 是扁平结构(无 function 包装层) | ★ Agent 必用 |
| tool_choice | 请求顶层 | auto / none / required(无"强制某函数"形式) | ★ 常用 |
| reasoning.effort | 请求顶层 reasoning | 推理强度:minimal/low/medium/high,GPT-5 推理模型专属 | ★ 推理场景 |
| max_output_tokens | 请求顶层 | 最大生成 token;output_ 前缀(区别于 Chat 的 completion_) | ★ 常用 |
| stream | 请求顶层 | 是否流式(类型化事件流,非同构 chunk) | ★ 常用 |
| text.format | 请求顶层 text | 结构化输出:json_object / json_schema | ★ 解析友好场景 |
| store | 请求顶层 | 是否服务端存储响应(默认 true) | ○ 隐私场景 |
| output[] + status | 响应 | 循环分支:completed 收尾 / failed 兜底 / in_progress 中转 | ★ 必读 |
| usage | 响应 | token 计数:input_tokens / output_tokens(含 reasoning_tokens) | ○ 推荐 |
// ── ① 请求:直接传字符串 input(极简入口)+ 扁平工具定义 ──
{
"model": "gpt-5",
"input": "跑一下测试",
"tools": [{"type": "function", "name": "run_tests", // 扁平!无 function 包装层
"parameters": {"type": "object",
"properties": {"suite": {"type": "string"}}}}]
}
// ── ② 响应:output 是 item 数组(含 reasoning + function_call 两种 item)──
{
"id": "resp_001",
"status": "completed",
"output": [
{"type": "reasoning", "summary": []}, // 推理链 item
{"type": "function_call", "call_id": "fc_abc", // call_id 命名
"name": "run_tests", "arguments": "{\"suite\":\"all\"}"}
],
"usage": {"input_tokens": 80, "output_tokens": 40, "total_tokens": 120}
}
// ── ③ 回填:两条 item(function_call 原样 + function_call_output)──
// 注意:与 Chat 相比没有 role:"tool"、没有 messages
{
"model": "gpt-5",
"input": [
{"type": "function_call", "call_id": "fc_abc",
"name": "run_tests", "arguments": "{\"suite\":\"all\"}"},
{"type": "function_call_output", "call_id": "fc_abc",
"output": "42 passed"}
]
}
// ── ④ 进阶:previous_response_id 续接(不用重传 history)──
{
"model": "gpt-5",
"previous_response_id": "resp_001", // 服务端已有 resp_001 的状态
"input": "那覆盖率呢?" // 只发新增输入
}
- 第 ① 段的 input 是字符串"跑一下测试"——Responses 允许这种极简入口。换成 Chat 协议就必须包成 messages: [{role:"user", content:...}],对比一下两边的"最短可工作请求"长度
- 把第 ③ 段 function_call item 的 call_id 改成 fc_xyz 再发,会发生什么?服务端会 400——这是 Chat 的 tool_call_id 配对的 Responses 等价坑
- 第 ④ 段没传 input 数组,只传了字符串"那覆盖率呢?"——服务端会拿这个字符串 + 上一轮的 resp_001 上下文续接。这是 Responses 相对 Chat 的最大成本优势(长程 Agent 不用每轮都重传全量 history)
python3 -c "
import json
chat = json.dumps({'model':'gpt-5','messages':[{'role':'user','content':'跑一下测试'}]}, ensure_ascii=False)
resp = json.dumps({'model':'gpt-5','input':'跑一下测试'}, ensure_ascii=False)
print(f'Chat {len(chat)} 字符 vs Responses {len(resp)} 字符')"
预期输出:Chat 70 字符 vs Responses 36 字符(已实跑验证)——字符串 input 入口几乎省一半。进阶:给两边各补一轮工具回填消息再测,观察差距如何随轮次累计(长程 Agent 的成本账)。
第三篇 · Anthropic Messages Schema
顶层 system + content 必为 block 数组 + max_tokens 必填
Messages 与 Chat 的三个结构差异:system 顶层不进 messages、content 必为数组(text + tool_use 可同轮并存)、max_tokens 必填。三个差异一起决定了协议严格度上限。
- system 顶层参数:可传 string 或 {type:"text", text, cache_control?} 数组
- max_tokens 必填:强制显式声明资源上限(OpenAI 侧全可选)
- cache_control: {type:"ephemeral"}:显式缓存断点,大 system / 工具列表的省钱利器
- 输入侧 block 联合:text / image / tool_result / document,tool_result 嵌在 user 消息里(结构性不同)
content_block_start/stop:流式块边界靠事件
Messages 的流式是显式状态机:每个 SSE 事件带 type 字段,content_block_start / stop 明确告诉解析器"这个 tool_use / thinking 到哪结束"——无需靠 index 拼凑。
- 事件族:message_start / content_block_start / content_block_delta / content_block_stop / message_delta / message_stop / ping
- delta 类型:text_delta / input_json_delta(工具入参分片) / thinking_delta(思维链) / signature_delta
- thinking block:思维链内容 + 防篡改 signature,回填多轮时必须原样保留
- ping 心跳:协议内建连接活性判断,不用自己造
完整 Schema + 实战核心子集 + 工具调用往返
前 2 张卡片拆开了 Messages 的两个剖面(content block 判别联合、流式事件状态机)。这一张把"完整 schema"贴回原样——按 Anthropic 官方 API 参考浓缩,包含请求体、消息与内容块、响应体、流式事件四组对象;然后用一张表划出"实战 80% 用到的 10 个字段";最后用工具调用最小环跑通——展示 Messages 与 Chat/Responses 的三个核心差异(system 顶层 / max_tokens 必填 / tool_result 嵌在 user 里)。
POST /v1/messages x-api-key: <api-key> // ★ 不用 Bearer,用自定义 x-api-key 头 anthropic-version: 2023-06-01 // ★ 显式版本头(必填) content-type: application/json
type MessagesRequest = {
model: string, // "claude-sonnet-4-5" / "claude-opus-4-1" ...
messages: MessageParam[], // 必填
max_tokens: number, // ★ 必填!强制显式声明资源上限
// —— system 是顶层参数(不进 messages)——
system?: string
| { type: "text", text: string, cache_control?: CacheControl }[],
// —— 采样 ——
temperature?: number, // 0~1
top_p?: number, top_k?: number,
stop_sequences?: string[],
// —— 工具(注意:tool_choice 是结构化对象,不是字符串枚举)——
tools?: ToolDefinition[],
tool_choice?: { type: "auto" } // 模型自决
| { type: "any" } // 必须调某个工具
| { type: "tool", name: string }, // 强制调指定工具
// —— 其他 ——
stream?: boolean,
metadata?: { user_id?: string }
}
type CacheControl = { type: "ephemeral" } // ★ 显式缓存断点
type MessageParam = {
role: "user" | "assistant", // ★ 只有两种 role(无 system / tool / developer)
content: string | ContentBlock[] // 空字符串不允许;块可为空数组占位
}
// —— 输入侧块(请求中可出现)——
type ContentBlock =
| { type: "text", text: string, cache_control?: CacheControl }
| { type: "image",
source: { type: "base64", media_type: "image/jpeg"|"image/png"|"image/gif"|"image/webp",
data: string }
| { type: "url", url: string } }
| { type: "tool_result", // ★ 工具结果:嵌在 user 消息里
tool_use_id: string, // 与 ToolUseBlock.id 配对
content?: string | ContentBlock[],
is_error?: boolean, // ★ schema 原生错误标记(区别于 Chat)
cache_control?: CacheControl }
| { type: "document", // PDF 支持
source: {...} }
// —— 输出侧块(响应中出现;thinking 块回填时必须原样保留)——
type OutputBlock =
| { type: "text", text: string }
| { type: "tool_use",
id: string, // "toolu_..."
name: string,
input: object } // ★ 已是对象,不是 JSON 字符串!
| { type: "thinking",
thinking: string, // 思维链内容
signature: string } // 防篡改签名(回填必须原样保留)
| { type: "redacted_thinking", data: string } // 加密思维链
type Message = {
id: string, // "msg_..."
type: "message",
role: "assistant",
model: string,
content: OutputBlock[], // ★ 必为数组(text + tool_use 可同轮并存)
stop_reason: "end_turn" | "max_tokens" | "stop_sequence"
| "tool_use" | "refusal" | "pause_turn",
stop_sequence: string | null,
usage: {
input_tokens: number,
output_tokens: number,
cache_creation_input_tokens?: number, // ★ 缓存计费一级公民
cache_read_input_tokens?: number
}
}
| 字段 | 位置 | 含义 | 必用? |
|---|---|---|---|
| model | 请求顶层 | 模型 ID(claude-sonnet-4-5) | ★ 必填 |
| messages | 请求顶层 | 对话历史;只有 user/assistant 两种 role | ★ 必填 |
| max_tokens | 请求顶层 | 最大生成 token;必填(Anthropic 严格度分水岭) | ★ 必填 |
| system | 请求顶层 | 系统指令(不进 messages,独立顶层参数) | ★ 常用 |
| tools | 请求顶层 | 工具定义;name/input_schema 扁平(无包装层) | ★ Agent 必用 |
| tool_choice | 请求顶层 | 结构化对象:{type:"auto"} / "any" / "tool",name | ★ 常用 |
| temperature | 请求顶层 | 0~1(OpenAI 0~2)——跨厂商代码要钳制范围 | ★ 常用 |
| stream | 请求顶层 | 是否流式;触发显式状态机事件流(content_block_start/stop + ping) | ★ 常用 |
| cache_control | 块级 / system 数组元素 / 工具定义 | 显式缓存断点;大 system / 工具列表的省钱利器 | ★ 长 system 必用 |
| stop_sequences | 请求顶层 | 停止序列(最多 4 个,命名与 OpenAI stop 不同) | ★ 常用 |
| x-api-key + anthropic-version | HTTP 头 | 认证头 + 版本头(必填,区别于 OpenAI 的 Authorization: Bearer) | ★ 必填 |
| stop_reason | 响应 | 循环分支:end_turn / tool_use / max_tokens / refusal | ★ 必读 |
| usage.cache_*_input_tokens | 响应 | 缓存命中/创建的 token 数(成本面板必看) | ○ 推荐 |
// ── ① 请求:max_tokens 必填 + system 顶层 + 扁平工具定义 ──
{
"model": "claude-sonnet-4-5",
"max_tokens": 1024, // ★ 必填
"system": "你是测试助手", // 顶层,不进 messages
"messages": [{"role": "user", "content": "跑一下测试"}],
"tools": [{"name": "run_tests", "description": "运行测试", // ★ 扁平!无 type/function 包装
"input_schema": {"type": "object",
"properties": {"suite": {"type": "string"}}}}]
}
// ── ② 响应:stop_reason = "tool_use";input 已是对象(不是 JSON 字符串)──
{
"id": "msg_001",
"type": "message",
"role": "assistant",
"content": [
{"type": "text", "text": "我先运行测试"},
{"type": "tool_use", "id": "toolu_xyz",
"name": "run_tests", "input": {"suite": "all"}} // ★ input 是对象!
],
"stop_reason": "tool_use",
"usage": {"input_tokens": 95, "output_tokens": 15}
}
// ── ③ 回填:assistant 原样回放 + user 内嵌 tool_result ──
// 注意:tool_result 嵌在 user 消息里(不是独立 role),用 tool_use_id 配对
{
"messages": [
{"role": "user", "content": "跑一下测试"},
{"role": "assistant", "content": [
{"type": "text", "text": "我先运行测试"},
{"type": "tool_use", "id": "toolu_xyz",
"name": "run_tests", "input": {"suite": "all"}}
]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_xyz", // tool_use_id 严格配对
"content": "42 passed"}
]}
]
}
- 把第 ① 段里的 max_tokens 删掉再发,会发生什么?(提示:Anthropic 侧 400 "max_tokens: Field required",对比 Chat 的可选)
- 第 ② 段 tool_use.input 已经是 {"suite": "all"} 对象——Chat/Responses 是 JSON 字符串 "{\"suite\":\"all\"}"。改协议时记得 json.loads 一下
- 第 ③ 段把 tool_use_id 改成 toolu_abc 再发会怎样?400——配对错误。Chat 报 tool_call_id、Responses 报 call_id、Messages 报 tool_use_id,三家命名不同但语义一致
- 如果模型在第 ② 段同时输出了 thinking 块(思维链),第 ③ 段回填时必须原样保留(含 signature)——否则服务端会判为篡改、拒绝续接
python3 -c "
req = {'model':'claude-opus-4-8','messages':[{'role':'user','content':'hi'}]}
required = ['model','max_tokens','messages']
missing = [k for k in required if k not in req]
print('400:', f'{missing[0]}: Field required' if missing else 'OK')"
预期输出:400: max_tokens: Field required(已实跑验证)——与 Anthropic 真实报错文案一致。给 req 补上 'max_tokens': 1024 再跑,输出 OK;再把 required 换成 OpenAI 的 ['model','messages'],体会「严格度分水岭」。
第四篇 · 三协议类型映射总表("翻译词典")
| 概念 | Chat Completions | Responses API | Messages API |
|---|---|---|---|
| 对话载体 | messages[] | input: string | Item[] | messages[] |
| 系统指令 | messages[0].role="system" | 顶层 instructions | 顶层 system |
| 消息角色 | system / developer / user / assistant / tool | 仅 message item 内 user / assistant / system / developer | 仅 user / assistant |
| 内容建模 | 扁平 message + 可选字段 | 判别联合 item | 判别联合 block |
| 工具定义 | {type:"function", function:{name, parameters}} 两层 | {type:"function", name, parameters} 扁平 | {name, input_schema} 无包装 |
| 模型发起调用 | message.tool_calls[].id | item{type:"function_call"}.call_id | block{type:"tool_use"}.id |
| 调用参数 | arguments(JSON 字符串) | arguments(JSON 字符串) | input(对象) |
| 结果回填 | 独立 {role:"tool", tool_call_id} | {type:"function_call_output", call_id} | user 内 tool_result block |
| 错误标记 | 无(content 文本自约定) | 无(output 文本自约定) | is_error: true |
| 长度上限 | max_completion_tokens(可选) | max_output_tokens(可选) | max_tokens(必填) |
| 结束信号 | finish_reason: stop/length/tool_calls | status: completed/failed | stop_reason: end_turn/max_tokens/tool_use |
| 流式形态 | 同构 chat.completion.chunk 流 | 类型化 response.* 事件 | 类型化 message/block/delta 事件 + ping |
| 流式块边界 | 无(按 index 聚合自拼) | output_item.added/done | content_block_start/stop |
| 缓存控制 | 自动(usage 明细报 cached) | 自动 | cache_control 显式断点 |
| 推理链暴露 | 不暴露(仅 reasoning_tokens 计数) | reasoning item(可加密续接) | thinking block(带签名回放) |
| 状态管理 | 无状态 | previous_response_id | 无状态 |
| token 计数字段 | prompt / completion_tokens | input / output_tokens | input / output_tokens + cache 两项 |
| 版本管理 | 无版本头 | 无版本头 | anthropic-version 头 |
第五篇 · 设计规律与权威来源
五条规律:从三家协议收敛方向看未来
三家后发 schema 都收敛到判别联合结构,工具往返配对键是共同的可靠性锚点——理解这五条对自研 Agent 协议 / 轨迹存储格式有直接指导。
- 判别联合是终局:Chat 扁平 message → Responses item、Messages block,两个阵营后发 schema 都收敛到 {type: ...} 自判别结构。自研直接选判别联合
- 工具往返配对键是可靠性锚点:三家共同的 id / call_id / tool_use_id 配对错就 400
- "字符串化 JSON" 两种立场:OpenAI 系把工具参数定为 JSON 字符串(流式可追加分片);Anthropic 定为结构化对象(类型系统友好)
- 严格度分水岭:max_tokens 必填(Anthropic)vs 全可选(OpenAI)——做兼容层时给 Anthropic 侧补默认上限
- 流式协议两代形态:同构 chunk 流(Chat,靠字段缺失表义)vs 类型化事件流(Responses / Messages,显式状态机 + 块边界 + 心跳)。新协议设计一律选后者
- 如果让你自研 Agent 轨迹存储格式,你会用判别联合 item 模型(仿 Responses)还是扁平 message(仿 Chat)?为什么?
官方文档索引:落地实现前必查
三家 schema 均在快速演进(尤其 Responses 与 thinking 相关结构),本文档为教学浓缩版,落地前以官方文档为准。
- OpenAI:platform.openai.com/docs/api-reference/chat / /responses / /guides/migrate-to-responses
- Anthropic:docs.anthropic.com/en/api/messages + platform.claude.com/docs 双域名同源;code.claude.com/docs 是 Claude Code 工具文档勿混
- MCP(工具生态层):modelcontextprotocol.io,Agent 怎么以标准化方式连接成百上千个外部工具
- OpenAI 部分参考页需登录开发者账号查看完整内容