docs: translate developer reference docs to English

Translate TESTING, toolcall-semantics and prompt-compatibility to English.
Keep DeepSeek role-marker tokens verbatim and fix the SSE note link to its
renamed ASCII path.
This commit is contained in:
omar
2026-06-02 22:20:10 +03:00
parent c01b8defb3
commit 0d3bcd4501
3 changed files with 350 additions and 353 deletions
+126 -128
View File
@@ -1,28 +1,26 @@
# DS2API 测试指南
# DS2API Testing Guide
语言 / Language: 中文 + English(同页)
Docs: [Overview](../README.MD) / [Architecture](./ARCHITECTURE.md) / [Deployment guide](./DEPLOY.md) / [API reference](../API.md)
文档导航: [总览](../README.MD) / [架构说明](./ARCHITECTURE.md) / [部署指南](./DEPLOY.md) / [接口文档](../API.md)
## Overview
## 概述 | Overview
DS2API provides two tiers of tests:
DS2API 提供两个层级的测试:
| 层级 | 命令 | 说明 |
| Tier | Command | Notes |
| --- | --- | --- |
| 单元测试(Go) | `./tests/scripts/run-unit-go.sh` | 不需要真实账号 |
| 单元测试(Node) | `./tests/scripts/run-unit-node.sh` | 不需要真实账号 |
| 单元测试(全部) | `./tests/scripts/run-unit-all.sh` | 不需要真实账号 |
| 端到端测试 | `./tests/scripts/run-live.sh` | 使用真实账号执行全链路测试 |
| Unit tests (Go) | `./tests/scripts/run-unit-go.sh` | No real account required |
| Unit tests (Node) | `./tests/scripts/run-unit-node.sh` | No real account required |
| Unit tests (all) | `./tests/scripts/run-unit-all.sh` | No real account required |
| End-to-end tests | `./tests/scripts/run-live.sh` | Full-chain test using a real account |
端到端测试集会录制完整的请求/响应日志,用于故障排查。
Node 单元测试脚本会先做 `node --check` 语法门禁,再以 `--test-concurrency=1` 串行执行测试文件,减少模块级共享状态带来的干扰。
The end-to-end suite records full request/response logs for troubleshooting.
The Node unit test script first runs a `node --check` syntax gate, then executes the test files serially with `--test-concurrency=1` to reduce interference from module-level shared state.
---
## PR 门禁 | PR Gates
## PR Gates
打开或更新 PR 前,按 `.github/workflows/quality-gates.yml` 的同等本地门禁执行:
Before opening or updating a PR, run the local equivalent of the gates in `.github/workflows/quality-gates.yml`:
```bash
./scripts/lint.sh
@@ -31,68 +29,68 @@ Node 单元测试脚本会先做 `node --check` 语法门禁,再以 `--test-co
npm run build --prefix webui
```
说明:
Notes:
- `./scripts/lint.sh` 会运行 Go 格式化检查和 `golangci-lint`;修改 Go 文件后仍建议先执行 `gofmt -w <files>`。
- `run-unit-all.sh` 串行调用 Go 与 Node 单元测试入口。
- `run-live.sh` 是真实账号端到端测试,适合作为发布或高风险改动后的补充验证,不属于每次 PR 的固定本地门禁。
- `./scripts/lint.sh` runs the Go format check and `golangci-lint`; after editing Go files it is still recommended to run `gofmt -w <files>` first.
- `run-unit-all.sh` invokes the Go and Node unit test entrypoints serially.
- `run-live.sh` is the real-account end-to-end test, suitable as supplementary verification after a release or high-risk change; it is not part of the fixed per-PR local gate.
---
## 快速开始 | Quick Start
## Quick Start
### 单元测试 | Unit Tests
### Unit Tests
```bash
./tests/scripts/run-unit-all.sh
```
```bash
# 或按语言拆分执行
# Or run per language
./tests/scripts/run-unit-go.sh
./tests/scripts/run-unit-node.sh
```
```bash
# 结构与流程门禁
# Structure and flow gates
./tests/scripts/check-refactor-line-gate.sh
./tests/scripts/check-node-split-syntax.sh
# 历史阶段门禁:阶段 6 手工烟测签字检查(默认读取 plans/stage6-manual-smoke.md)
# Legacy stage gate: stage 6 manual smoke sign-off check (defaults to plans/stage6-manual-smoke.md)
./tests/scripts/check-stage6-manual-smoke.sh
```
### 端到端测试 | End-to-End Tests
### End-to-End Tests
```bash
./tests/scripts/run-live.sh
```
**默认行为**:
**Default behavior**:
1. **Preflight 检查**:
- `go test ./... -count=1`(单元测试)
- `./tests/scripts/check-node-split-syntax.sh`(Node 拆分模块语法门禁)
1. **Preflight checks**:
- `go test ./... -count=1` (unit tests)
- `./tests/scripts/check-node-split-syntax.sh` (Node split-module syntax gate)
- `node --test tests/node/stream-tool-sieve.test.js tests/node/chat-stream.test.js tests/node/js_compat_test.js`
- `npm run build --prefix webui`(WebUI 构建检查)
- `npm run build --prefix webui` (WebUI build check)
2. **隔离启动**:复制 `config.json` 到临时目录,启动独立服务进程
2. **Isolated startup**: copy `config.json` to a temporary directory and start a standalone service process
3. **场景测试**:
- ✅ OpenAI 非流式 / 流式
- ✅ Claude 非流式 / 流式
- ✅ Admin API(登录 / 配置 / 账号管理)
3. **Scenario tests**:
- ✅ OpenAI non-streaming / streaming
- ✅ Claude non-streaming / streaming
- ✅ Admin API (login / config / account management)
- ✅ Tool Calling
- ✅ 并发压力测试
- ✅ Search 模型
- ✅ Concurrency stress test
- ✅ Search models
4. **结果收集**:继续执行所有用例(不中断),写入最终汇总
4. **Result collection**: continue running all cases (no early stop), then write the final summary
如果你只想跳过这些 preflight 检查,可以直接运行 `go run ./cmd/ds2api-tests --no-preflight`。
If you only want to skip these preflight checks, run `go run ./cmd/ds2api-tests --no-preflight` directly.
---
## CLI 参数 | CLI Flags
## CLI Flags
```bash
go run ./cmd/ds2api-tests \
@@ -106,61 +104,61 @@ go run ./cmd/ds2api-tests \
--keep 5
```
| 参数 | 说明 | 默认值 |
| Flag | Description | Default |
| --- | --- | --- |
| `--config` | 配置文件路径 | `config.json` |
| `--admin-key` | Admin 密钥 | `DS2API_ADMIN_KEY` 环境变量,回退 `admin` |
| `--out` | 产物输出根目录 | `artifacts/testsuite` |
| `--port` | 测试服务端口(`0` = 自动分配空闲端口) | `0` |
| `--timeout` | 单个请求超时秒数 | `120` |
| `--retries` | 网络/5xx 请求重试次数 | `2` |
| `--no-preflight` | 跳过 preflight 检查 | `false` |
| `--keep` | 保留最近几次测试结果(`0` = 全部保留) | `5` |
| `--config` | Config file path | `config.json` |
| `--admin-key` | Admin key | `DS2API_ADMIN_KEY` env var, fallback `admin` |
| `--out` | Artifact output root directory | `artifacts/testsuite` |
| `--port` | Test service port (`0` = auto-assign a free port) | `0` |
| `--timeout` | Per-request timeout in seconds | `120` |
| `--retries` | Retry count for network/5xx requests | `2` |
| `--no-preflight` | Skip preflight checks | `false` |
| `--keep` | Keep the most recent N test runs (`0` = keep all) | `5` |
---
## 自动清理 | Auto Cleanup
## Auto Cleanup
每次测试运行完成后,程序会自动扫描输出目录(`--out`),按时间排序保留最近 `--keep` 次运行的结果,超出部分自动删除。
After each test run completes, the program automatically scans the output directory (`--out`), keeps the most recent `--keep` runs sorted by time, and deletes the rest.
- 默认保留 **5** 次
- 设置 `--keep 0` 可关闭自动清理
- 被删除的旧运行目录会打印日志提示
- Keeps **5** runs by default
- Set `--keep 0` to disable auto cleanup
- Deleted old run directories are logged
---
## 产物结构 | Artifact Layout
## Artifact Layout
每次运行会创建一个以运行 ID 命名的目录:
Each run creates a directory named after the run ID:
```text
artifacts/testsuite/<run_id>/
├── summary.json # 机器可读报告
├── summary.md # 人类可读报告
├── server.log # 测试期间服务端日志
├── preflight.log # Preflight 命令输出
├── summary.json # machine-readable report
├── summary.md # human-readable report
├── server.log # server logs during the test
├── preflight.log # preflight command output
└── cases/
└── <case_id>/
├── request.json # 请求体
├── response.headers # 响应头
├── response.body # 响应体
├── stream.raw # 原始 SSE 数据(流式用例)
├── assertions.json # 断言结果
└── meta.json # 元信息(耗时、状态码等)
├── request.json # request body
├── response.headers # response headers
├── response.body # response body
├── stream.raw # raw SSE data (streaming cases)
├── assertions.json # assertion results
└── meta.json # metadata (elapsed time, status code, etc.)
```
---
## Trace 关联 | Trace Binding
## Trace Binding
每个测试请求自动注入 trace 信息,便于快速定位问题:
Each test request automatically injects trace information for quick troubleshooting:
| 位置 | 格式 |
| Location | Format |
| --- | --- |
| 请求头 | `X-Ds2-Test-Trace: <trace_id>` |
| 查询参数 | `__trace_id=<trace_id>` |
| Request header | `X-Ds2-Test-Trace: <trace_id>` |
| Query parameter | `__trace_id=<trace_id>` |
当用例失败时,`summary.md` 中会包含 trace ID。你可以快速搜索对应的服务端日志:
When a case fails, `summary.md` includes the trace ID. You can quickly search the corresponding server logs:
```bash
rg "<trace_id>" artifacts/testsuite/<run_id>/server.log
@@ -168,136 +166,136 @@ rg "<trace_id>" artifacts/testsuite/<run_id>/server.log
---
## 退出码 | Exit Code
## Exit Code
| 退出码 | 含义 |
| Exit code | Meaning |
| --- | --- |
| `0` | 所有用例通过 ✅ |
| `1` | 有用例失败 ❌ |
| `0` | All cases passed ✅ |
| `1` | Some cases failed ❌ |
可将测试集作为本地发布门禁使用(CI/CD 集成)。
The suite can be used as a local release gate (CI/CD integration).
---
## 安全提醒 | Sensitive Data Warning
## Sensitive Data Warning
⚠️ 测试集会存储**完整的原始请求/响应载荷**用于调试。
⚠️ The suite stores **full raw request/response payloads** for debugging.
- **不要**将 artifacts 目录上传到公开仓库
- **不要**在 Issue tracker 中分享未脱敏的 artifact 文件
- 如需分享日志,请先手动清除敏感信息(token、密码等)
- **Do not** upload the artifacts directory to a public repository
- **Do not** share un-redacted artifact files in an issue tracker
- If you need to share logs, manually clear sensitive information (tokens, passwords, etc.) first
---
## 常见用法 | Common Usage
## Common Usage
### 仅跑单元测试
### Run unit tests only
```bash
go test ./...
```
### 运行特定模块的单元测试
### Run unit tests for a specific module
```bash
# 运行 tool calls 相关测试(推荐用于调试 tool call 解析问题)
# Run tool-call tests (recommended for debugging tool-call parsing issues)
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/
# 运行单个测试用例
# Run a single test case
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/
# 运行 format 相关测试
# Run format tests
go test -v ./internal/format/...
# 运行 HTTP API 相关测试
# Run HTTP API tests
go test -v ./internal/httpapi/openai/...
```
### 调试 Tool Call 问题 | Debugging Tool Call Issues
### Debugging Tool Call Issues
当遇到 DeepSeek 工具调用解析问题时,可以使用以下方法:
When you hit a DeepSeek tool-call parsing issue, you can use the following approaches:
```bash
# 1. 运行 tool calls 相关的所有测试
# 1. Run all tool-call tests
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/
# 2. 查看测试输出中的详细调试信息
# 2. Inspect the detailed debug output in the test logs
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/ 2>&1
# 3. 检查具体测试用例的修复效果
# 测试用例位于 internal/toolcall/toolcalls_test.go,包含:
# - TestParseToolCallsWithDeepSeekHallucination: DeepSeek 典型幻觉输出
# - TestRepairLooseJSONWithNestedObjects: 嵌套对象的方括号修复
# - TestParseToolCallsWithMixedWindowsPaths: Windows 路径处理
# 3. Check the fix behavior for specific test cases
# Test cases live in internal/toolcall/toolcalls_test.go, including:
# - TestParseToolCallsWithDeepSeekHallucination: typical DeepSeek hallucinated output
# - TestRepairLooseJSONWithNestedObjects: bracket repair for nested objects
# - TestParseToolCallsWithMixedWindowsPaths: Windows path handling
```
### 运行 Node.js 测试
### Run Node.js tests
```bash
# 运行 Node 测试
# Run Node tests
node --test tests/node/stream-tool-sieve.test.js
# 或使用脚本
# Or use the script
./tests/scripts/run-unit-node.sh
```
### 跑端到端测试(跳过 preflight)
### Run end-to-end tests (skip preflight)
```bash
go run ./cmd/ds2api-tests --no-preflight
```
### 运行原始流仿真(独立工具)
### Run the raw-stream simulation (standalone tool)
```bash
./tests/scripts/run-raw-stream-sim.sh
```
说明:
- 该工具默认重放 `tests/raw_stream_samples/manifest.json` 声明的 canonical 样本,按上游 SSE 顺序做 1:1 仿真解析。
- 默认校验不出现 `FINISHED` 文本泄露,并要求存在结束信号。
- 默认**不**把 `raw accumulated_token_usage` 与本地解析 token 做强一致校验(当前实现以内容估算为准);如需强校验可显式加 `--fail-on-token-mismatch`。
- 每次运行都会把本地派生结果写入 `artifacts/raw-stream-sim/<run-id>/<sample-id>/replay.output.txt`,并输出结构化报告。
- 如果你有历史基线目录,可以通过 `--baseline-root` 让工具直接做文本对比。
- 更完整的协议级行为结构说明见 [DeepSeekSSE行为结构说明-2026-04-05.md](./DeepSeekSSE行为结构说明-2026-04-05.md)。
Notes:
- By default this tool replays the canonical samples declared in `tests/raw_stream_samples/manifest.json`, doing a 1:1 simulation parse following the upstream SSE order.
- By default it verifies there is no `FINISHED` text leak and requires an end signal to be present.
- By default it does **not** strictly compare `raw accumulated_token_usage` against locally parsed tokens (the current implementation estimates from content); add `--fail-on-token-mismatch` explicitly to enforce a strict check.
- Every run writes the locally derived results to `artifacts/raw-stream-sim/<run-id>/<sample-id>/replay.output.txt` and emits a structured report.
- If you have a historical baseline directory, use `--baseline-root` to have the tool do a direct text comparison.
- For a more complete protocol-level behavior description, see [deepseek-sse-behavior-2026-04-05.md](./deepseek-sse-behavior-2026-04-05.md).
### 对单个样本做回放比对
### Replay-compare a single sample
```bash
./tests/scripts/compare-raw-stream-sample.sh markdown-format-example-20260405-spacefix
```
说明:
- 该脚本会从 raw-only 样本目录读取 `upstream.stream.sse`。
- 回放结果会写入 `artifacts/raw-stream-sim/<run-id>/<sample-id>/`,便于直接查阅。
- 如果传入历史基线目录,脚本会自动对比当前回放输出和基线文本。
Notes:
- This script reads `upstream.stream.sse` from the raw-only sample directory.
- Replay results are written to `artifacts/raw-stream-sim/<run-id>/<sample-id>/` for easy inspection.
- If a historical baseline directory is provided, the script automatically compares the current replay output against the baseline text.
### 采集永久样本
### Capture a permanent sample
本地启动服务后,可以直接打:
After starting the service locally, you can call:
```bash
POST /admin/dev/raw-samples/capture
```
这个接口会把请求元信息和上游原始流写入 `tests/raw_stream_samples/<sample-id>/`,以后可以直接拿来做回放和字段分析。派生输出会在本地回放时再生成,不再落在样本目录里。
This endpoint writes the request metadata and the upstream raw stream to `tests/raw_stream_samples/<sample-id>/`, which can later be reused for replay and field analysis. Derived output is regenerated during local replay and is no longer stored in the sample directory.
### 从内存抓包查询并保存样本
### Query in-memory captures and save samples
如果问题刚刚在本地复现过,也可以先查当前进程内存里的抓包,再选择性落盘:
If the issue was just reproduced locally, you can first query the captures in the current process memory, then optionally persist them:
```bash
GET /admin/dev/raw-samples/query?q=广州&limit=10
GET /admin/dev/raw-samples/query?q=guangzhou&limit=10
POST /admin/dev/raw-samples/save
{"chain_key":"session:xxxx","sample_id":"tmp-from-memory"}
```
说明:
- `query` 会按 `chat_session_id` 把 `completion + continue` 归并成一条链,适合定位接续思考问题。
- `save` 支持用 `query`、`chain_key` 或 `capture_id` 选中目标。
- 生成的样本目录仍然是 `tests/raw_stream_samples/<sample-id>/`,可以直接喂给回放脚本。
Notes:
- `query` merges `completion + continue` into a single chain by `chat_session_id`, which is handy for locating continued-thinking issues.
- `save` supports selecting the target via `query`, `chain_key`, or `capture_id`.
- The generated sample directory is still `tests/raw_stream_samples/<sample-id>/` and can be fed directly to the replay script.
### 指定输出目录和超时
### Specify output directory and timeout
```bash
go run ./cmd/ds2api-tests \
@@ -305,10 +303,10 @@ go run ./cmd/ds2api-tests \
--timeout 60
```
### 在 CI 中使用
### Use in CI
```bash
# 确保 config.json 存在且包含有效测试账号
# Make sure config.json exists and contains a valid test account
./tests/scripts/run-live.sh
exit_code=$?
if [ $exit_code -ne 0 ]; then
+185 -186
View File
@@ -1,81 +1,81 @@
# API -> 网页对话纯文本兼容主链路说明
# API -> web-chat pure-text compatibility pipeline
文档导航:[总览](../README.MD) / [架构说明](./ARCHITECTURE.md) / [接口文档](../API.md) / [测试指南](./TESTING.md)
Docs: [Overview](../README.MD) / [Architecture](./ARCHITECTURE.md) / [API reference](../API.md) / [Testing guide](./TESTING.md)
> 本文档是 DS2API“把 OpenAI / Claude / Gemini 风格 API 请求兼容成 DeepSeek 网页对话纯文本上下文”的专项说明。
> 这是项目最重要的兼容产物之一。凡是修改消息标准化、tool prompt 注入、tool history 保留、文件引用、history split、下游 completion payload 组装等行为,都必须同步更新本文档。
> This document is the dedicated reference for how DS2API "compatibilizes OpenAI / Claude / Gemini style API requests into DeepSeek web-chat pure-text context".
> It is one of the project's most important compatibility products. Any change to message normalization, tool prompt injection, tool history retention, file references, history split, or downstream completion payload assembly must update this document in sync.
## 1. 核心结论
## 1. Core conclusion
DS2API 当前的核心思路,不是把客户端传来的 `messages`、`tools`、`attachments` 原样转发给下游。
DS2API's current core idea is not to forward the client's `messages`, `tools`, and `attachments` to the downstream as-is.
而是把这些高层 API 语义,统一压缩成 DeepSeek 网页对话更容易理解的三类输入:
Instead, it compresses those high-level API semantics into three kinds of input that DeepSeek web chat understands more easily:
1. `prompt`
一个单字符串,里面带有角色标记、system 指令、历史消息、assistant reasoning 标签、历史 tool call XML 等。
A single string containing role markers, system instructions, history messages, assistant reasoning tags, historical tool-call XML, etc.
2. `ref_file_ids`
一个文件引用数组,承载附件、inline 上传文件,以及必要时被拆出去的历史文件。
3. 控制位
例如 `thinking_enabled`、`search_enabled`、部分 passthrough 参数。
An array of file references carrying attachments, inline uploaded files, and, when necessary, history files that were split out.
3. Control flags
For example `thinking_enabled`, `search_enabled`, and some passthrough parameters.
也就是说,项目最重要的兼容动作,是把“结构化 API 会话”翻译成“网页对话纯文本上下文 + 文件引用”。
In other words, the project's most important compatibility action is translating a "structured API conversation" into a "web-chat pure-text context + file references".
## 2. 为什么这是核心产物
## 2. Why this is the core product
因为对下游来说,真正稳定的输入面不是 OpenAI/Claude/Gemini 的原生 schema,而是:
Because for the downstream, the truly stable input surface is not the native schema of OpenAI/Claude/Gemini, but:
- 一段连续的对话 prompt
- 一组可引用文件
- 少量开关位
- A continuous conversation prompt
- A set of referenceable files
- A few flags
这也是为什么很多表面上看像“协议兼容”的代码,最终都会收敛到同一类逻辑:
This is also why a lot of code that looks like "protocol compatibility" eventually converges to the same kind of logic:
- 先把不同协议的消息统一成内部消息序列
- 再把工具声明改写成 system prompt 文本
- 再把历史 tool call / tool result 改写成 prompt 可见内容
- 最后输出成 DeepSeek completion payload
- First normalize the messages of different protocols into an internal message sequence
- Then rewrite tool declarations into system prompt text
- Then rewrite historical tool calls / tool results into prompt-visible content
- Finally emit a DeepSeek completion payload
## 3. 统一心智模型
## 3. Unified mental model
当前主链路可以这样理解:
The current main pipeline can be understood like this:
```text
客户端请求
-> HTTP API surface(OpenAI / Claude / Gemini)
-> promptcompat 统一消息标准化
-> tool prompt 注入
-> DeepSeek 风格 prompt 拼装
-> 文件收集 / inline 上传 / history split(OpenAI 链路)
client request
-> HTTP API surface (OpenAI / Claude / Gemini)
-> promptcompat unified message normalization
-> tool prompt injection
-> DeepSeek-style prompt assembly
-> file collection / inline upload / history split (OpenAI path)
-> completion payload
-> 下游网页对话接口
-> downstream web-chat endpoint
```
标准化后的最终 prompt 统一限制为 131072 个 Unicode 字符;OpenAI、Claude、Gemini 三条链路都会在上游调用前返回 `context too long` 错误。
The final normalized prompt is uniformly limited to 131072 Unicode characters; all three paths (OpenAI, Claude, Gemini) return a `context too long` error before the upstream call.
对应的关键代码入口:
Key code entrypoints:
- OpenAI Chat / Responses:
- OpenAI Chat / Responses:
[internal/promptcompat/request_normalize.go](../internal/promptcompat/request_normalize.go)
- OpenAI prompt 组装:
- OpenAI prompt assembly:
[internal/promptcompat/prompt_build.go](../internal/promptcompat/prompt_build.go)
- OpenAI 消息标准化:
- OpenAI message normalization:
[internal/promptcompat/message_normalize.go](../internal/promptcompat/message_normalize.go)
- Claude 标准化:
- Claude normalization:
[internal/httpapi/claude/standard_request.go](../internal/httpapi/claude/standard_request.go)
- Claude 消息与 tool_use/tool_result 归一:
- Claude message and tool_use/tool_result normalization:
[internal/httpapi/claude/handler_utils.go](../internal/httpapi/claude/handler_utils.go)
- Gemini 复用 OpenAI prompt builder:
- Gemini reuses the OpenAI prompt builder:
[internal/httpapi/gemini/convert_request.go](../internal/httpapi/gemini/convert_request.go)
- DeepSeek prompt 角色标记拼装:
- DeepSeek prompt role-marker assembly:
[internal/prompt/messages.go](../internal/prompt/messages.go)
- prompt 可见 tool history XML:
- prompt-visible tool history XML:
[internal/prompt/tool_calls.go](../internal/prompt/tool_calls.go)
- completion payload:
- completion payload:
[internal/promptcompat/standard_request.go](../internal/promptcompat/standard_request.go)
## 4. 下游真正收到的东西
## 4. What the downstream actually receives
在“完成标准化后”,下游 completion payload 的核心形态是:
After normalization completes, the core shape of the downstream completion payload is:
```json
{
@@ -93,20 +93,20 @@ DS2API 当前的核心思路,不是把客户端传来的 `messages`、`tools`
}
```
重点是:
Key points:
- `prompt` 才是对话上下文主载体。
- `ref_file_ids` 只承载文件引用,不承载普通文本消息。
- `tools` 不会作为“原生工具 schema”直接下发给下游,而是被改写进 `prompt`。
- OpenAI Chat / Responses 原生走统一 OpenAI 标准化与 DeepSeek payload 组装;Claude / Gemini 会尽量复用 OpenAI prompt/tool 语义,其中 Gemini 直接复用 `promptcompat.BuildOpenAIPromptForAdapter`,Claude 消息接口在可代理场景会转换为 OpenAI chat 形态再执行。
- 客户端传入的 thinking / reasoning 开关会被归一到下游 `thinking_enabled`。Gemini `generationConfig.thinkingConfig.thinkingBudget` 会翻译成同一套 thinking 开关;关闭时即使上游返回 `response/thinking_content`,兼容层也不会把它当作可见正文输出。Claude surface 在流式请求且未显式声明 `thinking` 时,仍按 Anthropic 语义默认关闭;但在非流式代理场景,兼容层会内部开启一次下游 thinking,用于捕获“正文为空、工具调用落在 thinking 里”的情况,随后在回包前剥离用户不可见的 thinking block。
- 对 OpenAI Chat / Responses 的非流式收尾,如果最终可见正文为空,兼容层会优先尝试把思维链中的独立 `<tool_calls>...</tool_calls>` 结构当作真实工具调用解析出来。流式链路也会在收尾阶段做同样的 fallback 检测,但不会因为思维链内容去中途拦截或改写流式输出;thinking / reasoning 增量仍按原样先发,只有在结束收尾时才可能补发最终工具调用结果。只有正文为空且思维链里也没有可执行工具调用时,才继续按空回复错误处理。
- `prompt` is the main carrier of the conversation context.
- `ref_file_ids` carries only file references, not ordinary text messages.
- `tools` are not sent downstream as a "native tool schema"; they are rewritten into the `prompt`.
- OpenAI Chat / Responses natively go through the unified OpenAI normalization and DeepSeek payload assembly; Claude / Gemini reuse the OpenAI prompt/tool semantics as much as possible, where Gemini directly reuses `promptcompat.BuildOpenAIPromptForAdapter`, and the Claude messages endpoint, in proxyable scenarios, is converted to the OpenAI chat shape before execution.
- The thinking / reasoning flags passed by the client are normalized to the downstream `thinking_enabled`. Gemini `generationConfig.thinkingConfig.thinkingBudget` is translated to the same thinking flag; when disabled, even if the upstream returns `response/thinking_content`, the compatibility layer does not treat it as visible body output. On the Claude surface, for streaming requests without an explicit `thinking` declaration, Anthropic semantics still default it off; but in the non-streaming proxy scenario the compatibility layer internally enables downstream thinking once to capture the case of "empty body, tool call landing inside thinking", then strips the user-invisible thinking block before responding.
- For the non-streaming finalize of OpenAI Chat / Responses, if the final visible body is empty, the compatibility layer first tries to parse a standalone `<tool_calls>...</tool_calls>` structure in the chain of thought as a real tool call. The streaming path also performs the same fallback detection at finalize, but does not intercept or rewrite streaming output mid-flight because of chain-of-thought content; thinking / reasoning deltas are still emitted as-is first, and only at finalize may the final tool-call result be appended. Only when the body is empty and the chain of thought also has no executable tool call does it continue to be handled as an empty-reply error.
## 5. prompt 是怎么拼出来的
## 5. How the prompt is assembled
### 5.1 角色标记
### 5.1 Role markers
最终 prompt 使用 DeepSeek 风格角色标记:
The final prompt uses DeepSeek-style role markers:
- `<|begin▁of▁sentence|>`
- `<|System|>`
@@ -117,63 +117,63 @@ DS2API 当前的核心思路,不是把客户端传来的 `messages`、`tools`
- `<|end▁of▁sentence|>`
- `<|end▁of▁toolresults|>`
实现位置:
Implementation:
[internal/prompt/messages.go](../internal/prompt/messages.go)
### 5.2 thinking continuity 说明
### 5.2 Thinking continuity notes
如果启用了 thinking,会在最前面额外插入一个 system block,提醒模型:
When thinking is enabled, an extra system block is inserted at the very front to remind the model to:
- 继续既有会话,不要重开
- earlier messages 是 binding context
- 不要把最终回答只留在 reasoning 里
- Continue the existing session, not start over
- Treat earlier messages as binding context
- Not leave the final answer only in the reasoning
这部分不是客户端原始消息,而是兼容层主动补进去的连续性契约。
This part is not the client's original message; it is a continuity contract that the compatibility layer adds proactively.
### 5.3 相邻同角色消息会合并
### 5.3 Adjacent same-role messages are merged
在最终 `MessagesPrepareWithThinking` 中,相邻同 role 的消息会被合并成一个块,中间插入空行。
In the final `MessagesPrepareWithThinking`, adjacent messages with the same role are merged into a single block, with a blank line inserted between them.
这意味着:
This means:
- prompt 中看到的是“合并后的 role block”
- 不是客户端传来的逐条 message 原样排列
- What you see in the prompt is a "merged role block"
- Not the per-message arrangement passed by the client as-is
## 6. tools 为什么是“文本注入”,不是原生下发
## 6. Why tools are "text injection", not native passthrough
当前项目把工具能力视为“prompt 约束的一部分”。
The current project treats tool capability as "part of the prompt constraints".
具体做法:
Concretely:
1. 把每个 tool 的名称、描述、参数 schema 序列化成文本。
2. 拼成 `You have access to these tools:` 大段说明。
3. 再附上统一的 XML tool call 格式约束。
4. 把这整段内容并入 system prompt。
1. Serialize each tool's name, description, and parameter schema into text.
2. Assemble a large `You have access to these tools:` block.
3. Append a unified XML tool-call format constraint.
4. Merge this whole block into the system prompt.
工具调用正例仍只示范 canonical XML:`<tool_calls>` → `<invoke name="...">` → `<parameter name="...">`。
提示词会额外强调:如果要调用工具,工具块的首个非空白字符必须就是 `<tool_calls>`,不能只输出 `</tool_calls>` 而漏掉 opening tag。
正例中的工具名只会来自当前请求实际声明的工具;如果当前请求没有足够的已知工具形态,就省略对应的单工具、多工具或嵌套示例,避免把不可用工具名写进 prompt。
对执行类工具,脚本内容必须进入执行参数本身:`Bash` / `execute_command` 使用 `command`,`exec_command` 使用 `cmd`;不要把脚本示范成 `path` / `content` 文件写入参数。
The positive tool-call example still only demonstrates canonical XML: `<tool_calls>` → `<invoke name="...">` → `<parameter name="...">`.
The prompt additionally emphasizes: if a tool is to be called, the first non-whitespace character of the tool block must be `<tool_calls>`; it must not emit only `</tool_calls>` and drop the opening tag.
The tool names in the positive examples only come from tools actually declared in the current request; if the current request lacks enough known tool shapes, the corresponding single-tool, multi-tool, or nested examples are omitted to avoid writing unavailable tool names into the prompt.
For execution-type tools, the script content must go into the execution argument itself: `Bash` / `execute_command` use `command`, `exec_command` uses `cmd`; do not demonstrate the script as a `path` / `content` file-write argument.
OpenAI 路径实现:
OpenAI path implementation:
[internal/promptcompat/tool_prompt.go](../internal/promptcompat/tool_prompt.go)
Claude 路径实现:
Claude path implementation:
[internal/httpapi/claude/handler_utils.go](../internal/httpapi/claude/handler_utils.go)
统一工具调用格式模板:
Unified tool-call format template:
[internal/toolcall/tool_prompt.go](../internal/toolcall/tool_prompt.go)
这也是项目“网页对话纯文本兼容”的关键设计:
This is also the key design of the project's "web-chat pure-text compatibility":
- tools 对下游来说,本质上是 prompt 内规则
- 不是 native tool schema transport
- For the downstream, tools are essentially in-prompt rules
- Not a native tool schema transport
## 7. assistant 的 tool_calls / reasoning 如何保留
## 7. How assistant tool_calls / reasoning are retained
### 7.1 reasoning 保留方式
### 7.1 How reasoning is retained
assistant 的 reasoning 会变成一个显式标签块:
The assistant's reasoning becomes an explicit tagged block:
```text
[reasoning_content]
@@ -181,11 +181,11 @@ assistant 的 reasoning 会变成一个显式标签块:
[/reasoning_content]
```
然后再接可见回答正文。
followed by the visible answer body.
### 7.2 历史 tool_calls 保留方式
### 7.2 How historical tool_calls are retained
assistant 历史 `tool_calls` 不会保留成 OpenAI 原生 JSON,而会转成 prompt 可见的 XML:
The assistant's historical `tool_calls` are not retained as OpenAI native JSON; they are converted to prompt-visible XML:
```xml
<tool_calls>
@@ -195,72 +195,72 @@ assistant 历史 `tool_calls` 不会保留成 OpenAI 原生 JSON,而会转成
</tool_calls>
```
这也是当前项目里唯一受支持的 canonical tool-calling 形态;其他形态都会作为普通文本保留,不会作为可执行调用语法。
例外是 parser 会对一个非常窄的模型失误做修复:如果 assistant 输出了 `<invoke ...>` ... `</tool_calls>`,但漏掉最前面的 opening `<tool_calls>`,解析阶段会补回 wrapper 后再尝试识别。
This is also the only supported canonical tool-calling shape in the current project; all other shapes are kept as plain text and are not treated as executable call syntax.
The exception is that the parser fixes a very narrow model mistake: if the assistant emits `<invoke ...>` ... `</tool_calls>` but drops the leading opening `<tool_calls>`, the parsing stage restores the wrapper before attempting recognition.
这件事很重要,因为它决定了:
This matters because it determines that:
- 历史工具调用在 prompt 中是“可见文本历史”
- 不是“隐藏结构化元数据”
- Historical tool calls are "visible text history" in the prompt
- Not "hidden structured metadata"
实现位置:
Implementation:
[internal/prompt/tool_calls.go](../internal/prompt/tool_calls.go)
### 7.3 tool result 保留方式
### 7.3 How tool results are retained
tool / function role 的结果会作为 `<|Tool|>...<|end▁of▁toolresults|>` 进入 prompt。
Results of the tool / function role enter the prompt as `<|Tool|>...<|end▁of▁toolresults|>`.
如果 tool content 为空,当前会补成字符串 `"null"`,避免整个 tool turn 丢失。
If the tool content is empty, it is currently filled with the string `"null"` to avoid losing the entire tool turn.
## 8. files、附件、systemprompt 文件的实际语义
## 8. Actual semantics of files, attachments, and systemprompt files
这里要明确区分两类东西:
Here we must clearly distinguish two things:
1. 文本型 system prompt
例如 OpenAI `developer` / `system` / Responses `instructions` / Claude top-level `system`
这类会进入 `prompt`。
2. 文件型 systemprompt
例如通过附件、`input_file`、base64、data URL 上传的文件
这类不会直接内联进 `prompt`,而是进入 `ref_file_ids`。
1. Text-type system prompt
For example OpenAI `developer` / `system` / Responses `instructions` / Claude top-level `system`
These go into the `prompt`.
2. File-type systemprompt
For example files uploaded via attachment, `input_file`, base64, or data URL
These are not inlined directly into the `prompt`; they go into `ref_file_ids`.
OpenAI 文件相关实现:
OpenAI file-related implementation:
- inline/base64/data URL 上传:
- inline/base64/data URL upload:
[internal/httpapi/openai/files/file_inline_upload.go](../internal/httpapi/openai/files/file_inline_upload.go)
- 当 `runtime.disable_upstream_file_uploads=true` 时,显式 `/v1/files`、inline/base64 上传和 history split 的上游文件上传都会关闭;已有 `file_id` / `ref_file_ids` 仍按普通引用收集。
- 文件 ID 收集:
- When `runtime.disable_upstream_file_uploads=true`, upstream file uploads for explicit `/v1/files`, inline/base64 upload, and history split are all turned off; existing `file_id` / `ref_file_ids` are still collected as ordinary references.
- File ID collection:
[internal/promptcompat/file_refs.go](../internal/promptcompat/file_refs.go)
结论:
Conclusion:
- “systemprompt 文字”在 prompt 里
- “systemprompt 文件”通常只在 `ref_file_ids` 里
- "systemprompt text" is in the prompt
- "systemprompt files" are usually only in `ref_file_ids`
除非调用方自己把文件内容展开后再塞进 system/developer 文本,否则文件内容不会自动出现在 prompt 正文。
Unless the caller expands the file content themselves and stuffs it into the system/developer text, the file content does not automatically appear in the prompt body.
## 9. 多轮历史为什么不会一直完整内联在 prompt
## 9. Why multi-turn history is not always fully inlined in the prompt
history split 现在全局强制开启;旧配置中的 `history_split.enabled=false` 会被忽略。默认从第 2 个 user turn 起就可能触发,仍可通过 `history_split.trigger_after_turns` 调整触发阈值。
History split is now globally forced on; `history_split.enabled=false` in old configs is ignored. By default it may trigger from the 2nd user turn, and the trigger threshold can still be adjusted via `history_split.trigger_after_turns`.
相关实现:
Related implementation:
- 配置访问器:
- Config accessors:
[internal/config/store_accessors.go](../internal/config/store_accessors.go)
- 历史拆分:
- History split:
[internal/httpapi/openai/history/history_split.go](../internal/httpapi/openai/history/history_split.go)
触发后行为:
Behavior after triggering:
1. 旧历史消息被切出去。
2. 旧历史会被重新序列化成一个文本文件。
3. 真正上传的文件名固定是 `HISTORY.txt`。
4. 文件内容内部会使用 `IGNORE` 这层包装名来闭合 DeepSeek 官网原生文件标记。
5. 该文件上传后,其 `file_id` 会排在 `ref_file_ids` 最前面。
6. live prompt 只保留:
1. Old history messages are cut out.
2. The old history is re-serialized into a text file.
3. The actually uploaded file name is fixed as `HISTORY.txt`.
4. Inside the file content, the wrapper name `IGNORE` is used to close DeepSeek's native file markers.
5. After the file is uploaded, its `file_id` is placed first in `ref_file_ids`.
6. The live prompt keeps only:
- system / developer
- 最新 user turn 起的上下文
- context from the latest user turn onward
历史文件内容不是普通自由文本,而是用同一套角色标记再次序列化出的 transcript:
The history file content is not ordinary free text; it is a transcript re-serialized with the same set of role markers:
```text
[uploaded filename]: HISTORY.txt
@@ -272,59 +272,59 @@ history split 现在全局强制开启;旧配置中的 `history_split.enabled=
[file content begin]
```
所以“完整上下文”在当前实现里,其实通常分散在两处:
So the "full context" in the current implementation is usually split across two places:
- `prompt` 里的 live context
- `ref_file_ids` 指向的 history transcript file
- The live context in `prompt`
- The history transcript file pointed to by `ref_file_ids`
## 10. 各协议入口的差异
## 10. Differences between protocol entrypoints
### 10.1 OpenAI Chat / Responses
特点:
Characteristics:
- `developer` 会映射到 `system`
- Responses `instructions` 会 prepend 为 system message
- `tools` 会注入 system prompt
- `attachments` / `input_file` / inline 文件会进入 `ref_file_ids`
- history split 主要在这条链路里生效
- `developer` maps to `system`
- Responses `instructions` is prepended as a system message
- `tools` are injected into the system prompt
- `attachments` / `input_file` / inline files go into `ref_file_ids`
- History split mainly takes effect in this path
### 10.2 Claude Messages
特点:
Characteristics:
- top-level `system` 优先作为系统提示
- `tool_use` / `tool_result` 会被转换成统一的 assistant/tool 历史语义
- `tools` 同样会被并进 system prompt
- 常规执行通过 `internal/httpapi/claude/handler_messages.go` 转到 OpenAI chat 路径,模型 alias 会先解析成 DeepSeek 原生模型
- 当前代码里没有像 OpenAI 那样完整的 `ref_file_ids` 附件链路
- The top-level `system` takes priority as the system prompt
- `tool_use` / `tool_result` are converted into the unified assistant/tool history semantics
- `tools` are likewise merged into the system prompt
- Regular execution goes through `internal/httpapi/claude/handler_messages.go` to the OpenAI chat path, with the model alias resolved to a DeepSeek native model first
- The current code does not have a full `ref_file_ids` attachment path like OpenAI does
### 10.3 Gemini
特点:
Characteristics:
- `systemInstruction`、`contents.parts`、`functionCall`、`functionResponse` 会先归一
- tools 会转成 OpenAI 风格 function schema
- prompt 构建复用 OpenAI 的 `promptcompat.BuildOpenAIPromptForAdapter`
- 未识别的非文本 part 会被安全序列化进 prompt,并对二进制/疑似 base64 内容做省略或截断处理
- `systemInstruction`, `contents.parts`, `functionCall`, `functionResponse` are normalized first
- tools are converted to OpenAI-style function schemas
- prompt construction reuses OpenAI's `promptcompat.BuildOpenAIPromptForAdapter`
- Unrecognized non-text parts are safely serialized into the prompt, with binary/suspected-base64 content omitted or truncated
也就是说,Gemini 在“最终 prompt 语义”上,尽量和 OpenAI 保持一致。
In other words, Gemini stays as close to OpenAI as possible at the level of "final prompt semantics".
## 11. 一份贴近真实的最终上下文示意
## 11. A realistic example of the final context
假设用户发来一个多轮请求:
Suppose the user sends a multi-turn request:
- 有 system/developer 文本
- 有 tools
- 有一个文件型 systemprompt 附件
- 有历史 assistant tool call / tool result
- history split 已触发
- with system/developer text
- with tools
- with a file-type systemprompt attachment
- with historical assistant tool call / tool result
- history split already triggered
那么最终上下文更接近:
Then the final context looks closer to:
```json
{
"prompt": "<|begin▁of▁sentence|><|System|>continuity instructions...\\n\\n原 system / developer\\n\\nYou have access to these tools: ...<|end▁of▁instructions|><|User|>最新问题<|Assistant|>",
"prompt": "<|begin▁of▁sentence|><|System|>continuity instructions...\\n\\noriginal system / developer\\n\\nYou have access to these tools: ...<|end▁of▁instructions|><|User|>latest question<|Assistant|>",
"ref_file_ids": [
"file-history-ignore",
"file-systemprompt",
@@ -335,28 +335,28 @@ history split 现在全局强制开启;旧配置中的 `history_split.enabled=
}
```
这正是“API 转网页对话纯文本”的核心成果:
This is exactly the core result of "API to web-chat pure text":
- 大部分结构化语义被压进 `prompt`
- 文件保持文件
- 历史必要时拆文件
- Most structured semantics are compressed into the `prompt`
- Files stay files
- History is split into a file when necessary
## 12. 修改时必须同步本文档的场景
## 12. Scenarios that must update this document
只要触碰以下任一类行为,就必须在同一提交或同一 PR 中更新本文档:
Whenever you touch any of the following behaviors, you must update this document in the same commit or PR:
- 角色映射变更
- system / developer / instructions 合并规则变更
- assistant reasoning 保留格式变更
- assistant 历史 `tool_calls` 的 XML 呈现方式变更
- tool result 注入方式变更
- tool prompt 模板或 tool_choice 约束变更
- inline 文件上传 / 文件引用收集规则变更
- history split 触发条件、上传格式、`IGNORE` 包装格式变更
- completion payload 字段语义变更
- Claude / Gemini 对这套统一语义的复用关系变更
- Role mapping changes
- system / developer / instructions merge rule changes
- assistant reasoning retention format changes
- changes to the XML presentation of assistant historical `tool_calls`
- tool result injection method changes
- tool prompt template or tool_choice constraint changes
- inline file upload / file reference collection rule changes
- history split trigger conditions, upload format, or `IGNORE` wrapper format changes
- completion payload field semantics changes
- changes to how Claude / Gemini reuse this unified semantics
优先检查这些文件:
Check these files first:
- `internal/promptcompat/request_normalize.go`
- `internal/promptcompat/prompt_build.go`
@@ -375,9 +375,9 @@ history split 现在全局强制开启;旧配置中的 `history_split.enabled=
- `internal/prompt/tool_calls.go`
- `internal/promptcompat/standard_request.go`
## 13. 建议的最小验证
## 13. Suggested minimal verification
改动这条链路后,至少补齐或检查这些测试:
After changing this pipeline, at least add or check these tests:
- `go test ./internal/prompt/...`
- `go test ./internal/httpapi/openai/...`
@@ -385,22 +385,21 @@ history split 现在全局强制开启;旧配置中的 `history_split.enabled=
- `go test ./internal/httpapi/gemini/...`
- `go test ./internal/util/...`
如果改的是 tool call 相关兼容语义,还应同时检查:
If you change tool-call related compatibility semantics, also check:
- `go test ./internal/toolcall/...`
- `node --test tests/node/stream-tool-sieve.test.js`
## 14. 文档同步约定
## 14. Documentation sync convention
本文档是这条兼容链路的专项说明。
This document is the dedicated reference for this compatibility pipeline.
如果外部接口行为也变了,还应同步检查:
If external interface behavior also changes, also check:
- [API.md](../API.md)
- [API.md](../API.md)
- [docs/toolcall-semantics.md](./toolcall-semantics.md)
原则是:
The principle is:
- 内部主链路变化,至少更新本文档
- 外部可见契约变化,再同步更新 API 文档
- Internal main-pipeline changes: at least update this document
- Externally visible contract changes: also update the API docs
+39 -39
View File
@@ -1,12 +1,12 @@
# Tool call parsing semantics(Go/Node 统一语义)
# Tool call parsing semantics (unified Go/Node behavior)
本文档描述当前代码中的**实际行为**,以 `internal/toolcall`、`internal/toolstream` 与 `internal/js/helpers/stream-tool-sieve` 为准。
This document describes the **actual behavior** in the current code, with `internal/toolcall`, `internal/toolstream`, and `internal/js/helpers/stream-tool-sieve` as the source of truth.
文档导航:[总览](../README.MD) / [架构说明](./ARCHITECTURE.md) / [测试指南](./TESTING.md)
Docs: [Overview](../README.MD) / [Architecture](./ARCHITECTURE.md) / [Testing guide](./TESTING.md)
## 1) 当前唯一可执行格式
## 1) The only executable format today
当前版本只把下面这类 canonical XML 视为可执行工具调用:
The current version treats only the following canonical XML as an executable tool call:
```xml
<tool_calls>
@@ -16,60 +16,60 @@
</tool_calls>
```
约束:
Constraints:
- 必须有 `<tool_calls>...</tool_calls>` wrapper
- 每个调用必须在 `<invoke name="...">...</invoke>` 内
- 工具名必须放在 `invoke` 的 `name` 属性
- 参数必须使用 `<parameter name="...">...</parameter>`
- A `<tool_calls>...</tool_calls>` wrapper is required
- Each call must be inside `<invoke name="...">...</invoke>`
- The tool name must go in the `name` attribute of `invoke`
- Parameters must use `<parameter name="...">...</parameter>`
兼容修复:
Compatibility fix:
- 如果模型漏掉 opening `<tool_calls>`,但后面仍输出了一个或多个 `<invoke ...>` 并以 `</tool_calls>` 收尾,Go 解析链路会在解析前补回缺失的 opening wrapper。
- 这是一个针对常见模型失误的窄修复,不改变推荐输出格式;prompt 仍要求模型直接输出完整 canonical XML。
- If the model drops the opening `<tool_calls>` but still emits one or more `<invoke ...>` blocks ending with `</tool_calls>`, the Go parsing path restores the missing opening wrapper before parsing.
- This is a narrow fix for a common model mistake; it does not change the recommended output format. The prompt still requires the model to emit complete canonical XML directly.
## 2) 非 canonical 内容
## 2) Non-canonical content
任何不满足上述 canonical XML 形态的内容,都会保留为普通文本,不会执行。一个例外是上一节提到的“缺失 opening `<tool_calls>`、但 closing `</tool_calls>` 仍存在”的窄修复场景。
Any content that does not match the canonical XML shape above is kept as plain text and never executed. The one exception is the narrow "missing opening `<tool_calls>` but closing `</tool_calls>` still present" fix described in the previous section.
当前 parser 不把 allow-list 当作硬安全边界:即使传入了已声明工具名列表,XML 里出现未声明工具名时也会尽量解析并交给上层协议输出;真正的执行侧仍必须自行校验工具名和参数。
The current parser does not treat the allow-list as a hard security boundary: even when a list of declared tool names is passed in, an undeclared tool name appearing in the XML is still parsed best-effort and handed to the upper protocol layer for output; the actual execution side must still validate tool names and parameters itself.
## 3) 流式与防泄漏行为
## 3) Streaming and anti-leak behavior
在流式链路中(Go / Node 一致):
In the streaming path (consistent across Go / Node):
- canonical `<tool_calls>` wrapper 会进入结构化捕获
- 如果流里直接从 `<invoke ...>` 开始,但后面补上了 `</tool_calls>`,Go 流式筛分也会按缺失 opening wrapper 的修复路径尝试恢复
- 已识别成功的工具调用不会再次回流到普通文本
- 不符合新格式的块不会执行,并继续按原样文本透传
- fenced code block 中的 XML 示例始终按普通文本处理
- The canonical `<tool_calls>` wrapper enters structured capture
- If the stream starts directly from `<invoke ...>` but later closes with `</tool_calls>`, the Go streaming sieve also attempts recovery via the missing-opening-wrapper fix path
- Successfully recognized tool calls do not flow back into plain text
- Blocks that do not match the new format are not executed and continue to pass through as-is
- XML examples inside fenced code blocks are always treated as plain text
## 4) 输出结构
## 4) Output structure
`ParseToolCallsDetailed` / `parseToolCallsDetailed` 返回:
`ParseToolCallsDetailed` / `parseToolCallsDetailed` return:
- `calls`:解析出的工具调用列表(`name` + `input`)
- `sawToolCallSyntax`:检测到 canonical wrapper,或命中“缺失 opening wrapper 但可修复”的形态时会为 `true`
- `rejectedByPolicy`:当前固定为 `false`
- `rejectedToolNames`:当前固定为空数组
- `calls`: the parsed list of tool calls (`name` + `input`)
- `sawToolCallSyntax`: `true` when a canonical wrapper is detected, or when the recoverable "missing opening wrapper" shape is hit
- `rejectedByPolicy`: currently always `false`
- `rejectedToolNames`: currently always an empty array
## 5) 落地建议
## 5) Implementation guidance
1. Prompt 里只示范 canonical XML 语法。
2. 上游客户端仍应直接输出 canonical XML;DS2API 只对“closing tag 在、opening tag 漏掉”的常见失误做窄修复,不会泛化接受其他旧格式。
3. 不要依赖 parser 做安全控制;执行器侧仍应做工具名和参数校验。
1. The prompt only demonstrates the canonical XML syntax.
2. Upstream clients should still emit canonical XML directly; DS2API only applies the narrow fix for the common "closing tag present, opening tag missing" mistake and does not generally accept other legacy formats.
3. Do not rely on the parser for security control; the executor side must still validate tool names and parameters.
## 6) 回归验证
## 6) Regression testing
可直接运行:
Run directly:
```bash
go test -v -run 'TestParseToolCalls|TestProcessToolSieve' ./internal/toolcall ./internal/toolstream ./internal/httpapi/openai/...
node --test tests/node/stream-tool-sieve.test.js
```
重点覆盖:
Key coverage:
- canonical `<tool_calls>` wrapper 正常解析
- 非 canonical 内容按普通文本透传
- 代码块示例不执行
- The canonical `<tool_calls>` wrapper parses correctly
- Non-canonical content passes through as plain text
- Code-block examples are not executed