Translate TESTING, toolcall-semantics and prompt-compatibility to English. Keep DeepSeek role-marker tokens verbatim and fix the SSE note link to its renamed ASCII path.
18 KiB
API -> web-chat pure-text compatibility pipeline
Docs: Overview / Architecture / API reference / Testing guide
This document is the dedicated reference for how DS2API "compatibilizes OpenAI / Claude / Gemini style API requests into DeepSeek web-chat pure-text context". It is one of the project's most important compatibility products. Any change to message normalization, tool prompt injection, tool history retention, file references, history split, or downstream completion payload assembly must update this document in sync.
1. Core conclusion
DS2API's current core idea is not to forward the client's messages, tools, and attachments to the downstream as-is.
Instead, it compresses those high-level API semantics into three kinds of input that DeepSeek web chat understands more easily:
promptA single string containing role markers, system instructions, history messages, assistant reasoning tags, historical tool-call XML, etc.ref_file_idsAn array of file references carrying attachments, inline uploaded files, and, when necessary, history files that were split out.- Control flags
For example
thinking_enabled,search_enabled, and some passthrough parameters.
In other words, the project's most important compatibility action is translating a "structured API conversation" into a "web-chat pure-text context + file references".
2. Why this is the core product
Because for the downstream, the truly stable input surface is not the native schema of OpenAI/Claude/Gemini, but:
- A continuous conversation prompt
- A set of referenceable files
- A few flags
This is also why a lot of code that looks like "protocol compatibility" eventually converges to the same kind of logic:
- First normalize the messages of different protocols into an internal message sequence
- Then rewrite tool declarations into system prompt text
- Then rewrite historical tool calls / tool results into prompt-visible content
- Finally emit a DeepSeek completion payload
3. Unified mental model
The current main pipeline can be understood like this:
client request
-> HTTP API surface (OpenAI / Claude / Gemini)
-> promptcompat unified message normalization
-> tool prompt injection
-> DeepSeek-style prompt assembly
-> file collection / inline upload / history split (OpenAI path)
-> completion payload
-> downstream web-chat endpoint
The final normalized prompt is uniformly limited to 131072 Unicode characters; all three paths (OpenAI, Claude, Gemini) return a context too long error before the upstream call.
Key code entrypoints:
- OpenAI Chat / Responses: internal/promptcompat/request_normalize.go
- OpenAI prompt assembly: internal/promptcompat/prompt_build.go
- OpenAI message normalization: internal/promptcompat/message_normalize.go
- Claude normalization: internal/httpapi/claude/standard_request.go
- Claude message and tool_use/tool_result normalization: internal/httpapi/claude/handler_utils.go
- Gemini reuses the OpenAI prompt builder: internal/httpapi/gemini/convert_request.go
- DeepSeek prompt role-marker assembly: internal/prompt/messages.go
- prompt-visible tool history XML: internal/prompt/tool_calls.go
- completion payload: internal/promptcompat/standard_request.go
4. What the downstream actually receives
After normalization completes, the core shape of the downstream completion payload is:
{
"chat_session_id": "session-id",
"model_type": "default",
"parent_message_id": null,
"prompt": "<|begin▁of▁sentence|>...",
"ref_file_ids": [
"file-history",
"file-systemprompt",
"file-other-attachment"
],
"thinking_enabled": true,
"search_enabled": false
}
Key points:
promptis the main carrier of the conversation context.ref_file_idscarries only file references, not ordinary text messages.toolsare not sent downstream as a "native tool schema"; they are rewritten into theprompt.- OpenAI Chat / Responses natively go through the unified OpenAI normalization and DeepSeek payload assembly; Claude / Gemini reuse the OpenAI prompt/tool semantics as much as possible, where Gemini directly reuses
promptcompat.BuildOpenAIPromptForAdapter, and the Claude messages endpoint, in proxyable scenarios, is converted to the OpenAI chat shape before execution. - The thinking / reasoning flags passed by the client are normalized to the downstream
thinking_enabled. GeminigenerationConfig.thinkingConfig.thinkingBudgetis translated to the same thinking flag; when disabled, even if the upstream returnsresponse/thinking_content, the compatibility layer does not treat it as visible body output. On the Claude surface, for streaming requests without an explicitthinkingdeclaration, Anthropic semantics still default it off; but in the non-streaming proxy scenario the compatibility layer internally enables downstream thinking once to capture the case of "empty body, tool call landing inside thinking", then strips the user-invisible thinking block before responding. - For the non-streaming finalize of OpenAI Chat / Responses, if the final visible body is empty, the compatibility layer first tries to parse a standalone
<tool_calls>...</tool_calls>structure in the chain of thought as a real tool call. The streaming path also performs the same fallback detection at finalize, but does not intercept or rewrite streaming output mid-flight because of chain-of-thought content; thinking / reasoning deltas are still emitted as-is first, and only at finalize may the final tool-call result be appended. Only when the body is empty and the chain of thought also has no executable tool call does it continue to be handled as an empty-reply error.
5. How the prompt is assembled
5.1 Role markers
The final prompt uses DeepSeek-style role markers:
<|begin▁of▁sentence|><|System|><|User|><|Assistant|><|Tool|><|end▁of▁instructions|><|end▁of▁sentence|><|end▁of▁toolresults|>
Implementation: internal/prompt/messages.go
5.2 Thinking continuity notes
When thinking is enabled, an extra system block is inserted at the very front to remind the model to:
- Continue the existing session, not start over
- Treat earlier messages as binding context
- Not leave the final answer only in the reasoning
This part is not the client's original message; it is a continuity contract that the compatibility layer adds proactively.
5.3 Adjacent same-role messages are merged
In the final MessagesPrepareWithThinking, adjacent messages with the same role are merged into a single block, with a blank line inserted between them.
This means:
- What you see in the prompt is a "merged role block"
- Not the per-message arrangement passed by the client as-is
6. Why tools are "text injection", not native passthrough
The current project treats tool capability as "part of the prompt constraints".
Concretely:
- Serialize each tool's name, description, and parameter schema into text.
- Assemble a large
You have access to these tools:block. - Append a unified XML tool-call format constraint.
- Merge this whole block into the system prompt.
The positive tool-call example still only demonstrates canonical XML: <tool_calls> → <invoke name="..."> → <parameter name="...">.
The prompt additionally emphasizes: if a tool is to be called, the first non-whitespace character of the tool block must be <tool_calls>; it must not emit only </tool_calls> and drop the opening tag.
The tool names in the positive examples only come from tools actually declared in the current request; if the current request lacks enough known tool shapes, the corresponding single-tool, multi-tool, or nested examples are omitted to avoid writing unavailable tool names into the prompt.
For execution-type tools, the script content must go into the execution argument itself: Bash / execute_command use command, exec_command uses cmd; do not demonstrate the script as a path / content file-write argument.
OpenAI path implementation: internal/promptcompat/tool_prompt.go
Claude path implementation: internal/httpapi/claude/handler_utils.go
Unified tool-call format template: internal/toolcall/tool_prompt.go
This is also the key design of the project's "web-chat pure-text compatibility":
- For the downstream, tools are essentially in-prompt rules
- Not a native tool schema transport
7. How assistant tool_calls / reasoning are retained
7.1 How reasoning is retained
The assistant's reasoning becomes an explicit tagged block:
[reasoning_content]
...
[/reasoning_content]
followed by the visible answer body.
7.2 How historical tool_calls are retained
The assistant's historical tool_calls are not retained as OpenAI native JSON; they are converted to prompt-visible XML:
<tool_calls>
<invoke name="read_file">
<parameter name="path"><![CDATA[src/main.go]]></parameter>
</invoke>
</tool_calls>
This is also the only supported canonical tool-calling shape in the current project; all other shapes are kept as plain text and are not treated as executable call syntax.
The exception is that the parser fixes a very narrow model mistake: if the assistant emits <invoke ...> ... </tool_calls> but drops the leading opening <tool_calls>, the parsing stage restores the wrapper before attempting recognition.
This matters because it determines that:
- Historical tool calls are "visible text history" in the prompt
- Not "hidden structured metadata"
Implementation: internal/prompt/tool_calls.go
7.3 How tool results are retained
Results of the tool / function role enter the prompt as <|Tool|>...<|end▁of▁toolresults|>.
If the tool content is empty, it is currently filled with the string "null" to avoid losing the entire tool turn.
8. Actual semantics of files, attachments, and systemprompt files
Here we must clearly distinguish two things:
- Text-type system prompt
For example OpenAI
developer/system/ Responsesinstructions/ Claude top-levelsystemThese go into theprompt. - File-type systemprompt
For example files uploaded via attachment,
input_file, base64, or data URL These are not inlined directly into theprompt; they go intoref_file_ids.
OpenAI file-related implementation:
- inline/base64/data URL upload: internal/httpapi/openai/files/file_inline_upload.go
- When
runtime.disable_upstream_file_uploads=true, upstream file uploads for explicit/v1/files, inline/base64 upload, and history split are all turned off; existingfile_id/ref_file_idsare still collected as ordinary references. - File ID collection: internal/promptcompat/file_refs.go
Conclusion:
- "systemprompt text" is in the prompt
- "systemprompt files" are usually only in
ref_file_ids
Unless the caller expands the file content themselves and stuffs it into the system/developer text, the file content does not automatically appear in the prompt body.
9. Why multi-turn history is not always fully inlined in the prompt
History split is now globally forced on; history_split.enabled=false in old configs is ignored. By default it may trigger from the 2nd user turn, and the trigger threshold can still be adjusted via history_split.trigger_after_turns.
Related implementation:
- Config accessors: internal/config/store_accessors.go
- History split: internal/httpapi/openai/history/history_split.go
Behavior after triggering:
- Old history messages are cut out.
- The old history is re-serialized into a text file.
- The actually uploaded file name is fixed as
HISTORY.txt. - Inside the file content, the wrapper name
IGNOREis used to close DeepSeek's native file markers. - After the file is uploaded, its
file_idis placed first inref_file_ids. - The live prompt keeps only:
- system / developer
- context from the latest user turn onward
The history file content is not ordinary free text; it is a transcript re-serialized with the same set of role markers:
[uploaded filename]: HISTORY.txt
[file content end]
<|begin▁of▁sentence|><|User|>...<|Assistant|>...<|Tool|>...
[file name]: IGNORE
[file content begin]
So the "full context" in the current implementation is usually split across two places:
- The live context in
prompt - The history transcript file pointed to by
ref_file_ids
10. Differences between protocol entrypoints
10.1 OpenAI Chat / Responses
Characteristics:
developermaps tosystem- Responses
instructionsis prepended as a system message toolsare injected into the system promptattachments/input_file/ inline files go intoref_file_ids- History split mainly takes effect in this path
10.2 Claude Messages
Characteristics:
- The top-level
systemtakes priority as the system prompt tool_use/tool_resultare converted into the unified assistant/tool history semanticstoolsare likewise merged into the system prompt- Regular execution goes through
internal/httpapi/claude/handler_messages.goto the OpenAI chat path, with the model alias resolved to a DeepSeek native model first - The current code does not have a full
ref_file_idsattachment path like OpenAI does
10.3 Gemini
Characteristics:
systemInstruction,contents.parts,functionCall,functionResponseare normalized first- tools are converted to OpenAI-style function schemas
- prompt construction reuses OpenAI's
promptcompat.BuildOpenAIPromptForAdapter - Unrecognized non-text parts are safely serialized into the prompt, with binary/suspected-base64 content omitted or truncated
In other words, Gemini stays as close to OpenAI as possible at the level of "final prompt semantics".
11. A realistic example of the final context
Suppose the user sends a multi-turn request:
- with system/developer text
- with tools
- with a file-type systemprompt attachment
- with historical assistant tool call / tool result
- history split already triggered
Then the final context looks closer to:
{
"prompt": "<|begin▁of▁sentence|><|System|>continuity instructions...\\n\\noriginal system / developer\\n\\nYou have access to these tools: ...<|end▁of▁instructions|><|User|>latest question<|Assistant|>",
"ref_file_ids": [
"file-history-ignore",
"file-systemprompt",
"file-other-attachment"
],
"thinking_enabled": true,
"search_enabled": false
}
This is exactly the core result of "API to web-chat pure text":
- Most structured semantics are compressed into the
prompt - Files stay files
- History is split into a file when necessary
12. Scenarios that must update this document
Whenever you touch any of the following behaviors, you must update this document in the same commit or PR:
- Role mapping changes
- system / developer / instructions merge rule changes
- assistant reasoning retention format changes
- changes to the XML presentation of assistant historical
tool_calls - tool result injection method changes
- tool prompt template or tool_choice constraint changes
- inline file upload / file reference collection rule changes
- history split trigger conditions, upload format, or
IGNOREwrapper format changes - completion payload field semantics changes
- changes to how Claude / Gemini reuse this unified semantics
Check these files first:
internal/promptcompat/request_normalize.gointernal/promptcompat/prompt_build.gointernal/promptcompat/message_normalize.gointernal/promptcompat/tool_prompt.gointernal/httpapi/openai/files/file_inline_upload.gointernal/promptcompat/file_refs.gointernal/httpapi/openai/history/history_split.gointernal/promptcompat/responses_input_normalize.gointernal/httpapi/claude/standard_request.gointernal/httpapi/claude/handler_utils.gointernal/httpapi/gemini/convert_request.gointernal/httpapi/gemini/convert_messages.gointernal/httpapi/gemini/convert_tools.gointernal/prompt/messages.gointernal/prompt/tool_calls.gointernal/promptcompat/standard_request.go
13. Suggested minimal verification
After changing this pipeline, at least add or check these tests:
go test ./internal/prompt/...go test ./internal/httpapi/openai/...go test ./internal/httpapi/claude/...go test ./internal/httpapi/gemini/...go test ./internal/util/...
If you change tool-call related compatibility semantics, also check:
go test ./internal/toolcall/...node --test tests/node/stream-tool-sieve.test.js
14. Documentation sync convention
This document is the dedicated reference for this compatibility pipeline.
If external interface behavior also changes, also check:
The principle is:
- Internal main-pipeline changes: at least update this document
- Externally visible contract changes: also update the API docs