Files
ds2api/docs/prompt-compatibility.md
omar 0d3bcd4501 docs: translate developer reference docs to English
Translate TESTING, toolcall-semantics and prompt-compatibility to English.
Keep DeepSeek role-marker tokens verbatim and fix the SSE note link to its
renamed ASCII path.
2026-06-02 22:20:10 +03:00

18 KiB
Raw Permalink Blame History

API -> web-chat pure-text compatibility pipeline

Docs: Overview / Architecture / API reference / Testing guide

This document is the dedicated reference for how DS2API "compatibilizes OpenAI / Claude / Gemini style API requests into DeepSeek web-chat pure-text context". It is one of the project's most important compatibility products. Any change to message normalization, tool prompt injection, tool history retention, file references, history split, or downstream completion payload assembly must update this document in sync.

1. Core conclusion

DS2API's current core idea is not to forward the client's messages, tools, and attachments to the downstream as-is.

Instead, it compresses those high-level API semantics into three kinds of input that DeepSeek web chat understands more easily:

  1. prompt A single string containing role markers, system instructions, history messages, assistant reasoning tags, historical tool-call XML, etc.
  2. ref_file_ids An array of file references carrying attachments, inline uploaded files, and, when necessary, history files that were split out.
  3. Control flags For example thinking_enabled, search_enabled, and some passthrough parameters.

In other words, the project's most important compatibility action is translating a "structured API conversation" into a "web-chat pure-text context + file references".

2. Why this is the core product

Because for the downstream, the truly stable input surface is not the native schema of OpenAI/Claude/Gemini, but:

  • A continuous conversation prompt
  • A set of referenceable files
  • A few flags

This is also why a lot of code that looks like "protocol compatibility" eventually converges to the same kind of logic:

  • First normalize the messages of different protocols into an internal message sequence
  • Then rewrite tool declarations into system prompt text
  • Then rewrite historical tool calls / tool results into prompt-visible content
  • Finally emit a DeepSeek completion payload

3. Unified mental model

The current main pipeline can be understood like this:

client request
  -> HTTP API surface (OpenAI / Claude / Gemini)
  -> promptcompat unified message normalization
  -> tool prompt injection
  -> DeepSeek-style prompt assembly
  -> file collection / inline upload / history split (OpenAI path)
  -> completion payload
  -> downstream web-chat endpoint

The final normalized prompt is uniformly limited to 131072 Unicode characters; all three paths (OpenAI, Claude, Gemini) return a context too long error before the upstream call.

Key code entrypoints:

4. What the downstream actually receives

After normalization completes, the core shape of the downstream completion payload is:

{
  "chat_session_id": "session-id",
  "model_type": "default",
  "parent_message_id": null,
  "prompt": "<|begin▁of▁sentence|>...",
  "ref_file_ids": [
    "file-history",
    "file-systemprompt",
    "file-other-attachment"
  ],
  "thinking_enabled": true,
  "search_enabled": false
}

Key points:

  • prompt is the main carrier of the conversation context.
  • ref_file_ids carries only file references, not ordinary text messages.
  • tools are not sent downstream as a "native tool schema"; they are rewritten into the prompt.
  • OpenAI Chat / Responses natively go through the unified OpenAI normalization and DeepSeek payload assembly; Claude / Gemini reuse the OpenAI prompt/tool semantics as much as possible, where Gemini directly reuses promptcompat.BuildOpenAIPromptForAdapter, and the Claude messages endpoint, in proxyable scenarios, is converted to the OpenAI chat shape before execution.
  • The thinking / reasoning flags passed by the client are normalized to the downstream thinking_enabled. Gemini generationConfig.thinkingConfig.thinkingBudget is translated to the same thinking flag; when disabled, even if the upstream returns response/thinking_content, the compatibility layer does not treat it as visible body output. On the Claude surface, for streaming requests without an explicit thinking declaration, Anthropic semantics still default it off; but in the non-streaming proxy scenario the compatibility layer internally enables downstream thinking once to capture the case of "empty body, tool call landing inside thinking", then strips the user-invisible thinking block before responding.
  • For the non-streaming finalize of OpenAI Chat / Responses, if the final visible body is empty, the compatibility layer first tries to parse a standalone <tool_calls>...</tool_calls> structure in the chain of thought as a real tool call. The streaming path also performs the same fallback detection at finalize, but does not intercept or rewrite streaming output mid-flight because of chain-of-thought content; thinking / reasoning deltas are still emitted as-is first, and only at finalize may the final tool-call result be appended. Only when the body is empty and the chain of thought also has no executable tool call does it continue to be handled as an empty-reply error.

5. How the prompt is assembled

5.1 Role markers

The final prompt uses DeepSeek-style role markers:

  • <|begin▁of▁sentence|>
  • <|System|>
  • <|User|>
  • <|Assistant|>
  • <|Tool|>
  • <|end▁of▁instructions|>
  • <|end▁of▁sentence|>
  • <|end▁of▁toolresults|>

Implementation: internal/prompt/messages.go

5.2 Thinking continuity notes

When thinking is enabled, an extra system block is inserted at the very front to remind the model to:

  • Continue the existing session, not start over
  • Treat earlier messages as binding context
  • Not leave the final answer only in the reasoning

This part is not the client's original message; it is a continuity contract that the compatibility layer adds proactively.

5.3 Adjacent same-role messages are merged

In the final MessagesPrepareWithThinking, adjacent messages with the same role are merged into a single block, with a blank line inserted between them.

This means:

  • What you see in the prompt is a "merged role block"
  • Not the per-message arrangement passed by the client as-is

6. Why tools are "text injection", not native passthrough

The current project treats tool capability as "part of the prompt constraints".

Concretely:

  1. Serialize each tool's name, description, and parameter schema into text.
  2. Assemble a large You have access to these tools: block.
  3. Append a unified XML tool-call format constraint.
  4. Merge this whole block into the system prompt.

The positive tool-call example still only demonstrates canonical XML: <tool_calls> → <invoke name="..."> → <parameter name="...">. The prompt additionally emphasizes: if a tool is to be called, the first non-whitespace character of the tool block must be <tool_calls>; it must not emit only </tool_calls> and drop the opening tag. The tool names in the positive examples only come from tools actually declared in the current request; if the current request lacks enough known tool shapes, the corresponding single-tool, multi-tool, or nested examples are omitted to avoid writing unavailable tool names into the prompt. For execution-type tools, the script content must go into the execution argument itself: Bash / execute_command use command, exec_command uses cmd; do not demonstrate the script as a path / content file-write argument.

OpenAI path implementation: internal/promptcompat/tool_prompt.go

Claude path implementation: internal/httpapi/claude/handler_utils.go

Unified tool-call format template: internal/toolcall/tool_prompt.go

This is also the key design of the project's "web-chat pure-text compatibility":

  • For the downstream, tools are essentially in-prompt rules
  • Not a native tool schema transport

7. How assistant tool_calls / reasoning are retained

7.1 How reasoning is retained

The assistant's reasoning becomes an explicit tagged block:

[reasoning_content]
...
[/reasoning_content]

followed by the visible answer body.

7.2 How historical tool_calls are retained

The assistant's historical tool_calls are not retained as OpenAI native JSON; they are converted to prompt-visible XML:

<tool_calls>
  <invoke name="read_file">
    <parameter name="path"><![CDATA[src/main.go]]></parameter>
  </invoke>
</tool_calls>

This is also the only supported canonical tool-calling shape in the current project; all other shapes are kept as plain text and are not treated as executable call syntax. The exception is that the parser fixes a very narrow model mistake: if the assistant emits <invoke ...> ... </tool_calls> but drops the leading opening <tool_calls>, the parsing stage restores the wrapper before attempting recognition.

This matters because it determines that:

  • Historical tool calls are "visible text history" in the prompt
  • Not "hidden structured metadata"

Implementation: internal/prompt/tool_calls.go

7.3 How tool results are retained

Results of the tool / function role enter the prompt as <|Tool|>...<|end▁of▁toolresults|>.

If the tool content is empty, it is currently filled with the string "null" to avoid losing the entire tool turn.

8. Actual semantics of files, attachments, and systemprompt files

Here we must clearly distinguish two things:

  1. Text-type system prompt For example OpenAI developer / system / Responses instructions / Claude top-level system These go into the prompt.
  2. File-type systemprompt For example files uploaded via attachment, input_file, base64, or data URL These are not inlined directly into the prompt; they go into ref_file_ids.

OpenAI file-related implementation:

Conclusion:

  • "systemprompt text" is in the prompt
  • "systemprompt files" are usually only in ref_file_ids

Unless the caller expands the file content themselves and stuffs it into the system/developer text, the file content does not automatically appear in the prompt body.

9. Why multi-turn history is not always fully inlined in the prompt

History split is now globally forced on; history_split.enabled=false in old configs is ignored. By default it may trigger from the 2nd user turn, and the trigger threshold can still be adjusted via history_split.trigger_after_turns.

Related implementation:

Behavior after triggering:

  1. Old history messages are cut out.
  2. The old history is re-serialized into a text file.
  3. The actually uploaded file name is fixed as HISTORY.txt.
  4. Inside the file content, the wrapper name IGNORE is used to close DeepSeek's native file markers.
  5. After the file is uploaded, its file_id is placed first in ref_file_ids.
  6. The live prompt keeps only:
    • system / developer
    • context from the latest user turn onward

The history file content is not ordinary free text; it is a transcript re-serialized with the same set of role markers:

[uploaded filename]: HISTORY.txt
[file content end]

<|begin▁of▁sentence|><|User|>...<|Assistant|>...<|Tool|>...

[file name]: IGNORE
[file content begin]

So the "full context" in the current implementation is usually split across two places:

  • The live context in prompt
  • The history transcript file pointed to by ref_file_ids

10. Differences between protocol entrypoints

10.1 OpenAI Chat / Responses

Characteristics:

  • developer maps to system
  • Responses instructions is prepended as a system message
  • tools are injected into the system prompt
  • attachments / input_file / inline files go into ref_file_ids
  • History split mainly takes effect in this path

10.2 Claude Messages

Characteristics:

  • The top-level system takes priority as the system prompt
  • tool_use / tool_result are converted into the unified assistant/tool history semantics
  • tools are likewise merged into the system prompt
  • Regular execution goes through internal/httpapi/claude/handler_messages.go to the OpenAI chat path, with the model alias resolved to a DeepSeek native model first
  • The current code does not have a full ref_file_ids attachment path like OpenAI does

10.3 Gemini

Characteristics:

  • systemInstruction, contents.parts, functionCall, functionResponse are normalized first
  • tools are converted to OpenAI-style function schemas
  • prompt construction reuses OpenAI's promptcompat.BuildOpenAIPromptForAdapter
  • Unrecognized non-text parts are safely serialized into the prompt, with binary/suspected-base64 content omitted or truncated

In other words, Gemini stays as close to OpenAI as possible at the level of "final prompt semantics".

11. A realistic example of the final context

Suppose the user sends a multi-turn request:

  • with system/developer text
  • with tools
  • with a file-type systemprompt attachment
  • with historical assistant tool call / tool result
  • history split already triggered

Then the final context looks closer to:

{
  "prompt": "<|begin▁of▁sentence|><|System|>continuity instructions...\\n\\noriginal system / developer\\n\\nYou have access to these tools: ...<|end▁of▁instructions|><|User|>latest question<|Assistant|>",
  "ref_file_ids": [
    "file-history-ignore",
    "file-systemprompt",
    "file-other-attachment"
  ],
  "thinking_enabled": true,
  "search_enabled": false
}

This is exactly the core result of "API to web-chat pure text":

  • Most structured semantics are compressed into the prompt
  • Files stay files
  • History is split into a file when necessary

12. Scenarios that must update this document

Whenever you touch any of the following behaviors, you must update this document in the same commit or PR:

  • Role mapping changes
  • system / developer / instructions merge rule changes
  • assistant reasoning retention format changes
  • changes to the XML presentation of assistant historical tool_calls
  • tool result injection method changes
  • tool prompt template or tool_choice constraint changes
  • inline file upload / file reference collection rule changes
  • history split trigger conditions, upload format, or IGNORE wrapper format changes
  • completion payload field semantics changes
  • changes to how Claude / Gemini reuse this unified semantics

Check these files first:

  • internal/promptcompat/request_normalize.go
  • internal/promptcompat/prompt_build.go
  • internal/promptcompat/message_normalize.go
  • internal/promptcompat/tool_prompt.go
  • internal/httpapi/openai/files/file_inline_upload.go
  • internal/promptcompat/file_refs.go
  • internal/httpapi/openai/history/history_split.go
  • internal/promptcompat/responses_input_normalize.go
  • internal/httpapi/claude/standard_request.go
  • internal/httpapi/claude/handler_utils.go
  • internal/httpapi/gemini/convert_request.go
  • internal/httpapi/gemini/convert_messages.go
  • internal/httpapi/gemini/convert_tools.go
  • internal/prompt/messages.go
  • internal/prompt/tool_calls.go
  • internal/promptcompat/standard_request.go

13. Suggested minimal verification

After changing this pipeline, at least add or check these tests:

  • go test ./internal/prompt/...
  • go test ./internal/httpapi/openai/...
  • go test ./internal/httpapi/claude/...
  • go test ./internal/httpapi/gemini/...
  • go test ./internal/util/...

If you change tool-call related compatibility semantics, also check:

  • go test ./internal/toolcall/...
  • node --test tests/node/stream-tool-sieve.test.js

14. Documentation sync convention

This document is the dedicated reference for this compatibility pipeline.

If external interface behavior also changes, also check:

The principle is:

  • Internal main-pipeline changes: at least update this document
  • Externally visible contract changes: also update the API docs