Files
omar 0d3bcd4501 docs: translate developer reference docs to English
Translate TESTING, toolcall-semantics and prompt-compatibility to English.
Keep DeepSeek role-marker tokens verbatim and fix the SSE note link to its
renamed ASCII path.
2026-06-02 22:20:10 +03:00

9.9 KiB

DS2API Testing Guide

Docs: Overview / Architecture / Deployment guide / API reference

Overview

DS2API provides two tiers of tests:

Tier Command Notes
Unit tests (Go) ./tests/scripts/run-unit-go.sh No real account required
Unit tests (Node) ./tests/scripts/run-unit-node.sh No real account required
Unit tests (all) ./tests/scripts/run-unit-all.sh No real account required
End-to-end tests ./tests/scripts/run-live.sh Full-chain test using a real account

The end-to-end suite records full request/response logs for troubleshooting. The Node unit test script first runs a node --check syntax gate, then executes the test files serially with --test-concurrency=1 to reduce interference from module-level shared state.


PR Gates

Before opening or updating a PR, run the local equivalent of the gates in .github/workflows/quality-gates.yml:

./scripts/lint.sh
./tests/scripts/check-refactor-line-gate.sh
./tests/scripts/run-unit-all.sh
npm run build --prefix webui

Notes:

  • ./scripts/lint.sh runs the Go format check and golangci-lint; after editing Go files it is still recommended to run gofmt -w <files> first.
  • run-unit-all.sh invokes the Go and Node unit test entrypoints serially.
  • run-live.sh is the real-account end-to-end test, suitable as supplementary verification after a release or high-risk change; it is not part of the fixed per-PR local gate.

Quick Start

Unit Tests

./tests/scripts/run-unit-all.sh
# Or run per language
./tests/scripts/run-unit-go.sh
./tests/scripts/run-unit-node.sh
# Structure and flow gates
./tests/scripts/check-refactor-line-gate.sh
./tests/scripts/check-node-split-syntax.sh

# Legacy stage gate: stage 6 manual smoke sign-off check (defaults to plans/stage6-manual-smoke.md)
./tests/scripts/check-stage6-manual-smoke.sh

End-to-End Tests

./tests/scripts/run-live.sh

Default behavior:

  1. Preflight checks:

    • go test ./... -count=1 (unit tests)
    • ./tests/scripts/check-node-split-syntax.sh (Node split-module syntax gate)
    • node --test tests/node/stream-tool-sieve.test.js tests/node/chat-stream.test.js tests/node/js_compat_test.js
    • npm run build --prefix webui (WebUI build check)
  2. Isolated startup: copy config.json to a temporary directory and start a standalone service process

  3. Scenario tests:

    • ✅ OpenAI non-streaming / streaming
    • ✅ Claude non-streaming / streaming
    • ✅ Admin API (login / config / account management)
    • ✅ Tool Calling
    • ✅ Concurrency stress test
    • ✅ Search models
  4. Result collection: continue running all cases (no early stop), then write the final summary

If you only want to skip these preflight checks, run go run ./cmd/ds2api-tests --no-preflight directly.


CLI Flags

go run ./cmd/ds2api-tests \
  --config config.json \
  --admin-key admin \
  --out artifacts/testsuite \
  --port 0 \
  --timeout 120 \
  --retries 2 \
  --no-preflight=false \
  --keep 5
Flag Description Default
--config Config file path config.json
--admin-key Admin key DS2API_ADMIN_KEY env var, fallback admin
--out Artifact output root directory artifacts/testsuite
--port Test service port (0 = auto-assign a free port) 0
--timeout Per-request timeout in seconds 120
--retries Retry count for network/5xx requests 2
--no-preflight Skip preflight checks false
--keep Keep the most recent N test runs (0 = keep all) 5

Auto Cleanup

After each test run completes, the program automatically scans the output directory (--out), keeps the most recent --keep runs sorted by time, and deletes the rest.

  • Keeps 5 runs by default
  • Set --keep 0 to disable auto cleanup
  • Deleted old run directories are logged

Artifact Layout

Each run creates a directory named after the run ID:

artifacts/testsuite/<run_id>/
├── summary.json          # machine-readable report
├── summary.md            # human-readable report
├── server.log            # server logs during the test
├── preflight.log         # preflight command output
└── cases/
    └── <case_id>/
        ├── request.json      # request body
        ├── response.headers  # response headers
        ├── response.body     # response body
        ├── stream.raw        # raw SSE data (streaming cases)
        ├── assertions.json   # assertion results
        └── meta.json         # metadata (elapsed time, status code, etc.)

Trace Binding

Each test request automatically injects trace information for quick troubleshooting:

Location Format
Request header X-Ds2-Test-Trace: <trace_id>
Query parameter __trace_id=<trace_id>

When a case fails, summary.md includes the trace ID. You can quickly search the corresponding server logs:

rg "<trace_id>" artifacts/testsuite/<run_id>/server.log

Exit Code

Exit code Meaning
0 All cases passed ✅
1 Some cases failed ❌

The suite can be used as a local release gate (CI/CD integration).


Sensitive Data Warning

⚠️ The suite stores full raw request/response payloads for debugging.

  • Do not upload the artifacts directory to a public repository
  • Do not share un-redacted artifact files in an issue tracker
  • If you need to share logs, manually clear sensitive information (tokens, passwords, etc.) first

Common Usage

Run unit tests only

go test ./...

Run unit tests for a specific module

# Run tool-call tests (recommended for debugging tool-call parsing issues)
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/

# Run a single test case
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/

# Run format tests
go test -v ./internal/format/...

# Run HTTP API tests
go test -v ./internal/httpapi/openai/...

Debugging Tool Call Issues

When you hit a DeepSeek tool-call parsing issue, you can use the following approaches:

# 1. Run all tool-call tests
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/

# 2. Inspect the detailed debug output in the test logs
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/ 2>&1

# 3. Check the fix behavior for specific test cases
# Test cases live in internal/toolcall/toolcalls_test.go, including:
# - TestParseToolCallsWithDeepSeekHallucination: typical DeepSeek hallucinated output
# - TestRepairLooseJSONWithNestedObjects: bracket repair for nested objects
# - TestParseToolCallsWithMixedWindowsPaths: Windows path handling

Run Node.js tests

# Run Node tests
node --test tests/node/stream-tool-sieve.test.js

# Or use the script
./tests/scripts/run-unit-node.sh

Run end-to-end tests (skip preflight)

go run ./cmd/ds2api-tests --no-preflight

Run the raw-stream simulation (standalone tool)

./tests/scripts/run-raw-stream-sim.sh

Notes:

  • By default this tool replays the canonical samples declared in tests/raw_stream_samples/manifest.json, doing a 1:1 simulation parse following the upstream SSE order.
  • By default it verifies there is no FINISHED text leak and requires an end signal to be present.
  • By default it does not strictly compare raw accumulated_token_usage against locally parsed tokens (the current implementation estimates from content); add --fail-on-token-mismatch explicitly to enforce a strict check.
  • Every run writes the locally derived results to artifacts/raw-stream-sim/<run-id>/<sample-id>/replay.output.txt and emits a structured report.
  • If you have a historical baseline directory, use --baseline-root to have the tool do a direct text comparison.
  • For a more complete protocol-level behavior description, see deepseek-sse-behavior-2026-04-05.md.

Replay-compare a single sample

./tests/scripts/compare-raw-stream-sample.sh markdown-format-example-20260405-spacefix

Notes:

  • This script reads upstream.stream.sse from the raw-only sample directory.
  • Replay results are written to artifacts/raw-stream-sim/<run-id>/<sample-id>/ for easy inspection.
  • If a historical baseline directory is provided, the script automatically compares the current replay output against the baseline text.

Capture a permanent sample

After starting the service locally, you can call:

POST /admin/dev/raw-samples/capture

This endpoint writes the request metadata and the upstream raw stream to tests/raw_stream_samples/<sample-id>/, which can later be reused for replay and field analysis. Derived output is regenerated during local replay and is no longer stored in the sample directory.

Query in-memory captures and save samples

If the issue was just reproduced locally, you can first query the captures in the current process memory, then optionally persist them:

GET /admin/dev/raw-samples/query?q=guangzhou&limit=10
POST /admin/dev/raw-samples/save
{"chain_key":"session:xxxx","sample_id":"tmp-from-memory"}

Notes:

  • query merges completion + continue into a single chain by chat_session_id, which is handy for locating continued-thinking issues.
  • save supports selecting the target via query, chain_key, or capture_id.
  • The generated sample directory is still tests/raw_stream_samples/<sample-id>/ and can be fed directly to the replay script.

Specify output directory and timeout

go run ./cmd/ds2api-tests \
  --out /tmp/ds2api-test \
  --timeout 60

Use in CI

# Make sure config.json exists and contains a valid test account
./tests/scripts/run-live.sh
exit_code=$?
if [ $exit_code -ne 0 ]; then
  echo "Tests failed! Check artifacts for details."
  exit 1
fi