Translate TESTING, toolcall-semantics and prompt-compatibility to English. Keep DeepSeek role-marker tokens verbatim and fix the SSE note link to its renamed ASCII path.
317 lines
9.9 KiB
Markdown
317 lines
9.9 KiB
Markdown
# DS2API Testing Guide
|
|
|
|
Docs: [Overview](../README.MD) / [Architecture](./ARCHITECTURE.md) / [Deployment guide](./DEPLOY.md) / [API reference](../API.md)
|
|
|
|
## Overview
|
|
|
|
DS2API provides two tiers of tests:
|
|
|
|
| Tier | Command | Notes |
|
|
| --- | --- | --- |
|
|
| Unit tests (Go) | `./tests/scripts/run-unit-go.sh` | No real account required |
|
|
| Unit tests (Node) | `./tests/scripts/run-unit-node.sh` | No real account required |
|
|
| Unit tests (all) | `./tests/scripts/run-unit-all.sh` | No real account required |
|
|
| End-to-end tests | `./tests/scripts/run-live.sh` | Full-chain test using a real account |
|
|
|
|
The end-to-end suite records full request/response logs for troubleshooting.
|
|
The Node unit test script first runs a `node --check` syntax gate, then executes the test files serially with `--test-concurrency=1` to reduce interference from module-level shared state.
|
|
|
|
---
|
|
|
|
## PR Gates
|
|
|
|
Before opening or updating a PR, run the local equivalent of the gates in `.github/workflows/quality-gates.yml`:
|
|
|
|
```bash
|
|
./scripts/lint.sh
|
|
./tests/scripts/check-refactor-line-gate.sh
|
|
./tests/scripts/run-unit-all.sh
|
|
npm run build --prefix webui
|
|
```
|
|
|
|
Notes:
|
|
|
|
- `./scripts/lint.sh` runs the Go format check and `golangci-lint`; after editing Go files it is still recommended to run `gofmt -w <files>` first.
|
|
- `run-unit-all.sh` invokes the Go and Node unit test entrypoints serially.
|
|
- `run-live.sh` is the real-account end-to-end test, suitable as supplementary verification after a release or high-risk change; it is not part of the fixed per-PR local gate.
|
|
|
|
---
|
|
|
|
## Quick Start
|
|
|
|
### Unit Tests
|
|
|
|
```bash
|
|
./tests/scripts/run-unit-all.sh
|
|
```
|
|
|
|
```bash
|
|
# Or run per language
|
|
./tests/scripts/run-unit-go.sh
|
|
./tests/scripts/run-unit-node.sh
|
|
```
|
|
|
|
```bash
|
|
# Structure and flow gates
|
|
./tests/scripts/check-refactor-line-gate.sh
|
|
./tests/scripts/check-node-split-syntax.sh
|
|
|
|
# Legacy stage gate: stage 6 manual smoke sign-off check (defaults to plans/stage6-manual-smoke.md)
|
|
./tests/scripts/check-stage6-manual-smoke.sh
|
|
```
|
|
|
|
### End-to-End Tests
|
|
|
|
```bash
|
|
./tests/scripts/run-live.sh
|
|
```
|
|
|
|
**Default behavior**:
|
|
|
|
1. **Preflight checks**:
|
|
- `go test ./... -count=1` (unit tests)
|
|
- `./tests/scripts/check-node-split-syntax.sh` (Node split-module syntax gate)
|
|
- `node --test tests/node/stream-tool-sieve.test.js tests/node/chat-stream.test.js tests/node/js_compat_test.js`
|
|
- `npm run build --prefix webui` (WebUI build check)
|
|
|
|
2. **Isolated startup**: copy `config.json` to a temporary directory and start a standalone service process
|
|
|
|
3. **Scenario tests**:
|
|
- ✅ OpenAI non-streaming / streaming
|
|
- ✅ Claude non-streaming / streaming
|
|
- ✅ Admin API (login / config / account management)
|
|
- ✅ Tool Calling
|
|
- ✅ Concurrency stress test
|
|
- ✅ Search models
|
|
|
|
4. **Result collection**: continue running all cases (no early stop), then write the final summary
|
|
|
|
If you only want to skip these preflight checks, run `go run ./cmd/ds2api-tests --no-preflight` directly.
|
|
|
|
---
|
|
|
|
## CLI Flags
|
|
|
|
```bash
|
|
go run ./cmd/ds2api-tests \
|
|
--config config.json \
|
|
--admin-key admin \
|
|
--out artifacts/testsuite \
|
|
--port 0 \
|
|
--timeout 120 \
|
|
--retries 2 \
|
|
--no-preflight=false \
|
|
--keep 5
|
|
```
|
|
|
|
| Flag | Description | Default |
|
|
| --- | --- | --- |
|
|
| `--config` | Config file path | `config.json` |
|
|
| `--admin-key` | Admin key | `DS2API_ADMIN_KEY` env var, fallback `admin` |
|
|
| `--out` | Artifact output root directory | `artifacts/testsuite` |
|
|
| `--port` | Test service port (`0` = auto-assign a free port) | `0` |
|
|
| `--timeout` | Per-request timeout in seconds | `120` |
|
|
| `--retries` | Retry count for network/5xx requests | `2` |
|
|
| `--no-preflight` | Skip preflight checks | `false` |
|
|
| `--keep` | Keep the most recent N test runs (`0` = keep all) | `5` |
|
|
|
|
---
|
|
|
|
## Auto Cleanup
|
|
|
|
After each test run completes, the program automatically scans the output directory (`--out`), keeps the most recent `--keep` runs sorted by time, and deletes the rest.
|
|
|
|
- Keeps **5** runs by default
|
|
- Set `--keep 0` to disable auto cleanup
|
|
- Deleted old run directories are logged
|
|
|
|
---
|
|
|
|
## Artifact Layout
|
|
|
|
Each run creates a directory named after the run ID:
|
|
|
|
```text
|
|
artifacts/testsuite/<run_id>/
|
|
├── summary.json # machine-readable report
|
|
├── summary.md # human-readable report
|
|
├── server.log # server logs during the test
|
|
├── preflight.log # preflight command output
|
|
└── cases/
|
|
└── <case_id>/
|
|
├── request.json # request body
|
|
├── response.headers # response headers
|
|
├── response.body # response body
|
|
├── stream.raw # raw SSE data (streaming cases)
|
|
├── assertions.json # assertion results
|
|
└── meta.json # metadata (elapsed time, status code, etc.)
|
|
```
|
|
|
|
---
|
|
|
|
## Trace Binding
|
|
|
|
Each test request automatically injects trace information for quick troubleshooting:
|
|
|
|
| Location | Format |
|
|
| --- | --- |
|
|
| Request header | `X-Ds2-Test-Trace: <trace_id>` |
|
|
| Query parameter | `__trace_id=<trace_id>` |
|
|
|
|
When a case fails, `summary.md` includes the trace ID. You can quickly search the corresponding server logs:
|
|
|
|
```bash
|
|
rg "<trace_id>" artifacts/testsuite/<run_id>/server.log
|
|
```
|
|
|
|
---
|
|
|
|
## Exit Code
|
|
|
|
| Exit code | Meaning |
|
|
| --- | --- |
|
|
| `0` | All cases passed ✅ |
|
|
| `1` | Some cases failed ❌ |
|
|
|
|
The suite can be used as a local release gate (CI/CD integration).
|
|
|
|
---
|
|
|
|
## Sensitive Data Warning
|
|
|
|
⚠️ The suite stores **full raw request/response payloads** for debugging.
|
|
|
|
- **Do not** upload the artifacts directory to a public repository
|
|
- **Do not** share un-redacted artifact files in an issue tracker
|
|
- If you need to share logs, manually clear sensitive information (tokens, passwords, etc.) first
|
|
|
|
---
|
|
|
|
## Common Usage
|
|
|
|
### Run unit tests only
|
|
|
|
```bash
|
|
go test ./...
|
|
```
|
|
|
|
### Run unit tests for a specific module
|
|
|
|
```bash
|
|
# Run tool-call tests (recommended for debugging tool-call parsing issues)
|
|
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/
|
|
|
|
# Run a single test case
|
|
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/
|
|
|
|
# Run format tests
|
|
go test -v ./internal/format/...
|
|
|
|
# Run HTTP API tests
|
|
go test -v ./internal/httpapi/openai/...
|
|
```
|
|
|
|
### Debugging Tool Call Issues
|
|
|
|
When you hit a DeepSeek tool-call parsing issue, you can use the following approaches:
|
|
|
|
```bash
|
|
# 1. Run all tool-call tests
|
|
go test -v -run 'TestParseToolCalls|TestRepair' ./internal/toolcall/
|
|
|
|
# 2. Inspect the detailed debug output in the test logs
|
|
go test -v -run TestParseToolCallsWithDeepSeekHallucination ./internal/toolcall/ 2>&1
|
|
|
|
# 3. Check the fix behavior for specific test cases
|
|
# Test cases live in internal/toolcall/toolcalls_test.go, including:
|
|
# - TestParseToolCallsWithDeepSeekHallucination: typical DeepSeek hallucinated output
|
|
# - TestRepairLooseJSONWithNestedObjects: bracket repair for nested objects
|
|
# - TestParseToolCallsWithMixedWindowsPaths: Windows path handling
|
|
```
|
|
|
|
### Run Node.js tests
|
|
|
|
```bash
|
|
# Run Node tests
|
|
node --test tests/node/stream-tool-sieve.test.js
|
|
|
|
# Or use the script
|
|
./tests/scripts/run-unit-node.sh
|
|
```
|
|
|
|
### Run end-to-end tests (skip preflight)
|
|
|
|
```bash
|
|
go run ./cmd/ds2api-tests --no-preflight
|
|
```
|
|
|
|
### Run the raw-stream simulation (standalone tool)
|
|
|
|
```bash
|
|
./tests/scripts/run-raw-stream-sim.sh
|
|
```
|
|
|
|
Notes:
|
|
- By default this tool replays the canonical samples declared in `tests/raw_stream_samples/manifest.json`, doing a 1:1 simulation parse following the upstream SSE order.
|
|
- By default it verifies there is no `FINISHED` text leak and requires an end signal to be present.
|
|
- By default it does **not** strictly compare `raw accumulated_token_usage` against locally parsed tokens (the current implementation estimates from content); add `--fail-on-token-mismatch` explicitly to enforce a strict check.
|
|
- Every run writes the locally derived results to `artifacts/raw-stream-sim/<run-id>/<sample-id>/replay.output.txt` and emits a structured report.
|
|
- If you have a historical baseline directory, use `--baseline-root` to have the tool do a direct text comparison.
|
|
- For a more complete protocol-level behavior description, see [deepseek-sse-behavior-2026-04-05.md](./deepseek-sse-behavior-2026-04-05.md).
|
|
|
|
### Replay-compare a single sample
|
|
|
|
```bash
|
|
./tests/scripts/compare-raw-stream-sample.sh markdown-format-example-20260405-spacefix
|
|
```
|
|
|
|
Notes:
|
|
- This script reads `upstream.stream.sse` from the raw-only sample directory.
|
|
- Replay results are written to `artifacts/raw-stream-sim/<run-id>/<sample-id>/` for easy inspection.
|
|
- If a historical baseline directory is provided, the script automatically compares the current replay output against the baseline text.
|
|
|
|
### Capture a permanent sample
|
|
|
|
After starting the service locally, you can call:
|
|
|
|
```bash
|
|
POST /admin/dev/raw-samples/capture
|
|
```
|
|
|
|
This endpoint writes the request metadata and the upstream raw stream to `tests/raw_stream_samples/<sample-id>/`, which can later be reused for replay and field analysis. Derived output is regenerated during local replay and is no longer stored in the sample directory.
|
|
|
|
### Query in-memory captures and save samples
|
|
|
|
If the issue was just reproduced locally, you can first query the captures in the current process memory, then optionally persist them:
|
|
|
|
```bash
|
|
GET /admin/dev/raw-samples/query?q=guangzhou&limit=10
|
|
POST /admin/dev/raw-samples/save
|
|
{"chain_key":"session:xxxx","sample_id":"tmp-from-memory"}
|
|
```
|
|
|
|
Notes:
|
|
- `query` merges `completion + continue` into a single chain by `chat_session_id`, which is handy for locating continued-thinking issues.
|
|
- `save` supports selecting the target via `query`, `chain_key`, or `capture_id`.
|
|
- The generated sample directory is still `tests/raw_stream_samples/<sample-id>/` and can be fed directly to the replay script.
|
|
|
|
### Specify output directory and timeout
|
|
|
|
```bash
|
|
go run ./cmd/ds2api-tests \
|
|
--out /tmp/ds2api-test \
|
|
--timeout 60
|
|
```
|
|
|
|
### Use in CI
|
|
|
|
```bash
|
|
# Make sure config.json exists and contains a valid test account
|
|
./tests/scripts/run-live.sh
|
|
exit_code=$?
|
|
if [ $exit_code -ne 0 ]; then
|
|
echo "Tests failed! Check artifacts for details."
|
|
exit 1
|
|
fi
|
|
```
|