Promote v1.14.0-lx.3-rc.2 to a stable (Latest) release. Functionally
identical — no runtime change since the rc, only this changelog entry.
Payload: DNS command-multiplex (rc.1) + AWG re-graft onto wireguard-go
v0.0.5 with upstream merge (rc.2). Device-verified by owner.
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.
lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
(PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
point every L3-forwarded packet transits, incl. established flows that
bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
sing-tun); wireguard-go stays v0.0.5 with local submodule replace.
Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
SPEC.md rewritten to current (multiplex) architecture, no chronology.
HISTORY.md captures v1 standalone class-error, the field bug, rejected paths.
New project rule (README + CONSTITUTION 3.2): SPEC.md = current state first,
chronology/rationale of architecture changes go to HISTORY.md.
DNS stream now runs on the shared c.ctx via dispatchCommands (CommandDNS),
mirroring handleConnectionsStream: auto-reconnects with Connect(), dies with
the client, no per-stream Close()/OnError. Removes DnsQueryHandler,
DnsQuerySubscription and the standalone SubscribeDNSQueries client method.
DnsQuery/DnsAnswer/dnsQueryFromGRPC unchanged.
Add DNS as a first-class multiplexed command, uniform with CommandConnections:
- command.go: CommandDNS constant (next in iota)
- command_client.go: case CommandDNS in dispatchCommands; DNSIncludeAnswers
option field (like StatusInterval); WriteDNSQuery in CommandClientHandler
All three upstream touch-points wrapped in // lx:begin dns / // lx:end dns
(CONSTITUTION 3.3). SPEC 018 v2.
Runtime detour/selector rings crash the core (fatal stack overflow via
unbounded DialContext recursion); static rings are already rejected at start
by lintOutbound. Worked out the full event model (E1-E5) and topologyMu race
linearization, then adversarially verified it (7-agent workflow): deadlock and
false-positive attacks HOLD, but TOCTOU BREAKS — even a correct core guard is
not airtight without also covering the endpoint manager, Manager.Remove,
history side-channels, and the pointer-vs-tag graph divergence after a runtime
Create. Owner decision (2026-07-06): protection lives at the UI level (LxBox
validates before SelectOutbound); core stays a minimal delta to upstream.
No core code changed. SPEC is a design record + Roadmap row (status DEFERRED).
Split the lx go vet step into two passes so every lx-owned package keeps
the full analyzer set; only daemon/ and experimental/libbox/ (upstream
TriggerDebugCrash/TriggerGoPanic) drop the unsafeptr check.
Producer run populated musl-toolchain-cache with 4 arch assets; restore path
validated locally (asset name, gh download, tar layout under naiveproxy/src).
Status -> C, Roadmap updated.
snapshot.debian.org intermittently 503s during the musl sysroot build and
blocks releases (v1.14.0-lx.2-rc.1 failed twice on it). actions/cache also
misses across tag builds (ref-scoping). Add a producer workflow that uploads
the built toolchain to a musl-toolchain-cache release, and a restore step in
lx-release.yml that pulls it on cache-miss before falling back to
snapshot.debian.org. Both workflows are lx-owned; zero upstream diff.
Full audit of the LX delta (10 axes, adversarial verification): 32 findings,
27 confirmed, 24 fixed on branch lx-spec022-audit-fixes, 3 skipped by design
(#12/#17/#18). Records #19 resolution (SPEC 013 test kept — upstream ships none).
Exact-string match on quic-go's formatted CRYPTO_ERROR silently disabled the
Cloudflare Access hint on any error-text reformat. Match the inner
"tls: access denied" alert as a substring instead.
s3 pads only cookie-reply messages (paddings.cookie); s4 pads every transport
data packet (paddings.transport). Folding s3 into the MTU budget dropped MTU and
warned spuriously for an atypical s3>s4 config. Fix calc, comment, warning, docs.
The 'if c.reader == nil' fast path read reader without synchronising against the
RoundTrip goroutine's write (a data race -race flagged). Always receive on
created first; on an already-closed channel that is effectively free.
calculateIPv4Checksum summed a fixed 20 bytes, producing a wrong checksum after
TTL decrement when the header carried options (IHL>5). Take the real IHL*4 span;
validate IHL against the buffer before the read. Adds an IHL=6 test.
Pool() keyed URL-test history by the raw slot tag; history is stored under
RealTag(detour), so a nested-group member always reported Delay=0 in GetPool.
Resolve the slot tag and read under RealTag, matching seedPool/rebuildPool.
questionCache returned a fresh cached response without logging/emitting; only
the stale (optimistic) branch emitted. Mirror the Exchange path: fresh->cached,
stale->optimistic. Gated by HasSubscribers, so zero cost with no profiler.
SuspendAmneziaWG left idleAsleep untouched, so an endpoint idle-suspended
BEFORE the guard fired could be resurrected by the next dial (resumeOnDial
keys only on idleAsleep) — reintroducing the AmneziaWG-over-WireGuard kernel
hang the guard exists to prevent. Now clears idleAsleep under resumeMu so
the guard is ordered against a concurrent wake.
sendConnect's ReadFrame loop blocked forever on a peer that completes
TCP+TLS but never returns the CONNECT HEADERS, wedging the outbound under
o.runMu (Close hangs too). A ctx watcher now trips tlsConn's deadline on
timeout/cancel and is joined before the long-lived readLoop starts.
The main README's 'Features & status' table and Feature-configuration section
listed XHTTP/AWG/observability/round_robin but not the SPEC 021 MASQUE
(CONNECT-IP / WARP) outbound. Add it in both README.md and README.ru.md:
intro line, a Features table row (device-verified on Wi-Fi + LTE, h3/h2), and
a short config example with the network=transport footgun and the h2 fallback
for UDP:443-filtered networks. Links to docs-lx/lx-config §4 and SPECS/021.
The user-facing config guide covered XHTTP/AWG/urltest but not the SPEC 021
MASQUE (CONNECT-IP / WARP) outbound — only the internal SPECS/021 had it.
Add a full section (§4, renumbering Observability→§5, Validate→§6): feature
table row, field table, h3/h2 example, the network=transport footgun, dns-block
requirement, h3-vs-h2 guidance (UDP:443 filtering, cold-start), device-verified
status, and a link to SPECS/021/CONFIG.md. RU mirror kept in sync.
The 1.14 merge (6b63cee4) brought daemon/managed_service.go and
experimental/libbox/debug.go into vet scope; both crash Go ON PURPOSE via
*(*int)(unsafe.Pointer(uintptr(0)))=0 (TriggerDebugCrash/TriggerGoPanic),
which vet's unsafeptr analyzer flags. They are upstream files we don't edit,
and no lx-owned file uses unsafe at all, so disable just that analyzer.
Fixes the red lint job on every push since the merge.
A failed dial is an actionable error and should be visible where the success
(tunnel established, INFO) is — the LxBox core-log forwarder only surfaces
INFO+, so a DEBUG failure was invisible on-device. WARN makes established/
failed a symmetric, forwardable pair. The detailed dial phases (establishing,
udp-socket-up) stay DEBUG.
Promotion of the rc.1..rc.22 series to a non-prerelease tag (publishes as
Latest). Functionally rc.22 + the linux-mips-softfloat asset (#6). Section
header matches the tag exactly so lx-release.yml extracts it into the notes.
Rides out transient snapshot.debian.org 503s in the linux-musl toolchain
download/keyring steps (rc.22 failure cause). No code change; publish still
requires all builds.
rc.22 linux-musl jobs failed on 'get-clang.sh: 503 No healthy backends' from
snapshot.debian.org (transient mirror outage). get-clang.sh retries internally
but back-to-back, so a whole outage window fails all attempts. Wrap the two
Debian-fetching steps with an external retry + growing backoff:
- Download Chromium musl toolchain: 5 attempts, 30→60→120→240s
- Regenerate Debian keyring: 4 attempts, 20→40→80s
Conservative: publish still needs all builds (a real musl breakage still blocks
the release, only transient mirror flakes are ridden out). No code change.
Requested in #6 (Atheros AR9344). Chromium/cronet has no big-endian MIPS
toolchain, so the musl+naive path is impossible for this target; add it to
the plain cross-compile matrix instead as a pure-Go CGO_ENABLED=0 build —
statically linked (runs on musl/OpenWrt as-is), with with_naive_outbound
and with_purego dropped (purego has no mips port either). Everything else
matches the desktop tag set.
Verified locally: GOOS=linux GOARCH=mips GOMIPS=softfloat build with the
reduced tag set compiles clean; `file` reports ELF 32-bit MSB MIPS32,
statically linked.
Adds establish/handshake/success debug logs to the MASQUE dial path so a
stuck tunnel is diagnosable from /logs/core alone (motivated by the LxBox
§130 device case: h3 hung in QUIC handshake because inbound UDP:443 was
Log the tunnel-establish phases so a stuck dial is diagnosable from
/logs/core alone, without a goroutine dump:
- "establishing <h3|h2> tunnel to <server> (sni=...)" on start
- "udp socket up, starting QUIC handshake" (h3) — pinpoints whether a
hang is the socket or the handshake (inbound UDP:443 filtered → our
ClientHello left but no ServerHello came back)
- "tunnel established" on success, "tunnel failed: <err>" on failure
Motivated by a live device case (LxBox §130): h3 hung in the QUIC
handshake because inbound UDP:443 was filtered by the network while AWG
(UDP on a non-443 port) worked — invisible in logs before this, required
a pprof dump to locate. h2 (TCP:443) is the fix there.
Refs: SPEC 021
Complete masque outbound config reference verified against code:
full JSONC template (all params), minimal config, per-field table
(masque-specific + inherited DialerOptions), profile matrix, value
formats (duration/keys/ip), start-time validation, and common footguns
(network=transport not tcp/udp, dns block required, exit-IP changes on
reconnect, keepalive vs idle_timeout).
Rework the tunnel lifecycle around a *session (device + ipConn + closer +
ctx + activity counter), guarded by runMu with a generation guard.
- C1 (was HIGH): a dropped or suspended tunnel is now rebuilt on the next
dial. Previously 'running' latched true, so after the tunnel died every
DialContext short-circuited and dialed into a dead stack — permanent
blackhole. teardownSession clears o.sess so ensureSession rebuilds.
- C2: teardownSession closes ipConn, which unblocks the paired pump parked in
a blocking read (context cancellation alone can't interrupt it); no leaked
goroutine, no zombie half-open tunnel. Idempotent via sync.Once.
- B1: idleWatcher suspends the whole tunnel (gVisor netstack, pumps, QUIC
keepalive) after idle_timeout of no traffic; the next dial rebuilds it —
near-zero resident cost when idle.
- B4: idle_timeout (default 5m) and keep_alive_period (default 30s) are config
options; negative disables. A5: fail-fast mtu<=16000 on h2.
- D2: drop the dead congestion_control option field.
lifecycle_test.go covers the generation guard, idempotent teardown and
close-guard under -race.
Refs: SPEC 021 audit B1/B4/C1/C2/A5/D2
- A4/A5/B5: cap peer-declared capsule payloadLen at maxCapsulePayload (64KiB)
to prevent int-overflow/OOM on a hostile length; shrink recvCh 64->8 and the
receive window 1GiB->8MiB so real HTTP/2 flow-control provides backpressure
instead of unbounded RAM.
- B2: reuse a per-conn scratch for the outgoing capsule frame (tx pump is the
sole writer; writeData flushes before returning) instead of allocating per
packet.
Refs: SPEC 021 audit A4/A5/B2/B5
- A1: a single malformed/empty inbound datagram (unparseable context-ID
varint) no longer tears the tunnel down — drop-and-continue like the
sibling context-ID!=0 / bad-payload cases.
- A3: snapshot the IP header before composeDatagram mutates TTL/checksum, so
the ICMP 'packet too big' reply quotes the original datagram (RFC 1191/792).
- A2: drop the redundant second status check in ConnectTunnelH3 (2xx already
validated in dialCONNECTIP); removes the unused responseInfo carrier.
- B3: reuse a per-Conn scratch for the outgoing datagram (contextID + packet)
instead of allocating per packet — safe, quic-go copies the slice before
SendDatagram returns and the tx pump is the sole writer.
Refs: SPEC 021 audit A1/A2/A3/B3
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).
The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.
Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.
Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.
Refs: SPEC 021 TEST_PLAN.md
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.
Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
semantics differ from WARP's expected extended-CONNECT authority/headers
(SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.
Refs: SPEC 021 TEST_PLAN.md
The "GRO off + batch 8" idea (a global alternative to Down/Up) was measured on-device
and REJECTED, for three independent reasons (SPEC.md §14):
1. Wrong holder — the main android RAM holder is device.pool.messageBuffers
(PreallocatedBuffersPerPool=4096 × ~64KB ≈ 100MB), which does NOT depend on
BatchSize; the batch-sized bufsArrs held only ~14MB. Shrinking batch wouldn't
have touched the ~100MB.
2. Not deliverable — the LX_WG_NO_GRO env switch never reaches Go's os.Getenv on
Android (wrap.<pkg> prop shows in /proc/environ but not in the runtime's env
snapshot), forcing a hardcode.
3. Fragile — hardcoded batch=8 crashed at start (SIGABRT): device.BatchSize()=
max(bind,tun) clamped back to 128 via the TUN offload while msgsPool was 8, so
Send sliced out of range. Coherent only by also gating TUN offload across three
submodule layers.
Down/Up (rc.19) stays the only viable mechanism. Brings the experiment folder
(protocol + device heap snapshots + RESULT) into lx-1.14 for the record; the
experiment CODE stays on the lx-1.14-nogro-* branches, not merged.
Sync SPEC 002 with the code (commit c0bbb1c5): GET on a non-packet-up node no
longer hard-errors — it falls back to POST + WARN so one bad subscription node
doesn't fail the whole config. Updated the mode-gate wording in SPEC.md §verif,
PARAM_MAP.md (full rationale + the old error text it replaces), URL_PARSING.md
table, and IMPLEMENTATION_REPORT.md. header/cookie uplink outside packet-up
stays a hard error (no safe default).
A subscription node sometimes ships uplink_http_method=GET on a non-packet-up
node (auto/stream-up/stream-one). GET can only carry the uplink in packet-up
(other modes put the body in the request, which GET has none of), so the strict
check rejected it — failing the ENTIRE config over one bad outbound in a large
subscription (observed: initialize outbound[361] ... can be GET only in
packet-up mode → whole tunnel won't start).
Fall back to POST (the safe default that works in every mode) and log a WARN
instead of erroring, so the rest of the config still loads. POST is what the
node should have used; the fallback just makes one malformed remote node
self-healing rather than fatal. Kept strict for packet-up (GET honoured there).
Uses log.StdLogger() for the warning (the client transport layer gets no
logger in its constructor signature; threading one through would touch upstream
signatures). Test: TestUplinkGetFallsBackToPostOutsidePacketUp; removed the GET
case from TestValidationRejections. Verified on a real binary — check on a
stream-one+GET config now warns and passes (exit 0) instead of FATAL.
The base-version step (c4fd73cd) adds an `upstream` remote (SagerNet/sing-box)
so git-describe can see the v1.14.0-alpha.* tags. But `gh release create` without
--repo resolves the target repo from the remotes and picked `upstream` →
HTTP 403 "Resource not accessible by integration" against
api.github.com/repos/SagerNet/sing-box/releases (the token has no rights there).
This is why rc.19's builds all succeeded but publish failed, while rc.18 (before
the upstream remote existed) published fine.
Pin --repo "${{ github.repository }}" so publish always targets this fork
regardless of what remotes the earlier steps added.
rc.19 gates idle-suspend behind with_lx_idle_suspend (mobile-only) and records the
on-device Android verification: suspending 8 idle+unreachable WG endpoints freed
134MB of bufsArrs live heap (223.9→89.9MB, recv-workers 18→2), matching the
~8.4MB/worker model — ~10x the desktop delta, on the platform the feature targets.
Adds ANDROID_RESEARCH/live-baseline/ — a full pprof snapshot of a real production
config with the feature OFF (263MB bufsArrs, 56% CPU on GC at idle) and its ON
"after" counterpart, closing the RESEARCH.md device gap end-to-end (buffer pool
and GC cost measured together, not inferred). Credentials scrubbed.
Idle-suspend frees the recv-worker bufsArrs, which are ~8MB each only where
BatchSize=128 (Android/Linux) — on desktop BatchSize is small and the feature
saves almost nothing. Make that platform scope explicit in the build instead of
running the tick everywhere.
The idle-suspend tick now compiles only with the new `with_lx_idle_suspend` tag,
baked into the mobile AAR (build_libbox sharedTags) but NOT the desktop LX_TAGS.
Without the tag, a config that sets route.lx_idle_suspend fails fast at start
("rebuild with -tags with_lx_idle_suspend (mobile-only feature)") rather than a
silent no-op. The gate is a single function: reachability_lx.go carries the tick
under the tag, idle_suspend_stub_lx.go is the no-tag stub that errors, and
reachability_common_lx.go keeps InvalidateReachability (needed by the group
interface in every build). The dial hot path (resumeOnDial/stampActivity) and the
upstream group files are untouched — without the tick, idleAsleep is never set, so
resumeOnDial always takes its fast path.
Adds stub unit tests (option set → error, unset → no-op). Both build variants and
the full route/wireguard/group suites are green; gofmt/vet clean; desktop lx-check
passes without the tag. Docs (lx-config.md + ru, SPEC.md §3/§10) describe the tag.
The base-version derivation used `git describe --match v1.14.0-alpha.*` as the
primary source, but actions/checkout only fetches THIS repo's tags — the alpha
tags are SagerNet/sing-box (upstream) tags, absent in the CI clone. So `git
describe` found nothing and silently fell to the subject-grep fallback, which
resolves alpha.36 (alpha.37 was merged in a commit whose subject omits the
number). That's why rc.17/rc.18 notes shipped "base alpha.36" while a local
clone with upstream tags gets 37.
Fetch just the upstream v1.14.0-alpha.* tags before git describe so the primary
graph-based path works in CI. Both hand-fixed on the published releases; this
makes the next tag correct automatically.
Full write-up of the Android device run (CPH2411, Android 15, rc.18) in a
dedicated ANDROID_RESEARCH/ subfolder: README (report), METHOD (reproducible
procedure), RESULTS (per-scenario + heap A/B), and artifacts/ (raw evidence:
lx idle log lines, goroutine dumps, pprof heap .pb + top renders). No access
credentials anywhere.
Headline, now measured on the target platform: PopulatePools.func3 (the
bufsArrs holder from RESEARCH.md) inuse_space 223.93 -> 89.89 MB (-134 MB /
-60%), recv-workers 18 -> 2, on suspending 8 of 9 WG endpoints. = 16 workers x
~8.4 MB (BatchSize=128), matching the source model, ~10x the desktop RSS delta.
RESEARCH.md status + SPEC.md §12/§13 updated: the Android heap A/B gap is
closed (only the battery A/B remains deferred).
Device run on CPH2411 (Android 15, rc.18) via the LxBox app Debug API.
9 WG endpoints (1 real WARP reachable + 8 synthetic unreachable),
lx_idle_suspend=30s. All behaviors confirmed on-device: suspend fires,
reachable final stays up, wake-by-dial, no-flap, kill-switch.
Headline: PopulatePools.func3 (the bufsArrs holder from RESEARCH.md)
inuse_space 223.93 to 89.89 MB (-134 MB / -60%), recv-workers 18 to 2.
= 16 freed workers x ~8.4 MB (BatchSize=128), matching the model, ~10x
the desktop RSS delta. Closes the Android device-verification gap.
Device-verified idle-suspend: idle AND unreachable WG/AWG endpoints go Down to
free their recv-worker bufsArrs (the Android GC-heat holder), waking on next dial.
Opt-in via route.lx_idle_suspend; off by default. Notes cover the reachability
walk, the GRO reason for Down-over-smaller-batch, the concurrency fixes, and the
2026-07-01 live-run results (recv-workers 16→0, RSS −31%).
Selectively brings idle + unreachable WireGuard/AmneziaWG endpoints Down,
freeing their recv-worker bufsArrs (the measured GC-scan heat holder on
Android) and stopping their per-peer timers (battery), then wakes them lazily
on the next dial. Off by default (lx_idle_suspend absent/0 = zero overhead).
Includes the fix for the shipped tick iterating the wrong manager (it never
reached any endpoint — the feature was inert on a live box), the full
reachability walk (final/rule/selector Now/urltest pool/detour, event-driven
cached), and the rewritten as-built spec (SPEC.md) + research doc (RESEARCH.md)
+ test plan.
Live-verified: suspend/wake/probe-wake/re-sleep/no-flap/kill-switch across
selector, urltest pool, nested groups, AWG-guard, and the real production
config; resource A/B recv-workers 16->0, RSS -31% on desktop. 29 unit tests,
adversarially checked.
The idle-suspend feature is implemented, the tick bug is fixed, and every
reachability node type plus suspend/wake/probe/no-flap/kill-switch and the
resource A/B (recv-workers 16→0, RSS -31%) are live-verified. Reflect that in
the docs and give the folder clean roles:
- SPEC.md (was SPEC_idle_suspend_lever.md): rewritten from scratch in Russian
as the as-built implementation spec — Down/Up model, reachability walk +
event-driven cache, endpoint-side suspend/wake, the tick bug and its fix
(§11), full test coverage (§12, 29 units named), and what is deliberately
deferred (§13: Tier B netstack teardown, keys-safe BindUpdate path, on-device
battery/heap measurement). Old Tier-A "light sleep" design (never shipped)
removed.
- RESEARCH.md (was SPEC.md): the diagnostic root-cause doc keeps its unique
on-device heap A/B proof (holder = recv-worker bufsArrs) — renamed so its
role (research, not implementation spec) is unambiguous.
- TEST_PLAN_idle_suspend.md: all pass criteria checked, §RESULTS + edge-case
matrix + wake-latency series (cold ~50ms / warm ~36ms, +14-21ms ≈ 1 handshake;
far-server caveat) + production-config run.
Cross-references and section numbers updated across all three files.
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Code recon of the wireguard-go submodule disproved SPEC.md's original PRIMARY
lever (shrink StdNetBind.BatchSize() 128→8). It cannot be done without breaking
GRO receive:
- GRO-rx is ENABLED on android (UDP_GRO set with no android gate,
controlfns_linux.go:90-104; SPEC 010 gated only GSO-tx, not GRO-rx) → rxOffload=true.
- The GRO path splits one coalesced packet into up to 64 datagrams; readAt =
len(msgs) - IdealBatchSize/udpSegmentMaxDatagrams (bind_std.go:269) HARDCODES
IdealBatchSize=128, and getMessages() allocs a 128-slot array. Shrinking bufsArrs
to 8 either desyncs bufs(8) vs array(128) → OOB panic, or overflows the split
("splitting coalesced packet resulted in overflow", bind_std.go:565). GRO can't
be disabled (needed for download throughput, §010).
So the old claim "packet loss excluded, array just shorter" was wrong for the GRO
path. Lever 1 (and lever 2, which inherits the same idle-socket GRO problem) are
rejected. PRIMARY becomes lever 3 — Down idle+unreachable devices: BindClose ends
the recv-workers and frees bufsArrs whole, while the active node keeps batch=128 so
its GRO is intact. The "most expensive fallback" is in fact the only viable lever.
Updated: status line, the lever-candidates section (struck lever 1, promoted lever 3),
the fix-logic section (renamed + rewritten around Down), verification, and residual
risks (the fast-channel risk is gone; the new risk is handshake-on-wake). Implemented
on lx-spec020-idle-suspend; see SPEC_idle_suspend_lever.md §13 + TEST_PLAN.
Add TEST_PLAN_idle_suspend.md — build/config/commands/pass-criteria to
device-verify the shipped idle-suspend on a real run: suspend fires for
idle+unreachable WG/AWG endpoints, reachable ones never suspend, wake-on-dial,
the bufsArrs memory drop (pprof heap), and no flapping. Uses the user's
WARP/AWG + plain-WG nodes. Link it from SPEC_idle_suspend_lever.md §13.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Record the as-built decision so it is not re-derived:
- §13.1 what shipped (c55cf11e): idle-suspend via Down/Up, not light sleep
(source-verified holder is recv-worker bufsArrs; timersStop does not free it).
- §13.2 bind-swap investigation (the promised comment): BindUpdate resizes the
bind WITHOUT zeroing keys; key-zeroing lives only in Down/peer.Stop. So a
keys-safe wake is possible only while still Up — the shipped Down path can't.
- §13.3 the three reduced-bind paths (B=Down shipped / A=BindUpdate keys-safe /
Hybrid) with the GRO-off + max(bind,tun) gotchas.
- §13.4 recommendation: ship B, escalate to A/Hybrid only if device INFO logs
show handshake flapping hurts. Reduced-bind urltest wake deferred (low value
on path B since keys are already zeroed).
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.
- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.
Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
Complete RU translation of docs-lx/lx-config.md, mirroring its structure:
the §0 exhaustive "every field at a glance" example, all per-section field
tables (XHTTP v1+v2, AmneziaWG 2.0, id/ip/ib masquerade, urltest balancer),
examples and the build section. JSONC code is preserved; only prose and
inline comments are translated. Intra-doc #anchors are re-pointed to the
Russian heading slugs (all 6 verified to resolve); the §0 example validates
as JSON.
Cross-link both ways (en ↔ ru) and re-point README.ru.md's four lx-config
links to the Russian version. File name follows the README.ru.md convention
(.ru.md, not -ru.md).
Add a §0 kitchen-sink config carrying ALL 52 lx-added fields in one place —
XHTTP transport (26), AmneziaWG 2.0 endpoint incl. id/ip/ib (21), urltest
round_robin balancer (5) — each with its default and allowed values inline,
and mutually-exclusive / server-ignored fields flagged. Sourced by reading
option/*.go directly (not the prior doc), so it is complete.
This also surfaced that §1 documented only 7 of the 26 XHTTP fields (the v1
set); fill in the 19 missing v2 fields (session/seq placement, uplink-data
placement, X-Padding obfs family, packet-up tuning, accepted-but-ignored)
as grouped tables. Fix three code-vs-doc disagreements the extraction found:
- `mode: auto` resolves to stream-one on Reality (not always packet-up);
- s3/s4 are AWG 2.0 junk-size params (not "cookie-reply/transport" junk);
- h1-h4 unset spelling includes "" as well as 0; x_padding_bytes framing.
The default wire shape is unchanged; all v2 fields are opt-in. §0 example
validated as JSON. Russian translation (lx-config.ru.md) to follow.
docs/ is an upstream-owned tree (it arrives wholesale from SagerNet on
every rebase). Our three downstream docs lived inside it — lx-config.md,
lx-changelog.md, lx-release-runbook.md — mixing fork files into the
upstream surface against CONSTITUTION principle #1 (thin layer / minimal
diff). Move them to a dedicated root-level docs-lx/ so the boundary
between our docs and upstream's is explicit.
- git mv preserves history.
- Updated every reference (docs/lx-* -> docs-lx/lx-*): README.md/.ru.md,
SPECS/{003,004,005,009,020,README}, transport/wireguard/endpoint.go
comments, lx-ci.yml, and lx-release.yml (the release-notes extractor +
fallback URL now read docs-lx/lx-changelog.md).
- Fixed the now-relative links inside the moved files that pointed at
upstream docs/ siblings: lx-config.md -> ../docs/configuration/outbound/
urltest.md; lx-changelog.md -> ../docs/changelog.md (x2).
Verified: all relative + external links resolve, both workflows are valid
YAML, the release-notes awk path is docs-lx/, go vet clean on the touched
package. No release feature — folds into the next tag naturally.
Secondary design doc alongside the authoritative SPEC.md, NOT a replacement.
SPEC.md has on-device proof (heap A/B) that the GC-scan holder is bufsArrs of
recv-workers and its primary lever is shrinking StdNetBind.BatchSize() — that
stays authoritative.
This companion works out SPEC.md's 'lever 3' (suspend inactive devices) in
detail: a light variant (per-peer timersStop, keypairs/socket kept live, cheap
wake without handshake) plus a reachability walk (final + rules + active
selector/pool choices, generation-cached) that decides which devices are idle
AND unreachable from the active routing tree.
Banner up top flags where this doc's source-reading diverged from SPEC.md's
measurements (it guessed gvisor netstack; SPEC.md measured bufsArrs) and notes
light-suspend does NOT free bufsArrs — only Down/BatchSize does. Use as the
fallback design for lever 3, not a competing primary.
Spec only — no code changes.
The subject-grep base-detection missed alpha.37: it was merged in a commit
titled "Merge upstream/testing (bump version, fix linux ping)" with no
"alpha.37" in the subject (upstream tagged it after we merged), so the grep
found only alpha.36 and rc.17 notes shipped a stale base.
Make `git describe --match v1.14.0-alpha.*` the primary source — it reads
HEAD's ancestry in the commit graph, independent of merge-message wording —
and keep the subject-grep as the fallback for a fork checkout without
upstream tags. Verified locally: now resolves v1.14.0-alpha.37.
SPEC 014 dropped with_clash_api because LxBox (Android) drives the core
over the native libbox CommandClient, making the Clash REST server dead
weight in the AAR. But the drop landed in the shared Makefile.lx LX_TAGS,
which also feeds every desktop/CLI release build (mac/windows/linux-musl
via `make -s lx-print-tags`). A CLI binary has no CommandClient channel —
it is managed by external dashboards (yacd/MetaCubeXD) over the Clash REST
API — so every desktop release since rc.1 shipped with no way to manage
the core; a config with experimental.clash_api failed fast. CI stayed
green (lx-ci BASE_TAGS kept the tag), so it was invisible in CI.
Restore with_clash_api to the desktop LX_TAGS; leave build_libbox (AAR)
unchanged. The two tag sets now diverge by design: desktop = with Clash
API, AAR = without.
Verified: desktop binary builds with with_clash_api in Tags; `check`
accepts an experimental.clash_api config; the Clash REST server comes up
live (endpoints answer 401 security-middleware, not the stub's fail-fast).
Docs: Makefile.lx comment, SPEC 014 (§2/§3.1 scoped to AAR + new §3.4),
lx-release.yml tag comment + notes line, changelog rc.17.
The release-notes template hardcoded "base v1.14.0-alpha.35"; it went stale and
had to be hand-edited on rc.14, rc.15 and rc.16 (each was actually on alpha.36).
Resolve the base dynamically in the "Resolve tag" step: take the highest alpha.NN
named in any "Merge upstream" commit subject (robust on a fork without upstream
tags fetched), falling back to git describe against upstream alpha tags, then a
generic v1.14.x label. The notes line now interpolates steps.ver.outputs.base.
§8: validated transport JSON with all 14 new fields at non-default values,
the equivalent flat-camelCase vless:// URL, and a defaults table for the
toUri() omitempty logic. Fixture verified with sing-box check; mirrors
lx-test/config/xhttp_obfs_full.json.
Relocate docs/lx-xhttp-url-parsing.md -> SPECS/002-XHTTP_CLIENT_TRANSPORT/URL_PARSING.md
so all XHTTP docs live together with the spec. Fix internal/back links.
Scanned igareck/vpn-configs-for-russia, extracted+deduped 10 unique XHTTP
nodes, ran each through our with_xhttp binary. 4 alive — all downloaded 1MB,
traffic egressed via the server IP:
- 2x plain -> packet-up
- 2x reality -> stream-one (hu99.bearbeer.digital, bez3.stream-room.com)
The two reality nodes resolve auto->stream-one and work live, closing the
open stream-one live-verification TODO from task 011 (previously synthetic-only).
Other 6 nodes dead for server-side reasons (504, HTTP/1.1-not-H2, reset,
TLS hang) — our transport errored cleanly in every case.
Remaining live TODO: obfs/placement modes (no public node is configured for them).
Our XHTTP client is HTTP/2 only (http2.Transport); Xray supports H1/H2/H3.
h3-only nodes won't connect — flag for the link parser. Out of SPEC 002 scope
(separate future 'XHTTP over HTTP/3' task).
scMaxConcurrentPosts is a removed Xray knob (grep + GitHub code search
total:0 in current XTLS/Xray-core and sing-box-extended). Current Xray
serializes to one upload POST body in flight at a time, which our sequential
packet-up Write already matches, so the field is accepted for config/link
symmetry but ignored by the client.
- option: V2RayXHTTPOptions.ScMaxConcurrentPosts (json sc_max_concurrent_posts)
- PARAM_MAP: document as legacy/ignore tier with the real concurrency mechanism
(bounded pipe + WroteRequest serialization, server-side seq reorder)
- url-parsing doc: scMaxConcurrentPosts -> accept-but-ignore
- xhttp_obfs_full.json: include the field so check covers it
Verified: build/gofmt/vet clean, 16 unit tests pass, sing-box check passes
on all 3 xhttp configs incl the field.
Self-contained reference for the link parser: maps every vless://...type=xhttp
URL param (flat query + extra={...} JSON) to sing-box transport snake_case fields.
Covers TLS/Reality mapping, the extra-JSON number→"min-max" coercion, mode=auto
pass-through, path-with-query-tail, and ignored fields (scMaxConcurrentPosts,
server-only). Examples validated with sing-box check.
Implement all 12 client-relevant Xray/sing-box-extended XHTTP params on the
existing lean-native client (no Xray vendoring):
- session/seq placement (path|query|header|cookie) + keys
- uplink-data placement (body|auto|header|cookie, chunked base64) + key + chunk size
- uplink_http_method (upper-cased; GET only in packet-up)
- X-Padding obfs mode: placement (cookie|header|query|queryInHeader) + key/header +
method repeat-x | tokenish (HPACK-Huffman-tuned via golang.org/x/net/http2/hpack)
- packet-up tuning: sc_max_each_post_bytes (split), sc_min_posts_interval_ms (throttle)
4 server-only fields (server_max_header_bytes/no_sse_header/sc_max_buffered_posts/
sc_stream_up_server_secs) accepted but ignored by the client.
New files: transport/v2rayxhttp/{meta.go,xpadding.go}, xhttp_test.go.
Range fields use the "min-max" string form (no badoption.Range in sing).
Default (non-obfs) wire shape kept byte-identical to the live-verified v1
(x_padding='0' in Referer, session/seq on path, payload in body).
Verified: 16/16 unit tests, sing-box check on 3 configs incl full obfs,
go vet/gofmt/build (tagged+untagged) clean, negative (no with_xhttp) rejects.
Adversarial wire-protocol review against PARAM_MAP found no bugs.
Live test of a non-default mode against an Xray server remains an open TODO.
The rc.15 domain fix was confirmed on a real device: with the default
sticky_hash ["process","domain"] and no dest_ip workaround, browser traffic
spreads across the pool (on-device per-domain uniformity ~0.27 -> 0.95+).
Update README + lx-config.md status from "not yet device-verified" to
device-verified.
PRIMARY lever spelled out: bufsArrs size = bind.BatchSize() (128 on StdNetBind/android);
shrink it at one point (conn/bind_std.go:322, android branch -> 8/16). BindUpdate
(device.go:558) and getMessages() (bind_std.go:260) follow automatically, both recv
goroutines (v4+v6) covered, no packet loss (array stays full, just shorter), MaxSegmentSize
untouched (GRO intact). Effect 8MB->~0.5MB/recv = 176MB->~11MB at 11 devices. Only risk:
gigabit channel may lose throughput -> fall back to dynamic batch.
On-device throughput A/B (static arm64 curl, download via tunnel):
- baseline batch=128 = 10.7 MB/s median (~86 Mbps), stable 9.1-11.2.
- CPU under load: Syscall6 36%, scanobject 6%, crypto ~3%. The bottleneck is the
WARP channel + syscall overhead, NOT batch processing. GRO/batch only matters at
hundreds-of-Mbps/gigabit, so shrinking batch does NOT cost throughput on a typical
mobile/WARP channel (where the heat is reported).
Re-ranked the levers: PRIMARY is now the global smaller StdNetBind.BatchSize()
(128->8-16) — one point, no activity detection, cuts bufsArrs 8MB->~1MB/recv (176MB->
~11-22MB at 11 devices). Dynamic-batch and Down-idle drop to secondary. Caveat: verify
on a fast Wi-Fi/gigabit channel before release (batch may matter there). Also confirmed
heap scales linearly with live device count (11->269MB, 4->104MB, 1->0). No code changed.
The lx feature docs had drifted: README (en/ru) and docs/lx-config.md still said
"currently XHTTP + AWG2" and covered only SPEC 002/003/009 — the observability
layer (SPEC 014-018) and round_robin load balancing (SPEC 019) were undocumented
in the lx overview, and urltest.md still described the pre-rc.15 domain behaviour.
- docs/lx-config.md: new "## 3. round_robin load balancing" (mode/balancer,
pool/pool_tolerance/sticky_hash, ["none"] sentinel + badjson-[] caveat, slot-hash
binding, example, status) and "## 4. Observability (CommandClient extensions)"
(URLTestOutbound/GetRules/GetGroups/GetOutbounds/GetPool/SubscribeDNSQueries +
Connection.detourList, all behind with_lx_command); Validate&build -> ## 5.
- README.md / README.ru.md: broaden the stale "XHTTP + AWG2" framing; add feature
rows for observability and round_robin with honest status.
- docs/configuration/outbound/urltest.md: reconcile sticky_hash "domain" with the
rc.15 fix — domain reads metadata.Domain (survives domain->IP resolve), so it
works for normal sniffed domain traffic, not only literal-IP destinations; the
warning is reframed (domain works; dest_ip is an alternative).
Docs-only; no code change.
On-device A/B (Debug API /diag/pprof) settles it with high confidence:
- RoutineReceiveIncoming holds 180MB (61% cum); peek = 100% via sync.Pool.Get (in
worker hands, not the pool, not the channels).
- 11 live wireguard endpoints, 22 RoutineReceiveIncoming ALL on StdNetBind (batch=128),
0 on ClientBind. 22 x bufsArrs[128] x 64KB = 176MB ~= 180MB.
- batch=128 because WARP/AWG endpoints use a WireGuardListener dialer => StdNetBind
(endpoint.go:200-202), whose BatchSize()=128 on android.
- A/B: switching the active node WARP->home does NOT free memory (buffers do not sleep);
config 11 ep -> 1 ep gives 269MB -> 0 (= Iliya's workaround, reproduced via profile).
Both prior diagnoses were wrong: batch=1 (no, StdNetBind=128) and drain-on-Suspend of
channels/sync.Pool (misses; bufsArrs of live recv-workers holds it). MaxSegmentSize
2200->65535 is the volume trigger (x30 bytes), not the holder (downLocked/pools/channels/
batch identical 1.13<->1.14).
Fix = shrink batch for INACTIVE devices (naive lazy-bufsArrs impossible: StdNetBind
getMessages() is a fixed 128). Levers + a required download-throughput measurement
documented; pending lever choice.
On the client path BatchSize()=1 (client_bind/stackDevice/systemDevice all return 1),
so the old bufsArrs=128 => 8MB/device claim is wrong; bufsArrs is ~64KB. The 224MB
pprof attributes to PopulatePools is the sync.Pool.New alloc SITE, not the holder.
Real holders: (A) device.pool messageBuffers sync.Pool local+victim cache, (B) the 3
buffered device channels. scanobject 52% comes from the pointer-dense element/container
wrappers (4 of 5 WaitPools are scan-type), not the noscan [65535]byte arrays.
Fix rewritten to drain-on-Suspend: park RoutineReadFromTUN via a stackDevice.suspended
seam in Read (the lockless pool writer surviving Down), drain the 3 device channels
after Down, then ONE runtime.GC()+FreeOSMemory() per selector transition. PopulatePools
swap rejected (WaitPool.count underflow -> cond.Wait deadlock). slim-batch and the
route-graph refcount are dropped (not needed for heat). Added in-repo verification
(device/suspenddrain_test.go + HeapInuse bench) since on-device A/B is impossible.
Android 100% CPU / heat on configs with many WG/AWG endpoints in a
selector. Diagnosed via on-device pprof (§207): scan-bound GC over a
224MB live heap = wireguard-go Device.PopulatePools buffers, held by
~10 idle WG devices. Suspend()=Down() marks the device idle but does
NOT release pools/workers (only Close() does). A/B on device: dropping
spare WG endpoints removes the heat.
SPEC 020: SLIM idle devices (shrink maxBatchSize 128->1-4, do not Close
— keepalives + shared-node safety) gated by a route-reachability
refcount (rules + final + Now, not just selector).
SPEC 010: note our GRO split-brain patch is now upstream-native on
v0.0.3 (commit 24ea133); MaxSegmentSize=65535 must stay (GRO fuel) —
heat is fixed by device count/slimming, never by shrinking the buffer.
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.
1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
The router resolves a domain destination to an IP and overwrites
metadata.Destination before a group's DialContext runs, so destination.Fqdn
is empty when the balancer builds the key. stickyComponent("domain") read
that empty Fqdn, so a single process's key was process+NUL for every site
-> one fixed slot. On device this measured 28/1/1 across a 3-node pool
(uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
back to destination.Fqdn only for a direct dial. After: spread 0.95+.
2. living pool nodes could change slot index during a health-check, moving
sticky keys. balancePoolFirstLive compacted with a filtering append (a
transiently-dead slot shifted every later live node left); planTolerantPool
did delete(inPool, occupant) (an evicted-but-living node re-entered a later
slot, cascading); manual URLTest rebuild ran the tolerant planner even at
pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
dead/empty slots rewritten by index; dedicated planFirstLivePool for the
tolerance==0 rebuild).
3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
(badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
empty array to nil, indistinguishable from omitted, so the default always
applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].
Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
v2 superseded v1; keeping both as separate files (SPEC.md + SPEC_V2.md + the v1
TEST_REPORT) was just confusing. Delete the v1 SPEC and its TEST_REPORT (they remain in
git history) and rename SPEC_V2.md → SPEC.md as the one canonical doc. Drop the "v2"
suffix and stale "design not started" status from the header.
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.
Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.