SPEC.md rewritten to current (multiplex) architecture, no chronology.
HISTORY.md captures v1 standalone class-error, the field bug, rejected paths.
New project rule (README + CONSTITUTION 3.2): SPEC.md = current state first,
chronology/rationale of architecture changes go to HISTORY.md.
Runtime detour/selector rings crash the core (fatal stack overflow via
unbounded DialContext recursion); static rings are already rejected at start
by lintOutbound. Worked out the full event model (E1-E5) and topologyMu race
linearization, then adversarially verified it (7-agent workflow): deadlock and
false-positive attacks HOLD, but TOCTOU BREAKS — even a correct core guard is
not airtight without also covering the endpoint manager, Manager.Remove,
history side-channels, and the pointer-vs-tag graph divergence after a runtime
Create. Owner decision (2026-07-06): protection lives at the UI level (LxBox
validates before SelectOutbound); core stays a minimal delta to upstream.
No core code changed. SPEC is a design record + Roadmap row (status DEFERRED).
Producer run populated musl-toolchain-cache with 4 arch assets; restore path
validated locally (asset name, gh download, tar layout under naiveproxy/src).
Status -> C, Roadmap updated.
snapshot.debian.org intermittently 503s during the musl sysroot build and
blocks releases (v1.14.0-lx.2-rc.1 failed twice on it). actions/cache also
misses across tag builds (ref-scoping). Add a producer workflow that uploads
the built toolchain to a musl-toolchain-cache release, and a restore step in
lx-release.yml that pulls it on cache-miss before falling back to
snapshot.debian.org. Both workflows are lx-owned; zero upstream diff.
Full audit of the LX delta (10 axes, adversarial verification): 32 findings,
27 confirmed, 24 fixed on branch lx-spec022-audit-fixes, 3 skipped by design
(#12/#17/#18). Records #19 resolution (SPEC 013 test kept — upstream ships none).
Complete masque outbound config reference verified against code:
full JSONC template (all params), minimal config, per-field table
(masque-specific + inherited DialerOptions), profile matrix, value
formats (duration/keys/ip), start-time validation, and common footguns
(network=transport not tcp/udp, dns block required, exit-IP changes on
reconnect, keepalive vs idle_timeout).
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).
The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.
Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.
Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.
Refs: SPEC 021 TEST_PLAN.md
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.
Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
semantics differ from WARP's expected extended-CONNECT authority/headers
(SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.
Refs: SPEC 021 TEST_PLAN.md
The "GRO off + batch 8" idea (a global alternative to Down/Up) was measured on-device
and REJECTED, for three independent reasons (SPEC.md §14):
1. Wrong holder — the main android RAM holder is device.pool.messageBuffers
(PreallocatedBuffersPerPool=4096 × ~64KB ≈ 100MB), which does NOT depend on
BatchSize; the batch-sized bufsArrs held only ~14MB. Shrinking batch wouldn't
have touched the ~100MB.
2. Not deliverable — the LX_WG_NO_GRO env switch never reaches Go's os.Getenv on
Android (wrap.<pkg> prop shows in /proc/environ but not in the runtime's env
snapshot), forcing a hardcode.
3. Fragile — hardcoded batch=8 crashed at start (SIGABRT): device.BatchSize()=
max(bind,tun) clamped back to 128 via the TUN offload while msgsPool was 8, so
Send sliced out of range. Coherent only by also gating TUN offload across three
submodule layers.
Down/Up (rc.19) stays the only viable mechanism. Brings the experiment folder
(protocol + device heap snapshots + RESULT) into lx-1.14 for the record; the
experiment CODE stays on the lx-1.14-nogro-* branches, not merged.
Sync SPEC 002 with the code (commit c0bbb1c5): GET on a non-packet-up node no
longer hard-errors — it falls back to POST + WARN so one bad subscription node
doesn't fail the whole config. Updated the mode-gate wording in SPEC.md §verif,
PARAM_MAP.md (full rationale + the old error text it replaces), URL_PARSING.md
table, and IMPLEMENTATION_REPORT.md. header/cookie uplink outside packet-up
stays a hard error (no safe default).
rc.19 gates idle-suspend behind with_lx_idle_suspend (mobile-only) and records the
on-device Android verification: suspending 8 idle+unreachable WG endpoints freed
134MB of bufsArrs live heap (223.9→89.9MB, recv-workers 18→2), matching the
~8.4MB/worker model — ~10x the desktop delta, on the platform the feature targets.
Adds ANDROID_RESEARCH/live-baseline/ — a full pprof snapshot of a real production
config with the feature OFF (263MB bufsArrs, 56% CPU on GC at idle) and its ON
"after" counterpart, closing the RESEARCH.md device gap end-to-end (buffer pool
and GC cost measured together, not inferred). Credentials scrubbed.
Idle-suspend frees the recv-worker bufsArrs, which are ~8MB each only where
BatchSize=128 (Android/Linux) — on desktop BatchSize is small and the feature
saves almost nothing. Make that platform scope explicit in the build instead of
running the tick everywhere.
The idle-suspend tick now compiles only with the new `with_lx_idle_suspend` tag,
baked into the mobile AAR (build_libbox sharedTags) but NOT the desktop LX_TAGS.
Without the tag, a config that sets route.lx_idle_suspend fails fast at start
("rebuild with -tags with_lx_idle_suspend (mobile-only feature)") rather than a
silent no-op. The gate is a single function: reachability_lx.go carries the tick
under the tag, idle_suspend_stub_lx.go is the no-tag stub that errors, and
reachability_common_lx.go keeps InvalidateReachability (needed by the group
interface in every build). The dial hot path (resumeOnDial/stampActivity) and the
upstream group files are untouched — without the tick, idleAsleep is never set, so
resumeOnDial always takes its fast path.
Adds stub unit tests (option set → error, unset → no-op). Both build variants and
the full route/wireguard/group suites are green; gofmt/vet clean; desktop lx-check
passes without the tag. Docs (lx-config.md + ru, SPEC.md §3/§10) describe the tag.
Full write-up of the Android device run (CPH2411, Android 15, rc.18) in a
dedicated ANDROID_RESEARCH/ subfolder: README (report), METHOD (reproducible
procedure), RESULTS (per-scenario + heap A/B), and artifacts/ (raw evidence:
lx idle log lines, goroutine dumps, pprof heap .pb + top renders). No access
credentials anywhere.
Headline, now measured on the target platform: PopulatePools.func3 (the
bufsArrs holder from RESEARCH.md) inuse_space 223.93 -> 89.89 MB (-134 MB /
-60%), recv-workers 18 -> 2, on suspending 8 of 9 WG endpoints. = 16 workers x
~8.4 MB (BatchSize=128), matching the source model, ~10x the desktop RSS delta.
RESEARCH.md status + SPEC.md §12/§13 updated: the Android heap A/B gap is
closed (only the battery A/B remains deferred).
Device run on CPH2411 (Android 15, rc.18) via the LxBox app Debug API.
9 WG endpoints (1 real WARP reachable + 8 synthetic unreachable),
lx_idle_suspend=30s. All behaviors confirmed on-device: suspend fires,
reachable final stays up, wake-by-dial, no-flap, kill-switch.
Headline: PopulatePools.func3 (the bufsArrs holder from RESEARCH.md)
inuse_space 223.93 to 89.89 MB (-134 MB / -60%), recv-workers 18 to 2.
= 16 freed workers x ~8.4 MB (BatchSize=128), matching the model, ~10x
the desktop RSS delta. Closes the Android device-verification gap.
The idle-suspend feature is implemented, the tick bug is fixed, and every
reachability node type plus suspend/wake/probe/no-flap/kill-switch and the
resource A/B (recv-workers 16→0, RSS -31%) are live-verified. Reflect that in
the docs and give the folder clean roles:
- SPEC.md (was SPEC_idle_suspend_lever.md): rewritten from scratch in Russian
as the as-built implementation spec — Down/Up model, reachability walk +
event-driven cache, endpoint-side suspend/wake, the tick bug and its fix
(§11), full test coverage (§12, 29 units named), and what is deliberately
deferred (§13: Tier B netstack teardown, keys-safe BindUpdate path, on-device
battery/heap measurement). Old Tier-A "light sleep" design (never shipped)
removed.
- RESEARCH.md (was SPEC.md): the diagnostic root-cause doc keeps its unique
on-device heap A/B proof (holder = recv-worker bufsArrs) — renamed so its
role (research, not implementation spec) is unambiguous.
- TEST_PLAN_idle_suspend.md: all pass criteria checked, §RESULTS + edge-case
matrix + wake-latency series (cold ~50ms / warm ~36ms, +14-21ms ≈ 1 handshake;
far-server caveat) + production-config run.
Cross-references and section numbers updated across all three files.
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Code recon of the wireguard-go submodule disproved SPEC.md's original PRIMARY
lever (shrink StdNetBind.BatchSize() 128→8). It cannot be done without breaking
GRO receive:
- GRO-rx is ENABLED on android (UDP_GRO set with no android gate,
controlfns_linux.go:90-104; SPEC 010 gated only GSO-tx, not GRO-rx) → rxOffload=true.
- The GRO path splits one coalesced packet into up to 64 datagrams; readAt =
len(msgs) - IdealBatchSize/udpSegmentMaxDatagrams (bind_std.go:269) HARDCODES
IdealBatchSize=128, and getMessages() allocs a 128-slot array. Shrinking bufsArrs
to 8 either desyncs bufs(8) vs array(128) → OOB panic, or overflows the split
("splitting coalesced packet resulted in overflow", bind_std.go:565). GRO can't
be disabled (needed for download throughput, §010).
So the old claim "packet loss excluded, array just shorter" was wrong for the GRO
path. Lever 1 (and lever 2, which inherits the same idle-socket GRO problem) are
rejected. PRIMARY becomes lever 3 — Down idle+unreachable devices: BindClose ends
the recv-workers and frees bufsArrs whole, while the active node keeps batch=128 so
its GRO is intact. The "most expensive fallback" is in fact the only viable lever.
Updated: status line, the lever-candidates section (struck lever 1, promoted lever 3),
the fix-logic section (renamed + rewritten around Down), verification, and residual
risks (the fast-channel risk is gone; the new risk is handshake-on-wake). Implemented
on lx-spec020-idle-suspend; see SPEC_idle_suspend_lever.md §13 + TEST_PLAN.
Add TEST_PLAN_idle_suspend.md — build/config/commands/pass-criteria to
device-verify the shipped idle-suspend on a real run: suspend fires for
idle+unreachable WG/AWG endpoints, reachable ones never suspend, wake-on-dial,
the bufsArrs memory drop (pprof heap), and no flapping. Uses the user's
WARP/AWG + plain-WG nodes. Link it from SPEC_idle_suspend_lever.md §13.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Record the as-built decision so it is not re-derived:
- §13.1 what shipped (c55cf11e): idle-suspend via Down/Up, not light sleep
(source-verified holder is recv-worker bufsArrs; timersStop does not free it).
- §13.2 bind-swap investigation (the promised comment): BindUpdate resizes the
bind WITHOUT zeroing keys; key-zeroing lives only in Down/peer.Stop. So a
keys-safe wake is possible only while still Up — the shipped Down path can't.
- §13.3 the three reduced-bind paths (B=Down shipped / A=BindUpdate keys-safe /
Hybrid) with the GRO-off + max(bind,tun) gotchas.
- §13.4 recommendation: ship B, escalate to A/Hybrid only if device INFO logs
show handshake flapping hurts. Reduced-bind urltest wake deferred (low value
on path B since keys are already zeroed).
docs/ is an upstream-owned tree (it arrives wholesale from SagerNet on
every rebase). Our three downstream docs lived inside it — lx-config.md,
lx-changelog.md, lx-release-runbook.md — mixing fork files into the
upstream surface against CONSTITUTION principle #1 (thin layer / minimal
diff). Move them to a dedicated root-level docs-lx/ so the boundary
between our docs and upstream's is explicit.
- git mv preserves history.
- Updated every reference (docs/lx-* -> docs-lx/lx-*): README.md/.ru.md,
SPECS/{003,004,005,009,020,README}, transport/wireguard/endpoint.go
comments, lx-ci.yml, and lx-release.yml (the release-notes extractor +
fallback URL now read docs-lx/lx-changelog.md).
- Fixed the now-relative links inside the moved files that pointed at
upstream docs/ siblings: lx-config.md -> ../docs/configuration/outbound/
urltest.md; lx-changelog.md -> ../docs/changelog.md (x2).
Verified: all relative + external links resolve, both workflows are valid
YAML, the release-notes awk path is docs-lx/, go vet clean on the touched
package. No release feature — folds into the next tag naturally.
Secondary design doc alongside the authoritative SPEC.md, NOT a replacement.
SPEC.md has on-device proof (heap A/B) that the GC-scan holder is bufsArrs of
recv-workers and its primary lever is shrinking StdNetBind.BatchSize() — that
stays authoritative.
This companion works out SPEC.md's 'lever 3' (suspend inactive devices) in
detail: a light variant (per-peer timersStop, keypairs/socket kept live, cheap
wake without handshake) plus a reachability walk (final + rules + active
selector/pool choices, generation-cached) that decides which devices are idle
AND unreachable from the active routing tree.
Banner up top flags where this doc's source-reading diverged from SPEC.md's
measurements (it guessed gvisor netstack; SPEC.md measured bufsArrs) and notes
light-suspend does NOT free bufsArrs — only Down/BatchSize does. Use as the
fallback design for lever 3, not a competing primary.
Spec only — no code changes.
SPEC 014 dropped with_clash_api because LxBox (Android) drives the core
over the native libbox CommandClient, making the Clash REST server dead
weight in the AAR. But the drop landed in the shared Makefile.lx LX_TAGS,
which also feeds every desktop/CLI release build (mac/windows/linux-musl
via `make -s lx-print-tags`). A CLI binary has no CommandClient channel —
it is managed by external dashboards (yacd/MetaCubeXD) over the Clash REST
API — so every desktop release since rc.1 shipped with no way to manage
the core; a config with experimental.clash_api failed fast. CI stayed
green (lx-ci BASE_TAGS kept the tag), so it was invisible in CI.
Restore with_clash_api to the desktop LX_TAGS; leave build_libbox (AAR)
unchanged. The two tag sets now diverge by design: desktop = with Clash
API, AAR = without.
Verified: desktop binary builds with with_clash_api in Tags; `check`
accepts an experimental.clash_api config; the Clash REST server comes up
live (endpoints answer 401 security-middleware, not the stub's fail-fast).
Docs: Makefile.lx comment, SPEC 014 (§2/§3.1 scoped to AAR + new §3.4),
lx-release.yml tag comment + notes line, changelog rc.17.
§8: validated transport JSON with all 14 new fields at non-default values,
the equivalent flat-camelCase vless:// URL, and a defaults table for the
toUri() omitempty logic. Fixture verified with sing-box check; mirrors
lx-test/config/xhttp_obfs_full.json.
Relocate docs/lx-xhttp-url-parsing.md -> SPECS/002-XHTTP_CLIENT_TRANSPORT/URL_PARSING.md
so all XHTTP docs live together with the spec. Fix internal/back links.
Scanned igareck/vpn-configs-for-russia, extracted+deduped 10 unique XHTTP
nodes, ran each through our with_xhttp binary. 4 alive — all downloaded 1MB,
traffic egressed via the server IP:
- 2x plain -> packet-up
- 2x reality -> stream-one (hu99.bearbeer.digital, bez3.stream-room.com)
The two reality nodes resolve auto->stream-one and work live, closing the
open stream-one live-verification TODO from task 011 (previously synthetic-only).
Other 6 nodes dead for server-side reasons (504, HTTP/1.1-not-H2, reset,
TLS hang) — our transport errored cleanly in every case.
Remaining live TODO: obfs/placement modes (no public node is configured for them).
scMaxConcurrentPosts is a removed Xray knob (grep + GitHub code search
total:0 in current XTLS/Xray-core and sing-box-extended). Current Xray
serializes to one upload POST body in flight at a time, which our sequential
packet-up Write already matches, so the field is accepted for config/link
symmetry but ignored by the client.
- option: V2RayXHTTPOptions.ScMaxConcurrentPosts (json sc_max_concurrent_posts)
- PARAM_MAP: document as legacy/ignore tier with the real concurrency mechanism
(bounded pipe + WroteRequest serialization, server-side seq reorder)
- url-parsing doc: scMaxConcurrentPosts -> accept-but-ignore
- xhttp_obfs_full.json: include the field so check covers it
Verified: build/gofmt/vet clean, 16 unit tests pass, sing-box check passes
on all 3 xhttp configs incl the field.
Implement all 12 client-relevant Xray/sing-box-extended XHTTP params on the
existing lean-native client (no Xray vendoring):
- session/seq placement (path|query|header|cookie) + keys
- uplink-data placement (body|auto|header|cookie, chunked base64) + key + chunk size
- uplink_http_method (upper-cased; GET only in packet-up)
- X-Padding obfs mode: placement (cookie|header|query|queryInHeader) + key/header +
method repeat-x | tokenish (HPACK-Huffman-tuned via golang.org/x/net/http2/hpack)
- packet-up tuning: sc_max_each_post_bytes (split), sc_min_posts_interval_ms (throttle)
4 server-only fields (server_max_header_bytes/no_sse_header/sc_max_buffered_posts/
sc_stream_up_server_secs) accepted but ignored by the client.
New files: transport/v2rayxhttp/{meta.go,xpadding.go}, xhttp_test.go.
Range fields use the "min-max" string form (no badoption.Range in sing).
Default (non-obfs) wire shape kept byte-identical to the live-verified v1
(x_padding='0' in Referer, session/seq on path, payload in body).
Verified: 16/16 unit tests, sing-box check on 3 configs incl full obfs,
go vet/gofmt/build (tagged+untagged) clean, negative (no with_xhttp) rejects.
Adversarial wire-protocol review against PARAM_MAP found no bugs.
Live test of a non-default mode against an Xray server remains an open TODO.
PRIMARY lever spelled out: bufsArrs size = bind.BatchSize() (128 on StdNetBind/android);
shrink it at one point (conn/bind_std.go:322, android branch -> 8/16). BindUpdate
(device.go:558) and getMessages() (bind_std.go:260) follow automatically, both recv
goroutines (v4+v6) covered, no packet loss (array stays full, just shorter), MaxSegmentSize
untouched (GRO intact). Effect 8MB->~0.5MB/recv = 176MB->~11MB at 11 devices. Only risk:
gigabit channel may lose throughput -> fall back to dynamic batch.
On-device throughput A/B (static arm64 curl, download via tunnel):
- baseline batch=128 = 10.7 MB/s median (~86 Mbps), stable 9.1-11.2.
- CPU under load: Syscall6 36%, scanobject 6%, crypto ~3%. The bottleneck is the
WARP channel + syscall overhead, NOT batch processing. GRO/batch only matters at
hundreds-of-Mbps/gigabit, so shrinking batch does NOT cost throughput on a typical
mobile/WARP channel (where the heat is reported).
Re-ranked the levers: PRIMARY is now the global smaller StdNetBind.BatchSize()
(128->8-16) — one point, no activity detection, cuts bufsArrs 8MB->~1MB/recv (176MB->
~11-22MB at 11 devices). Dynamic-batch and Down-idle drop to secondary. Caveat: verify
on a fast Wi-Fi/gigabit channel before release (batch may matter there). Also confirmed
heap scales linearly with live device count (11->269MB, 4->104MB, 1->0). No code changed.
On-device A/B (Debug API /diag/pprof) settles it with high confidence:
- RoutineReceiveIncoming holds 180MB (61% cum); peek = 100% via sync.Pool.Get (in
worker hands, not the pool, not the channels).
- 11 live wireguard endpoints, 22 RoutineReceiveIncoming ALL on StdNetBind (batch=128),
0 on ClientBind. 22 x bufsArrs[128] x 64KB = 176MB ~= 180MB.
- batch=128 because WARP/AWG endpoints use a WireGuardListener dialer => StdNetBind
(endpoint.go:200-202), whose BatchSize()=128 on android.
- A/B: switching the active node WARP->home does NOT free memory (buffers do not sleep);
config 11 ep -> 1 ep gives 269MB -> 0 (= Iliya's workaround, reproduced via profile).
Both prior diagnoses were wrong: batch=1 (no, StdNetBind=128) and drain-on-Suspend of
channels/sync.Pool (misses; bufsArrs of live recv-workers holds it). MaxSegmentSize
2200->65535 is the volume trigger (x30 bytes), not the holder (downLocked/pools/channels/
batch identical 1.13<->1.14).
Fix = shrink batch for INACTIVE devices (naive lazy-bufsArrs impossible: StdNetBind
getMessages() is a fixed 128). Levers + a required download-throughput measurement
documented; pending lever choice.
On the client path BatchSize()=1 (client_bind/stackDevice/systemDevice all return 1),
so the old bufsArrs=128 => 8MB/device claim is wrong; bufsArrs is ~64KB. The 224MB
pprof attributes to PopulatePools is the sync.Pool.New alloc SITE, not the holder.
Real holders: (A) device.pool messageBuffers sync.Pool local+victim cache, (B) the 3
buffered device channels. scanobject 52% comes from the pointer-dense element/container
wrappers (4 of 5 WaitPools are scan-type), not the noscan [65535]byte arrays.
Fix rewritten to drain-on-Suspend: park RoutineReadFromTUN via a stackDevice.suspended
seam in Read (the lockless pool writer surviving Down), drain the 3 device channels
after Down, then ONE runtime.GC()+FreeOSMemory() per selector transition. PopulatePools
swap rejected (WaitPool.count underflow -> cond.Wait deadlock). slim-batch and the
route-graph refcount are dropped (not needed for heat). Added in-repo verification
(device/suspenddrain_test.go + HeapInuse bench) since on-device A/B is impossible.
Android 100% CPU / heat on configs with many WG/AWG endpoints in a
selector. Diagnosed via on-device pprof (§207): scan-bound GC over a
224MB live heap = wireguard-go Device.PopulatePools buffers, held by
~10 idle WG devices. Suspend()=Down() marks the device idle but does
NOT release pools/workers (only Close() does). A/B on device: dropping
spare WG endpoints removes the heat.
SPEC 020: SLIM idle devices (shrink maxBatchSize 128->1-4, do not Close
— keepalives + shared-node safety) gated by a route-reachability
refcount (rules + final + Now, not just selector).
SPEC 010: note our GRO split-brain patch is now upstream-native on
v0.0.3 (commit 24ea133); MaxSegmentSize=65535 must stay (GRO fuel) —
heat is fixed by device count/slimming, never by shrinking the buffer.
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.
1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
The router resolves a domain destination to an IP and overwrites
metadata.Destination before a group's DialContext runs, so destination.Fqdn
is empty when the balancer builds the key. stickyComponent("domain") read
that empty Fqdn, so a single process's key was process+NUL for every site
-> one fixed slot. On device this measured 28/1/1 across a 3-node pool
(uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
back to destination.Fqdn only for a direct dial. After: spread 0.95+.
2. living pool nodes could change slot index during a health-check, moving
sticky keys. balancePoolFirstLive compacted with a filtering append (a
transiently-dead slot shifted every later live node left); planTolerantPool
did delete(inPool, occupant) (an evicted-but-living node re-entered a later
slot, cascading); manual URLTest rebuild ran the tolerant planner even at
pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
dead/empty slots rewritten by index; dedicated planFirstLivePool for the
tolerance==0 rebuild).
3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
(badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
empty array to nil, indistinguishable from omitted, so the default always
applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].
Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
v2 superseded v1; keeping both as separate files (SPEC.md + SPEC_V2.md + the v1
TEST_REPORT) was just confusing. Delete the v1 SPEC and its TEST_REPORT (they remain in
git history) and rename SPEC_V2.md → SPEC.md as the one canonical doc. Drop the "v2"
suffix and stale "design not started" status from the header.
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.
Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:
- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
group -> empty. Additive proto/daemon/libbox, behind with_lx_command.
Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.
Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
Clarify the Now() cold-start tradeoff: variant B (write the fallback node straight
into selectedOutbound*) would eliminate the micro-gap entirely — Now() and DialContext
would read one field, so they can't diverge — at the cost of touching upstream's
selection logic (stub in the choice field + one extra Interrupt() on the first real
switch, which is no worse than any later latency switch). Variant A (Now() stays a
reader) was chosen purely for minimal upstream intrusion; its only cost is a negligible
micro-gap from two separate Select() calls racing on the first seconds. Documents the
path to B if the feature outgrows upstream's selectedOutbound* later.
Doc-only; rc.12 already shipped variant A, no retag.
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.
Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
Live run on 5 vless nodes (3 instances, one per mode): round_robin rotates strictly
across the live set and skips dead nodes; both sticky strategies pin deterministically;
bad config is rejected at start; -race clean on units and live. Feature is now
device-verified, not just isolated.
The run surfaced a config caveat (not a bug, by design): dest_ip is empty until the
destination is resolved, so a sticky key of only source_ip/dest_ip/dest_port collapses
to "" for domain traffic and pins everything to one node. Documented in urltest.md —
use `domain` in `hash` for domain-based traffic.
Add a `mode` to the urltest group so it can distribute traffic instead of only
picking the lowest-delay node, with optional per-flow stickiness.
- mode: least_test (default, unchanged) | round_robin (rotate across live nodes)
| least_connection (reserved, phase 2 — rejected at config time).
- round_robin selects once per connection over the tag-sorted live set (nodes with
a fresh URL-test result supporting the network); UDP/QUIC sessions stay on one
node; first usable outbound is the fallback when nothing is live. The legacy
selectedOutbound* cache path is untouched — balancing is a separate branch in
DialContext/ListenPacket.
- sticky {mode, timeout, cap, hash}: binds one flow to one node. hash components
process|domain|source_ip|dest_ip|dest_port concatenate in order; absent -> "",
all-empty key -> one fixed node (keyless flows never rotate). mode jumphash
(default, stateless consistent hash — ~1/n remap on node-set change) or ttlmap
(key->node table, lazy + ticker eviction, 2000 LRU cap, 10m TTL, dead-node re-pin).
Reuses the existing urltest health ticker/history as the single liveness source;
no new probing. Now() reports the last-picked tag in balanced modes.
Tests (go test -race, 15 cases): distribution, dead-node skip, all-dead fallback,
jumphash stability + empty-key fixed node, ttlmap stick/expire/cap/dead-repick,
key building, validation. The race detector caught a real bug in the sticky
sweeper (read t.ticker unlocked while close() nilled it) — fixed by passing the
channels into the goroutine, mirroring URLTestGroup.loopCheck.
Also folds the SPEC 016 connections-map mutex (ebf9cc07) into the rc.11 changelog
section, which had not yet shipped in a release.
Connections is the client-side CommandConnections accumulator. With 2+
subscribers (LxBox screenClient + profilerClient) one goroutine writes
connectionMap in ApplyEvents while another ranges it in Iterator →
"concurrent map iteration and map write" fatal error → SIGABRT of the
whole process (reproduced in ~20s under traffic, CPH2411/Android15).
Add access sync.Mutex; lock every public method touching
connectionMap/input/filtered: ApplyEvents, FilterState, SortBy*, Iterator.
- FilterState split into public (locks) + private filterState (no lock);
ApplyEvents calls the private one under its already-held lock
(sync.Mutex is not reentrant). Field filterState -> filterStateValue to
free the name for the method.
- evictClosedConnections stays lock-free: private, only called from
ApplyEvents under lock.
- Iterator returns a COPY of filtered — the gomobile caller walks it
after the Go call returns (lock released), so it must not read the
live slice a concurrent ApplyEvents/SortBy is rewriting.
This is the UI/command channel, not the data plane — uncontended lock
~20ns. LxBox per-client accumulators (§170) stay as the consumer scheme;
the mutex is class-correctness insurance against a 3rd consumer.
Verified: go test -race TestConnectionsConcurrentAccess (writer || 3
readers, 2000 rounds) green; go build ./... and -tags with_lx_command
green; gofmt clean.
Codify the rule: before cutting any lx release/prerelease tag, check whether
upstream/testing moved ahead of our last merge and, by default, merge it in
first — then build/gofmt/lx-check, then changelog, then tag.
- docs/lx-release-runbook.md: pre-release gate checklist, drift-check commands,
the manual `git merge upstream/testing` flow (replaces SPECS/004 auto-rebase
while upstream is v1.14.*-alpha), conflict zones (.pb.go, wireguard-go submodule,
build_libbox marker, observability files), and the one-liner sequence.
- SPECS/004 SPEC.md: pointer to the runbook + note that manual merge superseded
auto-rebase on this branch.
LxBox feedback: DnsQuery lacked which DNS server / outbound channel the query went
through. A DNS rule selects a server (matchDNS by action.Server), not an outbound;
the channel is the server's own detour, fixed at config time. Add to DnsQueryEvent:
- dnsServer/dnsServerType = transport.Tag()/Type() (transport is the Exchange param,
so available on all emit paths incl. failures);
- outbound = the server's detour tag (TransportAdapter.OutboundTag() from
DialerOptions.Detour), with a selector expanded to its live node via Now()
server-side (like Connection.Detour), empty on cached/optimistic.
Also gate event construction on HasSubscribers(): with no profiler attached the DNS
hot path builds nothing (no event/answers/outbound lookup) — previously every
resolution built an event just to be dropped for lack of a listener. The Now()
resolution therefore never touches the hot path.
Wire: additive proto fields + OutboundTag() on DNSTransport (embedded adapter
satisfies it). libbox DnsQuery.DNSServer/DNSServerType/Outbound(). Changelog rc.10.