B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.
Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.
With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.
Fix (upstream file, lx:tproxy_writeback_connect):
* CONNECT the write-back socket to the one peer it ever talks to. The kernel's
compute_score() rejects a connected socket for any other peer, and a
connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
reuseport selection outright — so it can no longer be handed a datagram it
will not read. Nothing about the reply changes: same spoofed source, same
single peer, Write instead of WriteTo. The unconnected path is kept verbatim
for a destination that cannot be bound (domain socksaddr).
* A failed cached write now CLOSES the socket instead of only dropping the
reference (upstream left the fd to the GC finalizer).
* TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
but the cache evicts lazily, so after the inbound is gone nothing wakes the
live sessions and each strands its write-back socket. Invisible upstream
(one close at shutdown); on this fork the engine is rebuilt on every apply,
so it was one stranded generation per apply.
Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.
The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.
Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
New balancer flag priority (option.URLTestBalancerOptions.Priority): the
pool is re-derived from CONFIG ORDER every health-check tick via
balancePoolPriority/planPriorityPool — the first live member owns slot 0,
so when the top node answers probes again traffic returns to it on the
next tick (30s failover interval). Probing walks top-down and stops at
the first live node, so the steady-state cost stays one probe per tick.
Replace-in-slot deliberately does not apply here: failover forces sticky
["none"], so relocating nodes across slots breaks no flow keys. Plain
round_robin/random paths are untouched.
failoverBalancer() now emits Priority:true; the KNOWN LIMITATION note and
the panel's "nothing brings it back" blurb are gone because the
limitation is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- random is a REAL urltest mode (lx SPEC 019 v2): uniform draw over LIVE
slots only, pool sized to every member; dead slots keep their place
(never-shrink) but are never picked, for random AND round_robin AND
sticky (degrade-to-live). All-dead pools fall back to Select.
- Globals.SweepInterval + Globals.GroupHealth master switch, resolved by
one pure function (model.SweepSchedule) shared by validator and apply;
unparseable is warned-and-ON, never silently off. ConfigureSweep no
longer resets the cursor on every cron reconcile (release blocker:
a ~6-min cycle was restarted every 60s and never completed).
- multi-WAN egress gateway: ubus netifd status -> uci static -> main
table; a gatewayless non-P2P egress warns CRITICAL instead of silently
blackholing the second uplink.
- endpoint resolver (route.default_domain_resolver): bootstrap-direct
clone of a named resolver, profile override beats globals.
- chains are composable: chain: hops flatten recursively, cycle-guarded,
entry egress lifts only at position 0 (fail-closed mid-path).
- group test publishes its scope so "measuring" lights only the cards a
run covers; health run is explicitly global (all_nodes).
- panel: biased-sample honesty (no ratio until a failure CAN be on
record), profiles auto-pin plate, sweep/GroupHealth settings UI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.
lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
(PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
point every L3-forwarded packet transits, incl. established flows that
bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
sing-tun); wireguard-go stays v0.0.5 with local submodule replace.
Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
Pool() keyed URL-test history by the raw slot tag; history is stored under
RealTag(detour), so a nested-group member always reported Delay=0 in GetPool.
Resolve the slot tag and read under RealTag, matching seedPool/rebuildPool.
SuspendAmneziaWG left idleAsleep untouched, so an endpoint idle-suspended
BEFORE the guard fired could be resurrected by the next dial (resumeOnDial
keys only on idleAsleep) — reintroducing the AmneziaWG-over-WireGuard kernel
hang the guard exists to prevent. Now clears idleAsleep under resumeMu so
the guard is ordered against a concurrent wake.
A failed dial is an actionable error and should be visible where the success
(tunnel established, INFO) is — the LxBox core-log forwarder only surfaces
INFO+, so a DEBUG failure was invisible on-device. WARN makes established/
failed a symmetric, forwardable pair. The detailed dial phases (establishing,
udp-socket-up) stay DEBUG.
Log the tunnel-establish phases so a stuck dial is diagnosable from
/logs/core alone, without a goroutine dump:
- "establishing <h3|h2> tunnel to <server> (sni=...)" on start
- "udp socket up, starting QUIC handshake" (h3) — pinpoints whether a
hang is the socket or the handshake (inbound UDP:443 filtered → our
ClientHello left but no ServerHello came back)
- "tunnel established" on success, "tunnel failed: <err>" on failure
Motivated by a live device case (LxBox §130): h3 hung in the QUIC
handshake because inbound UDP:443 was filtered by the network while AWG
(UDP on a non-443 port) worked — invisible in logs before this, required
a pprof dump to locate. h2 (TCP:443) is the fix there.
Refs: SPEC 021
Rework the tunnel lifecycle around a *session (device + ipConn + closer +
ctx + activity counter), guarded by runMu with a generation guard.
- C1 (was HIGH): a dropped or suspended tunnel is now rebuilt on the next
dial. Previously 'running' latched true, so after the tunnel died every
DialContext short-circuited and dialed into a dead stack — permanent
blackhole. teardownSession clears o.sess so ensureSession rebuilds.
- C2: teardownSession closes ipConn, which unblocks the paired pump parked in
a blocking read (context cancellation alone can't interrupt it); no leaked
goroutine, no zombie half-open tunnel. Idempotent via sync.Once.
- B1: idleWatcher suspends the whole tunnel (gVisor netstack, pumps, QUIC
keepalive) after idle_timeout of no traffic; the next dial rebuilds it —
near-zero resident cost when idle.
- B4: idle_timeout (default 5m) and keep_alive_period (default 30s) are config
options; negative disables. A5: fail-fast mtu<=16000 on h2.
- D2: drop the dead congestion_control option field.
lifecycle_test.go covers the generation guard, idempotent teardown and
close-guard under -race.
Refs: SPEC 021 audit B1/B4/C1/C2/A5/D2
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).
The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.
Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.
Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.
Refs: SPEC 021 TEST_PLAN.md
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.
Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
semantics differ from WARP's expected extended-CONNECT authority/headers
(SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.
Refs: SPEC 021 TEST_PLAN.md
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.
- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.
Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.
1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
The router resolves a domain destination to an IP and overwrites
metadata.Destination before a group's DialContext runs, so destination.Fqdn
is empty when the balancer builds the key. stickyComponent("domain") read
that empty Fqdn, so a single process's key was process+NUL for every site
-> one fixed slot. On device this measured 28/1/1 across a 3-node pool
(uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
back to destination.Fqdn only for a direct dial. After: spread 0.95+.
2. living pool nodes could change slot index during a health-check, moving
sticky keys. balancePoolFirstLive compacted with a filtering append (a
transiently-dead slot shifted every later live node left); planTolerantPool
did delete(inPool, occupant) (an evicted-but-living node re-entered a later
slot, cascading); manual URLTest rebuild ran the tolerant planner even at
pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
dead/empty slots rewritten by index; dedicated planFirstLivePool for the
tolerance==0 rebuild).
3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
(badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
empty array to nil, indistinguishable from omitted, so the default always
applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].
Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.
Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:
- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
group -> empty. Additive proto/daemon/libbox, behind with_lx_command.
Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.
Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.
Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
The lx-ci gofmt-lint step only checks files matching the lx-owned glob
(_xhttp|_awg|_lx.go|_command_lx); the new SPEC 019 files fell outside it. Rename
to the _lx.go convention so CI gofmt-checks them, and fix the changelog reference.
No code change.
Add a `mode` to the urltest group so it can distribute traffic instead of only
picking the lowest-delay node, with optional per-flow stickiness.
- mode: least_test (default, unchanged) | round_robin (rotate across live nodes)
| least_connection (reserved, phase 2 — rejected at config time).
- round_robin selects once per connection over the tag-sorted live set (nodes with
a fresh URL-test result supporting the network); UDP/QUIC sessions stay on one
node; first usable outbound is the fallback when nothing is live. The legacy
selectedOutbound* cache path is untouched — balancing is a separate branch in
DialContext/ListenPacket.
- sticky {mode, timeout, cap, hash}: binds one flow to one node. hash components
process|domain|source_ip|dest_ip|dest_port concatenate in order; absent -> "",
all-empty key -> one fixed node (keyless flows never rotate). mode jumphash
(default, stateless consistent hash — ~1/n remap on node-set change) or ttlmap
(key->node table, lazy + ticker eviction, 2000 LRU cap, 10m TTL, dead-node re-pin).
Reuses the existing urltest health ticker/history as the single liveness source;
no new probing. Now() reports the last-picked tag in balanced modes.
Tests (go test -race, 15 cases): distribution, dead-node skip, all-dead fallback,
jumphash stability + empty-key fixed node, ttlmap stick/expire/cap/dead-repick,
key building, validation. The race detector caught a real bug in the sticky
sweeper (read t.ticker unlocked while close() nilled it) — fixed by passing the
channels into the goroutine, mirroring URLTestGroup.loopCheck.
Also folds the SPEC 016 connections-map mutex (ebf9cc07) into the rc.11 changelog
section, which had not yet shipped in a release.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.