Commit Graph
225 Commits
Author SHA1 Message Date
omarandClaude Opus 5 8fd5c52488 fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.

Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.

With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.

Fix (upstream file, lx:tproxy_writeback_connect):
  * CONNECT the write-back socket to the one peer it ever talks to. The kernel's
    compute_score() rejects a connected socket for any other peer, and a
    connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
    reuseport selection outright — so it can no longer be handed a datagram it
    will not read. Nothing about the reply changes: same spoofed source, same
    single peer, Write instead of WriteTo. The unconnected path is kept verbatim
    for a destination that cannot be bound (domain socksaddr).
  * A failed cached write now CLOSES the socket instead of only dropping the
    reference (upstream left the fd to the GC finalizer).
  * TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
    but the cache evicts lazily, so after the inbound is gone nothing wakes the
    live sessions and each strands its write-back socket. Invisible upstream
    (one close at shutdown); on this fork the engine is rebuilt on every apply,
    so it was one stranded generation per apply.

Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.

The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.

Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:41:02 +03:00
omar a57717dabb health plan S7: docs, contract comments, SPEC 019 update
release / aarch64_cortex-a53 (push) Successful in 3m43s
release / x86_64 (push) Successful in 3m30s
release / apk aarch64_cortex-a53 (push) Successful in 5m11s
release / apk x86_64 (push) Failing after 5m8s
release / release apk (push) Has been skipped
release / release (push) Successful in 12s
- lx-changelog: health board + observatory + global probe + sub cache entry
- DECISIONS.md: D18 board vs delete-and-overlay, D19 observatory vs sweep, D20 global probe settings
- contract comments: urltest.go CheckOutbounds (fast circuit) + observatory.go loop (background circuit + freshness gate) document the two-circuit split
- SPEC 019: dial-error section updated - slots still not moved, but board verdict demotes dead slot on next pick + retry (§5.B); sticky/replace-in-slot/never-shrink invariants preserved
2026-07-24 18:30:48 +03:00
omar 4f0618515e health plan wave 2: alive-only selection+retry, observatory replaces sweep
- S3: Select() alive-only by verdict; dial failure marks board + retries <=3 within ctx; ListenPacket retries to first send; balancer slot liveness reads board verdict; testNodes marks-fail instead of delete; SPEC 019 slot invariants preserved; selector untouched
- S4: new observatory.go (reachability plan from rules, batch<=24/concurrency<=12/timeout 5s, freshness gate, cursor preserved on identical plan); probeplan BuildObservatoryPlan; health.go on board verdicts (TTL=max(3*interval,10min)); engine.dead overlay removed; sweep.go+probeall.go+TestAllNodes+/api/nodes/test removed (->404); GroupHealth.Used published; exit-test extended to chains; panel unused-badge + chain Test button; stats on board
2026-07-24 16:33:28 +03:00
omar fd698162c9 health plan wave 1: urltest health board, global probe settings, sub cache out of UCI
- S1: History gains LastOK/Delay/LastFail; MarkFailed/Verdict on storage (common/urltest/board_lx.go); StoreURLTestHistory preserves LastFail
- S2: per-group ProbeURL/ProbeInterval removed, globals only; sweep_interval drained as dead option; panel fields dropped
- S5: subscription nodes cached in /etc/shater/subs/<name>.json; UCI keeps manual nodes only; sub update writes cache file; panel PUT split; legacy from_sub migration
2026-07-24 13:33:23 +03:00
omarandClaude Fable 5 e60ad231f4 feat(group): failover fail-back — a revived higher-priority node re-takes the slot
New balancer flag priority (option.URLTestBalancerOptions.Priority): the
pool is re-derived from CONFIG ORDER every health-check tick via
balancePoolPriority/planPriorityPool — the first live member owns slot 0,
so when the top node answers probes again traffic returns to it on the
next tick (30s failover interval). Probing walks top-down and stops at
the first live node, so the steady-state cost stays one probe per tick.
Replace-in-slot deliberately does not apply here: failover forces sticky
["none"], so relocating nodes across slots breaks no flow keys. Plain
round_robin/random paths are untouched.

failoverBalancer() now emits Priority:true; the KNOWN LIMITATION note and
the panel's "nothing brings it back" blurb are gone because the
limitation is.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 18:31:17 +03:00
omarandClaude Fable 5 026bba904f feat: real random strategy over the live pool; sweep becomes an honest knob; multi-WAN egress gateway
- random is a REAL urltest mode (lx SPEC 019 v2): uniform draw over LIVE
  slots only, pool sized to every member; dead slots keep their place
  (never-shrink) but are never picked, for random AND round_robin AND
  sticky (degrade-to-live). All-dead pools fall back to Select.
- Globals.SweepInterval + Globals.GroupHealth master switch, resolved by
  one pure function (model.SweepSchedule) shared by validator and apply;
  unparseable is warned-and-ON, never silently off. ConfigureSweep no
  longer resets the cursor on every cron reconcile (release blocker:
  a ~6-min cycle was restarted every 60s and never completed).
- multi-WAN egress gateway: ubus netifd status -> uci static -> main
  table; a gatewayless non-P2P egress warns CRITICAL instead of silently
  blackholing the second uplink.
- endpoint resolver (route.default_domain_resolver): bootstrap-direct
  clone of a named resolver, profile override beats globals.
- chains are composable: chain: hops flatten recursively, cycle-guarded,
  entry egress lifts only at position 0 (fail-closed mid-path).
- group test publishes its scope so "measuring" lights only the cards a
  run covers; health run is explicitly global (all_nodes).
- panel: biased-sample honesty (no ratio until a failure CAN be on
  record), profiles auto-pin plate, sweep/GroupHealth settings UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 18:20:26 +03:00
Leadaxe 9dc758e43c Merge upstream/testing (L3-forwarding, snell, bridge) into lx-1.14
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.

lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
  (PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
  dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
  idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
  point every L3-forwarded packet transits, incl. established flows that
  bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
  restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
  sing-tun); wireguard-go stays v0.0.5 with local submodule replace.

Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
2026-07-08 15:10:38 +03:00
世界 c9690acdf1 Add windows bridge 2026-07-08 18:22:36 +08:00
世界 9ec7cc8cbe Improve bridge 2026-07-08 16:30:06 +08:00
世界 c7fe778cae Add bridge outbound 2026-07-08 00:34:26 +08:00
世界 f2dd4bfd75 Imrpove flow tracking & sniff action 2026-07-07 18:37:18 +08:00
世界 e6419b945f Add L3 forwarding support 2026-07-06 21:13:33 +08:00
世界 5af56d2cfd Add snell protocol 2026-07-05 12:47:42 +08:00
Leadaxe 8e4db706ed lx(urltest): read pool history under RealTag in Pool() (SPEC 022 #5)
Pool() keyed URL-test history by the raw slot tag; history is stored under
RealTag(detour), so a nested-group member always reported Delay=0 in GetPool.
Resolve the slot tag and read under RealTag, matching seedPool/rebuildPool.
2026-07-02 18:17:02 +03:00
Leadaxe c67bb9d8d4 lx(awg): keep guard-suspended AmneziaWG endpoint down (SPEC 022 #2)
SuspendAmneziaWG left idleAsleep untouched, so an endpoint idle-suspended
BEFORE the guard fired could be resurrected by the next dial (resumeOnDial
keys only on idleAsleep) — reintroducing the AmneziaWG-over-WireGuard kernel
hang the guard exists to prevent. Now clears idleAsleep under resumeMu so
the guard is ordered against a concurrent wake.
2026-07-02 18:06:58 +03:00
Leadaxe ad84c06699 fix(masque): log tunnel dial failure at WARN, not DEBUG
A failed dial is an actionable error and should be visible where the success
(tunnel established, INFO) is — the LxBox core-log forwarder only surfaces
INFO+, so a DEBUG failure was invisible on-device. WARN makes established/
failed a symmetric, forwardable pair. The detailed dial phases (establishing,
udp-socket-up) stay DEBUG.
2026-07-02 11:17:59 +03:00
Leadaxe bae1d95333 feat(masque): transport-phase debug logging (dial diagnostics)
Log the tunnel-establish phases so a stuck dial is diagnosable from
/logs/core alone, without a goroutine dump:
- "establishing <h3|h2> tunnel to <server> (sni=...)" on start
- "udp socket up, starting QUIC handshake" (h3) — pinpoints whether a
  hang is the socket or the handshake (inbound UDP:443 filtered → our
  ClientHello left but no ServerHello came back)
- "tunnel established" on success, "tunnel failed: <err>" on failure

Motivated by a live device case (LxBox §130): h3 hung in the QUIC
handshake because inbound UDP:443 was filtered by the network while AWG
(UDP on a non-443 port) worked — invisible in logs before this, required
a pprof dump to locate. h2 (TCP:443) is the fix there.

Refs: SPEC 021
2026-07-02 04:10:53 +03:00
Leadaxe 986d7658f7 feat(masque): self-healing tunnel + idle-suspend (stateless idle)
Rework the tunnel lifecycle around a *session (device + ipConn + closer +
ctx + activity counter), guarded by runMu with a generation guard.

- C1 (was HIGH): a dropped or suspended tunnel is now rebuilt on the next
  dial. Previously 'running' latched true, so after the tunnel died every
  DialContext short-circuited and dialed into a dead stack — permanent
  blackhole. teardownSession clears o.sess so ensureSession rebuilds.
- C2: teardownSession closes ipConn, which unblocks the paired pump parked in
  a blocking read (context cancellation alone can't interrupt it); no leaked
  goroutine, no zombie half-open tunnel. Idempotent via sync.Once.
- B1: idleWatcher suspends the whole tunnel (gVisor netstack, pumps, QUIC
  keepalive) after idle_timeout of no traffic; the next dial rebuilds it —
  near-zero resident cost when idle.
- B4: idle_timeout (default 5m) and keep_alive_period (default 30s) are config
  options; negative disables. A5: fail-fast mtu<=16000 on h2.
- D2: drop the dead congestion_control option field.

lifecycle_test.go covers the generation guard, idempotent teardown and
close-guard under -race.

Refs: SPEC 021 audit B1/B4/C1/C2/A5/D2
2026-07-02 02:38:28 +03:00
Leadaxe ac3d25b845 feat(masque): working h2 CONNECT-IP on WARP via manual x/net/http2 framer
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).

The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.

Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.

Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.

Refs: SPEC 021 TEST_PLAN.md
2026-07-02 01:55:26 +03:00
Leadaxe bd5d1e513b fix(masque): resolve domains before L3 dial; device-verified h3 on WARP
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.

Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
  edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
  semantics differ from WARP's expected extended-CONNECT authority/headers
  (SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.

Refs: SPEC 021 TEST_PLAN.md
2026-07-02 01:04:43 +03:00
Leadaxe 0f41d00ac6 feat(masque): SPEC 021 — MASQUE CONNECT-IP outbound (Cloudflare WARP)
CONNECT-IP (RFC 9484) outbound tunnelling whole IP packets over HTTP/3 and
HTTP/2, targeting Cloudflare WARP. One `type: masque` outbound with a
`profile` field (cloudflare default | standard) and `network` (h3 | h2).

- transport/masque/connectip: vendored connect-ip-go (client subset) ported
  onto sagernet/quic-go — no second quic-go, no external dependency.
- transport/masque: cloudflare/standard profiles (ECDSA pubkey pinning), h3
  ConnectTunnel (Extended CONNECT cf-connect-ip + advertiseDefaultRoute), and
  h2 capsule-DATAGRAM over stdlib net/http (no http fork needed).
- protocol/masque: adapter.Outbound reusing transport/wireguard gVisor
  stackDevice via NewDevice(System:false); lazy tunnel + two IP pumps;
  DialContext/ListenPacket/Close.
- constant/option/include wiring (with_quic + with_gvisor); graceful
  ErrGVisorNotIncluded without gvisor.

Key material (ECDSA priv/pub, ip/ipv6) is taken ready from config — WARP
device registration is done client-side (Dart), out of core scope.

Unit-tested (profiles, TLS pinning, EC key round-trip, capsule round-trip,
IPv4 checksum vector, prefix parsing, config decode via registry). NOT yet
device-verified against live WARP (needs real key material).

Refs: SPEC 021, SagerNet/sing-box#4000
2026-07-02 00:24:00 +03:00
Leadaxe 66116ac089 fix(spec-020): idle tick must iterate the endpoint manager
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.

Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.

Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
2026-07-01 03:08:09 +03:00
Leadaxe 239c0515ad test(spec-020): unit tests for reachability walk, idle logic, event cache 2026-07-01 00:25:04 +03:00
Leadaxe ace08f7aca feat(spec-020): event-driven reachability cache (replace per-tick walk)
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).

- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
  adapter.Router), registered into ctx in box.go, pulled by groups via
  service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
  atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
  recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
  their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
  for next tick rather than being lost, publishes under RWMutex. Starts dirty so
  the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
  legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
  onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.

Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
2026-06-30 22:16:49 +03:00
Leadaxe c55cf11efa feat(spec-020): idle WG/AWG endpoint suspend via Down/Up
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.

- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
  outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
  static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
  ~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
  live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
  distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
  without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
  stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.

Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
2026-06-30 21:21:53 +03:00
Leadaxe 6dc83fc109 Merge upstream/testing (bump version, fix linux ping) into lx-1.14 2026-06-30 01:57:50 +03:00
世界 3d734792ca Fix linux ping 2026-06-29 11:43:26 +08:00
Leadaxe a531879e02 fix(SPEC 019 v2): three sticky/pool bugs found by device verification
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.

1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
   The router resolves a domain destination to an IP and overwrites
   metadata.Destination before a group's DialContext runs, so destination.Fqdn
   is empty when the balancer builds the key. stickyComponent("domain") read
   that empty Fqdn, so a single process's key was process+NUL for every site
   -> one fixed slot. On device this measured 28/1/1 across a 3-node pool
   (uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
   back to destination.Fqdn only for a direct dial. After: spread 0.95+.

2. living pool nodes could change slot index during a health-check, moving
   sticky keys. balancePoolFirstLive compacted with a filtering append (a
   transiently-dead slot shifted every later live node left); planTolerantPool
   did delete(inPool, occupant) (an evicted-but-living node re-entered a later
   slot, cascading); manual URLTest rebuild ran the tolerant planner even at
   pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
   dead/empty slots rewritten by index; dedicated planFirstLivePool for the
   tolerance==0 rebuild).

3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
   (badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
   empty array to nil, indistinguishable from omitted, so the default always
   applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].

Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
2026-06-28 21:48:31 +03:00
Leadaxe 50efd8f335 fix(SPEC 019 v2): balancer.pool 0 = default, not error (rc.14)
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.

Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
2026-06-28 18:11:55 +03:00
Leadaxe 5997b1812a lx(1.14): SPEC 019 v2 — round_robin pool, lazy health-check, slot-hash sticky, GetPool
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:

- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
  replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
  full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
  per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
  found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
  living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
  per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
  from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
  the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
  group -> empty. Additive proto/daemon/libbox, behind with_lx_command.

Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.

Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
2026-06-28 17:50:43 +03:00
Leadaxe feab497fe4 lx(1.14): SPEC 019 Now() cold-start fallback to Select()
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.

Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
2026-06-28 11:00:17 +03:00
Leadaxe ffa3b65427 lx(1.14): rename urltest_balance -> _lx suffix for CI gofmt coverage
The lx-ci gofmt-lint step only checks files matching the lx-owned glob
(_xhttp|_awg|_lx.go|_command_lx); the new SPEC 019 files fell outside it. Rename
to the _lx.go convention so CI gofmt-checks them, and fix the changelog reference.
No code change.
2026-06-28 01:15:58 +03:00
Leadaxe 5ebff914fc lx(1.14): SPEC 019 urltest mode + sticky load-balancing
Add a `mode` to the urltest group so it can distribute traffic instead of only
picking the lowest-delay node, with optional per-flow stickiness.

- mode: least_test (default, unchanged) | round_robin (rotate across live nodes)
  | least_connection (reserved, phase 2 — rejected at config time).
- round_robin selects once per connection over the tag-sorted live set (nodes with
  a fresh URL-test result supporting the network); UDP/QUIC sessions stay on one
  node; first usable outbound is the fallback when nothing is live. The legacy
  selectedOutbound* cache path is untouched — balancing is a separate branch in
  DialContext/ListenPacket.
- sticky {mode, timeout, cap, hash}: binds one flow to one node. hash components
  process|domain|source_ip|dest_ip|dest_port concatenate in order; absent -> "",
  all-empty key -> one fixed node (keyless flows never rotate). mode jumphash
  (default, stateless consistent hash — ~1/n remap on node-set change) or ttlmap
  (key->node table, lazy + ticker eviction, 2000 LRU cap, 10m TTL, dead-node re-pin).

Reuses the existing urltest health ticker/history as the single liveness source;
no new probing. Now() reports the last-picked tag in balanced modes.

Tests (go test -race, 15 cases): distribution, dead-node skip, all-dead fallback,
jumphash stability + empty-key fixed node, ttlmap stick/expire/cap/dead-repick,
key building, validation. The race detector caught a real bug in the sticky
sweeper (read t.ticker unlocked while close() nilled it) — fixed by passing the
channels into the goroutine, mirroring URLTestGroup.loopCheck.

Also folds the SPEC 016 connections-map mutex (ebf9cc07) into the rc.11 changelog
section, which had not yet shipped in a release.
2026-06-28 01:10:17 +03:00
世界 b010759f45 Fix Cloudflared edge discovery ignoring configured resolver 2026-06-25 17:39:15 +08:00
世界 7e96370229 Fix group status updates broken by API service
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.

Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
2026-06-25 17:38:54 +08:00
世界 65b90e6f4b tailscale: Fix auth URL not refreshed after logout 2026-06-25 17:38:53 +08:00
世界 c2da94f252 platform: Add shell support for iOS 2026-06-25 17:38:52 +08:00
世界 dcbb8271f8 platform: Add tailscale device name and logout 2026-06-25 17:38:52 +08:00
世界 1ee172f6af tailssh: fix platform SFTP session teardown 2026-06-25 17:38:52 +08:00
世界 f18b01b7fe tailscale: Add tailssh server 2026-06-25 17:38:51 +08:00
世界 d0ef5c028d hysteria2: Add gecko obfs 2026-06-25 17:38:51 +08:00
世界 db0e33d74b Fix tailscale dns 2026-06-25 17:38:28 +08:00
世界 e97a1bec41 tailscale: Expose more peer info fields 2026-06-25 17:38:27 +08:00
世界 03499260df tailscale: Add runtime exit node API 2026-06-25 17:38:27 +08:00
世界 5d0b87d9bf tailscale: Fix handle peer DNS query 2026-06-25 17:38:26 +08:00
世界 23a676ac8a tailscale: Revert dialer deprecation and remove control_http_client 2026-06-25 17:38:25 +08:00
世界 4b55717612 Fix hysteria2 realm server 2026-06-25 17:38:12 +08:00
世界 d8f461cc5a Update hysteria2 realm 2026-06-25 17:38:12 +08:00
世界 094808a4fa Add hysteria2 realm service and support 2026-06-25 17:38:11 +08:00
世界 e8643485d2 Allow customizing TUN DNS mode and hijack interface DNS by default 2026-06-25 17:38:10 +08:00