Commit Graph
2723 Commits
Author SHA1 Message Date
Leadaxe c4fd73cdce ci(lx-release): fetch upstream alpha tags so git-describe base works in CI
The base-version derivation used `git describe --match v1.14.0-alpha.*` as the
primary source, but actions/checkout only fetches THIS repo's tags — the alpha
tags are SagerNet/sing-box (upstream) tags, absent in the CI clone. So `git
describe` found nothing and silently fell to the subject-grep fallback, which
resolves alpha.36 (alpha.37 was merged in a commit whose subject omits the
number). That's why rc.17/rc.18 notes shipped "base alpha.36" while a local
clone with upstream tags gets 37.

Fetch just the upstream v1.14.0-alpha.* tags before git describe so the primary
graph-based path works in CI. Both hand-fixed on the published releases; this
makes the next tag correct automatically.
2026-07-01 10:18:32 +03:00
Leadaxe 3ae3bbf203 docs(spec-020): ANDROID_RESEARCH — on-device idle-suspend report + artifacts
Full write-up of the Android device run (CPH2411, Android 15, rc.18) in a
dedicated ANDROID_RESEARCH/ subfolder: README (report), METHOD (reproducible
procedure), RESULTS (per-scenario + heap A/B), and artifacts/ (raw evidence:
lx idle log lines, goroutine dumps, pprof heap .pb + top renders). No access
credentials anywhere.

Headline, now measured on the target platform: PopulatePools.func3 (the
bufsArrs holder from RESEARCH.md) inuse_space 223.93 -> 89.89 MB (-134 MB /
-60%), recv-workers 18 -> 2, on suspending 8 of 9 WG endpoints. = 16 workers x
~8.4 MB (BatchSize=128), matching the source model, ~10x the desktop RSS delta.

RESEARCH.md status + SPEC.md §12/§13 updated: the Android heap A/B gap is
closed (only the battery A/B remains deferred).
2026-07-01 09:22:30 +03:00
Leadaxe 23193f5408 docs(spec-020): Android device RESULTS — heap A/B closes RESEARCH gap
Device run on CPH2411 (Android 15, rc.18) via the LxBox app Debug API.
9 WG endpoints (1 real WARP reachable + 8 synthetic unreachable),
lx_idle_suspend=30s. All behaviors confirmed on-device: suspend fires,
reachable final stays up, wake-by-dial, no-flap, kill-switch.

Headline: PopulatePools.func3 (the bufsArrs holder from RESEARCH.md)
inuse_space 223.93 to 89.89 MB (-134 MB / -60%), recv-workers 18 to 2.
= 16 freed workers x ~8.4 MB (BatchSize=128), matching the model, ~10x
the desktop RSS delta. Closes the Android device-verification gap.
2026-07-01 04:08:01 +03:00
Leadaxe 702f4dc0d4 docs(lx-changelog): rc.18 — SPEC 020 idle-suspend for WG/AWG endpoints
Device-verified idle-suspend: idle AND unreachable WG/AWG endpoints go Down to
free their recv-worker bufsArrs (the Android GC-heat holder), waking on next dial.
Opt-in via route.lx_idle_suspend; off by default. Notes cover the reachability
walk, the GRO reason for Down-over-smaller-batch, the concurrency fixes, and the
2026-07-01 live-run results (recv-workers 16→0, RSS −31%).
2026-07-01 03:16:16 +03:00
Leadaxe 937fc7f498 Merge SPEC 020: idle-suspend for WG/AWG endpoints (Down/Up)
Selectively brings idle + unreachable WireGuard/AmneziaWG endpoints Down,
freeing their recv-worker bufsArrs (the measured GC-scan heat holder on
Android) and stopping their per-peer timers (battery), then wakes them lazily
on the next dial. Off by default (lx_idle_suspend absent/0 = zero overhead).

Includes the fix for the shipped tick iterating the wrong manager (it never
reached any endpoint — the feature was inert on a live box), the full
reachability walk (final/rule/selector Now/urltest pool/detour, event-driven
cached), and the rewritten as-built spec (SPEC.md) + research doc (RESEARCH.md)
+ test plan.

Live-verified: suspend/wake/probe-wake/re-sleep/no-flap/kill-switch across
selector, urltest pool, nested groups, AWG-guard, and the real production
config; resource A/B recv-workers 16->0, RSS -31% on desktop. 29 unit tests,
adversarially checked.
2026-07-01 03:13:28 +03:00
Leadaxe d70d0d26b2 docs(spec-020): rewrite as-built spec, split research from implementation
The idle-suspend feature is implemented, the tick bug is fixed, and every
reachability node type plus suspend/wake/probe/no-flap/kill-switch and the
resource A/B (recv-workers 16→0, RSS -31%) are live-verified. Reflect that in
the docs and give the folder clean roles:

- SPEC.md (was SPEC_idle_suspend_lever.md): rewritten from scratch in Russian
  as the as-built implementation spec — Down/Up model, reachability walk +
  event-driven cache, endpoint-side suspend/wake, the tick bug and its fix
  (§11), full test coverage (§12, 29 units named), and what is deliberately
  deferred (§13: Tier B netstack teardown, keys-safe BindUpdate path, on-device
  battery/heap measurement). Old Tier-A "light sleep" design (never shipped)
  removed.
- RESEARCH.md (was SPEC.md): the diagnostic root-cause doc keeps its unique
  on-device heap A/B proof (holder = recv-worker bufsArrs) — renamed so its
  role (research, not implementation spec) is unambiguous.
- TEST_PLAN_idle_suspend.md: all pass criteria checked, §RESULTS + edge-case
  matrix + wake-latency series (cold ~50ms / warm ~36ms, +14-21ms ≈ 1 handshake;
  far-server caveat) + production-config run.

Cross-references and section numbers updated across all three files.
2026-07-01 03:08:23 +03:00
Leadaxe 66116ac089 fix(spec-020): idle tick must iterate the endpoint manager
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.

Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.

Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
2026-07-01 03:08:09 +03:00
Leadaxe c341d1286c docs(spec-020): reject lever 1 (smaller BatchSize) — breaks GRO; PRIMARY = Down (lever 3)
Code recon of the wireguard-go submodule disproved SPEC.md's original PRIMARY
lever (shrink StdNetBind.BatchSize() 128→8). It cannot be done without breaking
GRO receive:
- GRO-rx is ENABLED on android (UDP_GRO set with no android gate,
  controlfns_linux.go:90-104; SPEC 010 gated only GSO-tx, not GRO-rx) → rxOffload=true.
- The GRO path splits one coalesced packet into up to 64 datagrams; readAt =
  len(msgs) - IdealBatchSize/udpSegmentMaxDatagrams (bind_std.go:269) HARDCODES
  IdealBatchSize=128, and getMessages() allocs a 128-slot array. Shrinking bufsArrs
  to 8 either desyncs bufs(8) vs array(128) → OOB panic, or overflows the split
  ("splitting coalesced packet resulted in overflow", bind_std.go:565). GRO can't
  be disabled (needed for download throughput, §010).

So the old claim "packet loss excluded, array just shorter" was wrong for the GRO
path. Lever 1 (and lever 2, which inherits the same idle-socket GRO problem) are
rejected. PRIMARY becomes lever 3 — Down idle+unreachable devices: BindClose ends
the recv-workers and frees bufsArrs whole, while the active node keeps batch=128 so
its GRO is intact. The "most expensive fallback" is in fact the only viable lever.

Updated: status line, the lever-candidates section (struck lever 1, promoted lever 3),
the fix-logic section (renamed + rewritten around Down), verification, and residual
risks (the fast-channel risk is gone; the new risk is handshake-on-wake). Implemented
on lx-spec020-idle-suspend; see SPEC_idle_suspend_lever.md §13 + TEST_PLAN.
2026-07-01 00:50:59 +03:00
Leadaxe fc085bc838 docs(spec-020): live test plan for idle-suspend (Down/Up) + link from lever doc
Add TEST_PLAN_idle_suspend.md — build/config/commands/pass-criteria to
device-verify the shipped idle-suspend on a real run: suspend fires for
idle+unreachable WG/AWG endpoints, reachable ones never suspend, wake-on-dial,
the bufsArrs memory drop (pprof heap), and no flapping. Uses the user's
WARP/AWG + plain-WG nodes. Link it from SPEC_idle_suspend_lever.md §13.
2026-07-01 00:33:36 +03:00
Leadaxe 239c0515ad test(spec-020): unit tests for reachability walk, idle logic, event cache 2026-07-01 00:25:04 +03:00
Leadaxe ace08f7aca feat(spec-020): event-driven reachability cache (replace per-tick walk)
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).

- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
  adapter.Router), registered into ctx in box.go, pulled by groups via
  service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
  atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
  recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
  their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
  for next tick rather than being lost, publishes under RWMutex. Starts dirty so
  the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
  legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
  onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.

Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
2026-06-30 22:16:49 +03:00
Leadaxe 4dd9b4ffc9 docs(spec-020): implementation log §13 — Down shipped, bind-swap/keys, 3 paths
Record the as-built decision so it is not re-derived:
- §13.1 what shipped (c55cf11e): idle-suspend via Down/Up, not light sleep
  (source-verified holder is recv-worker bufsArrs; timersStop does not free it).
- §13.2 bind-swap investigation (the promised comment): BindUpdate resizes the
  bind WITHOUT zeroing keys; key-zeroing lives only in Down/peer.Stop. So a
  keys-safe wake is possible only while still Up — the shipped Down path can't.
- §13.3 the three reduced-bind paths (B=Down shipped / A=BindUpdate keys-safe /
  Hybrid) with the GRO-off + max(bind,tun) gotchas.
- §13.4 recommendation: ship B, escalate to A/Hybrid only if device INFO logs
  show handshake flapping hurts. Reduced-bind urltest wake deferred (low value
  on path B since keys are already zeroed).
2026-06-30 21:53:46 +03:00
Leadaxe c55cf11efa feat(spec-020): idle WG/AWG endpoint suspend via Down/Up
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.

- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
  outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
  static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
  ~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
  live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
  distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
  without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
  stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.

Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
2026-06-30 21:21:53 +03:00
Leadaxe dcc964dd9d docs(lx-config): add full Russian translation (lx-config.ru.md)
Complete RU translation of docs-lx/lx-config.md, mirroring its structure:
the §0 exhaustive "every field at a glance" example, all per-section field
tables (XHTTP v1+v2, AmneziaWG 2.0, id/ip/ib masquerade, urltest balancer),
examples and the build section. JSONC code is preserved; only prose and
inline comments are translated. Intra-doc #anchors are re-pointed to the
Russian heading slugs (all 6 verified to resolve); the §0 example validates
as JSON.

Cross-link both ways (en ↔ ru) and re-point README.ru.md's four lx-config
links to the Russian version. File name follows the README.ru.md convention
(.ru.md, not -ru.md).
2026-06-30 20:19:52 +03:00
Leadaxe 0134fcfc45 docs(lx-config): exhaustive "every field at a glance" example + fill XHTTP v2 fields
Add a §0 kitchen-sink config carrying ALL 52 lx-added fields in one place —
XHTTP transport (26), AmneziaWG 2.0 endpoint incl. id/ip/ib (21), urltest
round_robin balancer (5) — each with its default and allowed values inline,
and mutually-exclusive / server-ignored fields flagged. Sourced by reading
option/*.go directly (not the prior doc), so it is complete.

This also surfaced that §1 documented only 7 of the 26 XHTTP fields (the v1
set); fill in the 19 missing v2 fields (session/seq placement, uplink-data
placement, X-Padding obfs family, packet-up tuning, accepted-but-ignored)
as grouped tables. Fix three code-vs-doc disagreements the extraction found:
- `mode: auto` resolves to stream-one on Reality (not always packet-up);
- s3/s4 are AWG 2.0 junk-size params (not "cookie-reply/transport" junk);
- h1-h4 unset spelling includes "" as well as 0; x_padding_bytes framing.

The default wire shape is unchanged; all v2 fields are opt-in. §0 example
validated as JSON. Russian translation (lx-config.ru.md) to follow.
2026-06-30 20:12:39 +03:00
Leadaxe 38bda198ff docs: move lx docs out of upstream docs/ into docs-lx/
docs/ is an upstream-owned tree (it arrives wholesale from SagerNet on
every rebase). Our three downstream docs lived inside it — lx-config.md,
lx-changelog.md, lx-release-runbook.md — mixing fork files into the
upstream surface against CONSTITUTION principle #1 (thin layer / minimal
diff). Move them to a dedicated root-level docs-lx/ so the boundary
between our docs and upstream's is explicit.

- git mv preserves history.
- Updated every reference (docs/lx-* -> docs-lx/lx-*): README.md/.ru.md,
  SPECS/{003,004,005,009,020,README}, transport/wireguard/endpoint.go
  comments, lx-ci.yml, and lx-release.yml (the release-notes extractor +
  fallback URL now read docs-lx/lx-changelog.md).
- Fixed the now-relative links inside the moved files that pointed at
  upstream docs/ siblings: lx-config.md -> ../docs/configuration/outbound/
  urltest.md; lx-changelog.md -> ../docs/changelog.md (x2).

Verified: all relative + external links resolve, both workflows are valid
YAML, the release-notes awk path is docs-lx/, go vet clean on the touched
package. No release feature — folds into the next tag naturally.
2026-06-30 18:55:26 +03:00
Leadaxe ac94dcd4b6 docs(spec-020): add companion design for lever 3 (idle WG light-suspend)
Secondary design doc alongside the authoritative SPEC.md, NOT a replacement.
SPEC.md has on-device proof (heap A/B) that the GC-scan holder is bufsArrs of
recv-workers and its primary lever is shrinking StdNetBind.BatchSize() — that
stays authoritative.

This companion works out SPEC.md's 'lever 3' (suspend inactive devices) in
detail: a light variant (per-peer timersStop, keypairs/socket kept live, cheap
wake without handshake) plus a reachability walk (final + rules + active
selector/pool choices, generation-cached) that decides which devices are idle
AND unreachable from the active routing tree.

Banner up top flags where this doc's source-reading diverged from SPEC.md's
measurements (it guessed gvisor netstack; SPEC.md measured bufsArrs) and notes
light-suspend does NOT free bufsArrs — only Down/BatchSize does. Use as the
fallback design for lever 3, not a competing primary.

Spec only — no code changes.
2026-06-30 17:58:41 +03:00
Leadaxe e7b0bcdc75 ci(lx-release): derive base from git graph, not merge-subject text
The subject-grep base-detection missed alpha.37: it was merged in a commit
titled "Merge upstream/testing (bump version, fix linux ping)" with no
"alpha.37" in the subject (upstream tagged it after we merged), so the grep
found only alpha.36 and rc.17 notes shipped a stale base.

Make `git describe --match v1.14.0-alpha.*` the primary source — it reads
HEAD's ancestry in the commit graph, independent of merge-message wording —
and keep the subject-grep as the fallback for a fork checkout without
upstream tags. Verified locally: now resolves v1.14.0-alpha.37.
2026-06-30 13:38:04 +03:00
Leadaxe c35ccd0ea7 fix(build): restore with_clash_api on desktop/CLI — drop is AAR-only
SPEC 014 dropped with_clash_api because LxBox (Android) drives the core
over the native libbox CommandClient, making the Clash REST server dead
weight in the AAR. But the drop landed in the shared Makefile.lx LX_TAGS,
which also feeds every desktop/CLI release build (mac/windows/linux-musl
via `make -s lx-print-tags`). A CLI binary has no CommandClient channel —
it is managed by external dashboards (yacd/MetaCubeXD) over the Clash REST
API — so every desktop release since rc.1 shipped with no way to manage
the core; a config with experimental.clash_api failed fast. CI stayed
green (lx-ci BASE_TAGS kept the tag), so it was invisible in CI.

Restore with_clash_api to the desktop LX_TAGS; leave build_libbox (AAR)
unchanged. The two tag sets now diverge by design: desktop = with Clash
API, AAR = without.

Verified: desktop binary builds with with_clash_api in Tags; `check`
accepts an experimental.clash_api config; the Clash REST server comes up
live (endpoints answer 401 security-middleware, not the stub's fail-fast).

Docs: Makefile.lx comment, SPEC 014 (§2/§3.1 scoped to AAR + new §3.4),
lx-release.yml tag comment + notes line, changelog rc.17.
2026-06-30 13:23:53 +03:00
Leadaxe 51b3bd07a8 ci(lx-release): derive upstream base version for notes, stop hardcoding alpha.35
The release-notes template hardcoded "base v1.14.0-alpha.35"; it went stale and
had to be hand-edited on rc.14, rc.15 and rc.16 (each was actually on alpha.36).
Resolve the base dynamically in the "Resolve tag" step: take the highest alpha.NN
named in any "Merge upstream" commit subject (robust on a fork without upstream
tags fetched), falling back to git describe against upstream alpha tags, then a
generic v1.14.x label. The notes line now interpolates steps.ver.outputs.base.
2026-06-30 11:04:22 +03:00
Leadaxe e0f8d6d9bb docs(xhttp): add golden fixture to URL_PARSING for Android round-trip tests
§8: validated transport JSON with all 14 new fields at non-default values,
the equivalent flat-camelCase vless:// URL, and a defaults table for the
toUri() omitempty logic. Fixture verified with sing-box check; mirrors
lx-test/config/xhttp_obfs_full.json.
2026-06-30 02:15:37 +03:00
Leadaxe ab29eb1bb0 docs(lx-changelog): rc.16 — SPEC 002 v2 full XHTTP param support 2026-06-30 01:59:57 +03:00
Leadaxe 6dc83fc109 Merge upstream/testing (bump version, fix linux ping) into lx-1.14 2026-06-30 01:57:50 +03:00
Leadaxe 8851eec001 merge: SPEC 002 v2 — full client-side XHTTP param support
Merge lx-1.14-xhttp-full: full client-side support of extended XHTTP params
(session/seq/uplink-data placement+keys, uplink method, X-Padding obfs incl
tokenish/HPACK), accept-but-ignore for server-only + legacy scMaxConcurrentPosts,
URL-parsing guide for Android/launcher teams.

Verified: 16 unit tests, sing-box check on 3 configs, build/vet/gofmt clean,
default path live-confirmed on 4 real XHTTP nodes (packet-up + stream-one/reality,
2026-06-30 01:19:56 +03:00
Leadaxe 1330837400 docs(xhttp): move URL parsing guide into SPEC 002 folder
Relocate docs/lx-xhttp-url-parsing.md -> SPECS/002-XHTTP_CLIENT_TRANSPORT/URL_PARSING.md
so all XHTTP docs live together with the spec. Fix internal/back links.
2026-06-30 01:19:27 +03:00
Leadaxe 61c9301d13 docs(xhttp): live-verify default path on 4 real nodes (stream-one TODO closed)
Scanned igareck/vpn-configs-for-russia, extracted+deduped 10 unique XHTTP
nodes, ran each through our with_xhttp binary. 4 alive — all downloaded 1MB,
traffic egressed via the server IP:
- 2x plain  -> packet-up
- 2x reality -> stream-one (hu99.bearbeer.digital, bez3.stream-room.com)

The two reality nodes resolve auto->stream-one and work live, closing the
open stream-one live-verification TODO from task 011 (previously synthetic-only).
Other 6 nodes dead for server-side reasons (504, HTTP/1.1-not-H2, reset,
TLS hang) — our transport errored cleanly in every case.

Remaining live TODO: obfs/placement modes (no public node is configured for them).
2026-06-30 00:42:16 +03:00
Leadaxe c3e54a8141 docs(xhttp): note HTTP/3 (alpn=h3) limitation in URL parsing guide
Our XHTTP client is HTTP/2 only (http2.Transport); Xray supports H1/H2/H3.
h3-only nodes won't connect — flag for the link parser. Out of SPEC 002 scope
(separate future 'XHTTP over HTTP/3' task).
2026-06-30 00:36:06 +03:00
Leadaxe 7527467965 feat(xhttp): accept sc_max_concurrent_posts (legacy, ignored)
scMaxConcurrentPosts is a removed Xray knob (grep + GitHub code search
total:0 in current XTLS/Xray-core and sing-box-extended). Current Xray
serializes to one upload POST body in flight at a time, which our sequential
packet-up Write already matches, so the field is accepted for config/link
symmetry but ignored by the client.

- option: V2RayXHTTPOptions.ScMaxConcurrentPosts (json sc_max_concurrent_posts)
- PARAM_MAP: document as legacy/ignore tier with the real concurrency mechanism
  (bounded pipe + WroteRequest serialization, server-side seq reorder)
- url-parsing doc: scMaxConcurrentPosts -> accept-but-ignore
- xhttp_obfs_full.json: include the field so check covers it

Verified: build/gofmt/vet clean, 16 unit tests pass, sing-box check passes
on all 3 xhttp configs incl the field.
2026-06-30 00:35:31 +03:00
Leadaxe 31926aebd4 docs(xhttp): vless:// XHTTP URL parsing guide for Android/launcher teams
Self-contained reference for the link parser: maps every vless://...type=xhttp
URL param (flat query + extra={...} JSON) to sing-box transport snake_case fields.
Covers TLS/Reality mapping, the extra-JSON number→"min-max" coercion, mode=auto
pass-through, path-with-query-tail, and ignored fields (scMaxConcurrentPosts,
server-only). Examples validated with sing-box check.
2026-06-29 15:39:08 +03:00
Leadaxe 33ee291b17 feat(xhttp): full client-side support of extended XHTTP params (SPEC 002 v2)
Implement all 12 client-relevant Xray/sing-box-extended XHTTP params on the
existing lean-native client (no Xray vendoring):

- session/seq placement (path|query|header|cookie) + keys
- uplink-data placement (body|auto|header|cookie, chunked base64) + key + chunk size
- uplink_http_method (upper-cased; GET only in packet-up)
- X-Padding obfs mode: placement (cookie|header|query|queryInHeader) + key/header +
  method repeat-x | tokenish (HPACK-Huffman-tuned via golang.org/x/net/http2/hpack)
- packet-up tuning: sc_max_each_post_bytes (split), sc_min_posts_interval_ms (throttle)

4 server-only fields (server_max_header_bytes/no_sse_header/sc_max_buffered_posts/
sc_stream_up_server_secs) accepted but ignored by the client.

New files: transport/v2rayxhttp/{meta.go,xpadding.go}, xhttp_test.go.
Range fields use the "min-max" string form (no badoption.Range in sing).
Default (non-obfs) wire shape kept byte-identical to the live-verified v1
(x_padding='0' in Referer, session/seq on path, payload in body).

Verified: 16/16 unit tests, sing-box check on 3 configs incl full obfs,
go vet/gofmt/build (tagged+untagged) clean, negative (no with_xhttp) rejects.
Adversarial wire-protocol review against PARAM_MAP found no bugs.
Live test of a non-default mode against an Xray server remains an open TODO.
2026-06-29 15:23:57 +03:00
Leadaxe cafbe54603 docs(SPEC 002): full XHTTP param spec + PARAM_MAP (audit of Xray-core + sing-box-extended)
- Rename minimal v1 spec to SPEC_v1.md (history kept)
- New SPEC.md: full client-side support of 12 client-relevant XHTTP params
  (session/seq/uplink-data placement+keys, uplink method, X-Padding obfs incl tokenish)
- PARAM_MAP.md: per-param map of all 16 NekoBox+ fields (+2 client tuning extras),
  client-vs-server verdict, defaults, on-wire effect, impl notes
- 4 server-only fields (serverMaxHeaderBytes/noSSEHeader/scMaxBufferedPosts/
  scStreamUpServerSecs) documented as accept-but-ignore
- Verification corrections folded in (no Access-Control-Allow-Credentials;
  mode-gates + method upper-case live in extended option layer only)
2026-06-29 15:02:47 +03:00
世界 b2d51f6ec3 Bump version 2026-06-29 11:43:26 +08:00
世界 3d734792ca Fix linux ping 2026-06-29 11:43:26 +08:00
Leadaxe a26f2a6c3f docs: mark round_robin (SPEC 019) device-verified end to end
The rc.15 domain fix was confirmed on a real device: with the default
sticky_hash ["process","domain"] and no dest_ip workaround, browser traffic
spreads across the pool (on-device per-domain uniformity ~0.27 -> 0.95+).
Update README + lx-config.md status from "not yet device-verified" to
device-verified.
2026-06-29 02:43:41 +03:00
Leadaxe 6ad8ec91fb docs(SPEC 020): add fix logic — single-point StdNetBind.BatchSize() shrink on android
PRIMARY lever spelled out: bufsArrs size = bind.BatchSize() (128 on StdNetBind/android);
shrink it at one point (conn/bind_std.go:322, android branch -> 8/16). BindUpdate
(device.go:558) and getMessages() (bind_std.go:260) follow automatically, both recv
goroutines (v4+v6) covered, no packet loss (array stays full, just shorter), MaxSegmentSize
untouched (GRO intact). Effect 8MB->~0.5MB/recv = 176MB->~11MB at 11 devices. Only risk:
gigabit channel may lose throughput -> fall back to dynamic batch.
2026-06-29 02:43:09 +03:00
Leadaxe 3ef5717890 docs(SPEC 020): throughput measured — batch not the limit on mobile/WARP, primary lever = global smaller BatchSize
On-device throughput A/B (static arm64 curl, download via tunnel):
- baseline batch=128 = 10.7 MB/s median (~86 Mbps), stable 9.1-11.2.
- CPU under load: Syscall6 36%, scanobject 6%, crypto ~3%. The bottleneck is the
  WARP channel + syscall overhead, NOT batch processing. GRO/batch only matters at
  hundreds-of-Mbps/gigabit, so shrinking batch does NOT cost throughput on a typical
  mobile/WARP channel (where the heat is reported).

Re-ranked the levers: PRIMARY is now the global smaller StdNetBind.BatchSize()
(128->8-16) — one point, no activity detection, cuts bufsArrs 8MB->~1MB/recv (176MB->
~11-22MB at 11 devices). Dynamic-batch and Down-idle drop to secondary. Caveat: verify
on a fast Wi-Fi/gigabit channel before release (batch may matter there). Also confirmed
heap scales linearly with live device count (11->269MB, 4->104MB, 1->0). No code changed.
2026-06-29 02:25:55 +03:00
Leadaxe b28c689cb9 docs: document observability (SPEC 014-018) + round_robin (SPEC 019) across lx docs
The lx feature docs had drifted: README (en/ru) and docs/lx-config.md still said
"currently XHTTP + AWG2" and covered only SPEC 002/003/009 — the observability
layer (SPEC 014-018) and round_robin load balancing (SPEC 019) were undocumented
in the lx overview, and urltest.md still described the pre-rc.15 domain behaviour.

- docs/lx-config.md: new "## 3. round_robin load balancing" (mode/balancer,
  pool/pool_tolerance/sticky_hash, ["none"] sentinel + badjson-[] caveat, slot-hash
  binding, example, status) and "## 4. Observability (CommandClient extensions)"
  (URLTestOutbound/GetRules/GetGroups/GetOutbounds/GetPool/SubscribeDNSQueries +
  Connection.detourList, all behind with_lx_command); Validate&build -> ## 5.
- README.md / README.ru.md: broaden the stale "XHTTP + AWG2" framing; add feature
  rows for observability and round_robin with honest status.
- docs/configuration/outbound/urltest.md: reconcile sticky_hash "domain" with the
  rc.15 fix — domain reads metadata.Domain (survives domain->IP resolve), so it
  works for normal sniffed domain traffic, not only literal-IP destinations; the
  warning is reframed (domain works; dest_ip is an alternative).

Docs-only; no code change.
2026-06-29 02:23:49 +03:00
Leadaxe 3ceca12b98 docs(SPEC 020): holder PROVEN on repro — bufsArrs of live recv-workers, not channels
On-device A/B (Debug API /diag/pprof) settles it with high confidence:
- RoutineReceiveIncoming holds 180MB (61% cum); peek = 100% via sync.Pool.Get (in
  worker hands, not the pool, not the channels).
- 11 live wireguard endpoints, 22 RoutineReceiveIncoming ALL on StdNetBind (batch=128),
  0 on ClientBind. 22 x bufsArrs[128] x 64KB = 176MB ~= 180MB.
- batch=128 because WARP/AWG endpoints use a WireGuardListener dialer => StdNetBind
  (endpoint.go:200-202), whose BatchSize()=128 on android.
- A/B: switching the active node WARP->home does NOT free memory (buffers do not sleep);
  config 11 ep -> 1 ep gives 269MB -> 0 (= Iliya's workaround, reproduced via profile).

Both prior diagnoses were wrong: batch=1 (no, StdNetBind=128) and drain-on-Suspend of
channels/sync.Pool (misses; bufsArrs of live recv-workers holds it). MaxSegmentSize
2200->65535 is the volume trigger (x30 bytes), not the holder (downLocked/pools/channels/
batch identical 1.13<->1.14).

Fix = shrink batch for INACTIVE devices (naive lazy-bufsArrs impossible: StdNetBind
getMessages() is a fixed 128). Levers + a required download-throughput measurement
documented; pending lever choice.
2026-06-29 02:07:46 +03:00
Leadaxe 0d19fe2c66 docs(SPEC 020): rewrite — corrected heap holder (sync.Pool+channels, not bufsArrs/128), drain-on-Suspend fix
On the client path BatchSize()=1 (client_bind/stackDevice/systemDevice all return 1),
so the old bufsArrs=128 => 8MB/device claim is wrong; bufsArrs is ~64KB. The 224MB
pprof attributes to PopulatePools is the sync.Pool.New alloc SITE, not the holder.
Real holders: (A) device.pool messageBuffers sync.Pool local+victim cache, (B) the 3
buffered device channels. scanobject 52% comes from the pointer-dense element/container
wrappers (4 of 5 WaitPools are scan-type), not the noscan [65535]byte arrays.

Fix rewritten to drain-on-Suspend: park RoutineReadFromTUN via a stackDevice.suspended
seam in Read (the lockless pool writer surviving Down), drain the 3 device channels
after Down, then ONE runtime.GC()+FreeOSMemory() per selector transition. PopulatePools
swap rejected (WaitPool.count underflow -> cond.Wait deadlock). slim-batch and the
route-graph refcount are dropped (not needed for heat). Added in-repo verification
(device/suspenddrain_test.go + HeapInuse bench) since on-device A/B is impossible.
2026-06-29 00:40:05 +03:00
Leadaxe 7a1e61e047 docs(SPEC 020): multi-WG idle-buffer heat — root cause + slim-device fix plan
Android 100% CPU / heat on configs with many WG/AWG endpoints in a
selector. Diagnosed via on-device pprof (§207): scan-bound GC over a
224MB live heap = wireguard-go Device.PopulatePools buffers, held by
~10 idle WG devices. Suspend()=Down() marks the device idle but does
NOT release pools/workers (only Close() does). A/B on device: dropping
spare WG endpoints removes the heat.

SPEC 020: SLIM idle devices (shrink maxBatchSize 128->1-4, do not Close
— keepalives + shared-node safety) gated by a route-reachability
refcount (rules + final + Now, not just selector).

SPEC 010: note our GRO split-brain patch is now upstream-native on
v0.0.3 (commit 24ea133); MaxSegmentSize=65535 must stay (GRO fuel) —
heat is fixed by device count/slimming, never by shrinking the buffer.
2026-06-28 23:39:22 +03:00
Leadaxe a531879e02 fix(SPEC 019 v2): three sticky/pool bugs found by device verification
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.

1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
   The router resolves a domain destination to an IP and overwrites
   metadata.Destination before a group's DialContext runs, so destination.Fqdn
   is empty when the balancer builds the key. stickyComponent("domain") read
   that empty Fqdn, so a single process's key was process+NUL for every site
   -> one fixed slot. On device this measured 28/1/1 across a 3-node pool
   (uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
   back to destination.Fqdn only for a direct dial. After: spread 0.95+.

2. living pool nodes could change slot index during a health-check, moving
   sticky keys. balancePoolFirstLive compacted with a filtering append (a
   transiently-dead slot shifted every later live node left); planTolerantPool
   did delete(inPool, occupant) (an evicted-but-living node re-entered a later
   slot, cascading); manual URLTest rebuild ran the tolerant planner even at
   pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
   dead/empty slots rewritten by index; dedicated planFirstLivePool for the
   tolerance==0 rebuild).

3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
   (badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
   empty array to nil, indistinguishable from omitted, so the default always
   applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].

Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
2026-06-28 21:48:31 +03:00
Leadaxe ebf5c2d048 docs(SPEC 019): collapse to a single canonical SPEC.md (v2)
v2 superseded v1; keeping both as separate files (SPEC.md + SPEC_V2.md + the v1
TEST_REPORT) was just confusing. Delete the v1 SPEC and its TEST_REPORT (they remain in
git history) and rename SPEC_V2.md → SPEC.md as the one canonical doc. Drop the "v2"
suffix and stale "design not started" status from the header.
2026-06-28 18:42:21 +03:00
Leadaxe 50efd8f335 fix(SPEC 019 v2): balancer.pool 0 = default, not error (rc.14)
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.

Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
2026-06-28 18:11:55 +03:00
Leadaxe 08d1fecd5f Merge upstream/testing (dhcp dns, age report, go mod) into lx-1.14 2026-06-28 17:51:17 +03:00
Leadaxe 5997b1812a lx(1.14): SPEC 019 v2 — round_robin pool, lazy health-check, slot-hash sticky, GetPool
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:

- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
  replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
  full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
  per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
  found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
  living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
  per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
  from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
  the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
  group -> empty. Additive proto/daemon/libbox, behind with_lx_command.

Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.

Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
2026-06-28 17:50:43 +03:00
世界 b70df9668e Improve DHCP DNS server initialize
Avoid querying DHCP DNS servers on cellular networks
2026-06-28 16:51:41 +08:00
Leadaxe 6139a6e305 docs(SPEC 019): record variant B (cache fallback) as considered-and-rejected
Clarify the Now() cold-start tradeoff: variant B (write the fallback node straight
into selectedOutbound*) would eliminate the micro-gap entirely — Now() and DialContext
would read one field, so they can't diverge — at the cost of touching upstream's
selection logic (stub in the choice field + one extra Interrupt() on the first real
switch, which is no worse than any later latency switch). Variant A (Now() stays a
reader) was chosen purely for minimal upstream intrusion; its only cost is a negligible
micro-gap from two separate Select() calls racing on the first seconds. Documents the
path to B if the feature outgrows upstream's selectedOutbound* later.

Doc-only; rc.12 already shipped variant A, no retag.
2026-06-28 11:25:34 +03:00
Leadaxe a40f0b64e4 Merge upstream/testing (oom report logs) into lx-1.14 2026-06-28 11:00:33 +03:00
Leadaxe feab497fe4 lx(1.14): SPEC 019 Now() cold-start fallback to Select()
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.

Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
2026-06-28 11:00:17 +03:00
世界 53e87a7ea5 Add age support for report export 2026-06-28 14:17:41 +08:00