Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.
They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.
Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
- an explicit value in .config beats a `default` (same rule that kept our
`# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
- `select` is OR-ed in after the user value, so shater-core's
`DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
transitive kmods) back on their own.
Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.
Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.
The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:
config ALL_NONSHARED ... default ALL
config ALL_KMODS ... default ALL
config ALL ... default y
In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.
Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.
Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.
Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was
running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn —
none of which we ship) before it ever got near our four.
Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the
.config that ships inside the ImmortalWrt SDK tarball. That file is the
buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus
CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so
`make defconfig` re-selected every kernel module of the target as =m and
package/kernel/linux/compile — pulled in via shater-core's nft kmod
deps — packed the lot.
Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds":
start the .config EMPTY so kconfig can only pull in what our packages
actually select. Carried over from the SDK's .config, nothing more:
the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a
wrong guess here means silently cross-compiling for another arch),
CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and
CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one
makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel).
Also adds the diagnostics this lane never had, since a failed run leaves
a 27 MB log: the carried-over identity lines, the post-defconfig kmod
count and target readout, a hard check that all four of our packages
survived defconfig, an abort if the kmod count is back in the hundreds,
and du/df after compile.
opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR,
CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in
the signed Packages index can't mask an empty SHA256. Script survives -u
(all vars use :? or :- defaults).
- .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the
copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a
trap-based tmpdir cleanup, and set -euo pipefail.
bash -n clean on both.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Backend audit fixes (upstream-file edits wrapped in // lx: markers):
- experimental/libbox oom_report.go/report.go: OOM reports + configuration.json
(server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files
→ 0o700 / 0o600. [sec-perms]
- daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret
compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare.
[sec-consttime]
- service/oomkiller/timer.go: network-extension cleanupTriggered logic was
inverted, so FreeOSMemory was never called after a trigger; flip both
assignments so a trigger schedules the deferred free and the next poll runs +
clears it. [sec-oomcleanup]
- transport/v2rayxhttp/client.go (lx-native file): session id used math/rand →
crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface.
- daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a
goroutine + the ssh-agent fd on every closed session (second io.Copy blocked
on an idle agent Read forever); tie both copies + the session ctx to a
cancel that closes both ends. [sec-sshagent]
- daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to
1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate]
- route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in
select while stopIdleSuspend niled it after close (race + goroutine leak on
Close-during-tick); pass the stop channel to the loop by value.
go build ./... (default) and the D9 shaterd linux build (tags
with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,
with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer
warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/...
./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend
and v2rayxhttp with with_xhttp.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- README.md: new Russian product README (what/features/architecture
mermaid/install both feeds/build/repo layout/CI/upstream/docs/license)
- README.en.md: concise English mirror (root readme was previously English)
- README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme,
a competing Russian README) -> points to README.md + engine-fork docs
- docs-shater/README.md: folder index
Install commands copied verbatim from docs-shater/INSTALL.md; all links
verified against existing files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Remove untracked-quality artifacts accidentally committed during work
sessions (all authored downstream, unreferenced anywhere in code/docs/CI):
- 5 session screenshots in repo root (devices-after-copy-fix.png,
live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png)
- tmp/gen_linux_test (29 MB throwaway traffic-gen binary)
Guard against repeats: ignore /*.png (root screenshots) and /tmp/.
Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are
left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- drop c/Users/.../gen_linux_test (28MB binary accidentally committed in 129e31fbd)
- .gitignore: ignore .idea/ at any depth (shater/.idea from IDE)
- CLAUDE.md: orchestrator delegates to model fable
- bump shaterd/shater-core r2->r3, luci-app-shater r1->r2 for v0.2.1 release
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
release-apk now runs with if: !cancelled() so an unrelated arch build
failure (e.g. x86_64) does not block publishing the aarch64 apk feed.
download-artifact only fetches existing artifacts and the publish loop
already skips missing apkfeed-* dirs.
The filtered-query-log block (SegMeter + top blocked + live QueryLog) is
redundant with Insights. The stats poll stays - it still feeds the DNS
filtering and Groups modules.
The subscription form now carries exactly: name, URL, update interval,
fetch via (+detour when proxied), User-Agent, HWID, and extra headers.
Format, device identity, regex/proto/country filters, dedup and expiry-alert
knobs are gone from the form (still honoured from UCI; a save carries them
through untouched). Headers are edited as key-value rows and serialize to
the existing `Headers: []string` "Key: value" contract. Name is editable:
a rename rewrites FromSub on the sub's cached nodes and refuses collisions.
Rule.Kill ""/"default" used to drop the rule, letting its traffic fall
through to the broader rules below and finally the default route - a silent
leak of exactly the traffic the operator singled out. ruleKillFallback now
always returns an outbound: ""/"default"/"closed"/unrecognised block the
rule's traffic in place; only an explicit kill=open goes direct. The default
route exists solely for traffic no rule matched.
One physical device with several addresses (v4+v6, multiple leases) used to
show as several devices. Discover now folds addresses sharing a MAC into a
single row: new `ips` field lists every address primary-first, `ip` stays
the primary (most recent lease), state is the best among addresses.
MAC-less hosts remain one-per-IP. The panel shows the extra addresses as
secondary chips; naming keys the config entry by MAC whenever it is known.
The engine resolves domains for itself (node server names, urltest probes,
subscription/DoH fetches). Those queries carried an invalid client address
and still landed in every insights surface. Gate them out at the single
ingestion point (Aggregator.handleEvent): an event with an invalid or
loopback client is dropped before totals, top domains, per-server counts,
the timeline, and the query-log rings. Only LAN-client traffic is collected.
The built-in block-ads / ru-bypass / private rule bundles are gone:
model.Preset, Model.Presets, the `config preset` UCI section, its render,
the panel Preset type, and every fixture. The generate-side expansion was
already removed with the profile rewrite in the previous commit.
A profile is now a pure uplink-conditional rule switch: Name/Enabled/
Priority/MatchIface/Enable-DisableRules/EndpointResolver. The per-profile
DefaultTarget/DefaultEgress overrides and the profile-level schedule window
(SchedDays/SchedStart/SchedEnd/SchedUTCOffset) are removed from the model,
UCI parse/render, the generator, the WAN watcher, and the panel. Rule-level
scheduling is untouched.
The panel's uplink condition is now picked from a dropdown of the router's
UCI interfaces (GET /api/interfaces, same source as the egress picker);
stored interfaces missing from the live list render as stale chips.
generate/profile.go is rewritten here (applyProfilesAndPresets ->
applyProfiles), which also drops the generate-side preset-pack expansion;
the preset model/UCI/panel surface is removed in the next commit.
Audit of the run-51 logs showed actions/cache@v3.3.2 works on the act_runner
(cold: "Cache saved" x4; next job: "Cache restored" in ~2s, npm --fast skip,
usign/dl reused) and the sdk-cache mirror seeds correctly — but the single
biggest recurring cost was NOT cached: `scripts/feeds update -a` re-cloned
base+packages+luci+routing+telephony every run (~7.8 min warm x 4 SDK jobs on
the serial runner ≈ ~28 min/run wasted; github ~1 MB/s from this host).
Cache .cache/feeds/{opkg,apk} (workspace dir, actions/cache-persisted, visible
in the SDK container via --volumes-from) symlinked over the SDK's empty feeds/:
`feeds update` now git-fetches deltas (seconds) instead of full clones, always
checking out feeds.conf's pins. Fail-safe: any error on the cached checkouts
wipes the cache and clones fresh. Key by SDK release (feeds-opkg-24.10.4 /
feeds-apk-25.12.1) — stable across runs, invalidates on an SDK bump; both arch
jobs of a lane share one entry (identical pins, serial runner).
Steady-state warm run: ~60+ min -> ~20-22 min. Also documented in the workflow
header: never key a cache on github.sha — each cache SAVE stalls the act_runner
~3 min, so per-run-changing keys would add +3 min/entry every run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Builds were dominated by re-fetching the ImmortalWrt 25.12 SDK tarball
(~300 MB) every run, and a stalled downloads.immortalwrt.org transfer wedged
the apk job for 40+ min (plain `wget -q`, no timeout — same class as the
elfutils hang).
- New ci/fetch-sdk.sh (runner-side): cache -> our durable `sdk-cache` release
mirror -> upstream with a stall-kill (curl --speed-limit 64K --speed-time 60
--max-time 1800) + 3 retries + zstd-magic/size validation; seeds the mirror
best-effort (github.token, non-fatal) so cold runs never touch upstream again.
A 40-min hang is now impossible; the in-container fallback wget also gets
--timeout=60 --tries=3.
- actions/cache@v3.3.2 (last release on the OLD cache API that Gitea act_runner
implements; v4/v3.4.x use the new GitHub cache service) for: SDK tarball, SDK
dl/ sources (hash of package Makefiles; PKG_HASH re-verified so a stale cache
can't leak a wrong source), Go mod+build (go.sum), npm node_modules
(package-lock.json) with build-shaterd.sh --fast, apt archives, built usign.
Degrades safely if the cache server is off — the SDK mirror is independent.
- concurrency group release-${github.ref} cancel-in-progress so a re-dispatch
cancels the stale run instead of piling up (tags stay isolated).
Signing (usign/apk), both keys, per-arch publish, manual triggers, LOCALMIRROR
and the scoped 4-package collection are unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
byedpi (ciadpi) ships as a separate optional package; the panel offered the
`byedpi` egress type regardless, so selecting it without the package installed
created a dead, fail-closed egress. Now GET /api/status reports
`byedpi_installed` (exec.LookPath("ciadpi"), os.Stat fallback), and the egress
type picker disables the ByeDPI option with a hint when it's absent. Existing
byedpi egresses are never hidden or rewritten (config is sacred) — shown with an
amber warning and still round-trip on save; only NEW selection is blocked.
Unknown status (older daemon / fetch fail) => no gating.
Bump shaterd PKG_RELEASE 1 -> 2 (the SPA is embedded in the daemon binary).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On a live `apk add` / `opkg install`, shater-core's post-install hung forever
(observed on BananaWRT 25.12 at "Executing shater-core...post-install", child
`flock 1000` in locks_lock_inode_wait). Root cause: a USE_PROCD init sources
/lib/functions/procd.sh on every rc.common action, whose procd_lock takes a
BLOCKING exclusive flock on /var/lock/procd_<svc>.lock held until the process
exits. shater-cron re-execs itself as the eternal `loop`, so it held that lock
forever; base-files' default_postinst then ran `/etc/init.d/shater-cron enable`
synchronously inside the transaction, blocking on the flock while the package
manager waited on the postinst — a permanent deadlock.
Fix (two layers):
- shater-cron `loop()`: `exec 1000>&-` closes fd 1000 up front so the eternal
loop never holds the rc.common flock (no-op when procd_lock is absent).
- 30_shater-core: defer enable/restart into a detached (setsid + bounded)
background block that waits for apk/opkg to finish before touching init.d,
with all fds to /dev/null (a held stdout pipe would hang apk on EOF too).
Bump PKG_RELEASE 1 -> 2 so existing installs pick up the fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk lane runs the SDK on a bare debian:bookworm host, and the ImmortalWrt
25.12 SDK prerequisite check requires python3-distutils ("Checking
'python3-distutils'... failed. Prerequisite check failed." ->
.prereq-build Error 1), aborting before any package built. The opkg lane was
unaffected because the openwrt/sdk image ships the prereqs. Add
python3-distutils (and python3-setuptools defensively) to the host deps. apk
lane only; opkg untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two build-harness bugs surfaced once the SDK builds actually ran:
1. Permission denied writing the feed. ci/build-feed.sh creates $OUT as root on
the runner, but the openwrt/sdk container runs as the unprivileged `buildbot`
(uid 1000) — so `cp` of the .ipk into $OUT failed ("Permission denied"),
yielding 0 packages and then "usign signing failed" (nothing to sign). Set
`chmod 0777 "$OUT"` on the runner before docker run (a chmod from inside the
container, as buildbot, cannot fix a root-owned dir). The apk lane already
chmods $OUT from its root debian container, so it was unaffected.
2. Collecting the whole SDK. ci/sdk-build.sh did `find bin -name '*.ipk'`, which
swept up the hundreds of prebuilt kmod/base .ipk shipped in the SDK image —
bloating the feed and signing foreign kmods under our key. Collect strictly
our four by name (`<pkg>_*.ipk`) and require >=4. Applied the same narrowing
to ci/sdk-build-apk.sh (apk names carry no arch: `<pkg>-*.apk`), keeping the
"wrong SDK produced only .ipk" guard.
No change to the feed format/signing (usign/KEY_BUILD/shater-feed.pub, apk EC
key), the package set, or triggers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk (and opkg) SDK builds intermittently hung fetching build-time
sources like elfutils-0.192.tar.bz2 from sourceware.org: curl's
--connect-timeout covers only the TCP handshake, not a stalled mid-transfer,
so a slow upstream hangs the whole job (no --max-time in OpenWrt download.mk).
Set CONFIG_LOCALMIRROR=https://sources.cdn.openwrt.org in .config before
`make defconfig` in both ci/sdk-build-apk.sh and ci/sdk-build.sh so the SDK
tries the fast OpenWrt source CDN before each package's own PKG_SOURCE_URL —
fixes elfutils and any other flaky upstream. Mirror verified to hold the file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Same root cause as the apk jobs: scripts/build-shaterd.sh builds through a
go.mod `replace => ./submodules/wireguard-go` (AmneziaWG fork, bumped in
16a47b596), and actions/checkout does not fetch submodules by default, so
`go build` died with "reading submodules/wireguard-go/go.mod: no such file or
directory" in the opkg build jobs (x86_64 + aarch64_cortex-a53) as well. Init
only that one submodule — build-harness only, no change to the opkg feed
format/signing (usign/KEY_BUILD/shater-feed.pub) or package set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/build-shaterd.sh builds via a go.mod `replace => ./submodules/
wireguard-go` (the AmneziaWG-patched fork), so that submodule must exist or
`go build` dies with "reading submodules/wireguard-go/go.mod: no such file or
directory". actions/checkout does not fetch submodules by default. Init only
that one submodule (public GitHub URL; clients/apple+android are large and
unused) in the additive build-apk jobs — the opkg build jobs are left untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Additive next to the opkg/24.10 lane — nothing existing changed. The same 4
packages (shaterd, shater-core, luci-app-shater, byedpi) are built through the
official ImmortalWrt 25.12 apk-SDK and published as per-arch rolling releases
apk-latest-<arch> / apk-<tag>-<arch> (x86_64, aarch64_cortex-a53).
- ci/sdk-build-apk.sh: drives the 25.12 SDK inside debian:bookworm, compiles
.apk, then `apk mkndx --root T --keys-dir T/keys --allow-untrusted
--sign KEY --output packages.adb *.apk` — the exact form the OpenWrt 25.12
buildsystem uses (unsigned members, signed index).
- ci/build-feed-apk.sh: per-arch runner entrypoint (same --volumes-from and
artifact-order contract as ci/build-feed.sh).
- ci/gen-apk-key.sh: one-shot EC (prime256v1) keypair generator; private half
-> Gitea secret KEY_APK, public dist/shater-apk.pem committed.
- release.yml: additive build-apk / release-apk jobs; `on:` triggers untouched
(v* tags + workflow_dispatch); apk release tags deliberately non-`v*`.
- docs-shater/INSTALL.md section 6, .gitignore (out-apk/), dist/shater-apk.pem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The three dependent plaques (Keep file on flash / File size cap /
Download) only had their inner controls disabled — the rows still
looked live and hoverable. Field gains a disabled prop rendering
set-field--off: pointer-events none, opacity .45, grayscale, flattened
background — the whole plaque reads and behaves as switched off
(aria-disabled included). Wired to !logToFile on all three.
Verified live on the testbed: with the toggle off all three plaques are
inert and dimmed; flipping it back on restores them.
embed.FS carries no timestamps, so SPA responses went out with neither
Last-Modified nor ETag and browsers fell back to HEURISTIC caching — a
stale index.html kept showing the previous panel after a daemon upgrade
(user saw pre-c61cfe3a download buttons enabled with the toggle off).
index.html / SPA fallback / favicon / 404s => Cache-Control: no-cache
(revalidate every load); a HIT under assets/ (content-hashed by Vite)
=> public, max-age=31536000, immutable. Guarded by TestStaticCacheHeaders.
Verified live on the testbed: / and /settings no-cache, hashed asset
immutable, missing asset 404 no-cache; a plain reload now picks up the
new SPA.
Operator decision (supersedes ad9781bf): "Log file (downloadable)" off
must leave NO trace — delete the saved log files outright, and gray the
download buttons out while the file is off.
logsink: New and Reconfigure purge the active segment and the rotated
.1 whenever ToFile is off — at the old and new configured locations AND
both standard paths (a Persist flip must not leave a stale copy). A
daemon booting with the toggle off sweeps leftovers from a previous
life too.
panel: /api/log reverts to the pre-ad9781bf precedence (toggle off =>
syslog scrape / '# logging disabled'; segments are never served while
the file is off, even if a leftover exists). SPA: the three download
buttons are disabled when LogToFile is off; note/flash texts and the
?mock fixture say the files were deleted.
Verified on the docker-OpenWrt testbed via the panel: off+apply deletes
/var/log/shaterd.log* (and /etc/shater), buttons gray out; on+apply
starts a fresh file and downloads work again.
Flipping "Log file (downloadable)" off looked like it deleted the logs:
the file stayed on disk, but GET /api/log switched to the logread scrape
and the collected history became undownloadable (user report). The
toggle stops WRITING — it must not disown what was already collected.
New precedence: retained segments are streamed whenever they exist,
prefixed with a '# note: file logging is off …' line when the toggle is
off (even with syslog off too); the syslog-scrape and '# logging
disabled' fallbacks now speak only when nothing is retained. Settings
note/flash texts and the ?mock fixture updated to match.
Verified on the docker-OpenWrt testbed: with log_file=0 the download
returns the note + full history; re-enabling via the panel resumes
appending to the same file with nothing lost.
modernc.org/sqlite is the only pure-Go SQLite and costs ~3.5 MB in the
static shaterd link; the stats store never used anything SQL-specific —
it is a ring of two append-only streams with a monotonic seq cursor.
bbolt is already linked via experimental/cachefile, so the swap is free.
sqlitering.go -> boltring.go: buckets queries/conns keyed by 8-byte
big-endian seq (bbolt key order == cursor order), rows as JSON of the
existing LogEntry/ConnLogEntry structs, meta bucket carries the durable
per-stream HWM (same max-only monotonic semantics). The async writer
contract is untouched (writeCh 4096, drop counters, 256-row/500ms
batches, 30s retention tick). Disk cap: chunked oldest-first deletes
with the same hysteresis, then at most one bbolt Compact per pass
(sagernet/bbolt exports Compact) behind the same 110%+1MiB free-space
guard that gated VACUUM. A legacy SQLite-format stats.db (or any
unreadable file) is replaced in place with one warning; open failure
still falls back to the in-memory ring.
Zero user-visible change: the "sqlite" backend selector value and the
Snapshot.Backend string are kept verbatim. Tests ported assert-for-
assert plus new coverage: legacy-file replacement, overflow drops,
memRing parity round-trip, disk-cap convergence.
Router shaterd (linux/amd64): 28,004,478 -> 24,428,670 bytes (-3.58 MB);
modernc.org/* gone from go.mod/go.sum and the dep graph.
shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS
transport is never generated, and the slim shater/registry never
registers the transport, so the tag gated nothing in this binary.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
Upstream defect: acme.go is behind with_acme but acme_logger.go was not,
so go.uber.org/zap linked into every build even with ACME disabled. Only
acme.go references ACMELogWriter/ACMEEncoderConfig, so the twin gate is
behaviour-preserving; a with_acme build still compiles.
Marked lx:acme_logger_gate; upstream-PR candidate (drop the lx block on
rebase once merged). -94 KB on the router shaterd link.
The admin panel is shater's own web server and generate never emits a
clash_api service (shater/engine/engine.go pre-registers its own
dnstrack.Manager precisely because no api/clash_api observer exists on
the router). With include.Context gone the Clash server was already out
of the link; dropping the tag records the decision. Desktop/CLI LX_TAGS
keeps with_clash_api for external dashboards.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
The shater data plane is tproxy/redirect (netplane); generate never emits
a tun inbound, so the userspace gvisor netstack is unreachable code. With
the slim registry it was already dead-code eliminated by the linker —
dropping the tag makes the intent explicit and stops compiling ~3.6 MB of
gvisor sources into the build at all. A future tun inbound would fall
back to the system stack; re-add the tag if that ever lands.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
include.Context registers upstream's entire zoo — tor (bine), ssh, snell,
anytls, naive, masque, mdns/resolved, and the api service whose daemon
bridge links grpc+protobuf — none of which shater/generate ever emits.
shater/registry registers exactly what the generator can produce (tproxy/
redirect/direct/socks/http/mixed inbounds; direct/block/selector/urltest/
socks/http/ss/vmess/trojan/vless/shadowtls outbounds + hysteria2/tuic
behind with_quic; wireguard endpoint behind with_wireguard; tcp/udp/tls/
https/hosts/local/fakeip + DoQ/DoH3 DNS transports; xhttp + v2rayquic
transport blank imports), with build-tag stub twins so a tag-less
'go build ./...' stays green. Zero upstream diff.
Measured on linux/amd64 with the D9 router tag set: 47.05 MB -> 31.07 MB
raw (-34%); the unreachable gvisor stack and grpc/protobuf are dead-code
eliminated even before any tag changes. UPX --lzma artifact: 12.49 MB ->
~8.8 MB. Since a UPX-packed binary unpacks fully into anonymous pages,
the same ~16 MB comes off resident RAM on the router.
New "Daemon log" group on Settings (Faceplate): the LogLevel verbosity
select (relocated, honest note — "none" is a turn-down to panic-only, not
a true off; failures still alert), LogToFile / LogToSyslog / LogPersist
toggles, a validated LogMaxKB editor (128–8192), and three download
buttons (day / 3 days / everything) → downloadLog() fetches
GET /api/log?range=… with the session cookie, filename from
Content-Disposition, blob save. Honest warn plates: file-off = only a
slice of the syslog ring (ranges approximate); both-off = nothing is
written anywhere; flash vs tmpfs (lost on reboot, wears flash, ~33 MB
budget). Globals type gains LogToSyslog/LogToFile/LogPersist/LogMaxKB
1:1 with the backend; mock.ts mirrors the honesty contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The daemon's own log (engine + control-plane) went only to os.Stderr →
procd → the logread RAM ring: no file, no wall-clock timestamps, no size
cap, and "LogLevel=none" silenced ONLY the engine while the control-plane
kept writing at trace. So "download last day/3d/all", "limit the size" and
"fully turn it off" were all unmet.
New shater/logsink: one long-lived, atomically-reconfigurable Sink that
receives BOTH halves' byte streams, stamps every complete line with a UTC
RFC3339 wall clock (what makes date ranges real), and fans each line to a
size-capped 2-segment rotated file (ToFile) and/or the real os.Stderr
(ToSyslog). Both off = the line is dropped — the only true full silence.
Persistent path sits behind a stats-style disk-free guard (suspend+warn
once, auto-resume); tmpfs path is bounded by the cap itself. ANSI stripped
from the file copy only.
Wiring: control-plane via log.SetStdLogger over the sink; engine via a new
box.Options.DefaultLogWriter threaded into all three box.New sites
(apply/close-then-start/restore) by engine.SetDefaultLogWriter; live
reconfigure on every apply.Reconcile (SIGHUP / control socket / panel
apply) so panel changes take effect without a daemon restart.
controlLogLevel now makes the control-plane respect Globals.LogLevel
(silent vocab → panic-only; unknown → warn, mirroring generate).
Globals: LogToSyslog/LogToFile (default true), LogPersist (default false =
/var/log tmpfs; true = /etc/shater flash), LogMaxKB (default 2048, clamped
[128,8192]; 0 = default, not off — LogToFile is the off switch). UCI
parse/render/aliases + ValidateGlobals clamp-warn.
Endpoint GET /api/log?range=1d|3d|all (session-gated): streams the log line
by line, oldest segment first, filtered by the timestamp prefix; UTC
attachment filename. Honest fallbacks — file off + syslog on → a
"# note: … syslog ring only, ranges approximate" comment then a
`logread -e shater` scrape; both off → "# logging disabled". Unknown range
→ 400.
init.d: shater/shater-cron gate their `logger -t` status lines on
log_syslog so "logread off" is honest at the shell layer too.
Tests: logsink rotation-cap/timestamp/toggle-gating/engine→sink,
model round-trip + validate, endpoint session-gate/range/fallbacks.
VM-verified on QEMU (x86_64, OpenWrt 24.10): download+ranges, size-cap
rotation, file-off/full-off, persistent path, live reconfigure — all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A rule disabled in UCI but force-enabled by the active profile, with an
iface:/zone: source outside the tproxy-inbound set, got its engine route
rule but no nft divert — its traffic never entered the engine, and the
fail-closed forward drop and accept_local sysctls skipped the device too.
Root cause: generate applied profile enable/disable in its own
effectiveRules while netplane read raw Rule.Enabled. Fixed with one shared
resolver in the leaf model package (ResolveActiveProfile +
ApplyProfileRuleOverrides) that both the engine route plan and the nft
divert plan consult, so they can never disagree about which rules are in
force. applyLocked now threads a single now through generate + nft render +
sysctls, closing the schedule-boundary race between the two planes.
Verified: a profile-enabled iface rule now joins the divert set, the
per-rule tproxy emit, the fail-closed drop and the accept_local sysctl;
the inverse (profile-disabled) drops the device. Parity regression on the
existing generate profile tests stays green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The schedule evaluator called time.LoadLocation, but the router binary
embeds no tzdata and OpenWrt ships none — so LoadLocation always failed
and windows silently ran in UTC while the panel promised local time.
- Windows now anchor to SchedUTCOffset (minutes east of UTC), which the
panel captures from the editing browser on every schedule save; the
daemon evaluates now.UTC()+offset with no location database. This
sidesteps the weekly-recurring day-shift that a full local<->UTC
conversion cannot express in one window. SchedTZ is deleted (documented
in the removed-options list; old configs parse and drain it). DST is a
stated limitation (followed on re-save). generate/schedule.go collapses
from a second copy of the evaluator to a thin adapter over the model one.
- The iface-profile schedule was honored by the WAN watcher since
08d5d6cc, but generate warned "the watcher does not look at the schedule"
and the panel muted the editor with "the router ignores the schedule" —
both false. Warning and lie removed; the editor is live and labelled
"applies together with the uplink match".
- Stale fictions: FEATURES.md nftset/FakeIP-mode MVP line and the shipped
conffile's dead `option dns_mode 'nftset'` corrected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- The generator iterated dns_rules in raw slice order and never read
DNSRule.Order; first-match "top to bottom" was true only because the
panel pre-sorts. Now sorted by (Order, index) like route rules, so a
hand-edited UCI or any API client gets the declared order.
- dns_rule match_src silently dropped zone:/iface:/MAC entries (the
in-engine DNS plane matches source IPs only), which could widen a rule
to ALL clients or skip it entirely. Each dropped entry now warns, with
the consequence spelled out.
- BlockDoH :443 IP list was incomplete (no NextDNS anycast, no actual
cloudflare-dns.com 104.16.x, sparse v6). Extended across all listed
providers, now accepts anycast CIDRs, with a maintenance note that the
list is manual. The hostname NXDOMAIN + canary layers already cover
resolve-by-name; UI still says "well-known providers only".
- Panel: intercept-OFF copy no longer overstates the bypass (plaintext to
external resolvers is already hijacked by the D14 catch-all); allowlist
note gains the per-device-Block-wins caveat.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit found the Insights numbers were real but mislabelled:
- "Blocked" counted every NXDOMAIN, upstream timeout and zero-answer as a
block. Now blocked = strictly the engine's own filter verdict
(dnstrack.SourceFiltered: D15 blocklist + BlockDoH predefined-NXDOMAIN).
Failures (timeout/SERVFAIL-reject) become their own `failed` category;
the three counters are mutually exclusive and sum to Queries. The DNS
log "block" tag follows the same signal.
- Per-minute sparkline positioned buckets evenly by index over a sparse
slice, so "60 min" could span hours. Now points sit at their real
Bucket.Minute, gaps render as gaps, and the label states the actual
span + active-minute count instead of a fictional "last N min".
- Per-device domains skipped the LAN filter every other view applies, so
the router's own urltest/sub-fetch dials appeared as a phantom WAN-IP
device. Now folded into the `router` pseudo-device like the DNS log.
- Honest labels: "Outbounds/exits" -> "DNS lookups per exit"; top
domains/hosts meta "N tracked" -> "top N shown". Overview query log
shows the real per-device attribution, not the resolver tag; stale
"DNS events have no client IP" comments removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The whitelist accepted random but the human-readable "Supported:" tail
still named only four strategies — caught live on the VM where the model
and generate warnings disagreed about the supported set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- node_down is gone from AlertEventNames: nothing ever fired it, so a
channel subscribed to it was silence dressed as monitoring. An old
config's `list event 'node_down'` now warns as an unknown event and is
dropped. The accepted and emitted sets now coincide; the reserved-event
branch of ValidateAlerts stays as the guard against future divergence.
- model.ValidateGroups + KnownGroupStrategies: a typo'd strategy is
warned at validation time (was: silently built as least_test with only
a generate-time warning). Mirrors generate's warnGroupStrategy list.
- doc-comment honesty: random is a real engine mode (api.ts), sqlite
stats backend is a real persistent store (api.ts + model.go), resolver
type list gains tcp, pages/index.ts no longer claims Placeholder pages,
failover doc says fail-back exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New balancer flag priority (option.URLTestBalancerOptions.Priority): the
pool is re-derived from CONFIG ORDER every health-check tick via
balancePoolPriority/planPriorityPool — the first live member owns slot 0,
so when the top node answers probes again traffic returns to it on the
next tick (30s failover interval). Probing walks top-down and stops at
the first live node, so the steady-state cost stays one probe per tick.
Replace-in-slot deliberately does not apply here: failover forces sticky
["none"], so relocating nodes across slots breaks no flow keys. Plain
round_robin/random paths are untouched.
failoverBalancer() now emits Priority:true; the KNOWN LIMITATION note and
the panel's "nothing brings it back" blurb are gone because the
limitation is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- random is a REAL urltest mode (lx SPEC 019 v2): uniform draw over LIVE
slots only, pool sized to every member; dead slots keep their place
(never-shrink) but are never picked, for random AND round_robin AND
sticky (degrade-to-live). All-dead pools fall back to Select.
- Globals.SweepInterval + Globals.GroupHealth master switch, resolved by
one pure function (model.SweepSchedule) shared by validator and apply;
unparseable is warned-and-ON, never silently off. ConfigureSweep no
longer resets the cursor on every cron reconcile (release blocker:
a ~6-min cycle was restarted every 60s and never completed).
- multi-WAN egress gateway: ubus netifd status -> uci static -> main
table; a gatewayless non-P2P egress warns CRITICAL instead of silently
blackholing the second uplink.
- endpoint resolver (route.default_domain_resolver): bootstrap-direct
clone of a named resolver, profile override beats globals.
- chains are composable: chain: hops flatten recursively, cycle-guarded,
entry egress lifts only at position 0 (fail-closed mid-path).
- group test publishes its scope so "measuring" lights only the cards a
run covers; health run is explicitly global (all_nodes).
- panel: biased-sample honesty (no ratio until a failure CAN be on
record), profiles auto-pin plate, sweep/GroupHealth settings UI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two threads, both from the same question: does this setting do what it says?
## Health is per-group, because a dial path is per-group
Overview reported "119 up / 179 untested" over all nodes, and Nodes showed a
per-row ping. Both measured the wrong object. A group with an egress binding does
not dial the base node outbound at all — generate materialises per-member copies
(group-<name>-m<i>-<member>) and the group balances over those. So a node can be
alive direct and dead through the tunnel a group is bound to, and the panel said
"up". The same node in two groups with different egresses is two states that were
being collapsed into one number.
No new prober was needed: the engine already keeps a process-wide
urltest.HistoryStorage keyed by outbound tag, groups already probe their own
members into it, and stats already reads it — we simply projected it onto base
tags only. The copy tag carries the member NAME, so recovery needs no change to
generate. GET /api/groups/health now reports alive/dead/untested per group, with
an opt-in member list; the same summaries ride /api/stats so Overview needs no
extra poll.
Presentation is "alive / tested" with the untested remainder as a quiet aside,
never folded into dead: groups probe lazily and only while in use, so on a fresh
boot with a 376-node subscription almost everything is legitimately unmeasured,
and calling that "down" would scream catastrophe exactly when nothing is wrong.
Three things this exposed, all fixed here:
- TestAllNodes enumerated only om.Outbounds(), which by design excludes
endpoints. Every WireGuard/AmneziaWG node read "untested" forever no matter how
often the button was pressed — on a product whose driving requirement is AWG.
- Our ProbeFailDelay sentinel is gone from the engine's history entirely. It was
safe for least_test (slowest wins last) but round_robin's pool planner treats
any entry as alive, so a dead node could occupy the single slot of a failover
group — pinning failover to a corpse, which is the one thing it exists to
prevent. Failures now live in an engine-side overlay, invalidated by timestamp
against any later success; the engine's history holds measurements only.
- Writers now measure with the probe URL of the group that owns the tag. A manual
run used the global URL and overwrote a group's own measurement, leaving
least_test comparing latencies to different servers. Where one tag is claimed by
two groups with different URLs the ambiguity is inherent to the engine's keying,
so we use the neutral global URL and say so rather than picking a silent winner.
A scheduled sweep (engine/sweep.go, on by default) fills what nobody probes:
24 measurements per 10s tick, 12 in flight, skipping anything fresher than 5
minutes — ~5 min per full cycle on the production config. The freshness gate is
load-bearing beyond cost: testNodes skips a member whose history is younger than
the group's interval, so a sweep that kept refreshing would starve a failover
group's own 30s check. It is a layer under group-local probing, never a
replacement.
## Options that did not exist are deleted, not decorated
Audited every enumerated choice the panel offers against what this fork actually
implements (constant/, option/, protocol/group/, dns/), and split the results into
works / synonym / fiction. Fictions are removed outright — pre-release, so no
legacy path is kept for values nobody has.
Deleted: Group.Strategy random and leastload (both silently became least_test);
Egress.Type proxy and block (emitted no outbound at all — every binding dangled
and the traffic left over the plain WAN with the real IP); alert event node_down
(no emitter anywhere); Globals.DNSMode, Inbound.Sniff, Profile.ProbeURL/ProbeMode,
Node.XUDPConcurrency/XUDPProxyUDP443, Egress.Target.
Repaired instead of removed, because the engine could do them all along:
LogLevel "none" (asked for silence, got default verbosity — now LogOptions.Disabled);
Subscription.Format (the hint was stored, badge-rendered and ignored — the sniffing
parser always ran); Ruleset.Format (never read; the extension decided);
Group.Strategy failover (urltest + round_robin + pool 1 / tolerance 0 is exactly
"first working node in order" — verified through box.New with a sensitivity control).
Relabelled where the words lied: Single promised "first up" but a selector never
checks liveness; inbound "http" opens Mixed and answers SOCKS5 on the same port.
An unresolvable egress binding no longer fails open. It resolves to block, so the
bound traffic stops visibly instead of leaving with the real IP. Refusing the
config was the alternative and is worse: a dead engine under a closed kill-switch
blackholes the whole LAN over one mistyped name.
Also: Egress.Port no longer defaults to 1080 for every type. The parser invented
it, render persisted it, and the new "port is ignored" warning then fired on a
correctly written config — a warning on a healthy install is how a findings list
gets ignored.
## Geo data is no longer hardwired to one publisher
sing-geoip publishes country codes and nothing else — 238 files, all two-letter.
So "route Netflix around the tunnel" meant loading geoip-us: 159,125 prefixes and
~20 MB of kernel memory for something the netflix list does in 108 prefixes and
~14 KB. Provider selection is now a chain (generate/geosource.go): country codes
still resolve to SagerNet byte-identically, everything else to Loyalsoldier, and
metacubex adds AS<number> routing. Third-party .srs was verified to load with our
own reader (v1/v2 against our v5 ceiling) before any of this was built.
The ruleset preflight reads four header bytes over a ranged GET instead of HEAD,
so a rule-set whose format version we cannot parse degrades like an unreachable
one — that case would otherwise abort engine start, which is how the LAN goes down.
The status strip carried five pips — ENGINE active, UPTIME, CONFIG enabled,
DATA PLANE installed, KILL-SWITCH — and the user had to AND three of them
together to learn whether they were protected. `plane` and `engine_running`
already encode that, and more precisely than the booleans did. UPTIME duplicated
the Engine module's "running for"; KILL-SWITCH duplicated the module directly
below it. Collapsed to one derived line phrased in terms of traffic:
Protected — traffic from your network is going through the tunnel
Traffic blocked — the tunnel is down (hold, amber)
Not protected — traffic is going out directly (none + fail-closed, crit)
Not protected — running direct (none + fail-open, amber)
hold and none stay distinct: one is the kill-switch catching it, the other is
no safety net at all. Nothing was lost — every removed value still lives in the
module that owns it.
Findings are now routed by severity instead of all landing on the front page
(panel/src/findings.ts):
critical / warning -> Overview. Something needs attention.
info -> the page that owns the setting.
An info finding is a statement about the configuration: it never clears and asks
for nothing, so a permanent front-page entry only teaches people to skim the
list — which is how a real critical finding gets missed. The untunnelable note
now renders inside the Networks "Other traffic" section, beside the control it
describes. With nothing needing attention the section renders nothing at all.
Also fixed, found while auditing the rest of the labels: the Kill-switch module
read ARMED / policy: fail-closed with a green lamp even at plane=none — a
reassuring light directly beneath a readout saying nothing is protected. A
fail-closed setting is only armed if something is installed to enforce it, so it
now reads NOT IN EFFECT with a crit lamp and a "blocking now: no — nothing
installed" row; policy -> setting.
planeState.ts became the single source of the wording, and the plane banner was
dropped from Overview — it exists to carry the alarm to pages with no status
readout, and stacked under the new line it just said the same thing twice.
Verified against the live daemon on the bench: healthy, critical and hold states
all render correctly, console clean, note present on Networks and absent from
Overview.
.gitignore: MemPalace per-project files, added by the tooling.
Traffic TPROXY cannot carry (ICMP, IGMP, ESP/AH, GRE) was dropped for the whole
LAN regardless of routing. A box configured to tunnel only 8.8.8.8/32 still lost
ping to the entire internet, and with the shipped config RU addresses were
unpingable even though `ru-direct` sends them out unproxied — the very path where
TCP already exposes the real IP, so the drop prevented no leak at all.
The drop is now scoped to destinations the rules actually tunnel:
iifname "br-lan" meta l4proto != { tcp, udp } ip daddr @unt_d4_1 accept
iifname "br-lan" meta nfproto ipv4 drop
Destination sets come from the engine's already-parsed rule-sets via
ExtractIPSet(), so no .srs parsing and no second read of the bbolt cache the
engine holds locked. Rules are taken from the generated route rules, not the raw
model, so preset packs, WAN-profile overrides and schedules are all included.
Domain/geosite matchers are skipped when classifying: a packet with no stream
carries no domain, so such a rule can never apply to it.
Every policy line carries `l4proto != { tcp, udp }`, so no destination decision
can ever accept TCP/UDP — fail-closed is structurally untouched. Anything the
walk cannot prove direct (list not yet fetched, logical rule, unknown action,
inverted match) falls through to the drop and says so via an info finding.
No element cap: a continent-scale list loads in full. Measured on the bench with
geoip-us — 4s apply, 1.25 MB ruleset in 29.5k lines, ~27 MB RSS growth, engine
healthy. Cost is reported, not enforced; `untunnelable=direct` loads no sets.
Also fixed here, found while building it:
- plan warnings were computed and dropped, never reaching the operator; routing
them through the netplane channel was wrong (it marks everything critical by
construction), so they get their own info-level path
- nft ran with no timeout while holding the apply flock: one wedged invocation
would have deadlocked every later apply, reconcile and teardown. 60s cap; the
ruleset commits as a single netlink transaction, so killing it is safe
- set elements were emitted as one 3.1 MB line the lexer would hold as a single
token; now wrapped at 8 per line (identical to nft, readable when debugging)
- untunnelable copy still claimed ping never works; rewritten for the new
semantics across all three modes
panel: the theme switch read as a power toggle — it reused the component that
turns features on and off and sat inside the status cluster next to the ONLINE
lamp, so in light theme it looked like a switched-off appliance. Now a two-key
sun/moon selector, both states always visible (neither theme is an "off"), the
engaged key raised and lit by shading rather than accent colour, separated from
the indicators by a groove.
Verified on the OpenWrt bench: RU addresses ping, non-RU stay blocked, TCP routes
unchanged through the tunnel, DNS filtering and Block-DoH unaffected.
Group egress — for the case where the protocols themselves are DPI-blocked:
every node in the group dials ITS OWN server through the chosen egress (an
AmneziaWG tunnel, say), so the provider sees tunnel traffic instead of a VLESS
handshake. It binds the outgoing dial, not post-proxy traffic.
The binding is per-group, and that is the whole difficulty: group members are
SHARED outbounds, so two groups built from one subscription — one bound, one not
— would either leak the binding into the unbound group or fail to apply it. The
members of a bound group are therefore materialised as per-group copies
(group-<g>-m<i>-<member>), reusing the same rebuildNode the chain builder uses
for per-hop copies. Copies are made only when Egress is set, so an unbound group
over a 331-node subscription does not double the engine config. Copy tags are
checked against the node/group/egress/copy namespaces; a collision skips the
member with a warning rather than shadowing a real node. Precedence is chain hop
-> Node.Egress -> Group.Egress: a node pinned to a particular uplink was pinned
for a reason the group cannot know. A member whose copy cannot be built is
dropped rather than falling back to its unbound tag — falling back would leak
exactly the traffic the binding exists to hide.
Group test answers "what am I exiting through, and how fast": selected member,
latency, exit IP and country, via cloudflare.com/cdn-cgi/trace (country comes
free, so no GeoIP database on the router) with api.ipify.org as fallback. The
probe is pinned to the group's own outbound and refuses the direct outbound — a
direct answer would print the ISP's address and claim the tunnel works when it
does not. Measuring latency but failing to resolve the address stays ok=true
with an empty exit_ip; that is a working tunnel, not an error.
Uptime: /api/status gains started_unix + uptime_seconds, measured from process
start over a monotonic seam so an NTP step on an RTC-less router cannot be
reported as uptime. It is the daemon's uptime, not time since the last apply.
Also fixes: renaming an egress did not rewrite Group.Egress, silently dropping
the group back to the default route.
Verified on the testbed with the real 331-node subscription: two groups over one
subscription, one bound, one not — the bound group selected
group-auto-egress-m130-IE-trojan-141 while the unbound one selected the shared
IE-trojan-141, exit IP and country resolved for both, no group warnings.
Measured cost of binding a 331-node group: engine outbounds 335 -> 666, config
50 KB -> 114 KB, daemon RSS 62 MB -> 75 MB.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
I wrote the gap up as open while reviewing an agent report I had not yet seen;
the hosts/plain/AdBlock parse-and-compile path had in fact landed in the same
commit. Records the measurement that settles the disk question: StevenBlack's
2.4 MB of text compiles to 80873 domains in a 491 KB .srs, so it ships in the
production posture instead of being traded away for the 8 KB geosite list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
On a healthy router the warning set is reprinted by every reconcile — once a
minute from cron plus every hotplug event — so logread filled with the same
line forever and buried the warnings that matter (a blocklist that failed to
load, an interface the kill-switch does not cover, a missing data plane). On a
router logread is an in-memory ring buffer, so this also evicted the history
needed to investigate an incident.
Warnings are still returned in full by GET /api/status on every request; only
the logging is deduplicated, keyed on a fingerprint of the set. Message texts
carry volatile parts (free MiB on /overlay, compiled domain counts, the address
inside a network error), so digits are normalised for comparison only — the
logged and API-returned text is untouched. A restart reprints the full set, and
clearing the last warning logs one line saying so.
Also fixes the severity-to-syslog mapping: an [info] warning was being emitted
at WARN, so anyone filtering on WARN saw noise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Records the positions the audit changed: fail-closed must cover "engine never
started" (holding plane), engine start must not depend on the network (remote
rule-set preflight, with the deferred cache-seed fix noted), BlockDoH needs no
route-plane upstream exclusion (engine dials bypass route rules), DNSMode is
unimplementable and its control was removed, TPROXY's inability to carry
ICMP/IGMP/ESP/GRE is now an explicit 3-way policy, and fail-open degradations
must surface in the panel rather than only in logread.
Also flags the contradiction left open: D15 promises seeding StevenBlack/OISD/
AdGuard while blocklist source=url accepts only compiled .srs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A ground-up audit of the whole v0.2 stack by 8 parallel agents (DNS generate,
routing generate, model/parse/subscribe, netplane/apply/engine/alert, stats +
panel API, panel frontend, OpenWrt packaging), with every finding reproduced or
verified on the OpenWrt QEMU testbed. ~60 defects fixed, each with a regression
test that was checked to FAIL against the old behaviour.
RELEASE BLOCKERS
* Engine-start failure left the data plane ABSENT: with kill_switch=closed the
router silently degraded to a plain OpenWrt box — no tunnel, no filtering, no
kill-switch — while the panel looked healthy. Reproduced live. Now any
engine-start failure installs a fail-closed holding plane (forward blocked,
LAN-to-LAN and management preserved) and reports plane=hold/none.
* An unreachable remote rule-set aborted engine start entirely, so a router that
booted before its ISP link came up ended with a dead LAN and no way to recover.
Remote lists are now preflighted and skipped with a loud warning instead.
* `geosite:` in a routing rule hard-errored box.New — one legacy rule took the
whole LAN down. Same class: unvalidated CIDR / port / regexp, and marker-only
list entries ("." / "keyword:"). A lone `keyword:` also silently NXDOMAINed
the entire internet.
* Fail-closed drop only covered tproxy inbounds, not interfaces diverted by rule
sources — engine down leaked those networks to WAN in plaintext (4f618140 redux).
* UCI injection: a newline in a subscription-supplied node name broke out of the
line-oriented config and wrote attacker-controlled sections.
* Bootstrap deadlock: the daemon refused to start while disabled, but the panel
IS the daemon — a fresh install could never be configured from the UI.
SILENT FAILURES (the audit's main theme)
* per-device DNS block ignored the `suffix:` prefix — parental control that
quietly didn't block. Unknown `word:` prefixes now warn instead of vanishing.
* sqlite reused `seq` after retention wiped rows, stalling the live log forever.
* `after=` cursor returned the NEWEST rows, permanently skipping bursts.
* Stats emitted null arrays on a freshly booted router, blanking Overview.
* Alerts fired twice per incident; new_device alerts swallowed all but the first
device in a 60s window.
* Disabled subscription nodes were silently re-enabled on every refresh.
* Invalid Include/Exclude regexes failed OPEN, disabling the whole filter.
DEAD KNOBS — wired or honestly removed
ru-bypass preset (emitted an unsupported geoip: matcher) -> real geoip rule-set
Globals.ResolverFallback -> implemented via evaluate + match_response chain
Globals.DNSMode -> unimplementable by design; control removed, fake-IP
documented via a type=fakeip resolver instead
Rule.Kill -> implemented (default | closed | open)
Rule.Egress -> was read by nobody; multi-WAN binding silently no-op
ExpireAlertDays + quota -> subscription-userinfo parsed, persisted, alerted
StatsBackend hot-switch -> store is re-created on change
Inbound.Sniff -> documented as vestigial (sniffing is a route action)
NEW
* Globals.Untunnelable (block | icmp | direct): TPROXY can only carry TCP/UDP, so
ICMP/IGMP/ESP/GRE were dropped with no explanation — ping simply didn't work.
Now an explicit policy, defaulting to the previous behaviour, and explained in
the UI by consequence rather than by protocol.
* apply now surfaces its warnings through /api/status (severity/section/name), so
fail-open degradations are visible in the panel instead of only in logread.
* Panel: Networks page (which LAN networks are intercepted + inbound editor),
DNS-rules editor, subscription quota/expiry, plane banner and findings list.
* Control-socket client got per-verb timeouts — a wedged daemon used to pile up
one stuck `shaterd status` per minute until OOM.
* cache.db is now bounded (8 MiB, tmpfs fallback below 24 MiB free): on a 98 MB
rootfs with ~33 MB free it could otherwise grow past what an upgrade needs.
* OpenWrt packaging: nftables-json + ca-bundle deps, postinst restart on binary
upgrade, idempotent rt_tables seeding, cron gated correctly.
VERIFIED ON THE TESTBED
fail-closed holds with the engine frozen; offline boot now starts the engine;
RU destinations go direct while the rest goes through a node (per-connection
proof); ads NXDOMAIN with allowlist override; DoH blocked while the configured
upstream still resolves; sqlite history survives a daemon restart; a failed
apply restores the previous config without dropping the engine.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Devices lose per-device Proxy/Target (routing is expressed with ordinary
routing rules whose Source picker targets a device); a device is now pure
DNS policy: identity + Enabled + Block/Allow. The panel drops the
"Route through proxy" toggle and Exit picker, and gains inline rename
(pencil -> input; renaming an unmanaged device upserts it into managed).
New Globals.BlockDoH (uci block_doh, default off): engine-level block of
known public DoH resolvers so clients fall back to plaintext :53 that the
engine intercepts. DNS layer answers the DoH hostnames + the Firefox
canary use-application-dns.net with NXDOMAIN; route layer rejects :443
(tcp+udp, HTTP/3 covered) to the hostnames and dedicated resolver IPs.
Hostnames/IPs that are themselves configured upstream resolvers are
excluded with a warning (never the canary). Reject rules set
Method=default explicitly - a directly constructed "" bypasses the
UnmarshalJSON normalisation and panics the engine at first match
(found live on the VM, regression-tested).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Per-device DNS block/allow and the DNS filter only caught DNS that the
tproxy forward-divert steals (client -> external resolver). A client using
the router itself as DNS hit dnsmasq directly (fib daddr type local bypass)
and slipped every filter. New Globals.DNSIntercept (uci dns_intercept):
when set, nft diverts all LAN :53 (tcp+udp, v4+v6, source-IP preserved)
into the engine ABOVE the fib-local bypass, so even DNS addressed to the
router is hijacked and per-device rules apply to everyone. DoT/DoQ :853
stays rejected (clients fall back to plaintext); DoH :443 can't be
intercepted (stated in the UI). .lan + private reverse zones are forwarded
back to dnsmasq (127.0.0.1:53, direct detour) so local names still resolve.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A WAN uplink on a private DHCP address (provider double-NAT) was wrongly
offered as a LAN source. Interfaces() now tags each interface with its
firewall zone (one `uci export firewall` pass, reusing the zone scanner),
and the picker treats an interface as a LAN network when it has a subnet
and its zone is not wan* — falling back to the private-subnet heuristic
only when the zone is unknown. VPN tunnels self-filter (no subnet).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The blind free-text Source input in both rule forms is now SrcPicker: a
chip slot whose popover offers the router's real private subnets (from
/api/interfaces, host bits normalized to the network address), the
discovered devices by name (the bare IP is what's stored), and a
validated custom IP/CIDR input. Empty = "everyone · all LAN clients".
Existing hand-typed Src values classify back into device/network/custom
chips by value, never rewritten. Shared module cache: one
interfaces+devices fetch per session across all open forms.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Domain matching belongs to rulesets (that's what they are for) — the Match
picker is now rulesets only / ip-cidr / port, with the value input hidden
for rulesets-only. The edit form keeps a "Domain(s) — legacy" field ONLY
when a rule already carries free-text domains, so old rules stay visible
and clearable instead of silently preserved.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- model: Ruleset/Blocklist/Allowlist Category -> Categories []string
(uci `list category`; a legacy lone `option category` still reads as a
one-element list and migrates to the list form on the next write)
- generate: one remote .srs per category, tag rs-<name>-<category>
(bl-/al- for the DNS filter); a rule referencing the ruleset matches
every category's set; non-geo sources keep their old single tags
- panel status: rows gain `category`; tag->(name,category) resolved from
the model, not string parsing (names/categories may contain dashes)
- CatSuggest is now a chip multi-select: pick from the SagerNet base ->
chip with a status LED (green = from base/verified, amber = added
offline "anyway"), duplicates flash the existing chip, Backspace/×
remove, and free unpicked text never survives blur or save
- Routing/DNS forms save Categories (>=1 chip required); ruleset rows
show `geosite · youtube +2`, freshness groups per name (oldest wins,
Update now refreshes every category's tag)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The native <datalist> popup is an unstyleable browser widget that clashed
with the panel. CatSuggest is an instrument-styled readout docked flush
under the Category input: SAGERNET BASE · n shelf label, sunken dense mono
list, matched substring lit in the accent, LED bar on the active row,
green exact-match footer, amber not-in-base warning. Keyboard: arrows /
Enter / Esc; combobox ARIA; both themes via tokens; reduced-motion safe.
Fix along the way: the row is a flex container with a gap, so bare text
nodes around <mark> became separate flex items and the gap split the
category name itself — the name now renders inside one span.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- GET /api/ruleset/categories?source=geosite|geoip — daemon lists the real
SagerNet rule-set branch via the GitHub git-trees API (UA set, 24h in-memory
cache, stale-on-error); panel drives a native <datalist> on the Category
inputs (Routing ruleset form + DNS geosite blocklist), lazy one fetch per
source per session
- Routing rules gained Edit: inline form (all three matchers shown at once —
domains/IPs/port — so nothing is silently dropped), preserves Order/Enabled/
Kill/Egress and off-form fields verbatim, reorder/toggle/delete frozen while
editing, "no matchers — matches everything" hint for catch-alls
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- generate.GeoRuleSetURL(source, category) — single source of truth for the
SagerNet .srs URL (routing rulesets, DNS filter, and the checker)
- POST /api/ruleset/check: daemon-side HEAD (GET+Range fallback) existence
probe of the exact URL the engine would fetch; {ok} / {not_found} /
{network} — plain client, router-own output is never tproxy-diverted
- Panel: geosite/geoip Save now checks first (Checking…); not_found blocks
with a form error; network failure offers explicit "Save anyway" so an
offline router can still be configured; url/inline/file flows untouched
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Three features from the second feedback pass:
- Node.Egress: a node can dial its OWN upstream through a named egress
(DialerOptions.Detour on the node outbound/WG endpoint; inherited by
groups/chains/rules; fail-open on unknown egress). Panel: per-row "via"
expander + "via <name>" chip on Nodes.
- Chains: egress:<name> allowed as the ENTRY hop only (hop 0) — lifted into
the first hop's detour; mid/last egress hops warn+drop. Panel: entry-hop
optgroup + "entry" badge in the chain editor (Targets).
- Test all nodes: engine.TestAllNodes force-probes every node outbound
(concurrency 16, 5s timeout) into the shared urltest history; failures
stored as ProbeFailDelay=0xFFFF sentinel (slowest, never poisons
least_test) and surfaced as DOWN, not untested. POST/GET /api/nodes/test;
"Test all" button with N/M progress on Nodes.
- geosite/geoip rule-sets are LIVE: source=geosite|geoip + category emit
official SagerNet remote .srs rule-sets (24h auto-update, direct fetch),
for routing rulesets AND DNS block/allowlists; freshness + Update now UI
applies to them; "inert" badge removed. New Category field (model+UCI).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- Nodes/Overview: counters now tally live probe health (up/down/untested/off)
instead of counting Enabled as "up" (was: 329/329 up with 5 tested)
- Routing: Target select bucketed into optgroups (Groups/Chains/Interfaces-
egresses/Nodes-last) so egress:* is no longer buried under 300+ nodes
- Insights/Overview logs: single fmtClock(unix) helper, browser-local time in
BOTH logs (DNS log was server-UTC next to local-time connections)
- Overview query-log badge no longer says "waiting for engine stats" while
persisted rows are on screen
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The stats log rings move behind a logRing seam: memRing (extracted RAM ring,
byte-identical to before) + sqliteRing (new, modernc.org/sqlite v1.38.2 pure-Go,
CGO-free musl-static). backend=sqlite persists the query/conn LOG rows to
/etc/shater/stats.db (WAL, tmpfs fallback /tmp/shater-stats.db, own lock) via an
async batched writer off the DNS/conn hot path; seq = the PK (monotonic, resumes
from the persisted max after restart). Cursor reads = WHERE seq</> ? ORDER BY
seq DESC LIMIT. Retention: keep <= StatsRingSize rows/table + a StatsDiskLimitMB
disk cap (0=unlimited) with prune + wal_checkpoint/VACUUM. Aggregates
(top-domains/timeline/hosts/devices/node-health) stay in RAM (bounded, rebuild
fast) — only the unbounded LOGS persist. Open failure → warn + memRing fallback.
Globals.StatsDiskLimitMB (0=unlimited, intOptAlways round-trip); Settings shows
it when backend=sqlite. Size: +1.2MB UPX (10.5->11.7MB), static/musl OK.
Verified: build (router tags)/vet 0, go test + -race ok (sqliteRing cursor
parity, retention, disk-cap, PERSIST-across-reopen, async no-loss, memory
regression); panel tsc/build clean. VM: backend=sqlite → /etc/shater/stats.db
created, 75 conns logged, **survive daemon restart** (75 rows intact, seq
resumes 76->79), disk cap set 16MB. box.New applies.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Log rows gain a monotonic Seq (uint64, per-log counters on the Aggregator,
stamped under mu on append; survives box swaps + ring wrap). StatsStore read
API becomes cursor-based: Queries(LogQuery{Limit,Before,After})/Conns(...) —
neither cursor = newest Limit; Before=<seq> = next older page (seq<before);
After=<seq> = new rows (seq>after); always newest-first. Endpoints
/api/stats/{log,conns} accept limit/before/after (legacy n = limit, so Overview
is unchanged); bare array, rows carry seq (client derives newest/oldest).
Panel: Insights logs (Connections + DNS) now accumulate a seq-desc deduped
list — Load more APPENDS the next older page (not refetch-all, scroll
preserved, hides when exhausted), a ~1.5s after=<newest> poll PREPENDS new rows
(slide-in keyed by seq), a Pause/Live toggle buffers arrivals into an 'N new'
pill, ~3000-row DOM cap re-arms Load more, poll gated on backend!=off + tab
visible. Stable seq keys.
Verified: build (router tags)/vet 0, go test ok (seq monotonic bounded+
unlimited, before/after paging no overlap/gap, wrapped-ring, clamp), panel
tsc/build clean, ?mock drive (paginate+live+pause). VM live-verify pending
(ssh-manager MCP disconnected).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
First phase of a pluggable stats-storage backend. New stats.StatsStore
interface (Start/Close/Resubscribe/Snapshot/RecentQueries/RecentConns); the
existing in-memory *Aggregator implements it unchanged, plus a noopStore for
OFF that never subscribes (so the HasSubscribers-gated DNS emit path skips all
per-query work). stats.NewStore(backend,...) selects off->noop, memory->agg,
sqlite->agg+warn (persistent backend lands in Phase 3). Snapshot gains a
'backend' field (off|memory|sqlite = effective). Globals.StatsBackend
(off|memory|sqlite, default memory) via the KillSwitch string-enum pattern
(model/uci/render + round-trip). Daemon + panel decouple from *Aggregator to
the interface. Settings gets a 3-way Logging-backend Select; Insights shows an
honest 'logging is off' state when backend=off.
Verified: build (router tags)/vet 0, go test ok (noopStore contract, NewStore
selection, StatsBackend round-trip), VM box.New PASS (stats verb reports
backend=memory); panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
maxStatsLogN capped /api/stats/log and /api/stats/conns responses at 200, so an
unlimited (StatsRingSize=0) or large ring still returned only 200 rows. Raised
the safety ceiling to 5000 (RecentQueries/RecentConns still return only what's
buffered) and lifted the Insights Load-more ceiling 500->5000 (step +200) to
match, so a big/unlimited log can actually be paged out.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Retention size knobs (StatsRingSize / StatsTimelineMinutes / StatsMaxDomains)
now mean: 0 = UNLIMITED (no trim, grows with RAM), N = fixed limit. Query-log
AND connection-log rings gain a growable append-only mode when RingSize==0
(fixed-ring modulo path kept for N>0). Per-field 0 skips that aggregate's prune
(domains) / trim (timeline); RetentionDisabled stays the master switch.
Absent-vs-explicit-0 round-trip fixed: DefaultGlobals seeds safe bounded
defaults (200/60/5000) so an unset UCI option is never accidentally unlimited;
render intOptAlways writes these three fields even at 0 so an explicit 0
survives WriteUCI->ReadUCI; daemon passes no Config on a read error (→ bounded
defaults, not a zero-value=unlimited Config). Settings reframes the three
inputs as '0 = unlimited' with a per-field grows-with-memory warning.
Verified: build (router tags)/vet 0, go test ok (round-trip 0/200/5000;
RingSize=0 grows to 500 q+conn; bounded at 200; MaxDomains=0 keeps 6000;
Timeline=0 no trim), VM box.New PASS; panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The Connections/DNS LogShell only DISABLED 'Load more' when the page wasn't
full — so it stayed visible (and looked clickable) even with e.g. 9 rows. Now
the button is HIDDEN unless a full page came back (rows.length >= n && n < 500),
so it only appears when there may actually be more; disabled only while busy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The real 600-905px 'crossing' was NOT the SegMeter dots (fixed in e24a2c1c) but
the 3 QUERIES/BLOCKED/ALLOWED tiles forced 3-across in the ~227px first column
of the 3-col overview: each value's min-content (~86px) → 277px overflowed the
cell by ~50px, painting the 3rd tile's number ~17-30px into the sparkline.
(Round 1 missed it: Playwright's 15px scrollbar turned physical 901 into an
886 single-col layout, so the tight 3-col band was never measured.) Fix:
.ins-tiles flex-wrap + .ins-tile{flex:1 1 90px;min-width:0} → reflow 2+1 when
narrow, 3-across when wide, robust to any digit count, no viewport breakpoint.
Also: per-device header total wraps to its own line; Connections/DNS rows fit
their scroll box at <=560px. Scoped to Insights; verified 0 crossings 360-1440.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The prior min-width:0/max-width:100% capped the .segs BOX but the 28 flex
segments still painted outside it (their ~305px min-content spilled past the
right border at nearly every width). Now .ins-filter .segs uses overflow:hidden
(drops the min-content contribution + clips sub-pixel) + tighter gap, and the
Filtered meter passes segments={20} so dots fit their column without
compressing past the ~8px floor. Also fixed a secondary body h-scroll ≤404px:
.ins-overview single-col → minmax(0,1fr), .ins-tiles reflow to 2-col ≤400px,
.ins-grid minmax(min(100%,320px),1fr). Scoped to Insights — shared SegMeter
(Overview) untouched. Verified 0 overflow + no body h-scroll across 360–1440px.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
LogEntry.Device was always '' — but the client address IS on the DNS
resolution context. dnstrack.QueryEvent gains Client netip.Addr, populated at
all three dns/client_log.go emit sites from adapter.ContextFrom(ctx).Source.Addr
(same context processInfoFromContext already reads). stats deviceLabel: LAN
source (a.lanNets.isLAN) -> DHCP hostname or IP; loopback/non-LAN/unknown ->
'router' (the appliance's own urltest/sub/DoH lookups). Insights DNS log now
shows device -> domain · resolver · action (mirrors the Connections log), with
a dimmed 'router' chip for router-originated lookups; falls back to '—' on
older data.
Verified: root build (dns tree + box, router tags)/vet 0, go test ok
(deviceLabel: LAN+lease/LAN+IP/loopback->router), panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
User frontend polish pass: (1) Connections/DNS action buttons were glued —
now a left-grouped .ins-log-btns (gap) + right-aligned count + separating
groove. (2) DNS log now shares a LogShell (scroll body + actions) with
Connections so they look 1:1 (differing only in columns: time·domain·resolver·
action); dropped the old <QueryLog> ticker here (still used on Overview).
(3) 'Blocked' toggle was clipped to 'Blocke' — .ins-toggle overflow:hidden
collapsed its flex min-width; added flex:none/nowrap + header flex-wrap.
(4) Filtered SegMeter's 28 segments overflowed the module's right edge —
scoped min-width:0/max-width:100% under .ins-filter (SegMeter elsewhere
untouched). (5) tabular-nums, consistent spacing, removed dead .ins-logwrap/
.qrows + unused imports. tsc/build clean; verified ?mock at 1280 & 380px, no
horizontal body scroll.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The conn-log/top-hosts LAN filter required a non-empty Metadata.Inbound, but
tproxy connection events carry an empty Inbound on this engine — so it dropped
every client connection (live: 0 conns / empty top_hosts despite real traffic).
Now filters by source IP being inside a LAN subnet: devices.LANNets() (new
exported helper reusing the #10 /etc/config/network parser) + a 30s-cached
lanNetCache. Keeps LAN clients (192.168.1.77) and drops the router's own
WAN-side node dials (Source 10.0.2.x) — which are both RFC1918, so only the
iface config distinguishes them. Fallback (no config): private routable
sources. Loopback/unspecified/multicast always dropped.
Verified: build (router tags)/vet 0, go test ok (KEEP 192.168.1.77 / DROP
10.0.2.15 subnet test).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DNS events carry no client IP, so the query log couldn't show which device
went where. Now the stats aggregator folds connection events (trafficcontrol)
into: a connection ring (RecentConns → GET /api/stats/conns: src device ->
dest domain|IP + tcp/udp + sniffed proto + exit) and a top-hosts map keyed by
domain-else-IP (Snapshot.top_hosts) so raw-IP UDP/TCP destinations surface
(host==ip = the by-IP case). LAN-source filter (non-empty inbound + routable
src) keeps the router's own node/probe dials out. Insights gains a scrollable
Connections log (device->dest+proto, Refresh/Load more) + a Top-hosts section
(net/proto badge, IP tag for domain-less); the DNS log is relabelled
'DNS log · decisions'.
Verified: build (router tags)/vet 0, go test ok (LAN fold, IP-only host,
non-LAN skip, Closed adds bytes), panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The shared <QueryLog> is a fixed-height ticker (overflow:hidden) for Overview,
but on Insights it's a full paginated log — Load more fetched rows that were
clipped and invisible. Scoped override: .ins-logwrap .qrows now scrolls
(max-height min(60vh,540px), overflow-y:auto). Overview ticker unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A 25s watchActiveProfile loop (mirrors watchNewDevices) reads the active
default-route dev (ip route show default, lowest metric), matches it against
enabled profiles' MatchIface (via netplane.IfaceDevice, so UCI-name OR device
lists work), and pins the highest-Priority match into Globals.ActiveProfile +
Reconcile — generate already applies an explicit ActiveProfile, so no generate
change. Anti-flap (write only on change), no-op when no MatchIface profiles
exist, releases a stale iface-pin but preserves a manual non-iface pin,
fail-safe on every error. Pure helpers pickIfaceProfile/desiredActiveProfile/
parseDefaultRouteDev unit-tested (failover-flip, tie-break, stale-release).
Schedule-window check deferred (unexported in generate). VM: build + matcher
tests PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Shared foundation: Engine.HTTPClient(via) dials through a running-box outbound
(OutboundManager.Outbound(tag).DialContext); via->tag map direct/group:/node:/
egress:/chain:. #1: Alert gains Via + Fallback — notifier sends through the
chosen detour via an injected client factory (daemon wires eng.HTTPClient);
on detour failure retries direct iff Fallback (else surfaces error). Direct
stays the default + the always-available safety path (killswitch/apply_fail
should keep Via empty or set Fallback). #8: Subscription gains FetchDetour;
new POST /api/subscription/update {name} makes the DAEMON fetch a sub through
the tunnel (FetchVia=proxy → Fetch(sub, HTTPClient(FetchDetour))) → update →
WriteUCI → reconcile; panel gets a per-sub Detour picker (under Proxy) + an
Update-now button. CLI 'sub update' stays direct.
Verified: root+shater build (router tags)/vet 0, go test ok (via->tag map,
alert Via+Fallback fallback-to-direct, detour-fail-no-fallback drops, model
round-trip w/ via/fallback/fetch_detour), VM box.New + engine tests PASS.
panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DNS events carry no client IP, so per-device DOMAIN stats need connection
events. box.go now builds+registers the trafficcontrol.Manager + AppendTracker
UNCONDITIONALLY (moved out of the needObservable gate) — an in-process
connection observable with NO clash/api port opened (Emit is non-blocking, so
an unsubscribed tracker never stalls the hot path). Engine.ConnManager()
exposes it (box-owned; pointer changes each Apply swap). stats connLoop
subscribes (pointer-identity resubscribe like dnsLoop), folding
{Source.Addr, Domain||Destination.Fqdn} into deviceDomains (bounded 512
clients / 200 domains-each). Snapshot gains device_domains
[{ip,name,domains:[{domain,count}]}]; Insights shows a per-device domain view.
Configurable retention: Globals StatsRingSize/StatsTimelineMinutes/
StatsMaxDomains/StatsRetentionDisabled (0=built-in defaults 200/60/5000);
stats.New resolves them, RetentionDisabled skips all pruning (RAM-bounded);
Settings gains a Statistics-retention section with a disable-trim toggle.
Verified: root+shater build (router tags)/vet 0, go test ok (conn-event fold,
retention, TestEngineConnManagerWired drives a real proxied conn + asserts a
live ConnectionEventNew + pointer-change-on-swap), VM box.New still Applies
with the tracker wired. panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Globals reads the canonical 'loglevel' key; a natural 'log_level' misspelling
was silently ignored, so 'log_level=debug' produced no debug output (found
while diagnosing node-health telemetry on the VM). applyGlobals now accepts
log_level as an alias, loglevel keeping priority. +TestLogLevelAlias.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
New Insights nav page surfacing the /api/stats Snapshot that was collected but
never shown: traffic overview (queries/blocked/allowed + timeline sparkline +
filtered%), top domains ranked with blocked portion (all/blocked toggle —
replaces the anemic 'top blocked = none'), per-rule traffic bars (bytes/packets
per routing rule), per-endpoint (outbounds + resolvers by count), per-device
bytes, and the query log (block/proxy/pass) via getStatsLog pagination. Answers
the user's ask: how often & how much traffic goes to which domain/IP, per rule
and per endpoint. Pure frontend — all data already in the aggregator. Polls
getStats every 3s; honest empty states; responsive (no horizontal body scroll).
Reuses SegMeter/QueryLog/Led/Button. Wired via router ROUTES + App Page switch
+ pages/index. Verified: tsc --noEmit clean, npm build ok.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The Overview 'N/N up' was cosmetic (enabled-count/total). Now the engine
pre-registers a shared urltest.HistoryStorage in the box ctx (mirrors the
dnstrack.Manager pattern; box.go reuses a ctx-provided store), exposes it via
Engine.URLTestHistory(), and stats collectNodeHealth() reads per-node
Delay/alive into a new Snapshot.NodeHealth ([]{tag,delay_ms,alive,tested,
age_seconds}, tag==node name). Panel joins it by name: Overview shows an
honest alive/tested/total readout + status LED; Nodes rows get a latency chip
+ alive/down/untested LED (untested = node not in any probing group, shown
'—' not 'down'). Falls back to the old count when node_health is absent.
Only urltest/least_test groups populate history (selector/single/manual do
not); a failed probe deletes the entry, so tested=false conflates never-probed
and last-probe-failed — both reported untested, never a false 'down'.
Verified: go build (router tags)/vet 0, go test engine+stats ok (new
TestNodeHealthFromURLTestHistory + nil-safe test), panel tsc/build clean.
VM live-verify pending (ssh-manager MCP disconnected mid-session).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Post-test feedback batch, 4 parallel Opus agents:
#6 rollback-hide: apply.Status gains can_rollback (armed commit-confirm
snapshot OR engine.HasLastGood(), a new non-mutating engine probe). Apply/
Overview hide the Roll back button when nothing to revert; the 'Nothing to
roll back' dead-end is gone.
#10 devices: parseNeigh now captures the 'dev' token and drops non-LAN
rows (WAN device + a fail-open 10.0.2.0/24 slirp guard), so QEMU WAN IPs
10.0.2.2/.3 no longer masquerade as devices. Discovered gains network/iface
labels (IP matched against /etc/config/network subnets); UI shows 'LAN·br-lan'.
#5 egress iface picker: new GET /api/interfaces (netplane.Interfaces via
ubus network.interface dump); Targets EGRESS interface field is now a select
of real UCI interfaces (degrades to free-text when empty).
#11 WG/AWG import: parse.WGToURI serializes a *Proxy back to a canonical
wireguard:// URI (round-trips ParseWGConf, all AmneziaWG knobs); new
POST /api/import-wg converts a pasted .conf; Nodes add-node accepts a
multi-line [Interface] config and imports it as a node.
#7 nodes search/grouping: live search (name/proto/host/sub) + collapsible
per-subscription and Manual groups (large groups collapsed by default,
search auto-expands matches).
Verified: go build (router tags)/vet/test 0; panel tsc/build clean; new
tests TestCanRollback, TestEngineHasLastGood, TestDiscoverDropsWANNeigh,
TestWGToURIRoundTrip, import-wg + interfaces api tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Profiles/presets were modelled but ignored. Now buildRoute applies them to
an EFFECTIVE rule set (input model never mutated). Active profile: explicit
Globals.ActiveProfile (existing+enabled) always wins; else auto-select the
highest-Priority enabled profile whose schedule window holds (via b.now,
sharing scheduleWindowActive with rule schedules); iface/probe-conditioned
profiles are skipped by auto-select (warn, control-plane Phase-2b) but
honored when pinned. Overrides: EnableRules/DisableRules (disable wins),
DefaultTarget/DefaultEgress on Final (target wins = leak-safe). Preset packs
block-ads(15 domains->block)/ru-bypass(geoip:ru->direct, inert w/o geodata)/
private(RFC1918+ll+lo->direct) inject rules through the SAME rule loop,
Preset.Order/Target overridable. Fail-open throughout; profile/preset-free
models byte-identical (regression-guarded). Verified build/vet 0, host tests,
VM box.New (TestProfilePresetAppliesCleanly).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Routing page now manages rule-sets (domain/ipcidr match sources): add/edit/
delete with source inline(entries)|url|file|geosite, and a dst_ruleset
checkbox picker in the rule form so a rule matches one or more rulesets
(matcher chip shows 'ruleset: ..'). Deleting a ruleset strips it from every
referencing rule. Reuses the existing save->apply machinery; URL tokens
masked; honest empty state. Backend materialisation landed in 57343693.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
resolveChainExit collapsed a chain to its exit hop, losing the multi-hop.
Now buildChain materialises per-chain hop-outbound copies Detour-linked
backward (h_n exits, h_n.Detour=h_{n-1}, .. h1.Detour=direct) so traffic
traverses L1..Ln and egresses at Ln; a rule/egress routing chain:<name>
targets the chain ENTRY tag. Per-chain copies keep base node/group
outbounds standalone and preserve the 0xff loop-guard mark. Group hops =
chain-local urltest over detoured member copies; WireGuard hops = detoured
endpoint copies. 1-hop == that hop; undefined/empty/unresolvable warns +
rule skipped (never aborts box.New); only referenced chains materialise.
Verified: build/vet 0, host tests + VM box.New (TestChainMultiHopApplies).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Move dnsBlockRuleForSrc to devices_test.go (imports option/C, untagged) and
drop the now-unused constant import from devices_linux_test.go, fixing the
cross-compile of the previous test fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
TestDeviceRulesValidate used dnsRuleForSrc (first source match) which
returned the device's ALLOW rule (route action, emitted first) while
asserting Predefined — a false failure. The generation was correct
(Phase-6 verified block works E2E). Now asserts the predefined-NXDOMAIN
block rule for the src specifically via dnsBlockRuleForSrc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Reaches v0.1's config-surface parity for the backend-ready features. Nav
gains Targets, Profiles, Settings (Profiles is a placeholder until its
backend lands).
- Targets page (groups/chains/egresses — was UCI-only): GROUPS editor
(source subscription|manual, subscription/member-node pickers, strategy,
include/exclude/proto/country filters + dedup, probe url/interval); CHAINS
editor (ordered hop list from group:/node:, signal-path viz); EGRESSES
editor (type interface|proxy|direct|block|byedpi, interface/target/port +
native DPI preset off|fragment|record|spoof). All pickers derive from live
config.
- Settings page (globals — was UCI-only): enabled, log level, kill-switch,
DNS mode, IPv6, confirm timeout, panel port, health probe url/interval,
fwmark/table base (hex, advanced), read-only schema/active-profile.
- Nodes page: per-subscription options expander — update interval, fetch-via,
format, UA, HWID + device fields, extra headers, include/exclude/proto/
country filters, dedup, expire-alert days (secrets masked, reuses save
machinery).
- api.ts: full types (Subscription/Group filters, Chain, Ruleset, Preset,
Profile, Inbound, Globals.PanelPort) + Model slices; router nav.
All save→apply like the other pages. tsc clean; build ok (82 kB gzip).
Remaining for full parity (next waves): generate for rulesets/chains(real
multi-hop)/profiles/presets, and the Profiles + Rulesets panel pages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Lets the operator choose which proxy path DNS goes through — a group
(balancer/urltest), a chain, an interface/egress, a specific node, or
direct. The backend already resolved resolver.Detour via resolveTarget
(group:/chain:/egress:/node:/direct); this exposes it in the panel (the
RESOLVERS section was read-only).
- Per-resolver detour <select> built live from the Model: Direct + a
Groups optgroup (balancer) + Chains + Interfaces/egresses (with type) +
a Nodes optgroup. Current path rendered as 'via group/chain/node/
interface <name>' or 'direct'.
- Add resolver (name + type doh/dot/plain/tcp/local/fakeip + conditional
address + fakeip pool + detour), edit, delete (repoints/clears default+
fallback), and editable default/fallback role selects (were read-only).
- Stale/missing detour target stays selectable + flagged '(missing)';
legacy bare-name detours normalized against the catalog. Secrets masked.
save->apply banner like the other sections.
Verified: tsc --noEmit clean; npm run build ok (71 kB gzip JS); Playwright
(mock) — the detour select shows Direct/Group auto (balancer)/egresses/
Nodes optgroup, changing it + add/delete + default/fallback all work with
the save->apply banner.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A group with source=subscription produced NO members (groupMembers only
read the explicit g.Nodes list, empty for sub-backed groups), so the group
was skipped and any rule targeting it never routed — the whole proxy path
was dead on a subscription setup.
Fix: for source=subscription, gather members from all nodes where
FromSub==g.Subscription (enabled + emitted), in config order, deduped; apply
Include/Exclude name regexes (case-insensitive, bad pattern warns+ignored)
and FilterProto/FilterCountry/Dedup via parse.ParseShareLink +
FilterSpecFromGroup + ApplyFilters. Manual/single/'' sources unchanged.
Verified: generate tests (sub group gathers exactly its FromSub nodes not
others; include/exclude; bad-regex-ignored; proto filter; manual unchanged).
Found live on the VM: 376-node subscription group 'auto' was 'no usable
members, skipped' before this fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Makes shaterd sub update real (was a Phase-2b stub): fetch a subscription
URL, parse+filter into nodes, persist them, reconcile.
- shater/subscribe: Fetch(sub, client) — HTTP GET with UA (sub.UA or
Shater/0.2 default), HAPP-style x-hwid/x-device-* headers, raw Headers
override, 20s timeout, 8MiB cap, non-2xx/empty = error. UpdateSubscription
(pure): parse.ParseSubURIs re-serializes ALL formats (clash/xray/sing-box/
links) to canonical share-links, pairs each with its *Proxy, ApplyFilters
(Include/Exclude/proto/country/dedup), builds []Node FromSub=<name> (name
from #fragment, collision-suffixed), REPLACES only that sub's cache. Zero
usable nodes -> error + leave the cache intact (a provider hiccup never
empties the config).
- cmd/shaterd: 'sub update [<name>]' — fetch each enabled sub (or one),
fold in, WriteUCI, best-effort SIGHUP reconcile; works with/without the
daemon; fetch_via=proxy warns + falls back to direct (MVP).
- model: sub-cache nodes now persist in UCI (config node + from_sub/
fingerprint/stale) so the fetched set survives restarts and the panel sees
them; ReadUCI reads them back; manual nodes unaffected. (v0.1 used a
separate JSON cache; UCI persistence matches v0.2's model<->UCI design.)
Verified: subscribe+model unit tests (FromSub tagging, filters, zero-node
safety, cache replace-keep-others, name collisions, UA/header/non-2xx); VM
E2E with the real feed https://pro.qomar.pw/sub/... -> 'sub update' fetched
376 nodes (261 vless/71 ss/29 vmess/15 trojan), all from_sub='default',
persisted to UCI, and box.New/Apply accepted all 376 (status running/active).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Ports the v0.1 Gitea release flow to the v0.2 single-binary + 4-package
layout, so a tag publishes a signed opkg feed the routers install from.
- .gitea/workflows/release.yml: on tag v* (+ dispatch), matrix over
{x86_64, aarch64_cortex-a53}. Per arch: setup Go 1.24/Node 20/UPX ->
scripts/build-shaterd.sh (SPA-embedded shaterd, stages the .upx) ->
ci/build-feed.sh (OpenWrt SDK container builds all 4 packages ->
usign-signed Packages index). A release job merges both arches into one
signed feed + publishes the rolling 'latest'/tag release via the Gitea API.
- ci/sdk-build.sh: in-SDK build — add openwrt/ as the 'shater' feed, feeds
update/install, make package/{shaterd,shater-core,byedpi,luci-app-shater}/
compile (shaterd validates+installs the staged prebuilt; byedpi cross-
compiles from source). ci/make-index.sh: opkg Packages(.gz) + usign sign
with KEY_BUILD (keyfile umask 077, no secret hardcoded), verifiable by
dist/shater-feed.pub. ci/install-usign.sh + ci/gitea-release.sh ported.
- INSTALL.md: add the signed feed src/gz line + import dist/shater-feed.pub
to /etc/opkg/keys; apk (25.12) path noted.
Key kept: usign feed key 5ac4b177689cb8e0 (public dist/shater-feed.pub,
secret Gitea repo secret KEY_BUILD). Decision: opkg (24.10 uses opkg; apk
is 25.12) — matches the existing usign trust anchor.
Verified structurally (no live runner here): release.yml is valid YAML, all
ci/*.sh are bash -n clean, no hardcoded secrets, and every package name/
path/arch/artifact/secret reference cross-checks against openwrt/, scripts/
build-shaterd.sh, and dist/shater-feed.pub. Live-runner unknowns (full SDK
compile of the 4 packages, router-side signature verify) flagged in-agent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Makes the whole product installable — shater-core DEPENDS +shaterd, and
this is what resolves it.
- scripts/build-shaterd.sh: the release build. Builds the panel SPA
(npm ci && npm run build), copies panel/dist -> shater/panel/webroot
(the go:embed dir), cross-builds shaterd for amd64 + arm64 with the D9
router tag set (CGO_ENABLED=0, -checklinkname=0 -s -w, static ET_EXEC no
PT_INTERP), then UPX --lzma --best (D10) and stages the .upx into
openwrt/shaterd/files. Version from arg/SHATER_VERSION/git-describe.
Measured: amd64 40.3MB->10.4MB, arm64 37.6MB->8.5MB.
- openwrt/shaterd: prebuilt-binary package (npm+embed+UPX don't reproduce
cleanly in the SDK, so CI stages the artifact). Maps OpenWrt ARCH
(x86_64->amd64, aarch64->arm64 = both BPI routers) to files/shaterd-<a>.upx,
installs /usr/bin/shaterd. RSTRIP/STRIP disabled (the SDK strip would
corrupt the UPX binary); DEPENDS empty (static); errors clearly when no
artifact is staged. GPL-3.0-or-later.
- docs-shater/INSTALL.md: build + install order (shaterd -> shater-core ->
luci-app-shater, optional byedpi) + enable/apply.
- gitignore: dist/shaterd-*, openwrt/shaterd/files/*.upx, panel webroot.
Verified on the OpenWrt musl VM: dist/shaterd-amd64.upx (10.4MB) decompresses
into RAM + runs (shaterd status OK), serves the REAL embedded Faceplate SPA
at :8088 ('SPA embedded=true', real Vite index.html + assets — not the
placeholder).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Out-of-band notifications on key events, delivered DIRECT to the internet
(plain net/http, never via the proxy) so a kill-switch/engine-down alert
reaches Telegram even when the tunnel is down.
- model: Alert{Name,Enabled,Type(telegram|webhook),Token,ChatID,URL,
Events[]} + Model.Alerts; uci (case alert, list event) + render +
round-trip fixture.
- shater/alert: Notifier — 8s-timeout default-transport client, per-alert
async delivery (telegram sendMessage / webhook JSON POST), (event,title)
dedup within 60s so a flapping engine can't spam, panic-safe. Update()
swaps config on reconcile; TestFire() for the test verb.
- cmd/shaterd: builds the Notifier in run, Update()s it after each
reconcile; fires apply_fail+killswitch on an apply/reconcile error while
enabled (initial + SIGHUP paths); watchNewDevices polls devices.Discover
every 45s and fires new_device on an unseen MAC (skips the startup
baseline). 'alert test' verb sends a test to every enabled alert (reads
UCI, no live daemon needed). Events emitted now: killswitch|new_device|
apply_fail; node_down|sub_expiry reserved.
- panel: Alerts section in DNS.tsx — list (name/type/events, enable, delete)
+ add form; Token/URL NEVER shown in clear (masked, round-tripped). api.ts
Alert type + Model.Alerts.
Verified: model round-trip incl. Alert; notifier unit tests (subscribed
delivers correct JSON, unsubscribed/disabled deliver nothing, dedup
suppresses rapid dup); build/vet; panel tsc+build; VM — 'shaterd alert test'
with a webhook alert delivered a POST to a local sink.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
devices_linux_test.go (generate box.New validation of per-device rules) and
panel/devices_test.go (/api/devices gating) were authored with the Phase-6
backend (677a1dde) but not staged in that commit. No source change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Scheduled rules are now evaluated by the control-plane at gen/reconcile
time (the engine has no time match), so a rule is only active inside its
window.
- generate: builder gains an injectable 'now' (defaults time.Now); a
SchedEnabled rule outside its window is SKIPPED (was emitted
unconditionally with a deferred warning). schedule.go: scheduleActive
parses SchedDays (mon..sun, empty=all), SchedStart/End HH:MM in SchedTZ
(LoadLocation, fallback local), handles overnight windows (end<start ->
now>=start || now<end), all-day (empty/equal end). Invalid HH:MM ->
fail-OPEN (emit + warn) so a typo never silently drops protection.
- cmd/shaterd: real 'schedule due' verb -> SIGHUP reconcile (no-op when
down); generate re-evaluates + the config-hash gate rebuilds the engine
only when a window boundary was actually crossed (no churn between
boundaries — verified reconcile changed=false per tick).
- shater-cron: calls 'shaterd schedule due' each tick (cheap, hash-gated).
- panel Routing: add-rule form gains schedule controls (enable + mon..sun
day toggles + From/To time inputs); scheduled rules show a days+time chip.
Verified: schedule unit tests (weekday window emitted/skipped, overnight
active across midnight, all-day weekend, invalid-time fail-open); build/
vet; panel tsc+build; VM (box.New accepts a scheduled config; 'schedule
due' reconciles changed=false = no churn).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
In-process stats fed by the engine's DNS-query event stream + nft counters.
- engine: DNSQueryManager() accessor; engine.New pre-registers a stable
*dnstrack.Manager into e.ctx (box.New only creates one when an api/
clash_api observable is present, which the router config has none of, so
the manager would be nil — pre-registering keeps the DNS stream alive).
- shater/stats: Aggregator subscribes to dnstrack QueryEvents and maintains
bounded top-domains, allowed-vs-blocked (blocked = NXDOMAIN / 0.0.0.0 /
failed), a 60-min timeline, a 200-entry live query-log ring, per-server
counts; polls netplane.ListClients/ListCounters for per-device + per-rule
traffic (client IP -> DHCP hostname). Snapshot()/RecentQueries(); re-subs
on box swap; resilient when the engine is down.
- daemon: creates+starts the aggregator, Resubscribe() after each reconcile,
Close on SIGTERM; control-socket 'stats' verb returns the real snapshot.
- panel: GET /api/stats (snapshot) + GET /api/stats/log?n= (live log),
session-gated; Stats type + getStatsLog() in api.ts; Overview QueryLog now
polls the live log, plus a DNS-filtering module + blocked SegMeter + top-
blocked list. Honest empty states, no fabricated data.
- upstream (minimal, marked // lx/D15): dnstrack SourceFiltered +
emitFilteredResponse at the two DNS-filter predefined-block sites in
dns/router.go — filter blocks now feed the query stream (were invisible).
Verified: stats+panel unit tests; panel tsc+build; VM E2E — DNS traffic
from a netns client produced /api/stats totals (queries 16, blocked 6),
top_domains[blocked-ad.example blocked 6], per-device row, and /api/stats/log
rows with correct block/allow; stream survived a box swap (SIGHUP). Overview
screenshot shows the live query log + blocked stats.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The DNS filter blocked via reject/default, which sing-box answers with
REFUSED — non-standard for ad-blocking and contradicting the Blocklist
Response field (nxdomain|zero) + the code comments. Switch to sing-box's
predefined DNS action:
- Response nxdomain (default) -> predefined Rcode NXDOMAIN (RcodeNameError).
- Response zero -> predefined Answer A 0.0.0.0 (+ AAAA :: when ipv6), owner
'*.' so one record serves every domain in a many-domain rule-set. This
finally implements 'zero'; the old 'not supported' warning/fallback is
gone.
- Remote rule-set fetch: DownloadDetour (deprecated in sing-box 1.14) ->
HTTPClient{DialerOptions{Detour: direct}} — no deprecation at box.New.
Verified on the VM via the live in-engine DNSRouter.Exchange: an nxdomain
blocklist -> NXDOMAIN(3); a zero blocklist -> NOERROR(0) + A 0.0.0.0 (not
REFUSED, not NXDOMAIN); a url remote blocklist with http_client validates
with zero deprecation notices. box.New tests cover nxdomain/zero/zero-no-
ipv6/remote/live.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Completes Phase 4 and passes its gate.
- generate: emit experimental.cache_file (enabled, /etc/shater/cache.db,
/tmp fallback) so remote rule-sets persist + auto-update and megalists
stay RAM-sane. Daemon + uci-defaults create /etc/shater.
- cmd/shaterd: real 'blocklist update' verb — SIGHUP-reconcile the running
daemon so url/file rule-sets re-fetch (cache_file updates); no-op when
down. shater-cron fires it on the blocklist interval.
- engine (D16): cache_file's bbolt EXCLUSIVE lock broke the apply-swap —
the new box couldn't take the lock the old held, so every live reconcile
stalled ~10s then failed 'cache-file timeout' (edits silently ignored).
Apply now treats the cache-lock timeout as a swap conflict AND proactively
goes close-old-then-start-new when the incoming config shares the running
cache_file (sharesCacheFileLock), no stall. Regression test added.
Phase-4 gate PASSED on the OpenWrt VM (netns client, dns_filter on):
blocked-ad.example + doubleclick.net -> blocked (reject, no answer);
example.com -> resolves via the engine resolver; enabling an allowlist
entry + 'blocklist update' -> doubleclick.net resolves (allow overrides
block); a real geosite ads megalist (.srs) loaded and its domains blocked
with shaterd RSS ~39 MB (sane); 'blocklist update' reconciles cleanly.
Follow-ups (flagged, not blocking): sing-box reject returns REFUSED not
NXDOMAIN (comment says NXDOMAIN — a predefined-NXDOMAIN action would match
the usual ad-block convention); remote rule-set download_detour is
deprecated in sing-box 1.14 (works, rename before 1.16).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The DNS-filter control surface, wired to the Phase-4 backend.
- DNS FILTER master Toggle (Globals.DNSFilter) with a network-wide
ad/tracker-blocking label + DNS mode / default+fallback resolver readout.
- BLOCKLISTS: per-list source badge (inline/url/file/geosite), URL host
(token masked) / inline entry count, NXDOMAIN|0.0.0.0 reply, enable
toggle + delete, 'filter off' badge when a list is on but the master is
off. Add form (name + domains-textarea|URL + reply). Quick-add chips seed
StevenBlack / OISD / AdGuard with canonical URLs (dedup-guarded).
- ALLOWLISTS: same pattern (overrides blocklists).
- RESOLVERS: read-only display (name, type, host masked, detour) — editing
deferred, noted on-screen.
- save->apply split (putConfig -> toast + banner -> apply) like the other
pages; secrets never rendered; honest empty states.
api.ts promoted: Blocklist/Allowlist interfaces, Model.Blocklists/Allowlists,
Globals.DNSFilter (from the page's local decls). Wired into App.tsx + index.
Verified: tsc --noEmit clean; npm run build ok (62.8 kB gzip); Playwright
screenshot (mock) confirms Faceplate fidelity + the add/quick-add/empty
states. DNS is no longer a placeholder — 5 of 6 nav pages are live (Devices
awaits Phase 6).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
In-engine DNS blocklist/allowlist filtering built on sing-box's compiled
rule-set matcher (no custom megalist matcher, per D5/D15). Rides the same
in-engine DNS plane as the hijack-dns funnel (D14).
- model: Blocklist{Name,Enabled,Source(inline|file|url|geosite),URL,Path,
Entries,Response(nxdomain|zero),UpdateInterval} + Allowlist; Model gains
Blocklists/Allowlists; Globals.DNSFilter master enable (opt-in, default
off). uci parse + render; round-trip fixture extended.
- generate/dnsfilter.go: each enabled list -> a rule-set (inline for
Entries as domain_suffix so subdomains match; remote for url w/
DownloadDetour=direct so fetches don't blackhole under kill-switch;
local for file; geosite skipped inert when no geodata). DNS rules
prepended: allow FIRST (rule_set:[al-*] -> route to default resolver,
terminal, so allowlist overrides), block SECOND (rule_set:[bl-*] ->
reject NXDOMAIN). zero-response -> NXDOMAIN + warn (predefined answer
needs a per-query name a many-domain rule-set can't carry). Off/empty ->
emits nothing.
- shater-core config: commented StevenBlack/OISD/AdGuard blocklists +
allowlist examples + dns_filter note, all inert.
Verified: model round-trip + generate tests; VM box.New (Apply+Start) of
DNSFilter=true + inline blocklist + allowlist + tproxy + doh resolver ->
valid sing-box config. List fetch/compile + 'blocklist update' verb + the
block-a-domain-E2E gate are the next task (remote rule-sets self-fetch).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Closes two gaps flagged by the LuCI launcher:
- panel port was hard-coded :8088 on both sides. Now globals.panel_port
(0 = default 8088): model parses/renders it; cmd/shaterd binds
:<panel_port> when set, else falls back to SHATER_PANEL_ADDR (env can
still disable). apply.Status + shaterd status now report panel_port so
LuCI builds the Open-panel redirect from it (fallback 8088), not a
constant.
- apply.Status gains kill_switch (from globals) so the readout/LuCI can
show fail-closed vs open. LuCI dashboard adds a kill-switch LED
(closed=green, open=amber).
api.ts Status type + mock updated to the new json keys (kill_switch,
panel_port). Round-trip fixture updated (PanelPort) and still passes.
Verified: build+vet, model+panel tests, panel tsc, node --check dashboard;
VM — with option panel_port '8090' the daemon binds :8090 (not 8088),
status reports panel_port:8090 + kill_switch:closed, /api/status on :8090
returns 401 without a session (listening).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The reachability layer for the admin panel (ARCHITECTURE §2): LuCI — an
already-authenticated, ACL-checked session — mints a single-use handoff
token and opens the embedded panel with a session, so the panel needs no
login of its own.
- shaterd: new 'mint-token' verb — dials the running daemon's control
socket, prints its {token} JSON. Robust bridge for the rpcd plugin
(stock OpenWrt has no AF_UNIX client: busybox nc lacks -U, no socat).
No existing verb touched.
- luci-app-shater: client-JS LuCI app under Services —
* rpcd exec plugin /usr/libexec/rpcd/shater: ubus object 'shater' with
status (shaterd status passthrough) + mint_token (shaterd mint-token).
* acl.d: least-privilege (status=read.ubus, mint_token=write.ubus,
scoped to the shater object).
* view dashboard.js: 5s-poll LED status grid + prominent 'Open panel'
button — mint_token via rpc.declare, then open http://<host>:8088/?t=
<token>; button disabled + reason when the daemon is down; tab opened
inside the click gesture so popup blockers don't kill it.
* menu.d entry, luci.mk Makefile (LUCI_DEPENDS +shater-core +rpcd,
PKGARCH all, GPL-3.0-or-later), uci-defaults (chmod plugin +x, reload
rpcd, clear luci cache).
Verified on the OpenWrt VM: shaterd status/mint-token JSON; rpcd plugin
list/call paths; and the full handoff — GET /?t=<token> -> 303 + HttpOnly
SameSite=Strict cookie -> /api/status 200; no cookie -> 401; token reuse
-> no second cookie (single-use). sh -n clean, ACL/menu JSON valid.
Known gap (flagged): panel port is hard-coded :8088 (SHATER_PANEL_ADDR
default) — a future globals.panel_port UCI field would let both sides
share the source. apply.Status has no kill-switch field yet (view shows
the available fields, no fabrication).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Three parallel-built Faceplate pages wired into the shell (Overview was
already live; DNS/Devices stay placeholders until Phases 4/6):
- Nodes.tsx: node + subscription management. Node list with enable toggles,
protocol pills (from the share-link scheme), MANAGED/STALE tags, add-from-
share-link, delete; subscriptions add/toggle/delete. Secrets NEVER shown —
protocol + masked host only; sub URLs show host with 'token hidden'. Sub
on-demand-refresh is wired-but-disabled (needs a backend endpoint).
- Routing.tsx: first-match rule list on a 'signal bus' rail with order
steppers, matcher-summary chips, target chips (group/node/egress/direct/
block), enable toggles; add-rule form with a target picker derived from the
live config; catch-all rule visually distinguished as route Final.
- Apply.tsx: commit-confirm control room — live LEDs + config hash, Apply
with a ConfirmTimeout countdown + Confirm (else honest auto-rollback note),
Rollback with a consequences confirm, before->after hash readout.
All three consume shater/panel's API (getConfig/putConfig/apply/confirm/
rollback) with the save->apply separation, honest empty/error states, no
fabricated data. api.ts Rule widened with the full field set (Src/Dst*/Proto/
Kill/Sched*) so pages share the type; wired into App.tsx's page switch +
pages/index.ts.
Verified: tsc --noEmit clean; npm run build ok (59.5 kB gzip JS, still
react+react-dom only); Playwright screenshots of Nodes/Routing/Apply (mock
backend) confirm Faceplate fidelity + real API-driven rendering.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Replaces the static showcase with the real wired SPA on the fixed
Faceplate design. Stack stays react+react-dom only (52 kB gzip JS).
- src/api.ts: typed same-origin client (Status + Model sections from
shater/model), credentials:include, ApiError w/ 401 -> unauth state.
- src/session.ts: ?t=<token> handoff -> POST /api/session -> scrub token,
preserve hash route (matches ARCHITECTURE §2 / the panel server bridge).
- src/router.ts: ~25-line hash router (useSyncExternalStore), no deps.
- src/App.tsx: Faceplate shell — header (master status LED from /api/status
+ Clock + ThemeSwitch), engraved 6-tab nav (Overview live; Nodes/Routing/
DNS/Devices/Apply placeholders for the next page-agents), footer statusbar
(nft LED + engine hash), loading/ready/unauth/error state machine + 5s
status poll. Reuses the existing Faceplate components.
- src/pages/Overview.tsx: REAL data from /api/status + /api/config — engine/
config/data-plane/kill-switch LEDs, module cards with live counts, a
QueryLog polling /api/stats that degrades to an honest empty state (no
fabricated stream), and an Apply/Confirm/Rollback control row wired to the
endpoints with inline result + toast.
- src/mock.ts: ?mock dev fixture backend (no daemon needed); vite /api proxy
to 127.0.0.1:8088 otherwise.
Contract for the remaining pages: add pages/<Name>.tsx, export from
pages/index.ts, switch in App.tsx; read/write via api.ts; shell owns
header/nav/footer/auth.
Verified: tsc --noEmit clean; npm run build ok (52 kB gzip); Playwright
screenshots of Overview in light + dark + mobile confirm Faceplate fidelity,
real API-driven counts/LEDs, visible focus, reduced-motion respected, and
the apply flow (inline 'reconciled' + toast + hash refresh).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Lets the panel EDIT config, not just read it.
- model.RenderUCIExport(m): pure inverse of ParseUCIExport (Model ->
'uci export shater' text), field-for-field on the same uci keys, shipped
anonymous-section + option-name convention, uci single-quote escaping.
Round-trip invariant ParseUCIExport(RenderUCIExport(m))==m holds for a
rich fixture (every section type, lists, bools both ways, ints/hex,
embedded quote). Bools always emitted (missing != false for default-true
fields); sub-cache nodes (FromSub!='') skipped (runtime state, not UCI).
- model.WriteUCI(m): delete-then-import ('uci delete shater' -> 'uci import
shater' <text> -> 'uci commit shater') so it REPLACES rather than appends
(busybox uci import merges). Via the uciRunner seam (extended with
Import); only WriteUCI touches uci, Render is pure.
- panel PUT /api/config: session-gated, decodes a Model, light validation
(manual node must carry a URI -> 400), WriteUCI, returns {ok,applied:false}
— editing does NOT auto-apply; client calls POST /api/apply after.
GET/PUT method-dispatch on the same path.
Verified: round-trip + write-replace-idempotence + PUT handler unit tests,
and VM E2E (mint->session->GET config->PUT adds a node->200; uci export
shows the new anonymous config node, +1 exactly no duplication, migrate
re-parses cleanly; original config restored).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The HTTP server the admin panel + thin LuCI consume, running INSIDE the
shaterd daemon and sharing its single *apply.Applier (no second engine).
- shater/panel: net/http server (no framework deps). JSON API under /api:
GET /api/status (apply.Status + version), GET /api/config (current
model.ReadUCI), POST /api/apply (snapshot -> reconcile -> arm-rollback),
POST /api/confirm, POST /api/rollback, GET /api/stats (Phase-5 stub).
- Auth per ARCHITECTURE §2: LuCI mints a single-use short-TTL token over
the daemon's unix control socket (new 'mint-token' verb -> MintToken);
POST /api/session {token} validates+consumes it and sets an HttpOnly,
SameSite=Strict session cookie; all other /api routes require it (401
otherwise). Also a GET /?t= redirect bridge matching the §2 diagram.
In-memory token/session stores with expiry; no external deps.
- Serves the embedded Faceplate SPA (go:embed all:webroot; build copies
panel/dist -> shater/panel/webroot, gitignored w/ .gitkeep so it compiles
on a fresh checkout, placeholder page when unbuilt). SPA fallback; unknown
/api/* -> JSON 404, never index.html.
- cmd/shaterd: cmdRun starts the panel server in a goroutine (bind failure
log-and-continue like the control socket), closed on SIGTERM. Gated by
SHATER_PANEL_ADDR (default :8088; off=disabled) to avoid widening the UCI
contract now.
Note: config MUTATION endpoints are intentionally NOT added yet (uci/LuCI
is the writer); /api/apply operates on current UCI. Router build uses the
D9 tag set (with_purego forces a glibc PT_INTERP, unusable on musl).
Verified: unit tests (401 no-session, session->cookie, single-use replay
401, expired/invalid 401, SPA serve) + VM E2E on the OpenWrt VM (mint ->
session -> cookie -> /api/status 200 -> replay 401 -> / serves SPA).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
shater routes marked packets via 'ip route add local default dev lo table
<N>'. Local-delivery of a packet whose source is on a directly-connected
subnet requires accept_local=1 on the LAN INGRESS interface — rp_filter=0
alone is not enough. Sysctls() set lo.accept_local=1 but never the LAN
iface, so with the shipped set a real LAN client behind br-lan is BLOCKED
(engine up, table applied, yet the tproxy'd packet never reaches :12345 —
it escapes to the fail-closed forward drop and the LAN goes dark). The
Phase-2 gate only passed because a stale accept_local was left set on the
test VM.
Fix (scoped, least-privilege — not a global 'all' change): new
netplane.ApplyIfaceSysctls(m) sets, per enabled tproxy ingress device
(nftEnabledInboundDevs), net.ipv4.conf.<dev>.accept_local=1 and
.rp_filter=0 (rp_filter is MAX(all,iface), so the static all=0 can't
override an iface value of 1). Called from apply.applyLocked after
ApplySysctl, fail-closed. Static Sysctls() drop-in can't know device
names, so this is dynamic at apply time; hotplug reconcile re-applies.
Verified on the OpenWrt VM from the buggy baseline (br-lan.accept_local=0):
shater's own apply flips it to 1 and a netns LAN client that was blocked
(rc=4 'Operation not permitted') now reaches the internet (egress WAN IP).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The process behind a shater egress of type='byedpi': an optional, separate
OpenWrt package that ships ByeDPI (ciadpi) + a procd supervisor. shaterd's
generate emits a SOCKS5 outbound egress-<name> -> 127.0.0.1:<port> (Phase
2b-i); a ciadpi instance from this package listens on that port, applies
TCP/TLS desync, and goes DIRECT (no tunnel).
- Makefile: package byedpi, pinned upstream v0.17.3 (real PKG_HASH), MIT,
per-target (compiled C via SDK toolchain calling ciadpi's own make).
- init.d/byedpi: procd multi-instance (one ciadpi per enabled config
instance, 127.0.0.1:<port> + desync args), inert by default, respawn,
config-change reload, validation. sh -n clean.
- config/byedpi: default instance disabled, port 1080, a documented desync
preset. uci-defaults/40_byedpi enables the init.
- Kept SEPARATE from shater-core (byedpi egress is opt-in).
Verified E2E on the OpenWrt VM: musl-static ciadpi (146 KB, no PT_INTERP,
Alpine-built) proxies + desyncs (log: DESYNC_DISORDER); a netns LAN client
routed through a type='byedpi' egress reaches the internet direct via
ciadpi; kill-switch stays honest (SIGKILL shaterd -> client blocked).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Model ciadpi (ByeDPI) as 'just another egress' per D13: an egress of
type 'byedpi' emits a SOCKS5 outbound to 127.0.0.1:<port> (loop-guard
mark so ciadpi's own upstream isn't re-diverted), which a routing rule
targets. The desync happens inside ciadpi, so no native tls_* flags apply.
- model: Egress gains Port int (ciadpi listen port, default 1080); uci
parses option port; Type doc now lists byedpi.
- generate: type 'byedpi' egress -> C.TypeSOCKS outbound (version 5,
127.0.0.1:port) tagged egress-<name>.
The ciadpi binary + procd package + E2E is Phase 2b-ii (separate). Pure
Go half verified: unit tests + box.New validation of a byedpi egress on
the OpenWrt VM (router tags) PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
An egress gains an optional 'dpi' preset that surfaces sing-box's already-
compiled route-action desync fields, so a ruleset can go DIRECT + desynced
with no tunnel and no extra binary (the DPI-blocked-but-not-IP-blocked case):
fragment -> tls_fragment (split the TLS ClientHello record)
record -> tls_record_fragment (alternative; mutually exclusive w/ fragment)
spoof -> tls_spoof (decoy ClientHello; wrong-sequence default)
- model: Egress gains DPI string; uci.go parses option dpi.
- generate: a type 'direct' egress now emits a real 'egress-<name>' direct
outbound (loop-guard mark) so it resolves as a rule target at all — before
this a direct egress target referenced a non-existent outbound (latent bug).
buildRoute's applyDPI stamps the matching route-action flag on every rule
routed to a DPI egress; fragment<->record mutual exclusion enforced; a DPI
preset on a default/catch-all egress warns (route Final carries no action).
byedpi reserved for Phase-2b (external SOCKS egress, D13); unknown -> warn+off.
- openwrt example config + ROADMAP updated.
Verified: unit tests (fragment/record/spoof/none/unknown) + box.New validation
of fragment and spoof configs on the OpenWrt VM (router tag set) all PASS.
tls_spoof validates at box.New time on the linux/router build (raw sockets are
only touched at dial time).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The netplane dnsnat chain redirected LAN :53 to the router's :53 assuming
the engine answered there, but generate/engine create no :53 DNS server —
so on the VM the redirect landed on dnsmasq, which resolved via its WAN
upstream OUTSIDE the tunnel (DNS leak), leaving the engine's own resolvers
(built with anti-leak detours) unused.
Per D14, use the sing-box-native hijack instead of an nft redirect:
- generate/route.go: hijackDNSRule() after sniffRule — matches the sniffed
DNS protocol and steals the query into the engine's internal resolver
(C.RuleActionTypeHijackDNS), which routes each query through its detour.
- netplane/nft.go: drop the prerouting :53 accept and the whole dnsnat
chain so LAN :53 is diverted by the normal tproxy catch-all into the
engine, where hijack-dns answers it. :853 DoT reject kept (forces :53).
With :53 now diverted, DNS also fails closed when the engine is down.
Verified E2E on the OpenWrt VM: with dnsmasq STOPPED, the netns LAN client
still resolves public names (only the engine's hijack-dns could answer);
a WAN :53 forward counter stayed at 0 while the shater divert counter
climbed — real egress was the engine's encrypted DoH to 1.1.1.1:443. No
plaintext client DNS reaches the WAN.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A transparent proxy pins its tproxy inbound to a FIXED port, identical
across every apply, so the start-new-then-close-old swap ALWAYS collided
with the still-running old box on a live config change:
reconcile failed: start instance: start inbound/tproxy[in-lan]:
listen tcp4 0.0.0.0:12345: bind: address already in use
=> the new config silently never took effect. Only initial apply and
disabled->enabled worked (no old listener to clash with).
Fix: keep start-new-first (zero-downtime + old-instance protection when
ports don't clash), but on a listener bind conflict against a running old
instance, fall back to close-old-then-start-new (applyCloseFirst): close
the old box to free the port, build+start a fresh box for the new opts;
the fail-closed nft kill-switch covers the brief gap. If the fresh box
can't come up, restore the previous config; if restore also fails, leave
the engine stopped (instance=nil, never a closed box) — kill-switch keeps
the LAN safe. isAddrInUse matches both errors.Is(EADDRINUSE) and the
error text (box.New may flatten the errno); compiles cross-platform.
Verified on the OpenWrt VM: edit a running config + SIGHUP now logs
'reconcile OK (changed=true)' with the engine hash advancing and NO
'address already in use'.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The inet-shater data plane relied entirely on TPROXY delivering to the
engine socket. nftables `tproxy` with NO listening socket returns
NFT_BREAK: it aborts its own rule before the trailing `meta mark set
0x2000 accept`, so the packet is left UNMARKED, falls through prerouting
policy accept, reaches the forward hook, and fw4 masquerades it to WAN.
Result: with the engine dead and kill_switch=closed, LAN clients LEAKED
straight out the WAN (confirmed on the VM: wget succeeded, conntrack
showed the flow SNAT'd to the WAN IP). kill_switch=closed only set the
engine's internal route.Final=block, which is moot when the engine is
down — there was no nft-level fail-closed layer.
Fix: in closed mode the forward chain now drops LAN-ingress traffic that
reaches it bound for a public dst (an escape, since diverted traffic is
delivered locally and never traverses forward), after accepting mgmt /
interface+tunnel egress marks and LAN-to-LAN / link-local so the LAN and
router keep working. Open mode still falls through (documented fail-open).
Folds in the old ipv6-off-closed case. Regression test asserts the v4/v6
drop is present in closed mode and absent in open mode.
Verified E2E on the OpenWrt VM (netns LAN client -> tproxy -> engine ->
AmneziaWG WARP exit): engine up, client egresses 104.28.212.73 warp=on
(direct WAN = 45.131.214.140 warp=off); SIGKILL the engine -> client is
BLOCKED (no WAN leak, no SNAT'd conntrack), whereas before this fix the
same scenario leaked to WAN.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
ParseUCIExport ignored the UCI section name for inbound/subscription/node/
group/chain/egress/ruleset/rule, honoring only 'option name'. A named
section like `config node 'ss1'` therefore yielded an EMPTY name, so
generate could not resolve a rule target 'node:ss1' and silently skipped
the rule — under kill_switch closed that BLOCKS the traffic instead of
proxying it. Now every named section falls back to the section name
(firstNonEmpty(option name, section name)); an explicit 'option name'
still wins. Matches preset/profile/resolver, which already did this.
Regression test covers all three forms for node/inbound/group.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
shaterd run is the single procd-supervised process that holds the one
box.New engine + the inet-shater data plane; every other verb is a
short-lived process that SIGNALS it (SIGHUP or the unix control socket)
and never builds a second engine.
- shater/apply: Applier drives model->generate->engine swap->netplane
(nft+routing+sysctl) under a cross-process flock, fail-closed (engine
error aborts before netplane; netplane error keeps the kill-switch up).
Reconcile (SIGHUP), honest Teardown (SIGTERM), commit-confirm
Snapshot/Confirm/ArmRollback/Rollback, ACTIVE_FLAG gating. flock.go
no-op default + flock_unix.go syscall.Flock override (cross-platform).
- shater/cmd/shaterd: run/migrate/reconcile/apply/confirm/rollback/
status/nodes/stats + sub|ruleset|schedule Phase-2b no-op stubs.
Pidfile single-owner guard; reconcile cold-start no-op (fork-storm
guard); control socket at /var/run/shaterd.ctl for reply-bearing verbs.
Verified on the OpenWrt x86_64 musl VM with the D9 router tag set
(static ET_EXEC, no PT_INTERP): daemon stays inert while globals.enabled=0
(no inet-shater table), applies on SIGHUP (table + fwmark 0x2000->shater +
loop-guard mark 0xff + tproxy divert appear, status all true, engine
hash set), and does an honest teardown on SIGTERM (table/rule/active-flag/
pidfile all gone, process exits 0).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Generate(m)/GenerateWithWarnings(m) build a full option.Options from the
neutral model: outbounds (share-link + AWG-endpoint via shater/parse),
urltest/selector groups, tproxy/mixed inbounds, route rules with the
kill-switch Final gate (closed->block, open->direct), and DNS. Every
inbound/egress carries the netplane loop-guard RoutingMark.
Verified on the OpenWrt x86_64 musl VM: engine.New().Apply (box.New +
Start) accepts AND starts all 7 cases — ss+tproxy+killswitch-closed,
AmneziaWG endpoint, 2-node urltest group, killswitch-open->direct, all
reachable share-link protocols, white-box hy2/tuic/shadowtls mapping, and
bad-node-skipped. RoutingMark validates only on Linux, so the engine
suite is //go:build linux and runs on the VM; Windows/macOS get build+vet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DPI-bypass stays a per-ruleset egress choice, never a global toggle.
Evaluated zapret (NFQUEUE packet plane) vs ByeDPI (local SOCKS desync
proxy); chose ByeDPI because it *is* an egress and composes with our
routing model with zero conflict against the verified inet-shater TPROXY
plane. zapret explicitly rejected. Native tls_fragment/spoof (already
compiled in) stay as free complementary egress presets.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Defines the in-process daemon model: 'shaterd run' owns the box, CLI verbs signal
it (SIGHUP=reconcile, SIGTERM=teardown); shater/apply orchestrates generate+engine+
netplane under flock; openwrt/shater-core supervises shaterd. MVP scope for the gate.
Engine drives sing-box in-process (D11): Apply builds+validates via box.New,
swaps atomically (start new, close old), gates on a config hash (no churn on
unchanged apply), keeps the old instance running if a new config fails validation,
and supports Rollback to last-good. Instance()/Hash() expose state for stats later.
Integration test verifies apply/no-op/swap/invalid-keeps-old/rollback/close.
No go.mod changes.
Full map of the v0.1 xrayctl/shater-core internals + the sing-box option surface,
what ports verbatim (nft/routing plane, UCI model, parsers, subsystems) vs. what is
rewritten (generator, DNS, box lifecycle, stats), the v0.2 package layout, and the
wave plan. Authoritative reference for all Phase 2 agents.
Records the Phase 2 architecture: embed the engine in-process (apply = atomic
instance swap), rewrite only the generator (xray JSON -> sing-box options),
port parsers + nft/routing near-verbatim. Ship one binary shaterd.
AmneziaWG 2.0 proven E2E vs live Cloudflare WARP (warp=off->on through tunnel);
embedding via box.New verified on VM; router musl build ~9-11MB UPX. Next: Phase 2.
shater/cmd/shater-proto: drives the engine via the library API (box.New /
Start / Close) from our own Go main — the pattern the control-plane reuses.
Builds an option.Options in code (mixed inbound + direct outbound), proves the
data path with an in-process socks5 probe, and validates a tproxy+shadowsocks
variant through box.New. Verified musl-static on the x86_64 OpenWrt VM (egress
via the embedded engine), arm64 build proof, go vet clean. No go.mod changes.
Key API notes captured for Phase 2:
- ctx = include.Context(service.ContextWith(bg, deprecated.NewStderrManager(...)))
- option.Inbound/Outbound.Options MUST be a POINTER to the concrete struct.
- box.New constructs+validates every adapter; a single non-special outbound is
auto-selected as default route.
Orchestrator-laid foundation per CLAUDE.md: lightweight single-bundle SPA to be
embedded in the forked binary and served by the daemon on its own port.
- tokens.css: Faceplate tokens ported verbatim from docs-shater/DESIGN.md
(light/dark, prefers-color-scheme default + data-theme override both ways),
mono instrument voice, tabular-nums, focus-visible, reduced-motion floor.
- Placeholder App shell (component library <Faceplate>/<Module>/<Toggle>/<Led>/
<SegMeter>/<QueryLog> + pages land in Phase 3, delegated).
Builds clean: 145K dist (46KB gzip JS), tsc --noEmit passes.
Phase 1 VM findings: canonical LX_TAGS links glibc (naive/cronet/purego dlopen)
and won't run on musl OpenWrt; router build drops with_naive_outbound,with_purego
for a fully-static binary. Ship UPX-lzma (~9-11MB from ~40MB raw).
Upstream sing-box-lx already ships a docs/ mkdocs site; keep our project docs
separate and unambiguous in docs-shater/ (parallels upstream's docs-lx/).
Updated all references in README.md, CLAUDE.md, CONTEXT.md, ARCHITECTURE.md.
Foundation pivot. The complete, working, VM-verified xray-based project is
preserved on the `v0.1` branch; `main` is reset to a docs-first scaffold for
v0.2, which will be built as a FORK of sing-box-lx with our control-plane,
DNS filter, stats and admin panel embedded in the one binary.
- Preserve everything on branch v0.1 (pushed).
- Remove the v0.1 implementation + old design docs from main (recoverable from
v0.1); keep LICENSE, .gitignore, .gitattributes, dist/shater-feed.pub (feed
signing key 5ac4b177689cb8e0 carries over).
- License -> GPL-3.0 (sing-box is GPL-3.0).
- Add full project context so it survives compaction:
docs/CONTEXT.md (start here), DECISIONS.md, ARCHITECTURE.md, ROADMAP.md,
FEATURES.md, and a new README.
Engine/UI decisions (see docs/DECISIONS.md): fork sing-box-lx (AmneziaWG 2.0 +
broad protocols, GPL-3.0, library-first) and embed the whole product for tight
integration; keep the fork maintainable via an additive overlay (shater/, panel/,
openwrt/) rebased on upstream tags. UI = thin LuCI launcher + a separate admin
panel served by the daemon, entered via a short-lived token minted in the
authenticated LuCI session. Do NOT write a proxy engine from scratch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Promote v1.14.0-lx.3-rc.2 to a stable (Latest) release. Functionally
identical — no runtime change since the rc, only this changelog entry.
Payload: DNS command-multiplex (rc.1) + AWG re-graft onto wireguard-go
v0.0.5 with upstream merge (rc.2). Device-verified by owner.
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.
lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
(PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
point every L3-forwarded packet transits, incl. established flows that
bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
sing-tun); wireguard-go stays v0.0.5 with local submodule replace.
Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
SPEC.md rewritten to current (multiplex) architecture, no chronology.
HISTORY.md captures v1 standalone class-error, the field bug, rejected paths.
New project rule (README + CONSTITUTION 3.2): SPEC.md = current state first,
chronology/rationale of architecture changes go to HISTORY.md.
DNS stream now runs on the shared c.ctx via dispatchCommands (CommandDNS),
mirroring handleConnectionsStream: auto-reconnects with Connect(), dies with
the client, no per-stream Close()/OnError. Removes DnsQueryHandler,
DnsQuerySubscription and the standalone SubscribeDNSQueries client method.
DnsQuery/DnsAnswer/dnsQueryFromGRPC unchanged.
Add DNS as a first-class multiplexed command, uniform with CommandConnections:
- command.go: CommandDNS constant (next in iota)
- command_client.go: case CommandDNS in dispatchCommands; DNSIncludeAnswers
option field (like StatusInterval); WriteDNSQuery in CommandClientHandler
All three upstream touch-points wrapped in // lx:begin dns / // lx:end dns
(CONSTITUTION 3.3). SPEC 018 v2.
Runtime detour/selector rings crash the core (fatal stack overflow via
unbounded DialContext recursion); static rings are already rejected at start
by lintOutbound. Worked out the full event model (E1-E5) and topologyMu race
linearization, then adversarially verified it (7-agent workflow): deadlock and
false-positive attacks HOLD, but TOCTOU BREAKS — even a correct core guard is
not airtight without also covering the endpoint manager, Manager.Remove,
history side-channels, and the pointer-vs-tag graph divergence after a runtime
Create. Owner decision (2026-07-06): protection lives at the UI level (LxBox
validates before SelectOutbound); core stays a minimal delta to upstream.
No core code changed. SPEC is a design record + Roadmap row (status DEFERRED).
Split the lx go vet step into two passes so every lx-owned package keeps
the full analyzer set; only daemon/ and experimental/libbox/ (upstream
TriggerDebugCrash/TriggerGoPanic) drop the unsafeptr check.
Producer run populated musl-toolchain-cache with 4 arch assets; restore path
validated locally (asset name, gh download, tar layout under naiveproxy/src).
Status -> C, Roadmap updated.
snapshot.debian.org intermittently 503s during the musl sysroot build and
blocks releases (v1.14.0-lx.2-rc.1 failed twice on it). actions/cache also
misses across tag builds (ref-scoping). Add a producer workflow that uploads
the built toolchain to a musl-toolchain-cache release, and a restore step in
lx-release.yml that pulls it on cache-miss before falling back to
snapshot.debian.org. Both workflows are lx-owned; zero upstream diff.
Full audit of the LX delta (10 axes, adversarial verification): 32 findings,
27 confirmed, 24 fixed on branch lx-spec022-audit-fixes, 3 skipped by design
(#12/#17/#18). Records #19 resolution (SPEC 013 test kept — upstream ships none).
Exact-string match on quic-go's formatted CRYPTO_ERROR silently disabled the
Cloudflare Access hint on any error-text reformat. Match the inner
"tls: access denied" alert as a substring instead.
s3 pads only cookie-reply messages (paddings.cookie); s4 pads every transport
data packet (paddings.transport). Folding s3 into the MTU budget dropped MTU and
warned spuriously for an atypical s3>s4 config. Fix calc, comment, warning, docs.
The 'if c.reader == nil' fast path read reader without synchronising against the
RoundTrip goroutine's write (a data race -race flagged). Always receive on
created first; on an already-closed channel that is effectively free.
calculateIPv4Checksum summed a fixed 20 bytes, producing a wrong checksum after
TTL decrement when the header carried options (IHL>5). Take the real IHL*4 span;
validate IHL against the buffer before the read. Adds an IHL=6 test.
Pool() keyed URL-test history by the raw slot tag; history is stored under
RealTag(detour), so a nested-group member always reported Delay=0 in GetPool.
Resolve the slot tag and read under RealTag, matching seedPool/rebuildPool.
questionCache returned a fresh cached response without logging/emitting; only
the stale (optimistic) branch emitted. Mirror the Exchange path: fresh->cached,
stale->optimistic. Gated by HasSubscribers, so zero cost with no profiler.
SuspendAmneziaWG left idleAsleep untouched, so an endpoint idle-suspended
BEFORE the guard fired could be resurrected by the next dial (resumeOnDial
keys only on idleAsleep) — reintroducing the AmneziaWG-over-WireGuard kernel
hang the guard exists to prevent. Now clears idleAsleep under resumeMu so
the guard is ordered against a concurrent wake.
sendConnect's ReadFrame loop blocked forever on a peer that completes
TCP+TLS but never returns the CONNECT HEADERS, wedging the outbound under
o.runMu (Close hangs too). A ctx watcher now trips tlsConn's deadline on
timeout/cancel and is joined before the long-lived readLoop starts.
The main README's 'Features & status' table and Feature-configuration section
listed XHTTP/AWG/observability/round_robin but not the SPEC 021 MASQUE
(CONNECT-IP / WARP) outbound. Add it in both README.md and README.ru.md:
intro line, a Features table row (device-verified on Wi-Fi + LTE, h3/h2), and
a short config example with the network=transport footgun and the h2 fallback
for UDP:443-filtered networks. Links to docs-lx/lx-config §4 and SPECS/021.
The user-facing config guide covered XHTTP/AWG/urltest but not the SPEC 021
MASQUE (CONNECT-IP / WARP) outbound — only the internal SPECS/021 had it.
Add a full section (§4, renumbering Observability→§5, Validate→§6): feature
table row, field table, h3/h2 example, the network=transport footgun, dns-block
requirement, h3-vs-h2 guidance (UDP:443 filtering, cold-start), device-verified
status, and a link to SPECS/021/CONFIG.md. RU mirror kept in sync.
The 1.14 merge (6b63cee4) brought daemon/managed_service.go and
experimental/libbox/debug.go into vet scope; both crash Go ON PURPOSE via
*(*int)(unsafe.Pointer(uintptr(0)))=0 (TriggerDebugCrash/TriggerGoPanic),
which vet's unsafeptr analyzer flags. They are upstream files we don't edit,
and no lx-owned file uses unsafe at all, so disable just that analyzer.
Fixes the red lint job on every push since the merge.
A failed dial is an actionable error and should be visible where the success
(tunnel established, INFO) is — the LxBox core-log forwarder only surfaces
INFO+, so a DEBUG failure was invisible on-device. WARN makes established/
failed a symmetric, forwardable pair. The detailed dial phases (establishing,
udp-socket-up) stay DEBUG.
Promotion of the rc.1..rc.22 series to a non-prerelease tag (publishes as
Latest). Functionally rc.22 + the linux-mips-softfloat asset (#6). Section
header matches the tag exactly so lx-release.yml extracts it into the notes.
Rides out transient snapshot.debian.org 503s in the linux-musl toolchain
download/keyring steps (rc.22 failure cause). No code change; publish still
requires all builds.
rc.22 linux-musl jobs failed on 'get-clang.sh: 503 No healthy backends' from
snapshot.debian.org (transient mirror outage). get-clang.sh retries internally
but back-to-back, so a whole outage window fails all attempts. Wrap the two
Debian-fetching steps with an external retry + growing backoff:
- Download Chromium musl toolchain: 5 attempts, 30→60→120→240s
- Regenerate Debian keyring: 4 attempts, 20→40→80s
Conservative: publish still needs all builds (a real musl breakage still blocks
the release, only transient mirror flakes are ridden out). No code change.
Requested in #6 (Atheros AR9344). Chromium/cronet has no big-endian MIPS
toolchain, so the musl+naive path is impossible for this target; add it to
the plain cross-compile matrix instead as a pure-Go CGO_ENABLED=0 build —
statically linked (runs on musl/OpenWrt as-is), with with_naive_outbound
and with_purego dropped (purego has no mips port either). Everything else
matches the desktop tag set.
Verified locally: GOOS=linux GOARCH=mips GOMIPS=softfloat build with the
reduced tag set compiles clean; `file` reports ELF 32-bit MSB MIPS32,
statically linked.
Adds establish/handshake/success debug logs to the MASQUE dial path so a
stuck tunnel is diagnosable from /logs/core alone (motivated by the LxBox
§130 device case: h3 hung in QUIC handshake because inbound UDP:443 was
Log the tunnel-establish phases so a stuck dial is diagnosable from
/logs/core alone, without a goroutine dump:
- "establishing <h3|h2> tunnel to <server> (sni=...)" on start
- "udp socket up, starting QUIC handshake" (h3) — pinpoints whether a
hang is the socket or the handshake (inbound UDP:443 filtered → our
ClientHello left but no ServerHello came back)
- "tunnel established" on success, "tunnel failed: <err>" on failure
Motivated by a live device case (LxBox §130): h3 hung in the QUIC
handshake because inbound UDP:443 was filtered by the network while AWG
(UDP on a non-443 port) worked — invisible in logs before this, required
a pprof dump to locate. h2 (TCP:443) is the fix there.
Refs: SPEC 021
Complete masque outbound config reference verified against code:
full JSONC template (all params), minimal config, per-field table
(masque-specific + inherited DialerOptions), profile matrix, value
formats (duration/keys/ip), start-time validation, and common footguns
(network=transport not tcp/udp, dns block required, exit-IP changes on
reconnect, keepalive vs idle_timeout).
Rework the tunnel lifecycle around a *session (device + ipConn + closer +
ctx + activity counter), guarded by runMu with a generation guard.
- C1 (was HIGH): a dropped or suspended tunnel is now rebuilt on the next
dial. Previously 'running' latched true, so after the tunnel died every
DialContext short-circuited and dialed into a dead stack — permanent
blackhole. teardownSession clears o.sess so ensureSession rebuilds.
- C2: teardownSession closes ipConn, which unblocks the paired pump parked in
a blocking read (context cancellation alone can't interrupt it); no leaked
goroutine, no zombie half-open tunnel. Idempotent via sync.Once.
- B1: idleWatcher suspends the whole tunnel (gVisor netstack, pumps, QUIC
keepalive) after idle_timeout of no traffic; the next dial rebuilds it —
near-zero resident cost when idle.
- B4: idle_timeout (default 5m) and keep_alive_period (default 30s) are config
options; negative disables. A5: fail-fast mtu<=16000 on h2.
- D2: drop the dead congestion_control option field.
lifecycle_test.go covers the generation guard, idempotent teardown and
close-guard under -race.
Refs: SPEC 021 audit B1/B4/C1/C2/A5/D2
- A4/A5/B5: cap peer-declared capsule payloadLen at maxCapsulePayload (64KiB)
to prevent int-overflow/OOM on a hostile length; shrink recvCh 64->8 and the
receive window 1GiB->8MiB so real HTTP/2 flow-control provides backpressure
instead of unbounded RAM.
- B2: reuse a per-conn scratch for the outgoing capsule frame (tx pump is the
sole writer; writeData flushes before returning) instead of allocating per
packet.
Refs: SPEC 021 audit A4/A5/B2/B5
- A1: a single malformed/empty inbound datagram (unparseable context-ID
varint) no longer tears the tunnel down — drop-and-continue like the
sibling context-ID!=0 / bad-payload cases.
- A3: snapshot the IP header before composeDatagram mutates TTL/checksum, so
the ICMP 'packet too big' reply quotes the original datagram (RFC 1191/792).
- A2: drop the redundant second status check in ConnectTunnelH3 (2xx already
validated in dialCONNECTIP); removes the unused responseInfo carrier.
- B3: reuse a per-Conn scratch for the outgoing datagram (contextID + packet)
instead of allocating per packet — safe, quic-go copies the slice before
SendDatagram returns and the tx pump is the sole writer.
Refs: SPEC 021 audit A1/A2/A3/B3
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).
The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.
Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.
Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.
Refs: SPEC 021 TEST_PLAN.md
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.
Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
semantics differ from WARP's expected extended-CONNECT authority/headers
(SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.
Refs: SPEC 021 TEST_PLAN.md
The "GRO off + batch 8" idea (a global alternative to Down/Up) was measured on-device
and REJECTED, for three independent reasons (SPEC.md §14):
1. Wrong holder — the main android RAM holder is device.pool.messageBuffers
(PreallocatedBuffersPerPool=4096 × ~64KB ≈ 100MB), which does NOT depend on
BatchSize; the batch-sized bufsArrs held only ~14MB. Shrinking batch wouldn't
have touched the ~100MB.
2. Not deliverable — the LX_WG_NO_GRO env switch never reaches Go's os.Getenv on
Android (wrap.<pkg> prop shows in /proc/environ but not in the runtime's env
snapshot), forcing a hardcode.
3. Fragile — hardcoded batch=8 crashed at start (SIGABRT): device.BatchSize()=
max(bind,tun) clamped back to 128 via the TUN offload while msgsPool was 8, so
Send sliced out of range. Coherent only by also gating TUN offload across three
submodule layers.
Down/Up (rc.19) stays the only viable mechanism. Brings the experiment folder
(protocol + device heap snapshots + RESULT) into lx-1.14 for the record; the
experiment CODE stays on the lx-1.14-nogro-* branches, not merged.
Sync SPEC 002 with the code (commit c0bbb1c5): GET on a non-packet-up node no
longer hard-errors — it falls back to POST + WARN so one bad subscription node
doesn't fail the whole config. Updated the mode-gate wording in SPEC.md §verif,
PARAM_MAP.md (full rationale + the old error text it replaces), URL_PARSING.md
table, and IMPLEMENTATION_REPORT.md. header/cookie uplink outside packet-up
stays a hard error (no safe default).
A subscription node sometimes ships uplink_http_method=GET on a non-packet-up
node (auto/stream-up/stream-one). GET can only carry the uplink in packet-up
(other modes put the body in the request, which GET has none of), so the strict
check rejected it — failing the ENTIRE config over one bad outbound in a large
subscription (observed: initialize outbound[361] ... can be GET only in
packet-up mode → whole tunnel won't start).
Fall back to POST (the safe default that works in every mode) and log a WARN
instead of erroring, so the rest of the config still loads. POST is what the
node should have used; the fallback just makes one malformed remote node
self-healing rather than fatal. Kept strict for packet-up (GET honoured there).
Uses log.StdLogger() for the warning (the client transport layer gets no
logger in its constructor signature; threading one through would touch upstream
signatures). Test: TestUplinkGetFallsBackToPostOutsidePacketUp; removed the GET
case from TestValidationRejections. Verified on a real binary — check on a
stream-one+GET config now warns and passes (exit 0) instead of FATAL.
The base-version step (c4fd73cd) adds an `upstream` remote (SagerNet/sing-box)
so git-describe can see the v1.14.0-alpha.* tags. But `gh release create` without
--repo resolves the target repo from the remotes and picked `upstream` →
HTTP 403 "Resource not accessible by integration" against
api.github.com/repos/SagerNet/sing-box/releases (the token has no rights there).
This is why rc.19's builds all succeeded but publish failed, while rc.18 (before
the upstream remote existed) published fine.
Pin --repo "${{ github.repository }}" so publish always targets this fork
regardless of what remotes the earlier steps added.
rc.19 gates idle-suspend behind with_lx_idle_suspend (mobile-only) and records the
on-device Android verification: suspending 8 idle+unreachable WG endpoints freed
134MB of bufsArrs live heap (223.9→89.9MB, recv-workers 18→2), matching the
~8.4MB/worker model — ~10x the desktop delta, on the platform the feature targets.
Adds ANDROID_RESEARCH/live-baseline/ — a full pprof snapshot of a real production
config with the feature OFF (263MB bufsArrs, 56% CPU on GC at idle) and its ON
"after" counterpart, closing the RESEARCH.md device gap end-to-end (buffer pool
and GC cost measured together, not inferred). Credentials scrubbed.
Idle-suspend frees the recv-worker bufsArrs, which are ~8MB each only where
BatchSize=128 (Android/Linux) — on desktop BatchSize is small and the feature
saves almost nothing. Make that platform scope explicit in the build instead of
running the tick everywhere.
The idle-suspend tick now compiles only with the new `with_lx_idle_suspend` tag,
baked into the mobile AAR (build_libbox sharedTags) but NOT the desktop LX_TAGS.
Without the tag, a config that sets route.lx_idle_suspend fails fast at start
("rebuild with -tags with_lx_idle_suspend (mobile-only feature)") rather than a
silent no-op. The gate is a single function: reachability_lx.go carries the tick
under the tag, idle_suspend_stub_lx.go is the no-tag stub that errors, and
reachability_common_lx.go keeps InvalidateReachability (needed by the group
interface in every build). The dial hot path (resumeOnDial/stampActivity) and the
upstream group files are untouched — without the tick, idleAsleep is never set, so
resumeOnDial always takes its fast path.
Adds stub unit tests (option set → error, unset → no-op). Both build variants and
the full route/wireguard/group suites are green; gofmt/vet clean; desktop lx-check
passes without the tag. Docs (lx-config.md + ru, SPEC.md §3/§10) describe the tag.
The base-version derivation used `git describe --match v1.14.0-alpha.*` as the
primary source, but actions/checkout only fetches THIS repo's tags — the alpha
tags are SagerNet/sing-box (upstream) tags, absent in the CI clone. So `git
describe` found nothing and silently fell to the subject-grep fallback, which
resolves alpha.36 (alpha.37 was merged in a commit whose subject omits the
number). That's why rc.17/rc.18 notes shipped "base alpha.36" while a local
clone with upstream tags gets 37.
Fetch just the upstream v1.14.0-alpha.* tags before git describe so the primary
graph-based path works in CI. Both hand-fixed on the published releases; this
makes the next tag correct automatically.
Full write-up of the Android device run (CPH2411, Android 15, rc.18) in a
dedicated ANDROID_RESEARCH/ subfolder: README (report), METHOD (reproducible
procedure), RESULTS (per-scenario + heap A/B), and artifacts/ (raw evidence:
lx idle log lines, goroutine dumps, pprof heap .pb + top renders). No access
credentials anywhere.
Headline, now measured on the target platform: PopulatePools.func3 (the
bufsArrs holder from RESEARCH.md) inuse_space 223.93 -> 89.89 MB (-134 MB /
-60%), recv-workers 18 -> 2, on suspending 8 of 9 WG endpoints. = 16 workers x
~8.4 MB (BatchSize=128), matching the source model, ~10x the desktop RSS delta.
RESEARCH.md status + SPEC.md §12/§13 updated: the Android heap A/B gap is
closed (only the battery A/B remains deferred).
Device run on CPH2411 (Android 15, rc.18) via the LxBox app Debug API.
9 WG endpoints (1 real WARP reachable + 8 synthetic unreachable),
lx_idle_suspend=30s. All behaviors confirmed on-device: suspend fires,
reachable final stays up, wake-by-dial, no-flap, kill-switch.
Headline: PopulatePools.func3 (the bufsArrs holder from RESEARCH.md)
inuse_space 223.93 to 89.89 MB (-134 MB / -60%), recv-workers 18 to 2.
= 16 freed workers x ~8.4 MB (BatchSize=128), matching the model, ~10x
the desktop RSS delta. Closes the Android device-verification gap.
Device-verified idle-suspend: idle AND unreachable WG/AWG endpoints go Down to
free their recv-worker bufsArrs (the Android GC-heat holder), waking on next dial.
Opt-in via route.lx_idle_suspend; off by default. Notes cover the reachability
walk, the GRO reason for Down-over-smaller-batch, the concurrency fixes, and the
2026-07-01 live-run results (recv-workers 16→0, RSS −31%).
Selectively brings idle + unreachable WireGuard/AmneziaWG endpoints Down,
freeing their recv-worker bufsArrs (the measured GC-scan heat holder on
Android) and stopping their per-peer timers (battery), then wakes them lazily
on the next dial. Off by default (lx_idle_suspend absent/0 = zero overhead).
Includes the fix for the shipped tick iterating the wrong manager (it never
reached any endpoint — the feature was inert on a live box), the full
reachability walk (final/rule/selector Now/urltest pool/detour, event-driven
cached), and the rewritten as-built spec (SPEC.md) + research doc (RESEARCH.md)
+ test plan.
Live-verified: suspend/wake/probe-wake/re-sleep/no-flap/kill-switch across
selector, urltest pool, nested groups, AWG-guard, and the real production
config; resource A/B recv-workers 16->0, RSS -31% on desktop. 29 unit tests,
adversarially checked.
The idle-suspend feature is implemented, the tick bug is fixed, and every
reachability node type plus suspend/wake/probe/no-flap/kill-switch and the
resource A/B (recv-workers 16→0, RSS -31%) are live-verified. Reflect that in
the docs and give the folder clean roles:
- SPEC.md (was SPEC_idle_suspend_lever.md): rewritten from scratch in Russian
as the as-built implementation spec — Down/Up model, reachability walk +
event-driven cache, endpoint-side suspend/wake, the tick bug and its fix
(§11), full test coverage (§12, 29 units named), and what is deliberately
deferred (§13: Tier B netstack teardown, keys-safe BindUpdate path, on-device
battery/heap measurement). Old Tier-A "light sleep" design (never shipped)
removed.
- RESEARCH.md (was SPEC.md): the diagnostic root-cause doc keeps its unique
on-device heap A/B proof (holder = recv-worker bufsArrs) — renamed so its
role (research, not implementation spec) is unambiguous.
- TEST_PLAN_idle_suspend.md: all pass criteria checked, §RESULTS + edge-case
matrix + wake-latency series (cold ~50ms / warm ~36ms, +14-21ms ≈ 1 handshake;
far-server caveat) + production-config run.
Cross-references and section numbers updated across all three files.
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Code recon of the wireguard-go submodule disproved SPEC.md's original PRIMARY
lever (shrink StdNetBind.BatchSize() 128→8). It cannot be done without breaking
GRO receive:
- GRO-rx is ENABLED on android (UDP_GRO set with no android gate,
controlfns_linux.go:90-104; SPEC 010 gated only GSO-tx, not GRO-rx) → rxOffload=true.
- The GRO path splits one coalesced packet into up to 64 datagrams; readAt =
len(msgs) - IdealBatchSize/udpSegmentMaxDatagrams (bind_std.go:269) HARDCODES
IdealBatchSize=128, and getMessages() allocs a 128-slot array. Shrinking bufsArrs
to 8 either desyncs bufs(8) vs array(128) → OOB panic, or overflows the split
("splitting coalesced packet resulted in overflow", bind_std.go:565). GRO can't
be disabled (needed for download throughput, §010).
So the old claim "packet loss excluded, array just shorter" was wrong for the GRO
path. Lever 1 (and lever 2, which inherits the same idle-socket GRO problem) are
rejected. PRIMARY becomes lever 3 — Down idle+unreachable devices: BindClose ends
the recv-workers and frees bufsArrs whole, while the active node keeps batch=128 so
its GRO is intact. The "most expensive fallback" is in fact the only viable lever.
Updated: status line, the lever-candidates section (struck lever 1, promoted lever 3),
the fix-logic section (renamed + rewritten around Down), verification, and residual
risks (the fast-channel risk is gone; the new risk is handshake-on-wake). Implemented
on lx-spec020-idle-suspend; see SPEC_idle_suspend_lever.md §13 + TEST_PLAN.
Add TEST_PLAN_idle_suspend.md — build/config/commands/pass-criteria to
device-verify the shipped idle-suspend on a real run: suspend fires for
idle+unreachable WG/AWG endpoints, reachable ones never suspend, wake-on-dial,
the bufsArrs memory drop (pprof heap), and no flapping. Uses the user's
WARP/AWG + plain-WG nodes. Link it from SPEC_idle_suspend_lever.md §13.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Record the as-built decision so it is not re-derived:
- §13.1 what shipped (c55cf11e): idle-suspend via Down/Up, not light sleep
(source-verified holder is recv-worker bufsArrs; timersStop does not free it).
- §13.2 bind-swap investigation (the promised comment): BindUpdate resizes the
bind WITHOUT zeroing keys; key-zeroing lives only in Down/peer.Stop. So a
keys-safe wake is possible only while still Up — the shipped Down path can't.
- §13.3 the three reduced-bind paths (B=Down shipped / A=BindUpdate keys-safe /
Hybrid) with the GRO-off + max(bind,tun) gotchas.
- §13.4 recommendation: ship B, escalate to A/Hybrid only if device INFO logs
show handshake flapping hurts. Reduced-bind urltest wake deferred (low value
on path B since keys are already zeroed).
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.
- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.
Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
Complete RU translation of docs-lx/lx-config.md, mirroring its structure:
the §0 exhaustive "every field at a glance" example, all per-section field
tables (XHTTP v1+v2, AmneziaWG 2.0, id/ip/ib masquerade, urltest balancer),
examples and the build section. JSONC code is preserved; only prose and
inline comments are translated. Intra-doc #anchors are re-pointed to the
Russian heading slugs (all 6 verified to resolve); the §0 example validates
as JSON.
Cross-link both ways (en ↔ ru) and re-point README.ru.md's four lx-config
links to the Russian version. File name follows the README.ru.md convention
(.ru.md, not -ru.md).
Add a §0 kitchen-sink config carrying ALL 52 lx-added fields in one place —
XHTTP transport (26), AmneziaWG 2.0 endpoint incl. id/ip/ib (21), urltest
round_robin balancer (5) — each with its default and allowed values inline,
and mutually-exclusive / server-ignored fields flagged. Sourced by reading
option/*.go directly (not the prior doc), so it is complete.
This also surfaced that §1 documented only 7 of the 26 XHTTP fields (the v1
set); fill in the 19 missing v2 fields (session/seq placement, uplink-data
placement, X-Padding obfs family, packet-up tuning, accepted-but-ignored)
as grouped tables. Fix three code-vs-doc disagreements the extraction found:
- `mode: auto` resolves to stream-one on Reality (not always packet-up);
- s3/s4 are AWG 2.0 junk-size params (not "cookie-reply/transport" junk);
- h1-h4 unset spelling includes "" as well as 0; x_padding_bytes framing.
The default wire shape is unchanged; all v2 fields are opt-in. §0 example
validated as JSON. Russian translation (lx-config.ru.md) to follow.
docs/ is an upstream-owned tree (it arrives wholesale from SagerNet on
every rebase). Our three downstream docs lived inside it — lx-config.md,
lx-changelog.md, lx-release-runbook.md — mixing fork files into the
upstream surface against CONSTITUTION principle #1 (thin layer / minimal
diff). Move them to a dedicated root-level docs-lx/ so the boundary
between our docs and upstream's is explicit.
- git mv preserves history.
- Updated every reference (docs/lx-* -> docs-lx/lx-*): README.md/.ru.md,
SPECS/{003,004,005,009,020,README}, transport/wireguard/endpoint.go
comments, lx-ci.yml, and lx-release.yml (the release-notes extractor +
fallback URL now read docs-lx/lx-changelog.md).
- Fixed the now-relative links inside the moved files that pointed at
upstream docs/ siblings: lx-config.md -> ../docs/configuration/outbound/
urltest.md; lx-changelog.md -> ../docs/changelog.md (x2).
Verified: all relative + external links resolve, both workflows are valid
YAML, the release-notes awk path is docs-lx/, go vet clean on the touched
package. No release feature — folds into the next tag naturally.
Secondary design doc alongside the authoritative SPEC.md, NOT a replacement.
SPEC.md has on-device proof (heap A/B) that the GC-scan holder is bufsArrs of
recv-workers and its primary lever is shrinking StdNetBind.BatchSize() — that
stays authoritative.
This companion works out SPEC.md's 'lever 3' (suspend inactive devices) in
detail: a light variant (per-peer timersStop, keypairs/socket kept live, cheap
wake without handshake) plus a reachability walk (final + rules + active
selector/pool choices, generation-cached) that decides which devices are idle
AND unreachable from the active routing tree.
Banner up top flags where this doc's source-reading diverged from SPEC.md's
measurements (it guessed gvisor netstack; SPEC.md measured bufsArrs) and notes
light-suspend does NOT free bufsArrs — only Down/BatchSize does. Use as the
fallback design for lever 3, not a competing primary.
Spec only — no code changes.
The subject-grep base-detection missed alpha.37: it was merged in a commit
titled "Merge upstream/testing (bump version, fix linux ping)" with no
"alpha.37" in the subject (upstream tagged it after we merged), so the grep
found only alpha.36 and rc.17 notes shipped a stale base.
Make `git describe --match v1.14.0-alpha.*` the primary source — it reads
HEAD's ancestry in the commit graph, independent of merge-message wording —
and keep the subject-grep as the fallback for a fork checkout without
upstream tags. Verified locally: now resolves v1.14.0-alpha.37.
SPEC 014 dropped with_clash_api because LxBox (Android) drives the core
over the native libbox CommandClient, making the Clash REST server dead
weight in the AAR. But the drop landed in the shared Makefile.lx LX_TAGS,
which also feeds every desktop/CLI release build (mac/windows/linux-musl
via `make -s lx-print-tags`). A CLI binary has no CommandClient channel —
it is managed by external dashboards (yacd/MetaCubeXD) over the Clash REST
API — so every desktop release since rc.1 shipped with no way to manage
the core; a config with experimental.clash_api failed fast. CI stayed
green (lx-ci BASE_TAGS kept the tag), so it was invisible in CI.
Restore with_clash_api to the desktop LX_TAGS; leave build_libbox (AAR)
unchanged. The two tag sets now diverge by design: desktop = with Clash
API, AAR = without.
Verified: desktop binary builds with with_clash_api in Tags; `check`
accepts an experimental.clash_api config; the Clash REST server comes up
live (endpoints answer 401 security-middleware, not the stub's fail-fast).
Docs: Makefile.lx comment, SPEC 014 (§2/§3.1 scoped to AAR + new §3.4),
lx-release.yml tag comment + notes line, changelog rc.17.
The release-notes template hardcoded "base v1.14.0-alpha.35"; it went stale and
had to be hand-edited on rc.14, rc.15 and rc.16 (each was actually on alpha.36).
Resolve the base dynamically in the "Resolve tag" step: take the highest alpha.NN
named in any "Merge upstream" commit subject (robust on a fork without upstream
tags fetched), falling back to git describe against upstream alpha tags, then a
generic v1.14.x label. The notes line now interpolates steps.ver.outputs.base.
§8: validated transport JSON with all 14 new fields at non-default values,
the equivalent flat-camelCase vless:// URL, and a defaults table for the
toUri() omitempty logic. Fixture verified with sing-box check; mirrors
lx-test/config/xhttp_obfs_full.json.
Relocate docs/lx-xhttp-url-parsing.md -> SPECS/002-XHTTP_CLIENT_TRANSPORT/URL_PARSING.md
so all XHTTP docs live together with the spec. Fix internal/back links.
Scanned igareck/vpn-configs-for-russia, extracted+deduped 10 unique XHTTP
nodes, ran each through our with_xhttp binary. 4 alive — all downloaded 1MB,
traffic egressed via the server IP:
- 2x plain -> packet-up
- 2x reality -> stream-one (hu99.bearbeer.digital, bez3.stream-room.com)
The two reality nodes resolve auto->stream-one and work live, closing the
open stream-one live-verification TODO from task 011 (previously synthetic-only).
Other 6 nodes dead for server-side reasons (504, HTTP/1.1-not-H2, reset,
TLS hang) — our transport errored cleanly in every case.
Remaining live TODO: obfs/placement modes (no public node is configured for them).
Our XHTTP client is HTTP/2 only (http2.Transport); Xray supports H1/H2/H3.
h3-only nodes won't connect — flag for the link parser. Out of SPEC 002 scope
(separate future 'XHTTP over HTTP/3' task).
scMaxConcurrentPosts is a removed Xray knob (grep + GitHub code search
total:0 in current XTLS/Xray-core and sing-box-extended). Current Xray
serializes to one upload POST body in flight at a time, which our sequential
packet-up Write already matches, so the field is accepted for config/link
symmetry but ignored by the client.
- option: V2RayXHTTPOptions.ScMaxConcurrentPosts (json sc_max_concurrent_posts)
- PARAM_MAP: document as legacy/ignore tier with the real concurrency mechanism
(bounded pipe + WroteRequest serialization, server-side seq reorder)
- url-parsing doc: scMaxConcurrentPosts -> accept-but-ignore
- xhttp_obfs_full.json: include the field so check covers it
Verified: build/gofmt/vet clean, 16 unit tests pass, sing-box check passes
on all 3 xhttp configs incl the field.
Self-contained reference for the link parser: maps every vless://...type=xhttp
URL param (flat query + extra={...} JSON) to sing-box transport snake_case fields.
Covers TLS/Reality mapping, the extra-JSON number→"min-max" coercion, mode=auto
pass-through, path-with-query-tail, and ignored fields (scMaxConcurrentPosts,
server-only). Examples validated with sing-box check.
Implement all 12 client-relevant Xray/sing-box-extended XHTTP params on the
existing lean-native client (no Xray vendoring):
- session/seq placement (path|query|header|cookie) + keys
- uplink-data placement (body|auto|header|cookie, chunked base64) + key + chunk size
- uplink_http_method (upper-cased; GET only in packet-up)
- X-Padding obfs mode: placement (cookie|header|query|queryInHeader) + key/header +
method repeat-x | tokenish (HPACK-Huffman-tuned via golang.org/x/net/http2/hpack)
- packet-up tuning: sc_max_each_post_bytes (split), sc_min_posts_interval_ms (throttle)
4 server-only fields (server_max_header_bytes/no_sse_header/sc_max_buffered_posts/
sc_stream_up_server_secs) accepted but ignored by the client.
New files: transport/v2rayxhttp/{meta.go,xpadding.go}, xhttp_test.go.
Range fields use the "min-max" string form (no badoption.Range in sing).
Default (non-obfs) wire shape kept byte-identical to the live-verified v1
(x_padding='0' in Referer, session/seq on path, payload in body).
Verified: 16/16 unit tests, sing-box check on 3 configs incl full obfs,
go vet/gofmt/build (tagged+untagged) clean, negative (no with_xhttp) rejects.
Adversarial wire-protocol review against PARAM_MAP found no bugs.
Live test of a non-default mode against an Xray server remains an open TODO.
The rc.15 domain fix was confirmed on a real device: with the default
sticky_hash ["process","domain"] and no dest_ip workaround, browser traffic
spreads across the pool (on-device per-domain uniformity ~0.27 -> 0.95+).
Update README + lx-config.md status from "not yet device-verified" to
device-verified.
PRIMARY lever spelled out: bufsArrs size = bind.BatchSize() (128 on StdNetBind/android);
shrink it at one point (conn/bind_std.go:322, android branch -> 8/16). BindUpdate
(device.go:558) and getMessages() (bind_std.go:260) follow automatically, both recv
goroutines (v4+v6) covered, no packet loss (array stays full, just shorter), MaxSegmentSize
untouched (GRO intact). Effect 8MB->~0.5MB/recv = 176MB->~11MB at 11 devices. Only risk:
gigabit channel may lose throughput -> fall back to dynamic batch.
On-device throughput A/B (static arm64 curl, download via tunnel):
- baseline batch=128 = 10.7 MB/s median (~86 Mbps), stable 9.1-11.2.
- CPU under load: Syscall6 36%, scanobject 6%, crypto ~3%. The bottleneck is the
WARP channel + syscall overhead, NOT batch processing. GRO/batch only matters at
hundreds-of-Mbps/gigabit, so shrinking batch does NOT cost throughput on a typical
mobile/WARP channel (where the heat is reported).
Re-ranked the levers: PRIMARY is now the global smaller StdNetBind.BatchSize()
(128->8-16) — one point, no activity detection, cuts bufsArrs 8MB->~1MB/recv (176MB->
~11-22MB at 11 devices). Dynamic-batch and Down-idle drop to secondary. Caveat: verify
on a fast Wi-Fi/gigabit channel before release (batch may matter there). Also confirmed
heap scales linearly with live device count (11->269MB, 4->104MB, 1->0). No code changed.
The lx feature docs had drifted: README (en/ru) and docs/lx-config.md still said
"currently XHTTP + AWG2" and covered only SPEC 002/003/009 — the observability
layer (SPEC 014-018) and round_robin load balancing (SPEC 019) were undocumented
in the lx overview, and urltest.md still described the pre-rc.15 domain behaviour.
- docs/lx-config.md: new "## 3. round_robin load balancing" (mode/balancer,
pool/pool_tolerance/sticky_hash, ["none"] sentinel + badjson-[] caveat, slot-hash
binding, example, status) and "## 4. Observability (CommandClient extensions)"
(URLTestOutbound/GetRules/GetGroups/GetOutbounds/GetPool/SubscribeDNSQueries +
Connection.detourList, all behind with_lx_command); Validate&build -> ## 5.
- README.md / README.ru.md: broaden the stale "XHTTP + AWG2" framing; add feature
rows for observability and round_robin with honest status.
- docs/configuration/outbound/urltest.md: reconcile sticky_hash "domain" with the
rc.15 fix — domain reads metadata.Domain (survives domain->IP resolve), so it
works for normal sniffed domain traffic, not only literal-IP destinations; the
warning is reframed (domain works; dest_ip is an alternative).
Docs-only; no code change.
On-device A/B (Debug API /diag/pprof) settles it with high confidence:
- RoutineReceiveIncoming holds 180MB (61% cum); peek = 100% via sync.Pool.Get (in
worker hands, not the pool, not the channels).
- 11 live wireguard endpoints, 22 RoutineReceiveIncoming ALL on StdNetBind (batch=128),
0 on ClientBind. 22 x bufsArrs[128] x 64KB = 176MB ~= 180MB.
- batch=128 because WARP/AWG endpoints use a WireGuardListener dialer => StdNetBind
(endpoint.go:200-202), whose BatchSize()=128 on android.
- A/B: switching the active node WARP->home does NOT free memory (buffers do not sleep);
config 11 ep -> 1 ep gives 269MB -> 0 (= Iliya's workaround, reproduced via profile).
Both prior diagnoses were wrong: batch=1 (no, StdNetBind=128) and drain-on-Suspend of
channels/sync.Pool (misses; bufsArrs of live recv-workers holds it). MaxSegmentSize
2200->65535 is the volume trigger (x30 bytes), not the holder (downLocked/pools/channels/
batch identical 1.13<->1.14).
Fix = shrink batch for INACTIVE devices (naive lazy-bufsArrs impossible: StdNetBind
getMessages() is a fixed 128). Levers + a required download-throughput measurement
documented; pending lever choice.
On the client path BatchSize()=1 (client_bind/stackDevice/systemDevice all return 1),
so the old bufsArrs=128 => 8MB/device claim is wrong; bufsArrs is ~64KB. The 224MB
pprof attributes to PopulatePools is the sync.Pool.New alloc SITE, not the holder.
Real holders: (A) device.pool messageBuffers sync.Pool local+victim cache, (B) the 3
buffered device channels. scanobject 52% comes from the pointer-dense element/container
wrappers (4 of 5 WaitPools are scan-type), not the noscan [65535]byte arrays.
Fix rewritten to drain-on-Suspend: park RoutineReadFromTUN via a stackDevice.suspended
seam in Read (the lockless pool writer surviving Down), drain the 3 device channels
after Down, then ONE runtime.GC()+FreeOSMemory() per selector transition. PopulatePools
swap rejected (WaitPool.count underflow -> cond.Wait deadlock). slim-batch and the
route-graph refcount are dropped (not needed for heat). Added in-repo verification
(device/suspenddrain_test.go + HeapInuse bench) since on-device A/B is impossible.
Android 100% CPU / heat on configs with many WG/AWG endpoints in a
selector. Diagnosed via on-device pprof (§207): scan-bound GC over a
224MB live heap = wireguard-go Device.PopulatePools buffers, held by
~10 idle WG devices. Suspend()=Down() marks the device idle but does
NOT release pools/workers (only Close() does). A/B on device: dropping
spare WG endpoints removes the heat.
SPEC 020: SLIM idle devices (shrink maxBatchSize 128->1-4, do not Close
— keepalives + shared-node safety) gated by a route-reachability
refcount (rules + final + Now, not just selector).
SPEC 010: note our GRO split-brain patch is now upstream-native on
v0.0.3 (commit 24ea133); MaxSegmentSize=65535 must stay (GRO fuel) —
heat is fixed by device count/slimming, never by shrinking the buffer.
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.
1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
The router resolves a domain destination to an IP and overwrites
metadata.Destination before a group's DialContext runs, so destination.Fqdn
is empty when the balancer builds the key. stickyComponent("domain") read
that empty Fqdn, so a single process's key was process+NUL for every site
-> one fixed slot. On device this measured 28/1/1 across a 3-node pool
(uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
back to destination.Fqdn only for a direct dial. After: spread 0.95+.
2. living pool nodes could change slot index during a health-check, moving
sticky keys. balancePoolFirstLive compacted with a filtering append (a
transiently-dead slot shifted every later live node left); planTolerantPool
did delete(inPool, occupant) (an evicted-but-living node re-entered a later
slot, cascading); manual URLTest rebuild ran the tolerant planner even at
pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
dead/empty slots rewritten by index; dedicated planFirstLivePool for the
tolerance==0 rebuild).
3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
(badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
empty array to nil, indistinguishable from omitted, so the default always
applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].
Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
v2 superseded v1; keeping both as separate files (SPEC.md + SPEC_V2.md + the v1
TEST_REPORT) was just confusing. Delete the v1 SPEC and its TEST_REPORT (they remain in
git history) and rename SPEC_V2.md → SPEC.md as the one canonical doc. Drop the "v2"
suffix and stale "design not started" status from the header.
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.
Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:
- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
group -> empty. Additive proto/daemon/libbox, behind with_lx_command.
Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.
Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
Clarify the Now() cold-start tradeoff: variant B (write the fallback node straight
into selectedOutbound*) would eliminate the micro-gap entirely — Now() and DialContext
would read one field, so they can't diverge — at the cost of touching upstream's
selection logic (stub in the choice field + one extra Interrupt() on the first real
switch, which is no worse than any later latency switch). Variant A (Now() stays a
reader) was chosen purely for minimal upstream intrusion; its only cost is a negligible
micro-gap from two separate Select() calls racing on the first seconds. Documents the
path to B if the feature outgrows upstream's selectedOutbound* later.
Doc-only; rc.12 already shipped variant A, no retag.
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.
Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
The TEST_REPORT landed after the tag was cut, so the as-tagged notes still said
"not device-verified" and base alpha.35. Feature was live-verified on 5 vless nodes;
base is alpha.36 after the pre-rc merge. Published GitHub release notes edited to match.
Live run on 5 vless nodes (3 instances, one per mode): round_robin rotates strictly
across the live set and skips dead nodes; both sticky strategies pin deterministically;
bad config is rejected at start; -race clean on units and live. Feature is now
device-verified, not just isolated.
The run surfaced a config caveat (not a bug, by design): dest_ip is empty until the
destination is resolved, so a sticky key of only source_ip/dest_ip/dest_port collapses
to "" for domain traffic and pins everything to one node. Documented in urltest.md —
use `domain` in `hash` for domain-based traffic.
The lx-ci gofmt-lint step only checks files matching the lx-owned glob
(_xhttp|_awg|_lx.go|_command_lx); the new SPEC 019 files fell outside it. Rename
to the _lx.go convention so CI gofmt-checks them, and fix the changelog reference.
No code change.
Pre-rc.11 sync. Upstream changes: darwin local DNS refactored to a raw
mDNSResponder call, iOS deb upload fix, version bump. No overlap with lx files
(protocol/group, option, constant untouched).
Add a `mode` to the urltest group so it can distribute traffic instead of only
picking the lowest-delay node, with optional per-flow stickiness.
- mode: least_test (default, unchanged) | round_robin (rotate across live nodes)
| least_connection (reserved, phase 2 — rejected at config time).
- round_robin selects once per connection over the tag-sorted live set (nodes with
a fresh URL-test result supporting the network); UDP/QUIC sessions stay on one
node; first usable outbound is the fallback when nothing is live. The legacy
selectedOutbound* cache path is untouched — balancing is a separate branch in
DialContext/ListenPacket.
- sticky {mode, timeout, cap, hash}: binds one flow to one node. hash components
process|domain|source_ip|dest_ip|dest_port concatenate in order; absent -> "",
all-empty key -> one fixed node (keyless flows never rotate). mode jumphash
(default, stateless consistent hash — ~1/n remap on node-set change) or ttlmap
(key->node table, lazy + ticker eviction, 2000 LRU cap, 10m TTL, dead-node re-pin).
Reuses the existing urltest health ticker/history as the single liveness source;
no new probing. Now() reports the last-picked tag in balanced modes.
Tests (go test -race, 15 cases): distribution, dead-node skip, all-dead fallback,
jumphash stability + empty-key fixed node, ttlmap stick/expire/cap/dead-repick,
key building, validation. The race detector caught a real bug in the sticky
sweeper (read t.ticker unlocked while close() nilled it) — fixed by passing the
channels into the goroutine, mirroring URLTestGroup.loopCheck.
Also folds the SPEC 016 connections-map mutex (ebf9cc07) into the rc.11 changelog
section, which had not yet shipped in a release.
Connections is the client-side CommandConnections accumulator. With 2+
subscribers (LxBox screenClient + profilerClient) one goroutine writes
connectionMap in ApplyEvents while another ranges it in Iterator →
"concurrent map iteration and map write" fatal error → SIGABRT of the
whole process (reproduced in ~20s under traffic, CPH2411/Android15).
Add access sync.Mutex; lock every public method touching
connectionMap/input/filtered: ApplyEvents, FilterState, SortBy*, Iterator.
- FilterState split into public (locks) + private filterState (no lock);
ApplyEvents calls the private one under its already-held lock
(sync.Mutex is not reentrant). Field filterState -> filterStateValue to
free the name for the method.
- evictClosedConnections stays lock-free: private, only called from
ApplyEvents under lock.
- Iterator returns a COPY of filtered — the gomobile caller walks it
after the Go call returns (lock released), so it must not read the
live slice a concurrent ApplyEvents/SortBy is rewriting.
This is the UI/command channel, not the data plane — uncontended lock
~20ns. LxBox per-client accumulators (§170) stay as the consumer scheme;
the mutex is class-correctness insurance against a 3rd consumer.
Verified: go test -race TestConnectionsConcurrentAccess (writer || 3
readers, 2000 rounds) green; go build ./... and -tags with_lx_command
green; gofmt clean.
Codify the rule: before cutting any lx release/prerelease tag, check whether
upstream/testing moved ahead of our last merge and, by default, merge it in
first — then build/gofmt/lx-check, then changelog, then tag.
- docs/lx-release-runbook.md: pre-release gate checklist, drift-check commands,
the manual `git merge upstream/testing` flow (replaces SPECS/004 auto-rebase
while upstream is v1.14.*-alpha), conflict zones (.pb.go, wireguard-go submodule,
build_libbox marker, observability files), and the one-liner sequence.
- SPECS/004 SPEC.md: pointer to the runbook + note that manual merge superseded
auto-rebase on this branch.
Two nits surfaced by the lx-vs-upstream cleanliness audit (no runtime impact):
- box.go: the dnstrack registration comment said "service.FromContext" — the
§180 dead-stream signature. The actual readers use PtrFromContext (pairs with
MustRegisterPtr). Fixed the comment + noted why FromContext[*T] returns nil,
so a future debugger doesn't "fix" the readers back into §180.
- common/dnstrack/manager.go: removed the unused SourceRejected constant —
rejected resolutions are folded into SourceFailed at the emit site, so
"rejected" never reaches the wire. Replaced with a comment to prevent re-adding
an unreachable client case.
Audit verdict: code clean — no concurrency/wire/behaviour issues; dns/client.go
byte-identical to upstream, emits additive and subscriber-gated.
LxBox feedback: DnsQuery lacked which DNS server / outbound channel the query went
through. A DNS rule selects a server (matchDNS by action.Server), not an outbound;
the channel is the server's own detour, fixed at config time. Add to DnsQueryEvent:
- dnsServer/dnsServerType = transport.Tag()/Type() (transport is the Exchange param,
so available on all emit paths incl. failures);
- outbound = the server's detour tag (TransportAdapter.OutboundTag() from
DialerOptions.Detour), with a selector expanded to its live node via Now()
server-side (like Connection.Detour), empty on cached/optimistic.
Also gate event construction on HasSubscribers(): with no profiler attached the DNS
hot path builds nothing (no event/answers/outbound lookup) — previously every
resolution built an event just to be dropped for lack of a listener. The Now()
resolution therefore never touches the hot path.
Wire: additive proto fields + OutboundTag() on DNSTransport (embedded adapter
satisfies it). libbox DnsQuery.DNSServer/DNSServerType/Outbound(). Changelog rc.10.
DNS attribution was empty (0/119 on device): TUN+DNS hijack returns on a fast-path
(route.go:91/226) BEFORE matchRule, and searchProcessInfo — which fills
metadata.ProcessInfo — lives inside matchRule (:416). So fast-path DNS (most DNS on
a VPN) reached the SubscribeDNSQueries emit with nil ProcessInfo. Fix: call
r.searchProcessInfo(ctx, &metadata) before both fast-path hijacks (stream+packet);
idempotent + cached, one lookup per flow. Corrects SPEC 018 пункт 3 (the earlier
'cached attribution correct' claim checked ctx consistency, not that ProcessInfo
was populated before the resolve).
Also: DnsAnswer.rdata was the full RR string ('google.com. 29 IN A 1.2.3.4'); strip
the header prefix so clients get the bare value ('1.2.3.4' / CNAME target).
No proto/wire change. LxBox §180 needs no client change. Changelog rc.9.
SubscribeDNSQueries returned Unimplemented on device and emitted nothing: the
dnstrack.Manager is registered via MustRegisterPtr (key *dnstrack.Manager) but
read via service.FromContext[*dnstrack.Manager] (key **dnstrack.Manager), so the
lookup always found nil. Server -> Unimplemented; emit sites -> silent drop.
Fix all three readers to service.PtrFromContext[dnstrack.Manager] (the pair of
MustRegisterPtr, as trafficManager does in daemon/instance.go). Verified the
manager resolves to the exact pointer box.go registered. No proto/wire change;
rc.7 contract intact. LxBox §180 needs no client change.
Changelog rc.8.
The release-notes heredoc hardcoded 'base v1.13.13' and only AWG+XHTTP — stale
since the 1.14 migration, identical for every rc, and never reflecting what a tag
actually shipped (SPEC 014/015/017/018 were invisible). Now the 'What's new'
section is extracted from docs/lx-changelog.md for the current version (awk between
'#### vX' and the next '#### '), spliced via 'sed r' so changelog backticks/$()
stay inert (no command injection from doc prose). Base line fixed to alpha.35;
standing-features list updated with the CommandClient extensions.
Two doc-hygiene fixes from LxBox review (no logic change):
- Point the реализатор at 'Согласованная форма' as the binding contract; the
earlier 'Решение' proto sketch (Empty input, no failed/answers) is illustrative.
- Note that rcode=-1 ships as signed int32 (distinct on the wire from 65535,
verified); client must map -1 -> 'no answer' before any .toUInt().
Hijacked DNS (the norm on an Android VPN) is answered before a connection becomes
a traffic tracker, so DNS queries never reach the connections stream — the only
egress was the text log, which carries no app attribution. Add common/dnstrack
(a Subscriber[QueryEvent] mirror of trafficcontrol) emitting one event per
resolution from dns/client.go, attributed via adapter.ContextFrom(ctx).ProcessInfo
(same ctx on cache-hit and miss, so cached queries are attributed too).
Failures are first-class: timeout/loopback/rejected-cached/SERVFAIL-reject emit
failed=true + error + rcode=-1 (no response) — without this the stream is blind to
DNS failures, the primary throttling signal. CNAME chains preserved: with
includeAnswers, each event carries the full response.Answer in wire order (CNAME
hops + final A/AAAA, not filtered to IPs).
Wire: rpc SubscribeDNSQueries(SubscribeDNSQueriesRequest) returns (stream
DnsQueryEvent) + DnsAnswer; event-driven server stream (no ticker); libbox
SubscribeDNSQueries(includeAnswers, handler). Tag-less core -> Unimplemented.
Detour/Chain and other streams unchanged.
Docs: SPECS/018, lx-changelog rc.7.
chain omits the final outbound's own detour by design (upstream loop only
unwinds OutboundGroup via Now() and breaks on the first non-group), so a node
detouring through e.g. WARP never shows in the routing chain. Add Detour
[]string to TrackerMetadata, unwound from the final outbound's Dependencies()
(= its detour for a non-group outbound), descending into groups via Now()
against the same atomic snapshot, with a seen-guard against cycles.
Wire: additive 'repeated string detourList = 23' on the Connection proto
message (hand-applied to keep the generated diff minimal — no toolchain churn),
mapped in connectionToProto, surfaced on libbox Connection as Detour()
StringIterator. Chain / Clash-API unchanged.
Docs: SPECS/017, lx-changelog rc.6.
Parent the per-node delay test to the gRPC per-call ctx instead of the
long-lived boxService.ctx, so cancelling the call aborts the in-flight
dial before C.TCPTimeout without tearing down the connection. Restores
the granular per-node cancel the Clash API had implicitly via r.Context()
(there was never a cancelDelays endpoint).
Mass-cancel is unblocked client-side on the existing gomobile binding
via a separate ping CommandClient + Disconnect() (no native-surface
change, no server batch RPC) — closes the LxBox feedback.
Docs: SPEC 015 §3.6 (cancellation), SPEC 014 (#4240 deleted upstream →
seam-removal criterion switched to upstream-code), lx-changelog rc.5.
Ignore test/cache.db.
Connections (command_types.go:115) держит connectionMap/input/filtered без
синхронизации; ApplyEvents/evictClosedConnections/FilterState/SortBy*/Iterator
зовутся из разных gRPC-горутин (по одной на подписчика handleConnectionsStream).
≥2 подписчика CommandConnections → concurrent map iteration and map write →
fatal error → SIGABRT всего процесса.
Всплыло после CommandClient-миграции (раньше connections слушал ≤1 потребитель).
Обойдено клиент-стороной в LxBox §170 (per-client accumulator), но в ядре не
починено — третий потребитель вернёт краш. Фикс: sync.Mutex вокруг состояния.
Референс: LxBox docs/spec/tasks/170.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
Serialize probe rounds in startProber to eliminate unbounded fan-out of
fire-and-forget probe goroutines (up to 100/sec per direction), and close
HTTP/3 transports via transport.Close() in addition to CloseIdleConnections.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.
Three command-protocol additions completing the Clash-API -> CommandClient
migration (SPEC 015, behind with_lx_command):
- GetGroups / GetOutbounds: unary pull-snapshots over the existing readGroups()
and the SubscribeOutbounds builder. The CommandClient is push-only; if the
SubscribeGroups stream never opened (service not STARTED at subscribe) or broke,
the client had no cheap way to re-read group state and the main screen stayed
empty (tunnel connected, groups=[]). These getters close that gap without
recreating the whole client. Both needed: SubscribeGroups covers only in-group
nodes, endpoints (WG/AWG) + standalone outbounds appear only via the flat list.
Errors via status.Error (unary read convention, like GetRules).
- len<2 fix: readGroups() silently dropped groups with < 2 items (upstream commit
5bc0dfa9), hiding single-node selectors -- a regression vs Clash, whose /proxies
returned group.All() unfiltered. readGroups() is the single source feeding both
SubscribeGroups (startup broadcast) and GetGroups, so the fix covers both.
Handlers in started_service_command_lx{,_stub}.go behind with_lx_command;
client methods in command_client_command_lx.go reuse the existing gRPC->libbox
iterators. proto seam under // lx: marker, regenerated via pinned lx-proto.
E2E tests (test/command_lx_test.go) drive the public daemon API: single-node
group survives, flat list returned, not-started rejected. Both tag/no-tag builds
green; no-tag answers Unimplemented.
Split the original SPEC 014 by NATURE of change:
- 014 CLASH_API_TO_COMMANDCLIENT_MIGRATION (dir renamed) — the migration itself:
with_clash_api drop (rc.1) + box.go Android-start fix (rc.3). No RPC tech-spec.
- 015 COMMAND_PROTOCOL_RPC_EXTENSIONS — single home of all command-RPC work:
URLTestOutbound + GetRules (DONE, rc.2) + GetGroups + GetOutbounds + the len<2
readGroups bugfix (TODO, rc.4). All §3.6 class, with_lx_command.
Docs-only; shipped rc.2 code unchanged. 015 documents the pull-vs-push gap
(GetGroups/GetOutbounds) and the upstream len<2 group-drop defect, plus an
upstream-candidacy plan (§7): pull-getters + len<2 are clean upstream defects;
RPCs currently ship in lx-form, an upstream PR would need upstream-form (future).
Records the Android start fatal fixed in rc.3 (commit 029acd11): PlatformLogWriter
no longer forces the Clash server; observability served by the native
CommandClient. Self-contained in the feature SPEC — fix + WATCH
SagerNet/sing-box#4240 + the obligation to drop the // lx: box.go seam on the
next rebase if upstream resolves it.
Upstream box.go forced needClashAPI whenever PlatformLogWriter is set (always
on Android/libbox), because the Clash server was historically the only log/
traffic observer. With with_clash_api dropped (rc.1), that made every Android
start fatal: 'clash api is not included in this build' — even with no clash_api
in the config.
Split the concern behind a // lx: seam: PlatformLogWriter now requests
observability (Observable log factory + connection/traffic tracker), served by
the native CommandClient (SubscribeLog/SubscribeConnections), NOT the Clash
server. Only an explicit experimental.clash_api block still creates the Clash
server (and still fails fast without the tag). daemon is already nil-safe to a
missing clashServer, so Clash-mode degrades gracefully. Desktop unaffected.
Verified: core starts with no clash_api config; still fail-fast with one.
Clarify where a URLTestOutbound result surfaces, since the original §3.2 wording
only named OutboundGroupItem ("for nodes in groups") and omitted SubscribeOutbounds
— the actual channel that carries endpoint (WG/AWG/Tailscale) delay.
- §3.2: add a 3-row channel-map table (synchronous RPC response = any node;
SubscribeOutbounds = all outbounds AND all endpoints; SubscribeGroups = only
OutboundGroup members). All three share urlTestObserver via urlTestHistoryStorage.
- §3.2: sync the client signature to what shipped — (*URLTestOutboundResult, error)
with int32 timeout (gomobile can't bind the draft (uint16,string,error)); note why.
- §3.7 / §5: replace the narrow "history flows to OutboundGroupItem" line with the
SubscribeOutbounds + SubscribeGroups split.
Docs-only; no code or artifact change (rc.2 binaries unchanged).
Restore over the native libbox CommandClient what upstream only exposed through
the dropped Clash API: per-node delay testing and a route+DNS rule-table snapshot.
Both RPCs are a pure bridge (CONSTITUTION §3.6) gated by the with_lx_command tag.
- daemon/started_service.proto: URLTestOutbound + GetRules RPCs and messages under
the // lx:begin/end lx_command marker; regenerated .pb.go/_grpc.pb.go.
- daemon/started_service_command_lx.go (+ _stub.go): handlers behind with_lx_command,
stub twin returns codes.Unimplemented. URLTestOutbound resolves an outbound OR an
endpoint (no OutboundGroup assert), honours link+timeout, error-in-payload Variant B
(delay==0 && error=="" is success 0ms), history Store/Delete via group.RealTag.
GetRules returns route + DNS rules split by isDNS.
- adapter/dns.go + dns/router.go: new adapter.DNSRouter.Rules() getter (route Router
already had one), read under rulesAccess; both under // lx: markers.
- experimental/libbox/command_client_command_lx.go: CommandClient.URLTestOutbound
(*URLTestOutboundResult, error) and GetRules (RuleIterator, error) — gomobile-bindable
shapes (the SPEC's bare (uint16,string,error) does not bind); Variant B preserved.
- cmd/internal/build_libbox/main.go: with_lx_command into sharedTags (AAR).
- Makefile.lx: with_lx_command in LX_TAGS; pinned lx-proto/lx-proto-install targets
(protoc-gen-go v1.36.11, protoc-gen-go-grpc v1.5.1) for reproducible regeneration.
- lx-ci.yml: vet+gofmt cover the lx files; build-check proves both builds toggle the
stub marker.
- Collateral one-time pin normalisation of managed_service/v2rayapi/v2raygrpc .pb.go
(audited in SPEC §3.5).
docs(lx-changelog): v1.14.0-lx.1-rc.2. SPEC 014 → accepted.
CONSTITUTION: three owner principles (thin layer / follow upstream / build
what we+users need) wired into §1–2 as a priority hierarchy with a
necessary-not-sufficient lock; §3.1(а) rewritten from binary 'not in upstream'
to a three-prong test (needed by us/users / absent from OUR built channel /
cheaper than the alternative, with an auditable touched-file count); new §3.6
legalizes the 'libbox command-protocol extensions' change-class — handlers
gated by with_lx_command behind the proven started_service_usbip{,_stub}.go
pattern, .proto seam under a // lx: marker, .pb.go regenerated (never
hand-edited). §3.5 version bumped 1.13.13-lx.N → 1.14.0-lx.N.
SPEC 014: two CommandClient RPCs restoring what was lost when with_clash_api
was dropped — URLTestOutbound (per-node delay for outbound OR endpoint, custom
url + timeout, synchronous {delay,error}, all errors in payload) and GetRules
(route + DNS rule table snapshot). DNS-rules need a marked getter on
adapter.DNSRouter/dns.Router (route-only doesn't). Deterministic proto
regeneration (pinned protoc in Makefile.lx) is a mandatory deliverable. No
separate history RPC, no cancel-handle, no batch — those live in the client.
LxBox is moving to manage the core over the native libbox CommandClient
(group/url-test/select/connections streams), so the Clash REST API is dead
weight on the client. Drop with_clash_api from both the Android AAR
(build_libbox sharedTags) and the desktop LX_TAGS. A config referencing
experimental.clash_api now fails fast (no silent fallback); lx configs won't.
lx-release.yml: tags with an -rc.N / -alpha.N / -beta.N suffix now publish as
GitHub pre-releases (--prerelease), so an unverified build never displaces the
stable lx release as Latest.
First build on the upstream 1.14 base. The WG-endpoint GRO fix (010) lands at
the AmneziaWG v0.0.3 submodule source (no downstream guard), but the Android
download-stall path is NOT yet re-verified on hardware -- hence the -rc.1 tag.
Step 2 of 2 of the 1.14 migration. Points the submodule at e5feca7
(AmneziaWG 2.0 obfuscation re-grafted onto sagernet/wireguard-go v0.0.3).
Verification on lx-1.14:
- full sing-box build with lx tags (with_gvisor/quic/wireguard/utls/clash_api/xhttp/awg): OK
- submodule builds clean for linux/android/windows/darwin (library packages)
- transport/wireguard, protocol/wireguard, protocol/group, route/rule tests: green
- broad test (option/route/transport/common): green; gofmt + go vet clean
- binary runs on 1.14; package_name_regex config validates; awg2_basic + awg2_ranged validate
§010 android UDP_GRO guard dropped (v0.0.3 fixes split-brain at source) — pending
on-device re-verification before any release tag.
Full 1.14 migration, step 1 of 2 (sing-box repo layer). Three conflicts
resolved, all as predicted by the feasibility analysis:
- route/rule/rule_item_package_name_regex.go (add/add): took upstream's
canonical version (slices.ContainsFunc) — our lx.15 backport collapses
back into upstream, so the file no longer diverges going forward.
- route/rule_conds.go: kept our package_name_regex in isProcess{,DNS}Rule
and took upstream's new isNeighbor{,DNS}Rule additions.
- cmd/internal/build_libbox/main.go: kept lx with_xhttp/with_awg append and
the no-tailscale block; deliberately dropped upstream's new with_usbip
(server-side USB/IP, contradicts client-trim).
go.mod auto-merged: wireguard-go require bumped to v0.0.3, lx replace block
(=> ./submodules/wireguard-go) preserved. Submodule pointer unchanged here —
the AmneziaWG graft rebase onto v0.0.3 is step 2 (next commit). This commit
does NOT build yet (submodule still on the old wireguard-go base).
Backport upstream 1.14 feature 941ce58b onto the 1.13.13 base without the
full migration. Adds the package_name_regex rule item (regex match over
ProcessInfo.AndroidPackageNames) to route, DNS and headless rules.
- new route/rule/rule_item_package_name_regex.go (verbatim upstream) + unit test
- PackageNameRegex option field in RawDefaultRule/RawDefaultDNSRule/DefaultHeadlessRule
- item registration in NewDefault{,DNS,Headless}Rule with E.Cause(err, package_name_regex)
- package_name_regex added to isProcess{,DNS,Headless}Rule conds
The commit's RuleSetVersion5 hunk is intentionally NOT ported (that is 1.14
rule-set v5, unrelated; base is RuleSetVersion4). Full 1.14 migration deferred
to v1.14.0 stable. SPEC 013 + Roadmap entry.
builds (no-tags + lx-tags), go vet, gofmt and rule tests all green.
Folder names were NNN-T-S-NAME, so the status (S) letter forced a rename on
every status change — and refs to the full name went stale each time. One was
already broken in-tree (client.go pointed at 002-F-O-… while the folder was
002-F-C), and the submodule needed a cosmetic commit once already (010 O→C).
Make the number the only stable anchor:
- Rename all 12 folders NNN-T-S-NAME → NNN-NAME (git mv, history preserved).
- Type/status now live in a table header at the top of each SPEC.md (canon),
aggregated by the Roadmap in SPECS/README.md (added missing 010, 012).
- Fix every ref to the old full name: README(.ru), docs/lx-changelog,
docs/lx-config, intra-SPECS cross-links, PROBE.md git-apply path,
TASKS.md titles, and transport/v2rayxhttp/client.go:5 (also un-stales O).
- Rewrite the convention + Workflow in SPECS/README.md and the DoD ritual in
IMPLEMENTATION_PROMPT.md ("rename folder to …-C-…" → "set status in header
+ Roadmap").
- Bump submodule wireguard-go (0c0c10b): fix comments point at the new folder
name SPECS/010-WG_ENDPOINT_GRO_SPLIT_BRAIN. Comment-only, no behavior change.
Not touched: gro-probe.patch (historical diagnostic diff artifact; §010 closed,
no longer applied).
The symptom was seen on DIFFERENT nodes including WG, so "↑/↓0 download stall" is
an umbrella over the symptom, not one bug. Correcting the overconfident closure:
- §010 GRO fix lives entirely in the wireguard-go submodule (imported only by
transport/wireguard/) and gates UDP_GRO — it physically cannot affect VLESS/
reality (TCP) nodes. It closes the WG share of the symptom only.
- In the lx.12→lx.14 window, route/conn.go (shared relay) and the sing copy path
were unchanged; the only non-WG-relevant change is §011 (xhttp stream-one), which
applies only if the node uses xhttp. For VLESS+reality-direct there is NO code
change that explains the disappearance.
- "also hangs on VLESS" was never strictly confirmed (wlan0 encrypted), so the
non-WG share has no confirmed root cause — it currently just doesn't reproduce.
Status stays C (not reproducible). Refs SPECS/012.
On the same exit node (VLESS NL 154.83.159.64:8443) and same network where the
baseline zombie download-stall was caught, the bug no longer reproduces on core
1.13.13-lx.14 — neither normally nor under the probe with LX_CONN_TRACE=0 (code =
release, env set via Android wrap.<pkg> prop, verified in /proc/<pid>/environ).
Likely cause: the §010 GRO split-brain fix landed in lx.14 (lx.12→lx.14 bumped the
wireguard-go submodule 27290b6d→6513629). §010 was literally "no-detour WG-endpoint
killed download on android" via receive-side coalescing — the same symptom class.
The original "also hangs on VLESS" note (→ "not a §010 dup") was never strictly
confirmed (wlan0 encrypted, core-log silent on direction), so §012 is most likely
a manifestation of §010 rather than a separate VLESS bug.
Honest caveats: the counter-proof (run baseline on a pre-lx.14 core on the same
node) was not done; the bug was intermittent. Status is "not reproducible", not
"root cause proven". Folder renamed 012-B-O → 012-B-C.
Probe stays as history on branch lx-conn-trace-probe (b6d8c40a, not deleted) — if
the symptom returns, activate via wrap.<pkg> LX_CONN_TRACE=… per RUN-PLAN.md, but
first make the wrapper transparent (it currently silences ReadWaiter/copyDirect on
reality-download and could mask the bug).
Closes SPECS/012.
Synchronous dual tcpdump (tun0+wlan0, single phone clock, from SYN) localized the
stall inside the kernel on the download direction (remoteConn→conn): 777B arrived
on wlan0 but never reached the app on tun0. Root cause not yet confirmed from
inside the kernel — pcap + code-reading only.
SPEC documents symptom, the synchronous-pcap proof, ruled-out causes (MSS, exit
protocol, server/edge/node, RST storm, §010 GRO), and strictness caveats. Adds
22.06 corrections that retire false leads: conn.go:262 pointed at the UDP
canceler not the copy; run_core.log was a launcher error (no macOS `timeout`),
not a kernel log; the device ran the release core without lx changes (the
LX_TCP_RESPONSE_TIMEOUT prototype was never on it).
PROBE.md describes the LX_CONN_TRACE probe (byte read/write counters per copy
direction, periodic tick + final snapshot) that splits the fork: read=0 → above
copy (proxy decrypt); read>0,write=0 → tun write stall. instrumentation.patch is
the self-contained probe diff; RUN-PLAN.md is the on-device run procedure.
Probe code lands on a separate branch (lx-conn-trace-probe) for CI builds, not on
lx — it forces the buffered copy path (loses splice) and is diagnostic-only.
Refs SPECS/012 (status O — probe written, not yet run on device).
No-detour WireGuard-endpoint killed download on android: UDP_GRO was enabled and
rxOffload read true, but the GRO receive dispatcher in bind_std.go is gated on
GOOS=="linux" (android is not "linux") → a coalesced super-packet was read as one
datagram and corrupted the WG stream. Gate UDP_GRO + rxOffload behind !android
(TX/GSO untouched; non-android linux unchanged).
Confirmed on device (CPH2411/Android-15): pre-fix probe rxoffload=true+dispatch=
single; post-fix rxoffload=false, download 0.44→20.7 Mbps, on par with a control
node on the same LTE cell. Candidate #2 (silent handover) not needed.
Bumps wireguard-go submodule pin to 6513629 (fix, no probe). Probe instrumentation
was never on lx — it lived only on the temporary gro-probe-010/*-verify branches.
Closes SPECS/010.
No reality+xhttp node available; per owner decision the fix is accepted on
synthetic evidence (line-by-line Xray contract match, issue #5635, hiddify
parity, green unit tests + check + builds). Live against a real Xray server
remains an open TODO documented in the 011 REPORT — re-open if it diverges.
- SPECS/011 → status C (folder 011-B-C); REPORT carries an honest live caveat.
- 002 REPORT: stream-one marked fixed-by-011 (was 'known bug').
- SPECS/README roadmap: add 011 row, update 002.
Branch lx-xhttp-streamone; NOT merged into lx.
stream-one was sending <path>/<sessionId>; Xray's splithttp server routes the
bidirectional stream-one handler only on an empty sessionId, so the request must
target the bare normalized path. With the sessionId present the server took the
stream-down branch and the response body carried non-VLESS bytes → VLESS
'unknown version'. Now dialStreamOne uses requestURL() (bare <path>, no trailing
slash); stream-up/packet-up keep their sessionId/seq.
mode=auto now mirrors Xray: reality → stream-one, otherwise packet-up. Reality is
detected by runtime type name (reality_detect.go) with kTLS unwrapping, avoiding a
with_xhttp→with_utls compile dependency so with_xhttp builds without with_utls.
Refs SPECS/011. Synthetic-validated; live pending.
Manual workflow_dispatch builder for any branch/tag: target ∈
{android-aar, apple-xcframework, binary, linux-musl, all}, branch = any ref.
Mirrors lx-release build jobs but uploads artifacts instead of releasing.
Lives on lx (default branch) so `gh workflow run` can find it; each job
checks out the requested branch, so lx source is never required to build it.
gh workflow run lx-build.yml -f target=android-aar -f branch=<branch>
No project code touched — CI file only.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
Serialize probe rounds in startProber to eliminate unbounded fan-out of
fire-and-forget probe goroutines (up to 100/sec per direction), and close
HTTP/3 transports via transport.Close() in addition to CloseIdleConnections.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.
QUIC — revert to ONE Initial. The earlier i1+i2 "developing session" was
conceptually wrong: each DCID is a distinct QUIC connection, so two
Initials with different DCIDs read as two ABANDONED connections (more
anomalous to a DCID-tracking DPI, not less), and a real same-DCID
continuation is impossible (short header is device-blocked; a 1-RTT
packet before the server's reply is an invalid QUIC state). So ip=quic
now emits a single fragmented Initial; realism comes from the
browser-accurate ClientHello (ib → uTLS, device-confirmed working), not
from packet count. masqueI1I2 quic branch returns i2=""; the dead
masqueQUICSecondInitialCPS is removed; the i2-conflict guard is now
sip-only.
SIP — INVITE (i1) + matching 100 Trying (i2), one dialog (separate work):
both whole valid SIP messages sharing Via branch / From tag / Call-ID /
CSeq from a single newSIPDialog pass; pseudo user/host names. (Device
result: still times out on the WARP DPI — see memory; kept for other
providers.)
Both build tags (with_utls / no-utls) build & test green; gofmt/vet clean;
lx-build ok. Docs: SPEC §9 rewritten (multi-packet QUIC considered &
rejected), §10 scope, IMPLEMENTATION_REPORT R11 marked rejected; tests
updated (TestAwgIpcLinesQUICSingleInitial, NonSIPNoI2, SIPExplicitI2Conflict).
The Ib hint finally affects the wire: ip=quic + ib=chrome|firefox builds
the ClientHello with uTLS (github.com/metacubex/utls, the same lib Reality
uses) so the decoy carries a genuine browser JA3/JA4 instead of our
generic ClientHello.
- buildClientHello is now a dispatcher: ib=""/curl → buildGenericClientHello
(the ~294B device-proven CH, unchanged default; uTLS has no curl-QUIC fp),
ib=chrome/firefox → buildBrowserClientHello.
- quic_clienthello_utls_awg.go (with_awg && with_utls): UQUICClient in QUIC
mode with HelloChrome_120 / HelloFirefox_120, ALPN forced to h3, and the
PQ hybrid key_share (X25519MLKEM768, ~1.2KB) stripped so the CH fits one
Initial (reality_client.go pattern). TLSVersMin/Max pinned to 1.3 (QUIC
requirement). Result ~510-620B — a real late-2023 browser JA3.
- quic_clienthello_utls_stub_awg.go (with_awg && !with_utls): graceful
fallback to the generic CH when uTLS isn't built.
- The larger CH re-shapes fragmentation, but planFragmentsN cuts any length
and I1–I4 hold (verified). i2 (multi-packet) uses the same browser too.
WHY ib is optional / forward-looking: on the target DPI ip=quic already
passes on fragmentation alone (no fingerprint check), so the default ib=""
keeps the device-proven generic path; the uTLS CH is a knob against a
future JA3/JA4-classifying DPI and is not itself device-verified. Honest
caveats: JA3 matches a pre-PQ browser (no MLKEM key_share), and JA4 (sorts
+ ignores GREASE) is not fooled by it.
Test (quic_clienthello_utls_awg_test.go, with_utls): chrome/firefox yield
distinct larger ClientHellos that still decrypt + carry SNI + offset≠0
(I1); chrome has GREASE ciphers, firefox doesn't; ""/curl stay generic.
Both tag combos build & test green; gofmt/vet clean; lx-build ok;
sing-box check passes for ib=chrome. SPEC §4/§6/§7 + IMPLEMENTATION_REPORT
R12 + TASKS updated.
ip=dns no longer requires id: when absent, the QNAME is a generated
pronounceable pseudo-domain — consistent with ip=sip's pseudo-host
fallback, and it removes the hardcoded-default-beacon problem (every
default user would otherwise share one QNAME).
- pgDomainHost (pseudo_gen_awg.go): domain-only pseudo name (2-/3-level
LDH), NEVER an IP or a "sip." subdomain — a DNS query for a bare IP or
a sip-prefixed name is implausible, unlike pgHost (which sip uses and
where an IP host is fine). Per-build (baked into the <b> blob), not
per-packet: fresh between users/regenerations, removing the cross-user
signature; the QNAME is fixed within one node's packets (CPS can't do a
pronounceable variable-length name per packet).
- masque_awg.go dispatch: ip=dns with empty id → pgDomainHost(); a set id
is still LDH-validated. id is now REQUIRED only for quic (SNI).
Tests: TestMasqueI1DomainRequiredForQUICOnly (only quic errors on empty
id); TestMasqueI1DomainOptionalForNonQUIC now also checks dns-without-id
produces a valid query whose QNAME is a multi-label pseudo-domain (no IP).
Docs (SPEC/EXAMPLES/IMPLEMENTATION_REPORT/TASKS/README/lx-config): id
required only for quic. sing-box check: ip=dns without id now passes.
Bring the user-facing docs in line with the as-built 009 masquerade after
the lx.12 release:
- README.md / README.ru.md: feature table + masquerade section rewritten —
quic is the only device-proven profile on a real LTE/WARP DPI (~330 ms),
now multi-packet (i1+i2) with a randomized per-call layout; dns/stun/sip
are correct client-initiated requests but blocked as a protocol class to
the WARP edge (kept for other providers). id required for quic/dns only.
- docs/lx-config.md: quic = i1+i2 + randomized; stun = Binding Request (was
"Binding Success Response"); sip = INVITE+SDP (was "200 OK response");
profiles framed as client-initiated, not WireSock server responses.
- docs/lx-changelog.md (new): fork changelog (lx.11, lx.12). Kept separate
from upstream changelog.md so a rebase onto upstream stays conflict-free.
Docs only.
ip=quic now emits TWO independent fragmented QUIC Initials (i1 + i2), so
the decoy flow reads as a developing QUIC session (two session starts)
instead of a single opener — lowering the single-packet signature.
Device-verified: an explicit i1+i2 config brings the WARP tunnel up with
NO latency regression vs i1-only (~340ms), confirming the multi-packet
form is safe for the handshake.
- masqueQUICSecondInitialCPS (quic_initial_awg.go): a second full
fragmented Initial with its OWN fresh DCID. NOT a short-header (that
was device-blocked, commit 64ce4a47) and NOT a DCID-reuse 1-RTT (an
impossible QUIC state that reads anomalous) — two independent Initials
just look like two QUIC sessions starting, which a browser does
routinely.
- masqueI2 (masque_awg.go): dispatch — only ip=quic fills i2; dns/stun/sip
return "" (single-packet decoys).
- awgIpcLines (device_awg.go): wires masque i2 into the i2 slot; guards an
explicit user i2 alongside id/ip/ib as a conflict (mirrors the i1 guard).
Safe by construction: i1/i2 are separate UDP datagrams sent before the
independently-built MessageInitiation (send.go), so neither touches the
real handshake.
Tests (device_awg_test.go): ip=quic fills both i1 and i2 as valid
independent Initials (different DCID, both carry the SNI, first CRYPTO
offset≠0); non-quic leaves i2 empty; explicit-i2 conflict rejected.
Spec §9 updated from hypothesis to implemented + device-verified; §9.2
residual risks (retry head-of-line budget — i3..i5 kept empty), §9.3
what's verified vs deferred. IMPLEMENTATION_REPORT R11 added.
Full package green; gofmt/vet clean; lx-build ok; sing-box check passes
for ip=quic and rejects explicit-i2 conflict.
Add two analysis sections to the 009 spec and normalize naming to 009
(the feature lives here; the LxBox task 146 is only the upstream source
of requirements/device facts, not "our" number).
SPEC.md:
- §8 "Active probing — граница односторонней маскировки (гипотезы)":
H3 (high confidence) a one-sided client decoy is only as strong as
what the TARGET SERVER genuinely serves on that port; the dns/stun/sip
timeouts are consistent with three DPI models (passive
destination-reputation / protocol allowlist / active probing) — H1/H2
(active-probe wording) are medium-confidence, with the honest caveat
that ":2408 answers QUIC" is unverified (it's the WG port, not :443).
Falsifiable device tests T1–T3 (incl. T3: point ip=quic at a
non-QUIC-serving host → should time out, isolating the borrowed
responder from the QUIC bytes).
- §9 "Многопакетная QUIC-последовательность (гипотеза усиления)":
i1..i5 multi-packet design, with a line-by-line send.go proof it CANNOT
break the WARP handshake (decoys are separate UDP datagrams before the
independently-built MessageInitiation). Flags the critical regression
trap: a short-header i2 is exactly the construct commit 64ce4a47
deleted as device-blocked; DCID-reuse is likely a fingerprint, not a
win; retry amplifies head-of-line bytes. Hypothesis to device-test
(bar: i2 must be no worse than i1-only), NOT a shipping decision.
Naming: drop "§146 §N" cross-refs to the LxBox spec from kernel comments
(they don't resolve inside 009); keep the two honest external source
refs (the task file path + "источник device-фактов — LxBox-задача 146").
Also refresh the masque_awg.go header (profiles are now Initial / query /
Binding Request / INVITE, not the old short-header/response list).
Docs only; no code-logic change (build/tests/gofmt/vet green).
ip=sip emitted a `SIP/2.0 200 OK` response as the client's first,
unsolicited packet — a server-role packet in the client's slot, the same
wrong-direction anomaly the old STUN/DNS profiles had, and it was missing
the Contact/Max-Forwards a flow opener needs. Replace it with a SIP
INVITE request carrying an SDP offer — what a UA legitimately sends first
to start a call.
New sip_invite_awg.go (masqueSIPInviteCPS):
- request-line INVITE sip:<user>@<host> SIP/2.0 (method, not a status).
- Via(branch=z9hG4bK)/Max-Forwards:70/From(tag)/To(no tag yet)/Call-ID/
CSeq:N INVITE/Contact, Content-Type: application/sdp, exact
Content-Length, SDP body (v=0, m=audio, rtpmap PCMU/PCMA/telephone-event).
- Hybrid randomization (no cross-user signature): pronounceable user names
and (when id is empty) the host come from PseudoGen and are baked into
<b> at build time (unique between users); volatile tokens (branch /
From-tag / Call-ID / CSeq / SDP session-id+version) are per-packet
<rc>/<rd> of fixed width, so Content-Length stays exact.
New pseudo_gen_awg.go: pronounceable pseudo names / hosts / public IPs
(ported from the LxBox PseudoGen §127) — plausible without being a
hardcoded RFC beacon (bob@biloxi.com) or obvious garbage, and never a
private IP. crypto/rand, not seeded.
id is now OPTIONAL for sip (empty → pgHost()); required only for quic
(SNI) and dns (QNAME); stun ignores it. Removed masqueSIPResponseCPS.
HONEST STATUS: not device-tested on WARP, but expected to time out like
dns/stun — SIP to the datacenter WARP edge :2408 is the same
destination-class anomaly (SIP lives on :5060 / a SIP server). The INVITE
form fixes the direction anomaly of the old 200 OK but not the
destination one. QUIC remains the only proven WARP mechanism; sip is the
strictly-better, direction-clean form kept for other providers.
Tests: TestMasqueSIPResponseStructure → TestMasqueSIPInviteStructure +
TestMasqueSIPInviteNoID (request-line, To-without-tag, exact
Content-Length, names not hardcoded, no-id → pseudo-host). Validation
test updated (id required for quic/dns only). Docs updated.
ip=dns emitted an EDNS OPT *response* (QR=1) as the client's first,
unsolicited packet — a wrong-direction anomaly (a response is a
server-role packet), the same defect STUN had. Replace it with a client
DNS *query* (QR=0, QTYPE HTTPS/65): what a client legitimately sends
first. Only two wire changes from the old code — FLAGS 0x8180→0x0100 and
QTYPE 0x0001→0x0041 — everything else (encodeDNSName, OPT RR, 0xFDE9
cover option) reused. Renamed masqueDNSResponseCPS → masqueDNSQueryCPS;
TXID/cover stay fresh per packet (<r 2>/<r 40>).
DEVICE RESULT (honest): the DNS query also TIMED OUT on the target
LTE/WARP DPI, as the design predicted. Confirms the fundamental finding:
packet quality and direction (request vs response) are secondary — the
blocker is the (protocol + destination) pair. The DPI cuts DNS/STUN/SIP
to the WARP edge 162.159.x:2408 as a protocol class, because raw
DNS/STUN/SIP to a datacenter IP is itself anomalous (DNS lives on :53, a
resolver — not a datacenter edge). QUIC alone bypasses the destination
check: QUIC/HTTP3 legitimately goes anywhere (the whole HTTP/3 web), so
QUIC to a Cloudflare IP is expected traffic.
So QUIC remains the only proven mechanism on this provider. The DNS query
is committed as the strictly-better (direction-correct, RFC-clean) form
and kept — like stun/sip — for other providers whose DPI only checks
well-formedness, NOT protocol-to-destination. Marked not-confirmed-on-WARP
in SPEC / EXAMPLES / IMPLEMENTATION_REPORT / lx-config.
Test: TestMasqueDNSResponseStructure → TestMasqueDNSQueryStructure
(QR=0, QNAME round-trips, QTYPE HTTPS, OPT to end). Docs updated.
QUIC (the proven mechanism) hardened, and the dud STUN profile rebuilt
as the strongest possible shape for other providers — guided by device
A/B on the target LTE/WARP DPI (only QUIC passes there; STUN is blocked
as a protocol class regardless of packet quality).
QUIC — randomize the fragment layout per call + robustness knobs:
- planFragmentsN: random cut points (was the fixed etalon offsets).
- randomizedWirePlan: random out-of-order CRYPTO permutation, repaired so
the offset-0 fragment is never first; PING/PADDING woven into random
gaps; one flex PADDING run pins the payload to the length field. I1–I4
hold by construction (stress test: 300 random packets).
- quicGenParams knobs (default 6 frags / 2 PING / 1250B): fragment count,
PING count, datagram-size range — escalation without a code change if a
DPI ever starts keeping a reassembly buffer. Length field / payload are
recomputed from the chosen size.
- Removed the now-dead etalonWirePlan / etalonCutpoints / planFragments.
STUN — Binding Request instead of Success Response:
- New stun_request_awg.go: a full WebRTC connectivity check (USERNAME,
ICE-CONTROLLING, PRIORITY, SOFTWARE=libwebrtc, MESSAGE-INTEGRITY
HMAC-SHA1, FINGERPRINT CRC-32), fresh txn/ufrag/key per call.
- A response sent unsolicited as the client's first packet is a
wrong-direction anomaly; a request is what an ICE client sends first.
- Removed masqueSTUNResponseCPS (+ orphaned be32/stunSoftwareLen).
- HONEST: this did NOT pass the target DPI (Timeout, like the old
response) — that DPI blocks STUN to a datacenter Cloudflare IP as a
class. Kept as the best shape in case another provider's DPI only
checks well-formedness. QUIC stays the only proven mechanism.
Tests: TestQUICInitialRandomizedInvariants (80 samples, I1–I4 + offsets
differ), TestQUICInitialRobustnessKnobs (4/10/12 frags, variable size),
TestMasqueSTUNRequestStructure (type 0x0001, FINGERPRINT verifies,
USERNAME+MESSAGE-INTEGRITY present), TestMasqueSTUNRequestUniqueness.
Docs: SPEC.md / IMPLEMENTATION_REPORT.md updated incl. the device record.
QUIC ip=quic now generates an out-of-order fragmented Initial (146,
commit 64ce4a47), not a 1-RTT short header. Rewrite the 009 spec docs to
describe the as-built state only — no design history, no superseded
short-header rationale, no open forks.
- SPEC.md: rewritten — I1 CPS mechanism, fragmented QUIC Initial with
I1–I4 invariants, crypto, validation, file map; drops the revert note,
the S1–S4 fork, and the short-header design.
- IMPLEMENTATION_REPORT.md: rewritten as a register of decisions R1–R7,
each with rationale and code refs (file:line) + commits.
- TASKS.md: clean checklist of the current state (no amendment block,
no strikethrough); device-smoke on DPI marked passed.
- EXAMPLES.md: drop §146-amendment blocks and the dangling masque_quic_awg.go
reference; fix the wrong "ib selects a ClientHello profile" claim (ib does
not affect the bytes); drop the dead PLAN.md link.
- Delete PLAN.md and HANDOFF_PROMPT.md (process docs, obsolete post-impl).
Docs only; no code change.
The ip=quic masquerade emitted a QUIC 1-RTT short header, which was
empirically BLOCKED by a real LTE-operator DPI (device-proven A/B, LxBox
task §146). Replace it with a full out-of-order fragmented QUIC Initial
(RFC 9001): a realistic browser-shaped ClientHello (id as SNI) split
across 6 CRYPTO frames in a permuted wire order — first frame offset≠0,
offset-0 frame near the end, PING/PADDING interleaved. A line-rate DPI
grabs the first frame, assumes offset 0, parses garbage and fails open;
a real QUIC server reorders the frames normally. This reverses the old
short-header rationale (the ≥1200-byte Initial it called impossible is
exactly what RFC 9000 §14.1 mandates, and the short header lost on DPI).
- New: quic_initial_awg.go (varint encoder, fragment plan + I1–I4
invariants, RFC 9001 Initial assembly), quic_clienthello_awg.go
(realistic ~294B TLS 1.3 ClientHello), quic_crypto_awg.go (HKDF /
AES-128-GCM-XOR-nonce / header protection, mirrored byte-for-byte from
common/sniff qtls so the keys match the live sniffer).
- id is now REQUIRED for ip=quic (it becomes the ClientHello SNI):
required for quic/dns/sip, optional only for stun.
- A flex PADDING run pins the payload to the length field for any SNI
length, so a long (≤253B) valid domain no longer overflows generation.
- Deleted masque_quic_awg.go (masqueQUICShortHeaderCPS / quicFirstByte).
- CPS transport unchanged: whole encrypted Initial emitted as one <b>
blob; fresh DCID + TLS random + ephemeral x25519 baked in per call.
- Tests reverse-parse our own output (decrypt, frame-walk, reassemble,
SNI) — §5 control vectors, I1–I4, uniqueness, long-SNI regression.
Cross-checked: the live common/sniff QUIC sniffer parses our Initial
and classifies it as chromium.
- Docs: README/README.ru/lx-config + SPECS/009 updated (id required for
quic, fragmented-Initial mechanism, short-header marked superseded).
Device-smoke on the blocked LTE network is a manual gate (not CI):
tunnel up + real traffic through DPI, control node alongside.
Add a Masquerade id/ip/ib row to the feature table and an AmneziaWG sugar
subsection (quic/dns/stun/sip, id required only for dns/sip, ib quic-only) in
both English and Russian READMEs. Link SPECS/009 report + examples.
Add declarative masquerade fields id (domain) / ip (protocol) / ib (browser)
on a wireguard endpoint — WireSock-style sugar over the AmneziaWG I1 CPS string.
Profiles quic/dns/stun/sip generate a protocol-shaped decoy packet, ported in
structure from the open-source WireSock reference (amneziawg-proxy/src/
transform.rs, MIT).
Mechanism: I1 CPS only (S1-S4 padding is impossible against Cloudflare WARP,
the target this eases connecting to); the vendored wireguard-go submodule is
untouched. QUIC is a 1-RTT short header (no SNI/ClientHello/JA3), matching
WireSock — no false "byte-perfect"/"fingerprint" claims. id is required only for
dns/sip (it lands on the wire as QNAME / SIP host) and optional for quic/stun;
when set it is always LDH-validated (mirror of is_valid_sni_hostname) as a
security boundary against SIP/DNS injection. ip is mandatory whenever any of
id/ip/ib is set; id/ip/ib are mutually exclusive with an explicit i1.
Gated by with_awg (rejected with a clear error otherwise); empty id/ip/ib leave
the config byte-identical to upstream. Tests assert each profile parses back as
its protocol (not tautologies); every CPS spec was verified against the real
amneziawg-go newObfChain. sing-box check passes for all four profiles and
rejects the conflict/injection/bad-value cases. See SPECS/009.
SPEC/REPORT/TASKS updated: selector-in-the-middle is now covered by a
runtime selector-guard (suspend AWG consumers before the switch), no longer
an uncovered case. Two complementary guards: Start-guard (static chain) +
selector-guard (runtime selector switch).
Start-guard covers a static detour chain but stops at a selector (its
chosen member is runtime-resolved). This adds the runtime half: in
Selector.SelectOutbound, BEFORE committing the switch, if the new member
reaches a wireguard endpoint, walk up the reverse-dependency ledger
(OutboundManager.ConsumersOf) and SuspendAmneziaWG() every AmneziaWG
consumer of the group — device down, started=false. Suspending before
s.selected.Store closes the race: by the time the group points at the WG
member, the consumer is down and a reconnect fails with "not ready"
instead of sending a junk handshake into WireGuard.
New adapter.AmneziaWGSuspendable marker + OutboundManager.ConsumersOf let
protocol/group act without importing protocol/wireguard. Plain-WG and
non-AWG consumers are left untouched. Variant B throughout.
Refs #2
SPEC/REPORT/memory updated: lazy dialer-guard removed (unverifiable +
sync.Once stale), selector-in-the-middle is a known uncovered case pending
a reliable selector-switch hook.
The lazy DetourDialer guard (lx.8) never fired on device — the hang is in
Endpoint.Start, before any dial — and it is unverifiable in the LxBox UI
(detour targets real servers, not groups) and sync.Once-caches its verdict
so it can't catch a selector changing at runtime. Revert
common/dialer/{detour,dialer}.go to upstream and remove its test.
The Start-guard in protocol/wireguard (field-verified on lx.9) stays as the
sole guard. Selector-in-the-middle is now a known uncovered case.
Refs #2
Field feedback: the user-facing message named Android / kernel hang, but
the restriction is architectural — amneziawg over wireguard is not
supported, period. Reword both guards (Start + dialer) to that; keep the
why (Android hang) in code comments for developers.
Refs #2
lx.8 lazy-only guard didn't fire on device (hang in Start before dial,
logcat-proven). SPEC/REPORT updated: two echelons (Start-guard for direct
transitive detour chain, dialer-guard for selector-in-the-middle).
The lazy DetourDialer guard (lx.8) never fired on Android: an AWG node
whose detour reaches a wireguard endpoint hangs synchronously in
Endpoint.Start (peer-domain resolve over the detour + junk handshake),
before any dial. Proven by logcat — kernel stuck in Starting, no guard
error logged.
Add a Start-guard in protocol/wireguard.Endpoint.Start: walk the
transitive detour chain (OutboundManager + Dependencies); if it reaches a
type=wireguard endpoint, log and skip device startup (started stays false)
so the instance comes up and other outbounds keep working — variant B,
never abort start. Stops at selector/urltest groups (runtime target),
leaving that case to the lazy dialer guard, which stays as the second
echelon.
Refs #2
- 007 AWG_OVER_WIREGUARD_DETOUR_GUARD: SPEC/PLAN/TASKS/REPORT, status C
- 008 AWG_JUNK_PARAM_VALIDATION: SPEC/PLAN/TASKS/REPORT, status C
- SPECS/README roadmap rows for 007 (#2) and 008 (#3)
An AmneziaWG node with detour into any wireguard-based endpoint (plain WG
or AWG) ends up tunnelling AWG traffic inside WireGuard, which hangs the
kernel on Android. Guard it in DetourDialer.init() like the empty-direct
check: lazy error, so the instance still starts and other outbounds keep
working while this node fails every dial (variant B).
Owner-is-AWG flows in via dialer.Options.IsAmneziaWG; the target is matched
by Type()==wireguard, expanding selector/urltest groups recursively. Detour
into a non-wireguard outbound (vless, …) and WG->AWG stay allowed.
Fixes#2
amneziawg-go sizes junk packets rand(0..jmax-jmin)+jmin before each
handshake; jmin>jmax makes rand.Int's argument <=0 and panics in the
retransmit-timer goroutine. validateJunk rejects it in awgIpcLines so the
config fails at endpoint build / sing-box check instead of crashing later.
Only the crash case is guarded; jc/size inconsistency stays allowed
(harmless, keeps the diff minimal, avoids rejecting working configs).
Fixes#3
The 006 spec folder was committed under its -N- name, then renamed to -C-
on disk without removing the old paths from git, leaving a duplicate
SPEC/PLAN/TASKS under 006-F-N-. Drop it; 006-F-C- is the canonical set.
Cosmetic (SPECS docs only) — does not affect release artifacts.
Replace the hard 'exactly two features and nothing else' with a small
client-side feature set (currently XHTTP + AWG2). New features are
allowed over time if they meet the constitution criteria: not planned
upstream, isolated per §3.2-3.3 (new files, own build tag, marked
seams), full Spec Kit cycle. Mirrored in both READMEs and lx-config.md.
lx-ci linux_musl smoke green for amd64/arm64/armv7/mipsle-softfloat —
all statically linked, no libdl.so.2, naive preserved. mipsle+naive
built with musl static, no fallback needed.
Widening H1..H4 from uint32 to MagicHeader (005) changed the longest type
in the struct, so gofmt re-aligns the json tags. go vet didn't catch it;
the lx-ci gofmt check did (red since lx.6). Format-only, no behavior change.
lx-release.yml: new build_linux_musl job (amd64/arm64/armv7/mipsle) that
clones cronet-go, fetches the Chromium musl toolchain via cmd/build-naive,
and builds CGO_ENABLED=1 with with_musl (swapping with_purego) so libcronet
is linked statically — no libdl.so.2, runs on musl routers, naive kept.
Linux moves out of the desktop build job. Artifact names mirror upstream
arch suffixes (armv7, mipsle-softfloat) without the -musl suffix since
Linux ships a single (musl) variant.
lx-ci.yml: dispatch-only linux_musl smoke job runs the same pipeline
(build + verify statically-linked / no libdl) without publishing.
Keep NaiveProxy (upstream feature) by mirroring upstream build.yml's musl
path instead of dropping it. Closes the libdl.so.2 failure on AsusWRT
Merlin + adds linux-armv7. CI-only, no Go code. See issue #1.
awg2_ranged.json (fake keys) exercises ranged H1-H4 through sing-box
check; wire it into the positive and negative CI checks alongside
awg2_basic.json.
Route h1..h4 through the writeStr path as canonical spec strings
("N" or "N-M") instead of writeUint, re-validating each with the key
name in the error. Unset headers are omitted, so a plain WireGuard
endpoint still yields a byte-identical device config.
Add option.MagicHeader (string-based, comparable): a single uint32 or an
inclusive "N-M" range (AWG 2.0 ranged headers from awg2 exports).
- UnmarshalJSON accepts a JSON number (backward compatible with the prior
uint32 field) and a JSON string "N"/"N-M"; canonicalizes, 0 -> unset.
- MarshalJSON keeps type fidelity: single value -> number, range -> string.
- Spec() re-validates for options built in code (libbox/launcher) bypassing
JSON. string base keeps AmneziaWGOptions comparable so IsSet() still works.
H1..H4 change from uint32 to MagicHeader.
README claimed naive/cronet builds CGO-free on "every target"; that is now
inaccurate — the windows/386 legacy (Win7) build drops with_naive_outbound
because cronet-go has no windows/386. Corrected README EN/RU and added the
caveat to docs/lx-config.md §3.
Mirrors upstream build.yml: the windows/386 leg uses a Win7-patched Go
(.github/setup_go_for_windows7.sh — MetaCubeX/go reverts of the Win7
removals) so the binary runs on Windows 7. Drops with_naive_outbound for
this leg (cronet-go has no windows/386 build); the rest of LX_TAGS
compiles for 386. Archive: sing-box-<ver>-windows-386-legacy-windows-7.zip,
matching the launcher's singbox-launcher-win7-32 (also 386).
AmneziaWG s3/s4 prepend junk to every transport message, so a plain-WG
MTU overflows the path and data packets fail with EMSGSIZE while the
handshake still succeeds. On an AWG endpoint (max(s3,s4) > 0):
- when mtu is unset, default to the recommended 1280 (not upstream 1408)
- when mtu is set too high for a conservative 1492-byte (PPPoE) budget,
log an advisory warning: mtu <= 1492 - 28 - 32 - max(s3,s4)
Plain WireGuard is untouched. Docs: lx-config.md §2 + SPECS/003 report.
Upstream-file edit (// lx:no-tailscale marker): remove with_tailscale (+ts_omit_*)
from build_libbox sharedTags. Client fork has no tailscale endpoints and it is the
largest dependency in the APK; keeps the AAR aligned with the desktop LX_TAGS set.
- IMPLEMENTATION_REPORT: live test vs real Xray (3x-ui) — packet-up/auto pass
(handshake + DNS + HTTPS + 2 MB download); padding fix (x_padding in Referer);
stream-one has a known downlink-framing bug
- README (EN/RU), docs/lx-config.md, SPECS roadmap updated; 002 folder -O- -> -C-
Verified live against a real Xray (3x-ui) XHTTP server (VLESS + Reality):
- padding must be carried as x_padding=<zeros> inside the Referer header
(Xray default PlacementQueryInHeader), not a standalone X-Padding header —
the server validates x_padding length (default 100-1000) and replies 400 Bad
Request when it is missing/out of range.
- auto now maps to packet-up (validated working); stream-one has a known
downlink-framing bug ("unknown version") and must be selected explicitly.
packet-up/auto: handshake + DNS + HTTPS + 2 MB download all flow through the tunnel.
- README.md is now the lx README in English (GitHub renders it → an arriving
visitor immediately sees this is a thin sing-box fork with XHTTP + AmneziaWG 2.0)
- README.ru.md: Russian version; mutual language switcher in both
- drop the static README.sing-box.md copy (it would go stale) in favor of a link
to the live upstream sing-box README on GitHub
- docs/lx-config.md: config reference for XHTTP transport and AmneziaWG 2.0
endpoint (field tables + examples with placeholder keys)
- .github/workflows/lx-ci.yml: matrix over the two features
(baseline / with_xhttp / with_awg / full) + negative check that feature-off
rejects its config; vet job; cross-platform matrix {linux,darwin,windows}x
{amd64,arm64} building the full lx set with the merged AWG fork (submodules)
- link the config doc from SPECS/README
- replace github.com/sagernet/wireguard-go => ./submodules/wireguard-go
(Leadaxe/wireguard-go @27290b6: sagernet base + AmneziaWG obfuscation, 3-way merge)
- add S3/S4 padding to option.AmneziaWGOptions + device_awg.go IpcSet emitter
(AWG 2.x; server config carries s1/s2/s3/s4)
LIVE-VALIDATED against a real AmneziaWG 2.0 server: handshake initiation ->
received handshake response -> keepalive -> traffic egresses via the server.
AWG is now functional, not just config-valid. Secrets never committed.
Verified against XTLS/Xray-core splithttp source (no live server):
- sessionId now formatted as dashed UUID (was 32-char hex) to match uuid.New().String()
- documented version-dependent padding placement (current Xray: x_padding query
param in Referer; older: standalone X-Padding) for live-test reconciliation
build + check + vet green.
- mark registry refactor + xhttp constant complete (pushed earlier)
- SPEC §7: hiddify port pulls vendored common/xray/* + quic-go/http3; record
faithful-vendor (A) vs lean-native (B) decision; recommend A
- status N -> O (in progress)
Return SUCCESS with empty answers instead of an error when the
queried address family has no range configured. Reject configurations
where neither inet4_range nor inet6_range is set.
The TTL computation and assignment loops treat OPT record's Hdr.Ttl
as a regular TTL, but per RFC 6891 it encodes EDNS0 metadata
(ExtRCode|Version|Flags). This corrupts cached responses causing
systemd-resolved to reject them with EDNS version 255.
Also fix pointer aliasing: storeCache() stored raw *dns.Msg pointer
so subsequent mutations by Exchange() corrupted cached data.
- Skip OPT records in all TTL loops (Exchange + loadResponse)
- Use message.Copy() in storeCache() to isolate cache from mutations
Treat rule_set items as merged branches instead of standalone boolean
sub-items.
Evaluate each branch inside a referenced rule-set as if it were merged
into the outer rule and keep OR semantics between branches. This lets
outer grouped fields satisfy matching groups inside a branch without
introducing a standalone outer fallback or cross-branch state union.
Keep inherited grouped state outside inverted default and logical
branches. Negated rule-set branches now evaluate !(...) against their
own conditions and only reapply the outer grouped match after negation
succeeds, so configs like outer-group && !inner-condition continue to
work.
Add regression tests for same-group merged matches, cross-group and
extra-AND failures, DNS merged-branch behaviour, and inverted merged
branches. Update the route and DNS rule docs to clarify that rule-set
branches merge into the outer rule while keeping OR semantics between
branches.
Before 795d1c289, nested rule-set evaluation reused the parent rule
match cache. In practice, this meant these fields leaked across nested
evaluation:
- SourceAddressMatch
- SourcePortMatch
- DestinationAddressMatch
- DestinationPortMatch
- DidMatch
That leak had two opposite effects.
First, it made included rule-sets partially behave like the docs'
"merged" semantics. For example, if an outer route rule had:
rule_set = ["geosite-additional-!cn"]
ip_cidr = 104.26.10.0/24
and the inline rule-set matched `domain_suffix = speedtest.net`, the
inner match could set `DestinationAddressMatch = true` and the outer
rule would then pass its destination-address group check. This is why
some `rule_set + ip_cidr` combinations used to work.
But the same leak also polluted sibling rules and sibling rule-sets.
A branch could partially match one group, then fail later, and still
leave that group cache set for the next branch. This broke cases such
as gh-3485: with `rule_set = [test1, test2]`, `test1` could touch
destination-address cache before an AdGuard `@@` exclusion made the
whole branch fail, and `test2` would then run against dirty state.
795d1c289 fixed that by cloning metadata for nested rule-set/rule
evaluation and resetting the rule match cache for each branch. That
stopped sibling pollution, but it also removed the only mechanism by
which a successful nested branch could affect the parent rule's grouped
matching state.
As a result, nested rule-sets became pure boolean sub-items against the
outer rule. The previous example stopped working: the inner
`domain_suffix = speedtest.net` still matched, but the outer rule no
longer observed any destination-address-group success, so it fell
through to `final`.
This change makes the semantics explicit instead of relying on cache
side effects:
- `rule_set: ["a", "b"]` is OR
- rules inside one rule-set are OR
- each nested branch is evaluated in isolation
- failed branches contribute no grouped match state
- a successful branch contributes its grouped match state back to the
parent rule
- grouped state from different rule-sets must not be combined together
to satisfy one outer rule
In other words, rule-sets now behave as "OR branches whose successful
group matches merge into the outer rule", which matches the documented
intent without reintroducing cross-branch cache leakage.
PreMatch and full match phases each created a fresh InboundContext,
causing process search (expensive OS syscalls) to run twice per
connection. Use a freelru ShardedLRU cache with 200ms TTL to serve
the second lookup from cache.
Add fpm-based Alpine APK packaging alongside existing DEB/RPM/Pacman
packages. Alpine APKs use `linux` in the filename to distinguish from
OpenWrt APKs which use the `openwrt` prefix.
CCM: Fix 1M context detection - use prefix match for versioned
beta strings (e.g. "context-1m-2025-08-07") and include cache
tokens in the 200K threshold check per Anthropic billing docs.
OCM: Add GPT-5.4 family pricing (standard/priority/flex) with
extended context (>272K) premium pricing support. Add context
window tracking to usage combinations, mirroring CCM's pattern.
Update normalizeGPT5Model defaults to latest known models.
Support the OpenAI Responses WebSocket API (`wss://.../v1/responses`)
for bidirectional frame proxying with usage tracking.
Fix Codex CLI client config examples to use profiles and correct flags.
Update openai-go v3.24.0 → v3.26.0.
When clients (e.g. Node.js Anthropic SDK) explicitly set Accept-Encoding: gzip,
Go's http.Transport does not transparently decompress the response body, because
it only does so when it added the header itself. This causes CCM's json.Unmarshal
to receive raw gzip bytes, silently failing to parse usage data and leaving the
usage counter unchanged.
Fix: remove Accept-Encoding from the outgoing proxy request. Transport adds it
automatically and transparently decompresses response.Body before CCM reads it.
Wire compression (CCM→Anthropic) is preserved — Transport still negotiates gzip.
Only CCM→localhost path is affected; compression on loopback has no practical
benefit.
Move hardcoded build tags and ldflags from Makefile, Dockerfile, CI
workflows, and local build scripts into canonical files under release/:
- release/DEFAULT_BUILD_TAGS (Linux common archs, Darwin, Android)
- release/DEFAULT_BUILD_TAGS_WINDOWS (includes with_purego)
- release/DEFAULT_BUILD_TAGS_OTHERS (no with_naive_outbound)
- release/LDFLAGS (shared linker flags)
The cache deduplication in Client.Exchange uses a channel-based lock
per DNS question. Waiting goroutines blocked on <-cond without context
awareness, causing them to accumulate indefinitely when the owning
goroutine's transport call stalls. Add select on ctx.Done() so waiters
respect context cancellation and timeouts.
When bbolt encounters corrupted page data at runtime, it panics
instead of returning an error. Wrap all DB transactions with
recover to catch these panics, delete the corrupted database
file, and reopen a fresh one.
- Enable ECH for NaiveProxy outbound with DNS resolver integration
- Add query_server_name option to override domain for ECH HTTPS record queries
- Update cronet-go dependency and remove windows_386 support
Align dev-next-grpc with wip2 by adding UsePlatformWIFIMonitor()
to the new PlatformInterface, allowing platform clients to indicate
they handle WIFI monitoring themselves.
We mistakenly believed that `libresolv`'s `search` function worked correctly in NetworkExtension, but it seems only `getaddrinfo` does.
This commit changes the behavior of the `local` DNS server in NetworkExtension to prefer DHCP, falling back to `getaddrinfo` if DHCP servers are unavailable.
It's worth noting that `prefer_go` does not disable DHCP since it respects Dial Fields, but `getaddrinfo` does the opposite. The new behavior only applies to NetworkExtension, not to all scenarios (primarily command-line binaries) as it did previously.
In addition, this commit also improves the DHCP DNS server to use the same robust query logic as `local`.
Previously, the buffer was not reset within the response loop. If a packet
handle failed or completed, the buffer retained its state. Specifically,
if `ReadPacketFrom` returned `io.ErrShortBuffer`, the error was ignored
via `continue`, but the buffer remained full. This caused the next
read attempt to immediately fail with the same error, creating a tight
busy-wait loop that consumed 100% CPU.
Validates `buffer.Reset()` is called at the start of each iteration to
ensure a clean state for 'ReadPacketFrom'.
The cache lookup was performed before rule matching, using the caller's
strategy (usually AsIS/0) instead of the resolved strategy. This caused
cache misses when ipv4_only was configured globally but the cache lookup
expected both A and AAAA records.
Remove LookupCache and ExchangeCache from Router, as the cache checks
inside client.Lookup and client.Exchange already handle caching correctly
after rule matching with the proper strategy and transport.
The Chinese documentation incorrectly stated that the default value for the domain_strategy field in the direct outbound module is dns.strategy. The correct value should be inbound.domain_strategy, as specified in the English documentation. This commit corrects the Chinese documentation to align with the accurate behavior described in the English version.
Signed-off-by: Monica <1379531829@qq.com>
For historical reasons, sing-box's `domain_suffix` rule matches literal prefixes instead of the same as other projects.
This change modifies the behavior of `domain_suffix`: If the rule value is prefixed with `.`,
the behavior is unchanged, otherwise it matches `(domain|.+\.domain)` instead.
The `process_path` rule of sing-box is inherited from Clash,
the original code uses the local system's path format (e.g. `\Device\HarddiskVolume1\folder\program.exe`),
but when the device has multiple disks, the HarddiskVolume serial number is not stable.
This change make QueryFullProcessImageNameW output a Win32 path (such as `C:\folder\program.exe`),
which will disrupt the existing `process_path` use cases in Windows.
I read other rule_item_xxx.go files, they are all snake case. This description is showed on dashboard like yacd.
Signed-off-by: kkocdko <31189892+kkocdko@users.noreply.github.com>
This helps the daemon work better on IoT devices
like RaspberryPi.
According to systemd's documentation,
`network.target` means there has already been
a network manager started, but the network may
not be "up". On most PCs this does not matter
because the network will turn to "up" almost
immidiately. The IoT devices' network interface
may not be set up quickly enough, so they may
meet that the sing-box daemon is started before
network is ready, which results that sing-box
cannot find a working route. The workaround
of this is restarting sing-box daemon but it
absolutely is not the perfect solution.
As `network-online.target` must be triggered by
network manager after you configured it, I keep
`network.target` so there will be no change to
those who do not enabled proper trigger service
like `NetworkManager-wait-online.service`.
See also: https://systemd.io/NETWORK_ONLINE/
Refactor Authenticator interface to struct &
Update smux &
Update gVisor to 20231204.0 &
Update quic-go to v0.40.1 &
Update wireguard-go &
Add GSO support for TUN/WireGuard &
Fix router pre-start &
Fix bind forwarder to interface for systems stack
Enhanced the issue reporting templates for both English and Chinese versions by adding more structured and comprehensive guideline checkboxes. This aims to ensure contributors provide sufficient and beneficial information for reproducing and resolving issues, thereby improving the quality of reports and making issue tracking more efficient.
Remove the information on password generation for `2022-blake3-aes-128-gcm` cipher from the Server Example section in the shadowsocks.md file as it is no longer needed.
The old meaning is wrong. Correct the meaning according to the English documentation and the actual effect of the option.
Signed-off-by: 嫦悅 <lomombwlo@gmail.com>
description:Please provide the operating system version
validations:
required:true
- type:dropdown
attributes:
label:Installation type
description:Please provide the sing-box installation type
options:
- Original sing-box Command Line
- sing-box for iOS Graphical Client
- sing-box for macOS Graphical Client
- sing-box for Apple tvOS Graphical Client
- sing-box for Android Graphical Client
- Third-party graphical clients that advertise themselves as using sing-box (Windows)
- Third-party graphical clients that advertise themselves as using sing-box (Android)
- Others
validations:
required:true
- type:input
attributes:
description:Graphical client version
label:If you are using a graphical client, please provide the version of the client.
- type:textarea
attributes:
label:Version
description:If you are using the original command line program, please provide the output of the `sing-box version` command.
render:shell
- type:textarea
attributes:
label:Description
description:Please provide a detailed description of the error.
validations:
required:true
- type:textarea
attributes:
label:Reproduction
description:Please provide the steps to reproduce the error, including the configuration files and procedures that can locally (not dependent on the remote server) reproduce the error using the original command line program of sing-box.
validations:
required:true
- type:textarea
attributes:
label:Logs
description:|-
In addition, if you encounter a crash with the graphical client, please also provide crash logs.
For Apple platform clients, please check `Settings - View Service Log` for crash logs.
For the Android client, please check the `/sdcard/Android/data/io.nekohasekai.sfa/files/stderr.log` file for crash logs.
render:shell
- type:checkboxes
id:supporter
attributes:
label:Supporter
options:
- label:I am a [sponsor](https://github.com/sponsors/nekohasekai/)
- type:checkboxes
attributes:
label:Integrity requirements
description:|-
Please check all of the following options to prove that you have read and understood the requirements, otherwise this issue will be closed.
Sing-box is not a project aimed to please users who can't make any meaningful contributions and gain unethical influence. If you deceive here to deliberately waste the time of the developers, you will be permanently blocked.
options:
- label:I confirm that I have read the documentation, understand the meaning of all the configuration items I wrote, and did not pile up seemingly useful options or default values.
required:true
- label:I confirm that I have provided the server and client configuration files and process that can be reproduced locally, instead of a complicated client configuration file that has been stripped of sensitive data.
required:true
- label:I confirm that I have provided the simplest configuration that can be used to reproduce the error I reported, instead of depending on remote servers, TUN, graphical interface clients, or other closed-source software.
required:true
- label:I confirm that I have provided the complete configuration files and logs, rather than just providing parts I think are useful out of confidence in my own intelligence.
--body "Branch \`$BRANCH\` is rebased onto \`$TARGET\`, builds, and passes \`check\`. Auto-PR was blocked — enable Settings → Actions → General → \"Allow GitHub Actions to create and approve pull requests\", or open the PR by hand."
- **XHTTP** transport (\`with_xhttp\`) — Xray-compatible "splithttp", composes with Reality (use \`auto\`; \`stream-one\` has a known framing bug).
- **CommandClient extensions** (\`with_lx_command\`) — native libbox gRPC parity for the Clash API dropped from the **Android AAR**: URLTestOutbound, GetRules, GetGroups/GetOutbounds, Connection.Detour, SubscribeDNSQueries. (Desktop/CLI binaries keep \`with_clash_api\` for external dashboards.)
### Binaries
Drop-in \`sing-box\` for **darwin / windows** × {amd64, arm64}, plus a **Windows 7 (32-bit)** legacy build (\`sing-box-${{ steps.ver.outputs.version }}-windows-386-legacy-windows-7.zip\` — built with a Win7-patched Go; without naive/cronet, which has no windows/386 target).
**Linux — static musl builds for routers** (AsusWRT Merlin, OpenWrt, Keenetic): \`linux-amd64\`, \`linux-arm64\`, \`linux-armv7\`, \`linux-mipsle-softfloat\`. These are statically linked (no \`libdl.so.2\`/glibc dependency) and **keep NaïveProxy** — they run on musl routers where the previous dynamic builds failed with \`libdl.so.2: cannot open shared object file\`. See SPECS/006.
**Linux — big-endian MIPS** (OpenWrt \`mips_24kc\`, e.g. Atheros AR93xx): \`linux-mips-softfloat\` — pure-Go static build **without NaïveProxy** (Chromium/cronet has no big-endian MIPS toolchain); everything else matches the desktop tag set.
Each archive contains the \`sing-box\` binary (\`sing-box version\` reports \`${{ steps.ver.outputs.version }}\`). Verify downloads against \`SHA256SUMS\`.
### Android
\`libbox-${{ steps.ver.outputs.version }}.aar\` (+ \`libbox-legacy-…\` for SDK 21) — gomobile build of \`experimental/libbox\` with \`with_xhttp\`+\`with_awg\` enabled, for embedding in an Android app. \`Libbox.version()\` reports the lx version.
stale-issue-message:'This issue is stale because it has been open 60 days with no activity. Remove stale label or comment or this will be closed in 5 days'
`sing-box-lx` — **тонкий downstream** апстрима [SagerNet/sing-box](https://github.com/SagerNet/sing-box): upstream **плюс ровно две фичи** и ничего больше:
1.**XHTTP** — клиентский v2ray-транспорт (совместимость с Xray XHTTP).
Главная ценность проекта — **согласованность с upstream**. Любое изменение оценивается по тому, насколько легко оно переживёт ребейз на следующий тег upstream.
-`.github/workflows/lx-ci.yml` — build(lx tags) → version → `go vet` → `sing-box check` (linux/amd64; полная матрица — в 004).
-`lx-test/config/minimal.json` — валидный конфиг для `check` (mixed-in + direct-out). Положен в `lx-test/`, **не** в upstream `test/` (там отдельный Go-модуль).
-`AGENTS.md` — указатель для агентов (force-add: upstream его `.gitignore`-ит; новый файл → нулевой конфликт при ребейзе).
-`SPECS/**` — Spec Kit (CONSTITUTION, IMPLEMENTATION_PROMPT, README, задачи 001–004).
**Правок upstream-файлов: 0.**`constant/version.go`, `Makefile`, `.gitignore` — не тронуты.
## Проверки (DoD)
```
$ make -f Makefile.lx lx-version → 1.13.13-lx.1
$ make -f Makefile.lx lx-build → ./sing-box (28 MB)
Хранить в `Makefile` (переменная `LX_TAGS`) и продублировать в `SPECS/CONSTITUTION.md` при изменениях.
> **Обновлено в §004:** набор расширен до полного upstream feature-set (`release/DEFAULT_BUILD_TAGS`) + `with_purego` + наши две фичи, с обязательным `-checklinkname=0` в `LX_LDFLAGS`. Актуальный источник истины — `Makefile.lx` (`make -f Makefile.lx lx-print-tags`) и `SPECS/004`.
## 2. Изменяемые / новые файлы
| Файл | Тип | Изменения |
|------|-----|-----------|
| `Makefile` | new (или дополнение) | Цель `lx-build`: `go build -tags "$(LX_TAGS)" -ldflags "$(LX_LDFLAGS)" -o sing-box ./cmd/sing-box`; переменные `LX_TAGS`, `VERSION=…-lx.$(LX_BUILD)` |
| `lx-test/config/*.json` | new | Sample-конфиги для `sing-box check` (минимальный валидный, без фич — для 001) |
| `SPECS/001-.../IMPLEMENTATION_REPORT.md` | new | Отчёт |
> Версия: upstream хранит строку версии в `constant/version.go` (или собирается через ldflags в `cmd/sing-box`). Проверить фактический механизм и **задавать `-lx` суффикс через `-ldflags -X`**, не правя `constant/version.go` напрямую (иначе лишний `// lx:` дифф на каждый ребейз). Если upstream не поддерживает ldflags-override — тогда минимальная `// lx:` правка в `constant/version.go`.
## 3. Зона касания upstream
-В идеале **ноль** правок upstream-файлов (всё через новые файлы + ldflags).
- Допустимый минимум: одна `// lx:` строка в `constant/version.go`, если ldflags-override невозможен.
## 4. Порядок работ
1. Проверить механизм версии upstream (`constant/version.go`, `cmd/sing-box`).
2.`Makefile`с`LX_TAGS`/`LX_LDFLAGS`/`lx-build`.
3. Sample-конфиг + CI-скелет.
4. Прогнать DoD, заполнить отчёт.
## 5. Риски
- Версионный механизм upstream может не принимать ldflags-override — fallback на `// lx:` правку.
-`with_xhttp`/`with_awg` как несуществующие теги не ломают сборку (Go игнорирует неизвестные build-теги) — но файлов с этими тегами пока нет, это нормально.
Заложить скелет downstream'а`sing-box-lx`: remotes, рабочая ветка, build-теги, версия с `-lx`, конвенция маркеров `// lx:` и шаблон гейтинга. После задачи репозиторий — корректный «upstream + ноль фич», готовый принимать XHTTP (002) и AWG2 (003).
---
## 1. Проблема / контекст
`Leadaxe/sing-box-lx` — форк-зеркало upstream (родословная `SagerNet/sing-box` сохранена). Нужна повторяемая инфраструктура downstream'а, при которой каждое будущее изменение изолировано и ребейзопригодно (см. CONSTITUTION § 3).
- Ветка `lx` базируется на стабильном теге `v1.13.13`. **(сделано)**
- Default branch на GitHub = `lx`; шумные зеркальные ветки (`dependabot/*`, `dev-*`, `copilot/*`) — вне внимания (можно удалить с origin, не обязательно).
### 2.2 Build-теги
- Ввести **`with_xhttp`** и **`with_awg`** как опознаваемые теги проекта (фактический код — в 002/003). Зафиксировать **канонический набор тегов сборки lx** в одном месте (см. PLAN), переиспользуемый в DoD и CI.
- Инвариант: без `with_xhttp`/`with_awg` бинарь ведёт себя как upstream.
### 2.3 Версия
-`sing-box version` должен печатать суффикс **`-lx.N`** (напр. `1.13.13-lx.1`).
- Суффикс задаётся при сборке (ldflags), не хардкодом в исходниках upstream (минимальный дифф).
## v2 — полная клиентская поддержка параметров (2026-06-29)
**Статус:** реализация code-complete + все проверки зелёные; **дефолтный путь лайв-подтверждён на реальных нодах** (4 живых XHTTP-сервера, packet-up + stream-one/reality, скачивание 1 МБ); **лайв obfs/placement** — остаётся открытым TODO (нужен сервер с такой настройкой).
Клиентский XHTTP-транспорт, подход **lean-native** (на примитивах sing-box, минимум зависимостей) — реализован многоагентным workflow в изолированном worktree, влит в `lx` (коммиты `2d97ff56` registry/const + `d1b434fc` транспорт).
**Файлы (новые, если не указано иное):**
-`transport/v2ray/registry.go`, `// lx` в `transport/v2ray/transport.go`, константа в `constant/v2ray.go` — registry-рефактор (ранее).
-`option/v2ray_xhttp.go` — тип `V2RayXHTTPOptions` (Host, Path, Mode, Headers, padding).
-`option/v2ray_transport.go` — **единственная upstream-правка** (// lx): поле `XHTTPOptions` + xhttp-case в Marshal/Unmarshal.
-`transport/v2rayxhttp/{client,conn,register}.go` — клиент; `register.go` под `//go:build with_xhttp`.
-`include/v2rayxhttp.go` (`//go:build with_xhttp`) — blank-import для запуска `init()`.
-`lx-test/config/xhttp_reality.json` — VLESS+xhttp+reality для `check`.
-`go vet` (lx-теги) по `transport/v2rayxhttp`, `option`, `transport/v2ray` → чисто; `go build ./...` без тегов → ок; `gofmt` чисто.
- Негатив: бинарь **без**`with_xhttp` отвергает xhttp-конфиг (`unknown transport type: xhttp`). Невалидный mode → `v2ray-xhttp: unknown mode`. Все 4 mode конструируются.
## Зона касания upstream (ребейз)
Ровно **1 файл**: `option/v2ray_transport.go` (3 правки в // lx-маркерах). Реестр и весь пакет `v2rayxhttp` — новые файлы, конфликтов не дают.
## Лайв-тест (реальный Xray/3x-ui XHTTP-сервер)
Проверено против VLESS + Reality + `type=xhttp` ноды (панель 3x-ui):
- ✅ **packet-up** (и `auto` → packet-up): handshake + DNS + HTTPS (example.com 200) + скачивание 2 МБ @ ~2.1 МБ/с — трафик выходит через IP сервера.
- ❌ **stream-one**: `unknown version` — баг при чтении downlink-ответа (выбирается только явно). → **Исправлено в задаче 011** (корень: stream-one должен слать голый путь без sessionId; auto+reality → stream-one). Принято на синтетике, лайв отложен.
**Ключевой фикс (по исходникам Xray hub.go/config.go + лайв):** padding кладётся как `x_padding=<нули>` в **query внутри заголовка `Referer`** (Xray default `PlacementQueryInHeader`, key `x_padding`), а**не** отдельным `X-Padding`. Сервер валидирует длину `x_padding` (дефолт 100–1000) и без неё отвечает **400 Bad Request**. Плюс `mode=auto` переключён на **packet-up**. Коммит `5a398a5e`. Также ранее: `sessionId` → UUID-формат, path-layout `<path>/<sessionId>[/<seq>]` сверены.
## Остаточные пробелы
1.~~**stream-one** — баг framing downlink (`unknown version`)~~ → **исправлено в 011** (голый путь без sessionId; `auto`+reality → stream-one). Лайв-подтверждение — открытый TODO в 011.
2.**packet-up** без xmux/переиспользования соединений; **stream-up** не лайв-тестился.
3.`x_padding_bytes` — строка «min-max» (нет Range-типа в badoption); дефолт 100–1000.
## Дальше
- Лаунчер: маппинг `type=xhttp` (его задача 023 сейчас маппит в `httpupgrade`) → реальный xhttp-транспорт.
| `serverMaxHeaderBytes` | `http.Server{MaxHeaderBytes}` — лимит размера заголовков входящего запроса | `8192` | У client-only транспорта нет `http.Server` |
| `noSSEHeader` | Сервер не шлёт `Content-Type: text/event-stream` на stream-down GET | `false` (SSE шлётся) | Клиент не читает Content-Type — обрабатывает оба случая без кода |
| `scMaxBufferedPosts` | Ёмкость серверной очереди переупорядочивания upload-POST (packet-up) | `30` | Клиент не знает о глубине буфера сервера |
| `scStreamUpServerSecs` | Интервал (сек, Range) периодической записи `X`-padding в ответ stream-up | `{20,80}` | Клиент тихо отбрасывает эти байты (`io.Discard`) |
**Решение для реализации:** не реализуем. В конфиге — `expose-but-ignore` (принимаем поля, чтобы
server-образные конфиги не падали на парсинге), помечены как inbound-only. Альтернатива (просто
document-and-skip без полей в struct) тоже допустима; финальный выбор зафиксирован в SPEC §6.
- Встроенные типы регистрируются в `init()` (в `transport.go` или соседнем файле) — поведение для http/ws/quic/grpc/httpupgrade без изменений.
-`NewClientTransport` → `ctor, ok := clientRegistry[options.Type]`; нет — прежняя ошибка.
- XHTTP-конструктор регистрируется из пакета `v2rayxhttp` через `init()`**только** под `//go:build with_xhttp` (через проводящий файл, чтобы импорт пакета подтягивался лишь с тегом).
Конструктор XHTTP должен соответствовать сигнатуре `ClientConstructor` (см. upstream `transport.go`): `(ctx, dialer, serverAddr, options, tlsConfig) → (adapter.V2RayClientTransport, error)`. Опции достаются из `options.XHTTPOptions`.
-`mode=auto` в sing-box-портах исторически падает в `packet-up`, что ломало, напр., аплоад в Telegram ([hiddify#2082](https://github.com/hiddify/hiddify-app/issues/2082)) — задокументировать фактический выбор режима.
- Рефактор `switch`→registry должен **точно** сохранить семантику ошибок и nil-обработку (`options.Type == ""` → `nil, nil`).
Добавить **клиентский XHTTP-транспорт** (совместимость с Xray XHTTP) для VLESS/VMess/Trojan, встроив его через **registry-рефактор** диспетчера v2ray-транспортов, за build-тегом `with_xhttp`.
---
## 1. Проблема / контекст
- Upstream sing-box XHTTP не поддерживает и не планирует ([#3550](https://github.com/SagerNet/sing-box/issues/3550)). Сервера на Xray всё чаще только XHTTP (после депрекации части транспортов в Xray).
- В sing-box диспетчер v2ray-транспортов — **хардкод-`switch`** по `options.Type` в `transport/v2ray/transport.go`. Добавлять `case` на каждый ребейз — точка постоянных конфликтов.
## 2. Цель
VLESS/VMess/Trojan outbound с`transport.type = "xhttp"` поднимают рабочее соединение к XHTTP-серверу Xray, в т.ч. поверх **TLS/Reality**. Без тега `with_xhttp` тип `xhttp` отвергается с понятной ошибкой.
- Превратить выбор клиентского транспорта в **реестр**: `transport.RegisterClient(type, ClientConstructor)` + `map[string]ClientConstructor`, заполняемый при `init()`.
- Встроенные транспорты (`http`, `ws`, `quic`, `grpc`, `httpupgrade`) регистрируются как раньше (поведение идентично upstream).
-`NewClientTransport` ищет конструктор в реестре вместо `switch` (поведение для известных типов — без изменений; для неизвестных — та же ошибка `unknown transport type`).
-`option/v2ray_transport.go`: поле `XHTTPOptions XHTTPOptions` в `_V2RayTransportOptions` + тип `XHTTPOptions` (в новом файле `option/v2ray_xhttp.go`, чтобы минимизировать дифф основного файла; в `_V2RayTransportOptions` — одна `// lx:` строка).
### 3.4 TLS/Reality
-`tlsConfig` прокидывается в конструктор как у прочих транспортов → связка **XHTTP + Reality** работает без доп. кода. (XHTTP + XTLS-Vision несовместимы — ограничение протокола, не наше.)
## 4. Критерии приёмки
-`sing-box check -c` принимает VLESS + `transport.type=xhttp` + `tls.reality`.
- Реальный коннект к XHTTP-серверу Xray (ручная проверка), хотя бы `mode=stream-one` и `packet-up`.
- Сборка **без**`with_xhttp`: конфиг с`xhttp` → ошибка `unknown transport type: xhttp` (или эквивалент реестра).
-`go test ./transport/...`, `go vet ./...` зелёные.
- Ребейз-проверка: при следующем upstream-теге конфликты возможны **только** в `transport/v2ray/transport.go`, `constant/v2ray.go`, `option/v2ray_transport.go`.
## 7. Разведка порта и выбор подхода (добавлено по ходу)
**Что показал референс `hiddify/hiddify-sing-box` (`transport/v2rayxhttp/client.go`):** XHTTP в hiddify реализован НЕ поверх примитивов sing-box, а через **вендорённое поддерево Xray** под `common/xray/{buf,net,pipe,signal/done,uuid}` + зависимости `quic-go`, `http3`, `golang.org/x/net/http2`, абстракция `DialerClient`/`XmuxClient` и опции `option.V2RayXHTTPOptions{ V2RayXHTTPBaseOptions }`. Целевой интерфейс прост — `adapter.V2RayClientTransport = { DialContext(ctx) (net.Conn, error); Close() error }` — но реализация тянет много транзитивного кода и завязана на старую версию sing-box hiddify.
**Развилка подхода (зафиксировать перед кодом порта):**
- **(A) Faithful-vendor.** Перенести hiddify `common/xray/*` + пакет `v2rayxhttp` как **новые файлы** (namespaced), адаптировать импорты под v1.13.13. Плюс: максимальная совместимость с реальными XHTTP-серверами, проверенный код. Минус: больший footprint (но всё — новые файлы → **нулевая зона касания upstream**, что согласуется с CONSTITUTION). Тащит `quic-go`/`http3` (часть уже в go.mod sing-box).
- **(B) Lean-native.** Написать компактный XHTTP-клиент на примитивах sing-box (по образцу in-tree `transport/v2rayhttpupgrade`). Плюс: меньше кода, меньше зависимостей. Минус: больше оригинальной работы и риск несовпадения с Xray по краям (`mode=auto`, padding, xmux).
**Рекомендация:****(A)** — приоритет проекта №2 (корректность/совместимость) важнее объёма, а изоляция в новых файлах сохраняет ребейзопригодность. Footprint велик, но не увеличивает конфликтность ребейза.
**Обязательно для приёмки:** живой XHTTP-сервер (Xray) для end-to-end проверки — синтетического `sing-box check` недостаточно (XHTTP под активной разработкой, версии client↔server должны совпадать).
> ⚠️ В `extra` эти значения часто приходят **числом** (`"scMaxEachPostBytes":"1000000"`,
> `"scMinPostsIntervalMs":30.0`). Транспорт sing-box-lx ждёт **строку `"min-max"`** — превратить
> одиночное число `N` в строку `"N-N"` (или просто `"N"` — парсер примет и то, и то). Дробную часть
> у `30.0` отбросить → `"30"`.
### 2.5 Игнорируемые / серверные
| URL-параметр | Действие |
|--------------|----------|
| `scMaxConcurrentPosts` | **Accept-but-ignore.** Legacy-поле старого Xray (в текущем Xray/extended его нет — там 1 POST-тело за раз). Клиент sing-box-lx шлёт upload-POST последовательно (= текущий Xray). Можно влить как `sc_max_concurrent_posts` (принято, но не используется) — или опустить (см. §6). |
| `serverMaxHeaderBytes`, `noSSEHeader`, `scMaxBufferedPosts`, `scStreamUpServerSecs` | server-only. Можно влить как `server_max_header_bytes`/`no_sse_header`/`sc_max_buffered_posts`/`sc_stream_up_server_secs` (клиент их принимает, но игнорирует) — или просто опустить. |
| `fragment`, `fm`, `fragment=...` | TLS-фрагментация (Xray-специфика). **Не часть XHTTP.** Маппить в свою TLS-fragment-фичу, если есть; иначе опустить. |
| `flow` | Для XHTTP всегда пустой (vision несовместим). |
## 6. Известные ограничения клиента (что НЕ маппить)
-`scMaxConcurrentPosts` — legacy-поле (удалено из текущего Xray-core и sing-box-extended; там upload сериализован в 1 POST-тело за раз). Наш клиент тоже шлёт последовательно = текущий Xray. Поле принимается (`sc_max_concurrent_posts`), но игнорируется.
-`downloadSettings` (асимметричный download-транспорт) — не поддержан; `mode=auto`+reality+downloadSettings
у нас всё равно даст stream-one, не stream-up.
-`spx` (spiderX), Xray browser-dialer — нет аналога.
- **HTTP/3 (`alpn=h3` / QUIC).** Наш XHTTP-клиент работает поверх **HTTP/2** (`http2.Transport`). Xray
умеет H1/H2/H3. Ноды, помеченные `alpn=h3`, мы обслуживаем по H2 (если сервер допускает); если сервер
**требует строго h3** — коннект не встанет. Это архитектурное ограничение транспорта, вне SPEC 002
(отдельная будущая задача «XHTTP over HTTP/3»). Парсеру: `alpn` маппить как есть, но `h3`-only ноды
помечать как потенциально неработающие.
-`fragment` / `fm` (TLS-фрагментация Xray) — не часть XHTTP; маппить в свою TLS-fragment-фичу (если есть)
или опускать.
---
## 7. Чек-лист для интегратора
- [ ]`type=xhttp` распознаётся как XHTTP-транспорт.
- [ ]`extra` декодируется как URL-encoded JSON и вливается в transport.
- [ ] Числовые `sc*`-поля из `extra` → строка `"min-max"`.
Хронология вендоренного wireguard-go: базы графта, миграции, что менял upstream. Актуальное состояние — в [SPEC.md](SPEC.md); здесь только «как было раньше и почему переделали».
---
## Почему граф, а не прямой `replace` на amneziawg-go
Первая идея — подключить `amnezia-vpn/amneziawg-go` напрямую через `replace`. **Не работает:** amneziawg-go основан на *upstream* wireguard-go и не имеет sagernet-добавок (`Send(offset)`, `InputPacket`, `conn` reserved/control), на которых держится `transport/wireguard` sing-box. Прямой replace ломает сборку.
Решение — **3-way graft**: обфускация Amnezia накладывается поверх `sagernet/wireguard-go` (а не наоборот). Так контракт sing-box↔device остаётся sagernet'овским, обфускация аддитивна. Форк-модуль — `Leadaxe/wireguard-go-awg2-lx`.
## База графта: эволюция
| Дата | Submodule commit | Sagernet-база | wireguard-go версия | Контекст |
-`9de6dc3 Add batched InputPackets` + `2c27bbf FIx batched InputPackets` — новый батч-вход `InputPackets([]*InputPacketRef) []*InputPacketRef` (возвращает unmatched refs — для L3-forward, где нет пира → вызывающий строит ICMP-unreachable). **`InputPacket` (singular) НЕ удалён** — переписан на size-based буфер + backpressure-кап `maxQueuedInputPackets`.
-`8403cdb Rework outbound buffer management` — **`QueueOutboundElement.buffer` сменил тип `*[MaxMessageSize]byte` → `[]byte`** (size-based пул через `GetOutboundBuffer(n)`/`PutOutboundBuffer` из sing-аллокатора, вместо фиксированного `messageBuffers`-пула). Элемент-пулы `outboundElements*` перешли с`WaitPool` на `sync.Pool`. Добавлен `peer.queuedOutboundPackets atomic.Int32` (backpressure-счётчик).
-`57baac9 Add batched UDP I/O on Darwin` + `fcbb7c4 Coalesce UDP GSO segments` — новый `conn/msgx_darwin.go` (sendmsg_x/recvmsg_x), GSO-iovec coalescing в `bind_std.go`.
**Оценка риска для графа ДО работы** (по памяти) была завышена: «`buffer`-type change ломает все AWG-хуки в send.go — основная работа». **По факту оказалось иначе:**
**Итог re-graft (`git apply --3way` граф-diff'а на v0.0.5):**
- **15 из 16** граф-файлов легли **чисто**. Конфликт — **только `send.go`**, и **на одной строке**: upstream добавил `peer.queuedOutboundPackets.Add(-…)` там, где граф добавил пустую строку. Взяли upstream (backpressure нужен).
- **Почему `buffer`-type change НЕ сломал граф:** AWG-хуки уже везде работают с `elem.buffer` как со **срезом** (`buffer[:MessageTransportHeaderSize]`, сдвиг `buffer[i+padding]`), а не как с массивом-указателем. Переход `*[N]byte → []byte` для них прозрачен.
- **Почему upstream `InputPacket`/`InputPackets` встали verbatim:** граф `send.go` их **не трогает** (junk-логика графа — в `SendHandshakeInitiation`, а не в input-пути), поэтому конфликта не было — upstream-версии сохранились.
- **Почему `RoutineEncryption` сшилась без ручного weave:** при `MessageEncapsulatingTransportSize = 0` upstream-offset `buffer[METS:METS+HeaderSize]` схлопывается к графовому `buffer[:HeaderSize]`. Граф-версия (заголовок в начале, без финального encapsulating re-slice) наложилась как есть.
**Вывод:** несущий инвариант `MessageEncapsulatingTransportSize = 0` — то, что делает re-graft дешёвым: он нейтрализует единственную точку, где upstream и граф расходятся по layout буфера.
Сборка после re-graft: device/conn/tun на linux/android/windows/darwin ✅; полный sing-box CLI с LX_TAGS (Go 1.24.7) ✅; тесты `transport/wireguard` + `protocol/wireguard` зелёные ✅.
## MTU / EMSGSIZE — находка 2026-06-10
При лайв-тесте AWG2-узла рукопожатие проходило, но трафик не шёл: `sendmsg: message too long` (**EMSGSIZE**). Причина — `S3`/`S4`: junk дописывается к **каждому** transport-сообщению, и обфусцированный data-пакет перерастает path MTU (1500, DF). Handshake маленький — проходит; transport — нет. Plain WG к тому же серверу с `mtu 1420` работает (S-junk нет).
Эмпирика (тот же узел, менялся только `mtu`, `S3=S4=60`):
| mtu | результат |
|----:|-----------|
| 1420 | ❌ EMSGSIZE |
| 1380 | ✅ ~58 ms |
| 1280 | ✅ ~55 ms |
| 1200 | ✅ ~60 ms |
Результат — MTU-политика в текущем SPEC.md (auto-default 1280 + warn при превышении бюджета). Источник находки — заметка агента лаунчера (`singbox-launcher`). Это не баг ядра, а размерный оверхед S-junk.
## Безопасность
Секреты живого AWG-сервера **никогда** не попадали в репозитории — лайв-конфиг держался только в `/tmp` и затирался (`shred`). Репо-конфиг `lx-test/config/awg2_basic.json` — с фейк-ключами.
-`transport/wireguard/device_awg.go` (`//go:build with_awg`) — `awgIpcLines()` шлёт IpcSet-ключи `jc=/jmin=/jmax=/s1..s4=/h1..h4=/i1..i5=` в device; `device_stub_awg.go` без тега даёт явную ошибку при заданных AWG-полях.
**2. amneziawg-go активирован через merged-форк (главное достижение):**
- amneziawg-go основан на *upstream* wireguard-go и не имеет sagernet-добавок (`Send(offset)`, `InputPacket`, `conn` reserved/control), на которых держится `transport/wireguard`. Прямой `replace` ломает сборку.
- Ключевое упрощение: **`MessageEncapsulatingTransportSize = 0`** — нейтрализует 8-байтный headroom sagernet (sing-box-lx его не использует), и обфускация Amnezia встаёт чисто без weave-конфликтов в send-пути.
-`conn/tun/ipc` оставлены **чисто sagernet** (обфускация только в `device/`: новые `obf*.go`+`magic-header.go` + графты в `send/receive/device/uapi`).
## MTU при ненулевых S3/S4 (EMSGSIZE) — дополнение 2026-06-10
При лайв-тесте AWG2-узла рукопожатие проходило, но трафик не шёл: ядро спамило `failed to send data packets: … sendmsg: message too long` (**EMSGSIZE**). Причина — прямое следствие `s3`/`s4`: junk дописывается к **каждому transport-сообщению**, и обфусцированный data-пакет перерастает path MTU физического интерфейса (1500, DF). Handshake маленький — проходит; transport — нет. Plain WG к тому же серверу с `mtu 1420` работает (S-junk нет).
Бюджет: `mtu ≤ 1500 − 28 (UDP/IP) − 32 (WireGuard) − max(S3, S4)`. Для `S3=S4=60` → `mtu ≤ 1380`; рекомендуемый клиентский MTU AmneziaWG — **1280** (запас на PPPoE/вложенные туннели). Эмпирика (тот же узел/сервер, менялся только `mtu`):
- **auto-default**: при незаданном `mtu` на AWG-эндпоинте ставим рекомендованный **1280** вместо upstream-дефолта `1408` (который сам бы превышал бюджет и триггерил наш же warn).
- **warn**: при явно заданном `mtu` выше бюджета — предупреждение (handshake пройдёт, данные — нет). Path MTU зашит консервативно **1492** (PPPoE): `mtu ≤ 1492 − 28 − 32 − max(s3,s4)` → для `s3=s4=60` это `1372`. Эмпирический потолок выше (1380), т.к. тест шёл по реальному 1500-Ethernet; 1492 — запас под узкие пути.
- Проверено (`check`): AWG `s3=s4=60` без `mtu` → тихо (default 1280); `mtu=1420` → `WARN … consider mtu <= 1372`; plain WG без `mtu` → тихо (1408).
Подтверждение (amneziawg-go docs): рекомендуемый клиентский MTU 1280; если `Jmax` ≥ системного MTU — junk-пакет фрагментируется и теряется на узких путях. Это не баг ядра, а размерный оверхед S-junk. Источник находки — заметка агента лаунчера (`singbox-launcher`, 2026-06-10).
## Безопасность
Секреты сервера **никогда** не попадали в репозитории — лайв-конфиг держался только в `/tmp` и затёрт (`shred`). Репо `lx-test/config/awg2_basic.json` — с фейк-ключами.
## Зона касания upstream (ребейз)
sing-box-lx: `go.mod` (replace), `option/wireguard*`, `protocol/wireguard/endpoint.go`, `transport/wireguard/*` — всё `// lx`. Форк wireguard-go ребейзится отдельно на новый тег sagernet (повтор 3-way merge амнезии).
## Остаточное / дальше
- reserved-feature не применяет reserved-байты в obfuscated send (для plain-AWG не нужно — карта пуста).
- Можно перевести `replace` с submodule на pinned-pseudoversion (submodule достаточно).
- Лаунчер: AWG-поля (S1–S4, I1–I5) в визард + парсер `.conf`/awg-quick; рассматривает кламп MTU для AWG-узлов (первичная истина про оверхед `s3`/`s4` — здесь, см. раздел MTU).
AmneziaWG = WireGuard-девайс с расширенным конфигом. В sing-box девайс создаётся в `transport/wireguard` поверх `github.com/sagernet/wireguard-go`. Стратегия: **подменить модуль на `amneziawg-go`** (API-совместим с wireguard-go) и **под `with_awg`** прокидывать AWG-поля в строку конфигурации девайса; endpoint остаётся типом `wireguard`.
> Проверить: совпадает ли публичный API amneziawg-go (пакеты `device`, `conn`, `tun`) с тем, что импортирует `transport/wireguard`. Если расходится — минимальные `patches/` или адаптерный слой в новом файле.
| `lx-test/config/awg2_*.json` | **new** | Конфиги для `sing-box check` |
## 4. Зона касания upstream (для ребейза)
`go.mod`/`go.sum`, файл опций wireguard-endpoint, `protocol/wireguard/endpoint.go`, `transport/wireguard/*` (минимально). Девайс-логика и опции AWG — в **новых** файлах под тегом → основной конфликт только в `go.mod` и одной ветке endpoint.
## 5. Порядок работ
1. Submodule + `go.mod` replace; собрать обычный WG (без `with_awg`) — поведение upstream.
2. Сверить API amneziawg-go vs `transport/wireguard`; при необходимости `patches/`.
4.`device_awg.go` (формат `jc=/h1=/i1=…`) под `with_awg`; stub без тега.
5. Прокидка в endpoint; конфиги; `check`; ручной коннект к AWG2-серверу.
## 6. Риски
- **API-дрейф** amneziawg-go относительно версии wireguard-go, на которую завязан upstream (`v0.0.2-beta.1.0.20260224…`). Возможен лаг — фиксировать совместимый коммит сабмодуля, не «latest».
- **Регистр I1–I5** (uppercase) — silent ignore при ошибке; валидировать.
- Взаимодействие junk/CPS с `persistent_keepalive` и MTU — проверять на реальном сервере.
- Доменный `server` + FakeIP: может потребоваться override резолва (референс hoaxisr) — добавлять только при подтверждённой необходимости.
- **Junk** (`Jc`/`Jmin`/`Jmax`) — `Jc` случайных пакетов размером `rand(Jmin..Jmax)` перед handshake initiation.
- **Магические заголовки** (`H1–H4`) — подменяют 4-байтный тип сообщения (init/response/cookie/transport); в AWG 2.0 — диапазоны `"N-M"`, из которых значение генерируется на лету.
- **Размерный padding** (`S1/S2` — на handshake, `S3/S4` — на **каждый** transport-пакет).
- **CPS-пакеты** (`I1–I5`) — снимки реального протокола (напр. QUIC Initial, STUN), которые уходят вперемешку с handshake, имитируя посторонний трафик. `I1` — центральный (см. [SPEC 009](../009-WIRESOCK_MASQUERADE_PROFILES/SPEC.md) — декларативные masquerade-профили `ip=quic/sip/dns`, которые генерируют `I1`).
Upstream sing-box AWG не принимает ([#4045](https://github.com/SagerNet/sing-box/issues/4045), closed not-planned) — реализовано в форке.
## Архитектура (два слоя)
Обфускация живёт **в вендоренном wireguard-go** (submodule), а sing-box только пробрасывает параметры. Это ключевое разделение: контракт `transport/wireguard` ↔ device остаётся sagernet'овским, обфускация — аддитивна.
- **6 modified**: `device.go` (AWG-state: `junk`, `headers`, `paddings`, `ipackets [5]*obfChain`), `send.go` (junk + CPS + padding в handshake/transport-путях), `receive.go` (детект magic-header на входе), `cookie.go`/`noise-protocol.go`/`uapi.go` (типы сообщений через генератор, парсинг AWG-ключей в IpcSet).
**Ключевой инвариант — `MessageEncapsulatingTransportSize = 0`** ([device/noise-protocol.go](../../submodules/wireguard-go/device/noise-protocol.go)). Upstream держит 8-байтный headroom перед transport-заголовком (для `conn.Bind.Send()`-префикса). Граф его **обнуляет**: AWG-обфускация формирует префикс сама (junk/CPS уходят отдельными буферами через `SendBuffers`, а не через encapsulating-space). При `= 0` upstream-выражения вида `buffer[MessageEncapsulatingTransportSize+MessageTransportHeaderSize:]` схлопываются к графовому виду `buffer[MessageTransportHeaderSize:]` — поэтому большинство upstream-функций компонуются с графом **без ручного weave**. Это несущий инвариант re-graft (§ ниже).
**Что граф НЕ трогает:**`conn/`, `tun/` — чисто sagernet (берутся из upstream verbatim). Обфускация замкнута в `device/`.
### Слой 2 — sing-box (проброс параметров, всё `// lx`)
- **`option/wireguard_awg.go`** — `AmneziaWGOptions`: `Jc/Jmin/Jmax`, `S1–S4`, `H1–H4` (тип `MagicHeader` — строка `"N"` или диапазон `"N-M"`, JSON-совместим с прежним uint32), `I1–I5` (string, регистр сохраняется). Promoted-встроены в `WireGuardEndpointOptions`.
- **`transport/wireguard/device_awg.go`** (`//go:build with_awg`) — `awgIpcLines()` рендерит IpcSet-ключи `jc=/jmin=/jmax=/s1..s4=/h1..h4=/i1..i5=`, дописываемые к WireGuard-конфигу устройства. `device_stub_awg.go` (`//go:build !with_awg`) даёт явную ошибку при заданных AWG-полях.
- **`transport/wireguard/endpoint.go`** — MTU-политика для AWG (см. ниже).
- **`validateJunk`** — отвергает `jmin > jmax` до старта: `amneziawg-go` считает `rand(0..jmax-jmin)+jmin`, и `jmax < jmin` даёт `rand.Int`с аргументом `≤ 0` → **паника ядра**. Гардим только этот crash-кейс.
Регистрация endpoint остаётся `C.TypeWireGuard` (AWG = WG + доп. поля, отдельный тип не вводим).
## MTU-политика (следствие S3/S4)
`S3`/`S4` дописывают junk к **каждому** transport-сообщению → обфусцированный data-пакет перерастает path MTU физического интерфейса (1500, DF) → ядро спамит `sendmsg: message too long` (**EMSGSIZE**), handshake проходит, а трафик — нет.
Логика в [transport/wireguard/endpoint.go](../../transport/wireguard/endpoint.go) (gated `max(s3,s4) > 0`, plain WG нетронут):
- **auto-default**: при незаданном `mtu` на AWG-эндпоинте — рекомендованный **1280** вместо upstream-дефолта 1408.
- **warn**: при явном `mtu` выше бюджета — предупреждение (`pathMTU = 1492`, консервативно под PPPoE). Для `s3=s4=60` → `mtu ≤ 1372`.
Держать `Jmax` ниже системного MTU (иначе junk-пакет фрагментируется и теряется на узких путях). Подробности: `docs-lx/lx-config.md` §2 (MTU).
## Процедура re-graft (при бампе upstream wireguard-go)
Когда upstream `sagernet/wireguard-go` двигает версию, граф переносится на новую базу. **Не merge, а controlled 3-way apply** граф-diff'а:
1.**База**: submodule → новый sagernet-коммит.
2.**Apply graft**: `git diff <старая-база> <старый-graft> | git apply --3way`. По практике 15/16 файлов ложатся чисто; конфликтует обычно только `send.go` (плотный upstream-путь).
3.**Разрешить конфликты вручную**, порядок по риску: `cookie`→`device`→`noise-protocol`→`uapi`→`receive`→**`send.go`** (высший — junk/padding-хуки в hot-path).
4.**Сверить несущие инварианты**: `MessageEncapsulatingTransportSize = 0`; графовый `RoutineEncryption` (заголовок в начале буфера, без финального encapsulating re-slice); AWG-state поля в `device.go`.
5.**Проверки**: сборка `device/conn/tun` на linux/android/windows/**darwin** (darwin особо — там upstream добавляет платформенный batch-send), затем полный `sing-box`с LX_TAGS, `go test ./transport/wireguard/ ./protocol/wireguard/`, **device-verify** живого AWG-туннеля (junk/handshake/трафик).
История конкретных re-graft'ов (какие базы, что менял upstream) — в [HISTORY.md](HISTORY.md).
## Критерии готовности
-`sing-box check -c` принимает wireguard-endpoint c `jc/h1/i1…` под `with_awg`.
- Реальный коннект к AmneziaWG 2.0 (device-verify): `sending handshake initiation` → `received handshake response` → keepalive → трафик через сервер, с непустыми `Jc` и хотя бы одним `I1`.
- Сборка **без**`with_awg`: обычный WG как upstream; AWG-поля → явная ошибка.
- **sing-box-lx**: `go.mod` (replace + pin), `option/wireguard_awg.go` + `// lx`-поля в основной struct, `transport/wireguard/device_awg*.go`, MTU-блок в `transport/wireguard/endpoint.go`, проброс в `protocol/wireguard/endpoint.go` — всё `// lx`.
- **submodule wireguard-go**: ребейзится отдельно (см. процедуру re-graft), не входит в merge-зону основного репо кроме pin в `go.mod`.
- [SPEC 020](../020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md) — idle-suspend WG/AWG-устройств (Down/Up); опирается на стабильный device-API той же вендоренной базы.
- on tag `v*-lx.*` → `build` (6 desktop, tar.gz/zip) + `build_android` (2 AAR) → `release`: `SHA256SUMS` + GitHub Release с notes (база `v1.13.13` + фичи + `lx-print-tags` + строка про AAR). Версия из тега, `sing-box version` → `-lx.N`.
- **`v1.13.13-lx.3` опубликован** (Latest): 6 архивов + `libbox-1.13.13-lx.3.aar` + `libbox-legacy-1.13.13-lx.3.aar` + `SHA256SUMS` — всё зелёное. Этот прогон впервые вживую подтвердил тяжёлый путь (cross ×6 с naive/cronet/purego + gomobile AAR + publish).
- **Windows 7 (32-bit)** legacy-таргет: `windows/386` собирается **пропатченным Go** (`.github/setup_go_for_windows7.sh` — реверты удаления Win7 из `MetaCubeX/go`, как в upstream `build.yml`) и **без `with_naive_outbound`** (`cronet-go` не имеет windows/386 — build constraints исключают всё). Артефакт `sing-box-<ver>-windows-386-legacy-windows-7.zip` — под лаунчер-сборку `singbox-launcher-win7-32` (она тоже 386). Остальной `LX_TAGS` (gvisor/quic/xhttp/awg/…) под 386 компилируется — проверено.
- **Никогда не force-push'ит `lx`** — только новая ветка + PR/issue на ревью.
- Демо (`workflow_dispatch tag=v1.13.13`): `Pick target` → `Up to date?` → success, остальное skipped, **0 side-effects** (ни веток, ни PR, ни issue).
## Операционные настройки репозитория (критично для CI)
`gh api repos/OWNER/REPO/actions/permissions/workflow`:
- **`default_workflow_permissions: write`** — иначе `gh release create` падает с`403 Resource not accessible by integration` (это и был корень падений первых релизных прогонов lx.2). NB: «релиз для тега уже существует» — **другая** ошибка (`already exists`), не 403.
- **`can_approve_pull_request_reviews: true`** («Allow GitHub Actions to create and approve pull requests») — иначе авто-PR ребейза ботом блокируется (есть fallback в issue).
- Оба включены 2026-06-09.
## Зона касания upstream (ребейз)
Все lx-артефакты — **новые файлы**: `.github/workflows/lx-{ci,release,rebase}.yml`, `Makefile.lx`, `lx-test/config/`. Единственная правка upstream-файла — `// lx`-блок в `cmd/internal/build_libbox/main.go` (теги AAR). При ребейзе новые файлы переносятся как есть, блок в `build_libbox` — вручную по маркеру.
## Остаточное / дальше
- Старый релиз `v1.13.13-lx.1` можно удалить (предшествует XHTTP-фиксу / полным тегам / libbox; `lx.3` его замещает).
- Лаунчер (репо `singbox-launcher`, отдельно): маппинг `type=xhttp` → реальный xhttp (его задача 023 сейчас в httpupgrade); AWG-поля (Jc/S1–S4/H1–H4/I1–I5) в визард + парсер `awg.conf`; замена бандлового `bin/sing-box` на lx-релиз.
- (опц.) XHTTP `stream-one` framing-баг (`auto`/`packet-up` работают, не блокер).
| `.github/workflows/lx-release.yml` | расширение | on tag `v*-lx.*`: cross-build desktop (через `Makefile.lx`, без дублирования тегов) + job **`build_android`** (AAR) → zip/checksums/GitHub Release |
| `lx-test/config/xhttp_reality.json`, `awg2_basic.json` | из 002/003 | Используются в CI `check` |
## 2. Версия / ldflags
`LX_LDFLAGS = -X github.com/sagernet/sing-box/constant.Version=<upstream>-lx.<N> -checklinkname=0 -s -w -buildid=`. `<N>` — счётчик lx-релизов поверх upstream-тега. **`-checklinkname=0` обязателен** для полного набора тегов (`badlinkname` → `go:linkname` в `crypto/tls` через `common/badtls`; Go 1.24 блокирует без флага). AAR версионируется отдельно — `build_libbox` берёт `git describe`, поэтому в обоих workflow перед сборкой AAR создаётся/обновляется тег.
4.`lx-rebase.yml` (сначала `workflow_dispatch`, потом cron).
5. Демо-прогон ребейз-workflow на текущем теге.
## 5. Риски
- **`-checklinkname=0`** (РЕШЕНО): без него полный набор не линкуется (`badtls`/`crypto/tls`). Локально подтверждено для linux/amd64 и windows/arm64; остальные 4 таргета верифицирует CI-матрица.
- **`with_naive_outbound` через cronet** тянет prebuilt `cronet-go/lib/<os>_<arch>` — если под какой-то таргет prebuilt отсутствует, naive там не соберётся → дропнуть naive на этой платформе (или из набора целиком). Проверяет CI-матрица.
- **AAR-сборка**: требует NDK r28 + OpenJDK 17 + gomobile (`make lib_install`); `build_libbox.checkJavaVersion()` ждёт строго `openjdk 17`. Версия AAR = `git describe`, поэтому тег должен существовать в чекауте.
- Авто-ребейз на **alpha/beta** теги нежелателен — фильтровать только стабильные (`vX.Y.Z` без суффиксов).
-`git submodule` в CI — не забыть `--init --recursive` и pin (нужно и для `with_awg`, и для AAR).
Собрать воспроизводимый конвейер сборки/CI/релизов `sing-box-lx`: кросс-платформенные бинари `sing-box`**и Android `libbox.aar`** с клиентским feature-set (полный upstream минус серверные/AI-теги) + lx-фичами (`with_xhttp`/`with_awg`), версия `-lx.N`, и **авто-ребейз на новый upstream-тег**.
> **Процедура выпуска → [docs-lx/lx-release-runbook.md](../../docs-lx/lx-release-runbook.md).**
> Главное правило: **перед любым тегом проверить дрейф upstream и обычно смержить его себе, и только
> потом резать релиз/пререлиз.** На ветке `lx-1.14` авто-ребейз на стабильный тег (ниже) заменён
> ручным `git merge upstream/testing` — пока upstream на `v1.14.*-alpha`, стабильного тега нет, а
> rc-линия `vX-lx.1-rc.N` сама является форматом поставки.
---
## 1. Проблема / контекст
Реальная стоимость downstream'а — не первичная разработка, а N ребейзов в год и регулярные сборки на 3 платформы. Нужен конвейер, который ловит «фичи поломались об новый upstream» раньше пользователя и выпускает drop-in бинарь для лаунчера.
## 2. Требования
### 2.1 Сборка
- **Desktop-бинарь `sing-box`** (drop-in для лаунчера): цель `make -f Makefile.lx lx-build`, output `sing-box`.
- **Набор `LX_TAGS`** — upstream feature-set (`release/DEFAULT_BUILD_TAGS`) **минус нерелевантные клиенту**: `with_tailscale` (нет tailscale-endpoint'ов), `with_ccm`/`with_ocm` (прокси Claude Code / OpenAI Codex — серверные AI-сервисы), `with_acme` (серверный выпуск TLS-сертов). Итог = `gvisor/quic/dhcp/wireguard/utls/clash_api/naive_outbound + badlinkname/tfogo_checklinkname0`**+ `with_purego`** (CGO-free кросс-сборка `with_naive_outbound` через prebuilt cronet) **+ `with_xhttp,with_awg`**. `Makefile.lx` — единственный источник истины (`make -f Makefile.lx lx-print-tags`).
- **`LX_LDFLAGS` обязан содержать `-checklinkname=0`** — иначе `badlinkname`/`tfogo_checklinkname0` ломают линк (`common/badtls` использует `go:linkname` в `crypto/tls`, который Go 1.24 блокирует). Зеркалит upstream `build_libbox`.
- **Android `libbox.aar`**: `make lib_install && make lib_android` (gomobile, NDK r28 + OpenJDK 17). `with_xhttp`/`with_awg` зашиты в `cmd/internal/build_libbox` (lx:-блок) → попадают в `libbox.aar` (SDK 23) и `libbox-legacy.aar` (SDK 21). Набор тегов AAR = upstream mobile-set **минус `with_tailscale`** (как desktop — самая тяжёлая либа в APK; правка обёрнута `// lx:no-tailscale`) + наши две фичи (NDK/CGO-сборка, `with_purego` не нужен).
- Версия `vX.Y.Z-lx.N` через ldflags (из 001); для AAR — через `git describe` внутри `build_libbox`.
### 2.2 CI-матрица
- **Политика триггеров (стоимость per-commit ↓).** Doc-only коммиты (`**.md`/`docs/**`/`SPECS/**`/LICENSE) **не запускают CI** (`paths-ignore`). На каждый push/PR — **только дешёвые** job'ы `lint` + `build-check`. Тяжёлые `cross` (6 таргетов) и `android` (gomobile AAR) — **только вручную, на `workflow_dispatch`** (`gh workflow run lx-ci.yml --ref lx` или кнопка Actions → Run workflow); на push их нет. Полную кросс-сборку + обе AAR на каждый релиз-тег и так гарантирует `lx-release.yml`. Серия быстрых пушей отменяет устаревшие прогоны (`concurrency: cancel-in-progress`).
- **`lint`** (push/PR): `go vet` по lx-пакетам с полными тегами + `gofmt` только по lx-файлам (`v2rayxhttp|_xhttp|_awg`, не по всему дереву upstream).
- **`build-check`** (push/PR): один нативный build `with_xhttp,with_awg` + `sing-box check` XHTTP/AWG2-конфигов (должны пройти); затем tagless baseline-бинарь → `check minimal.json` (проходит) + negative-check (XHTTP/AWG2-конфиги без тегов отвергаются).
- **`cross`** (dispatch): `{linux, darwin, windows} × {amd64, arm64}`, full `LX_TAGS`, CGO=0 — проверка, что полный набор + `with_purego` кросс-собирается везде.
- **`android`** (dispatch): `make lib_android` (NDK r28 + JDK17 + gomobile) — libbox AAR собирается с lx-фичами.
- Все job'ы — с submodule (`submodules: recursive`) и `fetch-depth: 0` (для `-lx` версии через `git describe`).
3. Успех + сборка/`check` зелёные → пуш ветки `lx-rebase/<tag>` и **PR**; конфликт → **issue**с диффом `// lx:` зон.
- Никогда не пушить силой в `lx` автоматически — только через PR с ревью.
### 2.4 Релизы
- Тег `vX.Y.Z-lx.N` → артефакты: desktop-архивы (`sing-box`× 6 платформ) **+ `libbox-<ver>.aar` и `libbox-legacy-<ver>.aar`**, общий `SHA256SUMS`.
- Release notes: upstream-база + состояние фич (`with_xhttp`/`with_awg`) + полный `LX_TAGS` desktop-бинаря (через `lx-print-tags`) + строка про AAR.
## 3. Критерии приёмки
- CI зелёный: на push/PR — дешёвые `lint` + `build-check`; полная матрица (`cross`×6 с полным `LX_TAGS` + `android` AAR) — вручную на `workflow_dispatch`. Doc-only коммиты CI не триггерят; релиз-тег собирает всё через `lx-release.yml`.
- Артефакты собираются, бинарь называется `sing-box`, `version` → `-lx.N`.
- **libbox AAR собирается в CI (job `android`) и публикуется в Release**; `Libbox.version()` → `-lx.N`; конфиг с AWG2/XHTTP не падает с «support not built».
- Авто-ребейз workflow отрабатывает на `workflow_dispatch` (демо на текущем теге → «уже актуально» или PR).
## 4. Вне скоупа
- Подпись кода/нотаризация; публикация AAR в Maven/jitpack (отдаём только GitHub Release asset).
- Интеграция AAR в приложение-потребитель (LxBox) — задача на стороне приложения.
- Полностью автоматический мёрж ребейза (всегда ревью).
## 5. Ссылки
- [Build from source — sing-box](https://sing-box.sagernet.org/installation/build-from-source/)
- [x]**`workflow_dispatch` — тяжёлое (вручную):** `cross``{linux,darwin,windows}×{amd64,arm64}` на полном `LX_TAGS` + `android` (`make lib_android`); на push не запускаются
- [x] submodule init (`submodules: recursive`) + `fetch-depth: 0` во всех job'ах
- [x] Зелёный прогон `lint`+`build-check` на push (подтверждено); `cross`×6 + `android` AAR — зелёные в релизном прогоне v1.13.13-lx.3 (cronet/naive/purego на всех 6)
## Авто-ребейз
- [x]`lx-rebase.yml`: fetch upstream tags → выбрать новейший **стабильный** (`^v[0-9]+\.[0-9]+\.[0-9]+$`) → rebase в CI
`H1`–`H4` в конфиге wireguard-endpoint теперь принимают **диапазон** AWG 2.0 (`"43613244-384550127"`) наряду с одиночным числом (`1234567890`, обратная совместимость). Spec-строка доезжает до IpcSet, vendored `submodules/wireguard-go` (который уже умел диапазоны) поднимает обфусцированный handshake. Реальный awg2-экспорт (`seliv_for_awg2.conf`) импортируется без правок.
Изменения — **только в lx-собственных файлах** (`option/wireguard_awg.go`, `transport/wireguard/device_awg.go`, оба созданы в 003). Ноль новых касаний upstream, vendored wireguard-go не тронут.
- База `string` ⇒ `AmneziaWGOptions` остаётся comparable, `IsSet()` (`o != AmneziaWGOptions{}`) работает.
-`Spec()` повторно валидирует и отдаёт канон для IpcSet — ловит опции, собранные в коде (libbox/лаунчер) мимо JSON.
- Валидация (uint32, start ≤ end) и в `UnmarshalJSON`, и в `Spec()`; имя поля в ошибке даёт contextjson (`h1: invalid magic header …`) на парсе и `E.Cause(err, "h1")` в `awgIpcLines` на сборке device.
**Эмит (`transport/wireguard/device_awg.go`):**
-`h1..h4` через `writeStr`-путь (spec-строка), unset → не эмитится.
- Plain WG (без AWG-полей) по-прежнему даёт `""` → byte-identical конфиг.
- ✅ baseline (без `with_awg`) — `awg2_ranged.json` отклонён явной ошибкой «awg support not built».
- ✅ Правки только в lx-own файлах; новые `// lx:`-зоны не добавлялись.
## Приёмка (живой сервер)
Лайв-тест против awg2-сервера из `seliv_for_awg2.conf` (ranged H1–H4):
```
peer - sending handshake initiation
peer - received handshake response ← uapi принял ranged-строки, обфусцированный handshake прошёл
```
-`curl --socks5-hostname 127.0.0.1:21080 https://api.ipify.org` → `{"ip":"64.188.69.128"}` (выходной IP = endpoint сервера): трафик идёт сквозь туннель.
- 0 ошибок `message too long` / EMSGSIZE (MTU не задан → дефолт `1280` под s3/s4 из 003 / `f806e24f`).
- IpcError при выставлении h1..h4-строк — **нет**.
Секреты в репозиторий не попадали: лайв-конфиг и лог держались в `/tmp`, затёрты после теста. `lx-test/config/awg2_ranged.json` — с фейк-ключами.
## Зона касания при следующем ребейзе
Без изменений относительно 003: `option/wireguard_awg.go` и `transport/wireguard/device_awg.go` — **lx-собственные** файлы, в upstream их нет, конфликтов на ребейзе не дают. Vendored `submodules/wireguard-go` не тронут.
## Вне скоупа
- Перепин `app/android/libbox.version` в лаунчере — задача на стороне LxBox.
- AWG inbound/server — отдельная будущая задача (как в 003).
Диапазон уже понимает нижний слой (`submodules/wireguard-go`: `device/magic-header.go` + `device/uapi.go` case `"h1".."h4"` → `newMagicHeader("N"/"N-M")`). Задача — донести spec-строку от JSON-конфига до IpcSet. Меняются **только lx-собственные файлы** (`option/wireguard_awg.go`, `transport/wireguard/device_awg.go`) — ноль новых касаний upstream, ребейз-стоимость не растёт.
Тип `option.MagicHeader` — `string` с канонизацией при парсе:
-`MarshalJSON`: одиночное значение → JSON number (type-fidelity со старым `uint32`), диапазон → JSON string.
- База string ⇒ тип comparable, `IsSet()` (`o != AmneziaWGOptions{}`) работает; zero value `""` = unset.
- Повторная валидация в `awgIpcLines` с именем ключа в ошибке (`E.Cause(err, "h1")`) — покрывает и программно собранные опции (libbox/лаунчер мимо JSON).
| `SPECS/README.md` | docs | строка 005 в roadmap |
## 3. Зона касания upstream (для ребейза)
**Ничего нового.** Оба изменяемых Go-файла — lx-собственные (созданы в 003), upstream-файлы не трогаются, `// lx:`-зоны не расширяются. Vendored `submodules/wireguard-go` не меняется.
## 4. Порядок работ
1.`option`: тип `MagicHeader` + замена полей + тесты.
3. Фикстура `awg2_ranged.json`; `go build` (с тегами и без), `go vet`, `go test`, `sing-box check` обеих AWG-фикстур.
4. Docs + roadmap.
5. Лайв/handshake-приёмка (конфиг с реальными ключами — только во временных файлах, как в 003).
6. REPORT, статус C, тег `v1.13.13-lx.6`.
## 5. Риски
- **Совместимость marshal**: код, который сериализует опции обратно в JSON (`sing-box format`, экспорт из лаунчера), должен получить number для одиночных значений — закрыто MarshalJSON-логикой + тестом round-trip.
- **`"h1": 0` / `"0"`**: прежняя семантика «0 = не задано» (omitempty по zero value uint32) сохраняется канонизацией в `""`.
- **Ошибка без имени поля** при парсе JSON — закрыто дублирующей валидацией в `awgIpcLines` с ключом и тестом текста ошибки.
Поддержать **диапазонные magic headers** AmneziaWG 2.0 (`H1`–`H4` вида `N-M`) в конфиге sing-box-lx. Только прослойка option → IpcSet; протокольный слой (vendored `Leadaxe/wireguard-go`) уже умеет диапазоны.
---
## 1. Проблема / контекст
- Лаунчер (LxBox) начал импортировать реальные awg2-экспорты. Живые конфиги содержат H-поля в новом формате AWG 2.0 — **диапазон** вместо числа:
```ini
H1 = 43613244-384550127
H2 = 826869626-2105069164
```
- Vendored merged-форк `submodules/wireguard-go` **уже умеет** диапазоны: `device/magic-header.go` — `newMagicHeader(spec)` принимает `"N"` и `"N-M"` (uint32, start ≤ end); `device/uapi.go` case `"h1".."h4"` парсит value через `newMagicHeader`.
- Но прослойка ядра не пропускает диапазон:
- `option/wireguard_awg.go`: `H1..H4 uint32` — диапазон не выразить, JSON-строка `"h1": "N-M"` не анмаршалится;
- `transport/wireguard/device_awg.go`: `writeUint("h1", o.H1)` — в IpcSet уходит только одиночное число.
## 2. Цель
Конфиг wireguard-endpoint с `"h1": "43613244-384550127"` (и одиночными `"h1": 1234567890` как раньше) парсится, проходит `sing-box check`, и диапазонная spec-строка доезжает до uapi девайса без IpcError.
## 3. Требования
### 3.1 Опции (`option/wireguard_awg.go`)
- `H1..H4` → тип «число-или-диапазон» (`MagicHeader` на базе string):
- `UnmarshalJSON`: принимает **JSON number** (обратная совместимость — существующие конфиги с `"h1": 1234567890` читаются без изменений) и **JSON string** `"N"` / `"N-M"`;
- валидация: обе части — uint32, start ≤ end; мусор → явная ошибка с именем поля;
- `MarshalJSON`: одиночное значение → number (type-fidelity как раньше), диапазон → string;
- тип comparable — `IsSet()` (`o != AmneziaWGOptions{}`) продолжает работать.
- `h1..h4` эмитятся как spec-строка (`writeStr`-путь), unset → не эмитить.
- Гарантия «plain WG даёт byte-identical конфиг» сохраняется.
### 3.3 Документация (`docs-lx/lx-config.md`)
- `h1`–`h4`: `int | "min-max"` + пример с диапазоном; обновить таблицу и маппинг awg.conf.
## 4. Критерии приёмки
- Тесты: unmarshal number / string-число / диапазон / ошибки (start > end, > uint32, мусор); ipc-строки с диапазоном (`\nh1=43613244-384550127`); существующие AWG-фикстуры не ломаются.
- `sing-box check` принимает конфиг с ranged H1–H4.
- Лайв-тест против awg2-сервера с ranged-конфигом (инфраструктура lx-test из 003), либо хотя бы handshake-проверка, что uapi принимает выставленные строки без IpcError.
- Сборка без `with_awg`: поведение upstream, AWG-поля → явная ошибка (как раньше).
- `go vet`, тесты затронутых пакетов — зелёные.
## 5. Вне скоупа
- Vendored `submodules/wireguard-go` — **не трогать**, он уже умеет диапазоны.
- Перепин `app/android/libbox.version` в лаунчере — задача на стороне LxBox.
- Диапазоны для S1–S4/Jc — в формате AWG 2.0 их нет (только H1–H4).
- [x]`transport/wireguard/device_awg.go`: `h1..h4` → spec-строка через writeStr-путь; валидация с именем ключа; unset → не эмитить
- [x]`device_awg_test.go` (`with_awg`): `\nh1=43613244-384550127`; одиночные значения как раньше; plain WG → `""`
## Проверки
- [x]`lx-test/config/awg2_ranged.json` (фейк-ключи) + `sing-box check` обеих AWG-фикстур
- [x] Сборка с тегами и без; `go vet`; `go test` затронутых пакетов
- [x] Существующие AWG-фикстуры (`awg2_basic.json`) не ломаются
## Приёмка
- [x] Лайв/handshake-тест против awg2-сервера с ranged-конфигом (uapi принял строки без IpcError, handshake + трафик прошли); секреты — только в temp-файлах, затёрты
## Документация и закрытие
- [x]`docs-lx/lx-config.md`: `h1`–`h4``int | "min-max"`, no-overlap, пример с диапазоном, маппинг awg.conf
Релизные Linux-бинари переведены на **статическую musl-сборку с сохранением NaïveProxy**, добавлены роутерные арки. Закрывает [issue #1](https://github.com/Leadaxe/sing-box-lx/issues/1): `libdl.so.2: cannot open shared object file` на AsusWRT Merlin + отсутствие `linux-armv7`.
**Go-кода нет** — задача чисто инфраструктурная (CI). Единственная Go-правка в этой ветке — hotfix gofmt-выравнивания в `option/wireguard_awg.go` (хвост 005, см. ниже).
## Диагноз (эмпирически подтверждён)
`with_naive_outbound` тянет `cronet-go`. Прежний релиз собирался в режиме `with_purego` (`CGO_ENABLED=0`): `purego` на Linux содержит `//go:cgo_import_dynamic … "libdl.so.2"` → бинарь **динамический**, требует `libdl.so.2`. На glibc ок, на musl (роутеры) — падает до старта. Кросс-сборкой проверено: `file` → `dynamically linked`, `strings|grep libdl.so.2` → 1; без naive/purego → `statically linked`, 0.
naive **сохраняем** — это upstream-фича (`release/DEFAULT_BUILD_TAGS` содержит `with_naive_outbound`; `protocol/naive/outbound.go` без `lx:`-маркеров). Поэтому не дропаем, а используем третий режим cronet-go — `with_musl` (статический `libcronet.a` + musl-toolchain), как делает upstream `build.yml`.
## Что сделано
**`.github/workflows/lx-release.yml`** — новый job `build_linux_musl` (зеркало upstream musl-pipeline):
- Linux убран из desktop-job `build` (остаются darwin/windows/win7); `release.needs += build_linux_musl`; release-notes обновлены.
**`.github/workflows/lx-ci.yml`** — dispatch-only smoke-job `linux_musl` (те же 4 арки): полный musl-pipeline + build + verify static, **без публикации** — безопасная приёмка. Помечен «keep in sync with build_linux_musl».
**Нейминг** — по upstream-схеме арочных суффиксов (`armv7` = arm+`v`+GOARM; `mipsle-softfloat` = arch+GOMIPS), но **без суффикса `-musl`**: upstream добавляет его, т.к. собирает и glibc, и musl на арку; у нас Linux — единственный (musl) вариант, и `linux-arm64`/`linux-armv7` совпадают с ожиданием скриптов потребителей.
## Приёмка
- ✅ YAML валиден (`python yaml`), `actionlint` чист (rc=0) для обоих workflow.
- ✅ CI smoke (`lx-ci` workflow_dispatch, job `linux_musl`×4, run [27407702652](https://github.com/Leadaxe/sing-box-lx/actions/runs/27407702652)) — все 4 **success**, `file` → `statically linked`, `libdl.so.2=0`:
- ✅ Боевой релиз [v1.13.13-lx.7](https://github.com/Leadaxe/sing-box-lx/releases/tag/v1.13.13-lx.7) опубликован — 4 musl-арки + desktop (darwin/win/win7) + 2 AAR + SHA256SUMS.
- ✅ **Field-verified** репортером issue #1 на AsusWRT Merlin RT-AX (`linux/arm64`): ядро устанавливается, стартует, работает; `sing-box version` → `1.13.13-lx.7`, теги включают `with_naive_outbound,with_musl`, `CGO: enabled`. Подтверждение на реальном устройстве, не только в CI.
> **Нейминг — апдейт от потребителя:** репортер подтвердил, что суффикс `-musl` для его скрипта **некритичен** (берёт архив с суффиксом или без). То есть наш выбор «без `-musl`» валиден без оглядки на чужой скрипт — это просто следствие единственного варианта на арку. `-softfloat` у mipsle при этом **обязателен** (FP-ABI, не линковка): softfloat запускается на любом MIPS-роутере, hardfloat — только на чипах с FPU.
## Побочный hotfix (хвост 005)
`option/wireguard_awg.go`: расширение `H1..H4 uint32 → MagicHeader` сменило самый длинный тип в struct, gofmt перевыровнял json-теги. `go vet` это не ловит, поэтому ушло в lx.6 с красным дешёвым CI (`lint` job). Исправлено `gofmt -w`, format-only. Урок: прогонять `gofmt -l` на lx-owned файлах перед коммитом.
## Зона касания upstream (для ребейза)
`lx-release.yml` / `lx-ci.yml` — **lx-собственные** файлы (в upstream их нет) → конфликтов на ребейзе не дают. `.github/CRONET_GO_VERSION` — upstream-файл, **только читаем**. Паттерн musl-pipeline заимствован из upstream `build.yml` как референс.
## Вне скоупа
- Экзотика (`386`/`riscv64`/`loong64`/`mips64le`) — точечно по запросу; cronet-musl под `mips64le` нет.
- DEB/RPM/Pacman/OpenWrt-пакеты — не публикуем (только `.tar.gz`).
- naive на Win7 (windows/386) — физически невозможен (нет `cronet-go/lib/windows_386`).
- Перепин `libbox.version` в лаунчере — на стороне LxBox.
Один новый job в `lx-release.yml` — `build_linux_musl` — повторяющий upstream `build.yml` musl-секцию, но с нашим `LX_TAGS`. Существующий desktop-путь не ломаем: из старого `build` job убираем строки `linux/*`, остальное (darwin/windows/win7) остаётся на `Makefile.lx`/purego.
```
build (desktop, как сейчас минус linux): darwin×2, windows×2, win7-386
build_linux_musl (NEW): linux musl-static + naive: amd64, arm64, armv7, mipsle
**Go-кода нет.**`Makefile.lx` править не обязательно (теги берём из него же; musl-логика живёт в CI, т.к. требует Chromium-toolchain env, которого в Makefile не выразить переносимо).
## 4. Зона касания upstream (для ребейза)
`lx-release.yml` — **lx-собственный** файл (создан в 004), в upstream его нет → конфликтов на ребейзе не даёт. Паттерн заимствован из upstream `build.yml`, но как референс, не как правка upstream-файла. `.github/CRONET_GO_VERSION` — upstream-файл, мы его **только читаем** (pin уже совпадает с go.mod), не меняем.
## 5. Порядок работ
1. SPEC/PLAN/TASKS (done).
2. Реализовать `build_linux_musl` + почистить linux из `build` + `release.needs` + notes.
3. Локально: валидация YAML (actionlint/python-yaml). Полную musl-сборку локально не проверить — Chromium toolchain только в CI.
5. На зелёном — verify-шаги (static/libdl) в логах; по возможности запуск armv7 под qemu-user.
6. REPORT, статус C, ответ в issue #1. Боевой релиз — тегом `v1.13.13-lx.7`.
## 6. Риски
- **Локально не верифицируемо** — отладка только через CI; закладываем несколько прогонов. Кеш toolchain критичен для скорости.
- **Время/размер**: Chromium toolchain — гигабайты; musl-бинарь крупнее (вшит libcronet, +неск. МБ). Приемлемо для релиза по тегу (не на каждый push).
- **mipsle softfloat**: проверить, что toolchain build-naive поддерживает target `linux/mipsle` + `GOMIPS=softfloat`. Если cronet-musl/mipsle не соберётся — fallback: mipsle через `DEFAULT_BUILD_TAGS_OTHERS` (без naive, `CGO_ENABLED=0`, статика) как делает upstream для арок без cronet. Зафиксировать в REPORT.
- **keyring/sysroot download** может флапать (внешняя Chromium infra) — ретраи.
Публиковать **статические musl-бинари**`sing-box` под роутерные Linux-арки, **сохраняя NaïveProxy-outbound**. Закрывает [issue #1](https://github.com/Leadaxe/sing-box-lx/issues/1): нужен `linux-armv7`, и текущий `linux-arm64` не запускается на AsusWRT Merlin (`libdl.so.2: cannot open shared object file`).
---
## 1. Проблема / контекст
- Лаунчер-аудитория ставит ядро на роутеры (AsusWRT Merlin, OpenWrt, Keenetic) — это **musl**-окружения.
- Текущие релизные `linux-amd64/arm64` собраны в режиме `with_purego` (`CGO_ENABLED=0`). `purego` через `//go:cgo_import_dynamic … "libdl.so.2"` делает бинарь **динамическим** и вешает зависимость от `libdl.so.2`. На glibc-десктопе ок, на musl — загрузчик падает до старта. Проверено эмпирически: `file` → `dynamically linked`, `strings | grep libdl.so.2` → 1 совпадение.
-`linux-armv7` в релизе **отсутствует** вовсе.
- **NaïveProxy-outbound — штатная upstream-фича** (`release/DEFAULT_BUILD_TAGS` содержит `with_naive_outbound`; `protocol/naive/outbound.go` — upstream-код без `lx:`-маркеров). По CONSTITUTION (upstream + ровно 2 фичи, из upstream ничего не выкусываем) её **нельзя** дропать ради статики.
Без `with_awg`/`with_xhttp` поведение не меняется — это чисто сборочная задача (CI), **Go-кода нет**.
## 3. Требования
### 3.1 Механизм — по подобию upstream `build.yml`
- Сборка musl-варианта повторяет upstream: clone `cronet-go` по pin `.github/CRONET_GO_VERSION` (уже совпадает с`go.mod`: `2faf34666c2c`), regenerate Debian keyring, download Chromium **musl** toolchain через `go run ./cmd/build-naive --target=linux/<arch> --libc=musl download-toolchain`, выставить env (`… env >> $GITHUB_ENV`), затем `CGO_ENABLED=1 go build`с тегом `with_musl` — `libcronet.a` линкуется статически.
- **zig не используется** — официальный путь cronet-go (Chromium toolchain) надёжнее и совпадает с upstream.
### 3.2 Теги
- Брать `LX_TAGS` (Makefile.lx, single source of truth), заменить `with_purego` → `with_musl`. `with_naive_outbound`**остаётся**. Остальные фичи (`with_xhttp,with_awg,…`) без изменений.
- **Без суффикса `-musl`**: upstream добавляет его, т.к. собирает и glibc, и musl на арку; у нас на Linux единственный вариант — musl, поэтому суффикс избыточен и сломал бы ожидание скриптов (`linux-arm64`/`linux-armv7`).
- naive присутствует (тег `with_naive_outbound` в сборке; по возможности — функциональная проверка).
- armv7/mipsle запускаются на реальном/эмулированном musl-роутере (`sing-box version` без ошибки загрузчика).
- darwin/windows/win7/android-ассеты не изменились по составу.
- **Верификация — через CI** (`workflow_dispatch`): Chromium musl-toolchain (гигабайты) недоступен локально на macOS, поэтому локальной сборки musl нет — приёмка по прогону workflow.
## 5. Вне скоупа
- Экзотические арки (`386`, `riscv64`, `loong64`, `mips64le`) — добавляются точечно строкой матрицы по запросу; cronet-musl под `mips64le` вообще нет.
- DEB/RPM/Pacman/OpenWrt-пакеты (upstream их делает) — нам не нужны, публикуем `.tar.gz`.
- naive на Win7 — физически невозможен (нет `cronet-go/lib/windows_386`).
Практический how-to по полям маскировки фичи 009. Это сахар над AmneziaWG `i1`:
вместо ручной CPS-строки `i1=<b 0x...>` пишешь домен/протокол/браузер, а движок
сам собирает пакет-приманку нужного протокола и шлёт его как `i1` перед handshake.
> Требуется сборка с `with_awg`. Без тега любой `id`/`ip`/`ib` отвергается:
> `AmneziaWG (awg) support is not included in this build, rebuild with -tags with_awg`.
---
## 1. Три поля
| Поле | Имя | Значения | Обязательно |
|------|-----|----------|-------------|
| `id` | домен | LDH-хост (`www.google.com`, `ozon.ru`, `_dmarc.example.com`) | **обязателен только для `quic`** (SNI); опционален для `dns` (QNAME или псевдо-домен), `sip` (host или псевдо-host) и `stun` (игнорируется) |
# 010 — WG-endpoint без `detour` режет download на Android (GRO split-brain)
| Поле | Значение |
|------|----------|
| Тип | B (bug) — расследование |
| Статус | **C (closed)** — корень подтверждён на железе (probe v2.1: `rxoffload=true`+`dispatch=single`), фикс верифицирован (download 0.44→20.7 Mbps, вровень с контрольной нодой), вмержен в `lx` (submodule `fb8d8d8`). Кандидат №2 не понадобился. **Обновление 1.14:** наш патч БОЛЬШЕ НЕ НУЖЕН — при миграции на v0.0.3 (re-graft submodule) фикс стал upstream-родным: коммит upstream `24ea133 «conn: harmonize GOOS checks between "linux" and "android"»` добавил `\|\| runtime.GOOS == "android"` в gейты приёмного пути `conn/bind_std.go` (строки 206/215/267/323/458). Наш §010-guard при миграции DROPPED (memory `wg-1.14-migration-is-submodule-rebase`). **Следствие:** GRO на Android теперь полностью рабочий (включается в `controlfns_linux.go:104` без android-guard + разбирается upstream-кодом) → большой `MaxSegmentSize=65535` (`device/queueconstants_android.go`) ему нужен как топливо. ⚠️ Откат `MaxSegmentSize→2200` ради экономии памяти задушит GRO-производительность download (то самое, что §010 чинил). Память от multi-WG нагрева лечить числом устройств, НЕ размером буфера. |
| Зона | ядро `sing-box-lx` + submodule `wireguard-go` (`conn/`) |
---
## Симптом
WireGuard-**endpoint** на Android без `detour`: download почти мёртв при живом
upload. Тот же конфиг с `"detour": "direct"` на endpoint'е — download нормальный.
Асимметрия (download убит, upload жив) указывает на дефект **только на приёме (RX)**.
`ClientBind` (detour-путь) не вызывает `controlFns`/`supportsUDPOffload`, читает по
одной датаграмме (`BatchSize()=1`) — offload-машинерии нет вообще, поэтому путь
иммунен. Это и объясняет «`detour: direct` лечит».
**Детерминирована только мёртвая RX-ветка разбора** (split-путь на android никогда
не зовёт `splitCoalescedMessages`). А вот **активация** offload — нет: `rxOffload`
взводится, лишь если ядро android вернуло `UDP_GRO == 1` (см. п.1), и даже при
взведённом `rxOffload` баг **проявляется**, только когда ядро коалесит в моменте —
нужен плотный входящий поток (download). На редком трафике ядро отдаёт по пакету —
склейки нет — работает. Поэтому «баг в коде» = неразбираемый GRO **при условии**, что
GRO вообще включился; первое детерминировано, второе — рантайм-зависимо и требует
repro (или прямого замера `rxOffload`, см. шаг 0).
---
## Опровергнутые гипотезы (по коду — не повторять)
| # | Гипотеза | Почему отвергнута |
|---|----------|-------------------|
| 1 | MTU / фрагментация | Симптом асимметричный (RX-only); при MTU резало бы симметрично. |
| 2 | У no-detour нет network-strategy / умного выбора интерфейса | Endpoint и direct зовут **один** конструктор `dialer.NewWithOptions → NewDefault`; `networkStrategy` гейтится только на `AutoDetectInterface`/`platformInterface`/`!disableDefaultBind` — одинаково для обоих. На Android оба привязаны к интерфейсу через `ProtectFunc == AutoDetectInterfaceFunc` (`route/network.go:340,368`). Разница не в выборе интерфейса. |
| 3 | Флаг `DirectOutbound` влияет на dialer | `DirectOutbound` — write-only поле; в `NewDefault` не передаётся и нигде не читается. |
| 4 | «Голый `ListenPacket` без стратегии» (`client_bind.go:89`) — корень | Эта ветка — multi-peer / без явного endpoint. Single-peer + валидный endpoint даёт `isConnect=true` (`endpoint.go:208`) → `DialContext`, не `ListenPacket`. А на no-detour `ClientBind` вообще не используется. |
---
## Открытый второй кандидат (если фикс GRO не лечит полностью)
**Тихий хэндовер / смена IP без смены интерфейса.** Re-bind сокета на смене сети идёт
через `onPauseUpdated` → `device.Up()` → `BindUpdate()`. Но событие `NetworkWake`
эмитится только из `notifyInterfaceUpdate`, а мобильный монитор
(`experimental/libbox/monitor.go:95-98`) дедуплицирует по **Name+Index, игнорируя
Addresses**. Смена source-IP на том же интерфейсе → событие подавляется → сокет не
переоткрывается. Не путь `StdNetBind`-vs-`ClientBind`, но самостоятельный сетевой
кандидат — держать открытым.
> Замечание против самой GRO-версии, которое надо снять repro: если баг чисто в
> GOOS-логике, он должен бить download и на стабильной сети, не только на сотовой.
> Если на стабильной сети download жив — либо GRO коалесит по-разному на разных
> интерфейсах, либо GRO не единственная причина (тогда вес смещается к кандидату №2).
---
## План проверки (эксперимент, не релиз)
Один дискриминирующий замер, изолирующий именно RX:
0.**(Бесплатно, без сборки ядра) Сначала снять факт `rxOffload`.** Залогировать
фактический возврат `supportsUDPOffload(conn)` → `(txOffload, rxOffload)` на
целевом Android. Это дискриминирует всю GRO-гипотезу до любого патча:
-`rxOffload == false` → GRO на этом ядре не активируется, корень №1 **мёртв** без
repro; вес немедленно уходит на кандидат №2 (тихий хэндовер).
-`rxOffload == true` → GRO взведён, переходим к шагу 1 (изолировать именно RX).
1. Запатчить `conn/controlfns_linux.go`: пропускать `setsockopt(UDP_GRO)` при
`runtime.GOOS == "android"` (TX/GSO **не трогать**, чтобы `txOffload` оставался
`true` и эксперимент проверял только RX). Тогда `rxOffload` читается `false` и
приём идёт обычным путём.
2. Собрать ядро и воспроизвести no-detour WG-endpoint на **любом** Android
(эмулятор/устройство — баг детерминирован в коде, оператор не нужен), снять
download под плотным потоком.
3. Желательно — пакетная проверка: реально ли `recvmsg` отдаёт >MTU датаграммы на
этом сокете (подтверждает, что GRO коалесит).
Исходы:
- download починился при живом upload → корень = GRO-на-android **подтверждён**, и
минимальный фикс найден тем же шагом;
- не починился → переходим к кандидату №2 (тихий хэндовер).
---
## Решение
Пока **нет** — это таска-расследование. Код в ядро/submodule — только после repro.
**Кандидат на фикс (когда подтверждён):** гейтить `UDP_GRO`-setsockopt
(`controlfns_linux.go`) и/или чтение `rxOffload` (`features_linux.go`) за `!android`,
чтобы `StdNetBind` на android не объявлял offload, который не умеет разбирать. Это
откатывает android на «без offload» — поведение, идентичное рабочему detour-пути.
> NB: фикс «добавить RX self-disable в linux-ветку» был бы **no-op** — эта ветка на
> android мёртвый код. Корень = split-brain GOOS, а не отсутствие fallback.
---
## Acceptance (для будущего фикса)
- [ ] No-detour WG-endpoint на Android даёт download, сопоставимый с `detour: direct`.
- [ ] Фикс не ломает offload/производительность на «настоящем» Linux (не-android) —
гейт за `!android`, а не глобальное отключение.
- [ ] Регресс: plain WG **и** AmneziaWG endpoint; single-peer (`isConnect`) и
multi-peer (`ListenPacket`) пути.
- [ ] Юнит/интеграционный тест на coalesced-receive (сейчас отсутствует — поэтому
дефект и проскочил).
---
## Источники (проверено по живому коду)
-`transport/wireguard/endpoint.go:200-215` — развилка StdNetBind vs ClientBind.
**Дата:** 2026-06-21 · **Статус:** Complete (синтетика) · **Лайв:** ⚠️ НЕ ПРОГОНЯЛСЯ — нет доступа к reality+xhttp ноде; приёмка на синтетике по решению владельца · **База:** ветка `lx` (`1.13.13-lx.13`)
> **Honest caveat.** Фикс принят на основании: построчной сверки с исходниками Xray (контракт stream-one = голый путь / пустой sessionId), независимого подтверждения [issue #5635](https://github.com/XTLS/Xray-core/issues/5635), совпадения с портом hiddify, и зелёной синтетики (юнит-тесты URL-layout + reality-детект, `check`, сборки). Это **не** заменяет лайв против реального Xray-сервера. Лайв остаётся открытым TODO — при первом доступе к reality+xhttp ноде прогнать сценарии из раздела «Дальше» и, если что-то не так, переоткрыть задачу.
## Проблема (из жалобы)
`vless + reality + xhttp`, `mode:auto`, `path:/` работает на стороннем (Xray-логика) ядре, **не работает** на нашем. Две связанные первопричины, сверены построчно с исходниками Xray-core `transport/internet/splithttp` (`main`) и подтверждены [issue #5635](https://github.com/XTLS/Xray-core/issues/5635) + референс-портом hiddify:
1.**stream-one слал `sessionId` в пути.** Xray-сервер (`hub.go`) роутит stream-one (двунаправленный) ТОЛЬКО при пустом sessionId. Наш `dialStreamOne` строил `<path>/<sessionId>` → сервер уходил в stream-down ветку → downlink не-VLESS → VLESS `unknown version`. (Зафиксировано как known bug ещё в 002.)
2.**`mode=auto` всегда → packet-up.** Xray: auto + Reality → stream-one. У нас auto лип к packet-up (т.к. stream-one был сломан).
## Что сделано
Изменения **только** в пакете `transport/v2rayxhttp` (новый код — ребейз-зона = ∅, upstream не тронут).
- ⚠️ **Лайв против реального Xray reality+xhttp сервера** — НЕ выполнен (нет доступа к ноде). По решению владельца задача принята на синтетике (статус → **C**), лайв остаётся открытым TODO (см. «Дальше»). При первом доступе к ноде — прогнать и при расхождении переоткрыть.
## Ребейз-зона
**∅.** Все изменения — в новых файлах пакета `v2rayxhttp` и его правках (новый код фичи). Upstream-файлы (`transport/v2ray/transport.go`, `constant`, `option/v2ray_transport.go`) — не тронуты.
## Остаточные риски
- **Матч Reality по имени типа** (`reality_detect.go`) хрупок к переименованию `RealityClientConfig` в sing/upstream. Митигировано юнит-тестом (двойники с теми же именами) — но тест останется зелёным при переименовании реального типа, а лайв сломается. При ребейзе сверять имя типа в `common/tls/reality_client.go`.
- **stream-up** по-прежнему не лайв-тестился (как и в 002).
## Дальше (открытый TODO — лайв)
При первом доступе к реальной Xray reality+xhttp ноде:
1.`make -f Makefile.lx lx-build`; собрать аутбаунд с параметрами рабочей подписки (server/uuid/sni/pbk/sid/path), локальный socks/mixed на `127.0.0.1:2080`.
2.`mode:stream-one` — `curl -x socks5h://127.0.0.1:2080 https://api.ipify.org` должен вернуть IP сервера (handshake+DNS+HTTPS+download).
3.`mode:auto` на той же ноде — идентичный результат (резолв в stream-one).
4. Регрессия: `mode:packet-up` на packet-up-ноде по-прежнему работает.
5. Если ок — оставить как есть (уже C). Если расхождение — переоткрыть задачу.
Слияние в `lx` и релиз `-lx.N` — на усмотрение владельца (фикс на ветке `lx-xhttp-streamone`, в `lx` не влит).
sessionID в stream-one больше нигде не используется (можно убрать его генерацию для этой ветки, но проще оставить — он безвреден, в URL не идёт). Остальное (POST, pipe-body, streamConn late-binding) — без изменений: разведка подтвердила, что late-binding корректен, баг был только в URL.
Сейчас `case modeAuto, modePacketUp:` → `dialPacketUp`.
**Правка:** выделить `modeAuto` в отдельную ветку:
```go
casemodeAuto:
ifc.realityEnabled{
returnc.dialStreamOne(ctx,sessionID)
}
returnc.dialPacketUp(ctx,sessionID)
casemodePacketUp:
returnc.dialPacketUp(ctx,sessionID)
```
`c.realityEnabled` — новое bool-поле на `Client`, проставляется один раз в `NewClient`.
### 3.4 REALITY-детект в `NewClient` — БЕЗ межтеговой зависимости (РАЗВИЛКА)
`*tls.RealityClientConfig` / `*tls.KTLSClientConfig` — под `//go:build with_utls`; `v2rayxhttp` — под `with_xhttp`. Прямой `tlsConfig.(*tls.RealityClientConfig)` введёт жёсткую связь `with_xhttp → with_utls` и сломает сборку `with_xhttp` без `with_utls` (нарушение CONSTITUTION §3.2). Варианты:
- **(A) Матч по имени типа (рекомендуется).** В `NewClient`:
```go
realityEnabled := tlsConfigIsReality(tlsConfig)
// helper: reflect.TypeOf(unwrap(tlsConfig)).String() содержит "RealityClientConfig"
```
Разворачивать KTLS-обёртку: у `*KTLSClientConfig` встроено поле `Config Config` → если имя типа = KTLS, взять inner и проверить снова. Делать через рефлексию по имени поля/типа, **без** импорта with_utls-типов. Плюс: нулевая межтеговая связь, работает и для kTLS. Минус: матч по строке имени типа — хрупковато к переименованию upstream (митигируется тестом).
- **(B) Проброс флага из вызывающего слоя.** Добавить признак reality в `option.V2RayXHTTPOptions` или в сигнатуру конструктора. Минус: правка upstream-сигнатуры `ClientConstructor`/диспетчера — расширяет ребейз-зону, противоречит «новый код в новых файлах». Отклонено.
- **(C) Эвристика по ServerName/NextProtos.** Ненадёжно (reality неотличим от обычного uTLS по этим полям). Отклонено.
**Решение: (A)** — изолированный helper в `client.go` (или соседнем lx-файле пакета), детект по имени типа с разворачиванием KTLS, покрытый юнит-тестом. Если по ходу выяснится, что рефлексия по приватному полю KTLS недоступна — fallback: матчить и `RealityClientConfig`, и `KTLSClientConfig` по суффиксу имени (для не-Linux kTLS не используется, риск низкий).
---
## 4. Порядок работ
1. Ветка `lx/xhttp` от `lx` (по git-дисциплине; сейчас HEAD на `lx-gro-probe-010`).
2. `requestURL` bare-path ветка + `dialStreamOne` голый путь (фикс 3.1 главный — проверяем stream-one лайв сразу).
4. Тест-конфиги, `sing-box check`, лайв-проверка (stream-one + auto на reality-ноде; packet-up регрессия).
5. DoD, IMPLEMENTATION_REPORT, статус (шапка SPEC.md + Roadmap) → C.
---
## 5. Риски
- **Лайв-сервер обязателен.** Синтетического `check` мало — XHTTP под активной разработкой, нужна reality-нода Xray. stream-one лайв-валидируем явно; auto — на той же ноде убеждаемся, что резолвится в stream-one.
- **Матч по имени типа (3.4-A)** хрупок к переименованию `RealityClientConfig` в upstream/sing. Митигировать юнит-тестом, который при ребейзе сразу покраснеет.
- **h2 для stream-one обязателен** — наш h2-only транспорт это обеспечивает; не регрессируем packet-up (он тоже h2 поверх reality).
- Не сломать stream-up/packet-up URL (оставляют sessionId) — правим только пустую-elem ветку `requestURL` и только `dialStreamOne`.
Починить XHTTP-режим **`stream-one`** (сломан с момента 002) и привести **`mode=auto`** к поведению Xray, чтобы конфиги `vless + reality + xhttp + mode:auto` поднимались на нашем ядре «как есть».
Build-tag: `with_xhttp`. Scope: **client-only**.
---
## 1. Проблема / контекст
Жалоба (2026-06-21): простейший `vless + reality + xhttp`, `mode:auto`, `path:/` работает на стороннем ядре (Xray-логика), **не работает на нашем**. Это типовой конфиг из панелей/подписок.
Две связанные первопричины (обе сверены построчно с исходниками Xray-core `transport/internet/splithttp` ветки `main` и подтверждены [issue #5635](https://github.com/XTLS/Xray-core/issues/5635), а также референс-портом hiddify):
### 1.1 `stream-one` шлёт `sessionId` в пути — главный баг
В Xray `stream-one` — это **один POST на голый путь без sessionId**; сервер (`hub.go`) роутит режим по наличию sessionId:
Наш `dialStreamOne` ([transport/v2rayxhttp/conn.go:21](../../transport/v2rayxhttp/conn.go)) строит URL как `c.requestURL(sessionID)` → `<path>/<sessionId>`. Сервер парсит **непустой** sessionId, уходит в stream-down ветку (ждёт парный stream-up POST, которого нет), и в response.Body летят **не VLESS-байты**. VLESS-парсер читает первый байт как версию → `unknown version` (часто `0x58='X'` из HTTP-обвязки). Зафиксировано как «known bug» ещё в [002 IMPLEMENTATION_REPORT](../002-XHTTP_CLIENT_TRANSPORT/IMPLEMENTATION_REPORT.md).
### 1.2 `mode=auto` всегда → packet-up
Xray (`dialer.go`): `auto` → packet-up по умолчанию, **но если REALITY → stream-one** (если ещё и `downloadSettings` → stream-up). Решает только наличие REALITY/downloadSettings, не h2/h3.
Наш `DialContext` ([client.go:155](../../transport/v2rayxhttp/client.go)) намеренно лепит `auto` к packet-up (т.к. stream-one был сломан). После фикса 1.1 `auto` должен резолвиться как Xray: **REALITY → stream-one**, иначе packet-up. Тогда конфиг из жалобы работает без правок пользователя.
---
## 2. Цель
`vless/vmess/trojan` outbound с`transport.type=xhttp`, `mode:stream-one` — поднимает рабочее соединение к Xray XHTTP-серверу (handshake + DNS + HTTPS + загрузка), в т.ч. поверх Reality. `mode:auto` при включённом Reality резолвится в `stream-one` (как Xray). `mode:packet-up`/`stream-up` — без регрессий.
---
## 3. Требования
### 3.1 Фикс `stream-one` (главное)
-`stream-one` шлёт запрос на **голый нормализованный путь** (`<path>`), **без**`sessionId` в URL (и нигде — ни query, ни header). Метод — `POST` (как сейчас), тело — uplink-pipe, downlink — response.Body того же запроса (late-binding уже корректен — `streamConn.created`).
-`requestURL()` при **пустом** наборе элементов обязан вернуть голый `<path>`**без** trailing-slash. Сейчас `requestURL()` с пустым elem даёт `<path>/` (ловушка `strings.Join([], "/")==""` → `c.path + "/"`), что отличается от Xray/hiddify (`<path>`).
-`stream-up` и `packet-up` URL **не трогаем** — они законно используют sessionId (`conn.go:48,52,86,236`).
### 3.2 `mode=auto` как Xray
-`auto` + **Reality включён** → `stream-one`.
-`auto` + Reality выключен → `packet-up` (текущая совместимая ветка; download-settings у нас нет — stream-up в auto не выбираем).
- REALITY-признак определяется **в конструкторе** (`NewClient`), один раз, и сохраняется на `Client` (в `DialContext` tlsConfig вне области видимости).
-`*tls.RealityClientConfig`/`*tls.KTLSClientConfig` объявлены под `//go:build with_utls`, а пакет `v2rayxhttp` — под `with_xhttp`. **Запрещён прямой type-assert** на with_utls-типы из v2rayxhttp: это введёт жёсткую зависимость `with_xhttp → with_utls` и сломает сборку с `with_xhttp` без `with_utls` (нарушение §3.2 CONSTITUTION — фича за своим тегом).
- Детект REALITY делать **без прямой ссылки на with_utls-тип**: по имени конкретного типа (`reflect`/`fmt %T`, сопоставление с суффиксом `RealityClientConfig`) либо иным способом, не вводящим импорт-связь между тегами. Способ фиксируется в PLAN.
### 3.4 Что НЕ трогаем (доказано в разведке)
- **ALPN.** Наш форсинг `["h2"]` → uTLS добавляет `http/1.1` → `["h2","http/1.1"]`. Xray `decideHTTPVersion` для списка длиной ≠1 → h2; reality всегда h2. Форсинг benign, а h2 для bidirectional stream-one **обязателен**. Оставляем как есть.
- **Padding.** `x_padding` в query внутри `Referer` совпадает с Xray (client request direction). Не трогаем.
- **Submodule, server/inbound** — вне scope.
---
## 4. Критерии приёмки
-`sing-box check -c` принимает `vless + reality + xhttp + mode:stream-one` и `mode:auto` (есть/будет тест-конфиг в `lx-test/config`).
- **Лайв:** реальный Xray XHTTP-сервер (reality-нода) — `mode:stream-one` поднимает соединение (handshake + DNS + HTTPS-страница + загрузка); `mode:auto` на той же ноде даёт идентичный результат (резолвится в stream-one). packet-up/stream-up — без регрессий.
- Сборка **с**`with_xhttp`**без**`with_utls` — компилируется (изоляция тега не нарушена).
- Сборка без `with_xhttp` = поведение upstream (xhttp отвергается).
-`go vet` (lx-теги) и `gofmt -l` по затронутым файлам — чисто. `go build ./...` без тегов — ок.
- Ребейз-зона не расширяется: правки только в новых файлах пакета `v2rayxhttp` и (возможно) комментарий в `option/v2ray_xhttp.go`. Upstream-файлы — не трогаем.
Разбор для зомби-соединения (download, ↓0 на tun0):
| Снимок | Где застряло | Трактовка |
|---|---|---|
| `read=0 write=0` | **выше** copy | proxy (reality/vless) не отдал НИ ОДНОГО расшифрованного байта, хотя pcap показал зашифрованные 777B → дефект расшифровки/фрейминга, НЕ в `connectionCopy` |
| `read>0 write=0` | **запись в tun** | байты из upstream прочитаны, но не записаны/не флашатся в gVisor tun-сокет → подтверждает гипотезу SPEC «застряло в `remoteConn→conn`» |
| `read>0 write>0` | copy шёл | смотреть на `err`/`timeout` — копирование двигалось, причина в завершении/лаге, не в самом stuck |
Ключ к развилке pcap: обёртка стоит на **расшифрованном**`remoteConn`, поэтому
`read` — это plaintext. `read=0` при наличии зашифрованного трафика в pcap снимает
с`connectionCopy` подозрение и переводит расследование выше по стеку (proxy-слой).
---
## Механика (почему обёртка перехватывает, а не обходится)
Обёртка `lxTraceConn` (`route/conn_trace_lx.go`) встраивает `net.Conn`, считает байты
в `Read`/`Write`, и **намеренно НЕ реализует**`Upstream()` / `ReaderReplaceable()` /
`WriterReplaceable()` / `SyscallConn()`. Это критично — иначе её обойдут:
висит**. Для висящего зомби это и есть прямой снимок: серия `tick#1 write=0`,
`tick#2 write=0`… показывает застревание в реальном времени, разрывать НЕ нужно.
- `lx-trace download final: read=… write=… err=…` — один раз, при отвисании
(↓0→↓2820) или разрыве (`DELETE /connections/<id>`).
Искать в `/tmp/lx_trace_run.log` строки `lx-trace download` для нужного conn (сверить
с `destinationIP`/`sourcePort` из шага 4).
6. **Сопоставить** найденную строку с pcap-потоком по времени/объёму (тот же conn).
## Чтение результата (развилка)
| Снимок | Вывод | Следующий шаг расследования |
|---|---|---|
| `read=0 write=0` | proxy (reality/vless) не отдал plaintext, хотя pcap видел зашифрованные 777B | копать **выше** `connectionCopy`: расшифровка/фрейминг vless/reality на download |
| `read>0 write=0` | байты прочитаны из upstream, не записаны в tun | подтверждение гипотезы SPEC; копать запись в gVisor tun-сокет (флаш/блокировка) |
| `read>0 write>0` | copy двигался | смотреть `err`/`timeout`; причина в завершении/лаге, не в чистом stuck |
## Период тика
Тик уже встроен. Период задаётся значением `LX_CONN_TRACE`: `1`/`true` → 5s (дефолт);
`3s`/`1s`/`500ms` → этот период. Для висящего зомби 5s достаточно; если нужен более
плотный снимок прогресса — поставить `LX_CONN_TRACE=1s`. Помнить: чем чаще тик, тем
больше строк в core-логе.
## Не забыть
- `LX_CONN_TRACE` — только на время диагностики (теряется zero-copy splice, half-close
деградирует в Close). После прогона — выключить / поставить релизное ядро обратно.
- Метод, который сработал и который НЕ менять: синхронный tun0+wlan0 ОДНОЙ командой
| Статус | **C (closed) — НЕ удалось воспроизвести на lx.14.** Симптом наблюдался на РАЗНЫХ нодах, включая WG → это зонтик над «↓0»-сталлом, не один баг. WG-долю закрыл фикс **§010 GRO** (вошёл в lx.14, доказан). Для не-WG нод (VLESS/reality) §010 не применим, отдельного код-фикса нет — там симптом сейчас не воспроизводится без подтверждённого объяснения. См. раздел «Закрытие 22.06» ниже. |
| Зона | ядро `sing-box-lx` — `route/conn.go` (`connectionCopy`, общий relay); фикс WG-доли — submodule `wireguard-go` (GRO, §010, UDP/WG-only) |
| Связь | WG-доля симптома = [010-WG_ENDPOINT_GRO_SPLIT_BRAIN](../010-WG_ENDPOINT_GRO_SPLIT_BRAIN/SPEC.md) (closed, фикс в lx.14). §010 это UDP/WG-only → НЕ объясняет не-WG (VLESS) случаи; «виснет и на VLESS» строго не подтверждено (см. «Закрытие») |
| Артефакты | [PROBE.md](PROBE.md) — зонд (env-гейт `LX_CONN_TRACE`, не прогнан в боевом режиме); [instrumentation.patch](instrumentation.patch) — диф зонда; [RUN-PLAN.md](RUN-PLAN.md) — процедура прогона (на случай повторного появления) |
---
## Симптом
Жалоба: WhatsApp/Telegram «висят» — чаты/медиа не грузятся. В клиенте (LxBox
Conns) у приложения одно TCP-соединение, оно **зомби**: `↑517/629 ↓0` — ClientHello
Добавить rule-item **`package_name_regex`** (route / DNS / headless) — матчинг имени Android-пакета по регулярному выражению. Точечный бэкпорт апстрим-фичи 1.14 на стабильную базу 1.13.13 **без** полной миграции на 1.14.
Scope: **все платформы** (фича активна там, где заполняется `ProcessInfo.AndroidPackageNames`, т.е. Android). Build-tag: нет — встроена в ядро роутинга.
> **Обновление (2026-07-02, база 1.14.0-alpha.35, аудит SPEC 022 #19):** после миграции базы на 1.14 сам impl (`route/rule/rule_item_package_name_regex.go` + поле `option.RawDefaultRule.PackageNameRegex`) стал **нативным upstream** — бэкпорт-дельта по коду больше не нужна и растворилась в базе. НО LX-тест `route/rule/rule_item_package_name_regex_test.go` **сохраняется намеренно**: upstream своего теста для этого item не поставляет (проверено на базе и на `upstream/testing`), так что это единственное покрытие фичи. Тест изолирован в своём `_test.go`, тестирует стабильный публичный API (`NewPackageNameRegexItem`/`Match`) и ребейз-конфликтов не несёт.
---
## 1. Проблема / контекст
Запрос (2026-06-23): нужен `package_name_regex` в проекте. У апстрима поле существует **только с sing-box 1.14.0** (commit [`941ce58b`](https://github.com/SagerNet/sing-box/commit/941ce58b) «Add `package_name_regex` route, DNS and headless rule item»), в ветке 1.13.x его нет. На нашей базе уже есть `package_name` (точное совпадение, map-lookup) — но не regex-вариант.
Полная миграция 1.13.13→1.14 оценена отдельным feasibility-разбором как ~1,5–2 дня работы с главным риском в ребейзе AmneziaWG-подмодуля `wireguard-go` (база 506b763 → v0.0.3, ветки diverged 52/51, ручная переинсерция §010 android-GRO fix). Сама же фича `package_name_regex` — изолированный add в `route/rule`, **не трогает** ни один awg/xhttp/selector/build-tag файл и **не гейтится** новым build-тегом. Поэтому выбран точечный бэкпорт, а полная миграция отложена до выхода **v1.14.0 stable** (её штатно подхватит существующий `lx-rebase.yml`, который по дизайну исключает alpha/beta/rc).
Апстрим-коммит `941ce58b` дополнительно содержит хунк про `C.RuleSetVersion5` в `option/rule_set.go` — это часть отдельного rule-set v5 (1.14), **не относится** к фиче и **не переносится** (на нашей базе `RuleSetVersionCurrent = RuleSetVersion4`).
---
## 2. Цель
Правило роутинга / DNS-правило / headless-правило (rule-set) с полем `package_name_regex: ["^com\\.termux.*", ...]` матчит соединение, если хотя бы одно из имён пакетов в `metadata.ProcessInfo.AndroidPackageNames` удовлетворяет хотя бы одному из выражений. Семантика и сообщения об ошибках — идентичны апстрим-1.14.
---
## 3. Требования
### 3.1 Новый rule-item
- Файл `route/rule/rule_item_package_name_regex.go` — дословно апстрим-версия из `941ce58b`: `PackageNameRegexItem`с`[]*regexp.Regexp`, конструктор `NewPackageNameRegexItem([]string) (*PackageNameRegexItem, error)` (компиляция через `regexp.Compile`, ошибка `parse expression <i>`), `Match` по `AndroidPackageNames`, человекочитаемый `String()` (усечение до 3 выражений в описании).
### 3.2 Option-поля
-`PackageNameRegex badoption.Listable[string]`с тегом `json:"package_name_regex,omitempty"` — сразу после `PackageName` в трёх структурах: `option.RawDefaultRule` ([option/rule.go](../../option/rule.go)), `option.RawDefaultDNSRule` ([option/rule_dns.go](../../option/rule_dns.go)), `option.DefaultHeadlessRule` ([option/rule_set.go](../../option/rule_set.go)). Выравнивание struct-тегов — под существующий столбец каждого файла (gofmt-чисто).
### 3.3 Регистрация item в правилах
-В`NewDefaultRule` ([route/rule/rule_default.go](../../route/rule/rule_default.go)), `NewDefaultDNSRule` ([route/rule/rule_dns.go](../../route/rule/rule_dns.go)), `NewDefaultHeadlessRule` ([route/rule/rule_headless.go](../../route/rule/rule_headless.go)) — блок `if len(options.PackageNameRegex) > 0 { ... }` сразу после `PackageName`-блока, с проброской ошибки `E.Cause(err, "package_name_regex")`. Все три конструктора уже возвращают `error` на нашей базе — сигнатуры не меняются.
### 3.4 Cond-функции
-`isProcessRule` / `isProcessDNSRule` ([route/rule_conds.go](../../route/rule_conds.go)) и `isProcessHeadlessRule` ([route/rule/rule_set.go](../../route/rule/rule_set.go)) — добавить `|| len(rule.PackageNameRegex) > 0`, чтобы правило с одним лишь `package_name_regex` корректно классифицировалось как process-rule.
### 3.5 Что НЕ трогаем
- Хунк `RuleSetVersion5` из апстрим-коммита (см. §1) — **не переносим**.
-`go build ./...` без тегов и сборка с lx-тегами (`with_gvisor with_quic with_wireguard with_utls with_clash_api with_xhttp with_awg`) — ок. ✅
-`go vet ./route/... ./option/...` и `gofmt -l` по затронутым файлам — чисто. ✅
- Юнит-тест `route/rule/rule_item_package_name_regex_test.go` зелёный: матч префикса/якоря `$`, матч одного из нескольких пакетов, no-match, nil `ProcessInfo` (без паники), ошибка на невалидном выражении. ✅
- Ребейз-зона: фича — это новый файл + точечные правки в 6 файлах роутинга/опций; коллизий с awg/xhttp/selector/CI-кластерами нет (подтверждено feasibility-разбором).
---
## 5. Вне скоупа
- Полная миграция на 1.14 (отложена до v1.14.0 stable; отдельный feasibility-отчёт).
- [Документация route/rule#package_name_regex](https://sing-box.sagernet.org/configuration/route/rule/#package_name_regex) — «Match android package name using regular expression», since 1.14.0.
- Feasibility-разбор миграции 1.13.13→1.14 (этой сессии) — обоснование точечного бэкпорта вместо полного перехода.
| Тип | F (feature) — смена канала управления ядром (client-side) |
| Статус | A (accepted) — `with_clash_api` drop из AAR в `v1.14.0-lx.1-rc.1`; box.go-фикс в `rc.3`; десктоп-регрессия исправлена в `rc.17` (§3.4) |
**Переезд управления ядром с Clash API на нативный libbox CommandClient — на Android.** LxBox перестаёт использовать Clash REST API и переходит на нативный gRPC-канал `StartedService` (поверх unix-сокета). Из **AAR-сборки** убирается `with_clash_api` — отпадает HTTP-сервер Clash и связанный attack surface. **Десктоп/CLI сохраняют `with_clash_api`** (внешние дашборды ходят по Clash REST API; нативного CommandClient-канала у CLI нет) — см. §3.4.
Этот SPEC фиксирует **сам переезд и его последствия**. Доработки command-протокола, понадобившиеся, чтобы CommandClient заменил Clash API по функциональности (per-node delay, таблица правил, pull-снапшоты групп, фикс потери групп), вынесены в отдельный **[SPEC 015 — COMMAND_PROTOCOL_RPC_EXTENSIONS](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md)**.
Графические клиенты sing-box исторически управляли ядром через два канала: Clash REST API (`experimental.clash_api`, под build-tag `with_clash_api`) и нативный libbox **CommandClient**. Clash API — это слой совместимости со сторонними дашбордами; для **своего** клиента на устройстве разработчики предполагают именно CommandClient (gRPC поверх локального unix-сокета).
LxBox переходит на CommandClient как единственный канал управления, потому что:
- это нативный, более богатый интерфейс (closed-connections история, per-event дельты соединений, ProcessInfo раздельными полями, NQ/STUN/Tailscale-инструменты) — то, что Clash REST не покрывает;
- Clash API — это лишний HTTP-сервер в процессе и открытый локальный порт (attack surface), не нужный, когда клиент ходит по нативному каналу;
- убрав `with_clash_api`, мы уменьшаем дифф и размер AAR.
---
## 2. Цель
LxBox (Android) управляет ядром **только** через CommandClient; `with_clash_api` не входит в **AAR-сборку**. Конфиг, ссылающийся на `experimental.clash_api`, fail-fast с понятной ошибкой (а не молчаливо деградирует). Функциональный паритет с Clash API по нужным UI возможностям достигается доработками CommandClient — см. [SPEC 015](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md).
> **Важно (исправлено):** дроп `with_clash_api` относится **только к Android AAR**. Десктоп/CLI-бинари (mac/windows/linux) управляются внешними дашбордами (yacd/MetaCubeXD) **именно через Clash REST API** — нативного CommandClient-канала вне gomobile/libbox у них нет. Поэтому `with_clash_api` **остаётся** в десктоп `LX_TAGS`. Изначально (rc.1) тег был ошибочно убран и из десктоп-набора тоже — см. §3.4.
- Убрать `with_clash_api` из `sharedTags` ([cmd/internal/build_libbox/main.go](../../cmd/internal/build_libbox/main.go), `// lx:`-блок). **Десктоп `LX_TAGS` (`Makefile.lx`) тег сохраняет** — см. §3.4.
- Без тега (в AAR) подключается `include/clashapi_stub.go` — конфиг с`experimental.clash_api` получает `clash api is not included in this build, rebuild with -tags with_clash_api` (fail-fast, **не** молчаливый отказ).
- lx-конфиги на Android `clash_api` не используют — управление идёт через CommandClient.
- Сделано в `v1.14.0-lx.1-rc.1` (commit `57b5b5e5`) — но изначально ошибочно срезано и с десктопа, исправлено в §3.4.
### 3.2 Доработки CommandClient → SPEC 015
Нативный CommandClient беднее Clash API по ряду возможностей, нужных UI (per-node delay-тест, таблица правил, pull-снапшоты групп/узлов, баг потери одно-узловых групп). Все эти доработки — **в [SPEC 015](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md)** (класс §3.6, build-tag `with_lx_command`). `URLTestOutbound` и `GetRules` уже зашиплены (rc.2); `GetGroups`/`GetOutbounds` + фикс `len<2` — target rc.4. Здесь они только упоминаются как часть полного перехода; тех-спека — в 015.
синхронный), `cc_channel.dart:177` (`ccUrlTestOutbound` через MethodChannel).
---
# Ответ ядра — вариант #2 работает БЕЗ правок биндинга
**От:** команда ядра sing-box-lx
**К:** LxBox
**Дата проверки:** против `experimental/libbox/command_client.go` + `common/urltest/urltest.go` (HEAD ветки lx-1.14)
**Итог:** ваш блокер снимается вариантом #2 — отдельный ping-`CommandClient` + его `Disconnect()`. Правка биндинга (#1/#3) НЕ нужна для устранения «зомби». Разбор ниже.
## Прямой ответ на главный вопрос
> «Рвёт ли `CommandClient.disconnect()` уже-ушедшие в dial per-call тесты?»
**Да.** Цепочка проверена по коду:
1.`Disconnect()` ([command_client.go:295](../../experimental/libbox/command_client.go)) делает ДВЕ вещи: `c.cancel()` (отменяет общий `c.ctx`) **и**`c.grpcConn.Close()`.
2. После слоя 1 серверный хэндлер привязал тест к gRPC **per-call**`ctx` (`testCtx := ctx`, [started_service_command_lx.go](../../daemon/started_service_command_lx.go)). Обрыв клиентского вызова/транспорта → gRPC-Go рантайм отменяет серверный stream-ctx этого вызова (стандартный `grpc.NewServer`, никакой обёртки, отвязывающей ctx, нет — [daemon/server.go:16](../../daemon/server.go), [command_server.go:164](../../experimental/libbox/command_server.go)).
3. Этот ctx течёт в `urltest.URLTest(testCtx, …)` → в **оба** ctx-aware этапа: `detour.DialContext(ctx, …)` (TCP/proxy connect+handshake) и `client.Do(req.WithContext(ctx))` (HTTP HEAD) — [urltest.go:99,127](../../common/urltest/urltest.go). Отмена ctx обрывает уже-ушедший dial, не дожидаясь `C.TCPTimeout`.
Итог: dial **НЕ** доживает до timeout независимо от disconnect — он падает по отмене ctx. «Зомби» закрываются.
## Но: рвите ОТДЕЛЬНЫЙ ping-client, не общий
Нюанс, который надо учесть. `Disconnect()` через `c.cancel()` отменяет **общий**`c.ctx` ([command_client.go:32-33](../../experimental/libbox/command_client.go)) — один на ВСЕ вызовы этого инстанса. Если дёрнуть `Disconnect()` на вашем основном client'е, оборвутся и Connections/Groups/Status-стримы. Поэтому:
**Держите ОТДЕЛЬНЫЙ `CommandClient`-инстанс под масс-пинг.** Каждый `NewCommandClient` ([command_client.go](../../experimental/libbox/command_client.go)) поднимает СВОЙ `c.ctx`/`c.cancel` и СВОЙ `grpcConn` (`Connect()`/`ConnectWithFD` — [:239](../../experimental/libbox/command_client.go), [:266](../../experimental/libbox/command_client.go)). Значит:
-`pingClient.disconnect()` отменяет только per-call ctx ping-тестов, рвёт только ping-conn;
- остальные стримы целы.
Это **ровно ваш fallback #2**, и он реализуем на текущем `v1.14.0-lx.1`-биндинге как есть: `urlTestOutbound` уже экспонирован, `disconnect()` уже экспонирован. Новой нативной поверхности не требуется.
## Почему epoch-гейт сам по себе не закрывал зомби (и почему теперь закроет)
Ваш epoch-гейт гасит **применение** результатов на стороне UI, но `urlTestOutbound` — синхронный блокирующий gomobile-вызов, и до слоя 1 серверный тест был привязан к `boxService.ctx` (жил, пока жив сервис) — отменить его было нечем, кроме сноса всего client'а. Теперь, при отдельном ping-client: epoch-бамп (мгновенный UI) **+** `pingClient.disconnect()` (рвёт серверные dial'ы) = и UI, и ядро реагируют. Воркер-пул после disconnect просто получит ошибки на оставшихся вызовах — их и так гасит epoch-гейт.
Практически: при «отмене» зовите `pingClient.disconnect()` и поднимайте свежий `pingClient` под следующий прогон (либо реконнект того же). Стоимость — один short-lived conn на прогон масс-пинга, дёшево.
## Про варианты #1 / #3 (per-call handle / CancelToken)
Не отвергаем, но считаем **избыточными** для вашей задачи:
-#1/#3 дают гранулярность «отменить ОДИН узел из батча, оставив остальные». Для масс-отмены («отменить весь прогон») это не нужно — disconnect ping-client'а гасит весь батч разом.
- Цена #1/#3 — новый stateful слой в gomobile-поверхности (handle-реестр / `CancelToken`-тип, биндинг, версионирование AAR). Это та самая «новая подсистема», которую SPEC 015 §3.6 просит не плодить.
- Если позже появится UX «перепинговать только этот узел с отменой» — вернёмся к #1. Пока YAGNI.
## Что нужно от вас для подтверждения
Проверьте на устройстве сценарий: запустить масс-пинг (concurrency=10) на медленных/недоступных узлах → нажать «отмена» → убедиться, что (а) серверные dial'ы рвутся в пределах ~момента, не висят до `C.TCPTimeout`; (б) Connections/Groups-стримы основного client'а не мигают/не пересоздаются. Если (а) не подтвердится — пришлите лог, копнём транспортный слой конкретного outbound (теоретически отдельный outbound мог бы игнорировать ctx в своём dial — но `urltest.URLTest` зовёт его правильно).
## Сводка ответа
- Вопрос #2 (disconnect рвёт per-call ctx тестов) — **ДА**, проверено по коду. Реализуйте #2 сами.
- Условие: **отдельный** ping-`CommandClient`-инстанс (свой `c.ctx`/`c.cancel`/conn — уже так устроено), чтобы disconnect не задел другие стримы.
- Правки биндинга (#1/#3) — не требуются; держим #1 в запасе под будущий per-node UX.
| Тип | F (feature) — расширения libbox command-протокола (CONSTITUTION §3.6) |
| Статус | M (mixed) — `URLTestOutbound` + `GetRules` в `v1.14.0-lx.1-rc.2`; `GetGroups`/`GetOutbounds` + фикс `len<2` в `v1.14.0-lx.1-rc.4`; `URLTestOutbound` cancel-fix (ctx-bind, §3.6) — в коде, ждёт релиз-тега. Слой 2 (масс-отмена) разблокирован на клиенте без правок биндинга (отдельный ping-client + `Disconnect()`); LxBox-фидбэк закрыт |
**Единый дом всех доработок нативного libbox CommandClient** (gRPC `StartedService`), доведших его до минимума, на котором UI LxBox реально работает после переезда с Clash API (см. [SPEC 014](../014-CLASH_API_TO_COMMANDCLIENT_MIGRATION/SPEC.md) — сам переезд).
Пять RPC-доработок, все одного класса §3.6 (handler'ы за `with_lx_command` по образцу `daemon/started_service_usbip{,_stub}.go`, шов в `.proto` под `// lx:`-маркером, логика в `*_lx.go`):
**Апстрим-кандидатность (см. §7):**#3/#4 (pull-геттеры) и #5 (`len<2`) — чистые upstream-дефекты, независимые от lx, кандидаты в upstream-PR. #1/#2 сейчас живут в **lx-форме** (за `with_lx_command`, `// lx:`); для upstream-PR им нужна **upstream-форма** (без тега и маркеров) — это план §7, не сделано.
---
## 1. Проблема / контекст
LxBox переехал с Clash API на нативный CommandClient ([SPEC 014](../014-CLASH_API_TO_COMMANDCLIENT_MIGRATION/SPEC.md): `with_clash_api` удалён в `v1.14.0-lx.1-rc.1`). Но нативный CommandClient **беднее** Clash API сразу по нескольким возможностям, без которых UI неполноценен:
- **Per-node delay-тест.** Существующий `rpc URLTest` ([daemon/started_service.go](../../daemon/started_service.go)) принимает только **группу** (hard type-assert на `adapter.OutboundGroup`, отказ «outbound is not a group»), меряет захардкоженным `https://www.gstatic.com/generate_204`, без таймаута, результат уходит только в стрим. Clash `/proxies/{name}/delay?url=&timeout=` ([experimental/clashapi/proxies.go:187](../../experimental/clashapi/proxies.go)) умел: одиночный узел, произвольный URL, таймаут, синхронный ответ.
- **Таблица правил.** Clash `/rules` ([experimental/clashapi/rules.go:24](../../experimental/clashapi/rules.go)) отдавал `router.Rules()` как `{type, payload, proxy}`. В CommandClient аналога нет.
- **Pull-снапшот групп/узлов.** Clash был **pull**: `GET /proxies` — дёрнул в любой момент, получил снапшот. CommandClient — **push-only**: группы приходят лишь стримом `SubscribeGroups`/`SubscribeOutbounds`. Если стартовый push не доехал (стрим не открылся из-за гонки фаз сервиса) — перезапросить нечем (см. §3.3).
- **Баг потери групп.** `readGroups()` молча выбрасывает группы с 0–1 узлом (`len<2`) — нарушает Clash-паритет, прячет одно-узловые селекторы (см. §3.5).
Движок и хранилища **уже всё умеют** — `urltest.URLTest(ctx, link, detour)` принимает URL и `ctx`-deadline; `adapter.Router.Rules()` публичен; `s.readGroups()` существует. Не хватает только **проброса через gRPC** (§3.6 «мост»: проброс существующей возможности ядра, не новая подсистема).
По CONSTITUTION §3.1(а) фича легальна: (а1) нужна LxBox; (а2) недоступна в нашем канале — была в upstream лишь через вырезанный Clash API (либо отсутствует как pull); (а3) свой дифф (швы в `.proto`/интерфейсы + handler'ы в `_lx.go`) дешевле, чем вернуть весь HTTP-сервер `with_clash_api`.
---
## 2. Общая инфраструктура (§3.6 класс)
### 2.1 Build-tag `with_lx_command` (§3.6 п.3)
- Реальные handler'ы — за `//go:build with_lx_command`; файл-близнец `*_stub.go` за `//go:build !with_lx_command` возвращает `codes.Unimplemented` (образец: `daemon/started_service_usbip{,_stub}.go`).
- Без тега сборка поведенчески эквивалентна upstream (новые RPC не обслуживаются). Регистрация самих RPC в `service` сгенерирована из `.proto` всегда — гейтится только рукописный handler.
- Тег в `sharedTags` ([cmd/internal/build_libbox/main.go](../../cmd/internal/build_libbox/main.go), `// lx:`-блок) — иначе RPC не попадёт в `libbox.aar`. Для десктопа — в `LX_TAGS` (`Makefile.lx`).
### 2.2 Детерминированная регенерация proto (§3.6 п.5)
Генерация была невоспроизводима: `make proto` ([Makefile](../../Makefile)) шеллит системный `protoc` из PATH, `proto_install` ставит `protoc-gen-go@latest` / `protoc-gen-go-grpc@latest`.
-`*.pb.go` / `*_grpc.pb.go` — машинный вывод, руками не правятся, маркеров не несут; на ребейзе **регенерируются** из смерженного `.proto`, не мёржатся текстом.
- **Цена pinned-тулчейна (audited):** committed `.pb.go` исторически сгенерированы иным тулчейном, чем пин. Первая регенерация под пином — помимо `daemon/started_service.*` — косметически переписывает соседние generated-файлы того же `make proto`-набора: `daemon/managed_service.{pb,_grpc.pb}.go`, `experimental/v2rayapi/stats.{pb,_grpc.pb}.go`, `transport/v2raygrpc/stream.{pb,_grpc.pb}.go` (`(Enum)(0)`→`Enum(0)`, `status.Error`→`Errorf`, import-order). Разовая нормализация; дальше воспроизводимо.
### 2.3 CI-инвариант (§3.6 п.7)
В`lx-ci.yml` — обе сборки: **без**`with_lx_command` (компилируется, `*_stub.go` отдаёт `Unimplemented`, поведение = upstream) и **с** тегом (RPC обслуживается). Usbip-паттерн делает проверку дешёвой.
---
## 3. RPC
### 3.1 `URLTestOutbound` ✅ rc.2
Шов в [daemon/started_service.proto](../../daemon/started_service.proto), под маркером:
- **Резолв тега в ОБОИХ менеджерах:** сначала `boxService.outboundManager.Outbound(tag)`, при промахе `boxService.endpointManager.Get(tag)`. `adapter.Endpoint` встраивает `Outbound`/`N.Dialer` → endpoint передаётся в `urltest.URLTest` без обёрток. **Никакого** type-assert на `OutboundGroup`.
- **Модель ошибок — ВСЁ в payload (Вариант B), `status.Error` не используется для прикладных сбоев:**
| Ситуация | `delay` | `error` |
|----------|---------|---------|
| успех (в т.ч. 0 мс) | latency | `""` |
| узел не найден | 0 | `"outbound or endpoint not found: <tag>"` |
| тест провален (timeout/dial/bad-status) | 0 | `err.Error()` |
Handler **всегда** возвращает `(resp, nil)` — транспортный gRPC-error остаётся `nil`.
- **ИНВАРИАНТ для клиента:** источник истины — поле `error`. `delay` валиден ⟺ `error == ""`. `delay==0 && error==""` = **успех 0 мс**, НЕ ошибка (иначе воркер словит ложный фейл на быстром локальном узле).
- **История (запись):** при успехе `urlTestHistoryStorage.StoreURLTestHistory(group.RealTag(detour), {Time: now, Delay: delay})`; при ошибке `DeleteURLTestHistory(realTag)`. `RealTag(detour)` для одиночного outbound/endpoint = `detour.Tag()`.
**Куда попадает delay (карта каналов):**`Store/DeleteURLTestHistory` будят общий `urlTestObserver`, на который подписаны оба групповых стрима:
| Канал | Что отдаёт | Покрытие узлов |
|-------|-----------|----------------|
| **Синхронный ответ RPC** | измеренный delay немедленно | **любой** узел (outbound/endpoint). Единственный гарантированный канал; UI ручного пинга читает его. |
| **`SubscribeOutbounds`** | `GroupItem.UrlTestDelay`/`UrlTestTime` из истории | **все** outbound'ы И **все** endpoint'ы (WG/AWG/Tailscale). Канал, где delay endpoint'а появляется в стриме. |
| **`SubscribeGroups`** | то же из истории | **только** узлы внутри `OutboundGroup`. Одиночные outbound вне групп и **любые endpoint'ы НЕ попадают**. |
Клиентский метод — `experimental/libbox/command_client_command_lx.go`:
**Почему не `(uint16, string, error)`:** gomobile НЕ биндит ни три возврата, ни `uint16`/`uint32`. Возврат — struct-обёртка (как `*SystemProxyStatus`, геттеры `getDelay()`/`getError()`); параметр `timeout` (не `timeoutMs`); типы `int32`. Возвращаемый Go-`error` — **только транспортный сбой**; прикладной исход — в `Result.Error` (Вариант B). `timeout` в мс (`0` → дефолт).
Worker-pool (масс-пинг N узлов, concurrency=10) — **в клиенте/LxBox**, не в ядре: ядро меряет один узел синхронно и stateless. Отмена — см. **§3.6**: после ctx-b'a тест привязан к gRPC per-call `ctx`, отмена одного вызова рвёт один тест (per-node, как Clash); масс-отмена = отмена N pooled-вызовов на клиенте. (Раньше тут было «отмена = закрытие conn» — это следствие бага привязки к `boxService.ctx`, исправлено в §3.6.)
- **Route-rules:** `boxService.router.Rules()` ([adapter/router.go:21](../../adapter/router.go), публичный) → `{type, payload, action, isDNS:false}`. Поля — как Clash `rules.go`.
- **DNS-rules:** через **новый геттер** (см. §3.2.1) → `{..., isDNS:true}`. `adapter.DNSRule` встраивает `Rule` → те же `Type()/String()/Action()`.
#### 3.2.1 Правка апстрим-интерфейса для DNS-rules (цена «route + DNS»)
`adapter.Router` НЕ выставляет DNS-правила; `dns.Router.rules []adapter.DNSRule` приватно ([dns/router.go:44](../../dns/router.go)), `adapter.DNSRouter` ([adapter/dns.go:18](../../adapter/dns.go)) геттера не имеет. **+2 тронутых апстрим-файла**, оба за `// lx:`-маркером:
-`adapter/dns.go` — в `DNSRouter` добавить `Rules() []DNSRule`.
**Проблема (pull vs push).** Clash был **pull**: `GET /proxies` → снапшот групп в любой момент. CommandClient — **push**: группы приходят только потоком `SubscribeGroups`. Стрим шлёт начальный снапшот первым `Send` (`readGroups()` до `select`) — НО только если открылся: `waitForStarted` ([started_service.go](../../daemon/started_service.go)) отклоняет подписку с ошибкой, когда сервис не `STARTED`/`STARTING` (`IDLE`, `STOPPING`, `FATAL` при рестарте/реконнекте). Если стрим не открылся или порвался — **перечитать нечем**: клиент вынужден пересоздавать весь `screenClient` (тяжёлый `refreshScreen`, рвущий `SubscribeConnections`). **На устройстве подтверждено:** watchdog делает 2 ретрая через `refreshScreen`, группы остаются пустыми (`tunnel=connected`, трафик идёт, но `groups=[]`, `nodes=0`). Переподписка ≠ pull.
Handler (`started_service_command_lx.go`): под `serviceAccess.RLock` проверить `serviceStatus == STARTED` (как `GetRules`; для unary честнее вернуть ошибку сразу, чем `waitForStarted` ждать перехода фазы — клиент узнаёт причину немедленно), затем вызвать `s.readGroups()` и вернуть. Тело — калька подготовки из `SubscribeGroups`, но один `readGroups()` + `return` вместо цикла с`Send`. Ошибка при не-`STARTED` — `status.Error(codes.FailedPrecondition, ...)` (unary read-конвенция, как `GetRules`; НЕ Вариант-B).
Клиент: `func (c *CommandClient) GetGroups() (OutboundGroupIterator, error)` — переиспользует тот же итератор и `outboundGroupIteratorFromGRPC`-конвертер, что `SubscribeGroups`.
### 3.4 `GetOutbounds` ✅ rc.4
Та же дыра pull vs push для плоского списка узлов. `SubscribeGroups` покрывает лишь узлы внутри групп; **одиночные outbound и любые endpoint'ы** (WG/AWG) видны только через `SubscribeOutbounds` (см. карту §3.1). Поэтому нужен **и**`GetOutbounds`, не только `GetGroups`.
### 3.5 Фикс `len<2` в `readGroups()` ✅ rc.4 — upstream-баг
`started_service.go:495`:
```go
iflen(g.Items)<2{
continue// ← группа с 0 или 1 узлом молча выбрасывается
}
```
Группа с одной нодой (частый кейс — `proxy → один сервер`), пустая группа, или группа, где часть нод не зарезолвилась (`!isLoaded`) и осталась одна — **не попадают** в `Groups`. **Upstream-дефект** (пришёл с рефактором `5bc0dfa9` «platform: Refactoring libbox to use gRPC-based protocol»). Clash отдавал `group.All()`**без** фильтра по количеству → при переезде группы с 1 узлом исчезли (видимая регрессия).
`readGroups()` — единственный источник, его зовут `SubscribeGroups` (стартовый бродкаст) И будущий `GetGroups`. Фикс здесь покрывает **оба** пути.
**Фикс:** убрать условие `len(g.Items) < 2 { continue }` — группы любого размера валидны. (Если апстрим вводил его как анти-шум — обоснования в коде нет; Clash-паритет требует отдавать все.)
**Как было в Clash (для истории).** Отдельного «cancel»-эндпоинта в Clash API **не было** — миф про `cancelDelays` не подтверждается: `experimental/clashapi/proxies.go` выставлял только `GET /proxies/{name}/delay` (unary). Отмена была **per-request и неявная**: каждый delay-тест шёл отдельным HTTP-запросом, chi привязывал `r.Context()` к жизни этого запроса, а `getProxyDelay` оборачивал его в `context.WithTimeout(r.Context(), …)` и передавал в `urltest.URLTest(ctx, …)`. Клиент рвал/абортил HTTP-запрос → `r.Context()` отменялся → `DialContext` падал. «Отменить весь масс-пинг» = клиент абортил свои N запросов; единой серверной кнопки не существовало.
**Что сломалось при переезде (баг, не дизайн).** Первая реализация `URLTestOutbound` (rc.2) привязывала тест к **долгоживущему**`boxService.ctx`, игнорируя gRPC per-call `ctx` хэндлера:
```go
testCtx:=boxService.ctx// ← живёт, пока жив ВЕСЬ сервис
Поэтому отмена вызова на клиенте (gRPC CANCELLED / разрыв стрима) **не доходила** до dial: тест переживал вызов и дотухал сам. Единственным рычагом оставалось снести **всё** соединение (рвёт и Connections/Groups-стримы). Это и породило формулировку «отмена = закрытие conn» — следствие бага, а не намеренное решение.
**Фикс (слой 1, сделано в коде).** Привязать тест к gRPC per-call `ctx` — первому аргументу хэндлера, который gRPC отменяет автоматически при отмене вызова/разрыве стрима:
Это **дословный аналог** Clash (`r.Context()` → `testCtx`): отмена одного вызова обрывает один тест на его dial/handshake, **не трогая** остальные стримы. `boxService` ниже всё ещё нужен (резолв тега, история) — лишних переменных нет. Цена — 2 строки, нулевой риск. Тронут только `daemon/started_service_command_lx.go` (handler) + комментарий в `experimental/libbox/command_client_command_lx.go`.
**Отмена масс-пинга (слой 2, на клиенте) — через отдельный ping-client.** Уточнено по фидбэку LxBox: gomobile-биндинг **не экспонирует per-call cancel** — у`urlTestOutbound(tag, link, timeoutMs)` нет cancel-параметра/handle, а Go-`CommandClient` держит **один**`c.ctx` на все вызовы инстанса (`getClientForCall` → общий `c.ctx`, один `c.cancel`). Поэтому «per-call `call.cancel()`» **нереализуемо** на текущем биндинге — единственный рычаг caller'а — `Disconnect()`.
Решение **без правок биндинга**: держать **отдельный `CommandClient`-инстанс под масс-пинг** (`pingClient ≠ statusClient/screenClient`). Каждый `NewCommandClient` поднимает свой `c.ctx`/`c.cancel`/`grpcConn`, поэтому `pingClient.Disconnect()` (`c.cancel()` + `grpcConn.Close()`) отменяет **только** per-call ctx ping-тестов, не трогая другие стримы. Цепочка проверена: disconnect → gRPC рвёт серверные stream-ctx этого conn → `testCtx` (слой 1) отменяется → `urltest.URLTest` обрывает `DialContext`/`client.Do`**до**`C.TCPTimeout` (оба ctx-aware). Это закрывает «зомби»-тесты (до ~concurrency in-flight dial'ов), которые epoch-гейт сам по себе не гасил (он гасит применение результата в UI, не серверный dial). При отмене: epoch-бамп (UI) + `pingClient.disconnect()` (ядро) + свежий pingClient под следующий прогон.
**Серверный batch-`CancelURLTests` (отложено, см. §5).** Единый серверный RPC «отменить весь пинг одним вызовом» по токену сэкономил бы N cancel'ов по сети, но требует **stateful-реестра** в ядре (`map[token][]CancelFunc` + мьютекс + жизненный цикл токенов) — категория «новой подсистемы», которую §3.6 просит избегать («только мост»). Выгода (микросекунды против stateful-слоя) не оправдывает дифф к upstream. Отклонено **с условием**: вернуться, если профайл покажет, что N клиентских cancel'ов реально дороги. Не нужен и для масс-отмены: `pingClient.Disconnect()` гасит весь батч одним вызовом без серверного состояния.
**Статус слоя 2.** Разблокирован на текущем биндинге (вариант #2: отдельный ping-client + `Disconnect()`); правок нативной поверхности не требуется. LxBox-фидбэк (`CLIENT_FEEDBACK_urltest_cancel_binding.md`) закрыт. Остаётся on-device verify на стороне LxBox: масс-пинг недоступных узлов → отмена → серверные dial'ы рвутся до `C.TCPTimeout`, прочие стримы целы (критерий §4.5a). Гранулярный per-node cancel (#1/#3) — в запасе под будущий UX, YAGNI.
---
## 4. Критерии приёмки
1. (✅ rc.2) Сборка с`with_lx_command`: `URLTestOutbound` меряет outbound И endpoint (AWG/WG); кастомный `link`/`timeout` применяются; `error`-поле по таблице §3.1; `delay==0 && error==""` = успех 0 мс.
2. (✅ rc.2) `GetRules` возвращает route- и DNS-правила с `isDNS`-разделением; route-поля совпадают с Clash.
3. (✅ rc.4) `GetGroups` возвращает тот же снапшот, что первый `Send``SubscribeGroups`; вызываем в любой момент при `STARTED`, не трогая активные стримы; при не-`STARTED` — `status.Error`с причиной.
4. (✅ rc.4) `GetOutbounds` возвращает то же, что `SubscribeOutbounds` (все outbound'ы + endpoint'ы с delay из истории).
5. (✅ rc.4) После фикса `len<2`: группы с 1 узлом видны и в стартовом бродкасте `SubscribeGroups`, и в `GetGroups`.
5a. (§3.6) `URLTestOutbound` привязан к gRPC per-call `ctx`, не к `boxService.ctx`: отмена вызова (gRPC CANCELLED) обрывает один in-flight тест на dial, прочие стримы (Connections/Groups) живы; сервис не останавливается. Масс-отмену делает клиент (N pooled-вызовов). Серверный batch-RPC не вводится.
6. Сборка **без**`with_lx_command`: компилируется, все RPC → `codes.Unimplemented`, поведение = upstream.
7.`Makefile.lx` proto-таргет регенерирует `*.pb.go` воспроизводимо; сгенерированный код gofmt-чист, без `// lx:`-маркеров.
8. CI зелёный на обеих сборках (§2.3).
9.`with_lx_command` в `sharedTags` (AAR) и `LX_TAGS` (desktop).
10. Перечень тронутых общих файлов совпадает с фактическим диффом:
- **Швы под `// lx:`:** `daemon/started_service.proto`, `adapter/dns.go`, `dns/router.go`, `cmd/internal/build_libbox/main.go`. Фикс §3.5 — правка `daemon/started_service.go` (под `// lx:` или как upstream-багфикс, см. §7).
- **Косметическая регенерация под пин (разовая, §2.2):** `daemon/managed_service.*`, `experimental/v2rayapi/stats.*`, `transport/v2raygrpc/stream.*`.
- **Логика — в новых файлах:** `daemon/started_service_command_lx.go` + `_stub.go`, `experimental/libbox/command_client_command_lx.go`. Cancel-fix §3.6 правит первые два (handler: `testCtx := ctx`; client: комментарий) — те же файлы, без новых.
---
## 5. Вне скоупа
- Отдельный history-RPC (`GetURLTestHistory`) — отклонён: delay синхронен в ответе RPC, живая история течёт в `SubscribeOutbounds`/`SubscribeGroups`.
- Серверный batch-`CancelURLTests` RPC — **отложен** (§3.6): per-node cancel уже даёт gRPC call ctx (слой 1), масс-отмена — N вызовов на клиенте (слой 2); серверный stateful-реестр токенов противоречит §3.6 «только мост». Условие возврата — профайл-доказательство, что N клиентских cancel'ов дороги. (Раньше формулировка была «per-call cancel / batch отклонены: отмена через close conn» — устарела: close conn был следствием бага §3.6, не дизайном.)
-`timeout` в микросекундах / смена типа delay — отклонены: мс, `uint16`→`uint32`.
- DNS rule-set (headless) snapshot, server/inbound RPC — будущие SPEC.
- **Чистые upstream-дефекты, независимые от lx — кандидаты в upstream-PR как есть:**
- **#5`len<2`** — баг их же рефактора `5bc0dfa9`, ломает Clash-паритет. Однострочный фикс, подаётся напрямую.
- **#3/#4 pull-геттеры** (`GetGroups`/`GetOutbounds`) — структурная дыра их протокола (push-only, нет unary-аналога `Subscribe*`). Полезны любому клиенту CommandClient, не только нам.
- **lx-форма vs upstream-форма (#1/#2, #3/#4):** в нашем дереве все эти RPC живут в **lx-форме** — за `with_lx_command`, под `// lx:`-маркерами, handler'ы в `_lx.go`. Это правильно для форка (§3.6). Но для **upstream-PR** им нужна **upstream-форма**: без build-tag, без маркеров, handler'ы в самом `started_service.go`, RPC безусловно в `service`. Конвертация lx-форма → upstream-форма — **отдельная будущая задача** (когда/если решим подавать PR), здесь НЕ выполняется. `URLTestOutbound`/`GetRules` уже зашиплены в lx-форме (rc.2) и менять их форму ради PR — отдельное решение.
> Практически: реализуем #3/#4/#5 в lx-форме (target `rc.4`), как #1/#2. Upstream-PR — потом, отдельным заходом, с конвертацией формы. Эта секция фиксирует обязательство отслеживать кандидатов, чтобы дифф к апстриму со временем уменьшался (CONSTITUTION §2).
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.