Backend audit fixes (upstream-file edits wrapped in // lx: markers):
- experimental/libbox oom_report.go/report.go: OOM reports + configuration.json
(server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files
→ 0o700 / 0o600. [sec-perms]
- daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret
compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare.
[sec-consttime]
- service/oomkiller/timer.go: network-extension cleanupTriggered logic was
inverted, so FreeOSMemory was never called after a trigger; flip both
assignments so a trigger schedules the deferred free and the next poll runs +
clears it. [sec-oomcleanup]
- transport/v2rayxhttp/client.go (lx-native file): session id used math/rand →
crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface.
- daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a
goroutine + the ssh-agent fd on every closed session (second io.Copy blocked
on an idle agent Read forever); tie both copies + the session ctx to a
cancel that closes both ends. [sec-sshagent]
- daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to
1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate]
- route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in
select while stopIdleSuspend niled it after close (race + goroutine leak on
Close-during-tick); pass the stop channel to the loop by value.
go build ./... (default) and the D9 shaterd linux build (tags
with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,
with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer
warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/...
./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend
and v2rayxhttp with with_xhttp.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.
lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
(PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
point every L3-forwarded packet transits, incl. established flows that
bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
sing-tun); wireguard-go stays v0.0.5 with local submodule replace.
Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
Idle-suspend frees the recv-worker bufsArrs, which are ~8MB each only where
BatchSize=128 (Android/Linux) — on desktop BatchSize is small and the feature
saves almost nothing. Make that platform scope explicit in the build instead of
running the tick everywhere.
The idle-suspend tick now compiles only with the new `with_lx_idle_suspend` tag,
baked into the mobile AAR (build_libbox sharedTags) but NOT the desktop LX_TAGS.
Without the tag, a config that sets route.lx_idle_suspend fails fast at start
("rebuild with -tags with_lx_idle_suspend (mobile-only feature)") rather than a
silent no-op. The gate is a single function: reachability_lx.go carries the tick
under the tag, idle_suspend_stub_lx.go is the no-tag stub that errors, and
reachability_common_lx.go keeps InvalidateReachability (needed by the group
interface in every build). The dial hot path (resumeOnDial/stampActivity) and the
upstream group files are untouched — without the tick, idleAsleep is never set, so
resumeOnDial always takes its fast path.
Adds stub unit tests (option set → error, unset → no-op). Both build variants and
the full route/wireguard/group suites are green; gofmt/vet clean; desktop lx-check
passes without the tag. Docs (lx-config.md + ru, SPEC.md §3/§10) describe the tag.
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.
- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.
Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
DNS attribution was empty (0/119 on device): TUN+DNS hijack returns on a fast-path
(route.go:91/226) BEFORE matchRule, and searchProcessInfo — which fills
metadata.ProcessInfo — lives inside matchRule (:416). So fast-path DNS (most DNS on
a VPN) reached the SubscribeDNSQueries emit with nil ProcessInfo. Fix: call
r.searchProcessInfo(ctx, &metadata) before both fast-path hijacks (stream+packet);
idempotent + cached, one lookup per flow. Corrects SPEC 018 пункт 3 (the earlier
'cached attribution correct' claim checked ctx consistency, not that ProcessInfo
was populated before the resolve).
Also: DnsAnswer.rdata was the full RR string ('google.com. 29 IN A 1.2.3.4'); strip
the header prefix so clients get the bare value ('1.2.3.4' / CNAME target).
No proto/wire change. LxBox §180 needs no client change. Changelog rc.9.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.
Full 1.14 migration, step 1 of 2 (sing-box repo layer). Three conflicts
resolved, all as predicted by the feasibility analysis:
- route/rule/rule_item_package_name_regex.go (add/add): took upstream's
canonical version (slices.ContainsFunc) — our lx.15 backport collapses
back into upstream, so the file no longer diverges going forward.
- route/rule_conds.go: kept our package_name_regex in isProcess{,DNS}Rule
and took upstream's new isNeighbor{,DNS}Rule additions.
- cmd/internal/build_libbox/main.go: kept lx with_xhttp/with_awg append and
the no-tailscale block; deliberately dropped upstream's new with_usbip
(server-side USB/IP, contradicts client-trim).
go.mod auto-merged: wireguard-go require bumped to v0.0.3, lx replace block
(=> ./submodules/wireguard-go) preserved. Submodule pointer unchanged here —
the AmneziaWG graft rebase onto v0.0.3 is step 2 (next commit). This commit
does NOT build yet (submodule still on the old wireguard-go base).
Backport upstream 1.14 feature 941ce58b onto the 1.13.13 base without the
full migration. Adds the package_name_regex rule item (regex match over
ProcessInfo.AndroidPackageNames) to route, DNS and headless rules.
- new route/rule/rule_item_package_name_regex.go (verbatim upstream) + unit test
- PackageNameRegex option field in RawDefaultRule/RawDefaultDNSRule/DefaultHeadlessRule
- item registration in NewDefault{,DNS,Headless}Rule with E.Cause(err, package_name_regex)
- package_name_regex added to isProcess{,DNS,Headless}Rule conds
The commit's RuleSetVersion5 hunk is intentionally NOT ported (that is 1.14
rule-set v5, unrelated; base is RuleSetVersion4). Full 1.14 migration deferred
to v1.14.0 stable. SPEC 013 + Roadmap entry.
builds (no-tags + lx-tags), go vet, gofmt and rule tests all green.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.