The carrier behind this router's SIM refuses TCP/443 to 9.9.9.9 and 1.1.1.1 while
carrying everything else — measured with a positive control (ya.ru:443 and
77.88.8.8:53 connect, every sim-bypass node connects, those two are refused). The
configured resolvers go out DIRECT, not through the tunnel, so on that uplink DNS
resolved nothing: the vless server names did not resolve, the hop in front of
awgout never came up, and the whole chain died with it. One pair of global scalars
cannot be right for two uplinks; the object that knows which uplink is live is the
profile.
* config profile gains resolver_default, resolver_fallback and fetch_detour
beside endpoint_resolver. Empty = inherit, PER FIELD.
* globals.fetch_detour replaces `const filterFetchDetour = tagDirect`. Behind a
carrier whitelist `direct` is not the safe path, it is the path where the
source is refused forever and the list never loads.
* A subscription's fetch_via becomes an OVERRIDE, which gives it a third state.
ReadUCI used to parse an absent option as the literal "direct", so "chose
clear-text" and "never touched this row" were the same value. migrate2to3
performs the reinterpretation ONCE, in the open. Schema 2 -> 3.
* An unusable override falls back (resolvers to globals, fetch_detour to direct)
and says so at critical, naming profile, field, value and what is in force.
The panel was displaying globals while the engine used the profile's value; the
owner caught it. The field now keeps the STORED value with a separate line naming
what is in force, and the rule that answers "what is in force" moved to the daemon
(GET /api/config/effective) so it stops existing in two languages.
Cold start, by owner's requirement: rule-sets are read from the cache when the
source is unreachable instead of being dropped, and the subscription cache reader
is fixed. Its first fix was wrong and only Linux said so — mtime ties to the digit
because the kernel caches the stamp per tick, and this board has no RTC, so the
ordering can invert across a reboot. Replaced by a generation counter in the file.
Woke and closed a LAN-dark defect: wgdedup read only the deprecated, always-empty
DownloadDetour, never HTTPClient.Detour, so fetch_detour=node:<awg> made a node
used, the dedup pass did not know, merged it away, and left the rule-set pointing
at a tag box.Start could not resolve. Reproduced through a real box.New.
Also: ValidateProfiles had no caller; "applied from the cache" graded critical
though the list is in force; the auth matrix never walked /api/log or
/api/rules/reachability; the CLI and daemon disagreed about where a subscription
is fetched.
NOT fixed, stated rather than implied: the R5 preflight still probes direct, so a
list never yet fetched cannot bootstrap over the detour alone; the router's own
DNS on the SIM stays dead (dnscrypt-proxy bootstraps via blocked addresses).
Gate: bash scripts/run-tests.sh green, 7/7, privileged tests really ran.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The line naming a nonzero `go test` status was printed only when every
privileged test had produced a verdict — on the reasoning that a named FAILED
already explains the status. The case that actually happens is the opposite
one: the run dies at package level, so it names no test, so the loop above
prints MISSING for all of them, and the one line pointing at the real cause was
the one suppressed. A reader then goes hunting for three vanished tests instead
of at the build error above.
To be exact about what was and was not broken, because the framing matters: the
exit status was never SWALLOWED. priv_bad is set by the MISSING branch, so
FAILED is set and the gate fails either way — this was a diagnosis bug, not a
correctness one. What changes is whether the log says why.
Verified on the branch a green run never reaches, by driving the edited block
with all four (priv_rc, priv_bad) combinations: the new message appears only for
(1,1), the old one only for (1,0), and priv_bad/FAILED come out 1 in both. The
full gate is green with the change in, which covers the (0,0) path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
CI's `[4/7] go test -race -shuffle=on` failed the v0.2.23 gate on
TestRefreshObservatoryForcesOnePass ("force flag survived the forced pass").
The product is NOT at fault, and this was established rather than assumed.
WHAT ACTUALLY BROKE. ConfigureObservatory starts the ticker goroutine and its
first tick fires immediately — by contract, so an applied config gets its first
verdicts in seconds — and a plan change additionally nudges the loop into a
pass on purpose. That tick advances the cursor and consumes the force flag.
Two tests then read exactly those fields straight after a Configure, i.e. read
values another goroutine is entitled to rewrite in the same instant. Four
assertions, all racy:
observatory_test.go:78 identical-plan reconfigure reset the cursor to 2
observatory_test.go:86 changed-plan reconfigure kept the cursor at 2
observatory_test.go:176 after refresh: cursor=2 force=false
observatory_test.go:186 force flag survived the forced pass
The last one is the busy guard: with the loop's first tick still in flight the
test's hand-driven observatoryTickOnce is a silent no-op, so nothing clears the
flag it just raised.
NOT a cross-test dependency, and not a leaked goroutine — the direction was
measured, not guessed. Each test reproduces ALONE in the CI container at
`-count=3000`: 16/3000 and 7/3000, with all four messages. The earlier
`-count=80` in isolation was simply too few iterations; a loaded `-shuffle=on`
package run widens the window, which is why CI saw it and a laptop did not.
THE FIX is isolation, not a weakened assertion. detachObservatoryLoop stops the
goroutine and leaves a PLACEHOLDER stop channel behind, so the reconfigures
these tests make still run the whole state machine — plan rebuild, cursor
policy, nudge — with no second writer (ConfigureObservatory starts a loop only
when e.obs.stop is nil; e.obs.nudge is left nil and every send to it has a
default). quiesceObservatoryLoop, which four chain tests already used for the
same reason, is now that plus a cursor rewind.
Mutation-checked: with detachObservatoryLoop neutered the flake returns at
18/3000 and 6/3000 with the same four messages; restored, 20 consecutive
`-race -count=1 -shuffle=on` runs of the package are clean, as is the full
`scripts/run-tests.sh`.
TWO TESTS GAINED THE ABILITY TO FAIL. TestObservatoryTickStoppedEngine and
TestObservatoryTicksDuringManualRun assert `cursor != 0` after a hand-driven
tick — which the loop's own first pass had already satisfied for them, so they
held whether or not the tick under test did anything. The second one is the
worse case: it exists to forbid the tick deferring to a manual run, and the
busy guard could make the tick do nothing while its assertion still passed.
Both now quiesce first.
TWO NEW TESTS, for the contract the flake kept stumbling into without ever
asserting it — a refresh raised while a tick is in flight:
- TestRefreshDuringInFlightTickRunsAFullForcedPass parks the loop's first
pass inside a stub probe, so "in flight" is a fact rather than a hope,
raises force there, and requires a second full pass over jobs the polite
freshness gate would skip. TWO independent wakeups carry the request across
— the buffered nudge and the tick's deferred re-nudge — and that is
measured: disabling EITHER leaves the test green, disabling BOTH makes it
fail with "force is still raised" and 2 attempts instead of 4. So it
asserts the observable contract, not a mechanism, and says so.
- TestForcedPassChainsItsBatchesWithoutWaitingForTheTick pins what
observatoryTickOnce's defer claims and nothing held: a forced pass chains
its batches instead of spending a 10s tick each. THREE batches, because two
prove nothing — the loop's unconditional first tick pays for one and the
refresh's still-unconsumed nudge pays for the second, so a two-batch plan
finishes even with the chaining removed. Measured that way round first;
at three, removing the defer leaves 48 of 54 targets undialled.
obsSelectorFixture/obsWideSelectorFixture exist because obsFixture's urltest
members are SelfChecked and the observatory does not dial them at all — a stub
waiting on that plan would hang, not fail.
No product file is touched: shater/engine/observatory.go is byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Measured on the production router: a `config ruleset` of type=ipcidr holding
10.10.10.0/24, a rule pointing it at node:awghome, config_applied=true,
tunnel_rules=1, engine_running=true, ZERO warnings — and from a LAN client,
100% packet loss and no TCP. The rule was accepted, applied, reported healthy,
and could not fire.
The cause is one line of ordering. `ip daddr { 10.0.0.0/8, 172.16.0.0/12,
192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, ... } accept` sits ABOVE every
divert line in the prerouting chain, so the packet is accepted and handed to
plain routing before the engine — which holds the rule — ever sees it. That
default is right and stays: LAN-to-LAN, the router's own services and every
local plane must not be dragged through a tunnel, and a catch-all rule must
never quietly acquire them. What was wrong is that naming a subnet OUTRIGHT
could not override it, and that nothing said so.
So the divert for NAMED private destinations is emitted one line higher, and
"named" is deliberately narrow (netplane/coverage.go, privateRoutedPlan):
- the CIDR must be an ENTRY of an INLINE type=ipcidr rule-set — the only
destination list this stage can read;
- it must be CONTAINED in 10/8, 172.16/12 or 192.168/16. A prefix that merely
overlaps one (0.0.0.0/0, 10.0.0.0/7) is a catch-all that happens to include
private space, and does not acquire it;
- the referencing rule must be enabled and target node:/group:/chain:/egress:
or block. `direct` is not an override: it asks for what the bypass already
does, and diverting into the engine to reach the same verdict would be
strictly worse, because the engine's direct outbound follows the DEFAULT
route and LAN-to-LAN could be pushed out the WAN;
- it must not overlap a network this router itself carries;
- 127/8, 169.254/16, 224/4 and 255.255.255.255 are never taken.
THE SELF-AMPUTATION GUARD DISTINGUISHES A LAN FROM AN UPLINK, and that
distinction is the difference between a safety device and an obstacle. A
collision with one of our OWN networks (any zone that is not a WAN zone, plus
any interface whose zone is unknown) is refused by name — diverting it takes
the LAN away from the LAN and the operator finds out over the console. A
collision with an UPLINK subnet routes and discloses: ISPs hand out RFC1918
WANs routinely — this router's own gateway is 10.0.0.1 — and on a /8 uplink
every private subnet on earth "collides", so refusing there would disable the
feature on precisely the routers that want it, for a reason that would read as
a bug. Nothing of ours lives on the uplink subnet: `fib daddr type local`
already accepts the router's own addresses above these lines, and every divert
line is scoped to LAN ingress, so router-originated traffic never meets them.
PING IS HOW ANYONE CHECKS A ROUTE, and a TPROXY divert carries TCP and UDP
only — the kernel needs a socket and ICMP has not got one. Stopping there would
rebuild this same defect one protocol down: TCP succeeds, ping reports 100%
loss, and the operator concludes the route is broken. So with l3_tunnel on, the
L3 mark is stamped on ICMP bound for these destinations (again above the
bypass, which is the only reason it was not already happening) and the existing
`ip rule` delivers it into the engine's TUN, where the SAME route rules pick
the outbound and a WireGuard/AmneziaWG one carries it. The forward chain's
fail-closed drop excludes that mark, because unlike the tproxy legs the LAN-to-TUN
leg really does traverse forward and the `oifname "shater-l3*"` accept that
would rescue it sits four steps lower. With l3_tunnel OFF nothing is emitted,
nothing is claimed, and the rule is told so by name.
THE DOUBT ALWAYS FALLS BACK TO THE BYPASS. Failing to route a named subnet
costs a feature and shows up the moment it is tested; routing one we should not
have touched can take the router's own management network into a tunnel that
may not even be up. So an unreadable list, an inventory we could not enumerate,
and an address family we cannot check the router's own addresses in (IPv6 —
`ubus call network.interface dump` reports IPv4 only) all resolve to "leave it
on the bypass", and every one of them says so. Seven distinct sentences now
exist where there was silence: refused-for-our-own-network, refused-for-no-
inventory, reserved space, catch-all-does-not-acquire, IPv6-not-checkable,
uplink-overlap-disclosed, and ping-does-not-reach-with-l3_tunnel-off. The
eighth is the blind spot itself: an address list this plan never reads
(url/file type=ipcidr, or geoip whose category is not an ISO country code)
might contain private destinations, and that is disclosed unconditionally —
"warn on suspicion" is not available, because suspicion would mean reading the
list. It is graded `warning` rather than critical through a named marker in
apply/warnings.go: it describes a maybe, and a red that means "probably fine"
is how the next red stops being read.
generate.ruleSetTypeIsIPCIDR now delegates to netplane.IsIPCIDRRulesetType.
Two packages asking the same question of the same field must not each carry
their own list of spellings.
VERIFIED
- `bash scripts/run-tests.sh` green in full ("OK: the shipped tag set, on
linux, passes every test we own", exit 0), with the three privileged
^TestIntegration tests RAN by name.
- Every new test mutation-checked: 17 reverts, each failing the test that
covers it, by name.
- BOTH CONTROLS. Without an explicit naming, private space is still bypassed
(TestPrivateDestinationBypassIsStillTheDefault) and a catch-all still does
not take it; with it, the divert appears above the bypass. A test green in
both states would prove nothing.
- BYTE-FOR-BYTE. Two goldens, plain and L3, captured from a git worktree at
the PARENT commit — not from this code, which would only prove
self-consistency. A config that names no private subnet renders the
identical text, so the applier's idempotence check still sees no work.
- REAL NFTABLES. The rendered plane (both the tproxy and the ICMP/L3 shapes)
loads with `nft -f` on nftables 1.0.9 and the kernel holds the lines as
written; the instrument was shown able to REJECT a deliberately broken copy
of the same file.
NOT VERIFIED
- Nothing here has been run on the testbed or the router. Whether the packet
that now reaches the engine actually comes out of awghome is the owner's
acceptance test, not this commit's claim.
- Whether a named IPv6 ULA could be handled safely was not investigated
beyond establishing that the inventory cannot check it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
D29 removed byedpi, so the feed carries three packages. sdk-build-apk.sh was
changed to >=3; build-feed-apk.sh still demanded >=4 and killed both arch lanes
of v0.2.22 with `expected >=4 .apk … found 3`. Nothing was published from that
run. The comment now says the count is duplicated, because reading one script
was what made this look done.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The `byedpi` egress kind, the `openwrt/byedpi` package (`ciadpi`), the readiness
endpoint and the panel plate are gone. D13 is not deleted from DECISIONS.md; it
is REVERSED there, with the reason, because the reason is the whole point.
D13 adopted an external desync process on an observation: the engine's own
`tls_fragment`/`tls_record_fragment` were tried against a live ISP and did not
get through, so the method was judged too weak for anything past "just fragment
the ClientHello". The method was never tried. `common/tlsfragment` dropped a
number of labels equal to the number of DOTS in the name, and a name always has
one more label than it has dots — so the cut always landed inside the FIRST
label. `www.youtube.com` was split inside `www` and `youtube` went to the wire
in one piece, which is the word the DPI matches on. Of six blocked names exactly
one got through: `youtube.com`, the one whose first label IS the blocked word.
That defect is fixed (815011dfb, efb2177f4). With it fixed the built-in presets
do the job the external process was brought in to do, and the process is 100 KB
of binary, a second procd service, a second UCI file, a port that agreed with
our egress by hand-written comment only, a readiness prober, a five-state
service model and a panel plate — all to work around fifteen lines of ours.
So this is not "ByeDPI turned out to be bad". It is a good tool that turned out
not to be needed, and the reason we thought it was needed was ours.
A CONFIG THAT STILL SAYS `type 'byedpi'` IS THE PART THAT NEEDED WORK. Nothing
is migrated and nothing is rewritten: the kind stays unbuildable, therefore
fail-closed — no outbound, no mark, no `ip rule`, no routing table, so every
node, group and rule bound to it is blocked rather than released onto the plain
WAN. A migration to `direct` was considered and rejected: it is the only rewrite
that leaves the egress routing at all, and it would silently turn a blocked
egress into a live plain-WAN path with the router's real address — by an
upgrade, on a config nobody touched. `CurrentSchemaVersion` is therefore not
bumped either: no stored field changes meaning, and a bump would only make this
build's configs unreadable to an older daemon for no gain.
What changes is what the operator is TOLD. `model.RetiredEgressTypes` is a
closed, positive table read by BOTH `ValidateEgresses` and the generator (one
copy of the sentence, because two copies drift). It names the removal, denies
that it is a typo, says nothing is built and that the traffic is blocked rather
than leaked, names the replacement (`direct`/`interface` with `dpi 'record'`),
refuses to promise which preset defeats a given ISP, and says `apk del byedpi`.
The generic "unknown type" is still there and still says something different, on
purpose: "we took this kind away" and "you mistyped something" send an operator
to different places, and a value that was correct on the day it was written must
not be reported as a spelling mistake. The type list stays closed and positive —
`interface`, `direct`, the alias `tunnel` — and `EgressTypeKnown` does NOT admit
the retired kind: being told it was removed and having it work anyway is worse
than either alone.
`Egress.Port` goes with the kind: no surviving egress dials anything, so the
option is no longer parsed and drains out of /etc/config/shater on the next
render, the same way the deleted per-group probe_url/probe_interval did.
Tests, verified by mutation, each failing by name:
- drop the retired branch in `ValidateEgresses` -> the retired kind is
reported as "is not one of interface/direct" and
TestRetiredEgressTypeIsReportedByTheValidator fails on both spellings;
- drop it in the generator -> "unknown type \"byedpi\"" and
TestRetiredEgressTypeIsReportedByTheGenerator fails;
- the FAIL-OPEN mutation, which is the one that matters: let `byedpi` fall
into the `direct` arm and be a known type -> four tests fail, including the
two that check no outbound is emitted. A removal that quietly starts routing
the traffic it used to block, under a reassuring message, is the failure with
the worst consequence;
- the panel half: empty RETIRED_EGRESS_TYPES -> two egressEdit tests fail.
Controls beside the claims: `interface`, `direct`, the `tunnel` alias and the
empty synonym must still resolve, warn about nothing and emit an outbound
(TestSupportedEgressTypesAreUntouched), and never-supported values — `proxy`,
`block`, `wireguard`, `byedpi2`, `bye dpi`, `sorcery` — must NOT draw the
removal sentence, which names a replacement for something that never existed.
CI and docs: the feed loses its fourth package everywhere the four were named —
`apk upgrade shaterd shater-core luci-app-shater`, in CLAUDE.md, both READMEs,
INSTALL.md, the release body and `shaterd`'s own diag bundle. The version
exception (byedpi carried upstream's version, ours come from the git tag) is
gone with it, so ci/version.sh and ci/sdk-build-apk.sh no longer have an
exception to remember and the "expected >=4 of OUR .apk" collect check is now 3.
INSTALL.md §5.3 gains the half a feed cannot do: dropping the package from the
feed does not take it off a router it is already on, so `apk del byedpi` is
written down, with what it removes and why it is safe.
Panel: 368 tests -> 339. Deleted with the mechanism they covered:
byedpiReady.test.ts, byedpiAge.test.ts, byedpiRefusal.test.ts (34 tests);
egressEdit.test.ts gains 5 for the retired-type sentence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Follow-up to 815011dfb, which fixed WHICH label is cut but left "a cut in every
candidate label" as an unconditional rule. Measured on this tree, loopback peer,
product default fallbackDelay, one ClientHello per row:
cuts tls_fragment (*net.TCPConn) tls_fragment (proxy conn) tls_record_fragment
1 502 ms 500 ms <1 ms
2 1.004 s 1.001 s <1 ms
4 2.008 s 2.002 s <1 ms
8 4.015 s 4.003 s 539 us
21 10.540 s 10.509 s 525 us
So a cut in the PACKET modes costs half a second of connection setup, and it
costs that on BOTH branches — not only on the sleep path. writeAndWaitAck sleeps
the whole fallbackDelay whenever the ACK returns inside 20 ms (its "under
transparent proxy" case), and N.UnwrapReader reaches the *net.TCPConn only when
nothing in the chain transforms the stream, which a proxy protocol conn always
does. A proxied egress — every subscription node — therefore takes the flat
500 ms branch regardless of RTT. The number of labels is chosen by whoever picked
the hostname, and a 253-byte SNI is 85 of them: ~42 s of one connection's setup,
bought from the LAN.
In tls_record_fragment nothing waits: the ClientHello leaves in ONE write, split
into more records. 21 cuts cost 525 us and 105 bytes of record headers, and
1.1.1.1 completed the handshake with the ClientHello in 22 records in the same
77 ms it took with 2. That is the mode the field measurement was taken in, and
the mode where cutting every label was always affordable.
Hence two budgets rather than one rule: 1 cut for the packet modes, 4 for
record-only — the latter not a cost limit but a shape limit, since real names
carry one to three labels outside the public suffix and a hostile one must not
turn a ClientHello into 85 records no ordinary client emits.
One cut is enough because of WHERE it goes. Candidates are now ordered, most
worth cutting first, and first is the REGISTRABLE label — the one immediately
left of the public suffix. That is what a name-based blocklist keys on
("youtube" of youtube.com, www.youtube.com and studio.youtube.com alike,
"ytimg" of i9.ytimg.com, "example" of a.b.example.co.uk), and severing it also
breaks any match on the whole FQDN, so one cut covers both matchers. It is
chosen by STRUCTURE, from the public suffix list — not by length, which is the
same trap from the other side: in cdn-static-assets.youtube.com the longest
label is not the blocked one. The rest follow longest-first, on the argument
that among labels with no structural ranking a long one is likelier to be a
distinctive token than "www", "m" or "tv"; they are reached only when the budget
allows more, or when the registrable label is too short to cut.
The offset now comes from the label's MIDDLE THIRD. Every interior offset severs
the label, but one byte in leaves "outube" of "youtube" and a matcher keyed on a
substring still reads it. The draw stays random inside that third: a fixed point
would be a constant a middlebox vendor can special-case in one line, and this
whole family of tricks lives on making reassembly the only counter.
Also in this commit, and the reason it is not merely a tuning change: the panic
that shipped in v0.2.21 now has an instrument of its own.
TestWriteDoesNotPanicOnAServerNameChosenFromTheLAN drives real ClientHellos
carrying ".youtube.com", "youtube.com." (a legitimate FQDN with the root dot,
which curl and every browser will send), "..", an IP literal and non-ASCII bytes
through all three modes, and FuzzCutOffsets does the open half — 25.7 million
executions found nothing, and the fuzzer is shown able to find a planted defect
its seed corpus cannot reach, in one second. A hand-built ClientHello reaches
the shapes crypto/tls refuses to emit: a zero-length name, a 253-byte name, and
a server_name_list with a SECOND entry, which is why planning runs on
MyServerName.Length rather than on everything left in the extension.
Nine mutations, each failing by name with the numbers: the old dot arithmetic,
the old rand.Intn offset, the exact original expression (panic: invalid argument
to Intn, conn.go:208 <- Write conn.go:67), the empty-plan guard, the budget, the
priority order, the sort back into wire order, the first-entry truncation, the
middle third, and a one-byte corruption of a segment.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`splits[:len(splits)-strings.Count(serverName.ServerName, ".")]` is identically
`splits[:1]`: labels are always one more than dots, so the subtraction cancels
for EVERY name in existence. One label was ever cut, and it was the leftmost
one. On the provider measured from this router — which blocks by the name in
the handshake, proved by the same address answering for SNI www.google.com and
going silent for www.youtube.com — that is the whole observed table:
youtube.com cut inside "youtube" -> 301
m.youtube.com cut inside "m" -> blocked
tv.youtube.com cut inside "tv" -> blocked
www.youtube.com cut inside "www" -> blocked
music/studio.* cut inside the label in front -> blocked
The one name that worked is the one whose first label IS the blocked word. The
count subtracted must be the labels of the PUBLIC SUFFIX, not the dots of the
whole name: "com" is one, "co.uk" and "com.br" and "pp.ru" are two.
Second half of the same defect, and the reason the table above shows a cut
"inside m" at all: the offset was `rand.Intn(len(label))`, whose 0 is the
label's own boundary — the label goes out whole in the next segment, which is
not a cut, it is a segment boundary that happens to touch a label. For a
one-byte label 0 is the ONLY value it can take. Offsets are now drawn from
[1, len-1], so a cut always leaves a non-empty piece of the label on both
sides, and a label too short to have an interior offset carries no cut instead
of a fake one. That also closes the 1-in-7 hole in the case that WAS working:
youtube.com drew offset 0 once every seven connections and handed the name over
intact.
Two panics went with it, both reachable from the LAN, because route/conn.go
wraps the outbound with this and the ClientHello it fragments is the client's:
an empty label (SNI ".youtube.com" or the perfectly ordinary FQDN
"youtube.com.", where the suffix list declines to answer and the trailing empty
label survives) reached rand.Intn(0) — "panic: invalid argument to Intn", the
daemon and with it the router's proxying. And a plan with no cuts at all would
have indexed b[:splitIndexes[0]] on an empty slice; Write now writes the
ClientHello unchanged in that case, which is the only honest thing to do for a
name of one byte.
The classification is closed and errs toward MORE cutting: narrowing the label
set needs proof (a public suffix that really is a tail of the name), widening
needs none, so a trailing dot, an unmanaged TLD, a name that IS a public suffix
("com", "co.uk", "localhost") and an IP literal all keep every label rather
than fall silently into "cut nothing". When no label is long enough to cut, the
name itself is cut once — a matcher looking for the whole FQDN still fails
across that split.
Dropped with it: `splits[0] == "..."`, unreachable since strings.Split on "."
cannot produce a token containing a dot. And the plan now runs over the FIRST
entry of the server_name_list (MyServerName.Length) instead of everything left
in the extension, so a second entry cannot be fed to the public suffix list as
if it were part of the name.
Tests (cutplan_test.go, package-internal so the plan itself is visible) are
verified by mutation five ways: the old dot arithmetic, the old rand.Intn
offset, the removed empty-label guard, the removed empty-plan guard, and a
one-byte corruption of a segment. Each fails by name and with the numbers. The
controls: youtube.com — the case that already worked — must still be severed;
the reassembled segments must be byte-identical to the ClientHello in all three
modes (tls_fragment, tls_record_fragment, both), with the record framing
re-parsed rather than assumed; and Write must report len(b).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
GET /api/status no longer carries the readiness report — the daemon dropped it
with the cache behind it, after one probe was measured at 6.4 s on 16 enabled
instances behind a black hole while the panel polled that endpoint every 5 s
from every open tab and read the field NOWHERE. The Status type, the mock
fixture and every comment describing a cache, a background refresh or a 20 s
staleness rule now say what the daemon does: one endpoint, and it connects when
a human asks.
`disabled` covers two situations with opposite next actions: no instance is
enabled — how the package ships — and an instance that IS written and looks
enabled while /etc/init.d/byedpi refuses it (`port 'auto'`, `port '99999'`,
`enabled ' 1'`, `enabled 'TRUE'` — all four measured on the 25.12.1 testbed
against validate_data). The editor's fixed sentence said "that is how the
package ships" about a section the operator had typed themselves. The daemon
keeps its `problems` list off the wire, so `detail` is the ONLY carrier: the
refusal now shows that sentence verbatim plus a tail that says only what is
true of both — the consequence, never the fix.
byedpiRefusal moves to byedpiReady.ts beside the gate it explains, and its
table now EXCLUDES `disabled` from the type, so re-adding a fixed sentence for
it does not compile. `?mock&byedpi=rejected` reaches the second case in a
browser; `?mock&byedpi=noanswer` reaches "nothing has been measured", which is
now only a failed fetch — the fabricated cold-cache body is gone.
Also: two comments about `config_applied` that the daemon's pointer+omitempty
change made false — the removed "positively phrased so a naive client falls the
alarming way" rationale, and "absent means a daemon too old", which now also
means the offline `shaterd status` stub.
Tests (byedpiRefusal.test.ts, +10) verified by mutation both ways: a fixed
"that is how it ships" and a fixed "your typo" each fail, and the control
asserts the factory state still reads as the factory state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`config_applied: false` means "/etc/config/shater was read and REFUSED — what is
running is the PREVIOUS configuration, your edit is not in effect", and the panel
draws a critical band saying exactly that. The field was a plain bool, so that
alarm was the ZERO VALUE OF THE TYPE — and `shaterd status`'s offline stub, built
by a process that never applied anything, over a data plane that may have been
installed and enforcing for weeks, published it by simply never mentioning the
field. It is the config_readable defect returning in a new field, with the one
difference that decides the fix: config_readable can be MEASURED by the stub and
now is, while this one cannot be measured at all without a daemon.
So the field says nothing when nobody measured it. ConfigApplied becomes a *bool
with omitempty; the live Applier.Status() assigns a verdict on BOTH arms, so an
absent key can only come from something that is not a live status. That is the
same closed-set-plus-unknown shape `plane`, `traffic` and `daemon_answered`
already have, and the one panel/src/appliedConfig.ts already implements
(=== true / === false / else unknown). The Go doc claiming absence should read as
false is gone: it contradicted the only consumer, and the consumer was right.
The four fields around it (apply_error, apply_error_stage, apply_attempts,
apply_failed_since_unix) stay plain: they are qualified by config_applied the way
enabled/kill_switch/panel_port are qualified by config_readable, and their zero
values point at "nothing was refused" — the quiet side, not the alarm.
Also: the stub shipped `warnings: null` on its happy path while apply.Status
documents Warnings as always non-nil so a consumer can map over it
unconditionally.
Three states, distinguishable ON THE WIRE through one `shaterd status`, with the
control that would catch the opposite break (a build that omitted the key for a
real refusal, deleting the alarm from the product).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two defects found by looking at both sides of the byedpi readiness check at once.
1. GET /api/status carried the whole readiness report from a cache that a poll
refreshed in the background once the copy passed byedpiRefreshAfter = 3 s.
The panel shell polls that endpoint every 5 s, so EVERY poll started a
refresh: a PATH lookup, a read of /etc/config/byedpi, and one connect per
enabled instance, forever, per open tab, hidden ones included. The design
note rejected a background ticker because "a closed panel costs nothing" —
true, and silent about the open one it had become.
Measured, one enabled instance, twelve polls five seconds apart:
before 12 connects, 13 ciadpi PATH lookups per minute per tab
after 0 connects, 12 PATH lookups (one per poll, for byedpi_installed)
And nothing read it: `grep -rn '\.byedpi\b' panel/src` finds no consumer —
the readiness plate, the per-egress cross-check and the egress-type gate all
come from GET /api/byedpi. So the field is gone from the status response, and
with its only cached reader gone the cache went too, together with the
background goroutine, the staleness rules, the negative-age contract and
Server.Close's duty to wait for a probe. GET /api/byedpi still connects, on
the goroutine of the request that asked.
2. readByeDPIInstances claimed to mirror /etc/init.d/byedpi "exactly" and did
not. The init script validates each section with
'enabled:bool:0' 'port:port:1080' and refuses to start one whose validation
failed. Go read the port with strconv.Atoi and, on failure, KEPT the 1080
default — so `option port 'auto'` on an enabled instance became "an enabled
instance on 1080", and anything else accepting there produced state
"listening": the one state that unlocks the byedpi egress type, handed out
for a proxy that does not exist. `port '99999'` produced the second half:
"unknown" with a sentence asserting a connection attempt that never happened.
The same shape lived in `enabled`: strings.ToLower+TrimSpace read ' 1' and
'TRUE' as on, while the router starts neither (measured — the first is
refused by validation, the second normalises to an empty value so
`[ "$enabled" -eq 1 ]` never fires).
The parse is now a closed positive list, and its expectations were MEASURED
on the 25.12.1 testbed against /sbin/validate_data with the init script's own
spec rather than inferred from libvalidate's source:
enabled: absent/"" -> off; exactly 1|on|true|yes|enabled -> starts;
exactly 0|off|false|no|disabled -> off; anything else -> does not
start, and is REPORTED by section, option and value.
port: absent/"" -> 1080; plain decimal digits 1..65535 -> that port;
anything else -> NO port is assumed, the section is not counted as
a listener and nothing is dialled for it.
Deliberately narrower than libvalidate's `port` (which also takes a sign,
leading whitespace and, through an overflow, twenty digits): narrow declines
to call a working instance a listener and prints why, wide hands out a green
apply onto a port nothing is on.
Two further sentences that asserted actions that never happened, found while
fixing the above and not reported by the review: instances past
byedpiMaxInstances were never dialled yet fell into the "the connection attempt
neither succeeded nor was refused" clause, and that clause listed their ports
alongside genuinely inconclusive ones. "Not dialled" is now its own tally with
its own sentence, and each sentence names only the ports its own claim covers.
Every test here was checked by mutation, and each carries its control:
byedpi_initparity_test.go proves the instrument BOTH accepts a valid section
(state listening, against a real socket, in a world where every connect is
accepted) AND refuses every value the init script would not start, dialling
nothing for them; byedpi_pollcost_test.go measures the poll cost with a meter
shown counting a real probe in the same test, and keeps the probe-cost control
(16 black-holed ports = 6.4 s) that explains why it is off the poll path.
Gate: bash scripts/run-tests.sh green, including -race; ok shater/panel by name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The board is not one row per name. carryForward supersedes by (name, KIND) and
appends carried rows LAST, so a chain `x` and a node `x` both live on it — and
`new Map(results.map(r => [r.group, r]))` kept the last. The chain card showed
the node's milliseconds, exit address and verdict as its own end-to-end
measurement, unmarked. Attribution is now by kind (targetResult.ts), with
kind:'' and a missing kind as ordered last resorts.
A connection routed to the engine's `block` outbound was drawn as plain mono
text, indistinguishable from `nl-reality-1` — on the page where a DNS row about
the same host gets a crit rail and a BLOCK mark. It is the kill-switch's own
Final and a legitimate rule target, so the connection log now carries the same
outcome axis the DNS log has: killed / carried / no exit recorded, a crit rail
and a mark that survives the width where the exit column is dropped.
Insights.tsx held a raw NUL at byte 36359 — a template separator written as the
byte instead of the escape. `file` called the source binary and ripgrep, git grep
and every tree-wide search skipped it in silence. It is the escape now, and the
whole of panel/src is free of control bytes.
Three contract texts had drifted from the daemon: the searched-field list did not
mention `error` (fixed on the Go side, and there were two copies), the connection
hint named neither `proto` nor the chain hops, and rowMatches folded case with
toLowerCase() — Unicode-aware, where the daemon folds ASCII only, so a needle
could find rows in the panel that the router would never return.
And the four status fields the daemon started publishing: config_applied,
apply_error, apply_error_stage, apply_attempts, apply_failed_since_unix. A
refused configuration retried on a widening interval while `engine_running` was
true, the hash was the OLD config's and every warning described the OLD config.
engine_running is TRUE there and is not contradicted — the band says WHICH
configuration is running, and the hash row, the traffic default and the findings
list each say they are about that older one. Absent is not false: a daemon
without the field is `unknown` and raises nothing, because there is no evidence
its hash is stale.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
TestBridgeFragmentSweepIsPerCall claimed its probe used "an EXISTING key, not a
new one: the sweep must still run". It did not: the stale datagram carried IPv4
id 61 and the probe id 62, and fragKey includes the identification, so the probe
opened a NEW key — the one arrangement in which the sweep runs even when it runs
only on new keys. Moving r.sweep(now) inside the `entry == nil` branch left the
test green.
The probe is now the SECOND fragment of a datagram whose first fragment is
already cached, with the two entries opened half a fragTimeout apart so the
stale one is past its deadline and the live one is not (deadlines are set at
creation and never refreshed). Two assertions before the probe pin the setup:
the stale entry must still be there, and the live key must already exist — if a
later edit breaks either, the test says so instead of quietly proving nothing.
The released bytes are checked too, which is the half of the timeout this test
is about (the correctness half is already caught by TestBridgeFragmentTimeout).
Same sweep of TestBridgeFragmentMalformed, which had the same shape of hole: a
FIRST fragment carries MF=1 and can never complete a datagram, so `got != nil`
is unreachable whether the packet was refused or accepted, and "truncated
header" asserted only that. Every subtest now asserts on the cache, and a case
for the classic overread — a header claiming TotalLength 276 in a 28-byte
buffer — is added; its control is the aligned subtest already at the bottom.
Mutations (linux, -race): sweep moved into the new-key branch fails
SweepIsPerCall by name; clamping TotalLength to the buffer instead of refusing
fails the new malformed subtest — and, as predicted, leaves its `got != nil`
assertion silent. Control: moving the sweep after the entry lookup while keeping
it unconditional keeps every test green, so the test discriminates "per call",
not "the line moved". No production code changed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
backupBeforeChange was added with "any failure aborts the write", justified by
"the uci commit that follows writes the same filesystem, so whatever stops one
stops the other". That holds for a full or read-only /overlay and for nothing
else — and the existence probe is a stat, which also returns ENOTDIR (something
dropped a file where /etc/shater should be), EACCES, ELOOP. In that state
PUT /api/config answered 500, `sub update` exited non-zero and the profile
watcher stopped saving, PERMANENTLY: none of those causes clears itself. A
convenience added this wave must not be able to take the product away.
Two changes, both about not inferring what can be measured:
- The probe is not evidence. stat(dest) answers "is this transition already
captured?"; when it cannot answer, the copy is now ATTEMPTED and the attempt
is the measurement. Only "the filesystem will not take bytes" short-circuits
it.
- The failure is classified. filesystemRefusesWrites is a positive, CLOSED list
— ENOSPC, EROFS, EDQUOT, EIO — each a condition under which the uci commit
would fail too, so aborting only changes which error the operator reads and
ours names the cause. Everything else is about the backup's PATH and falls to
the recoverable side: the config is saved, and the missing undo is NAMED
through reportBackupProblem (same shape as subCacheLogf; model cannot import
logsink, which imports model) rather than skipped in silence.
TestWriteAbortsWhenTheBackupCannotBeWritten used a FILE where the backup
directory should be — that is ENOTDIR, the exact case that must no longer veto —
so it now injects ENOSPC at the copy, and the ENOTDIR case moved to
TestBackupPathFailureDoesNotVetoTheWrite. statBackup/writeBackupFile are seams
because the two deciding failures are the two a temp directory cannot produce.
Mutation-checked (linux, -race), each with the other half green: restoring "any
failure aborts" fails only the two carry-on tests; "nothing aborts" fails only
the two abort tests; restoring the old stat handling fails only the test that
pins "attempt the copy"; dropping ENOSPC from the list or adding ENOTDIR to it
fails the classifier test and the end-to-end tests that depend on it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
1. A SWITCHED-OFF SUBSCRIPTION CAN STILL GO OUT ON THE PLAIN WAN (blocker).
Three places had to agree about `enabled=0` and did not: UpdateSubscription
resolves by name and never reads it; cmdSubUpdate reads it only when no name
was given; warnings.go skipped disabled subscriptions entirely on the stated
premise that one "is never fetched". The premise was the false one, and the
per-row Fetch-now button added this wave posts exactly the named request.
Kept the behaviour, dropped the premise. Enabled means "include in the
automatic refresh" everywhere else in the system — MergeSubCaches loads a
disabled subscription's cached nodes unconditionally and they route traffic —
and a refusal here is worked around by enable/fetch/disable, which enrols the
sub in the 6-hourly sweep and is strictly worse. The automatic paths still
honour it (the nameless sweep, and shater-cron's own `en = 1` check). The
named path now says so on stderr and in the daemon log, and the leak finding
fires for disabled subscriptions with the WHEN clause corrected — "every
scheduled refresh" is false of a subscription no schedule touches.
2. refreshBootArmor DISARMED THE NEXT BOOT FROM A CONFIG THE DAEMON REFUSES.
Every call site is gated on readErr == nil and nothing else; ParseUCIExport
drops unknown options silently, so a config written by a newer build reads
clean, and with the divert set emptied by the parse RenderHoldNft returns ""
and the armor was REMOVED — with no log line at all, unlike the disarm one
branch above it. Measured: with the new gate removed, the armor really is
deleted. Now gated on the schema, and both removal paths are announced.
3. THE FIRST-BOOT DEADLOCK IS NAMED. Every subscription pulled through the
tunnel, the tunnel built from nodes only a fetch supplies, the caches gone:
the fetch waits for the tunnel and the tunnel waits for the fetch, forever,
with the LAN dark. The CLI refusal goes to /dev/null (shater-cron) and the
daemon line to a syslog `log_syslog='0'` switches off. It is now a critical
finding in /api/status, which survives both, with the state named and two
escapes — the free one first, the costly one priced.
4. A REJECTED CONFIGURATION WAS INVISIBLE, AND THE ENGINE CHURNED. Measured on
the stand: with a config the engine cannot accept on disk, cron retries every
60s and every attempt is a full engine swap, while status showed
engine_running=true, the OLD hash, the OLD warnings, and `grep -ci` for the
broken element returned 0. Invisible by construction: everything published
about a config is published by a SUCCESSFUL apply, and engineDownCause is
gated on the engine being down — here it is up.
Status gains config_applied / apply_error / apply_error_stage /
apply_attempts / apply_failed_since_unix, and a critical finding that says
the running configuration is a DIFFERENT one and names the reason. Reconcile
paces an identical retry (three free attempts, then doubling to a 15m cap);
any change to the configuration cancels the wait, and POST /api/apply is
deliberately not paced. The post-swap abort is deliberately NOT recorded —
it is already loud and its retry costs no swap.
Also, from review-by-seams:
- The netplane channel was graded critical wholesale over three distinguishable
states. `udp '0'` + closed is the kill switch doing what it was told and may
be exactly what was asked for; the leak and the total cut-off are not. The
first is now `warning` (not `info`: attentionFindings drops info, and the
blast radius is wider than the switch's name). Default stays critical, the
exception is a closed list, and netplaneprotoseverity_test.go pins it against
the REAL renderer so a rewording fails by name instead of drifting.
- devicefilter_severity_test.go carried a FOURTH unlinked copy of
DEVICE-FILTER-NOT-APPLIED and compared it with itself — the same shape as the
noGatewayFinding fixture this wave removed. apply's copies are one constant
now, and the real coupling is a test that runs generate and grades what comes
back. Mutation: renaming the tag in generate fails it by name; the two old
fixture tests survive that untouched, which is the whole point.
- The history-write failure was logged ABOVE the deduplication gate its own
call site documents eight lines below. At one cron reconcile a minute a
standing cause (full /overlay, an entry over the 128 KiB ceiling) wrote 1440
identical lines a day, and under log_persist=1 that many appends to flash —
the exact wear the history ring's own dedup exists to prevent. Now gated on
the message changing, cleared by a success. The Warning is still returned
every time; only the log had a repetition problem.
Every fix mutation-checked with the failure text recorded, and every one has a
control showing the instrument can still give the opposite answer: an enabled
subscription still fetches and keeps the scheduled wording; a legitimate disarm
still happens and is still logged; a healthy box raises no rejected state; an
ordinary netplane finding is still critical; a DIFFERENT history failure still
prints. One mutation (the history-dedup latch) SURVIVED its first test — the
counter matched the success path's Info line too — and the test was fixed.
shater/apply is green. shater/cmd/shaterd was green when run 20 minutes ago and
now fails to BUILD on shater/panel/byedpi.go, a neighbour's in-flight refactor;
the full gate run for the same reason cannot be completed on this tree right now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two holes, both in the direction the verb cannot afford: `shaterd diag` produces
the one text block a person SENDS somewhere.
1. resolver.address was on the allow-list, printed verbatim. For a DoH resolver
that field is a URL, and generate/dns.go's parseDoHAddress keeps and USES
u.Path — which is exactly where NextDNS, AdGuard and Control D carry the
account identifier. Whoever holds it reads and rewrites this household's DNS,
so it is a credential. The panel had always read it that way (DNS.tsx's
resolverAddr shows u.host and flags the rest); the disagreement was resolved
in favour of the side whose output goes to a stranger. Now: scheme and host
survive, userinfo/path/query/fragment do not, and a BARE address
("1.1.1.1", "dns.adguard.com:853") is still printed in full because it is
host and port and it is what the fault is read from.
The fix could not be "delete the key from the list": TestDiagMasking-
IsClosedOverTheWholeModel asserted the allow-listed fields come out
UNMASKED, so it actively pinned the leak. The transform lives in a second
closed table (diagMaskedForm), and the sweep now compares the masked render
against the raw one line by line, expecting either the plain mask or exactly
what that table declares.
2. The second layer collected its literals from `uci export shater` alone, and
that file does not hold this router's credentials. model/render.go never
writes a FromSub node; the several hundred subscription nodes live in
/etc/shater/subs/*.json, which keep.d/shater-core describes in its own words
as carrying "every node's credentials". The reachable path is not
hypothetical: parse/sharelink.go quotes a rejected node's USERINFO into its
error, generate/outbound.go warns it, apply/warnings.go logs it, and the last
32 KiB of that log is section six of the bundle — with LogToFile on by
default. diagSubCacheSecrets now reads those files by the same closed
positive-list rule (unknown JSON key => collected, so a field added to
model.Node tomorrow is covered), and a file it cannot read is NAMED in the
bundle instead of silently reducing the scrub.
Fixing the first half exposed the second: the log carried the userinfo, not
the whole URI, so a literal scrub of the URI walked past it. diagSecretParts
expands every refused value into its userinfo, username, password, query
values (encoded and decoded) and path. Not the fragment — in a share link
that is the node's display name, which is on the printable side.
The banner no longer says secrets are masked "throughout". It says what is
masked, and then names what is still in there: values under 8 characters (masked
in the config, not scrubbed elsewhere), list/ruleset URLs, and the limits of a
literal scrub.
Mutation-checked, each with the control that the instrument SEES the planted
secret in the unfixed output:
resolver.address back on the allow-list -> resolver test fails on the id
diagMaskAddress made the identity function -> transform test names the field
sub-cache literals withheld from the scrub -> log-scrub test fails
diagSecretParts reduced to the whole value -> log-scrub test fails
sub-cache safe list turned into a blocklist -> closure test fails
unreadable cache file swallowed -> honesty test fails
nil collector seam read as "nothing to do" -> honesty test fails
transform applied per section, not per key -> new-field test fails
one section given an open default -> sweep fails, by name
an allow-listed value over-masked -> sweep fails, by name
masked lines dropped entirely -> sweep's vacuity guard fires
scripts/run-tests.sh green (all 7 steps, -race included).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two defects of the same family: a state the plane produces and nobody names.
1. The divert is written per PROTOCOL, the fail-closed drop per INTERFACE.
A tproxy inbound with `option udp '0'` puts its device in the drop scope
(nftDivertRefs does not look at the flags, and must not: the drop is the
backstop for ESP/GRE/SCTP too) while emitting no UDP TPROXY line for it.
With kill_switch=closed every outbound UDP packet from that network is
dropped; with kill_switch=open the same packets leave the WAN in the clear.
TCP works, DNS works (dnsmasq answers it past the fib-local bypass), so it
presents as "some sites do not load", not as a firewall. Verified by
rendering: no second LAN is required, the shipped one-inbound shape does it.
The drop is NOT narrowed to match the divert. Doing so would turn
`option udp '0'` — which is how you kill QUIC so the engine can route by SNI
— into "UDP now bypasses the proxy", i.e. it would convert a QUIC-blocking
config into a QUIC-leaking one, and it would open a per-protocol hole in the
kill switch through a knob whose name says nothing about leaking. The plane
already takes the other decision one field over: with ipv6 off no v6 divert
is emitted and closed mode drops v6 anyway, deliberately and in writing.
So the state stays and is named instead, in three shapes (protocol dropped /
protocol leaked / both flags off), each naming the network, the option, the
kill-switch state and the concrete traffic that dies.
coverage.go could not have caught this: it skips covered[i.Device], and the
device IS covered. The new check is derived from the model alone and so runs
outside that file's Interfaces() gate.
2. networkList's open `default:` sent (tcp=0, udp=0) to "" — which the engine
reads as BOTH — so an inbound the plane feeds nothing acquired a listener for
everything. The four cases are now named and closed, "neither" is a second
return value rather than a synonym for "both", and a tproxy inbound that
carries no protocol is refused with a warning that also names the netplane
half: switching both flags off does not remove the network from the plane, it
removes the way out of it, so closed mode cuts that network off completely.
generate_test.go: the three linux fixtures that built a tproxy inbound with
model.Inbound's zero-value flags now spell TCP/UDP out. UCI defaults both to
true; only a Go-built model gets false, and only that fixture relied on it.
Gate green (bash scripts/run-tests.sh, exit 0). Every new test mutation-checked
in both directions: suppressing the warning fails 6 tests by name, and making it
fire unconditionally fails the controls. Rendered ruleset text is byte-identical
for all nine shapes dumped before/after — the only diff is added warning lines.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two things the panel asserted that the code does not do.
1. THE OFF SWITCH IS NOT A GATE ON FETCHING. The row badge for a disabled
subscription with `fetch_via=proxy` and no detour was drawn quiet and said
"Nothing is disclosed yet — this subscription is switched off … so it is never
fetched". False in all three places that could have contradicted it:
Applier.UpdateSubscription resolves a subscription BY NAME and has never read
Enabled; `shaterd sub update` consults Enabled only when no name is given; and
this panel's own per-row "Fetch now" — new in this wave, previously buried in the
collapsed Options panel — is disabled on `busy || fetching` and nothing else. One
click sent the router's real address to the feed host under a badge saying
nothing was disclosed.
The two halves of the old condition are not alike, so they stopped being one
state. NO URL is real and refused at the bottom (subscribe/fetch.go rejects an
empty URL before it builds a request) — that branch keeps its quiet badge. OFF is
amber, and its sentence says what the switch actually does: it stops the
scheduled refresh, and the button on the row asks for a fetch whatever the switch
says.
The daemon reached the same conclusion from its side in this wave — the fetch is
deliberately allowed and logged, and its finding now varies on Enabled — so the
badge's own summary over that finding varies the same way. "On every scheduled
refresh" printed over a switched-off row is the same lie inverted: it sends the
reader hunting a cron job that is not running.
2. AN INLINE LIST HAS NO ENTRY COUNT, AND "NOT PUBLISHED" IS NOT "EMPTY".
engine.go fills RuleSetStat.RuleCount from (*rule.RemoteRuleSet).RuleCount(), and
LocalRuleSet.RuleCount does not exist in the tree at all, so an inline list always
arrives with rule_count 0. The chip called a working parental-control list
"empty — nothing matches". It reads "size unknown" now: unlit, never green and
never the amber that says something is wrong. A mixed group is counted as a floor
("1,284+") instead of presenting a partial sum as the whole.
BOTH INSTRUMENTS WERE HOLDING THE LIE UP. subFetch.test.ts pinned the sentence
verbatim, and deviceLists.test.ts fixed `{remote:false, rule_count:3}` — a record
no router can produce, so its green light was wired to nothing. The mock carried
the same impossible state on three local rule-sets and fabricated a count on
update. All of them now match what the daemon sends.
Verified in ?mock (new `?subleak=paused|pausedapplied|nourl`), 390 and 1280, both
themes. Each fix reverted in turn with the failure text; controls both ways — a
remote list that really is empty still says "empty", and the two leaking states
are still told apart from the one that is genuinely quiet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The daemon changed under the panel in three places, and in each one the panel
kept drawing a screen that was right only by accident.
byedpi readiness carries `age_seconds` now, because GET /api/status stopped
probing: sixteen enabled instances on ports that neither accept nor refuse cost
6.41 s per poll, measured, and the Apply page polls up to 27 times a minute. The
report is served from a cache and every sentence in it is present tense, so the
panel stamps it. Negative is not an age — zero is the common answer (a loopback
connect finishes in microseconds) — so "not a measurement" is its own reading,
and a daemon too old to send the field is a third one: the reading is real, its
age is not reported. The cold first poll after a start says "not measured yet"
rather than "not determined": nobody has looked is a normal state of a router
that booted ten seconds ago, and it calls for a different sentence than an
instrument that looked and failed. Both keep the unlit lamp and both keep the
egress type locked.
A test run no longer wipes the board, so a card can show a reading from twenty
minutes ago beside one from a second ago. Which is which comes from `scope`, not
from comparing timestamps — the router has no RTC and a computed "n minutes ago"
would be fiction. A carried row says "earlier run" and is drawn as a qualifier;
an empty scope is "cannot attribute", never "everything is carried", because a
real run always covers at least one target.
A chain blocked at a hop was kept red by matching a fragment of the daemon's
error sentence — the last place prose decided anything here. It arrives as
`blocked_by` now, so the match is gone and the row names the hop.
The mock carried the old contract: it emptied the board on every run while a
comment claimed the daemon did too. It carries forward now, by (name, kind),
capped at 64, and `?mock&board=carried` lands on a finished board holding both
kinds of row. `?mock&byedpi=cold` and `?mock&byedpiage=N` reach the two states
the freshness rendering exists for.
Verified in ?mock at 390 and 1280, both themes, no horizontal scroll. Each of the
three fixes was reverted in turn and the tests named the failure; the controls
run the other way too — a helper that marked every row carried, or reddened every
row, fails just as loudly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
17843be5a measured six packages damaging the machine that runs the suite and
fixed four; model and generate were left because another agent held those trees.
Measured again on 2026-07-27 with the same instrument, scoped to the two
packages, and the diagnosis held EXACTLY:
CREATED /etc/shater/config.pre-unreadable.bak
CREATED /etc/shater/config.pre-v0.bak
CREATED /etc/shater/config.pre-v1.bak
CREATED /etc/shater/config.pre-v2.bak
MODIFIED /etc/shater/cache.db
model. Every test that reaches writeUCIWith or migrateWith goes through a
fakeUCI, and that seam is what makes the config they read and write a fake.
backupBeforeChange is the one part of the package that does NOT use it: it
os.ReadFile's liveConfigPath and writes into configBackupDir directly. On a dev
box neither exists and the function returns "nothing to copy"; on the testbed and
the router both exist, so the suite planted four bogus copies in the product's
state directory. Worse than litter: the function is create-ONCE per schema and
never overwrites, so a copy planted by a test SILENTLY PREVENTS the real
pre-migration copy that box was going to take.
generate. generate.go emits experimental.cache_file with Path: cacheFilePath(),
and every *_linux_test.go that hands a generated config to engine.Apply/box.New
opens that bbolt DB for writing. Per cache.go's own file comment that DB is a
SAFETY device, not an optimisation: with it, RemoteRuleSet.StartContext skips the
start-time fetch, so the daemon can come up before the WAN does. Rewriting it
from a test is rewriting the thing that keeps a reboot from taking the LAN down.
THE FIX is the one the other four packages already use, not a third one: a
TestMain per package pointing the product paths at a private os.MkdirTemp, plus a
test that still pins the SHIPPED value — because an isolation that leaves the
real decision untested has only moved the defect.
model: liveConfigPath/configBackupDir -> a private dir; liveConfigPath is
pointed at a path that does NOT exist, which is exactly the dev-box
case the function already documents, so every test that does not opt
into backupSandbox behaves precisely as before.
New TestConfigBackupPathsAreTheShippedOnes.
generate: cacheDirPersistent/cacheFilePersistent/cacheFileFallback -> a private
dir, and the persistent one is CREATED so the package keeps
exercising the branch the ROUTER takes. The fallback had to move too:
on a host without /etc/shater the decision lands on
/tmp/shater-cache.db, which is just as hardcoded and just as much the
product's. New TestCachePathsAreTheShippedOnes, which also pins that
the DB lives inside the directory the free-space checks measure —
path.Dir, not filepath.Dir, since the gate also runs on Windows.
No waiver was needed at shater/testguard: it follows
`cacheDirPersistent = filepath.Join(dir, ...)` back to os.MkdirTemp on its own.
Verified:
- the sweep, scoped to the two packages: the five paths above BEFORE, "CLEAN"
AFTER. Then the FULL scripts/check-test-fs-isolation.sh: 48 package
verdicts, "CLEAN: the whole suite ran and not one path under /etc /var /usr
/root /home /opt /srv /run /tmp changed."
- positive control: a planted test in shater/model that restores the real
paths and calls backupBeforeChange -> the sweep names
"CREATED /etc/shater/config.pre-v9.bak", then bisects to "PACKAGE
.../shater/model" and "TEST ....TestPlantedViolatorWritesTheRealBackup".
shater/testguard stayed GREEN with the violator in the tree, which is the
documented blind spot and the reason the dynamic half exists.
NOTE, learned from the first attempt: a create-ONCE violator is named by the
verdict but NOT by the bisect — seed_canaries only creates what is missing,
so the file the whole-suite run left behind makes the per-package re-run a
no-op ("no single package reproduced it"). The bisect can only name defects
that repeat.
- mutation, model: liveConfigPath -> /tmp/shater-live and configBackupDir ->
/tmp each fail the new test by name; dropping the "keep the older copy"
return fails TestBackupBeforeChangeKeepsTheFirstCopy ("the first copy was
overwritten by a later write"); removing the ErrNotExist early return fails
TestBackupBeforeChangeSkipsWhenThereIsNothingToCopy; removing the
backupBeforeChange call from writeUCIWith fails
TestWriteTakesTheBackupBeforeReplacingTheConfig ("the write took no backup").
- mutation, generate: cacheDirPersistent -> /tmp/shater and cacheFilePersistent
-> /etc/shater-cache/cache.db each fail the new test, the second one also on
the dir/file mismatch; cacheFilePath forced to tmpfs fails
TestCacheFallsBackWhenDirMissing's CONTROL, forced to persistent fails its
first half; cache_file Enabled=false fails TestCacheFileEmittedAndEnabled.
Green again after every revert.
- counts, declared vs executed (go test -list vs top-level verdicts, shipped
tags, linux): model 160/160, generate 395/395, 0 failures. generate's 3 skips
are the pre-existing CAP_NET_ADMIN TestIntegrationL3* trio, which [5/7] runs
and passes.
- scripts/run-tests.sh: GREEN end to end, exit 0, including [4/7] under -race
("OK [race] in 43s") and [5/7] RAN all three privileged tests. The four
TestByeDPI* races reported earlier no longer fire.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
GET /api/status called byedpiProbe on every request. On a healthy loopback that
is nothing, but the probe's cost lives in exactly the state it was written to
report honestly: a port that neither accepts nor refuses burns the full
byedpiDialTimeout, and byedpiMaxInstances of them burn 6.4 s. Measured, on this
tree:
1 enabled instance, live listener 0.45 ms
1 enabled instance, refused 0.33 ms
16 enabled, refused 3.6 ms
1 enabled, black-holed 400 ms
16 enabled, black-holed 6.41 s
The panel shell polls /api/status every 5 s on every page and the Apply page
adds its own 4 s poll, so the pathological state hung the panel for seconds at a
time precisely while an operator was trying to find out what was wrong. A check
that gets slow exactly when it matters is worse than one that is always slow.
The poll now reads a cache (byedpiReadiness.cached), refreshed asynchronously off
the same path: 15 ns per call, primed, and 20 back-to-back polls against 16
black-holed ports cost less than one probe. A background ticker was rejected —
it would dial on a router whose panel nobody has open — and so was blocking the
first poll to fill a cold cache, since that is the same 6.4 s hang, just rarer.
The cache is not allowed to lie:
- every served report carries age_seconds. Detail is written in the present
tense, and a present-tense sentence about a measurement taken some seconds
ago is a claim nobody checked;
- a report older than byedpiCacheMaxAge is NOT SERVED. It is replaced by an
explicit unknown with a negative age, so a panel that ignores the age fails
to an unlit lamp rather than to a stale "listening" unlocking an egress type
onto a port nothing is on;
- GET /api/byedpi still really connects. A re-check button answered from a copy
is a button that does nothing.
And it has an OWNER. The refresh runs a goroutine that dials; left as a package
variable it belonged to nobody, could not be awaited, and — as the race detector
showed — went on reading byedpiConfigPath / byedpiInstalled / byedpiDial after
whatever started it believed it was finished. The cache is now per-Server, with
stop() that forbids further refreshes and does not return while one is dialling,
called from Server.Close. The daemon already defers that Close, so the probe
cannot outlive the server.
byedpiDial became a seam alongside byedpiInstalled and byedpiConfigPath: the
timeout branch is the expensive one and the one a real loopback cannot be
provoked into, so without it neither the cost nor its removal could be shown.
Ten mutations, each killed by a named test: the probe back on the request path;
the cached copy claiming age 0; an over-age reading quoted anyway; a cold cache
returning a blank instead of an explicit unknown; a refresh that is not
single-flight; a late older probe overwriting a newer one; /api/byedpi answering
from the cache; the unmeasured report claiming the binary is absent; stop() not
waiting; Server.Close not stopping. The last one survived its first test — which
asserted a poll straight after Close did not dial, and passed with the stop
removed entirely because the reading was fresh and no poll was due — so the test
now advances the clock to make it due.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two separate honesty defects in the shared test board, both surfaced while
closing the byedpi readiness fix.
1. startTestRun replaced the results slice outright, so pressing Test on one
NODE blanked every group and chain card on the Targets screen, and testing a
group blanked the nodes. Nothing on screen explained it, because nothing had
happened to those targets — the daemon had thrown their readings away.
Earlier results are now carried forward for every target the new run does not
itself re-measure. The alternative, one board wiped per run, is simpler and
has no staleness question at all; it was rejected because it destroys
information the daemon still has. These are the OBSERVATORY's numbers, taken
by a prober that never stopped, and a group's reading does not become false
because somebody tested a node afterwards.
The staleness question it does raise was already answered: every result
carries tested_unix, the instant the OBSERVATION was taken, and GroupTestStatus
publishes this run's scope — so a carried row is identifiable as carried
without comparing timestamps, and drawn with its age. The board is capped at
groupTestCarryMax, evicting the oldest first; that cap is the only way a row
can leave without a newer one taking its place, and it is documented as such.
done/total still describe this run's targets only.
2. A chain whose exit was never dialled, because an earlier hop was probed and
did not answer, shipped that fact as prose only: source="" (correct — nothing
measured the exit) plus a sentence naming the hop. A client reading source
strictly filed it under "nobody looked", which is the wrong colour, so the
panel kept the row loud by matching a fragment of our error message — the
last place it read our prose to decide anything.
GroupTestResult now carries blocked_by: the 1-based hop index, 0 everywhere
else. It does NOT set source; nothing measured this target's own path, and
stamping an instrument on a measurement that never happened is exactly the lie
source was added to prevent. blocked_by>0 beside source="" is the complete
statement. Field and sentence are produced together in chainBlockedResult so
they cannot come to disagree.
Tests (grouptest_board_test.go), each verified by mutation:
- carrying forward is asserted WITH its control, that a run does replace the
rows it covers — "nothing disappeared" alone is also satisfied by a board
that stopped updating;
- target identity is (name, kind), so a group and a node of one name do not
evict each other, with the empty-kind wildcard pinned both ways;
- the cap drops the oldest end;
- blocked_by carries the hop, keeps source empty, and every other result
carries 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
shater/stats/store_test.go ended with
_ = os.Remove(statsFilePath())
and statsFilePath() is not a test path. It is THE product path: /etc/shater/stats.db
on every host where that directory exists, which is the testbed and the router. So
`go test ./shater/...` deleted the accumulated query and connection log of whatever
machine ran it. The test passed. It had always passed — damage done by a test is a
side effect, not a wrong answer, and no instrument in this tree could see one.
A filesystem sweep (the new scripts/check-test-fs-isolation.sh: seed a router-shaped
canary tree in a container, run the whole gated suite, diff) found it was not alone.
Six packages, by measurement, not by reading:
shater/stats DELETED /etc/shater/stats.db (the line above; also
TestComboBackendSwitchSequence opened and pruned the
live DB, which the delete had been hiding)
shater/logsink DELETED /etc/shater/shaterd.log and /var/log/shaterd.log —
New()/Reconfigure() purge BOTH product locations when
the file toggle is off, so Config.Path (which every test
here already set) never protected them. The daemon's own
log, the one an operator reads after an outage.
shater/apply DELETED /var/run/shater.active — the ONE token hotplug and cron
check before touching the data plane. Clearing it on a
live router makes both stand down on a box that is up.
holdstate_test.go's `t.Cleanup(os.Remove(ActiveFlag))`
was not a cleanup; it was the delete.
shater/panel REWROTE /etc/shater/stats.db — stats.NewStore("sqlite") from
TestStatsEndpointsAcrossBackends resolves the product
path too.
shater/model CREATED /etc/shater/config.pre-v{0,1,2}.bak, config.pre-unreadable.bak
shater/generate REWROTE /etc/shater/cache.db
The last two are NOT fixed here — another agent is working in those trees. Both are
one TestMain away: model already has liveConfigPath/configBackupDir as vars, and
generate already has cacheFilePersistent; what leaks is product code (backupBeforeChange,
the engine's cache_file) called from tests that do not redirect them.
THE FIX is the seam generate/cache.go and generate/ruleset.go already use — the path
becomes a package-level var that only tests assign — plus, in each case, a test that
still pins the SHIPPED value, because an isolation that leaves the real decision
untested has only moved the defect:
stats: statsDirPersistent/statsFilePersistent/statsFileFallback + the exported
SetPathsForTest (exported because shater/panel needs it from outside).
New TestStatsFilePathPrefersPersistentDir covers both branches.
logsink: PersistPath/TmpfsPath + a TestMain, since the hazard is in New(), which
every test calls. New TestLogPathsAreTheShippedOnes.
apply: ActiveFlag + the existing TestMain. New TestActiveFlagIsTheShippedPath,
which also records WHY /var/run: tmpfs, so a reboot clears it.
TestNewStoreSelection got stronger rather than weaker. Its "sqlite" case used to
accept "sqlite" OR "memory" because the real path might not open on this host — an
expected value that depended on the machine. At a private path there is no excuse:
a writable directory MUST report "sqlite", and a new control at an unopenable path
MUST report "memory" (the honest "persistence is not active" signal) without a crash.
TWO GUARDS, because one of them cannot see half of it:
shater/testguard/fsisolation_test.go — parses every _test.go under shater/ and
fails BY NAME when a filesystem-mutating call gets a path that is not PROVABLY
temp-rooted. Positive and closed: what it cannot prove is a failure, not a
default, which is the only rule that catches a path built by a function call.
It follows local vars, closures, filepath.Join/Sprintf/+, helper parameters via
their call sites, helper return values, and the save/override/restore idiom.
Four waivers, each keyed on file+function+callee, each with the reason printed on
every run, each a struct field traced by hand; a waiver that stops matching fails
the test as STALE. Runs inside [2/7] and [4/7] — no new gate step, no new minute.
Blind spot, stated: damage done by PRODUCT code a test merely calls (which is
exactly logsink, model and generate above).
scripts/check-test-fs-isolation.sh — the dynamic half, for that blind spot. It
refuses to run outside a container unless told twice, because its method is to
let the damage happen and then look, and it seeds/unseeds only what was missing.
Verified:
- mutation, task 1: statsFilePath forced to the fallback -> the new path test
fails ("with ... present = .../fallback-stats.db, want the persistent ...");
newPersistent forced to memory -> "Backend = \"memory\", want \"sqlite\"";
the fallback made to report "sqlite" -> "Backend = \"sqlite\", want \"memory\"".
Green again after each revert.
- mutation, the guard: the original os.Remove(statsFilePath()) put back -> named
at store_test.go:154 with "the path comes out of statsFilePath(), which this
check cannot follow"; a planted test writing "/etc/config/network" -> named as
a literal path; the walk pointed at one package -> its own <150-file control
fires ("reading a blank page"); a waiver matching nothing -> STALE WAIVER.
- control, the sweep: with a planted violator it reports DELETED /etc/shater/stats.db
and MODIFIED /etc/config/network; without it, those are gone and only the two
foreign packages remain. Its bisect named shater/stats.TestComboBackendSwitchSequence
on its own.
- counts, declared vs executed (go test -list against top-level verdicts):
stats 112/112, panel 121/121, apply 122/122, logsink 26/26, testguard 1/1,
0 skips, 0 failures.
- scripts/run-tests.sh: [1/7][2/7][3/7][5/7][6/7][7/7] green. [4/7] -race fails on
four TestByeDPI* in shater/panel — a data race between byedpi.go's background
probe and byedpi_test.go's forceByeDPIBinary cleanup, in another agent's
uncommitted work (shater/panel/byedpi_cache_test.go is untracked). Proven not
ours: a pristine HEAD tree carrying ONLY this commit's files passes -race over
all 35 packages, and the same run with -skip ^TestByeDPI is green on the live
tree too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The DNS log drew every row from `action`, which answers WHICH WAY the lookup
went — so a query that left through a detour and then timed out came back as an
accent-coloured `proxy` tag, blocked=false, nothing else said. The row that
describes the exact moment the tunnel broke was the most reassuring line on the
page. `actionTag()` also fell open (`return 'pass'`), so every value the panel
did not recognise — including every value a future daemon might add — rendered
green.
The daemon now carries the outcome as its own axis (LogEntry.Status/Error,
a794fbe37). This brings it to the screen.
Two axes, and the outcome leads. logRoute.dnsRowMark decides both in one place:
status → answered | failed | '' (not recorded), POSITIVE and CLOSED, with the
fallback on the recoverable side. `blocked` refines a recorded answer
into the fourth situation and is never allowed to invent one on a row
whose outcome was never written.
action → block | proxy | pass | unknown, the same discipline. The path stays
VISIBLE on a failed row and muted, because "it failed" and "it failed
in the tunnel" are different reports and the second one closes tickets.
Four situations, four looks: a plain answer has no rail; a filter block keeps its
crit rail and BLOCK tag; a failure takes an amber rail, an amber wash, a filled
FAILED chip, and its cause verbatim beside the rcode reading (-1 renders "no
response", anything else the code the server really sent); a not-recorded outcome
is dashed and faint and claims nothing. A failure with no recorded cause says
"cause not recorded" rather than showing an empty cell that reads as fine.
`error` joins the searched fields (the daemon searches it — q=timeout works) and
the hint under the box now names it. `status` stays out: q=failed must not sweep
up every failure while somebody is looking for a domain by that name.
Verified in ?mock at 390 and 1280, both themes, no horizontal scroll: the four
states are pairwise distinct in computed border/background/colour, and the two
chips take their own line on a narrow screen so the domain keeps 92px instead of
being pinned at its 30px minimum. 19 new tests, each shown to fail under 16
mutations of the code it covers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Three things the panel was saying that it had not established.
BYEDPI. The egress type unlocked on `status.byedpi_installed`, which is
LookPath("ciadpi") — "is the package installed", while the operator is asking
"will traffic sent here go anywhere". They come apart on the SHIPPED config: the
packaged /etc/config/byedpi is inert, so installing the package unlocked the
type, the egress went on 127.0.0.1:1080, the apply was green and nobody was
listening. The gate is now `byedpi.state === 'listening'` and nothing else
(byedpiReady.ts, the only place that decides it). GET /api/byedpi also carries
the per-egress port cross-check, so a mismatch is drawn on the row that has it,
naming both ports, in crit — the state where every other signal reads healthy.
`unknown` is neither answer. It keeps the type locked (a control that opens on
nothing established is the same defect wearing a new word) and it is never
painted as a refusal: dashed border, unlit lamp, "not determined", plus a
Re-check button so a dropped request is not a dead end.
ONE NODE. A freshly pasted node had no instrument — the group test reads the
observatory's board and the observatory only probes what the rules route
through, so the first question anyone asks answered "not routed by any enabled
rule". Every node row now has Test, over the same singleton run and the same
GET poll the Targets page uses.
The reading is classified on `source`, not on prose: measured-and-failed is red,
`source:""` is an unlit lamp and the faintest text on the row, because a
negative result that cannot be told from a check that never ran answers nothing.
One escalation survives, documented and narrow: a chain whose exit was never
reached because a hop it runs through WAS probed and failed. Targets keeps its
exact previous appearance while its instrument changes underneath.
The three refusals stay three facts — 404 the node is not in the saved config,
503 the config could not be read (an unknown, never a verdict about the node),
400 no name — with three tones and three sentences.
SUBSCRIPTION FETCH ROUTE. The panel's draft predicate drew amber "proxy · no
route" while the daemon now grades the same fact critical on the same row, so a
saved leaking subscription wore both, at two severities, about one thing.
subFetch.ts reconciles them: where the daemon has spoken it outranks the
prediction, in BOTH directions — including the dangerous one, where the
predicate is content ("via group:auto") and the daemon reports the leak anyway.
The prediction still speaks for a draft nothing has applied yet, and a
subscription that is switched off is not accused of a disclosure the daemon
deliberately does not report for it.
Fixtures reach every state: ?byedpi=<five states>|mismatch, ?nodetest=ok|dead|
unmeasured|400|404|503, ?subleak=draft|applied|divergent. The group-test fixture
also stopped being kinder than the daemon — engine.startTestRun replaces the
whole board, so refreshing one target really does blank the others.
42 tests, 8 mutations each killed by name, and both controls: the gate is shown
to open on `listening` and to stay shut on the other four, and "not checked" is
shown to be drawn differently from "did not answer" — the assertions fail if
either pair is ever drawn alike. Verified in the browser at 390 and 1280, both
themes, no horizontal scroll.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
handleConfigPut decodes model.Model with DisallowUnknownFields, so this is the
layer the missing fields bit at: not "the attachment is ignored" but "the whole
save is rejected with json: unknown field \"Blocklists\"", losing every
unrelated edit batched into the same PUT. Mutating the JSON name reproduces
that message exactly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The panel half already ships Device.Blocklists/Device.Allowlists and PUT
/api/config decodes with DisallowUnknownFields, so until now the first
attachment rejected the WHOLE save with `json: unknown field "Blocklists"`,
losing every other edit in it. This is the engine half.
A device now references `config blocklist` / `config allowlist` sections by
name (UCI: `list blocklist` / `list allowlist`, since `block`/`allow` already
mean the typed domains), which brings geosite categories and url-sourced lists
to parental control for free.
Two things here are constructions, not checks.
The tag a device's rule references comes from the accumulator that emitted the
rule-set, never from the list name. Rule-set tags resolve at engine START
(RuleSetItem.Start), so a name-derived tag passes box.New and fails box.Start —
and because both configs share one cache_file path, every apply on a live
engine takes the close-old-then-start-new branch, so the old box is already
gone when the new one refuses. That is no engine, a closed kill switch and a
dark LAN, from one mistyped list name. A reference that yields no tag emits no
rule at all; the emptiness is warned, tagged DEVICE-FILTER-NOT-APPLIED so the
panel grades it critical rather than guessing from prose.
Materialisation is a single memoised point shared by both consumers. Devices
are built before the network-wide filter, so materialising a shared list twice
would hand dedupeRuleSetTags an already-claimed tag — which it DROPS, silently
switching the network-wide filter off for that list. The mutation test for this
reproduces exactly that: DNS-FILTER-NOT-APPLIED, filtering nothing.
Order is the feature: typed allow, typed block, attached allow, attached block,
then the network filter. Otherwise a parent who types youtube.com into a
child's Block loses to whatever an attached geosite category permits, and the
panel draws a "Blocked" chip over a rule that does nothing. Typed and attached
matchers stay SEPARATE rules — rule_set AND-gates over the domain matchers, so
merging them would mean "the domain AND the list".
Attaching a list is itself the switch for that device: Enabled=0 means "does
not participate in the network-wide filter", not "dead", so the list still
loads and filters here — and generate says so instead of leaving it to be
discovered. A blocked name's reply comes from the LIST (Blocklist.Response),
so tier 4 is up to two rules; the typed tier keeps NXDOMAIN, having no owning
object to say otherwise.
Purely additive: an old config has neither list, parses to nil, and the
generated engine config is byte-for-byte what it was. No schema bump.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
TestCacheFallsBackWhenDirMissing guarded itself with
if fi, err := os.Stat(cacheDirPersistent); err == nil && fi.IsDir() {
t.Skipf("%s exists on this machine; ...")
}
i.e. its subject was the machine it happened to run on. The gate caught it as an
UNDECLARED SKIP in one container run and not in the next, with no change to the
code — and BOTH outcomes were green. Only the undeclared-skip check saw it at
all; every other instrument here reports `ok shater/generate` either way.
An order-dependent test proves nothing on the runs where it does run either,
because nobody can tell afterwards which runs those were.
The three cache paths become vars (production never assigns them, same seam
generate/ruleset.go already uses for listsDirOverride) and the test points them
at a temp tree. It now covers BOTH branches with no skip: an absent dir must
choose tmpfs, and — the control — a present one must choose the persistent
path. Without that second half the test is satisfied by a cacheFilePath that
returns the fallback unconditionally, which is exactly the regression the
persistent branch exists to prevent (a cache that never survives a reboot, so a
reboot before the WAN is up fails to start the engine and takes the LAN with it).
Verified:
- both halves killed by mutation (force persistent -> the first assertion
fails; force fallback -> the control fails), green again after revert;
- 5 x `go test -shuffle=on ./shater/generate/`: 474 verdicts and 3 skips
every time, TestCacheFallsBackWhenDirMissing PASS on all five, never SKIP;
- shater/generate: 385 declared func Test*, 385 top-level verdicts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
shater/stats/filter.go:156 searches ConnLogEntry.Error along with the six
fields the doc names, so `q=timeout` works and the contract said it did
not. Verified against the predicate, not against a report.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The -race step failed with
FAIL shater/netplane 600.019s
panic: test timed out after 10m0s
running tests: TestApplyIfaceSysctlsCoversRuleDivertedIface
over code that was neither hung nor wrong. Measured (golang:1.26, 32 cores):
shater/netplane is 1.971 s without -race and 663.762 s with it. A 337x factor
is not "-race is slower".
Nine of netplane's test files intercept nft/ip/ubus/uci/sysctl by re-exec'ing
the test binary as a no-op helper — the standard os/exec trick. Under -race
that child is ThreadSanitizer-instrumented, and TSan's atexit_sleep_ms DEFAULTS
TO 1000: every -race process sleeps a flat second before exiting, on no CPU.
~660 intercepted commands, one second each. The per-test times said so out
loud — 12.17 / 13.18 / 14.17 / 129.62 s — they were counting, not measuring.
Isolated, five runs each, of a `func main() {}` with nothing in it:
built plain 0.0014 s/run
built with -race 1.010 s/run
built with -race, sleep disabled 0.008 s/run
So the children now run with GORACE=atexit_sleep_ms=0, set once in a package
TestMain rather than in each of the nine fakes (they all build the child env as
append(os.Environ(), ...), so one assignment covers the ones written later too).
TSan reads GORACE at process init, long before TestMain, so the detector of the
test process itself is untouched; only the children see it, and they do nothing
but write a canned string and exit. Proven, not assumed: a deliberate data race
in netplane is still reported under -race with this in place.
shater/netplane 663.762 s -> 10.625 s (203 === RUN and 128 top-level
verdicts on both sides)
shater/devices 28.412 s -> 0.358 s (same disease, same cure)
gate [4/7] end to end: was a 600 s timeout, now 56 s
WHAT THE GATE ITSELF WAS MISSING. A deadline and a failed assertion both exit
non-zero, and this script printed the same "FAILED [race]: go test exited 1"
for both — so the reader could not tell "the product is wrong" from "nobody
knows yet". [2/7]/[4/7] now name a timeout as a TIMED OUT, list the tests that
were still running, print only the goroutine dump instead of a quarter megabyte
of PASS lines, and spell out the two opposite fixes (a block, or slowness that
must be MEASURED first). Verified both ways: a sleeping test reads TIMED OUT, a
t.Fatal still reads FAILED.
The deadline stays at go test's own 10m, now written down with the measurement
beside it, and stays there as the hang detector — the slowest package under
-race is 18.9 s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Measured on the stand, real LAN clients in netns behind veth, counters in a
separate nft table: in the state this band predicts, 0 packets left the WAN
across the whole run, against 27 in the control that differs only by one added
catch-all rule. Except for exactly 2 — both plaintext UDP/53. So the band's
detail ("nothing reaches the internet") was wrong by those two packets, and its
own DNS step, which calls that lookup the one thing that still leaves, was
right. One word: nothing ELSE reaches the internet.
The DNS step was also behind reality. It named only the lookups devices send to
the ROUTER, but the shipped dns_intercept='1' pulls a query aimed at a resolver
the device picked for itself into the engine too, answers it there, and it
leaves in the same clear UDP/53 — measured both ways, each producing its own
plaintext packet on the WAN. Encrypted DNS is not the way out either: :853 out
of the LAN measured connects=0, because the plan rejects it. The generator's own
critical warning (generate/dns.go) has said all of this for as long as it has
existed; only the panel had fallen behind it.
"takes ... and drops it" is untouched, and measured: the engine accepts on the
local tproxy socket in ~100 us even for an unreachable address and then closes,
so the client gets an immediate ECONNRESET rather than a hang. "Blocks" and
"ignores" would both be less accurate. Nothing is added about ping: the stand's
ICMP probe was 100% loss in BOTH states, so it proved nothing either way.
The test is the point. The two halves live fifteen lines apart and each reads
fine alone, so a wording fix does not survive the next editor. The new test
checks the INVARIANT instead: the band is flattened to clauses and no clause may
claim that nothing leaves while another names something that does. Its detector
is proved on a fabricated band first (a prior that cannot fire measures
nothing), and it asserts the no-resolver band really does contain a clause
admitting the leak, so the check cannot pass by finding neither half.
Mutations, all caught: detail back to "nothing reaches" -> the invariant fails
and prints both clauses verbatim; DNS step back to the router-only wording ->
the resolver test fails; DNS step stops admitting the leak -> two tests fail.
Control: with one resolver configured the DNS step is absent and no clause
claims anything leaves; emitting the step unconditionally fails that control.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
matchLog's field list is documented as positive and closed, and the new
`status` is deliberately outside it for the same reason `outbound_kind` is: it
is a fixed vocabulary word, so q=failed would silently match every failed row
while the operator was looking for text. The test row now carries a Status, so
the assertion is not vacuous — a filter block IS an answer, which is also the
Status/Error invariant this row models.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
PART A is the frozen v0.1 survey, so `CurrentSchemaVersion=1` was archaeology
that happened to be correct about the branch it describes — and directly
contradicted the live schema subsection thirty lines below, which says
`shaterd migrate` writes 2. Anyone skimming the file map for "what is the schema
version" got 1. Say whose number it is, name v0.2's (2, steps {0->1, 1->2}), and
name what migrate1to2 did, since that is what the reader is usually after.
Also documents the `shaterd migrate` reporting contract in PART B: the closed
classification, the two non-syslog channels a failure reaches the operator on
with globals.log_syslog=0, and why 30_shater-core still exits 0 after one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
dnstrack.QueryEvent carries Failed and Error; stats.LogEntry carried neither.
A SERVFAIL, a timeout, a loopback or a rejected-cached lookup was therefore
written into the query log with action "pass" — or, when the resolver that
timed out had a detour, with the flow-coloured "proxy" — blocked=false, and
nothing anywhere saying no answer was produced. The daemon already knew, one
event at a time: TotalStats.Failed is counted from that very fact in the same
function. The row threw it away, so the aggregate said "N failed" while every
row said everything was fine.
LogEntry gains two fields:
Status — closed vocabulary, "answered" | "failed" | "" (NOT RECORDED), same
discipline as OutboundKind/RuleKind. It is a separate axis rather than a
fourth Action value because Action says WHICH PATH the lookup took: a query
that went out through a detour and then timed out is action=proxy AND
status=failed, and folding the two would erase the one fact that says
whether the tunnel is what broke. It is also what an old panel would have
silently mapped back onto "pass" through its own open fallback.
Error — the producer's own cause text, verbatim, meaningful only when
Status=="failed". No grading is invented on top: three of the four causes
are fixed literals ("loopback", "rejected (cached)", "rejected") and the
fourth is the transport's err.Error(), which cannot be classified without
guessing. "failed" with an empty Error is honest and reachable — the lookup
failed and the cause was not recorded. What IS derivable stays derivable:
Rcode separates "no response at all" (-1) from "the server refused".
queryStatus is a closed POSITIVE list over the sources a producer emits; an
unlisted or zero Source falls to "" (not recorded), never to "answered". The
aggregate is untouched: blocked/failed are computed once in handleEvent and the
row is labelled from those same two values, so the counter and the row can
never disagree and nothing is counted twice.
Cost: LogEntry 152 -> 184 B on 64-bit (+6.4 KB at the default 200-row ring).
Status is a package constant, so its body costs nothing; Error is interned in
its OWN table (maxErrKeys=128, clamped to 160 B) rather than the rule table,
because the transport's error text embeds the queried name and a flood of
distinct causes would otherwise keep clearing the routing-text table.
Tests: every assertion mutation-checked, and the control is three-state — the
same instrument separates answered from blocked from failed, with the
aggregate pinned to identical totals across the change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Both call sites swallowed the result. /etc/uci-defaults/30_shater-core ran
`shaterd migrate >/dev/null 2>&1` — stdout, stderr AND the exit status gone, so
a refusal was indistinguishable from a success on the one screen the operator
who caused it was reading. /etc/init.d/shater logged, but with a single sentence
that described only one of the outcomes: "routing rules that still carry the
removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space
on /overlay and re-run". On a DOWNGRADE every clause of that is false — nothing
is disabled, /overlay is not the problem, and re-running never helps, because
the fix is to put the newer package back. A confident wrong diagnosis costs more
than no diagnosis.
The outcome is now classified with a CLOSED positive list — ok / downgrade /
unreadable / failed — and the last rung is the point of it: an unrecognised
failure says it is unrecognised and quotes the binary verbatim instead of being
reported as one of the causes we can name. `downgrade` is recognised by the
substring "newer than this build", which both model.migrateWith's refusal and
model.ErrSchemaTooNew contain; that seam is a contract and is now pinned.
log_syslog=0 is honoured, not worked around. It is a statement about the syslog
stream, not a request to be left uninformed, so failures go to two channels that
are not syslog: the script's own stderr (the operator's terminal on a hand-typed
restart; the package manager's output inside `apk add`), and
/etc/shater/migrate-failed on flash — written on failure, REMOVED on the first
success, so its absence is the honest all-clear. syslog gets the same line when
log_syslog allows it. A migration that SUCCEEDED stays routine.
uci-defaults still exits 0, deliberately: a uci-defaults script that does not is
kept and re-run at every boot, and this one re-runs a detached enable+restart of
shater/shater-cron plus a firewall reload — one recoverable failure would become
permanent boot-time churn, to carry a status nothing reads. The retry that
matters already exists in start_service, which runs the migration every start.
Found by mutation while writing the tests: reverting start_service's call site
left every other test green, because they all call shater_migrate directly. The
reporter would have been perfect and unreachable. TestInitScriptStartServiceUses-
TheReporter closes that.
Verified: sh -n and busybox ash -n on the target (ImmortalWrt 25.12.1 r37978),
the classifier exercised there under busybox ash against the real uci; six
mutations rolled back one at a time, each caught by name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
renderDiag took a version through its seam and then ignored it, reading
constant.Version directly — so the one field the dead-daemon test could have
pinned was the one field it could not see change. The bundle now prints what it
was given, and the test asserts the value and not just the heading.
Also names the cost the schema gate adds: model.readDiskState's own comment says
"this runs once per write", and it now also runs once per apply, i.e. once a
minute from shater-cron — one `uci export shater` fork and two parses of a few
kilobytes. It reads the DISK rather than m.Globals.SchemaVersion deliberately:
m need not have come from disk (rollbackTo hands in an in-memory snapshot), and
the question is about the file this build would have to live with.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`fetch_via=proxy` with an empty `fetch_detour` is not "proxy, details to
follow". Applier.HTTPClient hands "" to resolveVia, which passes it through
(it is not a `chain:` selector), engine.ViaToTag maps "" to the tag `direct`,
and the feed is dialled through the box's direct outbound — over the ordinary
WAN, with the router's real address, merely from inside the daemon process
rather than from the CLI. Nothing fails. The subscription provider, the party
`fetch_via=proxy` is chosen to hide from, sees that address on every
scheduled refresh.
The picker exists and defaults to Direct, so the state is not "unconfigured";
it is "configured, and silently equal to direct". The message opens on that.
Critical, by this file's own rule at the top — protection the operator
CONFIGURED is not in effect — and by consistency: criticalMarkers already
grades the identical disclosure critical when generate says it about DNS
("in the clear", "your provider sees", "leaves over the plain WAN with your
real IP address").
A detour that names nothing is a SEPARATE finding at `warning`, because it
has the opposite consequence: resolveVia or the engine refuse by name and
UpdateSubscription returns the error rather than falling back, so nothing is
disclosed — what breaks is the refresh, loudly. One sentence for both would
send the operator to fix the wrong thing. A bare name that is really an
egress or a chain gets its own text giving the spelling that resolves, rather
than a false "nothing answers to that name".
Filed under section `subscription` + the sub's own name, which the panel
already routes to that row (Nodes.tsx entityFindings/findingsByName) and to
Overview. The severity is part of that binding, not just the volume: `info`
is filtered out of entity routing on purpose, so it would never reach the
row — recorded at the constant.
Also completes the FetchDetour contract in model.go, which listed neither
`chain:X` — the form apply.resolveVia has a dedicated branch for — nor what
"" actually does.
Verified: 9 mutations, each reverting one part, each caught by a named test;
controls show the same instrument silent for a resolved detour, for
fetch_via=direct, and for a subscription that is disabled or has no URL.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
model.WriteUCI already refuses to write a config whose schema is newer than the
build, so a downgrade can no longer eat the file. What was still open was
RUNNING one. ParseUCIExport ignores options it does not recognise — silently —
so a v3 config read by a v2 build yields a Model with the v3 settings simply
absent. The engine starts perfectly happily and routes traffic by a policy
nobody wrote. Nothing said so: cmdRun never called Migrate(), /etc/init.d/shater
calls it, logs one daemon.err line on failure and starts us anyway, and that log
defaults to a tmpfs file globals.log_syslog='0' can switch off entirely.
REFUSE OR START — and why refuse. Both sides, weighed by "the default falls to
the recoverable side":
* REFUSE. With kill_switch=closed the fail-closed plane goes up and LAN->WAN
forwarding stops. Loud, immediate, impossible to miss. SSH, LuCI and the
panel stay reachable, nothing on disk changes, and reinstalling the build
the router ran ten minutes ago puts everything back exactly as it was. The
damage is an outage the operator caused themselves and can undo.
* START ANYWAY. Traffic the missing rules were meant to tunnel leaves through
the plain WAN with the router's real address on it, and nothing announces
it. That is not recoverable in the sense that matters — the disclosure has
already happened. It is the same choice `sub update` made when it was given
FAIL over a silent direct fetch.
So: refuse. But the daemon does NOT exit and does not crash-loop — a refusal
nobody can see would be the third bad option. It stays up, keeps serving the
panel and the control socket, and says why in three places:
1. apply.schemaDowngradeGate refuses every apply (step 0 of applyLocked), with
the engine-swap failure policy of step 2: a previous engine that IS running
a config this build understood is left alone; with no engine, holdLocked
installs the fail-closed plane — and honours kill_switch=open, which is the
operator's documented fail-open choice and may not be quietly overridden.
This is in applyLocked and not only in cmdRun on purpose: cron reconciles
once a minute, so a gate that only ran at startup would be bypassed sixty
seconds later.
2. Status carries the PAIR: schema_version (disk) and schema_supported
(model.CurrentSchemaVersion). Either alone is unreadable — the panel
already showed the disk version, and "v3" next to a build that understands
v2 looks entirely normal. The difference IS the fault. schema_supported is
a compile-time constant and is therefore set even on the offline stub, i.e.
on the daemon most likely not to be answering. A critical warning naming
the downgrade is computed at READ time, because in this state no apply can
succeed and "the warnings of the last successful apply" would be empty.
3. cmdRun consults model.Migrate() before reading the config (so a bare
`shaterd run` gets the gate too) and classifies the outcome with
CheckConfigWritable: ErrSchemaTooNew is the downgrade, anything else is an
ordinary migration failure and is NOT reported as one.
Only ErrSchemaTooNew blocks. ErrUnmigratedConfig — schema-v1 dst_domain/dst_ip
leftovers — must not: the init script documents starting anyway with those rules
disabled, and turning that into a blackout would be a regression.
Also in this pass, reported by the coordinator: standing_state_test.go's
noGatewayFinding claimed to be quoted from netplane "so the test breaks if that
warning is ever reworded". It cannot — the string never leaves this package and
netplane.noGatewayWarning is never called — and the claim was already false when
it was read: netplane's text has since gained "over IPv4" and an IPv6 clause
while every test here stayed green. A fixture that advertises a guarantee it
does not provide is worse than one that advertises nothing. The comment now says
what it is, and netplanechannel_test.go pins the two couplings that are real:
the severity comes from the CHANNEL (warningFromText(t, "interface",
SeverityCritical), no classify pass), so no rewording can demote it — with the
control that the same texts on the generate channel are NOT critical — while
Section/Name DO come from the `kind "name": ` prefix, asserted in both
directions. A genuine text link is one exported helper in netplane away and is
left to whoever owns that file.
Verified: 13 seeded mutations. Twelve killed by named assertions; the
thirteenth SURVIVED — the guard in schemaWriteVerdict could not be seen, because
on a build host model.CheckConfigWritable answers nil for everything, so the
test reported success whether the guard was there or not. The checker is now
injected and both directions of that guard are killed. Every schema assertion is
walked over all three relations (disk newer / equal / older), so nothing here
passes by always answering the same way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
A device could only carry hand-typed domains. It can now also reference the
`config blocklist` / `config allowlist` sections by name, which brings geosite
categories and url lists to parental control for free (Device.Blocklists /
Device.Allowlists — the Go half is landing separately).
The composition problem was the order. The engine decides a name in five steps —
allow typed, block typed, allow attached, block attached, network filter — so the
typed lane and the attached lane of ONE control are two steps apart, with the
other control's lane in between. Two controls therefore cannot show the order by
position. The card draws it instead: a five-stop rail, lit per step where this
device actually has something, and the same step number stamped on each lane
inside the two pickers.
ListPicker is a new component rather than a generalised SrcPicker: that one is
welded to useSrcOptions(), to CIDR validation, and to an empty state reading
"everyone · all LAN clients", which on a block list means the opposite of the
truth. It reuses SrcPicker.css and its whole interaction language.
Honesty, in three places it would otherwise have lied:
- a list chip reports what /api/ruleset/status says, not that someone attached
it. Never fetched reads "not loaded" in crit, an empty one "empty", one the
engine has not mentioned "load unknown" — dim, never green. A name the config
no longer has reads "no such list".
- attaching a list is itself the switch for that device, so a list with
Enabled=0 is NOT drawn as dead, and the DNS page's "configured but inactive"
is replaced by a sentence naming the devices still running it. A row for such
a list now reads "N devices only" instead of "off"/"inactive".
- an attached allow list is terminal, so it lifts the network blocklists off
everything it covers. Said in the picker at the moment of choosing and again
on the card.
cleanDomain demanded /^[a-z0-9.-]+$/, so a colon could not be typed and the
engine's own full: / suffix: / keyword: vocabulary was unreachable from the
panel. parseDomainEntry accepts them from a closed positive list and refuses, by
name, the three shapes the engine silently discards: an unknown `word:` prefix, a
marker with no value, and an IP. The keyword case gets its own message — an empty
keyword is strings.Contains(host, "") and would take the device off the internet.
The logic lives in src/deviceLists.ts with tests, since `node --test` cannot load
a .tsx. Each test was mutation-checked, and the load reading is shown giving both
a positive and a negative result.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
A2 — the panel unlocked the byedpi egress type on LookPath("ciadpi"), which
answers "is the package installed" while the operator is asking "will traffic
sent here go anywhere". Those come apart on the SHIPPED configuration: the
packaged /etc/config/byedpi is inert (enabled='0'), and the port is coordinated
between the two packages by comment only — nothing in the daemon had ever read
that file. Result: type unlocked, egress on 127.0.0.1:1080, apply green, nobody
listening.
shater/panel/byedpi.go now decides on three separate facts (binary, enabled
instances + their ports read from the conffile, a TCP connect to each) and
reports a CLOSED state: unknown | not_installed | disabled | not_listening |
listening. Only "listening" may gate the egress type. GET /api/byedpi adds the
per-egress port reconciliation, so a mismatch is NAMED with both numbers instead
of going quiet. Nothing overclaims: the check is a connect, not a SOCKS5
handshake, and every sentence says so. A connect that is neither accepted nor
refused is "unknown", never "no".
C4 — a just-added node had no instrument: the group test reads the observatory's
board, and the observatory only probes what the rules route through, so the one
question a fresh node exists to ask ("is it alive?") answered "no rule routes
through it". POST /api/groups/test now takes {"kind":"node"} and runs the SAME
instrument — same singleton, same runner, same result type, same status poll,
same exit_ip through the target's own outbound with the same refusal to answer
from `direct`. The only addition is one fallback: a node the observatory does not
cover is measured once, here, through probeOneInto (the observatory's own
dialler), recorded under its own tag alone. A node whose base tag is a plan STORE
ALIAS — the egress-bound-group case — is NOT dialled: the board already holds its
egress-path number, and a bare-WAN measurement filed there would be the same
poisoning one layer down.
Results gained kind (group|chain|node|"") and source (observatory|on-demand|""),
so "nobody measured this" is distinguishable from "measured and dead".
Also: PUT /api/config maps model.ErrSchemaTooNew to 409 beside ErrUnmigratedConfig.
A downgrade refusal is the guard working, fixed by the operator, not by us; 500
sent the reader to the daemon log.
Every new test was verified by mutation (14 mutations, each killed by name), and
each detector has a control: the byedpi probe is shown seeing a real loopback
listener AND its absence with nothing else changed, and the node test is shown
telling a live node from a dead one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two holes from the operations audit, and I can confirm both of its readings.
CONFIGURATION HISTORY. Applier.lastGood and Applier.snapshot are fields in
this process. A daemon restart or a reboot loses both, and Rollback with no
snapshot goes to rollbackEngineAndPlane, which re-reads the CURRENT
/etc/config/shater — that is, it re-asserts the config that broke. With
confirm_timeout at 0 by the owner's choice there is no auto-rollback either,
so "what did the working config look like?" had no answer at all once the
daemon had restarted. Nothing on this router kept one.
Every successful apply now files RenderUCIExport(m) — the existing pure
function, not a second serializer — into /etc/shater/history/<unix>-<version>.uci.
* DEDUPLICATED against the newest entry. shater-cron reconciles once a
minute and every reconcile runs applyLocked to completion, change or no
change, so a file per apply would be ~1440 identical writes a day onto
overlay flash and would fill the ring with twenty copies of one config
twenty minutes after the last real edit. One file is now one change.
* 20 files / 512 KiB total / 128 KiB per entry, hard ceilings, not defaults.
The shipped /etc/config/shater is 12.6 KB of which 459 bytes are actual
configuration; a loaded one renders to a few KiB up to low tens of KiB, so
twenty entries normally cost 50-200 KiB and the byte cap binds only for
inline entry lists. Against what this product already grants itself on the
same overlay — 4 MiB of compiled lists, an 8 MiB rule-set cache, a stats.db
defaulting to 64 MB — 512 KiB is a rounding error. An entry over the
per-entry ceiling is REFUSED rather than allowed to evict the whole ring,
and the refusal is reported.
* 0700 dir / 0600 files. Checked, not assumed: the Makefile installs
/etc/config/shater with INSTALL_CONF, i.e. 0600 root:root, and these files
carry the same node credentials and subscription URLs.
* A failure NEVER fails the apply, and is never swallowed: it becomes a
Warning folded into the set Status publishes (gather + append + finalize,
the seam abortAfterSwap already uses), so the panel says the history has
stopped instead of the directory quietly going stale.
* NOT kept across sysupgrade. The audit's premise that /etc/shater is in
keep.d is wrong — keep.d/shater-core lists four specific paths, not the
directory. Excluding it follows model.backupBeforeChange's existing
precedent for config.pre-v*.bak: the archive is held in RAM across the
flash and routinely ends up in cloud storage, and this is a local undo for
changes made on THIS box.
`shaterd diag`. The only thing this product could hand over was
GET /api/log?range=, served by the daemon — so in a crash loop the one channel
that does not need ssh dies with the process. `shaterd diag` prints version,
our packages from `apk list -I`, status, `nft list table inet shater`,
`ip rule`, the log tail and the configuration, as one block, collected entirely
by the short-lived process.
* It works with a DEAD daemon, which is the case it exists for. The status
section falls back to the same offline stub `shaterd status` prints and
LEADS with the fact that no daemon answered, so an empty-looking section
can never read as a healthy one. No section is ever silently absent: a
missing nft/ip/apk produces "NOT COLLECTED: <reason>", and `uci export`
failing falls back to the raw file and says so.
* Masking is a POSITIVE, CLOSED list of the fields that may be PRINTED
(diagSafeUCI), keyed by section type. Everything it does not name is
masked — an unknown option, an unknown section, and every field added to
model.Model after this build. That is the direction the open `default:`
lesson demands: the recoverable side is "hidden", not "shown".
TestDiagMaskingIsClosedOverTheWholeModel proves it by reflection over every
string the model can render, with the control that the same instrument sees
those values in the unmasked text.
* A second layer scrubs the refused literals from the WHOLE document, because
masking the config alone would only move the leak: the daemon prints a
subscription URL into its own log on a fetch failure.
* node.uri keeps its scheme and nothing else — "is this node vless or
wireguard" is most of the diagnosis and a protocol name is not a secret.
Verified: 14 seeded mutations, every one killed by a named assertion (dedupe
removed, prune removed, ceiling removed, 0600->0644, 0700->0755, failure
swallowed, warning not folded, version not sanitized; allow-list defaulting to
ALLOWED, uri scheme dropped, scrub removed, failed sections made absent, stub
banner removed, masking removed). Every check is paired with its control — the
ring tests assert the newest entry is present and correct, so "the old one is
gone" cannot be satisfied by a ring that silently stopped writing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Four fixes on one path — the one a person actually walks when a site does not
open: Insights -> DNS log -> Connections. Three of them were fields the daemon
already put on the wire and the panel dropped on the floor.
1. Connections shows the routing record. ConnLogEntry gains rule_kind/rule/chain,
with the discipline stats.go wrote them under: "" is NOT RECORDED and can
never be drawn as "no rule matched". Three states, three readings, and a
CONTROL test that fails if any two of them render alike. The outbound path is
printed rule-named-tag first, dialling-outbound last (the wire order is the
reverse).
2. The DNS log says where the lookup left. outbound_kind is a closed four:
detour (tag named) / default (the resolver names no detour -> the query went
out the plain WAN, past the tunnel; marked amber) / local (cache, optimistic
answer, filter block: nothing egressed) / "" (not recorded). An unrecognised
value falls to "not recorded", the recoverable side, not to one of the answers.
3. Both logs take q=. The daemon filters inside the store on the same walk as the
cursor, so limit counts MATCHING rows. The searched fields are named under the
box, because a POSITIVE CLOSED list is also a statement about what is NOT
searched: no ports, no rule_kind, no outbound_kind — q=default matching every
default-egress row would be a trap wearing a filter costume. logRoute mirrors
filter.go exactly so the ?mock backend finds and misses what hardware does.
4. A TRUNCATED page is not the end of the log. A filtered walk is budgeted
(MaxFilterScan); a page that ended on that budget is short for a reason that
has nothing to do with how much data exists. X-Stats-Log-Truncated is now read
and the state is NAMED — an amber "Scan stopped" plate, the empty text saying
"not the end of the log" instead of "nothing found", the count line refusing
to say "all loaded", and the daemon resume cursor behind a button. The cursor
matters twice: a truncated page can have ZERO rows, so there is no row seq to
page from, and the live tail now advances on rows EXAMINED rather than rows
matched — a filtered after= poll that matched nothing used to rescan the same
window every tick forever.
Also: .fp-select gets max-width:100% + min-width:0. A <select> shrink-wraps to
its widest option and, as a flex item, refuses to shrink below it: the geo
provider label measured 501px in a 375px viewport and gave the PAGE a horizontal
scrollbar (scrollWidth 559 vs clientWidth 375, measured). Settings.css and
Networks.css each carried a narrow copy of this fix; the component is the right
place. Verified on an isolated harness with no page-local CSS: bare select
overflows a 320px row at 438px, adding the class alone brings it to 320/320.
Tests: 34 new, every one mutation-verified — unrecorded folded into default /
into local, rowMatches returning true unconditionally, outbound_kind added to the
searched fields, historyExhausted ignoring truncated, logEndNote drawing both
situations with one sentence, logCountLabel saying "all loaded" on an incomplete
scan. Each revert reproduced its own failure text. Both search directions are
covered (finds / does not find), which is what catches a filter that matches
everything. Browser-checked at 390 and 1280 in both themes, no horizontal scroll;
the six-click resume walk from "scan stopped" to "Nothing in the log matches" was
exercised live in ?mock.
NOT verified: no hardware or VM run — the truncated state was exercised against
the mock backend, whose scan budget is 60 rows where the daemon uses 20000.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Measured on the production BPI-R3: egress `ewan` was carrying the entire
household's traffic (chain default, plane full, verdict tunnel, 15 hours up)
while the panel showed, at CRITICAL, "this egress CANNOT REACH ANYTHING outside
its own subnet — every node, group and rule bound to it will fail to connect".
table 8208 held `default via 10.0.0.1 dev eth1`; the IPv4 half was perfect. eth1
holds one address, fe80::.../64, and the ISP publishes no IPv6, so
`ip -6 route show default` is empty router-wide. The -6 pass found no nexthop
and one family-agnostic text declared the whole egress dead.
Two defects in one line. A per-family fact was stated as an absolute, and the
absence of an optional ISP feature was graded as an outage — in the loudest
register this codebase has, on a channel apply grades critical wholesale. Red
that stands for fifteen hours over a healthy router is not a warning.
IPv4 stays loud and unchanged in substance: an uplink with no IPv4 nexthop
carries nothing. It now scopes its consequence to IPv4 and says outright that
it is not describing IPv6.
IPv6 splits on one piece of evidence — does the device hold a global IPv6
address? If it does not, IPv6 is simply not provisioned on this link: nothing
is broken, nothing leaks (the v6 mark keeps its own table and its unreachable
floor, so it cannot fall through to main), and there is nothing the operator
can do because the missing thing is upstream. Silent. If it does, IPv6 is
configured and the nexthop is missing anyway — a real fault, still critical,
now scoped to IPv6. Silence requires positive evidence: a failed address read
makes us louder, never quieter.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Three things the divert plane got wrong once more than one tproxy inbound
exists, and one it got wrong all along.
Per-rule diverts read `tcp`/`udp`/`tproxy_port` off the FIRST enabled tproxy
inbound and applied them to every device the plan touches. With one LAN — the
shipped shape — first and owner are the same section and nothing showed. With
two, `option udp '0'` on the second inbound was ignored (UDP diverted anyway,
into another section's listener), `option udp '1'` was ignored the other way
(no per-rule UDP line at all, so the rule's counter never ticked for UDP and
Insights showed a rule that appeared never to match), and `option tproxy_port`
pointed at the wrong listener. Same for the dns_intercept :53 lines, which sit
above the fib-local bypass. Each ingress device now resolves to the inbound
that OWNS it; a device no inbound claims still falls back to the primary,
because that is the only listener its traffic can reach. Verified
byte-identical output for every single-inbound shape against the pre-change
renderer.
Rule counters: a counter exists only for a rule the plane emitted a divert
line for, and it only emits them from SOURCE selectors — so a rule written by
domain or ruleset never appears in RuleTraffic at all, and absence there could
not be told apart from "carried nothing". It cannot be measured: which rule a
packet matches is decided inside the engine after the divert, where nftables
cannot see it. So no counter is invented. Instead the plane says which rules it
can measure (RuleMeasures) and what the numbers it does have actually mean —
an upper bound, not the rule's traffic — and the two are pinned to the rendered
ruleset in both directions. Counters are now declared BY the emitting line, so
a rule whose fragments were all dropped no longer leaves a counter attached to
nothing, reading a confident, permanent, false 0 B.
untunnelable_egress could resolve, pass validation and still mark nothing when
the plan has no LAN ingress device — while apply's note, gated on the same
binding succeeding, told the operator that IPsec/GRE/SCTP now leave through it.
The gate is right (the marking rule has no safe unscoped form), the silence was
not; bound-but-inert is now named.
stats/panel do NOT consult RuleMeasures yet — wiring it is a change outside
this package, and the gap is still visible to an operator today.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
A downgrade ate the configuration silently, and permanently. Migrate() refuses a
newer schema, but nothing on the write paths calls it: the daemon starts
regardless, ParseUCIExport reads the options it knows and drops the rest, and
WriteUCI replaces the WHOLE package. So an older build rewrote /etc/config/shater
with only what it understood. The second half is what made it unrecoverable —
schema_version round-tripped through the Model, so the rewritten file claimed the
OLDER version, and a newer build put back afterwards saw cur == CurrentSchemaVersion
and migrated nothing. Nobody had to be at the keyboard for any of it: the profile
watcher looks every 25 s and `sub update` runs from cron every 6 h, and both
persist through WriteUCI.
The mechanism for refusing already existed and already worked in the other
direction (ErrUnmigratedConfig + CheckConfigWritable); this is its second caller,
not new machinery.
- guardSchemaDowngrade refuses the write and the panel's pre-flight when the
config on disk is newer than this build, naming both versions and the way back
(put the newer package on again — the config is untouched). ErrSchemaTooNew so
a caller can answer 409 instead of 500.
- withDiskSchema takes schema_version from the DISK, never from the caller. A PUT
body that omits it sends 0, and a rendered 0 is an OMITTED option: the version
would have vanished and the next `shaterd migrate` would replay every step. A
body claiming 99 would have locked the box out of its own panel.
- backupBeforeChange copies the live config to /etc/shater/config.pre-v<schema>.bak
before the first migration and before the first write — once per schema version,
write-then-rename. A failed backup aborts: the `uci commit` that follows writes
to the same filesystem, so refusing costs nothing that was not already lost, and
best-effort-and-carry-on is the silent skip we keep paying for.
- The reverse direction is fenced by a test: an unmigrated v1 config still refuses
a rule-changing write as ErrUnmigratedConfig, still allows one that leaves the
rules alone, and still keeps its `list dst_domain` and its v1 stamp.
The shipped /etc/config/shater now says what "conffile" actually buys — values
across a package upgrade, not comments across the first write, which happens
without an operator — and the annotated file is installed a second time as
/usr/share/shater/config.sample, where nothing rewrites it.
INSTALL.md gains the downgrade procedure. Measured on the testbed VM (ImmortalWrt
25.12.1 r37978, apk-tools 3.0.5) against the real apk-v0.2.9/v0.2.10 feeds in an
isolated --root sandbox: `apk upgrade <named>` does not downgrade at all;
`apk add <pkg>=<ver>` does, and leaves a pin in world that a later upgrade obeys;
`apk upgrade -a` downgrades too but took four unrelated packages with it.
Tests in shater/model/schemadowngrade_test.go; every assertion checked by mutation
(8 mutations, each killed a named test) and every refusal paired with a control
that accepts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Five things the panel knew and did not say, each one a state where the
screen read healthier than the router was.
Rule.Kill was typed, round-tripped and drawn nowhere. `open` sends a
rule's traffic out direct — around the kill-switch, with the real
address — when its target cannot be built, and such a rule looked
exactly like one that fails closed. It now has an editor beside Target
and an amber mark on the row; the fail-closed default draws nothing, so
the two states are not priced alike. An unreadable value is its own
state: it blocks, like the daemon, and the picker re-surfaces it
verbatim rather than rewriting a value it never showed.
Alert channels were write-once for Type/Token/ChatID/URL/Events, so
fixing a typo meant deleting the channel and going back to BotFather for
a token you already owned. Add and edit are now one form. The token box
starts empty and the caption says what empty means — keep, never clear —
because the panel refuses to show the secret and a save may only clear a
field the editor could show. Same rule covers a type switch: the other
kind's settings stay stored and unused.
The add-rule form pre-filled Target=direct. An untouched form is a rule
with no matchers, i.e. the default route, so one press put the whole LAN
on the plain WAN. `block` would only have swapped the leak for an
outage; the recoverable default here is no default, so the form refuses
and asks.
The empty state said "all traffic follows the default route" without
naming it. On a fresh install that route is `block` — the LAN has no
internet — and this is the page the kill-switch alarm sends people to.
Both it and the lead now name the route in force.
The interception board was computed from the config alone and lit `lan`
green over a stopped engine. Green now needs the engine up AND the full
plane; a hold plane blocks rather than carries, and unknown is an unlit
socket.
Also: four rungs of the untunnelable copy claimed traceroute works. It
prints `* * *` and no hops on every setting — the wording is now
apply/warnings.go's own udpTracerouteFacts, said once.
Tests: killPolicy / alertEdit / defaultRoute / intercept, 38 cases, each
mutation-checked (16 mutants, all caught). Browser-verified at 390 and
1280, no horizontal overflow.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Five groups of real daemon settings had no control in the panel, so the
only way to change them was to edit the config over SSH. Each one is now
editable where it belongs, and each editor is built so it cannot lose a
field it declines to display.
DNS list refresh intervals. Blocklist.UpdateInterval was hard-coded to
"24h" in two places and shown nowhere, while the row displayed the
interval the ENGINE reported — a readout dressed as a knob.
Allowlist.UpdateInterval did not exist in the panel at all. It matters
because an allowlist is how a blocklist false positive gets corrected: one
pinned to a day delivers the fix up to a day after the site broke.
Geo data. GeoProvider and the four URL fields are consumed for real
(generate.SetGeoProvider, /api/ruleset/categories) and the panel USES the
data they pick, while offering no way to choose it. New Settings group
with the closed five-provider list, the custom {category} templates and
the two category indexes. Only `custom` reads the templates, so only
`custom` renders them; an unknown provider is preserved and marked rather
than silently rewritten to auto on page load.
Subscription filters. Include/Exclude/FilterProto/FilterCountry/Dedup,
Format, ExpireAlertDays and the three device-identity headers are now
editable — the same five filters a group already offered over its members,
applied one step earlier. 376 nodes can become the four Dutch ones without
SSH. ExpireAlertDays keeps its three states (blank = the 3-day default,
"off" = -1) instead of being flattened.
Edit-after-create. Blocklists, allowlists and resolvers could be
configured only at creation; a typo in a URL meant delete and rebuild, and
deleting a resolver clears whichever global slot it filled. Every one now
has a row editor. `file` and `geosite` sources are offered when a list
already IS one, so opening a list the panel cannot create never becomes a
way to destroy it.
Stale local type copies. DNS.tsx and Settings.tsx carried local
Blocklist/Allowlist/Globals extensions whose comments claimed api.ts did
not type those fields; api.ts had typed them for a long time. Deleted —
the note was an invitation to declare the next field twice. (Egress.Target
was already gone.)
Along the way, three defects the work surfaced:
* parseDomains cut comments per TOKEN, so pasting "# ads and trackers"
contributed ads, and, trackers as three real blocked domains. Cut per
line now.
* FetchVia=proxy with no FetchDetour resolves to the tag `direct`
(engine.ViaToTag), so the feed is pulled over the plain WAN and the
provider logs the router real address — the one thing `proxy` is
chosen to hide. The row said "via proxy" for it. It now says
"proxy - no route" and the editor carries an amber explanation. The
picker also gained chains, which apply.resolveVia supports for real
and the picker excluded with a comment that misdescribed the contract.
* Adding a subscription only saved it. apply does not fetch, and
shater-cron is inert unless globals.enabled=1 AND the service is live,
so on a router not yet switched on nothing would ever fill it — and
the only Update button sat at the bottom of a collapsed panel. Adding
now fetches, reported separately from the save, and every row carries
Fetch now. A row with no nodes says what to press.
Nodes also gained the forward link nothing had: a node is not something a
routing rule can point at, and no page said so.
The merges live in subEdit.ts / dnsListEdit.ts / geoProvider.ts because
`node --test` cannot mount JSX. The rebuilt shapes return Complete<T>, so
a field added to api.ts fails the build in the function that has to decide
about it; the subscription merge extends instead, because five of its
fields are provider-reported state no control can show.
Verified: npm run build green; 173 tests pass; 18 mutations each killed a
named test and a probe field added to Allowlist broke the build inside
nextAllowlist; zero Cyrillic in panel/src; Chromium at 390 and 1280 with
no horizontal overflow (the detector caught a real 559px select spill at
390 before the fix).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The shipped config is `enabled '0'` + `kill_switch 'closed'` with nothing
applied. Every lamp in the panel was derived from what is INSTALLED and none
from whether anything was MEANT to be, so a package that installed exactly as
designed showed a crit master lamp ("Engine down"), a crit kill-switch module
("NOT IN EFFECT") and three crit pips on Apply — at a person who had not done
anything yet. Red that fires on a correct installation is red nobody reads by
the time something is actually wrong.
serviceIntent() is the missing question, and every readout that used to answer
from the installed state now asks it first: off ⇒ unlit socket and a word that
says why; on ⇒ every alarm exactly as before. A positive `off` only — an
unreadable configuration stays `unknown` and keeps its crit, because that is
the state where the LAN really is cut off.
applyRisk() is the other half. Applying an empty config with the service on
and the kill-switch closed sets route.final = block, and the tproxy divert for
the shipped `lan` inbound is installed — so every TCP connection and UDP flow
from the LAN is handed to the engine and dropped. The panel read that state
perfectly once it existed and said nothing before, with confirm_timeout at 0,
so the most dangerous apply this router does ran with no auto-rollback. The
band names the outcome, the missing rollback and the fix, and does not block
the apply.
Insights had a short-circuit for this exact job that never fired: it was gated
on logging being off, and the shipped backend is memory. Ten sections drew ten
well-mannered "nothing yet" states and not one named the switch.
Every test is mutation-checked, and each one is paired with the control that
proves the instrument can still produce the alarm.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`fetch_via=proxy` on a subscription means "pull this feed through the tunnel",
and it is set for exactly one reason: the provider is blocked, or the owner does
not want the provider (and every hop to it) learning the router's real address.
The panel honoured it — Applier.UpdateSubscription resolves fetch_detour against
the running engine — so testing it once from the browser showed it working. The
CLI verb did not: it logged one daemon.warn line and fetched DIRECT.
/etc/init.d/shater-cron calls exactly that verb, so every scheduled refresh and
the fetch-at-boot went out unproxied, and the only trace was a syslog line in a
log globals.log_syslog='0' switches off.
The CLI cannot do this fetch itself — only one process may own the engine — so
it now DELEGATES: a new control-socket verb `sub update <name>` runs the very
same Applier.UpdateSubscription the panel's Refresh button calls. One
implementation, so the two paths cannot drift again.
With no daemon to ask, the subscription FAILS (exit 1) instead of falling back.
The refusal is recoverable — shater-cron does not stamp the item, so it retries
after RETRY_SECS and the already-cached nodes keep working — where a silent
direct fetch is not: the disclosure has already happened. Direct subscriptions
are untouched and still need no daemon at all.
Order is load-bearing: the direct pass and its UCI write run first, then the
delegated ones, because the daemon re-reads UCI and writes back userinfo
counters a later write from this process would silently drop.
Also in this file, reported by the LuCI agent: the offline stub of `shaterd
status` published config_readable=false after a SUCCESSFUL read, telling every
consumer to disbelieve three values it had just read correctly (LuCI worked
around it by reading the field only when a daemon answered), and swallowed a
FAILED read with no trace — the inverted lie apply.Status() was fixed for, in
the one situation that matters most: a full /overlay where "not enabled" tells
the owner they switched it off themselves while the fail-closed plane holds the
LAN shut. Both halves now mirror the live path exactly.
Tests are mutation-verified in both directions, with an instrument that gives a
positive reading for BOTH "went through the tunnel" and "went direct" — a live
origin server and a live control socket in every case, so neither zero is an
artifact of the other endpoint being absent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
sing-tun's classifyReturn answers `returnPass` for any IP fragment, so the l3
return path never judges one. On the WireGuard endpoint a passed packet still
reaches the endpoint's own tun stack; the bridge has no second consumer — both
deliverReturn and the batch read loops offer a packet to each attached return
path and then drop whatever nobody claimed. A fragmented answer coming back
through a bridge outbound was therefore lost outright, 100% of the time.
Fragments do arrive: the return direction is fragmented by the LOCAL kernel
(conntrack defragments at PREROUTING for the NAT lookup, the output path
re-fragments to the bridge TUN's 1500-byte MTU honouring IPCB frag_max_size).
packet.go's fixReturnChecksum already recognises a fragment and declines to
touch it — the path was known to carry them.
frag_reassembly.go is a deliberate sibling of transport/wireguard/
frag_reassembly.go: same algorithm, same ceilings (64 datagrams, 1 MiB, 5 s,
non-refreshed deadline, partial overlap poisons the key), so collapsing the two
into one shared package later is mechanical. They are not shared today only
because the seam that would host the shared type — transport/wireguard/port.go
and its test suite — is owned by other work in flight.
Windows is deliberately untouched: there a fragment never reaches deliver() at
all, because classifyInbound needs a transport header to decide ours/not-ours
and WinDivert reinjects the rest into the host stack. Different function,
different defect, platform we do not ship.
protocol/tailscale gets a comment, not a fix: the one ReturnPackets call that
package makes carries BuildUnreachable replies, which are synthesised whole and
can never be fragments, and the real tunnel return path is upstream
tstun.Wrapper.Write, ahead of every seam this tree owns.
Verified: 23 tests, all 14 seeded mutations killed (including "seam removed" on
both the portable and the Linux batch loop), -race clean on linux/amd64 in
docker and on windows/amd64. The darwin seam is compile- and vet-checked only.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
C1. handleEvent had QueryEvent.Outbound in its hand, used it only to compute
action(), and dropped it. The query log could name the resolver that answered
and not the channel that resolver's own packets took — the one fact an
anti-leak `detour` on a resolver exists to control.
LogEntry now carries Outbound + OutboundKind, on the ConnLogEntry.RuleKind
discipline: "" is reserved for NOT RECORDED, so the three states that all have
an empty tag stay distinct — "detour" (tag recorded), "default" (the resolver
names none, so its packets take the plain WAN), "local" (cache/optimistic/
filter block: nothing egressed at all). An unrecognised Source falls to
unrecorded, the recoverable side. Rows from older builds decode to unrecorded
and are therefore still distinguishable from a recorded no-detour row.
No rule name is attached, and that is deliberate: the DNS path has strictly
less to work with than the connection path did. A DNS *rule* picks a SERVER,
not an outbound, and the event carries no rule identity at all — only the
transport's tag. Inventing one would be a forgery.
Cost, measured: LogEntry 120 -> 152 B (+32 B/row, two string headers on
aarch64). +6.4 KB at the default ring of 200, +160 KB at 5000. Tag bodies go
through the existing intern table (maxRuleKeys=512, shared with the rule text).
C2. /api/stats/log and /api/stats/conns take q=<substring>, applied INSIDE the
store on the same walk as the seq cursor. It has to be there: Limit is applied
by the store, so post-filtering a returned page would hand back 3 rows of a
50-row page and call it a page. Substring, not regex — nothing a client can
type costs more than a linear scan.
Pagination stays honest. A filtered walk must examine rows it will not return,
so it is bounded (MaxFilterScan=20000) — and a page that stopped on that bound
is short for a reason that has nothing to do with how much data exists. That is
reported: LogPage.Truncated + ScanCursor, surfaced as X-Stats-Log-Truncated and
X-Stats-Log-Cursor. Unfiltered requests are untouched: no budget, never
truncated, same walk as before.
The logRing seam now returns LogPage/ConnPage instead of (rows, pending) so the
truncation state cannot be dropped on the floor between the ring and the API.
Tests: all mutation-verified (7 reverts, each reproduced with its message),
including the copying-variant control for the intern table — strings.Clone
passes an equality check and fails the identity check the test actually makes.
Filter coverage is both-directions (finds / does not find) on both backends,
with mem-vs-bolt parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
untunnelablePolicyWarnings opened its egress branch on
`Globals.UntunnelableEgress != ""` alone. netplane refuses far more than a
typo: UntunnelableEgressBinding fails CLOSED for any name that does not
resolve to an interface/tunnel egress WITH a device — nothing is marked in
prerouting, no forward-chain accept is rendered, addEgressRouting installs
no rule and no table, and the `untunnelable` policy decides everything by
itself. The note nevertheless opened with "...now leave through egress
"x": the kernel routes them out that interface", about a carrier that does
not exist; it even printed `(device )` once the device was interpolated.
The tail hedged the case thirty lines later, and the first sentence is what
gets read.
The branch is now gated on netplane's OWN verdict, called rather than
re-derived (apply imports netplane, so unlike model.ValidateUntunnelableEgress
there is no copy to keep in lockstep). That needs the whole model, so
collectWarnings/gatherWarnings/untunnelablePolicyWarnings take *model.Model
instead of model.Globals.
When the option is set and unbound, the note now LEADS with that fact and
then gives the ordinary policy text, because that is exactly what the router
is doing. The bound branch drops "a name that matches no interface/tunnel
egress" from its failure list — that case can no longer arrive there — and
names the device it resolved to.
Two further claims found while checking the rest of the file against the code:
- the `icmp` rung promised "IPTV and VPN passthrough work only toward
addresses your rules route directly". Multicast crosses this router under
NO setting (the stream is WAN-side inbound; a client's outbound multicast
is UDP, which untunnelableFilter structurally cannot match), and the other
three rungs all say so. One true clause was carrying one false one — the
same sentence the `block` note records having removed for being false.
- "\"icmp\" drops it, excepting only ping/echo" understated a leak. `icmp` is
the one rung that walks the destination plan and it ACCEPTS raw ESP/AH/GRE
toward provably-direct destinations, so the operator was told it was
contained while it left with the router's real address.
Ratcheted by untunnelable_egress_honesty_test.go, each assertion with a
control: the bound and unbound halves are walked in one pass, and the IPTV
and `icmp` checks fail if the matrix ever stops producing the notes they
read. The traceroute matrix grew a third egress value (set-and-bound,
set-and-unbound) so the bound branch keeps being walked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The dashboard told a live daemon from a dead one by `plane: ""` — a side
effect of the offline stub being a zero value, not a promise anyone made.
`shaterd status` now states it: daemon_answered, true on the live branch and
false on the stub. The detector reads the field first and keeps the plane test
only as the fallback for the non-atomic update window (new luci-app-shater,
old shaterd). When the two disagree the field wins; a stub carrying a plane
word must still read as "no daemon answered".
Both lists are positive and closed. A daemon_answered that is not exactly
true/false is not a verdict and falls through; a plane word this build does
not know lands in unknown. Nothing lights green or amber on a guess, and the
launcher button is still never disabled.
config_readable was already on the wire and nothing here read it. With it
false, enabled/kill_switch/panel_port are zero values: "inert (disabled)" and
a green "closed (fail-closed)" were being rendered out of placeholders, in the
one situation — a full /overlay, an interrupted commit — where the fail-closed
plane has the LAN cut off and the owner is told they did it to themselves.
Those rows now say "not known", a Configuration row carries the daemon's own
reason and its don't-switch-anything-off warning, and an absent nft table is
no longer softened to amber by an `enabled` nobody could read.
The field is consulted ONLY when a daemon answered: the offline stub reads UCI
directly and never sets ConfigReadable, so its false is a zero value while its
enabled/panel_port ARE real reads. Taking it at face value would put "could
not be read" on screen for a readable file. Same reasoning drops plane,
traffic and hash on the stub branch — the contract calls them placeholders.
panel_port is CONFIGURED, not bound: shaterd logs a panel bind failure and
carries on, and SHATER_PANEL_ADDR can switch the server off while the port is
still reported. Nothing measures a listener, so the hint, the tooltip and the
new Panel port row say the port is configured rather than checked, and its
lamp stays unlit even on a healthy router.
tests/status-readout.test.js grows the new cases and now runs under gate step
[7/7]. Mutation-checked four ways against copies: dropping the
daemon_answered branches fails 7 assertions by name, dropping the plane
fallback 5, reading config_readable without the daemon gate 5, and rendering
the placeholders as readings 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
This change was already in the working tree when this session started; it is
committed here because it is load-bearing and an uncommitted load-bearing file
is a trap.
netplane.L3SlotFor asks the kernel through `ip link show` and reclaims through
`ip link del`. golang:1.26 ships no iproute2, so in the docker re-exec lane
every slot read as FREE, TestIntegrationL3StaleSlotIsReclaimed stood itself
down rather than pass while proving the opposite of what it claims, and [5/7]
then failed the gate — correctly, since this environment HAS root and
/dev/net/tun and the capability guard is therefore not what skipped it.
Installing it is also what made the concurrent-namespace defect visible at all
(see 06c04c157): with no `ip` on PATH, no `ip link del` was ever issued and the
two test binaries that were destroying shater/generate's TUN looked innocent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`go test` runs package binaries CONCURRENTLY and every one of them shares the
host's network namespace. netplane.L3SlotFor is destructive by design — it
DELETES a candidate slot it finds occupied rather than waiting for it — and
netplane.removeL3Devices deletes both slots unconditionally. Two test binaries
reached those for real:
shater/engine l3slot_test.go calls l3RetargetForNext for its return value
shater/apply Applier.Teardown -> netplane.TeardownRouting -> removeL3Devices
Measured with an `ip` shim on PATH inside the gate container: apply.test issued
9 `ip link del shater-l3a` + 9 `ip link del shater-l3b` per run, engine.test one
per l3slot test — into the namespace where shater/generate's privileged tests
were holding a live TUN. From the other side that is
post-start inbound/tun[l3-in]: starting TUN interface: find tun interface: Link not found
no [shater-l3a shater-l3b] device exists after a successful Start
i.e. an intermittently red [2/7]/[4/7] in a package that did nothing wrong,
while [5/7] — which runs only `^TestIntegration`, so neither binary reaches the
slot code — passed the very same test seconds later. It only became visible when
iproute2 was installed into the gate container: without `ip` every slot read as
free and no deletion was ever issued.
Not a product defect. shaterd is one process with one engine; the running
generation's slot is excluded before anything is deleted, and nothing else on
the router calls L3SlotFor.
The kernel is faked rather than the CHOICE: making the engine's tests stub the
slot answer would delete the only place the ENGINE checks that the running
generation's slot is excluded, which is the invariant the production outage
violated. netplane.L3StubKernelForTest points the two kernel operations at an
in-memory set; engine and apply install it from TestMain (forget-proof, unlike a
per-test helper whose omission fails in a different package on some runs only).
netplane's TestL3StubKernelTakesTheSlotChoiceOffTheKernel is the control, in
both directions: stubbed, nothing reaches the exec seam; restored, the same call
does.
Mutation: with the engine TestMain reverted, the generate binary's
TestIntegrationL3* failed 8 of 8 runs beside a loop of the engine binary; with
it, 0 of 8. With L3StubKernelForTest degraded to a no-op, the control fails
naming the three escaped `ip` calls.
Also: the DoH3 ownership test's control now retries.
requireInstrumentFindsPackedQuery packed a query into a pooled buffer, released
it and demanded the scan find it — but under -race sync.Pool.Put drops one
object in four on purpose, so the control failed 18 of 60 measured runs and took
the whole -race pass down with it. Its sibling control in the same file already
retried for exactly this reason. The claim is existential ("this instrument CAN
find a released buffer"), so one success out of 32 proves it and nothing is
diluted; 0 of 60 after. What it does not buy is stated in the code: the VERDICT
is still a 3-in-4 detector under -race, which is the safe direction, and the
non-race pass runs the same test as a certainty.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
An egress type was read by two halves that never call each other. netplane's
EgressDevice accepted `tunnel`, so addEgressRouting gave it a mark, an `ip rule`,
a routing table with an unreachable floor and a prerouting mark bypass, and
`untunnelable_egress` (D26) carried ESP/AH/GRE/IGMP/SCTP out of it by kernel
routing with the engine nowhere in the path. generate's outbound switch had never
heard of `tunnel`: default arm, no outbound, so every node, group and rule bound
to the same egress was fail-closed. One name, two answers.
Refusing `tunnel` would have broken the half that works to match the half that
does not — D26's kernel egress is shipped and verified, and the generator's
refusal is already loud and fail-closed. `tunnel` is not a distinct kind either:
the data plane treats it identically to `interface` in every line that mentions
it, and the panel's own `interface` label already reads "out a specific WAN or
tunnel". So it is an ALIAS, and it is folded to `interface` ONCE, at the config
boundary (Model.NormalizeEgressTypes, called by ParseUCIExport/ReadUCI). Teaching
the generator a second string would have left two strings for the next consumer
to forget; after the fold there is one.
- model: CanonicalEgressType / EgressTypeKnown / KnownEgressTypes — a closed,
positive registry, plus NormalizeEgressTypes on the load path. An unrecognised
type is left as written, never defaulted: substituting `direct` for a typo
would send traffic somewhere nobody asked for.
- model: ValidateEgresses now NAMES an unknown type at validate time. Until now
the only notice was a generator warning raised while building an engine config,
which said nothing about the data plane — and the two disagreed anyway.
- netplane: EgressDevice and the prerouting mgmt-bypass consult the registry
instead of carrying their own copies of the rule. The bypass now keys off
EgressDevice, so a device-kind egress with no interface no longer gets an
accept for a mark addEgressRouting never installs.
- panel: the egress editor cleared Interface/Port/DPI for every type it had no
branch for — including types it renders no field for — so opening an egress it
labels "(unknown)", changing only the NAME and saving deleted its `interface`.
On a `tunnel` egress that silently unbound untunnelable_egress and dropped the
ESP/GRE carrier back to policy. A save may now only clear a field the editor
was in a position to show.
- panel: the unknown-type hint said "This engine builds no outbound for that
type", which was false for the one unknown type anybody had — the data plane
was building it a routing table at that moment. It now names both halves and
states what saving does.
Tests: TestEgressTypeMeansTheSameInBothHalves runs one table of written types
through the real boundary and then asks netplane AND generate, requiring one
verdict (external test package: generate imports netplane, so nothing inside
netplane can import generate). Mutation-checked both ways — dropping the fold
fails on `tunnel`; restoring the old EgressDevice string test reproduces the
historical split with "generate emitted outbound egress-probe = false ... want
true". Panel: egressEdit.test.ts, mutation-checked by restoring the
unconditional clear (Interface undefined, want 'wg0').
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The l3_tunnel default flip (164b703a7) turned 32 ORDINARY tests in
shater/generate red — the whole CI — because every fixture with a tproxy inbound
now generates the `l3-in` TUN and engine.Apply then wants /dev/net/tun, which the
act_runner LXC guest does not have. Three PRIVILEGED tests failed too, on a host
that DOES have the device.
The proposed fix was to move the TUN inbound out of generate and have the engine
add it at apply time. Refuted, on three grounds:
- it does not fix the 32. Thirty of them fail inside engine.Apply, not box.New;
the engine adding the inbound leaves them exactly as red, unless the l3_tunnel
signal travels OUTSIDE option.Options — and then
- the hash gate stops seeing it. Apply's fast path is a hash of the options; a
decision that is not in them makes toggling l3_tunnel a no-op reconcile, i.e.
the device stays up with the option off, or never comes up with it on;
- and the `icmp "tunnel"` warning cannot move. It needs the model, and the panel
reads it out of GenerateWithWarnings. Leaving it in a package that no longer
makes the decision it explains is a lie generator by construction.
What the failures actually were was contention. Measured under `docker run
--cap-add NET_ADMIN --device /dev/net/tun`: run alone, all three privileged tests
PASS; run as a package, all three FAIL — and one fails by finding a `shater-l3`
device that a DNS-filter test created. There are two L3 slots and they are global
to the process. So the fix is that the engine instrument in this suite does not
open a kernel device it does not own: withoutL3Ingress, one helper, applied at
applyAndClose and at the six other call sites.
Nothing is skipped, and the ingress does not lose coverage — it gains some:
- TestL3TunnelChangesNothingButTheTunInbound (ordinary, portable) proves the
default config MINUS the l3-in inbound is byte-identical, through the engine's
own marshaller, to the l3_tunnel=0 config. That is what lets the 32 Starts keep
speaking for the default config instead of merely for a config near it;
- TestL3TunInboundIsAcceptedByBoxNew (ordinary) puts the registry half of the
privileged test on a gate that can actually run it: a slim registry that loses
tun.RegisterInbound now fails on EVERY CI run with `type not found: tun`
instead of only where /dev/net/tun exists. That regression changes no generated
byte and costs a LAN-wide outage on the router;
- TestIntegrationL3StaleSlotIsReclaimed (privileged) covers what a RESTART finds:
an engine with l3Device == "" next to a device it did not open. It must take
the other slot, leave that one alone, and RECLAIM it on the next apply. The
occupied slot is held by a second live engine, not planted with `ip tuntap
add` — a planted device is PERSISTENT and therefore attachable, and the
planted version of this test passed with netplane.L3SlotFor's reclaim loop
deleted, i.e. proved nothing.
generate's placeholder device name is now longer than IFNAMSIZ allows. box.New
accepts it (measured), so the emitted config is still one the engine can
validate; Start refuses it and creates NO device. A caller that builds a box from
generate's output without going through engine.Apply therefore fails at once and
visibly, instead of quietly creating `shater-l3` — the one name every generation
wants, and the intermittent TUNSETIFF EBUSY that netplane/l3.go exists to refuse.
The "leaked TUN" in the sentinel's message was not a leak. Instrumented: Close
returns in ~300 µs with ZERO open /dev/net/tun fds (control: 1 fd immediately
before Close), and the device survives 3.8-4.6 s longer purely as the kernel's
deferred unregister_netdevice. On the stand (ImmortalWrt 25.12.1 r37978, kernel
6.12.94 — the router's revision) the same test takes 0.10 s, so the lag is a
nested-netns container artefact. l3GoneTimeout goes 5s -> 20s: a leak is
unbounded, so the longer budget costs one slow failure and gives up no
sensitivity.
Verification. CONTROL, the criterion that matters: without /dev/net/tun
`ok shater/generate` (was 32 failures). With `--device /dev/net/tun --cap-add
NET_ADMIN`: green, privileged tests really ran. On local_openwrt, cross-built
with the shipped tags: the WHOLE package green with every privileged test
executed, no contamination. `go build ./...`, `go vet ./shater/...` clean.
Mutation-verified, each reverted after: shortening the placeholder fails
TestL3PlaceholderCannotBecomeAKernelDevice by name; making withoutL3Ingress a
no-op brings back exactly 32 failures; gating a second config change on
l3_tunnel, and stripping nothing in the comparison, each fail
TestL3TunnelChangesNothingButTheTunInbound; removing tun.RegisterInbound fails
TestL3TunInboundIsAcceptedByBoxNew with the right hint; deleting L3SlotFor's
reclaim loop fails TestIntegrationL3StaleSlotIsReclaimed with the production
error verbatim (`TUNSETIFF: device or resource busy`); l3GoneTimeout at 1ms still
fires the leak sentinel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Five untunnelable notes told the operator that a plain `traceroute` works,
"still follows your rules", or that the hops it prints are the tunnel's path.
Measured on the production router: it prints `* * *` and nothing else, under
every rung of the ladder — `direct` included — with the L3 ingress on or off.
There is no mechanism that could print a hop. The UDP probe is diverted by
tproxy and delivered LOCALLY to the engine's socket; local delivery is not
forwarding, so the TTL is never decremented and no router on the path is
provoked into a time-exceeded. The engine opens its own connection with a
fresh TTL, and an ICMP error raised against that has no way back to the
client's datagram. `traceroute -I` and Windows `tracert` are ICMP echo and do
work — that half of the text was true and is kept.
One shared udpTracerouteFacts now carries the symptom, the cause and the way
out, so the panel cannot fork the claim; netplane/untunnelable.go states the
same fact in the same terms.
Second correction in the same notes: the outbounds that carry an echo are not
just WireGuard/AmneziaWG. generate/route.go's l3Target is exhaustive by
adapter registration — a wireguard/AWG node AND the direct outbound behind
`direct` or an interface egress. In the commonest configuration here that is
most of the address space, and those pings answer out of the ordinary uplink
with its real address. The old text let an operator conclude either
"tunnelled" or "dropped"; it was neither.
traceroute_honesty_test.go is the ratchet: an exhaustive matrix over policy x
kill switch x L3 x egress, asserting the retired sentences never return and
that any note mentioning a trace carries the shared facts verbatim — with a
control that fails if the matrix stopped mentioning tracing at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Four defects, all of the same family: something the plane relies on stops being
true and nothing says so.
1. One failed `uci -q export firewall` opened a hole AND switched off the alarm
for it. nftZoneDevices answered nil on a read failure — the same answer as an
empty zone — so a rule with `src: zone:lan` produced no divert line, no
fail-closed drop and no accept_local; and uncoveredNetworkWarnings, whose job
is to report exactly that, ran the same command, got the same nil and stayed
silent. The read now carries its error: renderNft refuses under a closed
kill-switch (same contract as an unusable device name) and warns under an
open one, and the coverage check names the blindness itself.
2. RoutingPresent did not check the fail-closed floor its Apply twin installs.
addEgressRouting/addL3Routing install three things per binding; the presence
checks knew two. A floor that failed to install once was never retried, and
the table fell through to `main` the first time its device went down. The
checklist test grows clause (e) so the next mark cannot repeat it.
3. A flow established before the divert plane existed bypassed it for life:
confirmed by conntrack while nothing diverted it, offloaded to fw4's
flowtable, steered by netdev-ingress ahead of our prerouting hook and
refreshed by its own packets. On the divert going from ABSENT to PRESENT —
not on every apply — the TCP/UDP entries of flows forwarded from the divert
devices' subnets are dropped, so they re-derive their path. Not a flush: the
router's own addresses and LAN-to-LAN are excluded, so SSH, LuCI and the panel
survive. Measured on the stand: 3 client flows cut, the live SSH session and
the router's own connections untouched; `conntrack` CLI confirmed absent
there, which is why this is ctnetlink.
4. The untunnelable text claimed Linux/macOS traceroute "still prints hops". It
prints none, under any policy: the UDP probe is delivered locally by tproxy,
local delivery does not decrement TTL, and no router raises time-exceeded.
`traceroute -I` is what works.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The l3_tunnel default flip (164b703a7) turned 51 tests in shater/generate red.
Two premises had changed, and each is repaired where it broke rather than at the
assertion:
- ~43 fixtures build an engine-topology model with no inbounds at all and assert
"this config produces no diagnostics". On the seeded-ON default such a model
earns an honest `icmp "tunnel"` warning: the L3 ingress is fed only by the
tproxy divert plane, and a model with no tproxy inbound raises none. The
warning is TRUE of those fixtures — they are not routers. So they now say they
run neither router-wide plane (nonDNSGlobals became plainGlobals, and gained
the same treatment for l3_tunnel that D24 gave dns_intercept), and every
"no warnings" assertion keeps its original strength instead of being loosened
to "no warnings except this one".
- 8 assertions counted len(opts.Inbounds). The subject of every one of them is
how many TPROXY LISTENERS survive a guard, and a total that also counts a
synthetic inbound answers a different question — one whose right number
changes whenever an unrelated global flips. They count tproxy listeners now,
and while there they gained the assertion the count was standing in for: that
the SURVIVOR of the clash guard is the first-declared listener, and that two
distinct ports keep the ports their nft diverts aim at.
TestL3TunnelOffEmitsNoTunInbound had lost its meaning rather than its fixture.
It read the default and asserted "off", so after the flip it was pinning
DefaultGlobals, not l3_tunnel. It now sets the opt-out explicitly and says why
the opt-out has to keep working, and TestL3TunnelOnByDefaultEmitsTunInbound
pins the other direction — that a model which never mentions l3_tunnel gets the
ingress — which nothing in this package did.
TestSniffIsNotAnInboundField asserted "exactly 1 inbound" purely so it could
index ins[0]. It checks every emitted listener now and counts what it checked,
so the guarantee that assertion was really providing (the loop ran) survives
without a count that any future synthetic inbound breaks for no reason.
The warning text is rewritten. "l3_tunnel is on but no tproxy inbound is
enabled" accused the reader of a choice they no longer made: since the flip it
is the default, and a message that reads as "you turned this on" sends them
hunting for a switch they never touched. It now says what is not happening, that
the ingress is on by default, and names BOTH exits — a tproxy inbound restores
it, `option l3_tunnel '0'` says the router does not want it — because which one
is right is a fact about their router the generator cannot know.
model/dnsintercept_test.go had the blindness its l3 twin documented: a plain
strings.Contains is satisfied by `#option dns_intercept '1'`, and the parse half
cannot tell either, because a commented option falls back to the seed, which
since D24 is also true. A config shipping the option commented out would have
passed both halves while giving a fresh install no visible option to flip. The
check is line-wise and comment-aware now, and its "config unreadable" branch is
a Fatal instead of a Skip — a guard that skips itself is how one ends up
reporting ok while guarding nothing.
Mutation-verified, each reverted after: seeding L3Tunnel=false fails the
default test by name; removing the l3_tunnel guard fails the opt-out test;
stripping either exit from the warning fails TestL3TunnelWithoutTproxySkipped;
setting a legacy SniffEnabled on the tproxy listener fails the sniff test;
disabling the listen-clash guard fails TestDuplicateTproxyPortSkipped; freezing
the tproxy port at the default fails TestMultiLanDistinctTproxyPortsBothKept;
commenting out the shipped dns_intercept fails the shipped-config test (and the
parse half stayed silent, which is the blindness).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Three holes, one shape: work that reads as coverage and is not.
1. shater/apply's TestApplyInstallsHoldWhenEngineFailsToStart — the only
end-to-end test between "the engine died" and "the LAN forwards to the
WAN in the clear" — asserted nothing. It broke the engine by pointing a
rule-set at /nonexistent/nope.srs and stood itself down with t.Skip when
that failed to break anything; it stopped breaking anything once
LocalRuleSet.reloadFile began treating an unreadable file as empty.
Measured in golang:1.26: the skip fired unconditionally and the package
still printed `ok shater/apply`.
It now injects the failure at the engineApply seam — the branch under
test is applyLocked's, and a particular cause that stops causing retires
the test silently — and COUNTS the seam calls, so applyLocked ceasing to
go through it fails by name instead of quietly asserting something else.
Everything else stays real: the model, generate, the kill-switch
decision, netplane.RenderHoldNft, the latch, Status. New companion
TestEngineApplyReallyFailsWithoutStarting is the control that the real
engine.Apply can fail with the engine left stopped, so the simulated
state is one this fork can be in.
Mutation-checked both ways: drop the holdLocked call from applyLocked and
the test fails with "0 holding planes were installed, want 1"; bypass the
seam and it fails with "the engine-swap seam ran 0 times, want exactly 1".
2. warnings_test.go had two of the same genre. The len(genWarnings)==0
t.Skip is now a t.Fatal — an unloadable blocklist must always warn, and a
generate that stops saying so is the W7 regression, not a reason to stand
down. TestStatusWarningsAlwaysNonNil pins readConfig itself: its
"zero warnings" assertion was true on a build host only because the
config read failed SILENTLY, so once that failure started publishing a
critical warning the same line meant two different things in two
environments.
3. The gate could not see any of it. It now runs the suites with -v and
matches every `--- SKIP` against SKIP_DECLARED; an undeclared skip fails
BY NAME, a declared one prints its reason on every run. check_skips
proves its own instrument first (no `=== RUN` line => the check was
reading a blank page), and it also reports on a suite that failed
elsewhere, so a red tree cannot become a hiding place. -v costs no test
time (38/25/24 s plain vs 38/24/24 s, warm) — only output, which is
filtered on a green run.
Also closes the same hole one language over: [6/7] requires every non-Go
test file in the tree to be claimed by a named runner, and [7/7] runs the
ones this gate owns with a verdict by name. openwrt/luci-app-shater/tests/
status-readout.test.js — 24 assertions over the one screen an operator
reaches while the LAN is cut off — was executed by nothing at all, and
[1/7] could not report it because `go list` is its instrument. The non-Go
suites run on the HOST before the docker re-exec, so the local loop really
executes them rather than printing "did not run" every time; where there is
no node at all they are named and the notice replaces the closing banner.
Controls, all run and reverted: a planted t.Skip is caught and named; a
planted failing .test.js is caught and named; an unclaimed test file is
caught and named.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
l3_tunnel was opt-in, and "off" had no honest win left in it. Off, a LAN ping
is decided by `untunnelable` alone and every rung is a drop (block) or a
disclosure (icmp/direct send the echo out of the WAN with the client's real
address). "Ping works" was never the state where ping was tunnelled — it was
the state where ping was leaking. On, an L3-capable outbound carries the echo
and one that is not drops it honestly: adapter.JudgeFlow returns ActionDrop for
an ICMP flow whose outbound is not a tun.Port, so no reply is forged. The price
is a standing TUN + gVisor netstack, ~2 MB RSS, and it is stated where the
option is.
The switch stays. It is a real answer on a 32/64 MB device and when bisecting
whether the L3 ingress is what broke a box — but it is now a WARNED answer:
ValidateGlobals says what the off state does to ping and names the policy that
takes over. Two combinations also changed meaning and are now reported:
untunnelable=icmp is no longer "block plus working ping" (the prerouting L3
mark claims every ICMP packet before the forward chain the echo accept lives
in, and a LAN host's ICMP errors are marked in with them and dropped in the
TUN), and the existing =direct report gains a sibling rather than standing
alone.
The fw4 seeding was the second half of the same problem. The divert set spans
every LAN inbound and every iface:/zone: rule source, but 30_shater-core seeded
a forwarding into shater_l3 for `lan` only — so on a multi-zone router ICMP
from the other zones is marked, routed, accepted by `inet shater`, and dropped
by fw4's zone policy with nothing in any log. Every zone gets a forwarding now,
guarded by a scan of the actual src/dest pairs so a re-run adds nothing. Every
zone including an uplink, because guessing which zones hold clients is wrong
somewhere and a superfluous entry authorises nothing: accept_to_shater_l3 is
`oifname "shater-l3*" accept`, and the only thing that routes a packet into
that device is our own fwmark rule.
scripts/testbed-lao.sh builds the second LAN zone this needs to be visible at
all. It is not installed by the package — that is the whole opt-in mechanism.
Verified on local_openwrt (ImmortalWrt 25.12.1 r37978): three runs of the
seeder leave exactly one forwarding per zone (lan/wan/lao) and no existing
section altered; deleting the lao forwarding removes `jump accept_to_shater_l3`
from chain forward_lao and re-seeding restores it; with the idempotency guard
disabled two runs produce nine forwardings instead of three.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Review found the hole and it is real. My previous fix gave the pooled buffer to
the transport and released it when the transport closed the body, on the grounds
that "http3.Transport closes the request body on every path, hence the Once".
That sentence is true about how many times the body is closed and says nothing
about when — the failure mode this project keeps writing down.
Verified against the pinned quic-go: on every error path RoundTripOpt
(http3/transport.go:167-173) closes the body the moment doRequest returns, and
doRequest (http3/client.go:338-341) waits only on the request-CANCELLATION
watchdog — close(reqDone); <-done — never on the goroutine writing the body.
Nothing in quic-go joins that goroutine. So Close is not a handoff point, and
the sync.Once stopped a double Release while doing nothing about a read after
one.
One correction to the review's severity, since it changes what we tell people:
on the failure path the bytes do not reach the resolver. Every ReadResponse
error branch (http3/stream.go:325, :336, :343, :363) calls str.CancelWrite
BEFORE RoundTripOpt closes the body, so what the writer reads out of the
recycled buffer is thrown at a cancelled stream. The disclosure primitive is the
success path only; the failure path is a read of somebody else's memory, which
is undefined behaviour and a -race finding, and not shippable either.
Fixed by not sharing at all: Pack() into memory the body owns. The alternative —
a lock around Read and Close — would also be correct and was rejected because it
keeps a released-but-referenced object alive, and that is now twice in one day
that an assumption about quic-go's internal lifetimes has been wrong.
The cost is negative, measured rather than assumed: Pack is 87 ns/op at 64 B and
1 alloc against 108 ns/op at 64 B and 1 alloc for the pooled version, because
buf.NewSize allocates the Buffer struct itself — the same 64 bytes — and then
adds Get/Put on top. The pool was never saving an allocation here.
The failure path cannot be caught on the wire, so the new test pins the cause:
a query tagged with a random needle, an exchange that fails (server never
answers; context already cancelled), then the pool drained on the goroutine
RoundTripOpt ran on, demanding the needle is not there. Mutations run without
-race: restoring pooledRequestBody fails both subtests 5/5, and blunting the
scan trips its control. -race is a separate pass, green at -count=3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two review findings on the fragment reassembler.
1. A whole datagram could vanish. entry.total was assigned before addRange
was asked, so a last fragment (MF=0) whose range was already covered by
MF=1 fragments answered fragInsertDuplicate and returned nil — while the
entry was already complete(). Nothing re-examined it, because every later
fragment is a duplicate too, so it died at its deadline with all its bytes
present. A duplicate now falls through to the completion check: it
contributes no bytes (held bytes still win) but it does contribute the
total length. This is what the documented first-wins policy always
implied; the code just did not do it.
The sender needed is non-conforming, so the old behaviour was safe rather
than exploitable — but it contradicted the comment three screens up, and
that comment is the next reader's only defence.
Also closed positively: a last fragment declaring an end BELOW the bytes
already held now poisons the datagram instead of quietly never completing.
2. The 5 s timeout was not a memory ceiling and the comment said it was.
sweep ran only when a NEW key was created, so once fragmented traffic
stopped, up to fragMaxEntries entries stayed resident indefinitely.
Both halves are fixed, and the honest one is the comment. sweep now runs
on EVERY fragment — an O(64) scan on a path that is already the rare one —
which releases residue as soon as any fragment arrives instead of waiting
for an unrelated new datagram. That still does not cover total silence, so
fragTimeout now documents the guarantee the code actually keeps: bounded
by fragMaxEntries/fragMaxTotalBytes at all times, released on the next
fragment, NOT "freed within 5 s".
No timer, deliberately: it would need a goroutine with a lifecycle tied to
something returnDeviceWrapper has no teardown hook for, and a goroutine
that must be stopped and might not be is a failure this project has
already paid for — to reclaim at most ~1.1 MiB that only exists after
fragmented traffic has already happened. What bounds growth is the byte
and entry ceiling; this timeout's job is correctness, and for that a
check driven by the arriving fragment is exact.
The now-unreachable per-key deadline check is removed rather than left as
dead defence in depth.
16 mutations, all red. M15 (duplicate returns early again) reds only the
buggy case while the control and the poison case stay green, so the test is
shown able to see both an assembled datagram and a lost one. M17 (sweep back
inside the new-key branch) reds the new test while both old timeout subtests
stay green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The rollback added for the "failed import commits the deletion" defect went
through migrate.go's staged(), which drops the revert's error on the floor
(`_ = u.Revert("shater")`). That is defensible where staged() lives — a
migration that cannot revert leaves a half-migrated config, wrong but visible —
and it is not defensible here, because the delta this path stages STARTS WITH A
DELETE OF THE WHOLE PACKAGE. A revert that silently does not take leaves that
delete in /tmp/.uci, the caller is told only "import failed" and believes
nothing happened, and the next `uci commit shater` from any process publishes
an EMPTY /etc/config/shater. The guard reintroduced the exact loss it was
added to prevent.
writeUCIWith now uses its own revertStagedWrite, which reports both failures.
migrate.go's staged() is untouched: changing its signature to suit this caller
would rewrite a contract three migration paths depend on, for a hazard those
paths do not have.
The wrapped error names the CONSEQUENCE and the one command that clears it
("a staged DELETE ... will publish it ... run `uci revert shater` NOW"), not
just the fact — "revert failed" tells an operator nothing about what it costs.
ErrStagedWriteStuck makes it machine-detectable, so a caller can tell "your
change did not happen" from "your change did not happen and this router is one
unrelated `uci commit` away from an empty config".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
files/ is not installed wholesale — every path in Package/shater-core/install is
explicit — so the keep.d file added alongside it would never have reached a
router. sysupgrade's "keep settings" walks /lib/upgrade/keep.d/*, and without
this entry /etc/shater/subs does not survive a flash: the restored box has its
rules and its groups and no nodes for them to point at, and the only repair is
`sub update`, which needs the internet the tunnel was going to provide.
/etc/config/shater needs no entry — it is a package conffile and sysupgrade
already keeps it that way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Nine places where the panel asserted more than it could know. Each was
checked against the daemon before being changed, and the two that a test
can reach are pinned by tests proven with a mutation.
MULTICAST IPTV WAS AN INSTRUCTION, AND IT WAS WRONG. The `direct` rung
said "Ping, multicast IPTV, and connecting to a VPN ... all work", so
someone who wanted IPTV read it and moved to the most open setting on the
ladder — the one that also lets a client's ESP/GRE past the proxy — and
still had no IPTV. The stream is UDP; every rule the policy emits carries
`meta l4proto != { tcp, udp }`, and the fail-closed forward chain accepts
only the RFC1918/link-local daddr sets, with no 224.0.0.0/4 among them.
The daemon says so itself in the note drawn a few pixels below. IPTV is
now stated once, and it says it does not work.
THE `block` COST LINE WAS UNCONDITIONAL, and three settings contradict
it: an open kill-switch (no drops are emitted at all), Globals.L3Tunnel
(ICMP is marked into the engine's TUN before the forward chain) and
Globals.UntunnelableEgress (ESP/AH/GRE/SCTP are routed out a named
device). The last two were not in the panel's `Globals` type, so the page
could not have been honest about them even in principle; they were added
rather than papered over with a vaguer sentence, and the copy is now
derived from all three.
THE KILL-SWITCH WAS READ WITH `=== 'closed'`. The daemon decides with
!EqualFold(TrimSpace(v), "open") and `Status.kill_switch` is the raw UCI
string, so `'Closed'`, `' closed '` and `''` — all of which BLOCK on the
router — drew OPEN, amber, "Nothing is meant to be blocked", and through
protectionState downgraded a plane-less router from crit to amber. One
normaliser now, `planeState.killSwitchClosed`, used by all five callers
that had their own spelling of it.
AN UNREADABLE CONFIG IS NOT "TURNED OFF". `enabled`, `kill_switch` and
`panel_port` are sourced from the config and are placeholders when it
could not be read (new `config_readable`). That happens on a full
/overlay or an interrupted `uci commit` — exactly when the fail-closed
plane has the LAN cut off on purpose — and the daemon publishes
plane:"hold" with enabled:false. Checking `!enabled` first rendered
"Turned off", amber, no alarm, and pointed at a Settings page backed by
the same unreadable file. The check now comes first, carries the daemon's
"do not turn anything off to fix it", and the kill-switch readout refuses
to name a policy it could not read instead of printing ARMED from "".
Also: the holding plane promises "no client TRAFFIC reaches the WAN", not
"nothing" — DNS to the router still goes to the ISP in the clear, by
design, so the daemon can recover; the stats backend is bbolt, not SQLite,
and reclaims space by rebuilding the file, not by a VACUUM that does not
exist (and skips it when the disk cannot fit the copy); the lock screen
sent people to System → shater when the menu entry is admin/services/shater,
which is the one instruction the product gives to someone who has just
lost access; and the panel port is configured, not confirmed — a failed
listen is only a log line.
RULESET.FORMAT WAS DESTROYED BY RENAMING A LIST. The edit form rebuilt
the object from its own controls and has no control for `Format`, so the
value could only be restored over SSH. It decides how a `file` list is
parsed and stops a `url` .srs being read as text; without it the list
matches nothing, the rule stops firing, and the traffic falls silently
through to the next rule. Carried now for the two sources the generator
consults it for. The same class of loss is made loud elsewhere: the two
other rebuild sites return `Complete<T>`, so adding a field to `Inbound`
or `DNSRule` fails the build in the function that has to decide.
Egress.Target is deleted: it is not in the Go model, so the "which egress
points at this node" branches could never fire, and had anything ever put
a string on it PUT would have rejected the whole write under
DisallowUnknownFields.
One layout fix on the way past: at 390px the policy plate's grid column
was sized by the select's longest option, so the sentence beside it was
clipped mid-word — which is how a line about what leaks loses its second
half.
Verified: npm run build + tsc clean; 57 tests pass; mutation-checked by
restoring the old comparison, the old check order and the old rebuild in
turn, each time watching the matching tests fail with the exact inverted
reading; browser-checked at 390 and 1280 against the mock, which now
reproduces `?ks=Closed` and `?cfg=unreadable` verbatim instead of
normalising them out of existence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two things the box could not survive, both silent.
BACKUPS CARRIED NOTHING. No shater package put a single entry in
/lib/upgrade/keep.d, so "keep settings" and LuCI Backup took /etc/config/shater
(a conffile) and nothing else. Everything the product knows besides UCI lives in
/etc/shater: the entire node inventory (subs/*.json, hundreds of nodes on the
live router), the boot-armor arm token, the compiled blocklists. Restored onto a
new router the config looked complete and had no nodes to route to — and the
repair, `sub update`, needs the internet the tunnel was supposed to provide.
keep.d/shater-core keeps subs/, boot.nft, lists/ and alert-state.json, and names
what it refuses and why: stats.db is history bounded only by stats_disk_limit_mb
(0 = unlimited) and the archive is built in RAM; cache.db is sing-box's cache and
a stale one is worse than none; shaterd.log is a log carrying the query history
of the box it came from.
THE WATCHDOG COULD NOT SEE A CRASH LOOP. /etc/init.d/shater respawns every 5s,
forever; shater-cron escalated only after five consecutive ticks where `pidof`
found nothing. A daemon dying seconds into startup is back before the next
60s sample, so the counter reset every time — while the fail-closed plane held
the LAN shut and the panel, served by that daemon, never came up.
The tick's sleep is now spent sampling the daemon's identity (via its pidfile,
not `pidof`, which also matches the CLI verbs this loop runs) every 5s. A tick in
which 3 different daemons lived is churn; two such ticks in a row is the verdict.
A legitimate bounce replaces the daemon once and is announced twice over
(RESTART_FLAG up, ACTIVE_FLAG down), either of which discards the tick.
The action is the one the operator already chose: kill_switch=open stops the
stack, exactly as the dead-daemon path does; kill_switch=closed — and an absent
or unrecognised value, which is the documented default — reports at daemon.crit
and leaves the decision to the person, naming the command that opens the LAN.
Also drops the ruleset loop from shater_run_due. `shaterd ruleset update` has
never existed; it exited 0, so the loop stamped every url rule-set as freshly
updated and fired a reconcile for work that never happened. Now that it exits
non-zero the same loop would emit ~288 syslog lines a day per rule-set instead.
The comment says who does own the refresh, and where the gap that is left is.
Verified: sh -n and busybox `ash -n`; the pure detector driven with synthetic
sample streams under busybox ash (13 cases); shater_sample_pid against a real
/proc with a live process named shaterd as the positive control; and the whole
chain end to end against a real 2s-lifetime crash loop. Each threshold and each
veto is pinned by a mutation that makes the gate fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
On the production router every configuration change with l3_tunnel=1 killed the
engine and held the LAN down, three times in a row:
19:10:33 reconcile failed: start inbound/tun[l3-in]: open tun: TUNSETIFF: device or resource busy
19:14:02 start instance failed and could not restore previous config; engine stopped
19:14:38 reconcile failed: TUNSETIFF: device or resource busy
A new generation had to open the device the outgoing one still held. That alone
is a failed apply; what made it an outage is that the recovery path rebuilds the
PREVIOUS config, which named the same device — so the rescue failed for exactly
the reason it was needed. A recovery path must not depend on the resource whose
contention it is recovering from.
The device is now one of two slots, chosen by the ENGINE at box-build time, on a
copy of the options taken AFTER the hash — so the stored config stays canonical
and a no-op reconcile is still a no-op. It cannot be chosen in generate: generate
runs every minute and its output is what Apply hashes, so an alternating name
there would rebuild the engine once a minute forever.
Rotation alone was NOT enough, and that was measured, not reasoned: the two-slot
build survived five applies of five kinds and then failed on 4 of 10 back-to-back
changes with the original outage in full, because a retired generation keeps its
TUN until its budgeted Close finishes. So an occupied non-current slot is now
DELETED rather than waited for — the running generation's slot is excluded first
and never touched, every other slot belongs to a box that is carrying nothing.
No bounded wait: waiting on an asynchronous kernel teardown is the race this
design removes.
The firewall never learns which slot is live — our accepts and the fw4 zone match
`shater-l3*`, verified to validate AND load on ImmortalWrt 25.12.1 / nftables
1.1.6, so the ruleset is byte-identical across a swap. Routers seeded by a
pre-slot build are migrated in place, or fw4 would silently resume dropping the
forward.
A2: turning the feature off left the device, the ip rule and table 8200 behind —
addL3Routing returned early instead of tearing down, and nothing else owns that
device. The disabled branch and TeardownRouting now remove all three.
Two smaller lies found while proving this, both measured: `ip -6 route flush`
does not take a non-unicast route, so the fail-closed floor survived and the next
add answered `File exists` — reported as a CRITICAL "this table has no floor,
traffic can leave over the plain WAN" on every apply, about a floor that was
right there; and teardown left it behind. Fixed both.
Verified on local_openwrt (ImmortalWrt 25.12.1, kernel 6.12.94 — the router's
revision) before and after, with binaries built from the same tree: the pre-fix
binary reproduces the outage and the leftovers; the fixed one survives all five
apply kinds and 12 back-to-back changes and leaves nothing behind. Ten reverted
mutations, each shown failing. See D28.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The tracker has carried the matched route rule and the outbound chain since
upstream (common/trafficcontrol/tracker.go Rule/Chain); nothing in shater/ ever
read them, so "why did this connection go out that exit" was unanswerable from
the log and cost hours per report.
ConnLogEntry gains RuleKind/Rule/Chain. Rule is the engine rule text, not the
model rule name: nothing survives generation that ties an emitted option.Rule
back to the /etc/config/shater rule it came from, and a guessed name would be
worse than none. RuleKind keeps the two empty cases apart — "default" is a
recorded fact (nothing matched, took route.Final), "" means not recorded at all,
which is what an old persisted row decodes to.
Both fields are interned, so the ring pays 56 B/row of headers instead of a
private copy of text that is identical across every connection one rule matched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
@
Three ways this package published calm over a router that was not doing what
its config said. All three are the inverted failure: not an error raised when
things are fine, but silence when they are not.
1. Critical policy-routing findings were erased by the next no-op reconcile.
applyDataPlaneLocked set routeWarnings only on the full path; applyLocked
published the set unconditionally, so a minute later the fast path replaced
it with one that no longer contained the finding. Neither surviving finding
("this egress CANNOT REACH ANYTHING outside its own subnet", "table could
not be given a fail-closed floor") makes RoutingPresent false, so nothing
brought it back: zero findings, plane full, green, over an egress carrying
nothing. The comment on the gate claimed the previous set stood; it did not.
planeOutcome now distinguishes "nothing was found" from "nothing was
checked" (routeMeasured, written only by measuredRouting), and applyLocked
carries the last MEASUREMENT forward across the fast path. A re-measurement
still retires a finding, so this is not a latch.
2. An unreadable configuration was published as enabled=false. The panel tests
!enabled before plane and renders "Turned off", amber, no alarm, "turn it on
in Settings" — over a LAN the boot armor had cut off, pointing at a settings
page backed by the same unreadable file. Status now carries config_readable
and config_error, plus a critical finding in section "config".
3. The reason the engine failed to start existed nowhere. holdLocked logged it
and called no publisher, and Warnings carries the last SUCCESSFUL apply — so
plane="hold" with an empty findings list was a normal state of the product.
The cause is recorded and published at read time while the engine is down,
so it self-clears when the engine comes up; the boot-time arm is a warning,
a real failure is critical.
Each fix is mutation-checked, and the route-warning test carries its control:
it sees a live finding, sees it survive the fast path, and sees a re-measured
clean state retire it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
generate/outbound.go resolved the bind device with netplane.IfaceDevice,
whose empty-name fallback is "br-lan" — correct for an INBOUND with no
network, a black hole for an egress. netplane.EgressDevice returns "" for
the same egress on purpose (it calls br-lan "catastrophic here"), so
addEgressRouting installed no `ip rule` and no routing table for that
egress's mark, and the prerouting marking and the forward-chain accept
skipped it too.
The outbound was therefore emitted with SO_BINDTODEVICE=br-lan and a
routing mark nothing routed: every node, group and rule bound to that
egress dialled public addresses out of the LAN bridge. Not a leak — the
bind pins the socket to the LAN — but a total, silent black hole, with the
panel showing a configured, applied egress and no findings at all. The
`if dev == "" { dev = eg.Interface }` line that stood there read as a
guard against exactly this and could never execute: IfaceDevice never
returns "".
- generate now calls netplane.EgressDevice — the data plane's own
resolution — so a bind can no longer name a device the routing was never
installed for, and ` eth1 ` binds what the netplane routes. A device-less
egress emits NO outbound and is reported; every reference to it then
resolves through egressDetourOrBlock to tagBlock, so the traffic is
blocked rather than sent out over the plain WAN.
- model.ValidateEgresses reports the same egress on the config channel
(netplane's own skip is silent), built on model.EgressHasDevice — the
model-side twin of EgressDevice, which ValidateUntunnelableEgress now
shares so the two model resolutions cannot drift either.
- TestEgressDeviceResolutionParity runs one table through
netplane.EgressDevice and model.EgressHasDevice and requires one verdict,
the same treatment TestUntunnelableEgressResolutionLockstep gave the
earlier validator/data-plane divergence.
Also: the UntunnelableEgress comment claimed "the panel says which, at
apply time, from whether the device is point-to-point". It does not. The
operator-facing text states both possibilities and declines to claim
either, there is no UI for the option, and isPointToPoint is consulted
only to warn that a gateway-less device can reach nothing. Said so, so the
next implementer does not read a described feature as a built one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`shaterd status` fabricated a status when the daemon was unreachable and
exited 0. The stub is the same struct, printed by the same marshaller, so the
only thing that distinguished it was `plane` being "" — a value a live
Applier.Status() cannot emit. luci-app-shater was forced to key its "daemon
down" verdict off exactly that side effect, and filling `plane` in the stub for
any reason would have silently turned "dead" into "fine" on that page.
Both branches now carry an explicit "daemon_answered" boolean, and the offline
branch exits 1. The field is ADDITIVE and spliced in, not re-marshalled: every
existing key keeps its name, value and position (including plane:"" — still
emitted deliberately so dashboard.js keeps working until it moves onto the new
field), and a newer daemon's unknown fields are relayed untouched.
model.writeUCIWith committed the staged package DELETION when the import that
was supposed to refill it failed: /etc/config/shater came out empty, the caller
saw only "WriteUCI: import: ...", the next ReadUCI reported Enabled=false and
the next reconcile tore the plane down. Both error paths now revert through
migrate.go's staged() instead — the same idiom, for the same reason.
`shaterd ruleset update` printed a note and exited 0. shater-cron runs it with
output discarded and, on a zero exit, stamps the ruleset as freshly updated and
sets changed=1, so every source=url ruleset was permanently "just updated" by a
verb that fetched nothing. notImpl now exits 1 (not 2 — a caller must be able to
tell an unimplemented verb from an unknown one).
pidfilePath becomes a var so the daemon-answered / daemon-absent split is
testable without writing to the real /var/run, mirroring ctlPath.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
sing-tun's forwardReturn.classifyReturn refuses to judge a fragment
(flow_parse.go sets `fragment` for IPv4 MF/offset and for an IPv6
fragment extension header; flow_dispatch.go:703 answers returnPass), so
a fragmented answer coming back through a WireGuard/AmneziaWG endpoint
falls through to the endpoint's own tun stack instead of the l3 return
path, and the LAN client never sees it.
Measured on the live router: `ping -c3 -s 1400` through an AWG tunnel
with MTU 1280 is 100% loss while the WAN capture shows 3 x (1312 + 208)
in both directions — the far host answers, the peer fragments the answer
to fit the tunnel, the fragments die in classifyReturn. `-s 56` is 3/3
and PMTUD with DF works end to end, so only the fragmented return is
broken.
sing-tun is pinned upstream with no `replace`, but the fix does not need
to live there: every decrypted packet passes returnDeviceWrapper.Write
before it is offered to ReturnPackets. Reassemble there and
classifyReturn gets a whole datagram.
Hard ceilings, because this runs on a 128-256 MB router: 64 concurrent
datagrams, 1 MiB of held bytes, 64 disjoint ranges per datagram, 65535
bytes per datagram, 5 s to complete (timer starts at the first fragment
and is never refreshed). Over any ceiling evicts oldest-first.
Overlap policy: a range contained in one already held is a duplicate and
is ignored (first-wins, deterministic) because benign networks do
retransmit; any PARTIAL overlap poisons the datagram until its deadline.
No conforming fragmenter emits one, and every historical hole in this
area comes from a reassembler that tried to resolve the conflict.
The MTU of shater-l3 is untouched (65535 on purpose) and sing-tun is
untouched.
14 mutations run against the tests; each turns at least one test red,
including the two that first survived (a stale-head reuse the sweep was
covering for, and a fast-path copy).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Every install recipe walked the reader through `shaterd apply` + `shaterd
confirm` as if commit-confirm were armed. It is not: DefaultGlobals() never
seeds ConfirmTimeout, the shipped config carries confirm_timeout '0', and
ArmRollback returns at once on a non-positive timeout. A reader following the
README believed an apply that cut their SSH would undo itself. It would not.
README/README.en/INSTALL now arm it in the recipe and say what 0 means; the
apply-flow diagram gained the edge it always took on a stock box.
The boot armor was documented nowhere at all (`grep -rli armor --include=*.md`
returned zero) while shipping enabled and blocking LAN->WAN on every boot.
INSTALL 4 now says what it is, why SSH/LuCI stay up on purpose, every condition
under which it refuses to arm, and how to switch it off.
Also removed or corrected, each checked against the code, not inherited:
* MASQUE/CONNECT-IP is advertised in both READMEs and absent from parse,
generate and model -- registry names it among the types deliberately left
unregistered. Dropped, with the fork-vs-product distinction spelled out.
The inverse too: Hysteria2/TUIC/XHTTP were tagged [T1] while shipped under
with_quic/with_xhttp; ShadowTLS is generate+registry only, no parser.
* `direct (flow-offload on)` -- no offload/flowtable/flow_offloading anywhere
in openwrt/, shater/ or panel/src. The product does not do this.
* shater-core deps were two releases stale in two places, one of which vouched
for a config.buildinfo check that never covered kmod-tun. Ruling narrowed to
what was actually checked.
* PORTING's "Full schema" -- the shipped config points at it -- was missing
l3_tunnel and untunnelable_egress (UCI is their only path; the panel does not
show them) and the blocklist/allowlist/device/alert sections, while listing a
`config preset` that ReadUCI has no branch for.
* ARCHITECTURE had no L3 ingress and no kernel egress at all, though both are
[MVP] and one creates an fw4 zone in the user's firewall config. New 3a.
* nftset-for-routing in the DNS diagram: that is the v0.1 mechanism, gone in v0.2.
* CONTEXT described a pre-Phase-1 repo and a 24.10.3 testbed. The testbed is
ImmortalWrt 25.12.1 r37978-cd0a06bfd3fd (read off the box), which is not a
detail: .apk does not install on 24.10 at all.
* The gate existed and no .md mentioned it. README/README.en/CONTEXT now do.
* release.yml's header still described publishing as either/or after the rolling
pointer became unconditional. Comment only.
* Shipped /etc/config/shater: schema_version '1' against CurrentSchemaVersion=2;
a pointer to a dns_filter line that was not in the globals block (added, '0');
and `option sniff '1'` on the inbound -- an option the model deliberately does
not have, which the first panel save would have silently washed out.
* lx-changelog pointed at a D25 heading that does not exist.
* ROADMAP 2b and 5 were done and unmarked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two suspicions, both put to a test rather than to a reading. Both were real, and
neither was the leak the suspicion named — both are objects released while still
in use.
roundTripHTTP3Race ran both racers on one cancellable context and cancelled it
before returning the WINNER. quic-go and net/http reset a request's stream when
its context dies, so the caller got a response whose body stopped mid-read:
H3_REQUEST_CANCELLED (local) (read 2687 of 65536 bytes). That path is taken
whenever there is no cached HTTP/3 connection and the request is replayable —
the first request to every host, and every one after an idle close. Each racer
now has a context of its own; losers are cancelled where everything used to be,
and the winner's cancel travels with its body.
DoH3's Exchange packed the query into a POOLED buffer and released it the moment
RoundTrip returned. But http3 writes the request body on a goroutine of its own
and returns as soon as the response HEADERS arrive — the body is still being
read. With the window held open the query on the wire diverges from the query we
packed at exactly offset 8192, quic-go's copy-buffer size: everything past that
was the next pool user's memory, sent to the resolver. Not a slowdown — a data
race and a small memory-disclosure primitive. The buffer now goes back when the
transport closes the body, which http3 does on every path, and can do twice.
Both files diverge from upstream again, hours after 0a6689b29 made them
byte-identical on purpose. Upstream carries the second defect in
dns/transport/https.go too; that file is outside this audit and is named in D27
so the next person finds it instead of rediscovering it.
sing-quic moves v0.6.2-0.20260525051024 -> v0.6.4-0.20260709034545. quic.go is
byte-identical across the two, so this neither duplicates nor retires the
packet-conn ownership fix — quic-go still does not own the socket. What it does
carry is the other half of the family we took only half of: clientConn.Close in
tuic/, hysteria/ and hysteria2/ now sets a past write deadline, word for word
the fix v2rayquic already had. We ship tuic and hysteria2. Cost, measured:
+256 KiB exactly on the stripped aarch64 binary and six indirect modules for a
realm port-mapping path nothing we generate can reach.
Tests are mutation-checked: reverting each fix makes them fail, with the text
quoted above. The DoH3 test carries its own control — it first proves the pool
does hand a released buffer back and that poisoning it lands, because a clean
result from an instrument that cannot produce a dirty one proves nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`running` changed meaning on 2026-07-26 (a8970b8ac): it was a hardcoded
true and is now the ENGINE's liveness (apply.go `Running: engineUp`).
dashboard.js was last touched on 15 July and stayed in the old epoch, so
a dead engine made the page report "Daemon (shaterd): not running" in
red, advise "start the Shater service first" — the service was running —
and DISABLE the button to the panel, which is the one place the config
can be fixed. The holding plane keeps management reachable on purpose
(netplane/nft.go: "The operator can always get in to fix the config");
LuCI was the only thing taking that guarantee away.
Daemon liveness is now derived from the wire, not from `running`. "The
ubus call returned" is not enough either: `shaterd status` EXITS 0 WITH
A FABRICATED STATUS when the daemon is unreachable (cmdStatus offline
stub), and that stub is the apply.Status zero value plus a UCI read — so
it carries enabled/table/kill_switch but leaves `plane` at "", a value
no live daemon emits. A known plane word is the positive proof a daemon
answered; an explicit empty one is proof none did. Everything else —
{} from a failed call, {"error":...} from the plugin (also what a live
but WEDGED daemon produces), a pre-`plane` daemon — is unknown, and
unknown is an unlit lamp, never green. The launcher button is never
disabled again: a mint that fails already reports itself.
"Interception: active" is gone. apply.go says of `active`, verbatim:
"Never render it as 'we are proxying'" — it is the run latch that gates
hotplug and cron, it stays raised while the engine is down and the LAN
is blocked, and this page painted it green next to two more green lamps
in exactly that state. It is now "Service latch", and its lamp reports
only whether the latch agrees with globals.enabled. The row that was
missing is `plane`: full / hold (LAN->WAN BLOCKED) / none. `traffic` is
shown too, because plane=full is not "tunnelled" — a `default -> direct`
router has a full plane and no tunnel at all.
The rpcd plugin's status docstring listed five fields of fourteen and
had done since before half of them existed; it now describes the real
shape and the two fields that are easy to misread.
tests/status-readout.test.js runs the derivation against six recorded
status shapes with no browser and no router. Mutation-checked: reverting
to `st.running` fails 14 assertions including the operator-visible
"not responding - start the Shater service" over a live daemon;
restoring the "Interception: active" row fails 9; putting
openBtn.disabled back fails 1 by name; opening the closed plane list
fails 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
A rule pointing at node:awgout, which was already the first hop of the
default-route chain, took the house off the internet for two minutes.
The pass saw one private key materialised twice, kept the copy that
sorted first alphabetically, and fail-closed everything that routed
through the other one — which happened to be the default route for all
traffic.
The mechanism was right and the framing was wrong. The physical limit is
one DEVICE per key, not one mention per key. Two copies that build the
same device — same key, same peers, same address/MTU/AWG parameters and
the same dialer — are one device written down twice, and there is nothing
for them to fight over. Those are now MERGED: one survives and every
reference to the others is rewritten to it, silently. That makes the
shape the owner wanted expressible: one chain using awgout as an
intermediate hop and another using it as a terminal, both entering over
the same egress, coexisting on one device.
Identity is the marshalled options blob rather than a hand-picked field
list, so a field added to WireGuardEndpointOptions or DialerOptions later
reads as "different" instead of being silently merged.
Only a real incompatibility — different detour, different peers,
different device parameters — is still two devices, and then:
- the survivor is chosen by WEIGHT, not by tag order: reachability from
route.Final (the default route) dominates, breadth of use breaks
ties, tag order only settles a true tie;
- the warning names the consequence. "Everything that routed through X
is fail-closed" is equally true of a stray test rule and of the whole
house's default route, and that is what the operator read it as. It
now says which of the three it is, measured on the finished config:
the default route is dead, or it survives via another path, or it
never touched the lost copy.
A merge must not rename away the subscription fetch detour: that
reference lives in the model and is resolved against the running box, so
this pass cannot rewrite it. Such tags win the survivor slot outright,
which costs nothing since every copy in a class is the same device.
Tests: identical copies coexist on one device; a real incompatibility
keeps the default-route copy even when it sorts last and says so; the
warning does not announce an outage when the default route survives
through a group, and does announce one when it dead-ends behind a
surviving exit; no duplication at all is a no-op. All seven mutations
(merge off, weight off, member-dedup off, pin off, detour-following off,
consequence collapsed, plus a positive control) fail the suite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`shaterd apply` exists for one reason: snapshot the last-good, apply, and arm
an automatic rollback so a change that costs you access to the router undoes
itself. It answered `{"changed":false}` and not one word about that.
On the live router (2026-07-26) that was a trap. The operator edited UCI, ran
`uci commit`, the `config.change` reload trigger had already restarted the
daemon, and the fresh daemon applied the new config on startup. By the time
`apply` ran there was nothing left to apply — and the last-good it snapshotted
as the ROLLBACK TARGET was the newly applied config itself. The watcher was
armed onto the very configuration it was meant to protect against: firing it
would have restored exactly what was already loaded. No safety net, no word
said, house offline.
The verb now answers the question it exists to answer, in a closed vocabulary:
rollback_armed true ONLY when a window was armed AND its target differs
from what is running. An armed watcher pointing at the
running config is not a net and is not reported as one.
reason applied | already-applied | nothing-to-apply | disabled |
commit-confirm-off | config-unreadable | apply-failed
message the same thing in the operator's words, never empty.
The two "nothing moved" cases are told apart where they CAN be: an
/etc/config/shater mtime later than this daemon's start, with the running
config already matching it, can only mean a reconcile beat this command to it
(reason=already-applied). Where they cannot — the `uci commit` reload trigger
is stop+start, so it moves the daemon's start past the edit — the text says
so instead of reading as success: no net, harmless if you changed nothing,
unprotected if you did, and shaterd cannot tell which.
Two silent holes surface as a side effect, both previously reported as plain
success: `confirm_timeout=0` (the SHIPPED DEFAULT in
openwrt/shater-core/files/etc/config/shater) makes ArmRollback a no-op, and a
failed post-apply ReadUCI skips the arming entirely.
Arming behaviour is byte-for-byte unchanged — this only makes its absence
visible. A real safeguard for the already-applied case is separate work.
Tests are mutation-verified three ways: reverting classifyApply to the old
{changed,error} fails 11 tests; blinding the mtime discriminator fails exactly
the discriminating one (and falls back to the honest ambiguous text); making
sameConfig always report "different" fails every invariant that forbids
claiming a net over an identical target.
NOT verified on hardware: local_openwrt was held by another agent, so the
control-socket round trip and the real mtime/daemon-start comparison have not
been exercised on a router.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`fetch_detour=chain:<X>` never worked. engine.ViaToTag maps "chain:X" to the
bare tag "X", but the generator materialises a chain as one wrapper per hop —
chain-<X>-h1..chain-<X>-hN — and routes into the LAST one. The lookup missed and
the update failed with "unknown outbound tag".
It failed CLOSED, so the feed was never pulled over the plain WAN by this path.
But the miss had a sharp edge: when a node or group happened to share the
chain's name, the lookup HIT it, and the subscription was fetched through a
completely different outbound with nothing said.
Applier.HTTPClient now resolves chain: before the engine sees it, against the
tags the RUNNING box actually holds (outbounds unioned with endpoints — a WG hop
is an endpoint and Outbounds() does not list those), mirroring the generator:
the highest-indexed chain-<X>-h<i> wrapper is the entry, and a chain that
flattens to one hop IS that hop. Every other via form is passed through
untouched.
The case the generator cannot serve is named rather than papered over: chains
are built lazily, only for a chain some enabled rule/egress/DNS detour targets,
and a fetch detour is not one of those references — so a chain nothing else
points at has no outbounds at all. That, and every other miss, is an explicit
refusal wrapping engine.ErrOutboundUnknown (the panel already maps it to 400).
Never a fall back to direct: that would put the feed and the owner's real
address on the plain WAN, which is the thing fetch_via=proxy is set to avoid.
Tests are mutation-checked. Pre-fix behaviour resolves "work"/"solo" and kills
every chain case; first-hop-instead-of-last, member-copies-count-as-hops,
dropped pass-through, and a silent direct fallback each kill their own test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`.dns-filter-note` under the endpoint-resolver readout is a DIRECT child of
`.dns-filter-card`, so it is a grid item. With no explicit span it auto-placed
into column 1 — the toggle's `auto` track — and sized that track to its own
max-content: 237px at 390px, 322px at 1280px. That left the `1fr` copy column
with 0px, so "Network-wide ad & tracker blocking" laid out one word per line
and spilled 2px past the viewport, scrolling the whole page sideways on a
phone. On desktop the same cause parked the 52px toggle in a 322px column,
270px away from the copy it labels.
Measured at 390px: documentElement.scrollWidth 377 vs clientWidth 375. With
`grid-column: 1 / -1` on the footnote: 375/375, and the track list goes from
`237px 0px` to `52px 185px`. Cancelling just that one declaration in the live
DOM puts 377/375 and `237px 0px` straight back, so nothing else contributes.
Verified with playwright over 320/360/375/390/414/430/480/560/640/720/768/
1024/1280/1440: zero horizontal overflow at every width, with every rule
editor open, all three master toggles flipped, every source tab, and every
resolver type. No `overflow-x: hidden` anywhere — the page does not scroll
sideways because nothing overflows, not because the symptom is hidden.
Focus rings and prefers-reduced-motion re-checked and unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The old file pinned every subagent to fable — which broke the moment that
quota ran out mid-session — and spent half its length on panel scaffolding
that has been done for weeks. It said nothing about the test gate, the
testbed, or the hardware router, so none of that reached a subagent unless
it was retyped by hand into the brief.
What is new is not advice, it is the list of things whose absence cost a
day each: a test must be mutation-checked or it is decoration; an
instrument with no control proves nothing; a subagent must be told it may
refute the orchestrator, because the best results this project has had
arrived exactly that way; a formally-true sentence that reads as "it works"
is still a lie.
Skills are now a table mapping this project's areas to the skills that
cover them, with the rule that they are invoked BEFORE the work rather
than after something failed to run, and that every brief must name them —
a subagent cannot see this conversation and will not guess they exist.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
shater-l3 was created at 1420, the WireGuard payload budget, copied one
layer too far out. It bought nothing: what actually goes into the tunnel
is sized by sing-tun's forwardToPort against Port.PortMTU(), which
already fragments to the outbound MTU without DF and answers a
well-formed `fragmentation needed` quoting it with DF. All 1420 did was
make the KERNEL split every packet above 1392 bytes of payload on its
way into the device -- and a fragment is the one thing sing-tun will not
judge. Dispatch returns on parsed.fragment before calling JudgeFlow, the
fragments reach the gVisor stack, it reassembles them, and the ICMP
forwarder's installFlow demands an unspecified port address that a
WireGuard endpoint never has. So it declined and answered the echo
itself. `ping -s 1392` honest, `ping -s 1393` a lie, and only for the
outbounds the feature exists for.
65535 rather than merely "large": no IP datagram can exceed it, so the
kernel cannot fragment at this device for any packet ever. Anything
smaller leaves a band open and re-opens the class. It is also sing-box's
own default TUN MTU on Linux.
Memory was measured, not argued. Three paired runs of the integration
test under -test.memprofilerate=1 allocate 5.41/5.48/5.47 MB at 65535
against 5.76/5.46/5.70 MB at 1420, and a -diff_base profile puts every
difference in netlink interface enumeration. Nothing in the read path
scales with the MTU: gVisor reads through fdbased.BufConfig, which
sing-tun pins to one 65535-byte view regardless. I predicted a ~1.8 MB
saving from GSO switching off above 49152 and was wrong -- protocol/tun
turns GSO back on at StartStateStart whenever a FlowOutbound exists, so
the GRO scaffolding is there at both values. The corrected reasoning is
in the constant's comment so the next reader does not redo the mistake.
The integration test now reads the MTU back off the real kernel device,
which is the assertion the value exists for: a kernel that clamped it
would restore the forgery without changing a generated byte.
D25's KNOWN HOLE block is replaced with what is genuinely left. Chiefly:
a big non-DF ping does not start WORKING, it starts failing HONESTLY --
classifyReturn declines fragments on the way back too, so the packet
really leaves, the far host really answers, and the reply is not NAT'd
home. And a client that fragments on the wire itself is still uncovered;
that is the nft carve-out's job, with a warning that conntrack defrag
may reassemble in prerouting and leave such a rule unable to match.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The Proto picker was a closed list of the two transports and the ten
sniffed L7 labels, and anything else drew "<value> — never matches".
The engine now routes ICMP by rule (Rule.Proto accepts icmp, icmpv4,
icmpv6), so a working ping rule was rendered as a dead one and could not
be created here at all — the operator had to hand-edit /etc/config/shater
and then watch the panel call the result broken.
Adds a third group, "Layer 3". All three spellings are offered: they are
not synonyms — icmpv4/icmpv6 pin the rule's ip_version — so hiding the
narrowing would both strand a capability outside the UI and silently
widen such a rule the first time someone edited it here.
The doc comment no longer claims the list IS generate/route.go's
sniffedProtocols; only the middle group is. ICMP goes to the emitted
rule's `network`, never to `protocol`, which is the whole reason it never
matched as a sniffed label.
An unknown value is still kept and offered as written, but the
never-matches flag is now judged on the lower-cased value, the way the
engine judges it — a hand-written `ICMP` is a live rule, not an inert one.
Verified: npm run build clean (tsc --noEmit + vite build); an icmp rule
added through the panel renders as a plain "PROTO icmp" chip; no
horizontal overflow at 360px.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The panel offers fwmark_base and table_base as free hex fields under
"Advanced" and nothing has ever checked them. What makes that more than a
footgun is that the derived values are invisible from the number typed: the
L3 mark is base+0x80, so 0x7f lands it exactly on 0xff — the loop-guard mark
the engine stamps on its OWN traffic — and `ip rule fwmark 0xff lookup 8200`
then captures everything the engine sends and routes it into the engine's
TUN. The router loses the internet the moment l3_tunnel is switched on, for
a reason nothing on screen connects to a collapsed section. fwmark_base 0xff
had produced the same failure since long before the L3 offset existed.
table_base is worse and got the same treatment: its derived values can land
on the kernel's own table ids, and teardown does `ip route flush table <n>`.
It is count-sensitive (egress #i uses base+0x10+i), so the check takes the
egresses rather than living in ValidateGlobals.
Written as "derive every value this layout produces, then look for
duplicates and reserved ids" rather than as a blacklist, so a future offset
is covered by construction. The layout constants are duplicated from
netplane (the import only runs one way) and pinned by netplane's
TestMarkLayoutConstantsLockstep.
Warn-only, like every check in this file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Two halves of the same omission.
1. A fwmark lookup that finds an empty table does not fail — it falls
through to main. Every mark-driven table now gets an `unreachable
default` at the maximum metric: it loses to any real default route while
one exists, it has no device so the kernel never garbage-collects it, and
it turns "lookup failed, try main" into "lookup succeeded: unreachable".
The fallthrough stops depending on somebody reading a warning at the
moment an interface goes down. Deliberately not gated on the kill-switch:
that switch decides whether traffic may escape the tunnel, while an egress
binding is a statement about WHICH UPLINK, and silently substituting a
different one is not what "fail open" was meant to permit.
RoutingPresent's "does this table have a default route" test is tightened
in the same breath, or the floor would answer it and turn the safety net
into a blindfold.
2. RoutingPresent had never heard of addL3Routing. This is the same defect
its own comment describes as already caught twice ("a presence check must
cover everything its Apply counterpart installs"), committed a third time
— and its trigger needs no interface to go down: editing a node URI
restarts the engine, the kernel destroys shater-l3 and takes `default dev
shater-l3 table 8200` with it, the rendered nft text is unchanged, so the
fast-path skipped ApplyRouting forever and LAN ping stayed dead until
someone restarted the daemon.
TestRoutingPresentSeesL3Table, TestEgressTableGetsFailClosedFloor and
TestEveryStampedMarkIsRoutedAndVerified all fail on the code they replace.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The forward chain let untunnelable-egress traffic past the kill-switch on
the strength of its fwmark alone. `ip rule fwmark X lookup N` does not
deliver the packet to table N, it delivers the LOOKUP there — and a lookup
that finds nothing falls through to main. So when the egress interface goes
down and the kernel garbage-collects its default route, every non-TCP/UDP
packet from the LAN is still stamped, still accepted here (above the
fail-closed drop), and leaves out the plain WAN with the router's real
address. Nothing we render changes, so no apply runs and nothing notices.
Ordinary egress traffic never had this hole: the engine binds those sockets
to the device, and a dead device fails the socket. The untunnelable-egress
path is made of nothing but a mark, so the accept now carries the second
opinion instead — `meta mark X oifname "dev"`, strictly narrower than either
half, true only when the routing did what the mark asked. The comment being
replaced argued correctly that oifname ALONE would be too loose, then drew
from that the conclusion that oifname should be dropped rather than added.
Same conjunction in the holding plane, where it is theory (that plane stamps
nothing) but where a bare mark accept has no business sitting.
Also folds the egress device resolution into one EgressDevice(), because the
binding and model.ValidateUntunnelableEgress had already drifted: the
validator trimmed the interface name and the binding did not, so `option
interface ' '` gave a panel saying "the option is ignored" over a data
plane that was marking packets for a table nobody built.
TestUntunnelableEgressAcceptIsBoundToItsDevice and
TestUntunnelableEgressResolutionLockstep fail on the code they replace.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
The drop that keeps a ping from reading as tunnelled lived in
preMatchFlow, overriding the pre-declared continueResult. That covered
every exit of THAT function and none of the walk above it: the
prepareMatchMetadata error return (which arrived later, with the shared
metadata refactor), the sniff bail-outs, and the default: arm of the
rule-action switch all returned PreMatchContinue on their own.
adapter.JudgeFlow maps Continue to tun.ActionAccept, and sing-tun answers
Accept by rewriting Echo into EchoReply itself -- the exact forgery this
delta exists to remove. Narrow paths, but paths.
PreMatch is now a funnel over the renamed preMatch walk, so the guard
sits on the single return value and cannot be outgrown by a new exit.
PreMatchBypass joins the drop: sing-tun implements ActionBypass on the
nfqueue plane only, so on the TUN path it lands in the same default: arm
as Accept and forges too.
Every ICMP case has an explicit TCP/UDP twin; the JudgeFlow mapping
table is pinned outright, including the one fix that must NOT be made
there -- refusing ActionFlow for a port whose address is not unspecified
would drop every ping through WireGuard/AWG, because the forward
dispatcher and the ICMP forwarder share that function with identical
arguments and only the latter needs an unspecified address.
That leaves a real hole open, now named in D25 rather than papered over:
a FRAGMENTED echo to a WireGuard/AWG outbound is still answered by the
router. The dispatcher returns before asking for a verdict at all when
the packet is a fragment, and the reassembled packet reaches the ICMP
forwarder, whose installFlow demands the unspecified address a WireGuard
endpoint never has. The two fixes that would close it both live outside
pre-match and are written down; the Consequence paragraph is scoped
until one lands.
The stack comment in generate/inbound.go repeated the "only gvisor
really forwards ICMP" argument that D25 itself retracts -- both stacks
run the same ForwardDispatcher first. Brought in line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
`icmp` fell through ruleMatchers' proto switch into RawDefaultRule.Protocol —
the SNIFFED-L7 field, compared against what the sniffers labelled a connection.
Nothing ever labels a flow "icmp" (PreMatch skips the sniff action for an ICMP
flow outright), so the rule was structurally valid and permanently dead. That
made the whole L3 ingress unusable on a real config: with no way to write "ICMP
goes here", every ping fell to the catch-all, which resolves to the chain's last
hop — a group of VLESS nodes that cannot carry layer 3 at all.
icmp is a NETWORK. NetworkItem.Match is a map lookup over metadata.Network, and
adapter.JudgeFlow sets that to N.NetworkICMP for BOTH ICMPv4 and ICMPv6 (one
case covers both protocol numbers), so there is exactly one network value and it
covers both families. `icmpv4`/`icmpv6` narrow that same network with an
ip_version item instead of inventing a second one: metadata.IPVersion comes from
the destination address, and an ICMPv6 packet always has an IPv6 destination —
no false positives, no false negatives.
An ICMP rule that cannot fire is not a dead setting: ICMP has no fall-through,
so route.preMatchFlow DROPS it. Four ways to get that silently are now reported:
l3_tunnel off (nothing enters the engine at all), icmpv6 with ipv6 off (neither
the nft mark nor the TUN address exists), a port matcher next to it (JudgeFlow
zeroes both ports), and a target that cannot carry layer 3 — decidable from the
model, because the capability is fixed by the outbound TYPE: only wireguard/AWG
endpoints and the direct outbound behind direct/interface egresses declare
N.NetworkICMP. A mixed group gets its own text (the answer follows group.Now()),
`block` gets none (dropping the ping IS the policy), and an unresolved target
gets none either (ruleKillFallback already said the louder thing).
Wording stays clear of shater/apply's criticalMarkers on purpose: a failed ping
is fail-CLOSED, and a cosmetic alarm is how the real one stops being read.
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.
Two halves.
Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.
Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.
The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
Transient, to be deleted when omp/work merges. Everything meant to outlive the
merge is already in D25/D26 and the lx changelog; this file is the part that is
only useful while the branch is still a branch — the verification commands, the
testbed recipe, what was proven on hardware and what was not, and the six files
that will conflict on rebase.
D26's "no port-like selector" line disposes of NAT-based forwarding and nothing
else, and read alone it says "impossible" — which is false and would be
re-derived at the cost of another research pass. The endpoint is protocol-blind
in both directions, so ESP could ride it untouched with the client's own source
address and no NAT whatsoever. That was declined for two reasons worth naming:
lx-owned code in the forward hot path, and a server-side AllowedIPs prerequisite
that turns a router option into a deployment contract.
D26 writes down where the engine's boundary actually is, because the intuitive
answer is wrong and someone will look for it again: the WG/AWG forward path
never consults gVisor in either direction, so the limit is sing-tun's
ForwardDispatcher — its parser and its port-shaped NAT — and the kernel egress
was chosen because it clears that limit without a line of new hot-path code, not
because userspace "cannot". Tailscale documents the same boundary for their
userspace mode and is quoted as corroboration, with the caveat that ours sits at
the dispatcher rather than the stack.
D25 said two things that do not survive checking, and both are corrected in
place rather than left for the next reader to trip over. It blamed the netstack
for the ICMP-echo ceiling; that was the dispatcher. And it called `stack: gvisor`
mandatory because the system stack fakes ping — the system stack runs the very
same dispatcher first and only forges an echo for packets the dispatcher
declined, so gvisor is a deliberate choice (already linked via with_wireguard,
and the combination the integration test exercises), not a necessity.
The operator note says what the option buys and refuses to call an egress a
tunnel on its own say-so: with a WireGuard device it is one, with a second WAN
the destination sees that uplink's address. It also says what the option does
not fix — multicast IPTV stays broken — and that IPsec through NAT-T is ordinary
UDP that never needed any of this.
ESP, AH, GRE, IGMP and SCTP cannot enter the engine, and the reason is not the
one that looks obvious. A WireGuard or AmneziaWG endpoint forwards straight past
its gVisor stack — WritePackets reads the IP version and the destination address
and hands the raw bytes to the device, and the return path offers every
decrypted packet back before the stack sees it. WireGuard would carry ESP today
if anything handed it one. What refuses is sing-tun's ForwardDispatcher: its
parser recognises TCP, UDP and ICMP echo, and its NAT wants a port-shaped
selector that ESP, AH and GRE do not have. The retracted rationale is corrected
where it was written down, not quietly dropped.
So these protocols go to the kernel instead. untunnelable_egress names an
interface or tunnel egress; prerouting stamps that egress's OWN mark on
everything that is not TCP or UDP, and addEgressRouting has already bound that
mark to a table whose default route leaves via the device. Every protocol works
because nothing in the path has to understand any of them. No new mark, no new
table, no new code in the hot path.
Whether that is a tunnel depends on the device, and nothing here claims
otherwise: a WireGuard interface is one, a second WAN is a different uplink
whose real address the far end sees.
The wide `!= { tcp, udp }` filter is safe here and stays banned for the L3
ingress, for the same reason stated in both places: there the receiver is a
dispatcher that knows four protocols, here it is the kernel. ICMP is claimed by
the L3 ingress first when both are on. The local plane keeps its exclusions —
router-addressed traffic, private destinations, ICMPv6 ND/RA — and with IPv6 off
the marking is scoped to v4, because addEgressRouting installs no v6 rule then
and a marked v6 packet would fall into the main table.
An interface egress with an empty `interface` no longer resolves: IfaceDevice
defaults to br-lan, so it passed the binding while addEgressRouting skipped it —
mark set, no rule, straight past a closed kill switch and out the default WAN.
Both gated tests stand an engine up on shater-l3. Run together, the second met
`TUNSETIFF: device or resource busy` and failed for a reason that had nothing to
do with what it asserts — the first had closed its box and yielded while
unregister_netdevice was still catching up. Each passed alone, which is the
shape of a fixture bug that gets rediscovered rather than fixed.
The poll that already guarded the first test is now a shared helper both call.
It stays a poll rather than a sleep for the reason it always was: the removal is
usually immediate and a fixed wait would be either flaky or slow.
An interface egress is a direct outbound carrying BindInterface and a routing
mark, and direct builds its ICMP port from the very same dialer control — so
ping routed at that egress leaves through that device, marked, like every other
packet bound to it. Nothing said so. Both halves of that sentence are one
`common.Cast[*dialer.DefaultDialer]` away from being false: if the dialer ever
stops being a DefaultDialer, icmpPort is nil, PreMatchFlow declines, and ping
through the egress degrades to a drop without a single generated byte changing.
The gated test asserts the live outbound, not the config, because that is where
the cast happens.
The failure the codegen half guards is worse than a broken ping: losing
BindInterface or the mark does not stop the echo, it sends it out the main table
over the plain WAN with the real address, which is the one thing an egress
exists to prevent.
byedpi is a SOCKS outbound and cannot be a tun.Port, so ICMP aimed at it is
dropped. That is the honest end of l3-honest-drop and it is pinned too, because
the alternative the TUN stack offers is a forged reply.
Measured on a throwaway harness in a container: peak RSS of a process that
brought the engine up went from ~26 MB to ~28 MB with l3_tunnel on, three
paired runs. It is x86_64, idle, with an empty ICMP NAT table, so it stays
listed as unverified for the router — an indicative figure is more useful than
silence only if it says loudly what it is not.
D25 writes down the reasoning that is expensive to reconstruct: why a TUN rather
than TPROXY, why the interface is its own with auto_route off, why gvisor is
mandatory rather than preferred, and why the ceiling is ICMP echo — a boundary
in sing-tun's flow parser and gVisor's protocol set, not an unfinished edge of
ours. It also records what carries layer 3 and what does not, that masque could
and does not, and the two things still unproven: the live-router path end to
end, and what a second gVisor NIC costs in memory on the hardware.
D17 gains one line: its claim that TPROXY cannot carry ICMP is still true, and
is no longer the end of the story.
Both nft tables run and a drop in either one wins, so our forward accept for
shater-l3 decides nothing on its own: fw4 sees a device in no zone and drops the
forward, and the feature fails with exactly the symptom it was built to fix —
ping does not work, and nothing says why.
The zone names the device directly rather than a network. fw4 resolves a zone's
networks through netifd, and a proto-none interface for a device the daemon
creates is never up and contributes nothing, so list network would compile to an
empty device set. list device compiles to a plain iifname/oifname match that is
valid before the TUN exists and starts matching the moment shaterd creates it,
with no firewall reload at enable time.
It is seeded unconditionally, not gated on l3_tunnel: uci-defaults run once, and
a zone naming an absent device is inert. Gating it would mean the option could
be switched on and never take effect. The sections are named so a re-run is a
no-op instead of a second zone, and kmod-tun joins DEPENDS because /dev/net/tun
is not on a stock image.
Kernel TPROXY needs a socket to hand a packet to, so it moves TCP and UDP and
nothing else. Everything else reached the forward chain and met the untunnelable
policy, whose best answer was "let it out with your real address" and whose
default was "drop it" — so on a stock install ping simply did not work, and the
setting that fixed it did so by leaking.
The engine has been able to do better for a while: sing-tun's ForwardDispatcher
does real ICMP forwarding with NAT on the echo id, and a WireGuard or AmneziaWG
endpoint is a tun.Port that carries the packet for real. What was missing was a
way in, because nothing on the router could hand it an IP packet.
l3_tunnel (opt-in, off by default) adds one: the generator emits an "l3-in" TUN
inbound and prerouting fwmarks LAN ICMP into it. The interface is its own and
auto_route is off, so the main routing table is never touched and the fwmark
plus addL3Routing's ip rule are the only entrance — the TPROXY plane is byte for
byte what it was. gvisor is not a preference: the system stack forges echo
replies locally, which is the very thing this is meant to end.
Only icmp and ipv6-icmp are ever marked, and only after the local plane is out
of the way — the router itself, private destinations, and ICMPv6 ND/RA, which
mean nothing off-link and take v6 down if one neighbour probe is tunnelled.
ESP, AH, GRE, IGMP and SCTP are deliberately left alone: sing-tun's parser and
gVisor's stack know no such protocol, so marking them would black-hole the
traffic while looking like a feature. They stay with the untunnelable policy,
which also keeps its say over what happens if the ip rule fails to install.
Ping and Windows tracert now cross the tunnel; IPv6 traceroute shows only the
destination, because the return path recognises TimeExceeded for v4 alone.
PreMatchContinue is not "fall back to the ordinary route" the way it is for TCP
and UDP. An ICMP flow has no ordinary route: the TUN stack takes the packet back
and answers the echo itself, swapping the addresses and writing a reply
(sing-tun stack_gvisor_icmp.go). So a ping routed to any outbound that cannot
carry layer 3 — every proxy protocol; only adapter.FlowOutbound can — came back
successful, and the operator read a working tunnel off a packet that was never
sent.
That is worse than the packet loss it replaced. Loss is a fault the operator can
see and chase; a forged reply is a fault that reports itself as health, and it
reports it on the one tool anyone reaches for first.
preMatchFlow now overrides continueResult once, at the top, for
N.NetworkICMP. One hunk covers every exit that used to fall through — no such
outbound, a group whose selection is gone, an outbound whose Network() omits
icmp, an outbound that is not a FlowOutbound — and keeps the diff to three lines
against a function upstream will keep editing. JudgeFlow carries the same
verdict in its !isPort branch, because FlowOutbound and tun.Port are separate
interfaces and drift between them must not reopen the forgery.
TCP and UDP are untouched, and the test pins that as hard as it pins the drop.
The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:
case $action in restart|reload) keep;; *) DISARM;; esac
`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:
13:28 /etc/shater/boot.nft present
---- reboot
18s at_S22: NO_TABLE armor_file=NO_FILE
It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.
Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.
Also closed, found while proving the above:
* Every restart left the LAN in the clear for 80-90ms. The exit path was
`Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
transactions with no `inet shater` between them, leaving fw4's
`lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
& Apply. TeardownExiting arms first under the apply lock and skips the
delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
REPLACES the table, so the kernel never observes its absence.
35k-sample instrument: 7 and 6 no-table hits before, 0 across three
runs after.
* SaveBootArmor fsynced the payload but not the directory, so a power cut
could lose the rename that publishes it -- a boot with no armor and no
error anywhere.
`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.
Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
mock.ts was a static import and the mock switch was read from the query string at
runtime, so the bundle that ships inside the daemon carried a complete fictional
router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive,
without a single request to the daemon. The only tell was a line in the footer.
That is worse than any wrong number — there is no data at all and nothing says
so. It is out of the production bundle now, which is 21 kB smaller for it.
Unknown state stopped reading as good news in two more places. The kill-switch
tile treated an absent plane as armed, because the check was "not none" and
undefined satisfies it — the contract in the API types says the opposite. And the
apply page announced "daemon auto-rolled back" from its own timer, while the
daemon, seeing the state generation move, disarms and says it is NOT rolling back
in the log only.
Alerts moved to Settings. They are about the kill switch, apply failures, new
devices and subscription expiry, and they lived at the bottom of the DNS page,
while Settings mentioned them in prose with nothing to click.
Findings truncation is visible now: the notice that says how many were suppressed
arrives as info, and the attention list keeps only critical and warning, so past
fifty findings the operator saw forty-nine and no hint of the rest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.
They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.
Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.
The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.
And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.
The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DialEarly with a packet conn the caller made sets a flag that means quic-go does
not own it: closing the transport only stops reading from the socket. Neither DNS
transport closed it. On the QUIC one it was closed on a failed handshake and
never on success, so every redial — idle timeout, retry error, engine reload —
left a UDP socket for the life of the process. On the HTTP/3 one the library
drives its own reconnects, so the leak compounds without anything in our code
looking wrong.
That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on
every reconnect without closing the previous one. Both are now owned by a watcher
tied to the connection's own context, so the socket lives exactly as long as the
connection does.
This matters more than it did last week: the shipped resolvers are DoH, and DNS
is intercepted by default now, so the whole network's query stream rides this
path on a router with 512 MB.
The same upstream commit fixes both halves. We had taken the v2ray half and not
the DNS one — the third time this session a paired fix arrived half-applied, and
the first of those cost a day of debugging. These two files are now byte-identical
to upstream so a rebase cannot reopen it.
Also from that family: websocket and httpupgrade leaked their conn on failed
handshakes, and a QUIC stream's Close did not release a blocked write.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Booting with the armor loaded, or restarting through the handoff, left the status
saying the LAN was not being held while it was being dropped. Transient after a
successful apply, but permanent on the unreadable-config path — and there the
apply-failure alert words itself "traffic is NOT being blocked" at the exact
moment it is. That sends the operator to fix something that is not broken, past
the protection that is holding.
The table cannot be identified from here — netplane exposes no read-back and nft
does not keep comments — but identifying it is the wrong question. Holding does
not claim the holding plane is the object in the kernel; it claims the engine is
down and forwarded traffic is being dropped. A leftover full ruleset does that
too: with no engine socket the tproxy statement breaks its own rule before the
accept, so the packet reaches the forward chain unmarked and meets the primary
drop. What decides it is whether the last applied config was enabled and
fail-closed, which is exactly what the boot armor's presence already means.
So it is derived at read time rather than latched. A latch set from an inference
would have to be remembered in order to be cleared, which is the trap the active
flag already taught us. ArmHold also stops deferring to a table it cannot
inspect and installs its own render instead — the honest answer to "do not claim
a foreign table blindly" is to make it ours, and a fresh render beats a snapshot
that predates an interface rename.
Also closes the last of the detour transports: the subscription fetch took a
client and dropped it, and the exits that leak are the error ones, retried by
cron forever against a broken feed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.
The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.
The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.
An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.
A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Splitting the log sink made its writes asynchronous, so a download could miss
the last lines still in the queue — silently, with a successful response. Those
are the lines the operator came for: a log is downloaded to find out what just
happened.
The panel is handed a barrier, not the sink: a func() set once at startup, the
same shape as the reconfigure hook and the stats setter already in the tree. It
cannot write, reconfigure or close, so it stays a consumer, and nothing about
the sink's type reaches it.
The wait is bounded at the sink's own control budget and enforced on the panel
side, so a wedged writer cannot turn the download into the new place the daemon
gets stuck — the very thing the async split was for. Past the bound the handler
serves what is on disk. With no barrier installed the path behaves as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The header could not render "offline": it keyed on a field the daemon pinned to
true, so a dead engine behind a fail-closed plane showed a pulsing green lamp.
Health now needs both signals to agree before it reads as up, and a negative
from either is enough to say down — which is honest against the field that was
already honest, and stays honest now that the other one is too.
Deleting the last catch-all rule was described as "traffic will fall through to
the next rule" on the very row the page badges as the default route. What
happens instead is the kill switch: closed, the network loses the internet;
open, it leaves with the real address. The dialog now says which, by reading the
saved setting, and the toggle asks the same question — the generator only emits
enabled rules, so switching it off is the same event.
The master switch tore the whole plane down without a word, while deleting a
rule-set got a confirmation. Deleting a node or a resolver claimed to remove it
"from the config" without mentioning what still points at it, though the
reference finder was already there and used for renames.
Every Apply button armed the auto-rollback, and only one page said so. The
window is now recorded where all of them pass through, carried in a band under
the nav on every route, and persisted — so the countdown and the keep button
survive a reload, which is what made the window unconfirmable before. Overview's
Confirm button is gone rather than gated: Confirm cannot fail, so a permanently
live button could only ever report success.
Blocklists printed "filtering" from two config checkboxes without asking whether
the list had ever loaded — while the daemon grades a failed load critical. They
now show what the rule-set rows already showed, and say "not loaded — nothing
blocked" when that is the truth.
Also: the clock read UTC while every timestamp rendered in the browser's zone,
so the router appeared to have started in the future; the rule counter on
Overview counted saved rules rather than the ones in force, unlike the routing
page; and the hop badge counted the entry egress the rail below it does not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Changing any stats knob on the persistent backend deleted months of history.
The replacement store was opened before the outgoing one was closed, so it hit
the first one's flock, timed out — and the open path treated ANY error as
corruption and unlinked the file. Unlink of an open file succeeds on Linux, so
the new ring opened an empty database while the panel was still told the
backend had not changed. The store now hands its resources over before asking
for them again, and deletion is gated on an allow-list of real corruption
signals; a busy, unreadable or read-only file degrades to the in-RAM ring and is
left alone.
The holder's reads were unguarded in a subtler way, caught only after the gate
failed twice: the accessor took the read lock, returned the pointer and released
it, so the call ran outside. A reader could hold a store the swap then closed and
be served its empty answer — an empty page presented as data. The accessor is
gone entirely, along with the possibility of handing out an unguarded reference.
Readers still do not block each other; the swap now waits out reads already in
flight, which is a page at most.
The urltest group published its chosen node through two plain fields written by
the prober and read on every dial and every panel poll — while the selector next
door does the same job atomically. They are one value now, so TCP and UDP can no
longer be read as a mismatched pair. Nothing had ever dialled through a group
while it was probing, which is why the detector had never seen it; a test now
does, and reproduces it deterministically against the old shape.
Close on a group whose ticker had already stopped returned before closing its
channel, and Touch would then arm a fresh loop nothing could stop. Reached by
pressing Test in the panel and applying a config within the next two minutes: the
orphan kept failing probes against a cancelled context and writing forged dead
verdicts into the board the live generation selects from. Close is now final.
The log sink held one mutex across a blocking write. Under procd stderr is a
pipe, so a reader that stopped draining wedged everything that logs — engine,
panel handlers, signal loop — while the process still answered a signal. It is
split: a front that assembles lines and a writer that owns the destinations,
joined by a bounded queue that drops and counts rather than blocking. Proven by
restoring the old shape: the package deadlocks for the full ten-minute timeout,
parked exactly where the field symptom said.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Running was the constant true. The panel builds its header from it, so the
"offline" branch was unreachable code: with the engine dead and the LAN behind
a fail-closed hold, the operator saw a pulsing green lamp and, on the page
people open to fix things, "engine: running". The honest field sat beside it,
documented as the honest answer to are-we-proxying, and was read nowhere.
running now means shater is running: the daemon answered and its engine has a
started instance. active stays what it always was and is documented as such —
the "meant to be running" latch that gates hotplug and cron, not a health
signal. It is deliberately not cleared on hold, because the cron loop gates on
it and clearing it would switch off the reconcile that brings the engine back.
Two paths published nothing and so left the previous config's verdict standing
for as long as the fault lasted. A rollback with no snapshot re-applied the
engine and the plane and never touched the traffic verdict, so a router rolled
back to a direct default kept reporting the tunnel. And an apply that failed in
the netplane stage had already swapped the engine, then returned before every
publisher, so status described the config that was no longer running — and the
next reconcile, seeing an unchanged hash, failed the same way and published
nothing again. Both now publish, with an unknown verdict: after a no-snapshot
rollback the engine runs options this process does not hold, and guessing from
UCI would describe the config we rolled away from.
The severity classifier had drifted from the texts production emits. Markers
were compared case-sensitively against wording that had since changed, and the
entity pattern could not match a message beginning with an upper-case tag —
so a blocklist that failed to load graded as a warning while a typo in its URL
graded critical, and the panel's banner, which only lights for criticals, stayed
dark for the outage. RULESET-NOT-APPLIED and DNS-FILTER-NOT-APPLIED are now read
as the structural markers their producer documents them to be, so severity no
longer depends on wording at all. Five markers that matched no living text are
deleted; three protection-section texts drop to warning, because a blocklist
that is stale but still blocking lights the alarm on most reconciles behind a
flaky link, and an alarm that is always on is how the real one goes unread.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The block policy collapsed into direct whenever routing's final target was
direct — the common "tunnel only what is blocked, everything else direct"
shape. So ICMP, ESP, AH, GRE, IGMP and SCTP left with the client's real
address under the setting whose own field doc promises "nothing ever leaves
with the client's real IP", including a standing VPN on the real address,
which is exactly what the middle rung exists to separate out.
Both ends of the ladder now short-circuit before the plan is consulted and
neither may consult it: direct accepts everything, block emits no line at all
and lets the fail-closed drops the caller writes next do the work.
A rule scoped by source could also widen the other family: emit() skipped a
family whose destination list was empty but not one whose source list was, so
a rule carrying only IPv6 source prefixes rendered an IPv4 line with no
ip saddr clause — an accept for every IPv4 host on the LAN. The two halves now
read "scoped" the same way the catch-all collapse already did.
No destination plan is built for block at all now. It is the shipped default,
and a geoip-backed plan is ~159 000 prefixes pushed into kernel memory and the
ruleset text for a policy that cannot use them.
The operator-facing texts said IPTV works. It does not, on any of the three
rungs: inbound multicast is never matched by these rules and a client's
outbound multicast UDP dies at the fail-closed guard regardless. Saying
otherwise invited trading the ESP/GRE block away for nothing. What actually
stops working under block is stated instead, and precisely: raw ESP/AH and
GRE, but not IPsec through NAT or any UDP VPN, which are ordinary tunnelled
traffic.
TestOnlyPinnedAddressIsTunnelled is how this hid: it asserted, on the default
policy, that an exception line was emitted, and read that as the feature
working. It was block rendering direct. Its render assertions move to icmp,
where they mean something.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The posture was inverted. A client using the DHCP-supplied resolver — the router
itself — was NOT intercepted: dnsmasq answered and forwarded to the ISP in the
clear, so the filter, the blocklists, the per-device rules and BlockDoH were all
inert for exactly the clients that did nothing wrong. A client that hardcoded
8.8.8.8 to route around us WAS intercepted, by the catch-all. Meanwhile the
docs promised no DNS leaks. The default now matches the promise.
Turning it on crosses a threshold that was already dangerous for anyone with two
resolvers. Above one transport, a node's domain server address stops being
resolved by the transport directly and goes through the client DNS plane
instead — so a blocklist entry, a block_doh NXDOMAIN or any dns_rule can answer
your own node's hostname, and one sloppy line in an ad list stops being an ad
that got through and becomes a tunnel that never comes up.
So the fix is gated on having two or more transports, not on the intercept
toggle: resolver_default plus resolver_fallback always reached that threshold,
long before this change. When no endpoint_resolver is configured the plane now
carries a bootstrap server — the default resolver cloned with its detour
dropped, keeping its type, so a DoH default stays DoH and only the tunnel hop
goes. An explicit endpoint_resolver still wins.
This is not a restore of the previous behaviour and the comment says so: at one
transport the dialer used the default resolver WITH its detour, so a lone
DoH-through-the-tunnel resolver was already a bootstrap loop. It is strictly
better than what came before.
Existing installs keep whatever they set — the config file is a conffile and is
never replaced — and an explicit dns_intercept '0' survives the render-parse
round trip, which a default-true bool otherwise makes easy to lose.
The no-resolver warning stays, and no default resolver is shipped to silence it:
a placeholder would remove the sentence without moving a single query, and the
panel would then say a resolver was configured while nothing was filtered. Its
wording is corrected instead — .lan keeps working through the built-in local
transport, which the old text denied.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A resolver whose detour no longer resolved fell back to "the default outbound",
which is not a default at all — it is a plain system socket. Every other place
in this generator fails such a reference closed, with an essay explaining why,
and wgdedup rewrites the very same field to block when it drops an endpoint. One
field, two opposite policies, and which one applied depended on whichever code
noticed the breakage first. A resolver detoured through a node the operator
switched off therefore handed the whole network's query stream to the ISP in the
clear, while the kill switch held the traffic itself.
It now fails closed, and the warning says what that means: the resolver answers
nothing, and if it is the default one, name resolution stops network-wide until
the target is restored. A dns_rule naming a missing resolver used to be dropped
whole, sending exactly the names the operator singled out to a resolver they did
not choose; it keeps its matchers and answers NXDOMAIN instead. Not a reject
action — one built in Go with an unset Method panics the engine at match time.
RoutingPresent never looked at per-egress rules or tables, and applyLocked skips
the whole routing stage on its word. So an egress table wiped by an ifdown was
never restored: the marked traffic fell through to main and left over the plain
WAN, permanently, with plane full and no warnings. It now verifies each binding
it installed, recording intent rather than outcome so a broken egress keeps the
plane reported absent and heals when the interface returns.
addEgressRouting discarded every ip error, so an egress that failed to install
reported success and the panel drew it green. Failures are now critical warnings
naming the egress, the device and what ip said — but still warnings, because
returning would abort the apply and punish the household for one bad uplink.
Also anchors the fwmark check: with a small fwmark_base the main mark is a
literal prefix of the first egress mark, so a substring match could answer "the
main rule is installed" while looking at an egress rule.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guard refused to start an AmneziaWG endpoint whose detour chain reached a
WireGuard one, and refused silently: not an error, just started=false, after
which every dial failed with "WireGuard is not ready yet". A selector hook went
further and suspended an already-working node the moment its group switched to a
WireGuard member.
It existed because AmneziaWG inside WireGuard hung the kernel on Android. We do
not ship Android, upstream dropped the guard once the cause was gone, and the
cure landed here yesterday — the ClientBind reserved-gate plus the submodule pin
that carries its twin. So the tree held both the cure and the prohibition on
using it, and the configuration simply did not come up while looking like a node
that "just does not work".
Also takes the two fixes that belong with it. ClientBind.conn was read on a
lock-free fast path and written under a mutex; upstream found that race with the
same end-to-end test we wrote yesterday, so we had taken one half of a pair
again. And the outer WireGuard UDP socket forced DF, unlike direct, hysteria and
tuic — with encapsulation the datagram regularly exceeds the path MTU and the
kernel drops it instead of fragmenting, a symptom indistinguishable from the bug
we spent yesterday on.
The race needed its own test: the existing e2e run did not flag it under -race
even at -count=15. Eight goroutines over both connect branches reproduce it
deterministically, naming the lock-free read and the guarded write.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The fork had a full suite and no CI that ran it. Upstream's test workflows
trigger on stable/testing/unstable; this repo only has main. And Gitea does not
read .github/workflows at all once .gitea/workflows exists, so those files were
decoration here. 115 of the 116 test files under shater/** had never executed in
CI even once, which is how TestDNSFilterRemoteBlocklistHTTPClient stayed red
across two published releases without anyone noticing.
The gate is a job inside release.yml that build-apk needs, because a separate
workflow cannot block another one. It runs the suite under the shipped tag set,
on Linux — 6 of 7 test files in transport/wireguard and 12 in shater/generate
compile only there or only under those tags, and those are exactly the files
covering AmneziaWG.
Three guards stop it from passing by running nothing, which is the failure this
whole change is about. The tag set may only ADD test files, never remove one.
Every package go list says has tests must appear as "ok <pkg>" in the output, so
a suite that collapses to "no test files" fails instead of passing. And the
panel run counts its test files first, because node --test exits 0 with "pass 0"
when the glob matches nothing.
The publish step used to exit 0 having published nothing: its assertions all
live inside a loop over artifacts, so an empty directory ran the body zero times
and reported success. It now counts what it published and fails on zero.
Verified by extracting the shipped step text and running it against stubs: empty
artifacts gives exit 0 before and exit 10 after; the rolling-release readback
still fires its own exit 14.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The wireguard-go submodule is pinned to 7d15f33, which lives on lx-awg2-v005.
.gitmodules named `lx` — a separate line, 42 commits one way and 131 the other,
with no common recent history.
That is a loaded gun rather than a cosmetic mismatch. `lx` has no hasReserved()
gate in conn/bind_std.go at all, so a single `git submodule update --remote`
would move the pin there and silently restore the defect fixed yesterday: the
bind shreds the AmneziaWG magic header of every transport packet, handshakes
complete, no data moves, and no chain containing an AmneziaWG node carries
traffic. It would also drop the padding-overrun fix and the v0.0.5 re-graft.
Nothing about the checked-out tree changes — the pin is untouched. Only the
branch a --remote update would follow.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The short-circuit left the chain's exit tag untested, because the exit is itself
a hop and every hop behind the break was rewritten that way. A stand run caught
it — the test lives in a file that does not compile on the dev host, so nothing
local could have.
That is not neutral silence. selectExcluding ranks untested ABOVE dead and says
so in its own comment: with no fresh-alive member, an untested one is a better
bet than a known-dead one. Leaving a provably broken path untested is therefore
a positive preference for it over a path we merely know is dead.
The two readings answer different questions and now differ on purpose. Is this
hop's own node alive — unknown behind a break, so the card keeps untested and
blocked_by. Can this chain carry traffic — known, no, because the hop in front
of it was probed and did not answer. The board carries that second answer, which
is the one selection, the freshness gate and the manual test all read.
The exit verdict is derived, not dialled: it records the consequence of a probe
that did happen one hop earlier, and it is re-derived every pass, so the moment
the blocker answers the walk reaches the exit again and the next verdict there is
a real measurement.
Also keeps a routed group warm. Its checker used to stop on the idle timeout and
nothing filled in behind it, so a rule that fires rarely would show untested
while being in force and pay a cold probe on the first real request. The gate
that adds this work answers false when it does not know — the mirror of the one
that withholds work, so plain sing-box keeps the lifecycle it always had.
And the tls-spoof suite now skips without tcpdump instead of failing sixteen
times: a missing tool is not measured, not broken. The same distinction this
commit is about.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A member of a chain hop wrapper was reachable by two probers: ours, from the
observatory plan, and sing-box's, from the urltest group the wrapper actually
is. Two independent readings of one node can disagree, and then neither can be
trusted — which is worse than the wasted dial.
The group's own checker is the right owner. A hop wrapper's members are the
per-chain copies, each carrying the previous hop as its detour, so that checker
already travels the chain prefix — the path the traffic takes. The plan now
records who dials each target and the observatory skips the ones a live checker
owns, keeping only what no group covers: node hops, the AmneziaWG endpoint,
selector members, and the members of groups that have been stood down.
The jobs stay in the plan rather than being deleted, and that is load-bearing:
the short-circuit reads the plan as the map of which tags measure which hop, so
deleting a urltest hop's members would erase that hop from the map and quietly
stop it blocking anything — on exactly the chains the feature exists for.
The short-circuit therefore moves to the group as well, through a ProbeGate the
engine implements: a scheduled check asks whether the path in front of it is up
before dialling, while an explicit check is never refused. Nothing is stored —
the gate recomputes from the live board every call — and Touch still arms the
ticker even while blocked, because a hop that refuses to tick has nothing left
to notice its own recovery. The gate answers yes whenever it does not know:
refusing on missing information is how a system talks itself into silence.
Two grounds now exist for a group not to probe and they must not be merged:
stood down means no rule reaches it at all, blocked means the path in front is
down right now. Both doc comments say so and name the chain hop wrapper as the
case where the difference bites.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hop probes were independent, so every hop was dialled whether or not the path
to it existed. A hop is dialled THROUGH the hops above it, so when hop 2 had no
live member left, the probe for hop 3 failed at hop 2 and hop 3 was recorded
dead. Dead means "we tested this and it did not work" — but nothing was learnt
about hop 3 at all. One broken hop painted the whole chain dead and pointed the
operator at the wrong place, and every one of those probes was a dial with a
timeout down a path already known to be broken.
Chain jobs now run in path order and the walk stops at the first hop that reads
dead. Hops below it are not dialled at all and are reported untested with
blocked_by naming the hop that stopped the walk — the honest answer, since
nothing was measured.
Nothing latches. There is no blocked flag: the gate is a fresh read of the
health board at every hop of every pass, and the cursor rewinds to the top each
cycle, so the first dead hop is never behind a break and is always retried. The
moment it answers, the rest of the chain runs in that same pass. Only a positive
dead blocks; untested never does, or a cold start would never open.
Blocked hops are rewritten rather than annotated, because board records do not
vanish when the prober stops dialling — they age out on their own TTL, and the
worst version of that is a stale dead pointing at a hop that may be fine.
The exit tag is exactly what stops being dialled, so the group test would have
waited out its full deadline and then reported "not reached yet" about a chain
it already knew was down. It now names the blocking hop immediately, gated on
the same freshness watermark so a break seen before the request cannot
short-circuit a pass that may be about to find that hop alive.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The chain card gave a single verdict, so a dead hop was invisible: the operator
saw "the chain is unhealthy" and had to guess which of four hops to look at.
Meanwhile a group used only inside a chain showed "unused" next to a live
alive/dead count, which reads as a diagnosis when it only means nothing measures
it on that path.
Render the hops as a rail that severs below the first dead one, so which hop is
answered before a word is read, and split the two "not routed" messages into the
routing fact and the explicit non-fact. The group one names the case directly: a
group used only as a hop inside a chain reads unused here on purpose, and its
real health is on that chain's card.
Also fixes a bug this would otherwise have shipped: the readout painted every
ok:false in the critical colour, so "not routed" would have rendered as a fault
— the exact lie being removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A node reached only as a chain hop was being measured twice, and the reading
the panel showed was the wrong one. On a router in Russia that is not a cosmetic
difference: a node the chain carries fine behind a WireGuard hop is dead when
dialled straight out of the WAN, so the group card read "0 of 2 alive" while
that very group was carrying every packet.
Two dial paths existed outside the observatory plan. URLTestGroup.PostStart
warmed up every urltest group at box start whether or not any rule reached it,
and the panel's Test button reached URLTest.DialContext, whose first act is
Touch() — arming a ticker that re-swept those groups directly every probe
interval for the next thirty minutes. Both wrote under the BASE node tag, and
both dialled the base outbound, which carries no chain detour at all.
The observatory was never the liar: its plan roots come from the rules, and a
chain hop copy is stored only under its own tag, so no plan job could ever
write under a base tag. The fix is therefore to remove the other two paths, not
to touch the plan.
TestGroups now asks the observatory for an out-of-turn pass and reports what it
measured; a target no enabled rule routes to is not dialled at all and says so.
Unused urltest groups stand down their own self-check via a new SelfCheck option
(nil keeps today's behaviour, so every existing config is unchanged). The one
direct dial left is the exit-address lookup, which has no other possible source
— it now runs only for a target that is both routed and already read alive, so
it travels the routed path and never touches an unused group.
Chain hop wrappers are probed as measurements of their own and surfaced as
chains[].hops[], because "which hop is dead" is the question an operator has and
the chain-level verdict cannot answer it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The unit tests pin the reserved-byte gate on each side in isolation, which
would still pass if the two halves disagreed about when to apply it. This wires
two real wireguard-go devices together over loopback UDP through ClientBind on
both ends — the bind the detour path actually uses — configures ranged h1-h4
plus s4 and junk, and asserts an inner IP packet reaches the peer's TUN.
It is red against the unconditional clear and green with the gate, so it covers
the failure the field hit rather than the code we happened to write. Tagged
with_awg, so it runs under the shipped router tag set.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous suppression compared each line with the one before it, which the
field never obliges. A dead chain makes the engine cycle the same message
across three outbound tags, so no two identical lines are adjacent: on the
router it produced 854 daemon lines in a ~760-line syslog ring and exactly one
summary, all while claiming "repeated 1 time". The rest of the system's log —
netifd, dnsmasq, the kernel — was evicted anyway.
Track a bounded table of open series keyed by the existing repeat key instead.
The first copy of a key prints; further copies inside its window are counted
whatever arrives in between; the window end emits one summary per key. The
summary now names its message, because several can close at once and "last
message" would simply be false under interleaving.
The table holds 256 keys and evicts the least recently seen, never silently: an
evicted series with a pending count prints its summary on the way out, marked
so the truncation is visible. Close, Reconfigure and any fatal flush every open
series first — a dying daemon may never reach Close.
TestRepeatAlternatingNotSuppressed asserted that A B A B must never be
collapsed. That assertion was the bug. It is replaced by a stronger one: the
messages get separate series, separate summaries and separate counts, so
distinct events still never fold into a single number.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every config apply built a new box and left the old one alive. The engine's own
log gives it away: inside a single shaterd process, lines carried uptime
counters half an hour apart in the same second, and a live router was found
running four generations at once. A process restart cleared it, so the leak
accrued purely on re-apply.
That is not just wasted memory on a 512 MB box. Each surviving generation keeps
its WireGuard devices up, and two devices sharing one private key evict each
other at the peer — so the leak reproduced the duplicate-device defect between
generations, underneath the deduplication that only reasons about one config.
Retirement now has a hard budget: 5s, which is exactly sing-box's own
C.StopTimeout (past which upstream already calls a stop excessive) and stays
under C.FatalStopTimeout. It is paid after the replacement is serving and only
on an apply that changed something, so a no-op reconcile stays free.
A close that blows the budget is ABANDONED, not waited on, and the apply is
still reported as the success it is — the new box is built, started and
carrying traffic, and failing there would abort the netplane stage and leave a
stale ruleset over a healthy engine. The stuck instance is surfaced through
PendingCloses() into `shaterd status` and the panel, and clears itself if the
shutdown ever completes. Repeated applies over a stuck close no longer stack:
the abandoned generation is remembered, not re-created.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An AmneziaWG node worked standalone and died the moment it was placed behind
an egress or a chain hop: the handshake completed, the peer answered, and then
not one byte of data ever arrived. The peer never confirmed the session, so it
re-handshook every 15 seconds, forever.
ClientBind cleared bytes 1-3 of every datagram on receive and stamped them on
send, unconditionally. Those bytes are Cloudflare's "reserved" field. They are
also where AmneziaWG puts the upper three bytes of its little-endian uint32
magic header, so zeroing them collapses the value to its low byte, which falls
outside every h1-h4 range and makes the peer classify the packet as an unknown
type and drop it silently.
Handshakes survived because s1/s2 padding pushes their magic past byte 3 — the
clear only scribbled on the random junk prefix. Transport packets have s4 = 0,
so their magic starts at byte 0 and took the hit. That asymmetry is the whole
signature: session up locally, zero data through.
Only the detour path was affected, because Endpoint.Start picks StdNetBind when
the dialer exposes WireGuardControl (no detour) and ClientBind otherwise. The
gate had already landed in StdNetBind; ClientBind was its untouched twin. The
two implement one contract and are now commented as the pair they are, so the
next fix cannot again land on one side only.
Measured on the box: h4 spans 0x60728123-0x60728155, so zeroing bytes 1-3
leaves 35..85 — the captured transport packet began with 56, while a node
without a detour carried a correct 0x6b039798 at the same moment.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With two catch-all rules both enabled in UCI and a WAN profile enabling one
and disabling the other, the panel drew BOTH switches on while the engine
ran only one chain. GET /api/config is right to return the raw model — that
is the desired state the panel PUTs back — but Routing.tsx read the row
state and the active count from it too, so the interface claimed a setting
was in force when it was not. Same defect class as the Protected badge.
/api/rules/reachability now carries the effective flag and, where the active
profile changed the outcome, its name and direction. The annotation is a
DIFF of ApplyProfileRuleOverrides output against desired state rather than a
second reading of the profiles name lists, so profile logic is not
duplicated and cannot drift — an unmigrated rule the profile is forbidden to
enable produces no diff and gets no badge, with nothing here needing to know
about LegacyDst.
In the UI the two states stay separate: the switch remains the only carrier
of desired state and still writes UCI, while the effective state drives the
dimmed row, the badge, the banner and the header count. Mirroring the
effective state into the switch would be worse than the original bug — the
operator would be toggling someone elses control, and the profiles decision
would be written back as their own choice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
A broken outbound makes the engine repeat one line about once a second —
370 copies in six minutes. The routers syslog ring holds ~760 lines, so
within minutes it evicts the history of every other subsystem and our own
startup lines with it. Diagnosing the WireGuard duplication above required
restarting the service purely to catch the first seconds of a boot.
Collapse runs into "last message repeated N times". The comparison key is
level + text with the uptime field dropped: comparing whole lines would
suppress only same-second bursts, because that counter ticks. The per
connection "[id duration]" group is deliberately KEPT in the key — those ids
are distinct connections, and folding "50 connections failed" into one count
would be a worse lie than the flood. Window 5s, so a standing fault keeps
being reported instead of looking like a frozen log. fatal/panic are never
suppressed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two
releases, which made the whole package exit non-zero no matter what the code
did — a real regression would have drowned in the familiar red.
The cause is not the packages no-network fetcher stub, as it first appears.
ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION
alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so
the fixture fell into the TEXT-list path: downloaded by generates own
fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which
every assertion below then contradicted. No stub content could fix that; the
stub decides the lists contents, not the rule-sets type.
Give the URL the .srs suffix the test always meant it to have, so the engine
fetches the compiled set itself through the direct outbound. No assertion is
weakened and the no-network stub stays in place.
Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags):
338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
A node may be copied freely by this package: a per-chain hop copy and a
per-group egress copy are rebuilt from the share-link so each can carry its
own Detour. For vless that is right — a copy is another TCP client. For
WireGuard it is not: each emitted endpoint is a real device holding the
nodes private key, and a peer keeps exactly ONE session per public key.
Two devices from one key evict each other continuously, and with keepalive
on both the loop never settles: NEITHER passes traffic.
buildOutboundsAndEndpoints emits the base endpoint for every enabled node
whether or not anything references it, so a WG node used only as a chain hop
always produced two devices. That is what any chain containing a WG node
looks like — every such chain was permanently dead.
Observed on the box: two UDP sockets from shaterd to the same peer port, the
servers peer endpoint flapping between them, +32 bytes/min through the
tunnel and every hop failing with "context deadline exceeded".
Deduplicate once on the assembled options, which catches all three producer
paths by construction. Duplicates are DELETED, not merely unreferenced:
box.New starts every endpoint regardless of reachability, so a leftover
would still bring its device up and still fight for the session. Dangling
references go to block, never to direct — a consumer whose tunnel just
disappeared must stop, not fall out onto the plain WAN.
Subscription fetch detours seed the reachability walk (they are direct
references like any rule), mirroring engine.ViaToTag exactly, with a
tripwire test against drift. A config that genuinely needs two devices for
one key keeps one and fail-closes the rest with a critical warning.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
Four defects, all found by the owner on the live router, all of the same
family: something declared itself working while it was not.
WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave
"create WireGuard device: gVisor is not included in this build". The router
tag set carried with_wireguard and with_awg but not with_gvisor, so
sing-tun compiled its stub instead of the netstack every WireGuard device
needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving
requirement", so this was a broken promise, not a trim.
The tag itself was the small half. The tag set was the ONE build
configuration nothing in the repo tested: TestAmneziaWGEndpoint passes
because tests build with the full upstream tags. So the set now lives in
one file (scripts/router-tags.sh) and two guards hold it to the feature
list -- a static check that needs no tags, no Linux and no network (so the
next such gap fails on the developer's machine), and a behavioural one that
constructs every declared protocol through box.New UNDER THE SHIPPED TAGS,
where skipping is forbidden. Removing the tag now fails with the feature
name, the missing tag, and why: "Either add the tag back, or stop declaring
the feature -- those are the only two honest options." Cost: +2.8 MB raw,
+0.6-0.7 MB packed per arch. D23; D9 corrected.
THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came
from plane === 'full', which reports whether the data plane is installed --
nft table, policy routing, live engine -- and says nothing about where the
traffic goes. On a config with one `default -> direct` rule and no groups
the plane is fully installed and every packet leaves in the clear, so the
worst possible state rendered as the reassuring one.
The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the
moment they reach the engine, not from the model: buildRoute changes the
answer (a scheduled rule outside its window is never emitted, only the last
condition-less rule reaches Final, an unresolved target is rewritten by
ruleKillFallback), and re-deriving it anywhere else is a second
implementation that will drift -- model/reachability.go exists because two
already did. Four verdicts, not three: `blocked` is separate because under
a closed kill-switch with no catch-all nothing leaks, and calling that
"going out directly" is a lie in the alarm direction. Rider: Overview's
defaultTarget printed the highest-Order enabled rule as the default; a rule
becomes Final by having no conditions, whatever its Order.
"PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE
(B15). Once the browser suppresses dialogs, window.confirm returns false
immediately, so all 15 confirmations across 7 pages read as "cancelled" and
silently did nothing, with no way to recover from inside the panel. Replaced
with an in-app dialog the browser cannot mute: focus trapped and parked on
Cancel, Esc and veil cancel, focus returned to the opener, crit styling for
destructive commits. useConfirm() throws if the provider is missing rather
than falling back to a quiet false -- the failure mode being fixed.
HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed,
so a feed's nodes of those types vanished. The real landmine was one layer
up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so
without schemePrefixes the links were gone before any parser ran and the
fix would have looked complete. Undeliverable parameters are refused when
the node cannot work or would be less secure than the link asked (obfs,
pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it
survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot
produce a QUIC TLS config, and that fails at dial time, not at box.New.
Also: nodes added by hand can be named and renamed. The name is the
outbound tag, so a rename rewrites every reference in one PUT -- rule
targets, group members, chain hops, detours -- in the spelling each already
uses, and is refused outright when a group answers to the same bare name.
Subscription nodes state why they cannot be renamed instead of hiding the
control.
go build, go vet, go test ./shater/... (13 packages), panel npm run build
and npm test (13/13) all green. NOT yet verified on hardware.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and
main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has
no `opkg` binary at all. The 24.10 lane was building and signing a feed no
device could consume.
Removed jobs `build` and `release` with the scripts only they called
(ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh)
and the usign trust anchor dist/shater-feed.pub. A committed public key is
an instruction: it invites the old install path for a feed that is no longer
produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD
still holds the secret half, and a usign secret contains its own public half,
so the identity is reconstructible if a 24.10 device ever needs serving.
D7 is marked SUPERSEDED by the new D22 rather than deleted.
Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from
2026-07-24 while every tag run published its versioned release correctly.
The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest
(workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run
never touched the rolling pointer. Asset replacement was never the problem;
ci/gitea-release.sh already deletes before recreating. A router pinned to
the rolling URL sat on 0.2.0 while `apk update` reported success: silent
staleness, the failure mode this repo keeps having to close.
The rolling pointer is now published on EVERY run, tag runs included, and a
new assert reads the release back over the API afterwards: our three
tag-versioned packages at the built version plus the index and the key must
be present (exit 13), and no package asset at any other version may survive
(exit 14). Same class of check as sdk-build-apk.sh's package-version assert,
added for the same reason -- the previous failure mode was silent.
KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references
it. Docs state plainly that mini_router is deliberately pinned to a
versioned URL and that the hand-edit per release is the price of pinning.
Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can
no longer install our packages. Its 25.12 rebuild is in flight separately.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
Code review of a8ef887c5 + 244b7c419 ("a rule's destination is a rule-set,
and nothing else") found that the change rested on a comment that was not
true. ParseUCIExport dropped dst_domain/dst_ip on the strength of "the
migration is re-run on every load"; model.Migrate() actually runs only from
`shaterd migrate`, i.e. the service init and uci-defaults. The daemon's run
path, the SIGHUP reconcile and the panel's config write never migrate.
So an uncommitted migration (a full /overlay is the documented way that
happens) turned `list dst_domain 'bank.ru'` + `target direct` into a rule
with NO matchers, which IS the spelling of a catch-all: generate points
route.Final at it and the LAST such rule wins. One failed `uci commit` sent
every packet on the router out the plain WAN, silently.
Rule.LegacyDst is the tripwire. It is non-empty exactly when the config
still carries the removed options, and three locks hang off it:
- ParseUCIExport holds such a rule DISABLED. Chosen over "make IsCatchAll
false" alone, which only covers matcher-less rules: `dst_domain` plus a
`src` was never a catch-all, and routing it without its destination
would still have sent a whole subnet direct.
- IsCatchAll returns false for it, so it can never own route.Final even
if something hands its Enabled bit back.
- ApplyProfileRuleOverrides refuses to enable it (a profile with
`list enable_rule` would otherwise have defeated the parser).
ValidateRules reports it through the existing warning channel, before the
Enabled gate, so the one message explaining the outage is not suppressed by
the fact that caused it. The init script logs a failed migration to syslog
instead of discarding its exit code and stderr.
The write path had none of this. PUT /api/config decodes a Model straight
from the request body and render.go wrote `enabled` from it, so a panel
save erased the operator's lists (as did the subscription cron, which
re-renders the whole package), and a crafted body with Enabled:true and no
LegacyDst put a live matcher-less rule on disk -- the same whole-router
leak, re-entered from the other side. WriteUCI now reads DISK state and
refuses a rule-changing write over an unmigrated config (409, not 500);
non-rule writers pass and legacyDstOpts carries the options across so cron
preserves them; withDiskLegacyDst takes the field from disk so a fabricated
one can never reach the renderer.
Migration hardening: an entry list that migrates to nothing no longer has
its legacy option deleted (that made "matches nothing" silently become
"matches everything"); a hand-written rule-set whose name collides is no
longer allowed to swallow the entries; delete failures propagate instead of
bumping schema_version past them forever; every error path reverts the
staged uci delta so another process's commit cannot flush a half-migration.
untunnelable.go follows the destination out of the rule: a rule whose
rule-sets are known to match by name is still skipped by the ping/IPTV/VPN
plan, as its v1 form was. D21 documents the AND->OR widening for the
engine's TCP/UDP path; it does not follow that a leak-guard should widen
itself during an upgrade, and with target=direct that meant previously
tunnelled ICMP leaving with the client's real address. Inline rule-sets are
now read from the options, so an engine that has not started yet no longer
costs the operator their ping.
Rule-set vocabulary: `full:`/`suffix:`/`keyword:`/`regexp:` in a text list
fetched by URL were dropped with no diagnostic at all (normaliseListDomain
rejects any token with a colon) -- not "reported as an unknown prefix".
Unifying was rejected: published filter lists are full of colon-bearing
syntax, and a third-party `regexp:` is compiled into the router's matcher
and run per query. The difference stands and is paid for in diagnostics,
per list, on every generate. D21 gains the source/vocabulary table.
Panel: the add form warns about a matcher-less rule exactly as the edit
form does, from one shared predicate; its isCatchAll matches the daemon's
new one; an unmigrated rule reads as held-off rather than merely switched
off. The comment promising a "New list" button that D21 rejected is gone.
go build ./..., go vet ./shater/..., go test ./shater/... (13 packages) and
panel `npm run build` are green. NOT yet verified on hardware.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.
The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.
The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.
isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.
Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.
`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.
THE BARE-ENTRY TRAP, and why the migration is not a copy
A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).
`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.
`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.
THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)
Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.
Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.
ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.
Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.
untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.
Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.
Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.
With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.
Fix (upstream file, lx:tproxy_writeback_connect):
* CONNECT the write-back socket to the one peer it ever talks to. The kernel's
compute_score() rejects a connected socket for any other peer, and a
connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
reuseport selection outright — so it can no longer be handed a datagram it
will not read. Nothing about the reply changes: same spoofed source, same
single peer, Write instead of WriteTo. The unconnected path is kept verbatim
for a destination that cannot be bound (domain socksaddr).
* A failed cached write now CLOSES the socket instead of only dropping the
reference (upstream left the fd to the GC finalizer).
* TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
but the cache evicts lazily, so after the inbound is gone nothing wakes the
live sessions and each strands its write-back socket. Invisible upstream
(one close at shutdown); on this fork the engine is rebuilt on every apply,
so it was one stranded generation per apply.
Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.
The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.
Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.
Cause: `restart` is not synchronised end to end.
* procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
`stop; start`, and the `service delete` ubus call returns as soon as
SIGTERM has been SENT. `start_service` therefore re-adds the instance
(and runs `shaterd migrate`) while the outgoing `shaterd run` is still
executing its honest teardown.
* The successor's only defence was `daemonAlive()` -> exit(1), leaning on
procd's `respawn 3600 5 0` to try again five seconds later. That is a
blind retry, not synchronisation: it neither knows nor waits for the
teardown, and it turns every restart into a logged crash plus a
five-second hole with no data plane.
* `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
engine.Close of a several-hundred-outbound box flushes cache.db to
flash before the netplane teardown even starts — aborting the teardown
at an arbitrary point and leaving the plane HALF removed.
* Nothing in the tree ever touched conntrack, so flows that crossed one
of those windows kept entries formed against a plane that no longer
exists. For UDP there is no handshake to resynchronise on and every
retry merely refreshes the entry, so the flow stays wedged for as long
as the client keeps asking — a flow-scoped, permanent failure that no
ruleset rebuild can reach.
* RoutingPresent() reported "plane intact" from the ip RULE alone, while
ApplyRouting installs a rule AND a `local default dev lo` route removed
by two independent commands. A teardown interrupted between them was
therefore invisible, applyLocked's fast-path skipped ApplyRouting
forever, and no reconcile could repair it.
Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):
* init: `start_service` waits for a live predecessor pidfile to clear
before opening the instance, so restart == stop + pause + start. Zero
cost at boot. term_timeout 10 -> 30 so an honest teardown is never
killed halfway.
* daemon: the single-owner guard WAITS for the predecessor (bounded,
60s) instead of exiting 1; it still refuses if the budget expires.
* netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
a blanket flush would drop the admin's own SSH/LuCI sessions) called
on every plane transition: after a ruleset loads, after the table is
removed, and once more in applyLocked when the whole plane (table +
policy routing + sysctls) is assembled.
* netplane: RoutingPresent() now verifies both halves it installs.
Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The report claimed the order=20 `default` shadowed the order=100 one and sent all
unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute
does not emit a condition-less rule as a match-all route rule: it sets
route.Final and continues, so the LAST condition-less rule by order wins, and it
can never shadow a rule that has conditions (those are emitted ahead of Final
regardless of order).
For the config on the router this inverts the conclusion: traffic IS going
through the proxy (order=100 -> group:auto is the live default) and the dead knob
is the order=20 `direct` one. Severity downgraded from high to medium
accordingly — a dead setting, not a leak.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Routing page drew every condition-less rule as "default route · final",
so a config with two of them showed two identical claims and no hint that
only the last one is the default the router uses.
A superseded rule now loses those marks — it keeps its real Order in the rail
instead of the "·" that means final — and gains a "never applies" badge plus
a line naming the rule that beat it and what to do about it: give this one a
condition, or delete one of the two. Warn semantics throughout (--amber,
dashed frame, dimmed target chip): orange is the ACTIVE state on this
faceplate, and a rule the router ignores is the opposite of active.
Verdicts come from GET /api/rules/reachability and are keyed by the rule's
index in Rules, never by name — the config that prompted this had two rules
both called `default`. They are re-fetched after every save, and a verdict
whose echoed name/order no longer matches the row is dropped rather than
shown, so the window between an optimistic edit and the refetch cannot badge
a working rule.
Rule rows were also keyed by name in React, which silently collapses two rows
that share one; the key now carries the model index.
The mock fixture gains a second condition-less rule so `?mock` renders the
state, and mock.getRulesReachability derives its verdicts from the live
fixture config rather than hard-coding them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:
* two condition-less rules retire each other, and the LAST one by Order
wins, so an earlier "default -> direct" is dead while looking live;
* a condition-less rule can NEVER retire a rule that HAS conditions —
those are emitted ahead of Final whatever their Order.
A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.
model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.
Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.
Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.
Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.
GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.
Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.
ci/version.sh is now the single source of truth. It derives the version
from `git describe`:
tag `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
off-tag build -> nearest tag + PKG_RELEASE=<commits since it> + 1
no tag/no git -> 0.0.0-r1 (below everything ever published)
Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.
The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.
The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.
byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.
Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:
daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
[ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...
`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.
Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:
* control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
DisableColors at its false zero value. It now comes from
controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
lives in an untagged file (same split as profilewatch.go) so it is
unit-testable off the linux target.
* engine (shater/generate): the generated option.LogOptions never set
DisableColor, so box.New built a colouring formatter over the shared
sink. logOptions() now sets it from the same TTY gate (seam:
logColorAllowed).
logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.
The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.
Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as 9dc954029 / aec82d444.
`nodes` now reads model.ReadUCI() — `uci export shater` merged with the
per-subscription JSON caches — which is literally the call GET
/api/config serves and generate builds the engine from, so the verb
cannot drift from the panel or from the running engine: there is no
second assembly here to drift. Both ends use the same nodesJSON():
the daemon answers over the control socket (like `stats`), and the CLI
falls back to reading the same on-disk state when no daemon is running
(like `status`). A read failure goes to stderr with a non-zero exit
instead of printing `[]`, so an empty list on stdout now means one thing.
Output is a purpose-built view rather than raw model.Node: the share-link
URI is a credential and CLI output ends up in tickets and cron mail, so
the view reports what the link decodes to (protocol/server/port) plus the
model's own facts (enabled/sub/egress/stale/fingerprint). Nodes whose URI
does not parse are still listed, with the reason in `parse_error` — the
engine skips exactly those, and hiding them would be the same lie smaller.
cmdReadStub keeps `stats`, where the default IS the truth (nothing was
counted without an engine), and now says so in its doc comment.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the
previous install, install from the signed apk feed, verify the default state,
restore a working config with 315 subscription nodes, then exercise the data
plane, panel API, config lifecycle, resilience and DNS.
74 PASS. Findings (detailed separately): two catch-all `default` rules where the
first sends all unspecific traffic direct and makes the second unreachable;
`shaterd nodes` is a stub returning [] while usage promises the node list; DNS to
the router LAN address dies after `service shater restart` (stop+pause+start is
fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI
colour codes reach syslog.
Also records the four-iteration CI hunt that ended in the green apk lane.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.
The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:
config PACKAGE_kmod-mlx5-core
tristate
default m
No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.
The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.
Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.
They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.
Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
- an explicit value in .config beats a `default` (same rule that kept our
`# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
- `select` is OR-ed in after the user value, so shater-core's
`DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
transitive kmods) back on their own.
Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.
Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.
The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:
config ALL_NONSHARED ... default ALL
config ALL_KMODS ... default ALL
config ALL ... default y
In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.
Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.
Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.
Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was
running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn —
none of which we ship) before it ever got near our four.
Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the
.config that ships inside the ImmortalWrt SDK tarball. That file is the
buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus
CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so
`make defconfig` re-selected every kernel module of the target as =m and
package/kernel/linux/compile — pulled in via shater-core's nft kmod
deps — packed the lot.
Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds":
start the .config EMPTY so kconfig can only pull in what our packages
actually select. Carried over from the SDK's .config, nothing more:
the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a
wrong guess here means silently cross-compiling for another arch),
CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and
CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one
makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel).
Also adds the diagnostics this lane never had, since a failed run leaves
a 27 MB log: the carried-over identity lines, the post-defconfig kmod
count and target readout, a hard check that all four of our packages
survived defconfig, an abort if the kmod count is back in the hundreds,
and du/df after compile.
opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR,
CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in
the signed Packages index can't mask an empty SHA256. Script survives -u
(all vars use :? or :- defaults).
- .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the
copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a
trap-based tmpdir cleanup, and set -euo pipefail.
bash -n clean on both.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Backend audit fixes (upstream-file edits wrapped in // lx: markers):
- experimental/libbox oom_report.go/report.go: OOM reports + configuration.json
(server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files
→ 0o700 / 0o600. [sec-perms]
- daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret
compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare.
[sec-consttime]
- service/oomkiller/timer.go: network-extension cleanupTriggered logic was
inverted, so FreeOSMemory was never called after a trigger; flip both
assignments so a trigger schedules the deferred free and the next poll runs +
clears it. [sec-oomcleanup]
- transport/v2rayxhttp/client.go (lx-native file): session id used math/rand →
crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface.
- daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a
goroutine + the ssh-agent fd on every closed session (second io.Copy blocked
on an idle agent Read forever); tie both copies + the session ctx to a
cancel that closes both ends. [sec-sshagent]
- daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to
1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate]
- route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in
select while stopIdleSuspend niled it after close (race + goroutine leak on
Close-during-tick); pass the stop channel to the loop by value.
go build ./... (default) and the D9 shaterd linux build (tags
with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,
with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer
warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/...
./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend
and v2rayxhttp with with_xhttp.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- README.md: new Russian product README (what/features/architecture
mermaid/install both feeds/build/repo layout/CI/upstream/docs/license)
- README.en.md: concise English mirror (root readme was previously English)
- README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme,
a competing Russian README) -> points to README.md + engine-fork docs
- docs-shater/README.md: folder index
Install commands copied verbatim from docs-shater/INSTALL.md; all links
verified against existing files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Remove untracked-quality artifacts accidentally committed during work
sessions (all authored downstream, unreferenced anywhere in code/docs/CI):
- 5 session screenshots in repo root (devices-after-copy-fix.png,
live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png)
- tmp/gen_linux_test (29 MB throwaway traffic-gen binary)
Guard against repeats: ignore /*.png (root screenshots) and /tmp/.
Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are
left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- drop c/Users/.../gen_linux_test (28MB binary accidentally committed in 129e31fbd)
- .gitignore: ignore .idea/ at any depth (shater/.idea from IDE)
- CLAUDE.md: orchestrator delegates to model fable
- bump shaterd/shater-core r2->r3, luci-app-shater r1->r2 for v0.2.1 release
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
release-apk now runs with if: !cancelled() so an unrelated arch build
failure (e.g. x86_64) does not block publishing the aarch64 apk feed.
download-artifact only fetches existing artifacts and the publish loop
already skips missing apkfeed-* dirs.
The filtered-query-log block (SegMeter + top blocked + live QueryLog) is
redundant with Insights. The stats poll stays - it still feeds the DNS
filtering and Groups modules.
The subscription form now carries exactly: name, URL, update interval,
fetch via (+detour when proxied), User-Agent, HWID, and extra headers.
Format, device identity, regex/proto/country filters, dedup and expiry-alert
knobs are gone from the form (still honoured from UCI; a save carries them
through untouched). Headers are edited as key-value rows and serialize to
the existing `Headers: []string` "Key: value" contract. Name is editable:
a rename rewrites FromSub on the sub's cached nodes and refuses collisions.
Rule.Kill ""/"default" used to drop the rule, letting its traffic fall
through to the broader rules below and finally the default route - a silent
leak of exactly the traffic the operator singled out. ruleKillFallback now
always returns an outbound: ""/"default"/"closed"/unrecognised block the
rule's traffic in place; only an explicit kill=open goes direct. The default
route exists solely for traffic no rule matched.
One physical device with several addresses (v4+v6, multiple leases) used to
show as several devices. Discover now folds addresses sharing a MAC into a
single row: new `ips` field lists every address primary-first, `ip` stays
the primary (most recent lease), state is the best among addresses.
MAC-less hosts remain one-per-IP. The panel shows the extra addresses as
secondary chips; naming keys the config entry by MAC whenever it is known.
The engine resolves domains for itself (node server names, urltest probes,
subscription/DoH fetches). Those queries carried an invalid client address
and still landed in every insights surface. Gate them out at the single
ingestion point (Aggregator.handleEvent): an event with an invalid or
loopback client is dropped before totals, top domains, per-server counts,
the timeline, and the query-log rings. Only LAN-client traffic is collected.
The built-in block-ads / ru-bypass / private rule bundles are gone:
model.Preset, Model.Presets, the `config preset` UCI section, its render,
the panel Preset type, and every fixture. The generate-side expansion was
already removed with the profile rewrite in the previous commit.
A profile is now a pure uplink-conditional rule switch: Name/Enabled/
Priority/MatchIface/Enable-DisableRules/EndpointResolver. The per-profile
DefaultTarget/DefaultEgress overrides and the profile-level schedule window
(SchedDays/SchedStart/SchedEnd/SchedUTCOffset) are removed from the model,
UCI parse/render, the generator, the WAN watcher, and the panel. Rule-level
scheduling is untouched.
The panel's uplink condition is now picked from a dropdown of the router's
UCI interfaces (GET /api/interfaces, same source as the egress picker);
stored interfaces missing from the live list render as stale chips.
generate/profile.go is rewritten here (applyProfilesAndPresets ->
applyProfiles), which also drops the generate-side preset-pack expansion;
the preset model/UCI/panel surface is removed in the next commit.
Audit of the run-51 logs showed actions/cache@v3.3.2 works on the act_runner
(cold: "Cache saved" x4; next job: "Cache restored" in ~2s, npm --fast skip,
usign/dl reused) and the sdk-cache mirror seeds correctly — but the single
biggest recurring cost was NOT cached: `scripts/feeds update -a` re-cloned
base+packages+luci+routing+telephony every run (~7.8 min warm x 4 SDK jobs on
the serial runner ≈ ~28 min/run wasted; github ~1 MB/s from this host).
Cache .cache/feeds/{opkg,apk} (workspace dir, actions/cache-persisted, visible
in the SDK container via --volumes-from) symlinked over the SDK's empty feeds/:
`feeds update` now git-fetches deltas (seconds) instead of full clones, always
checking out feeds.conf's pins. Fail-safe: any error on the cached checkouts
wipes the cache and clones fresh. Key by SDK release (feeds-opkg-24.10.4 /
feeds-apk-25.12.1) — stable across runs, invalidates on an SDK bump; both arch
jobs of a lane share one entry (identical pins, serial runner).
Steady-state warm run: ~60+ min -> ~20-22 min. Also documented in the workflow
header: never key a cache on github.sha — each cache SAVE stalls the act_runner
~3 min, so per-run-changing keys would add +3 min/entry every run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Builds were dominated by re-fetching the ImmortalWrt 25.12 SDK tarball
(~300 MB) every run, and a stalled downloads.immortalwrt.org transfer wedged
the apk job for 40+ min (plain `wget -q`, no timeout — same class as the
elfutils hang).
- New ci/fetch-sdk.sh (runner-side): cache -> our durable `sdk-cache` release
mirror -> upstream with a stall-kill (curl --speed-limit 64K --speed-time 60
--max-time 1800) + 3 retries + zstd-magic/size validation; seeds the mirror
best-effort (github.token, non-fatal) so cold runs never touch upstream again.
A 40-min hang is now impossible; the in-container fallback wget also gets
--timeout=60 --tries=3.
- actions/cache@v3.3.2 (last release on the OLD cache API that Gitea act_runner
implements; v4/v3.4.x use the new GitHub cache service) for: SDK tarball, SDK
dl/ sources (hash of package Makefiles; PKG_HASH re-verified so a stale cache
can't leak a wrong source), Go mod+build (go.sum), npm node_modules
(package-lock.json) with build-shaterd.sh --fast, apt archives, built usign.
Degrades safely if the cache server is off — the SDK mirror is independent.
- concurrency group release-${github.ref} cancel-in-progress so a re-dispatch
cancels the stale run instead of piling up (tags stay isolated).
Signing (usign/apk), both keys, per-arch publish, manual triggers, LOCALMIRROR
and the scoped 4-package collection are unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
byedpi (ciadpi) ships as a separate optional package; the panel offered the
`byedpi` egress type regardless, so selecting it without the package installed
created a dead, fail-closed egress. Now GET /api/status reports
`byedpi_installed` (exec.LookPath("ciadpi"), os.Stat fallback), and the egress
type picker disables the ByeDPI option with a hint when it's absent. Existing
byedpi egresses are never hidden or rewritten (config is sacred) — shown with an
amber warning and still round-trip on save; only NEW selection is blocked.
Unknown status (older daemon / fetch fail) => no gating.
Bump shaterd PKG_RELEASE 1 -> 2 (the SPA is embedded in the daemon binary).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On a live `apk add` / `opkg install`, shater-core's post-install hung forever
(observed on BananaWRT 25.12 at "Executing shater-core...post-install", child
`flock 1000` in locks_lock_inode_wait). Root cause: a USE_PROCD init sources
/lib/functions/procd.sh on every rc.common action, whose procd_lock takes a
BLOCKING exclusive flock on /var/lock/procd_<svc>.lock held until the process
exits. shater-cron re-execs itself as the eternal `loop`, so it held that lock
forever; base-files' default_postinst then ran `/etc/init.d/shater-cron enable`
synchronously inside the transaction, blocking on the flock while the package
manager waited on the postinst — a permanent deadlock.
Fix (two layers):
- shater-cron `loop()`: `exec 1000>&-` closes fd 1000 up front so the eternal
loop never holds the rc.common flock (no-op when procd_lock is absent).
- 30_shater-core: defer enable/restart into a detached (setsid + bounded)
background block that waits for apk/opkg to finish before touching init.d,
with all fds to /dev/null (a held stdout pipe would hang apk on EOF too).
Bump PKG_RELEASE 1 -> 2 so existing installs pick up the fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk lane runs the SDK on a bare debian:bookworm host, and the ImmortalWrt
25.12 SDK prerequisite check requires python3-distutils ("Checking
'python3-distutils'... failed. Prerequisite check failed." ->
.prereq-build Error 1), aborting before any package built. The opkg lane was
unaffected because the openwrt/sdk image ships the prereqs. Add
python3-distutils (and python3-setuptools defensively) to the host deps. apk
lane only; opkg untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two build-harness bugs surfaced once the SDK builds actually ran:
1. Permission denied writing the feed. ci/build-feed.sh creates $OUT as root on
the runner, but the openwrt/sdk container runs as the unprivileged `buildbot`
(uid 1000) — so `cp` of the .ipk into $OUT failed ("Permission denied"),
yielding 0 packages and then "usign signing failed" (nothing to sign). Set
`chmod 0777 "$OUT"` on the runner before docker run (a chmod from inside the
container, as buildbot, cannot fix a root-owned dir). The apk lane already
chmods $OUT from its root debian container, so it was unaffected.
2. Collecting the whole SDK. ci/sdk-build.sh did `find bin -name '*.ipk'`, which
swept up the hundreds of prebuilt kmod/base .ipk shipped in the SDK image —
bloating the feed and signing foreign kmods under our key. Collect strictly
our four by name (`<pkg>_*.ipk`) and require >=4. Applied the same narrowing
to ci/sdk-build-apk.sh (apk names carry no arch: `<pkg>-*.apk`), keeping the
"wrong SDK produced only .ipk" guard.
No change to the feed format/signing (usign/KEY_BUILD/shater-feed.pub, apk EC
key), the package set, or triggers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk (and opkg) SDK builds intermittently hung fetching build-time
sources like elfutils-0.192.tar.bz2 from sourceware.org: curl's
--connect-timeout covers only the TCP handshake, not a stalled mid-transfer,
so a slow upstream hangs the whole job (no --max-time in OpenWrt download.mk).
Set CONFIG_LOCALMIRROR=https://sources.cdn.openwrt.org in .config before
`make defconfig` in both ci/sdk-build-apk.sh and ci/sdk-build.sh so the SDK
tries the fast OpenWrt source CDN before each package's own PKG_SOURCE_URL —
fixes elfutils and any other flaky upstream. Mirror verified to hold the file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Same root cause as the apk jobs: scripts/build-shaterd.sh builds through a
go.mod `replace => ./submodules/wireguard-go` (AmneziaWG fork, bumped in
16a47b596), and actions/checkout does not fetch submodules by default, so
`go build` died with "reading submodules/wireguard-go/go.mod: no such file or
directory" in the opkg build jobs (x86_64 + aarch64_cortex-a53) as well. Init
only that one submodule — build-harness only, no change to the opkg feed
format/signing (usign/KEY_BUILD/shater-feed.pub) or package set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/build-shaterd.sh builds via a go.mod `replace => ./submodules/
wireguard-go` (the AmneziaWG-patched fork), so that submodule must exist or
`go build` dies with "reading submodules/wireguard-go/go.mod: no such file or
directory". actions/checkout does not fetch submodules by default. Init only
that one submodule (public GitHub URL; clients/apple+android are large and
unused) in the additive build-apk jobs — the opkg build jobs are left untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Additive next to the opkg/24.10 lane — nothing existing changed. The same 4
packages (shaterd, shater-core, luci-app-shater, byedpi) are built through the
official ImmortalWrt 25.12 apk-SDK and published as per-arch rolling releases
apk-latest-<arch> / apk-<tag>-<arch> (x86_64, aarch64_cortex-a53).
- ci/sdk-build-apk.sh: drives the 25.12 SDK inside debian:bookworm, compiles
.apk, then `apk mkndx --root T --keys-dir T/keys --allow-untrusted
--sign KEY --output packages.adb *.apk` — the exact form the OpenWrt 25.12
buildsystem uses (unsigned members, signed index).
- ci/build-feed-apk.sh: per-arch runner entrypoint (same --volumes-from and
artifact-order contract as ci/build-feed.sh).
- ci/gen-apk-key.sh: one-shot EC (prime256v1) keypair generator; private half
-> Gitea secret KEY_APK, public dist/shater-apk.pem committed.
- release.yml: additive build-apk / release-apk jobs; `on:` triggers untouched
(v* tags + workflow_dispatch); apk release tags deliberately non-`v*`.
- docs-shater/INSTALL.md section 6, .gitignore (out-apk/), dist/shater-apk.pem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The three dependent plaques (Keep file on flash / File size cap /
Download) only had their inner controls disabled — the rows still
looked live and hoverable. Field gains a disabled prop rendering
set-field--off: pointer-events none, opacity .45, grayscale, flattened
background — the whole plaque reads and behaves as switched off
(aria-disabled included). Wired to !logToFile on all three.
Verified live on the testbed: with the toggle off all three plaques are
inert and dimmed; flipping it back on restores them.
embed.FS carries no timestamps, so SPA responses went out with neither
Last-Modified nor ETag and browsers fell back to HEURISTIC caching — a
stale index.html kept showing the previous panel after a daemon upgrade
(user saw pre-c61cfe3a download buttons enabled with the toggle off).
index.html / SPA fallback / favicon / 404s => Cache-Control: no-cache
(revalidate every load); a HIT under assets/ (content-hashed by Vite)
=> public, max-age=31536000, immutable. Guarded by TestStaticCacheHeaders.
Verified live on the testbed: / and /settings no-cache, hashed asset
immutable, missing asset 404 no-cache; a plain reload now picks up the
new SPA.
Operator decision (supersedes ad9781bf): "Log file (downloadable)" off
must leave NO trace — delete the saved log files outright, and gray the
download buttons out while the file is off.
logsink: New and Reconfigure purge the active segment and the rotated
.1 whenever ToFile is off — at the old and new configured locations AND
both standard paths (a Persist flip must not leave a stale copy). A
daemon booting with the toggle off sweeps leftovers from a previous
life too.
panel: /api/log reverts to the pre-ad9781bf precedence (toggle off =>
syslog scrape / '# logging disabled'; segments are never served while
the file is off, even if a leftover exists). SPA: the three download
buttons are disabled when LogToFile is off; note/flash texts and the
?mock fixture say the files were deleted.
Verified on the docker-OpenWrt testbed via the panel: off+apply deletes
/var/log/shaterd.log* (and /etc/shater), buttons gray out; on+apply
starts a fresh file and downloads work again.
Flipping "Log file (downloadable)" off looked like it deleted the logs:
the file stayed on disk, but GET /api/log switched to the logread scrape
and the collected history became undownloadable (user report). The
toggle stops WRITING — it must not disown what was already collected.
New precedence: retained segments are streamed whenever they exist,
prefixed with a '# note: file logging is off …' line when the toggle is
off (even with syslog off too); the syslog-scrape and '# logging
disabled' fallbacks now speak only when nothing is retained. Settings
note/flash texts and the ?mock fixture updated to match.
Verified on the docker-OpenWrt testbed: with log_file=0 the download
returns the note + full history; re-enabling via the panel resumes
appending to the same file with nothing lost.
modernc.org/sqlite is the only pure-Go SQLite and costs ~3.5 MB in the
static shaterd link; the stats store never used anything SQL-specific —
it is a ring of two append-only streams with a monotonic seq cursor.
bbolt is already linked via experimental/cachefile, so the swap is free.
sqlitering.go -> boltring.go: buckets queries/conns keyed by 8-byte
big-endian seq (bbolt key order == cursor order), rows as JSON of the
existing LogEntry/ConnLogEntry structs, meta bucket carries the durable
per-stream HWM (same max-only monotonic semantics). The async writer
contract is untouched (writeCh 4096, drop counters, 256-row/500ms
batches, 30s retention tick). Disk cap: chunked oldest-first deletes
with the same hysteresis, then at most one bbolt Compact per pass
(sagernet/bbolt exports Compact) behind the same 110%+1MiB free-space
guard that gated VACUUM. A legacy SQLite-format stats.db (or any
unreadable file) is replaced in place with one warning; open failure
still falls back to the in-memory ring.
Zero user-visible change: the "sqlite" backend selector value and the
Snapshot.Backend string are kept verbatim. Tests ported assert-for-
assert plus new coverage: legacy-file replacement, overflow drops,
memRing parity round-trip, disk-cap convergence.
Router shaterd (linux/amd64): 28,004,478 -> 24,428,670 bytes (-3.58 MB);
modernc.org/* gone from go.mod/go.sum and the dep graph.
shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS
transport is never generated, and the slim shater/registry never
registers the transport, so the tag gated nothing in this binary.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
Upstream defect: acme.go is behind with_acme but acme_logger.go was not,
so go.uber.org/zap linked into every build even with ACME disabled. Only
acme.go references ACMELogWriter/ACMEEncoderConfig, so the twin gate is
behaviour-preserving; a with_acme build still compiles.
Marked lx:acme_logger_gate; upstream-PR candidate (drop the lx block on
rebase once merged). -94 KB on the router shaterd link.
The admin panel is shater's own web server and generate never emits a
clash_api service (shater/engine/engine.go pre-registers its own
dnstrack.Manager precisely because no api/clash_api observer exists on
the router). With include.Context gone the Clash server was already out
of the link; dropping the tag records the decision. Desktop/CLI LX_TAGS
keeps with_clash_api for external dashboards.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
The shater data plane is tproxy/redirect (netplane); generate never emits
a tun inbound, so the userspace gvisor netstack is unreachable code. With
the slim registry it was already dead-code eliminated by the linker —
dropping the tag makes the intent explicit and stops compiling ~3.6 MB of
gvisor sources into the build at all. A future tun inbound would fall
back to the system stack; re-add the tag if that ever lands.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
include.Context registers upstream's entire zoo — tor (bine), ssh, snell,
anytls, naive, masque, mdns/resolved, and the api service whose daemon
bridge links grpc+protobuf — none of which shater/generate ever emits.
shater/registry registers exactly what the generator can produce (tproxy/
redirect/direct/socks/http/mixed inbounds; direct/block/selector/urltest/
socks/http/ss/vmess/trojan/vless/shadowtls outbounds + hysteria2/tuic
behind with_quic; wireguard endpoint behind with_wireguard; tcp/udp/tls/
https/hosts/local/fakeip + DoQ/DoH3 DNS transports; xhttp + v2rayquic
transport blank imports), with build-tag stub twins so a tag-less
'go build ./...' stays green. Zero upstream diff.
Measured on linux/amd64 with the D9 router tag set: 47.05 MB -> 31.07 MB
raw (-34%); the unreachable gvisor stack and grpc/protobuf are dead-code
eliminated even before any tag changes. UPX --lzma artifact: 12.49 MB ->
~8.8 MB. Since a UPX-packed binary unpacks fully into anonymous pages,
the same ~16 MB comes off resident RAM on the router.
New "Daemon log" group on Settings (Faceplate): the LogLevel verbosity
select (relocated, honest note — "none" is a turn-down to panic-only, not
a true off; failures still alert), LogToFile / LogToSyslog / LogPersist
toggles, a validated LogMaxKB editor (128–8192), and three download
buttons (day / 3 days / everything) → downloadLog() fetches
GET /api/log?range=… with the session cookie, filename from
Content-Disposition, blob save. Honest warn plates: file-off = only a
slice of the syslog ring (ranges approximate); both-off = nothing is
written anywhere; flash vs tmpfs (lost on reboot, wears flash, ~33 MB
budget). Globals type gains LogToSyslog/LogToFile/LogPersist/LogMaxKB
1:1 with the backend; mock.ts mirrors the honesty contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The daemon's own log (engine + control-plane) went only to os.Stderr →
procd → the logread RAM ring: no file, no wall-clock timestamps, no size
cap, and "LogLevel=none" silenced ONLY the engine while the control-plane
kept writing at trace. So "download last day/3d/all", "limit the size" and
"fully turn it off" were all unmet.
New shater/logsink: one long-lived, atomically-reconfigurable Sink that
receives BOTH halves' byte streams, stamps every complete line with a UTC
RFC3339 wall clock (what makes date ranges real), and fans each line to a
size-capped 2-segment rotated file (ToFile) and/or the real os.Stderr
(ToSyslog). Both off = the line is dropped — the only true full silence.
Persistent path sits behind a stats-style disk-free guard (suspend+warn
once, auto-resume); tmpfs path is bounded by the cap itself. ANSI stripped
from the file copy only.
Wiring: control-plane via log.SetStdLogger over the sink; engine via a new
box.Options.DefaultLogWriter threaded into all three box.New sites
(apply/close-then-start/restore) by engine.SetDefaultLogWriter; live
reconfigure on every apply.Reconcile (SIGHUP / control socket / panel
apply) so panel changes take effect without a daemon restart.
controlLogLevel now makes the control-plane respect Globals.LogLevel
(silent vocab → panic-only; unknown → warn, mirroring generate).
Globals: LogToSyslog/LogToFile (default true), LogPersist (default false =
/var/log tmpfs; true = /etc/shater flash), LogMaxKB (default 2048, clamped
[128,8192]; 0 = default, not off — LogToFile is the off switch). UCI
parse/render/aliases + ValidateGlobals clamp-warn.
Endpoint GET /api/log?range=1d|3d|all (session-gated): streams the log line
by line, oldest segment first, filtered by the timestamp prefix; UTC
attachment filename. Honest fallbacks — file off + syslog on → a
"# note: … syslog ring only, ranges approximate" comment then a
`logread -e shater` scrape; both off → "# logging disabled". Unknown range
→ 400.
init.d: shater/shater-cron gate their `logger -t` status lines on
log_syslog so "logread off" is honest at the shell layer too.
Tests: logsink rotation-cap/timestamp/toggle-gating/engine→sink,
model round-trip + validate, endpoint session-gate/range/fallbacks.
VM-verified on QEMU (x86_64, OpenWrt 24.10): download+ranges, size-cap
rotation, file-off/full-off, persistent path, live reconfigure — all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A rule disabled in UCI but force-enabled by the active profile, with an
iface:/zone: source outside the tproxy-inbound set, got its engine route
rule but no nft divert — its traffic never entered the engine, and the
fail-closed forward drop and accept_local sysctls skipped the device too.
Root cause: generate applied profile enable/disable in its own
effectiveRules while netplane read raw Rule.Enabled. Fixed with one shared
resolver in the leaf model package (ResolveActiveProfile +
ApplyProfileRuleOverrides) that both the engine route plan and the nft
divert plan consult, so they can never disagree about which rules are in
force. applyLocked now threads a single now through generate + nft render +
sysctls, closing the schedule-boundary race between the two planes.
Verified: a profile-enabled iface rule now joins the divert set, the
per-rule tproxy emit, the fail-closed drop and the accept_local sysctl;
the inverse (profile-disabled) drops the device. Parity regression on the
existing generate profile tests stays green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The schedule evaluator called time.LoadLocation, but the router binary
embeds no tzdata and OpenWrt ships none — so LoadLocation always failed
and windows silently ran in UTC while the panel promised local time.
- Windows now anchor to SchedUTCOffset (minutes east of UTC), which the
panel captures from the editing browser on every schedule save; the
daemon evaluates now.UTC()+offset with no location database. This
sidesteps the weekly-recurring day-shift that a full local<->UTC
conversion cannot express in one window. SchedTZ is deleted (documented
in the removed-options list; old configs parse and drain it). DST is a
stated limitation (followed on re-save). generate/schedule.go collapses
from a second copy of the evaluator to a thin adapter over the model one.
- The iface-profile schedule was honored by the WAN watcher since
08d5d6cc, but generate warned "the watcher does not look at the schedule"
and the panel muted the editor with "the router ignores the schedule" —
both false. Warning and lie removed; the editor is live and labelled
"applies together with the uplink match".
- Stale fictions: FEATURES.md nftset/FakeIP-mode MVP line and the shipped
conffile's dead `option dns_mode 'nftset'` corrected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- The generator iterated dns_rules in raw slice order and never read
DNSRule.Order; first-match "top to bottom" was true only because the
panel pre-sorts. Now sorted by (Order, index) like route rules, so a
hand-edited UCI or any API client gets the declared order.
- dns_rule match_src silently dropped zone:/iface:/MAC entries (the
in-engine DNS plane matches source IPs only), which could widen a rule
to ALL clients or skip it entirely. Each dropped entry now warns, with
the consequence spelled out.
- BlockDoH :443 IP list was incomplete (no NextDNS anycast, no actual
cloudflare-dns.com 104.16.x, sparse v6). Extended across all listed
providers, now accepts anycast CIDRs, with a maintenance note that the
list is manual. The hostname NXDOMAIN + canary layers already cover
resolve-by-name; UI still says "well-known providers only".
- Panel: intercept-OFF copy no longer overstates the bypass (plaintext to
external resolvers is already hijacked by the D14 catch-all); allowlist
note gains the per-device-Block-wins caveat.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit found the Insights numbers were real but mislabelled:
- "Blocked" counted every NXDOMAIN, upstream timeout and zero-answer as a
block. Now blocked = strictly the engine's own filter verdict
(dnstrack.SourceFiltered: D15 blocklist + BlockDoH predefined-NXDOMAIN).
Failures (timeout/SERVFAIL-reject) become their own `failed` category;
the three counters are mutually exclusive and sum to Queries. The DNS
log "block" tag follows the same signal.
- Per-minute sparkline positioned buckets evenly by index over a sparse
slice, so "60 min" could span hours. Now points sit at their real
Bucket.Minute, gaps render as gaps, and the label states the actual
span + active-minute count instead of a fictional "last N min".
- Per-device domains skipped the LAN filter every other view applies, so
the router's own urltest/sub-fetch dials appeared as a phantom WAN-IP
device. Now folded into the `router` pseudo-device like the DNS log.
- Honest labels: "Outbounds/exits" -> "DNS lookups per exit"; top
domains/hosts meta "N tracked" -> "top N shown". Overview query log
shows the real per-device attribution, not the resolver tag; stale
"DNS events have no client IP" comments removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The whitelist accepted random but the human-readable "Supported:" tail
still named only four strategies — caught live on the VM where the model
and generate warnings disagreed about the supported set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- node_down is gone from AlertEventNames: nothing ever fired it, so a
channel subscribed to it was silence dressed as monitoring. An old
config's `list event 'node_down'` now warns as an unknown event and is
dropped. The accepted and emitted sets now coincide; the reserved-event
branch of ValidateAlerts stays as the guard against future divergence.
- model.ValidateGroups + KnownGroupStrategies: a typo'd strategy is
warned at validation time (was: silently built as least_test with only
a generate-time warning). Mirrors generate's warnGroupStrategy list.
- doc-comment honesty: random is a real engine mode (api.ts), sqlite
stats backend is a real persistent store (api.ts + model.go), resolver
type list gains tcp, pages/index.ts no longer claims Placeholder pages,
failover doc says fail-back exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New balancer flag priority (option.URLTestBalancerOptions.Priority): the
pool is re-derived from CONFIG ORDER every health-check tick via
balancePoolPriority/planPriorityPool — the first live member owns slot 0,
so when the top node answers probes again traffic returns to it on the
next tick (30s failover interval). Probing walks top-down and stops at
the first live node, so the steady-state cost stays one probe per tick.
Replace-in-slot deliberately does not apply here: failover forces sticky
["none"], so relocating nodes across slots breaks no flow keys. Plain
round_robin/random paths are untouched.
failoverBalancer() now emits Priority:true; the KNOWN LIMITATION note and
the panel's "nothing brings it back" blurb are gone because the
limitation is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- random is a REAL urltest mode (lx SPEC 019 v2): uniform draw over LIVE
slots only, pool sized to every member; dead slots keep their place
(never-shrink) but are never picked, for random AND round_robin AND
sticky (degrade-to-live). All-dead pools fall back to Select.
- Globals.SweepInterval + Globals.GroupHealth master switch, resolved by
one pure function (model.SweepSchedule) shared by validator and apply;
unparseable is warned-and-ON, never silently off. ConfigureSweep no
longer resets the cursor on every cron reconcile (release blocker:
a ~6-min cycle was restarted every 60s and never completed).
- multi-WAN egress gateway: ubus netifd status -> uci static -> main
table; a gatewayless non-P2P egress warns CRITICAL instead of silently
blackholing the second uplink.
- endpoint resolver (route.default_domain_resolver): bootstrap-direct
clone of a named resolver, profile override beats globals.
- chains are composable: chain: hops flatten recursively, cycle-guarded,
entry egress lifts only at position 0 (fail-closed mid-path).
- group test publishes its scope so "measuring" lights only the cards a
run covers; health run is explicitly global (all_nodes).
- panel: biased-sample honesty (no ratio until a failure CAN be on
record), profiles auto-pin plate, sweep/GroupHealth settings UI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two threads, both from the same question: does this setting do what it says?
## Health is per-group, because a dial path is per-group
Overview reported "119 up / 179 untested" over all nodes, and Nodes showed a
per-row ping. Both measured the wrong object. A group with an egress binding does
not dial the base node outbound at all — generate materialises per-member copies
(group-<name>-m<i>-<member>) and the group balances over those. So a node can be
alive direct and dead through the tunnel a group is bound to, and the panel said
"up". The same node in two groups with different egresses is two states that were
being collapsed into one number.
No new prober was needed: the engine already keeps a process-wide
urltest.HistoryStorage keyed by outbound tag, groups already probe their own
members into it, and stats already reads it — we simply projected it onto base
tags only. The copy tag carries the member NAME, so recovery needs no change to
generate. GET /api/groups/health now reports alive/dead/untested per group, with
an opt-in member list; the same summaries ride /api/stats so Overview needs no
extra poll.
Presentation is "alive / tested" with the untested remainder as a quiet aside,
never folded into dead: groups probe lazily and only while in use, so on a fresh
boot with a 376-node subscription almost everything is legitimately unmeasured,
and calling that "down" would scream catastrophe exactly when nothing is wrong.
Three things this exposed, all fixed here:
- TestAllNodes enumerated only om.Outbounds(), which by design excludes
endpoints. Every WireGuard/AmneziaWG node read "untested" forever no matter how
often the button was pressed — on a product whose driving requirement is AWG.
- Our ProbeFailDelay sentinel is gone from the engine's history entirely. It was
safe for least_test (slowest wins last) but round_robin's pool planner treats
any entry as alive, so a dead node could occupy the single slot of a failover
group — pinning failover to a corpse, which is the one thing it exists to
prevent. Failures now live in an engine-side overlay, invalidated by timestamp
against any later success; the engine's history holds measurements only.
- Writers now measure with the probe URL of the group that owns the tag. A manual
run used the global URL and overwrote a group's own measurement, leaving
least_test comparing latencies to different servers. Where one tag is claimed by
two groups with different URLs the ambiguity is inherent to the engine's keying,
so we use the neutral global URL and say so rather than picking a silent winner.
A scheduled sweep (engine/sweep.go, on by default) fills what nobody probes:
24 measurements per 10s tick, 12 in flight, skipping anything fresher than 5
minutes — ~5 min per full cycle on the production config. The freshness gate is
load-bearing beyond cost: testNodes skips a member whose history is younger than
the group's interval, so a sweep that kept refreshing would starve a failover
group's own 30s check. It is a layer under group-local probing, never a
replacement.
## Options that did not exist are deleted, not decorated
Audited every enumerated choice the panel offers against what this fork actually
implements (constant/, option/, protocol/group/, dns/), and split the results into
works / synonym / fiction. Fictions are removed outright — pre-release, so no
legacy path is kept for values nobody has.
Deleted: Group.Strategy random and leastload (both silently became least_test);
Egress.Type proxy and block (emitted no outbound at all — every binding dangled
and the traffic left over the plain WAN with the real IP); alert event node_down
(no emitter anywhere); Globals.DNSMode, Inbound.Sniff, Profile.ProbeURL/ProbeMode,
Node.XUDPConcurrency/XUDPProxyUDP443, Egress.Target.
Repaired instead of removed, because the engine could do them all along:
LogLevel "none" (asked for silence, got default verbosity — now LogOptions.Disabled);
Subscription.Format (the hint was stored, badge-rendered and ignored — the sniffing
parser always ran); Ruleset.Format (never read; the extension decided);
Group.Strategy failover (urltest + round_robin + pool 1 / tolerance 0 is exactly
"first working node in order" — verified through box.New with a sensitivity control).
Relabelled where the words lied: Single promised "first up" but a selector never
checks liveness; inbound "http" opens Mixed and answers SOCKS5 on the same port.
An unresolvable egress binding no longer fails open. It resolves to block, so the
bound traffic stops visibly instead of leaving with the real IP. Refusing the
config was the alternative and is worse: a dead engine under a closed kill-switch
blackholes the whole LAN over one mistyped name.
Also: Egress.Port no longer defaults to 1080 for every type. The parser invented
it, render persisted it, and the new "port is ignored" warning then fired on a
correctly written config — a warning on a healthy install is how a findings list
gets ignored.
## Geo data is no longer hardwired to one publisher
sing-geoip publishes country codes and nothing else — 238 files, all two-letter.
So "route Netflix around the tunnel" meant loading geoip-us: 159,125 prefixes and
~20 MB of kernel memory for something the netflix list does in 108 prefixes and
~14 KB. Provider selection is now a chain (generate/geosource.go): country codes
still resolve to SagerNet byte-identically, everything else to Loyalsoldier, and
metacubex adds AS<number> routing. Third-party .srs was verified to load with our
own reader (v1/v2 against our v5 ceiling) before any of this was built.
The ruleset preflight reads four header bytes over a ranged GET instead of HEAD,
so a rule-set whose format version we cannot parse degrades like an unreachable
one — that case would otherwise abort engine start, which is how the LAN goes down.
The status strip carried five pips — ENGINE active, UPTIME, CONFIG enabled,
DATA PLANE installed, KILL-SWITCH — and the user had to AND three of them
together to learn whether they were protected. `plane` and `engine_running`
already encode that, and more precisely than the booleans did. UPTIME duplicated
the Engine module's "running for"; KILL-SWITCH duplicated the module directly
below it. Collapsed to one derived line phrased in terms of traffic:
Protected — traffic from your network is going through the tunnel
Traffic blocked — the tunnel is down (hold, amber)
Not protected — traffic is going out directly (none + fail-closed, crit)
Not protected — running direct (none + fail-open, amber)
hold and none stay distinct: one is the kill-switch catching it, the other is
no safety net at all. Nothing was lost — every removed value still lives in the
module that owns it.
Findings are now routed by severity instead of all landing on the front page
(panel/src/findings.ts):
critical / warning -> Overview. Something needs attention.
info -> the page that owns the setting.
An info finding is a statement about the configuration: it never clears and asks
for nothing, so a permanent front-page entry only teaches people to skim the
list — which is how a real critical finding gets missed. The untunnelable note
now renders inside the Networks "Other traffic" section, beside the control it
describes. With nothing needing attention the section renders nothing at all.
Also fixed, found while auditing the rest of the labels: the Kill-switch module
read ARMED / policy: fail-closed with a green lamp even at plane=none — a
reassuring light directly beneath a readout saying nothing is protected. A
fail-closed setting is only armed if something is installed to enforce it, so it
now reads NOT IN EFFECT with a crit lamp and a "blocking now: no — nothing
installed" row; policy -> setting.
planeState.ts became the single source of the wording, and the plane banner was
dropped from Overview — it exists to carry the alarm to pages with no status
readout, and stacked under the new line it just said the same thing twice.
Verified against the live daemon on the bench: healthy, critical and hold states
all render correctly, console clean, note present on Networks and absent from
Overview.
.gitignore: MemPalace per-project files, added by the tooling.
Traffic TPROXY cannot carry (ICMP, IGMP, ESP/AH, GRE) was dropped for the whole
LAN regardless of routing. A box configured to tunnel only 8.8.8.8/32 still lost
ping to the entire internet, and with the shipped config RU addresses were
unpingable even though `ru-direct` sends them out unproxied — the very path where
TCP already exposes the real IP, so the drop prevented no leak at all.
The drop is now scoped to destinations the rules actually tunnel:
iifname "br-lan" meta l4proto != { tcp, udp } ip daddr @unt_d4_1 accept
iifname "br-lan" meta nfproto ipv4 drop
Destination sets come from the engine's already-parsed rule-sets via
ExtractIPSet(), so no .srs parsing and no second read of the bbolt cache the
engine holds locked. Rules are taken from the generated route rules, not the raw
model, so preset packs, WAN-profile overrides and schedules are all included.
Domain/geosite matchers are skipped when classifying: a packet with no stream
carries no domain, so such a rule can never apply to it.
Every policy line carries `l4proto != { tcp, udp }`, so no destination decision
can ever accept TCP/UDP — fail-closed is structurally untouched. Anything the
walk cannot prove direct (list not yet fetched, logical rule, unknown action,
inverted match) falls through to the drop and says so via an info finding.
No element cap: a continent-scale list loads in full. Measured on the bench with
geoip-us — 4s apply, 1.25 MB ruleset in 29.5k lines, ~27 MB RSS growth, engine
healthy. Cost is reported, not enforced; `untunnelable=direct` loads no sets.
Also fixed here, found while building it:
- plan warnings were computed and dropped, never reaching the operator; routing
them through the netplane channel was wrong (it marks everything critical by
construction), so they get their own info-level path
- nft ran with no timeout while holding the apply flock: one wedged invocation
would have deadlocked every later apply, reconcile and teardown. 60s cap; the
ruleset commits as a single netlink transaction, so killing it is safe
- set elements were emitted as one 3.1 MB line the lexer would hold as a single
token; now wrapped at 8 per line (identical to nft, readable when debugging)
- untunnelable copy still claimed ping never works; rewritten for the new
semantics across all three modes
panel: the theme switch read as a power toggle — it reused the component that
turns features on and off and sat inside the status cluster next to the ONLINE
lamp, so in light theme it looked like a switched-off appliance. Now a two-key
sun/moon selector, both states always visible (neither theme is an "off"), the
engaged key raised and lit by shading rather than accent colour, separated from
the indicators by a groove.
Verified on the OpenWrt bench: RU addresses ping, non-RU stay blocked, TCP routes
unchanged through the tunnel, DNS filtering and Block-DoH unaffected.
Group egress — for the case where the protocols themselves are DPI-blocked:
every node in the group dials ITS OWN server through the chosen egress (an
AmneziaWG tunnel, say), so the provider sees tunnel traffic instead of a VLESS
handshake. It binds the outgoing dial, not post-proxy traffic.
The binding is per-group, and that is the whole difficulty: group members are
SHARED outbounds, so two groups built from one subscription — one bound, one not
— would either leak the binding into the unbound group or fail to apply it. The
members of a bound group are therefore materialised as per-group copies
(group-<g>-m<i>-<member>), reusing the same rebuildNode the chain builder uses
for per-hop copies. Copies are made only when Egress is set, so an unbound group
over a 331-node subscription does not double the engine config. Copy tags are
checked against the node/group/egress/copy namespaces; a collision skips the
member with a warning rather than shadowing a real node. Precedence is chain hop
-> Node.Egress -> Group.Egress: a node pinned to a particular uplink was pinned
for a reason the group cannot know. A member whose copy cannot be built is
dropped rather than falling back to its unbound tag — falling back would leak
exactly the traffic the binding exists to hide.
Group test answers "what am I exiting through, and how fast": selected member,
latency, exit IP and country, via cloudflare.com/cdn-cgi/trace (country comes
free, so no GeoIP database on the router) with api.ipify.org as fallback. The
probe is pinned to the group's own outbound and refuses the direct outbound — a
direct answer would print the ISP's address and claim the tunnel works when it
does not. Measuring latency but failing to resolve the address stays ok=true
with an empty exit_ip; that is a working tunnel, not an error.
Uptime: /api/status gains started_unix + uptime_seconds, measured from process
start over a monotonic seam so an NTP step on an RTC-less router cannot be
reported as uptime. It is the daemon's uptime, not time since the last apply.
Also fixes: renaming an egress did not rewrite Group.Egress, silently dropping
the group back to the default route.
Verified on the testbed with the real 331-node subscription: two groups over one
subscription, one bound, one not — the bound group selected
group-auto-egress-m130-IE-trojan-141 while the unbound one selected the shared
IE-trojan-141, exit IP and country resolved for both, no group warnings.
Measured cost of binding a 331-node group: engine outbounds 335 -> 666, config
50 KB -> 114 KB, daemon RSS 62 MB -> 75 MB.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
I wrote the gap up as open while reviewing an agent report I had not yet seen;
the hosts/plain/AdBlock parse-and-compile path had in fact landed in the same
commit. Records the measurement that settles the disk question: StevenBlack's
2.4 MB of text compiles to 80873 domains in a 491 KB .srs, so it ships in the
production posture instead of being traded away for the 8 KB geosite list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
On a healthy router the warning set is reprinted by every reconcile — once a
minute from cron plus every hotplug event — so logread filled with the same
line forever and buried the warnings that matter (a blocklist that failed to
load, an interface the kill-switch does not cover, a missing data plane). On a
router logread is an in-memory ring buffer, so this also evicted the history
needed to investigate an incident.
Warnings are still returned in full by GET /api/status on every request; only
the logging is deduplicated, keyed on a fingerprint of the set. Message texts
carry volatile parts (free MiB on /overlay, compiled domain counts, the address
inside a network error), so digits are normalised for comparison only — the
logged and API-returned text is untouched. A restart reprints the full set, and
clearing the last warning logs one line saying so.
Also fixes the severity-to-syslog mapping: an [info] warning was being emitted
at WARN, so anyone filtering on WARN saw noise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Records the positions the audit changed: fail-closed must cover "engine never
started" (holding plane), engine start must not depend on the network (remote
rule-set preflight, with the deferred cache-seed fix noted), BlockDoH needs no
route-plane upstream exclusion (engine dials bypass route rules), DNSMode is
unimplementable and its control was removed, TPROXY's inability to carry
ICMP/IGMP/ESP/GRE is now an explicit 3-way policy, and fail-open degradations
must surface in the panel rather than only in logread.
Also flags the contradiction left open: D15 promises seeding StevenBlack/OISD/
AdGuard while blocklist source=url accepts only compiled .srs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A ground-up audit of the whole v0.2 stack by 8 parallel agents (DNS generate,
routing generate, model/parse/subscribe, netplane/apply/engine/alert, stats +
panel API, panel frontend, OpenWrt packaging), with every finding reproduced or
verified on the OpenWrt QEMU testbed. ~60 defects fixed, each with a regression
test that was checked to FAIL against the old behaviour.
RELEASE BLOCKERS
* Engine-start failure left the data plane ABSENT: with kill_switch=closed the
router silently degraded to a plain OpenWrt box — no tunnel, no filtering, no
kill-switch — while the panel looked healthy. Reproduced live. Now any
engine-start failure installs a fail-closed holding plane (forward blocked,
LAN-to-LAN and management preserved) and reports plane=hold/none.
* An unreachable remote rule-set aborted engine start entirely, so a router that
booted before its ISP link came up ended with a dead LAN and no way to recover.
Remote lists are now preflighted and skipped with a loud warning instead.
* `geosite:` in a routing rule hard-errored box.New — one legacy rule took the
whole LAN down. Same class: unvalidated CIDR / port / regexp, and marker-only
list entries ("." / "keyword:"). A lone `keyword:` also silently NXDOMAINed
the entire internet.
* Fail-closed drop only covered tproxy inbounds, not interfaces diverted by rule
sources — engine down leaked those networks to WAN in plaintext (4f618140 redux).
* UCI injection: a newline in a subscription-supplied node name broke out of the
line-oriented config and wrote attacker-controlled sections.
* Bootstrap deadlock: the daemon refused to start while disabled, but the panel
IS the daemon — a fresh install could never be configured from the UI.
SILENT FAILURES (the audit's main theme)
* per-device DNS block ignored the `suffix:` prefix — parental control that
quietly didn't block. Unknown `word:` prefixes now warn instead of vanishing.
* sqlite reused `seq` after retention wiped rows, stalling the live log forever.
* `after=` cursor returned the NEWEST rows, permanently skipping bursts.
* Stats emitted null arrays on a freshly booted router, blanking Overview.
* Alerts fired twice per incident; new_device alerts swallowed all but the first
device in a 60s window.
* Disabled subscription nodes were silently re-enabled on every refresh.
* Invalid Include/Exclude regexes failed OPEN, disabling the whole filter.
DEAD KNOBS — wired or honestly removed
ru-bypass preset (emitted an unsupported geoip: matcher) -> real geoip rule-set
Globals.ResolverFallback -> implemented via evaluate + match_response chain
Globals.DNSMode -> unimplementable by design; control removed, fake-IP
documented via a type=fakeip resolver instead
Rule.Kill -> implemented (default | closed | open)
Rule.Egress -> was read by nobody; multi-WAN binding silently no-op
ExpireAlertDays + quota -> subscription-userinfo parsed, persisted, alerted
StatsBackend hot-switch -> store is re-created on change
Inbound.Sniff -> documented as vestigial (sniffing is a route action)
NEW
* Globals.Untunnelable (block | icmp | direct): TPROXY can only carry TCP/UDP, so
ICMP/IGMP/ESP/GRE were dropped with no explanation — ping simply didn't work.
Now an explicit policy, defaulting to the previous behaviour, and explained in
the UI by consequence rather than by protocol.
* apply now surfaces its warnings through /api/status (severity/section/name), so
fail-open degradations are visible in the panel instead of only in logread.
* Panel: Networks page (which LAN networks are intercepted + inbound editor),
DNS-rules editor, subscription quota/expiry, plane banner and findings list.
* Control-socket client got per-verb timeouts — a wedged daemon used to pile up
one stuck `shaterd status` per minute until OOM.
* cache.db is now bounded (8 MiB, tmpfs fallback below 24 MiB free): on a 98 MB
rootfs with ~33 MB free it could otherwise grow past what an upgrade needs.
* OpenWrt packaging: nftables-json + ca-bundle deps, postinst restart on binary
upgrade, idempotent rt_tables seeding, cron gated correctly.
VERIFIED ON THE TESTBED
fail-closed holds with the engine frozen; offline boot now starts the engine;
RU destinations go direct while the rest goes through a node (per-connection
proof); ads NXDOMAIN with allowlist override; DoH blocked while the configured
upstream still resolves; sqlite history survives a daemon restart; a failed
apply restores the previous config without dropping the engine.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Devices lose per-device Proxy/Target (routing is expressed with ordinary
routing rules whose Source picker targets a device); a device is now pure
DNS policy: identity + Enabled + Block/Allow. The panel drops the
"Route through proxy" toggle and Exit picker, and gains inline rename
(pencil -> input; renaming an unmanaged device upserts it into managed).
New Globals.BlockDoH (uci block_doh, default off): engine-level block of
known public DoH resolvers so clients fall back to plaintext :53 that the
engine intercepts. DNS layer answers the DoH hostnames + the Firefox
canary use-application-dns.net with NXDOMAIN; route layer rejects :443
(tcp+udp, HTTP/3 covered) to the hostnames and dedicated resolver IPs.
Hostnames/IPs that are themselves configured upstream resolvers are
excluded with a warning (never the canary). Reject rules set
Method=default explicitly - a directly constructed "" bypasses the
UnmarshalJSON normalisation and panics the engine at first match
(found live on the VM, regression-tested).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Per-device DNS block/allow and the DNS filter only caught DNS that the
tproxy forward-divert steals (client -> external resolver). A client using
the router itself as DNS hit dnsmasq directly (fib daddr type local bypass)
and slipped every filter. New Globals.DNSIntercept (uci dns_intercept):
when set, nft diverts all LAN :53 (tcp+udp, v4+v6, source-IP preserved)
into the engine ABOVE the fib-local bypass, so even DNS addressed to the
router is hijacked and per-device rules apply to everyone. DoT/DoQ :853
stays rejected (clients fall back to plaintext); DoH :443 can't be
intercepted (stated in the UI). .lan + private reverse zones are forwarded
back to dnsmasq (127.0.0.1:53, direct detour) so local names still resolve.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A WAN uplink on a private DHCP address (provider double-NAT) was wrongly
offered as a LAN source. Interfaces() now tags each interface with its
firewall zone (one `uci export firewall` pass, reusing the zone scanner),
and the picker treats an interface as a LAN network when it has a subnet
and its zone is not wan* — falling back to the private-subnet heuristic
only when the zone is unknown. VPN tunnels self-filter (no subnet).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The blind free-text Source input in both rule forms is now SrcPicker: a
chip slot whose popover offers the router's real private subnets (from
/api/interfaces, host bits normalized to the network address), the
discovered devices by name (the bare IP is what's stored), and a
validated custom IP/CIDR input. Empty = "everyone · all LAN clients".
Existing hand-typed Src values classify back into device/network/custom
chips by value, never rewritten. Shared module cache: one
interfaces+devices fetch per session across all open forms.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Domain matching belongs to rulesets (that's what they are for) — the Match
picker is now rulesets only / ip-cidr / port, with the value input hidden
for rulesets-only. The edit form keeps a "Domain(s) — legacy" field ONLY
when a rule already carries free-text domains, so old rules stay visible
and clearable instead of silently preserved.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- model: Ruleset/Blocklist/Allowlist Category -> Categories []string
(uci `list category`; a legacy lone `option category` still reads as a
one-element list and migrates to the list form on the next write)
- generate: one remote .srs per category, tag rs-<name>-<category>
(bl-/al- for the DNS filter); a rule referencing the ruleset matches
every category's set; non-geo sources keep their old single tags
- panel status: rows gain `category`; tag->(name,category) resolved from
the model, not string parsing (names/categories may contain dashes)
- CatSuggest is now a chip multi-select: pick from the SagerNet base ->
chip with a status LED (green = from base/verified, amber = added
offline "anyway"), duplicates flash the existing chip, Backspace/×
remove, and free unpicked text never survives blur or save
- Routing/DNS forms save Categories (>=1 chip required); ruleset rows
show `geosite · youtube +2`, freshness groups per name (oldest wins,
Update now refreshes every category's tag)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The native <datalist> popup is an unstyleable browser widget that clashed
with the panel. CatSuggest is an instrument-styled readout docked flush
under the Category input: SAGERNET BASE · n shelf label, sunken dense mono
list, matched substring lit in the accent, LED bar on the active row,
green exact-match footer, amber not-in-base warning. Keyboard: arrows /
Enter / Esc; combobox ARIA; both themes via tokens; reduced-motion safe.
Fix along the way: the row is a flex container with a gap, so bare text
nodes around <mark> became separate flex items and the gap split the
category name itself — the name now renders inside one span.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- GET /api/ruleset/categories?source=geosite|geoip — daemon lists the real
SagerNet rule-set branch via the GitHub git-trees API (UA set, 24h in-memory
cache, stale-on-error); panel drives a native <datalist> on the Category
inputs (Routing ruleset form + DNS geosite blocklist), lazy one fetch per
source per session
- Routing rules gained Edit: inline form (all three matchers shown at once —
domains/IPs/port — so nothing is silently dropped), preserves Order/Enabled/
Kill/Egress and off-form fields verbatim, reorder/toggle/delete frozen while
editing, "no matchers — matches everything" hint for catch-alls
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- generate.GeoRuleSetURL(source, category) — single source of truth for the
SagerNet .srs URL (routing rulesets, DNS filter, and the checker)
- POST /api/ruleset/check: daemon-side HEAD (GET+Range fallback) existence
probe of the exact URL the engine would fetch; {ok} / {not_found} /
{network} — plain client, router-own output is never tproxy-diverted
- Panel: geosite/geoip Save now checks first (Checking…); not_found blocks
with a form error; network failure offers explicit "Save anyway" so an
offline router can still be configured; url/inline/file flows untouched
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Three features from the second feedback pass:
- Node.Egress: a node can dial its OWN upstream through a named egress
(DialerOptions.Detour on the node outbound/WG endpoint; inherited by
groups/chains/rules; fail-open on unknown egress). Panel: per-row "via"
expander + "via <name>" chip on Nodes.
- Chains: egress:<name> allowed as the ENTRY hop only (hop 0) — lifted into
the first hop's detour; mid/last egress hops warn+drop. Panel: entry-hop
optgroup + "entry" badge in the chain editor (Targets).
- Test all nodes: engine.TestAllNodes force-probes every node outbound
(concurrency 16, 5s timeout) into the shared urltest history; failures
stored as ProbeFailDelay=0xFFFF sentinel (slowest, never poisons
least_test) and surfaced as DOWN, not untested. POST/GET /api/nodes/test;
"Test all" button with N/M progress on Nodes.
- geosite/geoip rule-sets are LIVE: source=geosite|geoip + category emit
official SagerNet remote .srs rule-sets (24h auto-update, direct fetch),
for routing rulesets AND DNS block/allowlists; freshness + Update now UI
applies to them; "inert" badge removed. New Category field (model+UCI).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
- Nodes/Overview: counters now tally live probe health (up/down/untested/off)
instead of counting Enabled as "up" (was: 329/329 up with 5 tested)
- Routing: Target select bucketed into optgroups (Groups/Chains/Interfaces-
egresses/Nodes-last) so egress:* is no longer buried under 300+ nodes
- Insights/Overview logs: single fmtClock(unix) helper, browser-local time in
BOTH logs (DNS log was server-UTC next to local-time connections)
- Overview query-log badge no longer says "waiting for engine stats" while
persisted rows are on screen
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The stats log rings move behind a logRing seam: memRing (extracted RAM ring,
byte-identical to before) + sqliteRing (new, modernc.org/sqlite v1.38.2 pure-Go,
CGO-free musl-static). backend=sqlite persists the query/conn LOG rows to
/etc/shater/stats.db (WAL, tmpfs fallback /tmp/shater-stats.db, own lock) via an
async batched writer off the DNS/conn hot path; seq = the PK (monotonic, resumes
from the persisted max after restart). Cursor reads = WHERE seq</> ? ORDER BY
seq DESC LIMIT. Retention: keep <= StatsRingSize rows/table + a StatsDiskLimitMB
disk cap (0=unlimited) with prune + wal_checkpoint/VACUUM. Aggregates
(top-domains/timeline/hosts/devices/node-health) stay in RAM (bounded, rebuild
fast) — only the unbounded LOGS persist. Open failure → warn + memRing fallback.
Globals.StatsDiskLimitMB (0=unlimited, intOptAlways round-trip); Settings shows
it when backend=sqlite. Size: +1.2MB UPX (10.5->11.7MB), static/musl OK.
Verified: build (router tags)/vet 0, go test + -race ok (sqliteRing cursor
parity, retention, disk-cap, PERSIST-across-reopen, async no-loss, memory
regression); panel tsc/build clean. VM: backend=sqlite → /etc/shater/stats.db
created, 75 conns logged, **survive daemon restart** (75 rows intact, seq
resumes 76->79), disk cap set 16MB. box.New applies.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Log rows gain a monotonic Seq (uint64, per-log counters on the Aggregator,
stamped under mu on append; survives box swaps + ring wrap). StatsStore read
API becomes cursor-based: Queries(LogQuery{Limit,Before,After})/Conns(...) —
neither cursor = newest Limit; Before=<seq> = next older page (seq<before);
After=<seq> = new rows (seq>after); always newest-first. Endpoints
/api/stats/{log,conns} accept limit/before/after (legacy n = limit, so Overview
is unchanged); bare array, rows carry seq (client derives newest/oldest).
Panel: Insights logs (Connections + DNS) now accumulate a seq-desc deduped
list — Load more APPENDS the next older page (not refetch-all, scroll
preserved, hides when exhausted), a ~1.5s after=<newest> poll PREPENDS new rows
(slide-in keyed by seq), a Pause/Live toggle buffers arrivals into an 'N new'
pill, ~3000-row DOM cap re-arms Load more, poll gated on backend!=off + tab
visible. Stable seq keys.
Verified: build (router tags)/vet 0, go test ok (seq monotonic bounded+
unlimited, before/after paging no overlap/gap, wrapped-ring, clamp), panel
tsc/build clean, ?mock drive (paginate+live+pause). VM live-verify pending
(ssh-manager MCP disconnected).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
First phase of a pluggable stats-storage backend. New stats.StatsStore
interface (Start/Close/Resubscribe/Snapshot/RecentQueries/RecentConns); the
existing in-memory *Aggregator implements it unchanged, plus a noopStore for
OFF that never subscribes (so the HasSubscribers-gated DNS emit path skips all
per-query work). stats.NewStore(backend,...) selects off->noop, memory->agg,
sqlite->agg+warn (persistent backend lands in Phase 3). Snapshot gains a
'backend' field (off|memory|sqlite = effective). Globals.StatsBackend
(off|memory|sqlite, default memory) via the KillSwitch string-enum pattern
(model/uci/render + round-trip). Daemon + panel decouple from *Aggregator to
the interface. Settings gets a 3-way Logging-backend Select; Insights shows an
honest 'logging is off' state when backend=off.
Verified: build (router tags)/vet 0, go test ok (noopStore contract, NewStore
selection, StatsBackend round-trip), VM box.New PASS (stats verb reports
backend=memory); panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
maxStatsLogN capped /api/stats/log and /api/stats/conns responses at 200, so an
unlimited (StatsRingSize=0) or large ring still returned only 200 rows. Raised
the safety ceiling to 5000 (RecentQueries/RecentConns still return only what's
buffered) and lifted the Insights Load-more ceiling 500->5000 (step +200) to
match, so a big/unlimited log can actually be paged out.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Retention size knobs (StatsRingSize / StatsTimelineMinutes / StatsMaxDomains)
now mean: 0 = UNLIMITED (no trim, grows with RAM), N = fixed limit. Query-log
AND connection-log rings gain a growable append-only mode when RingSize==0
(fixed-ring modulo path kept for N>0). Per-field 0 skips that aggregate's prune
(domains) / trim (timeline); RetentionDisabled stays the master switch.
Absent-vs-explicit-0 round-trip fixed: DefaultGlobals seeds safe bounded
defaults (200/60/5000) so an unset UCI option is never accidentally unlimited;
render intOptAlways writes these three fields even at 0 so an explicit 0
survives WriteUCI->ReadUCI; daemon passes no Config on a read error (→ bounded
defaults, not a zero-value=unlimited Config). Settings reframes the three
inputs as '0 = unlimited' with a per-field grows-with-memory warning.
Verified: build (router tags)/vet 0, go test ok (round-trip 0/200/5000;
RingSize=0 grows to 500 q+conn; bounded at 200; MaxDomains=0 keeps 6000;
Timeline=0 no trim), VM box.New PASS; panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The Connections/DNS LogShell only DISABLED 'Load more' when the page wasn't
full — so it stayed visible (and looked clickable) even with e.g. 9 rows. Now
the button is HIDDEN unless a full page came back (rows.length >= n && n < 500),
so it only appears when there may actually be more; disabled only while busy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The real 600-905px 'crossing' was NOT the SegMeter dots (fixed in e24a2c1c) but
the 3 QUERIES/BLOCKED/ALLOWED tiles forced 3-across in the ~227px first column
of the 3-col overview: each value's min-content (~86px) → 277px overflowed the
cell by ~50px, painting the 3rd tile's number ~17-30px into the sparkline.
(Round 1 missed it: Playwright's 15px scrollbar turned physical 901 into an
886 single-col layout, so the tight 3-col band was never measured.) Fix:
.ins-tiles flex-wrap + .ins-tile{flex:1 1 90px;min-width:0} → reflow 2+1 when
narrow, 3-across when wide, robust to any digit count, no viewport breakpoint.
Also: per-device header total wraps to its own line; Connections/DNS rows fit
their scroll box at <=560px. Scoped to Insights; verified 0 crossings 360-1440.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The prior min-width:0/max-width:100% capped the .segs BOX but the 28 flex
segments still painted outside it (their ~305px min-content spilled past the
right border at nearly every width). Now .ins-filter .segs uses overflow:hidden
(drops the min-content contribution + clips sub-pixel) + tighter gap, and the
Filtered meter passes segments={20} so dots fit their column without
compressing past the ~8px floor. Also fixed a secondary body h-scroll ≤404px:
.ins-overview single-col → minmax(0,1fr), .ins-tiles reflow to 2-col ≤400px,
.ins-grid minmax(min(100%,320px),1fr). Scoped to Insights — shared SegMeter
(Overview) untouched. Verified 0 overflow + no body h-scroll across 360–1440px.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
LogEntry.Device was always '' — but the client address IS on the DNS
resolution context. dnstrack.QueryEvent gains Client netip.Addr, populated at
all three dns/client_log.go emit sites from adapter.ContextFrom(ctx).Source.Addr
(same context processInfoFromContext already reads). stats deviceLabel: LAN
source (a.lanNets.isLAN) -> DHCP hostname or IP; loopback/non-LAN/unknown ->
'router' (the appliance's own urltest/sub/DoH lookups). Insights DNS log now
shows device -> domain · resolver · action (mirrors the Connections log), with
a dimmed 'router' chip for router-originated lookups; falls back to '—' on
older data.
Verified: root build (dns tree + box, router tags)/vet 0, go test ok
(deviceLabel: LAN+lease/LAN+IP/loopback->router), panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
User frontend polish pass: (1) Connections/DNS action buttons were glued —
now a left-grouped .ins-log-btns (gap) + right-aligned count + separating
groove. (2) DNS log now shares a LogShell (scroll body + actions) with
Connections so they look 1:1 (differing only in columns: time·domain·resolver·
action); dropped the old <QueryLog> ticker here (still used on Overview).
(3) 'Blocked' toggle was clipped to 'Blocke' — .ins-toggle overflow:hidden
collapsed its flex min-width; added flex:none/nowrap + header flex-wrap.
(4) Filtered SegMeter's 28 segments overflowed the module's right edge —
scoped min-width:0/max-width:100% under .ins-filter (SegMeter elsewhere
untouched). (5) tabular-nums, consistent spacing, removed dead .ins-logwrap/
.qrows + unused imports. tsc/build clean; verified ?mock at 1280 & 380px, no
horizontal body scroll.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The conn-log/top-hosts LAN filter required a non-empty Metadata.Inbound, but
tproxy connection events carry an empty Inbound on this engine — so it dropped
every client connection (live: 0 conns / empty top_hosts despite real traffic).
Now filters by source IP being inside a LAN subnet: devices.LANNets() (new
exported helper reusing the #10 /etc/config/network parser) + a 30s-cached
lanNetCache. Keeps LAN clients (192.168.1.77) and drops the router's own
WAN-side node dials (Source 10.0.2.x) — which are both RFC1918, so only the
iface config distinguishes them. Fallback (no config): private routable
sources. Loopback/unspecified/multicast always dropped.
Verified: build (router tags)/vet 0, go test ok (KEEP 192.168.1.77 / DROP
10.0.2.15 subnet test).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DNS events carry no client IP, so the query log couldn't show which device
went where. Now the stats aggregator folds connection events (trafficcontrol)
into: a connection ring (RecentConns → GET /api/stats/conns: src device ->
dest domain|IP + tcp/udp + sniffed proto + exit) and a top-hosts map keyed by
domain-else-IP (Snapshot.top_hosts) so raw-IP UDP/TCP destinations surface
(host==ip = the by-IP case). LAN-source filter (non-empty inbound + routable
src) keeps the router's own node/probe dials out. Insights gains a scrollable
Connections log (device->dest+proto, Refresh/Load more) + a Top-hosts section
(net/proto badge, IP tag for domain-less); the DNS log is relabelled
'DNS log · decisions'.
Verified: build (router tags)/vet 0, go test ok (LAN fold, IP-only host,
non-LAN skip, Closed adds bytes), panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The shared <QueryLog> is a fixed-height ticker (overflow:hidden) for Overview,
but on Insights it's a full paginated log — Load more fetched rows that were
clipped and invisible. Scoped override: .ins-logwrap .qrows now scrolls
(max-height min(60vh,540px), overflow-y:auto). Overview ticker unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A 25s watchActiveProfile loop (mirrors watchNewDevices) reads the active
default-route dev (ip route show default, lowest metric), matches it against
enabled profiles' MatchIface (via netplane.IfaceDevice, so UCI-name OR device
lists work), and pins the highest-Priority match into Globals.ActiveProfile +
Reconcile — generate already applies an explicit ActiveProfile, so no generate
change. Anti-flap (write only on change), no-op when no MatchIface profiles
exist, releases a stale iface-pin but preserves a manual non-iface pin,
fail-safe on every error. Pure helpers pickIfaceProfile/desiredActiveProfile/
parseDefaultRouteDev unit-tested (failover-flip, tie-break, stale-release).
Schedule-window check deferred (unexported in generate). VM: build + matcher
tests PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Shared foundation: Engine.HTTPClient(via) dials through a running-box outbound
(OutboundManager.Outbound(tag).DialContext); via->tag map direct/group:/node:/
egress:/chain:. #1: Alert gains Via + Fallback — notifier sends through the
chosen detour via an injected client factory (daemon wires eng.HTTPClient);
on detour failure retries direct iff Fallback (else surfaces error). Direct
stays the default + the always-available safety path (killswitch/apply_fail
should keep Via empty or set Fallback). #8: Subscription gains FetchDetour;
new POST /api/subscription/update {name} makes the DAEMON fetch a sub through
the tunnel (FetchVia=proxy → Fetch(sub, HTTPClient(FetchDetour))) → update →
WriteUCI → reconcile; panel gets a per-sub Detour picker (under Proxy) + an
Update-now button. CLI 'sub update' stays direct.
Verified: root+shater build (router tags)/vet 0, go test ok (via->tag map,
alert Via+Fallback fallback-to-direct, detour-fail-no-fallback drops, model
round-trip w/ via/fallback/fetch_detour), VM box.New + engine tests PASS.
panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DNS events carry no client IP, so per-device DOMAIN stats need connection
events. box.go now builds+registers the trafficcontrol.Manager + AppendTracker
UNCONDITIONALLY (moved out of the needObservable gate) — an in-process
connection observable with NO clash/api port opened (Emit is non-blocking, so
an unsubscribed tracker never stalls the hot path). Engine.ConnManager()
exposes it (box-owned; pointer changes each Apply swap). stats connLoop
subscribes (pointer-identity resubscribe like dnsLoop), folding
{Source.Addr, Domain||Destination.Fqdn} into deviceDomains (bounded 512
clients / 200 domains-each). Snapshot gains device_domains
[{ip,name,domains:[{domain,count}]}]; Insights shows a per-device domain view.
Configurable retention: Globals StatsRingSize/StatsTimelineMinutes/
StatsMaxDomains/StatsRetentionDisabled (0=built-in defaults 200/60/5000);
stats.New resolves them, RetentionDisabled skips all pruning (RAM-bounded);
Settings gains a Statistics-retention section with a disable-trim toggle.
Verified: root+shater build (router tags)/vet 0, go test ok (conn-event fold,
retention, TestEngineConnManagerWired drives a real proxied conn + asserts a
live ConnectionEventNew + pointer-change-on-swap), VM box.New still Applies
with the tracker wired. panel tsc/build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Globals reads the canonical 'loglevel' key; a natural 'log_level' misspelling
was silently ignored, so 'log_level=debug' produced no debug output (found
while diagnosing node-health telemetry on the VM). applyGlobals now accepts
log_level as an alias, loglevel keeping priority. +TestLogLevelAlias.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
New Insights nav page surfacing the /api/stats Snapshot that was collected but
never shown: traffic overview (queries/blocked/allowed + timeline sparkline +
filtered%), top domains ranked with blocked portion (all/blocked toggle —
replaces the anemic 'top blocked = none'), per-rule traffic bars (bytes/packets
per routing rule), per-endpoint (outbounds + resolvers by count), per-device
bytes, and the query log (block/proxy/pass) via getStatsLog pagination. Answers
the user's ask: how often & how much traffic goes to which domain/IP, per rule
and per endpoint. Pure frontend — all data already in the aggregator. Polls
getStats every 3s; honest empty states; responsive (no horizontal body scroll).
Reuses SegMeter/QueryLog/Led/Button. Wired via router ROUTES + App Page switch
+ pages/index. Verified: tsc --noEmit clean, npm build ok.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The Overview 'N/N up' was cosmetic (enabled-count/total). Now the engine
pre-registers a shared urltest.HistoryStorage in the box ctx (mirrors the
dnstrack.Manager pattern; box.go reuses a ctx-provided store), exposes it via
Engine.URLTestHistory(), and stats collectNodeHealth() reads per-node
Delay/alive into a new Snapshot.NodeHealth ([]{tag,delay_ms,alive,tested,
age_seconds}, tag==node name). Panel joins it by name: Overview shows an
honest alive/tested/total readout + status LED; Nodes rows get a latency chip
+ alive/down/untested LED (untested = node not in any probing group, shown
'—' not 'down'). Falls back to the old count when node_health is absent.
Only urltest/least_test groups populate history (selector/single/manual do
not); a failed probe deletes the entry, so tested=false conflates never-probed
and last-probe-failed — both reported untested, never a false 'down'.
Verified: go build (router tags)/vet 0, go test engine+stats ok (new
TestNodeHealthFromURLTestHistory + nil-safe test), panel tsc/build clean.
VM live-verify pending (ssh-manager MCP disconnected mid-session).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Post-test feedback batch, 4 parallel Opus agents:
#6 rollback-hide: apply.Status gains can_rollback (armed commit-confirm
snapshot OR engine.HasLastGood(), a new non-mutating engine probe). Apply/
Overview hide the Roll back button when nothing to revert; the 'Nothing to
roll back' dead-end is gone.
#10 devices: parseNeigh now captures the 'dev' token and drops non-LAN
rows (WAN device + a fail-open 10.0.2.0/24 slirp guard), so QEMU WAN IPs
10.0.2.2/.3 no longer masquerade as devices. Discovered gains network/iface
labels (IP matched against /etc/config/network subnets); UI shows 'LAN·br-lan'.
#5 egress iface picker: new GET /api/interfaces (netplane.Interfaces via
ubus network.interface dump); Targets EGRESS interface field is now a select
of real UCI interfaces (degrades to free-text when empty).
#11 WG/AWG import: parse.WGToURI serializes a *Proxy back to a canonical
wireguard:// URI (round-trips ParseWGConf, all AmneziaWG knobs); new
POST /api/import-wg converts a pasted .conf; Nodes add-node accepts a
multi-line [Interface] config and imports it as a node.
#7 nodes search/grouping: live search (name/proto/host/sub) + collapsible
per-subscription and Manual groups (large groups collapsed by default,
search auto-expands matches).
Verified: go build (router tags)/vet/test 0; panel tsc/build clean; new
tests TestCanRollback, TestEngineHasLastGood, TestDiscoverDropsWANNeigh,
TestWGToURIRoundTrip, import-wg + interfaces api tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Profiles/presets were modelled but ignored. Now buildRoute applies them to
an EFFECTIVE rule set (input model never mutated). Active profile: explicit
Globals.ActiveProfile (existing+enabled) always wins; else auto-select the
highest-Priority enabled profile whose schedule window holds (via b.now,
sharing scheduleWindowActive with rule schedules); iface/probe-conditioned
profiles are skipped by auto-select (warn, control-plane Phase-2b) but
honored when pinned. Overrides: EnableRules/DisableRules (disable wins),
DefaultTarget/DefaultEgress on Final (target wins = leak-safe). Preset packs
block-ads(15 domains->block)/ru-bypass(geoip:ru->direct, inert w/o geodata)/
private(RFC1918+ll+lo->direct) inject rules through the SAME rule loop,
Preset.Order/Target overridable. Fail-open throughout; profile/preset-free
models byte-identical (regression-guarded). Verified build/vet 0, host tests,
VM box.New (TestProfilePresetAppliesCleanly).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Routing page now manages rule-sets (domain/ipcidr match sources): add/edit/
delete with source inline(entries)|url|file|geosite, and a dst_ruleset
checkbox picker in the rule form so a rule matches one or more rulesets
(matcher chip shows 'ruleset: ..'). Deleting a ruleset strips it from every
referencing rule. Reuses the existing save->apply machinery; URL tokens
masked; honest empty state. Backend materialisation landed in 57343693.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
resolveChainExit collapsed a chain to its exit hop, losing the multi-hop.
Now buildChain materialises per-chain hop-outbound copies Detour-linked
backward (h_n exits, h_n.Detour=h_{n-1}, .. h1.Detour=direct) so traffic
traverses L1..Ln and egresses at Ln; a rule/egress routing chain:<name>
targets the chain ENTRY tag. Per-chain copies keep base node/group
outbounds standalone and preserve the 0xff loop-guard mark. Group hops =
chain-local urltest over detoured member copies; WireGuard hops = detoured
endpoint copies. 1-hop == that hop; undefined/empty/unresolvable warns +
rule skipped (never aborts box.New); only referenced chains materialise.
Verified: build/vet 0, host tests + VM box.New (TestChainMultiHopApplies).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Move dnsBlockRuleForSrc to devices_test.go (imports option/C, untagged) and
drop the now-unused constant import from devices_linux_test.go, fixing the
cross-compile of the previous test fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
TestDeviceRulesValidate used dnsRuleForSrc (first source match) which
returned the device's ALLOW rule (route action, emitted first) while
asserting Predefined — a false failure. The generation was correct
(Phase-6 verified block works E2E). Now asserts the predefined-NXDOMAIN
block rule for the src specifically via dnsBlockRuleForSrc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Reaches v0.1's config-surface parity for the backend-ready features. Nav
gains Targets, Profiles, Settings (Profiles is a placeholder until its
backend lands).
- Targets page (groups/chains/egresses — was UCI-only): GROUPS editor
(source subscription|manual, subscription/member-node pickers, strategy,
include/exclude/proto/country filters + dedup, probe url/interval); CHAINS
editor (ordered hop list from group:/node:, signal-path viz); EGRESSES
editor (type interface|proxy|direct|block|byedpi, interface/target/port +
native DPI preset off|fragment|record|spoof). All pickers derive from live
config.
- Settings page (globals — was UCI-only): enabled, log level, kill-switch,
DNS mode, IPv6, confirm timeout, panel port, health probe url/interval,
fwmark/table base (hex, advanced), read-only schema/active-profile.
- Nodes page: per-subscription options expander — update interval, fetch-via,
format, UA, HWID + device fields, extra headers, include/exclude/proto/
country filters, dedup, expire-alert days (secrets masked, reuses save
machinery).
- api.ts: full types (Subscription/Group filters, Chain, Ruleset, Preset,
Profile, Inbound, Globals.PanelPort) + Model slices; router nav.
All save→apply like the other pages. tsc clean; build ok (82 kB gzip).
Remaining for full parity (next waves): generate for rulesets/chains(real
multi-hop)/profiles/presets, and the Profiles + Rulesets panel pages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Lets the operator choose which proxy path DNS goes through — a group
(balancer/urltest), a chain, an interface/egress, a specific node, or
direct. The backend already resolved resolver.Detour via resolveTarget
(group:/chain:/egress:/node:/direct); this exposes it in the panel (the
RESOLVERS section was read-only).
- Per-resolver detour <select> built live from the Model: Direct + a
Groups optgroup (balancer) + Chains + Interfaces/egresses (with type) +
a Nodes optgroup. Current path rendered as 'via group/chain/node/
interface <name>' or 'direct'.
- Add resolver (name + type doh/dot/plain/tcp/local/fakeip + conditional
address + fakeip pool + detour), edit, delete (repoints/clears default+
fallback), and editable default/fallback role selects (were read-only).
- Stale/missing detour target stays selectable + flagged '(missing)';
legacy bare-name detours normalized against the catalog. Secrets masked.
save->apply banner like the other sections.
Verified: tsc --noEmit clean; npm run build ok (71 kB gzip JS); Playwright
(mock) — the detour select shows Direct/Group auto (balancer)/egresses/
Nodes optgroup, changing it + add/delete + default/fallback all work with
the save->apply banner.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A group with source=subscription produced NO members (groupMembers only
read the explicit g.Nodes list, empty for sub-backed groups), so the group
was skipped and any rule targeting it never routed — the whole proxy path
was dead on a subscription setup.
Fix: for source=subscription, gather members from all nodes where
FromSub==g.Subscription (enabled + emitted), in config order, deduped; apply
Include/Exclude name regexes (case-insensitive, bad pattern warns+ignored)
and FilterProto/FilterCountry/Dedup via parse.ParseShareLink +
FilterSpecFromGroup + ApplyFilters. Manual/single/'' sources unchanged.
Verified: generate tests (sub group gathers exactly its FromSub nodes not
others; include/exclude; bad-regex-ignored; proto filter; manual unchanged).
Found live on the VM: 376-node subscription group 'auto' was 'no usable
members, skipped' before this fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Makes shaterd sub update real (was a Phase-2b stub): fetch a subscription
URL, parse+filter into nodes, persist them, reconcile.
- shater/subscribe: Fetch(sub, client) — HTTP GET with UA (sub.UA or
Shater/0.2 default), HAPP-style x-hwid/x-device-* headers, raw Headers
override, 20s timeout, 8MiB cap, non-2xx/empty = error. UpdateSubscription
(pure): parse.ParseSubURIs re-serializes ALL formats (clash/xray/sing-box/
links) to canonical share-links, pairs each with its *Proxy, ApplyFilters
(Include/Exclude/proto/country/dedup), builds []Node FromSub=<name> (name
from #fragment, collision-suffixed), REPLACES only that sub's cache. Zero
usable nodes -> error + leave the cache intact (a provider hiccup never
empties the config).
- cmd/shaterd: 'sub update [<name>]' — fetch each enabled sub (or one),
fold in, WriteUCI, best-effort SIGHUP reconcile; works with/without the
daemon; fetch_via=proxy warns + falls back to direct (MVP).
- model: sub-cache nodes now persist in UCI (config node + from_sub/
fingerprint/stale) so the fetched set survives restarts and the panel sees
them; ReadUCI reads them back; manual nodes unaffected. (v0.1 used a
separate JSON cache; UCI persistence matches v0.2's model<->UCI design.)
Verified: subscribe+model unit tests (FromSub tagging, filters, zero-node
safety, cache replace-keep-others, name collisions, UA/header/non-2xx); VM
E2E with the real feed https://pro.qomar.pw/sub/... -> 'sub update' fetched
376 nodes (261 vless/71 ss/29 vmess/15 trojan), all from_sub='default',
persisted to UCI, and box.New/Apply accepted all 376 (status running/active).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Ports the v0.1 Gitea release flow to the v0.2 single-binary + 4-package
layout, so a tag publishes a signed opkg feed the routers install from.
- .gitea/workflows/release.yml: on tag v* (+ dispatch), matrix over
{x86_64, aarch64_cortex-a53}. Per arch: setup Go 1.24/Node 20/UPX ->
scripts/build-shaterd.sh (SPA-embedded shaterd, stages the .upx) ->
ci/build-feed.sh (OpenWrt SDK container builds all 4 packages ->
usign-signed Packages index). A release job merges both arches into one
signed feed + publishes the rolling 'latest'/tag release via the Gitea API.
- ci/sdk-build.sh: in-SDK build — add openwrt/ as the 'shater' feed, feeds
update/install, make package/{shaterd,shater-core,byedpi,luci-app-shater}/
compile (shaterd validates+installs the staged prebuilt; byedpi cross-
compiles from source). ci/make-index.sh: opkg Packages(.gz) + usign sign
with KEY_BUILD (keyfile umask 077, no secret hardcoded), verifiable by
dist/shater-feed.pub. ci/install-usign.sh + ci/gitea-release.sh ported.
- INSTALL.md: add the signed feed src/gz line + import dist/shater-feed.pub
to /etc/opkg/keys; apk (25.12) path noted.
Key kept: usign feed key 5ac4b177689cb8e0 (public dist/shater-feed.pub,
secret Gitea repo secret KEY_BUILD). Decision: opkg (24.10 uses opkg; apk
is 25.12) — matches the existing usign trust anchor.
Verified structurally (no live runner here): release.yml is valid YAML, all
ci/*.sh are bash -n clean, no hardcoded secrets, and every package name/
path/arch/artifact/secret reference cross-checks against openwrt/, scripts/
build-shaterd.sh, and dist/shater-feed.pub. Live-runner unknowns (full SDK
compile of the 4 packages, router-side signature verify) flagged in-agent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Makes the whole product installable — shater-core DEPENDS +shaterd, and
this is what resolves it.
- scripts/build-shaterd.sh: the release build. Builds the panel SPA
(npm ci && npm run build), copies panel/dist -> shater/panel/webroot
(the go:embed dir), cross-builds shaterd for amd64 + arm64 with the D9
router tag set (CGO_ENABLED=0, -checklinkname=0 -s -w, static ET_EXEC no
PT_INTERP), then UPX --lzma --best (D10) and stages the .upx into
openwrt/shaterd/files. Version from arg/SHATER_VERSION/git-describe.
Measured: amd64 40.3MB->10.4MB, arm64 37.6MB->8.5MB.
- openwrt/shaterd: prebuilt-binary package (npm+embed+UPX don't reproduce
cleanly in the SDK, so CI stages the artifact). Maps OpenWrt ARCH
(x86_64->amd64, aarch64->arm64 = both BPI routers) to files/shaterd-<a>.upx,
installs /usr/bin/shaterd. RSTRIP/STRIP disabled (the SDK strip would
corrupt the UPX binary); DEPENDS empty (static); errors clearly when no
artifact is staged. GPL-3.0-or-later.
- docs-shater/INSTALL.md: build + install order (shaterd -> shater-core ->
luci-app-shater, optional byedpi) + enable/apply.
- gitignore: dist/shaterd-*, openwrt/shaterd/files/*.upx, panel webroot.
Verified on the OpenWrt musl VM: dist/shaterd-amd64.upx (10.4MB) decompresses
into RAM + runs (shaterd status OK), serves the REAL embedded Faceplate SPA
at :8088 ('SPA embedded=true', real Vite index.html + assets — not the
placeholder).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Out-of-band notifications on key events, delivered DIRECT to the internet
(plain net/http, never via the proxy) so a kill-switch/engine-down alert
reaches Telegram even when the tunnel is down.
- model: Alert{Name,Enabled,Type(telegram|webhook),Token,ChatID,URL,
Events[]} + Model.Alerts; uci (case alert, list event) + render +
round-trip fixture.
- shater/alert: Notifier — 8s-timeout default-transport client, per-alert
async delivery (telegram sendMessage / webhook JSON POST), (event,title)
dedup within 60s so a flapping engine can't spam, panic-safe. Update()
swaps config on reconcile; TestFire() for the test verb.
- cmd/shaterd: builds the Notifier in run, Update()s it after each
reconcile; fires apply_fail+killswitch on an apply/reconcile error while
enabled (initial + SIGHUP paths); watchNewDevices polls devices.Discover
every 45s and fires new_device on an unseen MAC (skips the startup
baseline). 'alert test' verb sends a test to every enabled alert (reads
UCI, no live daemon needed). Events emitted now: killswitch|new_device|
apply_fail; node_down|sub_expiry reserved.
- panel: Alerts section in DNS.tsx — list (name/type/events, enable, delete)
+ add form; Token/URL NEVER shown in clear (masked, round-tripped). api.ts
Alert type + Model.Alerts.
Verified: model round-trip incl. Alert; notifier unit tests (subscribed
delivers correct JSON, unsubscribed/disabled deliver nothing, dedup
suppresses rapid dup); build/vet; panel tsc+build; VM — 'shaterd alert test'
with a webhook alert delivered a POST to a local sink.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
devices_linux_test.go (generate box.New validation of per-device rules) and
panel/devices_test.go (/api/devices gating) were authored with the Phase-6
backend (677a1dde) but not staged in that commit. No source change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Scheduled rules are now evaluated by the control-plane at gen/reconcile
time (the engine has no time match), so a rule is only active inside its
window.
- generate: builder gains an injectable 'now' (defaults time.Now); a
SchedEnabled rule outside its window is SKIPPED (was emitted
unconditionally with a deferred warning). schedule.go: scheduleActive
parses SchedDays (mon..sun, empty=all), SchedStart/End HH:MM in SchedTZ
(LoadLocation, fallback local), handles overnight windows (end<start ->
now>=start || now<end), all-day (empty/equal end). Invalid HH:MM ->
fail-OPEN (emit + warn) so a typo never silently drops protection.
- cmd/shaterd: real 'schedule due' verb -> SIGHUP reconcile (no-op when
down); generate re-evaluates + the config-hash gate rebuilds the engine
only when a window boundary was actually crossed (no churn between
boundaries — verified reconcile changed=false per tick).
- shater-cron: calls 'shaterd schedule due' each tick (cheap, hash-gated).
- panel Routing: add-rule form gains schedule controls (enable + mon..sun
day toggles + From/To time inputs); scheduled rules show a days+time chip.
Verified: schedule unit tests (weekday window emitted/skipped, overnight
active across midnight, all-day weekend, invalid-time fail-open); build/
vet; panel tsc+build; VM (box.New accepts a scheduled config; 'schedule
due' reconciles changed=false = no churn).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
In-process stats fed by the engine's DNS-query event stream + nft counters.
- engine: DNSQueryManager() accessor; engine.New pre-registers a stable
*dnstrack.Manager into e.ctx (box.New only creates one when an api/
clash_api observable is present, which the router config has none of, so
the manager would be nil — pre-registering keeps the DNS stream alive).
- shater/stats: Aggregator subscribes to dnstrack QueryEvents and maintains
bounded top-domains, allowed-vs-blocked (blocked = NXDOMAIN / 0.0.0.0 /
failed), a 60-min timeline, a 200-entry live query-log ring, per-server
counts; polls netplane.ListClients/ListCounters for per-device + per-rule
traffic (client IP -> DHCP hostname). Snapshot()/RecentQueries(); re-subs
on box swap; resilient when the engine is down.
- daemon: creates+starts the aggregator, Resubscribe() after each reconcile,
Close on SIGTERM; control-socket 'stats' verb returns the real snapshot.
- panel: GET /api/stats (snapshot) + GET /api/stats/log?n= (live log),
session-gated; Stats type + getStatsLog() in api.ts; Overview QueryLog now
polls the live log, plus a DNS-filtering module + blocked SegMeter + top-
blocked list. Honest empty states, no fabricated data.
- upstream (minimal, marked // lx/D15): dnstrack SourceFiltered +
emitFilteredResponse at the two DNS-filter predefined-block sites in
dns/router.go — filter blocks now feed the query stream (were invisible).
Verified: stats+panel unit tests; panel tsc+build; VM E2E — DNS traffic
from a netns client produced /api/stats totals (queries 16, blocked 6),
top_domains[blocked-ad.example blocked 6], per-device row, and /api/stats/log
rows with correct block/allow; stream survived a box swap (SIGHUP). Overview
screenshot shows the live query log + blocked stats.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The DNS filter blocked via reject/default, which sing-box answers with
REFUSED — non-standard for ad-blocking and contradicting the Blocklist
Response field (nxdomain|zero) + the code comments. Switch to sing-box's
predefined DNS action:
- Response nxdomain (default) -> predefined Rcode NXDOMAIN (RcodeNameError).
- Response zero -> predefined Answer A 0.0.0.0 (+ AAAA :: when ipv6), owner
'*.' so one record serves every domain in a many-domain rule-set. This
finally implements 'zero'; the old 'not supported' warning/fallback is
gone.
- Remote rule-set fetch: DownloadDetour (deprecated in sing-box 1.14) ->
HTTPClient{DialerOptions{Detour: direct}} — no deprecation at box.New.
Verified on the VM via the live in-engine DNSRouter.Exchange: an nxdomain
blocklist -> NXDOMAIN(3); a zero blocklist -> NOERROR(0) + A 0.0.0.0 (not
REFUSED, not NXDOMAIN); a url remote blocklist with http_client validates
with zero deprecation notices. box.New tests cover nxdomain/zero/zero-no-
ipv6/remote/live.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Completes Phase 4 and passes its gate.
- generate: emit experimental.cache_file (enabled, /etc/shater/cache.db,
/tmp fallback) so remote rule-sets persist + auto-update and megalists
stay RAM-sane. Daemon + uci-defaults create /etc/shater.
- cmd/shaterd: real 'blocklist update' verb — SIGHUP-reconcile the running
daemon so url/file rule-sets re-fetch (cache_file updates); no-op when
down. shater-cron fires it on the blocklist interval.
- engine (D16): cache_file's bbolt EXCLUSIVE lock broke the apply-swap —
the new box couldn't take the lock the old held, so every live reconcile
stalled ~10s then failed 'cache-file timeout' (edits silently ignored).
Apply now treats the cache-lock timeout as a swap conflict AND proactively
goes close-old-then-start-new when the incoming config shares the running
cache_file (sharesCacheFileLock), no stall. Regression test added.
Phase-4 gate PASSED on the OpenWrt VM (netns client, dns_filter on):
blocked-ad.example + doubleclick.net -> blocked (reject, no answer);
example.com -> resolves via the engine resolver; enabling an allowlist
entry + 'blocklist update' -> doubleclick.net resolves (allow overrides
block); a real geosite ads megalist (.srs) loaded and its domains blocked
with shaterd RSS ~39 MB (sane); 'blocklist update' reconciles cleanly.
Follow-ups (flagged, not blocking): sing-box reject returns REFUSED not
NXDOMAIN (comment says NXDOMAIN — a predefined-NXDOMAIN action would match
the usual ad-block convention); remote rule-set download_detour is
deprecated in sing-box 1.14 (works, rename before 1.16).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The DNS-filter control surface, wired to the Phase-4 backend.
- DNS FILTER master Toggle (Globals.DNSFilter) with a network-wide
ad/tracker-blocking label + DNS mode / default+fallback resolver readout.
- BLOCKLISTS: per-list source badge (inline/url/file/geosite), URL host
(token masked) / inline entry count, NXDOMAIN|0.0.0.0 reply, enable
toggle + delete, 'filter off' badge when a list is on but the master is
off. Add form (name + domains-textarea|URL + reply). Quick-add chips seed
StevenBlack / OISD / AdGuard with canonical URLs (dedup-guarded).
- ALLOWLISTS: same pattern (overrides blocklists).
- RESOLVERS: read-only display (name, type, host masked, detour) — editing
deferred, noted on-screen.
- save->apply split (putConfig -> toast + banner -> apply) like the other
pages; secrets never rendered; honest empty states.
api.ts promoted: Blocklist/Allowlist interfaces, Model.Blocklists/Allowlists,
Globals.DNSFilter (from the page's local decls). Wired into App.tsx + index.
Verified: tsc --noEmit clean; npm run build ok (62.8 kB gzip); Playwright
screenshot (mock) confirms Faceplate fidelity + the add/quick-add/empty
states. DNS is no longer a placeholder — 5 of 6 nav pages are live (Devices
awaits Phase 6).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
In-engine DNS blocklist/allowlist filtering built on sing-box's compiled
rule-set matcher (no custom megalist matcher, per D5/D15). Rides the same
in-engine DNS plane as the hijack-dns funnel (D14).
- model: Blocklist{Name,Enabled,Source(inline|file|url|geosite),URL,Path,
Entries,Response(nxdomain|zero),UpdateInterval} + Allowlist; Model gains
Blocklists/Allowlists; Globals.DNSFilter master enable (opt-in, default
off). uci parse + render; round-trip fixture extended.
- generate/dnsfilter.go: each enabled list -> a rule-set (inline for
Entries as domain_suffix so subdomains match; remote for url w/
DownloadDetour=direct so fetches don't blackhole under kill-switch;
local for file; geosite skipped inert when no geodata). DNS rules
prepended: allow FIRST (rule_set:[al-*] -> route to default resolver,
terminal, so allowlist overrides), block SECOND (rule_set:[bl-*] ->
reject NXDOMAIN). zero-response -> NXDOMAIN + warn (predefined answer
needs a per-query name a many-domain rule-set can't carry). Off/empty ->
emits nothing.
- shater-core config: commented StevenBlack/OISD/AdGuard blocklists +
allowlist examples + dns_filter note, all inert.
Verified: model round-trip + generate tests; VM box.New (Apply+Start) of
DNSFilter=true + inline blocklist + allowlist + tproxy + doh resolver ->
valid sing-box config. List fetch/compile + 'blocklist update' verb + the
block-a-domain-E2E gate are the next task (remote rule-sets self-fetch).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Closes two gaps flagged by the LuCI launcher:
- panel port was hard-coded :8088 on both sides. Now globals.panel_port
(0 = default 8088): model parses/renders it; cmd/shaterd binds
:<panel_port> when set, else falls back to SHATER_PANEL_ADDR (env can
still disable). apply.Status + shaterd status now report panel_port so
LuCI builds the Open-panel redirect from it (fallback 8088), not a
constant.
- apply.Status gains kill_switch (from globals) so the readout/LuCI can
show fail-closed vs open. LuCI dashboard adds a kill-switch LED
(closed=green, open=amber).
api.ts Status type + mock updated to the new json keys (kill_switch,
panel_port). Round-trip fixture updated (PanelPort) and still passes.
Verified: build+vet, model+panel tests, panel tsc, node --check dashboard;
VM — with option panel_port '8090' the daemon binds :8090 (not 8088),
status reports panel_port:8090 + kill_switch:closed, /api/status on :8090
returns 401 without a session (listening).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The reachability layer for the admin panel (ARCHITECTURE §2): LuCI — an
already-authenticated, ACL-checked session — mints a single-use handoff
token and opens the embedded panel with a session, so the panel needs no
login of its own.
- shaterd: new 'mint-token' verb — dials the running daemon's control
socket, prints its {token} JSON. Robust bridge for the rpcd plugin
(stock OpenWrt has no AF_UNIX client: busybox nc lacks -U, no socat).
No existing verb touched.
- luci-app-shater: client-JS LuCI app under Services —
* rpcd exec plugin /usr/libexec/rpcd/shater: ubus object 'shater' with
status (shaterd status passthrough) + mint_token (shaterd mint-token).
* acl.d: least-privilege (status=read.ubus, mint_token=write.ubus,
scoped to the shater object).
* view dashboard.js: 5s-poll LED status grid + prominent 'Open panel'
button — mint_token via rpc.declare, then open http://<host>:8088/?t=
<token>; button disabled + reason when the daemon is down; tab opened
inside the click gesture so popup blockers don't kill it.
* menu.d entry, luci.mk Makefile (LUCI_DEPENDS +shater-core +rpcd,
PKGARCH all, GPL-3.0-or-later), uci-defaults (chmod plugin +x, reload
rpcd, clear luci cache).
Verified on the OpenWrt VM: shaterd status/mint-token JSON; rpcd plugin
list/call paths; and the full handoff — GET /?t=<token> -> 303 + HttpOnly
SameSite=Strict cookie -> /api/status 200; no cookie -> 401; token reuse
-> no second cookie (single-use). sh -n clean, ACL/menu JSON valid.
Known gap (flagged): panel port is hard-coded :8088 (SHATER_PANEL_ADDR
default) — a future globals.panel_port UCI field would let both sides
share the source. apply.Status has no kill-switch field yet (view shows
the available fields, no fabrication).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Three parallel-built Faceplate pages wired into the shell (Overview was
already live; DNS/Devices stay placeholders until Phases 4/6):
- Nodes.tsx: node + subscription management. Node list with enable toggles,
protocol pills (from the share-link scheme), MANAGED/STALE tags, add-from-
share-link, delete; subscriptions add/toggle/delete. Secrets NEVER shown —
protocol + masked host only; sub URLs show host with 'token hidden'. Sub
on-demand-refresh is wired-but-disabled (needs a backend endpoint).
- Routing.tsx: first-match rule list on a 'signal bus' rail with order
steppers, matcher-summary chips, target chips (group/node/egress/direct/
block), enable toggles; add-rule form with a target picker derived from the
live config; catch-all rule visually distinguished as route Final.
- Apply.tsx: commit-confirm control room — live LEDs + config hash, Apply
with a ConfirmTimeout countdown + Confirm (else honest auto-rollback note),
Rollback with a consequences confirm, before->after hash readout.
All three consume shater/panel's API (getConfig/putConfig/apply/confirm/
rollback) with the save->apply separation, honest empty/error states, no
fabricated data. api.ts Rule widened with the full field set (Src/Dst*/Proto/
Kill/Sched*) so pages share the type; wired into App.tsx's page switch +
pages/index.ts.
Verified: tsc --noEmit clean; npm run build ok (59.5 kB gzip JS, still
react+react-dom only); Playwright screenshots of Nodes/Routing/Apply (mock
backend) confirm Faceplate fidelity + real API-driven rendering.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Replaces the static showcase with the real wired SPA on the fixed
Faceplate design. Stack stays react+react-dom only (52 kB gzip JS).
- src/api.ts: typed same-origin client (Status + Model sections from
shater/model), credentials:include, ApiError w/ 401 -> unauth state.
- src/session.ts: ?t=<token> handoff -> POST /api/session -> scrub token,
preserve hash route (matches ARCHITECTURE §2 / the panel server bridge).
- src/router.ts: ~25-line hash router (useSyncExternalStore), no deps.
- src/App.tsx: Faceplate shell — header (master status LED from /api/status
+ Clock + ThemeSwitch), engraved 6-tab nav (Overview live; Nodes/Routing/
DNS/Devices/Apply placeholders for the next page-agents), footer statusbar
(nft LED + engine hash), loading/ready/unauth/error state machine + 5s
status poll. Reuses the existing Faceplate components.
- src/pages/Overview.tsx: REAL data from /api/status + /api/config — engine/
config/data-plane/kill-switch LEDs, module cards with live counts, a
QueryLog polling /api/stats that degrades to an honest empty state (no
fabricated stream), and an Apply/Confirm/Rollback control row wired to the
endpoints with inline result + toast.
- src/mock.ts: ?mock dev fixture backend (no daemon needed); vite /api proxy
to 127.0.0.1:8088 otherwise.
Contract for the remaining pages: add pages/<Name>.tsx, export from
pages/index.ts, switch in App.tsx; read/write via api.ts; shell owns
header/nav/footer/auth.
Verified: tsc --noEmit clean; npm run build ok (52 kB gzip); Playwright
screenshots of Overview in light + dark + mobile confirm Faceplate fidelity,
real API-driven counts/LEDs, visible focus, reduced-motion respected, and
the apply flow (inline 'reconciled' + toast + hash refresh).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Lets the panel EDIT config, not just read it.
- model.RenderUCIExport(m): pure inverse of ParseUCIExport (Model ->
'uci export shater' text), field-for-field on the same uci keys, shipped
anonymous-section + option-name convention, uci single-quote escaping.
Round-trip invariant ParseUCIExport(RenderUCIExport(m))==m holds for a
rich fixture (every section type, lists, bools both ways, ints/hex,
embedded quote). Bools always emitted (missing != false for default-true
fields); sub-cache nodes (FromSub!='') skipped (runtime state, not UCI).
- model.WriteUCI(m): delete-then-import ('uci delete shater' -> 'uci import
shater' <text> -> 'uci commit shater') so it REPLACES rather than appends
(busybox uci import merges). Via the uciRunner seam (extended with
Import); only WriteUCI touches uci, Render is pure.
- panel PUT /api/config: session-gated, decodes a Model, light validation
(manual node must carry a URI -> 400), WriteUCI, returns {ok,applied:false}
— editing does NOT auto-apply; client calls POST /api/apply after.
GET/PUT method-dispatch on the same path.
Verified: round-trip + write-replace-idempotence + PUT handler unit tests,
and VM E2E (mint->session->GET config->PUT adds a node->200; uci export
shows the new anonymous config node, +1 exactly no duplication, migrate
re-parses cleanly; original config restored).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The HTTP server the admin panel + thin LuCI consume, running INSIDE the
shaterd daemon and sharing its single *apply.Applier (no second engine).
- shater/panel: net/http server (no framework deps). JSON API under /api:
GET /api/status (apply.Status + version), GET /api/config (current
model.ReadUCI), POST /api/apply (snapshot -> reconcile -> arm-rollback),
POST /api/confirm, POST /api/rollback, GET /api/stats (Phase-5 stub).
- Auth per ARCHITECTURE §2: LuCI mints a single-use short-TTL token over
the daemon's unix control socket (new 'mint-token' verb -> MintToken);
POST /api/session {token} validates+consumes it and sets an HttpOnly,
SameSite=Strict session cookie; all other /api routes require it (401
otherwise). Also a GET /?t= redirect bridge matching the §2 diagram.
In-memory token/session stores with expiry; no external deps.
- Serves the embedded Faceplate SPA (go:embed all:webroot; build copies
panel/dist -> shater/panel/webroot, gitignored w/ .gitkeep so it compiles
on a fresh checkout, placeholder page when unbuilt). SPA fallback; unknown
/api/* -> JSON 404, never index.html.
- cmd/shaterd: cmdRun starts the panel server in a goroutine (bind failure
log-and-continue like the control socket), closed on SIGTERM. Gated by
SHATER_PANEL_ADDR (default :8088; off=disabled) to avoid widening the UCI
contract now.
Note: config MUTATION endpoints are intentionally NOT added yet (uci/LuCI
is the writer); /api/apply operates on current UCI. Router build uses the
D9 tag set (with_purego forces a glibc PT_INTERP, unusable on musl).
Verified: unit tests (401 no-session, session->cookie, single-use replay
401, expired/invalid 401, SPA serve) + VM E2E on the OpenWrt VM (mint ->
session -> cookie -> /api/status 200 -> replay 401 -> / serves SPA).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
shater routes marked packets via 'ip route add local default dev lo table
<N>'. Local-delivery of a packet whose source is on a directly-connected
subnet requires accept_local=1 on the LAN INGRESS interface — rp_filter=0
alone is not enough. Sysctls() set lo.accept_local=1 but never the LAN
iface, so with the shipped set a real LAN client behind br-lan is BLOCKED
(engine up, table applied, yet the tproxy'd packet never reaches :12345 —
it escapes to the fail-closed forward drop and the LAN goes dark). The
Phase-2 gate only passed because a stale accept_local was left set on the
test VM.
Fix (scoped, least-privilege — not a global 'all' change): new
netplane.ApplyIfaceSysctls(m) sets, per enabled tproxy ingress device
(nftEnabledInboundDevs), net.ipv4.conf.<dev>.accept_local=1 and
.rp_filter=0 (rp_filter is MAX(all,iface), so the static all=0 can't
override an iface value of 1). Called from apply.applyLocked after
ApplySysctl, fail-closed. Static Sysctls() drop-in can't know device
names, so this is dynamic at apply time; hotplug reconcile re-applies.
Verified on the OpenWrt VM from the buggy baseline (br-lan.accept_local=0):
shater's own apply flips it to 1 and a netns LAN client that was blocked
(rc=4 'Operation not permitted') now reaches the internet (egress WAN IP).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The process behind a shater egress of type='byedpi': an optional, separate
OpenWrt package that ships ByeDPI (ciadpi) + a procd supervisor. shaterd's
generate emits a SOCKS5 outbound egress-<name> -> 127.0.0.1:<port> (Phase
2b-i); a ciadpi instance from this package listens on that port, applies
TCP/TLS desync, and goes DIRECT (no tunnel).
- Makefile: package byedpi, pinned upstream v0.17.3 (real PKG_HASH), MIT,
per-target (compiled C via SDK toolchain calling ciadpi's own make).
- init.d/byedpi: procd multi-instance (one ciadpi per enabled config
instance, 127.0.0.1:<port> + desync args), inert by default, respawn,
config-change reload, validation. sh -n clean.
- config/byedpi: default instance disabled, port 1080, a documented desync
preset. uci-defaults/40_byedpi enables the init.
- Kept SEPARATE from shater-core (byedpi egress is opt-in).
Verified E2E on the OpenWrt VM: musl-static ciadpi (146 KB, no PT_INTERP,
Alpine-built) proxies + desyncs (log: DESYNC_DISORDER); a netns LAN client
routed through a type='byedpi' egress reaches the internet direct via
ciadpi; kill-switch stays honest (SIGKILL shaterd -> client blocked).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Model ciadpi (ByeDPI) as 'just another egress' per D13: an egress of
type 'byedpi' emits a SOCKS5 outbound to 127.0.0.1:<port> (loop-guard
mark so ciadpi's own upstream isn't re-diverted), which a routing rule
targets. The desync happens inside ciadpi, so no native tls_* flags apply.
- model: Egress gains Port int (ciadpi listen port, default 1080); uci
parses option port; Type doc now lists byedpi.
- generate: type 'byedpi' egress -> C.TypeSOCKS outbound (version 5,
127.0.0.1:port) tagged egress-<name>.
The ciadpi binary + procd package + E2E is Phase 2b-ii (separate). Pure
Go half verified: unit tests + box.New validation of a byedpi egress on
the OpenWrt VM (router tags) PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
An egress gains an optional 'dpi' preset that surfaces sing-box's already-
compiled route-action desync fields, so a ruleset can go DIRECT + desynced
with no tunnel and no extra binary (the DPI-blocked-but-not-IP-blocked case):
fragment -> tls_fragment (split the TLS ClientHello record)
record -> tls_record_fragment (alternative; mutually exclusive w/ fragment)
spoof -> tls_spoof (decoy ClientHello; wrong-sequence default)
- model: Egress gains DPI string; uci.go parses option dpi.
- generate: a type 'direct' egress now emits a real 'egress-<name>' direct
outbound (loop-guard mark) so it resolves as a rule target at all — before
this a direct egress target referenced a non-existent outbound (latent bug).
buildRoute's applyDPI stamps the matching route-action flag on every rule
routed to a DPI egress; fragment<->record mutual exclusion enforced; a DPI
preset on a default/catch-all egress warns (route Final carries no action).
byedpi reserved for Phase-2b (external SOCKS egress, D13); unknown -> warn+off.
- openwrt example config + ROADMAP updated.
Verified: unit tests (fragment/record/spoof/none/unknown) + box.New validation
of fragment and spoof configs on the OpenWrt VM (router tag set) all PASS.
tls_spoof validates at box.New time on the linux/router build (raw sockets are
only touched at dial time).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The netplane dnsnat chain redirected LAN :53 to the router's :53 assuming
the engine answered there, but generate/engine create no :53 DNS server —
so on the VM the redirect landed on dnsmasq, which resolved via its WAN
upstream OUTSIDE the tunnel (DNS leak), leaving the engine's own resolvers
(built with anti-leak detours) unused.
Per D14, use the sing-box-native hijack instead of an nft redirect:
- generate/route.go: hijackDNSRule() after sniffRule — matches the sniffed
DNS protocol and steals the query into the engine's internal resolver
(C.RuleActionTypeHijackDNS), which routes each query through its detour.
- netplane/nft.go: drop the prerouting :53 accept and the whole dnsnat
chain so LAN :53 is diverted by the normal tproxy catch-all into the
engine, where hijack-dns answers it. :853 DoT reject kept (forces :53).
With :53 now diverted, DNS also fails closed when the engine is down.
Verified E2E on the OpenWrt VM: with dnsmasq STOPPED, the netns LAN client
still resolves public names (only the engine's hijack-dns could answer);
a WAN :53 forward counter stayed at 0 while the shater divert counter
climbed — real egress was the engine's encrypted DoH to 1.1.1.1:443. No
plaintext client DNS reaches the WAN.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
A transparent proxy pins its tproxy inbound to a FIXED port, identical
across every apply, so the start-new-then-close-old swap ALWAYS collided
with the still-running old box on a live config change:
reconcile failed: start instance: start inbound/tproxy[in-lan]:
listen tcp4 0.0.0.0:12345: bind: address already in use
=> the new config silently never took effect. Only initial apply and
disabled->enabled worked (no old listener to clash with).
Fix: keep start-new-first (zero-downtime + old-instance protection when
ports don't clash), but on a listener bind conflict against a running old
instance, fall back to close-old-then-start-new (applyCloseFirst): close
the old box to free the port, build+start a fresh box for the new opts;
the fail-closed nft kill-switch covers the brief gap. If the fresh box
can't come up, restore the previous config; if restore also fails, leave
the engine stopped (instance=nil, never a closed box) — kill-switch keeps
the LAN safe. isAddrInUse matches both errors.Is(EADDRINUSE) and the
error text (box.New may flatten the errno); compiles cross-platform.
Verified on the OpenWrt VM: edit a running config + SIGHUP now logs
'reconcile OK (changed=true)' with the engine hash advancing and NO
'address already in use'.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
The inet-shater data plane relied entirely on TPROXY delivering to the
engine socket. nftables `tproxy` with NO listening socket returns
NFT_BREAK: it aborts its own rule before the trailing `meta mark set
0x2000 accept`, so the packet is left UNMARKED, falls through prerouting
policy accept, reaches the forward hook, and fw4 masquerades it to WAN.
Result: with the engine dead and kill_switch=closed, LAN clients LEAKED
straight out the WAN (confirmed on the VM: wget succeeded, conntrack
showed the flow SNAT'd to the WAN IP). kill_switch=closed only set the
engine's internal route.Final=block, which is moot when the engine is
down — there was no nft-level fail-closed layer.
Fix: in closed mode the forward chain now drops LAN-ingress traffic that
reaches it bound for a public dst (an escape, since diverted traffic is
delivered locally and never traverses forward), after accepting mgmt /
interface+tunnel egress marks and LAN-to-LAN / link-local so the LAN and
router keep working. Open mode still falls through (documented fail-open).
Folds in the old ipv6-off-closed case. Regression test asserts the v4/v6
drop is present in closed mode and absent in open mode.
Verified E2E on the OpenWrt VM (netns LAN client -> tproxy -> engine ->
AmneziaWG WARP exit): engine up, client egresses 104.28.212.73 warp=on
(direct WAN = 45.131.214.140 warp=off); SIGKILL the engine -> client is
BLOCKED (no WAN leak, no SNAT'd conntrack), whereas before this fix the
same scenario leaked to WAN.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
ParseUCIExport ignored the UCI section name for inbound/subscription/node/
group/chain/egress/ruleset/rule, honoring only 'option name'. A named
section like `config node 'ss1'` therefore yielded an EMPTY name, so
generate could not resolve a rule target 'node:ss1' and silently skipped
the rule — under kill_switch closed that BLOCKS the traffic instead of
proxying it. Now every named section falls back to the section name
(firstNonEmpty(option name, section name)); an explicit 'option name'
still wins. Matches preset/profile/resolver, which already did this.
Regression test covers all three forms for node/inbound/group.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
shaterd run is the single procd-supervised process that holds the one
box.New engine + the inet-shater data plane; every other verb is a
short-lived process that SIGNALS it (SIGHUP or the unix control socket)
and never builds a second engine.
- shater/apply: Applier drives model->generate->engine swap->netplane
(nft+routing+sysctl) under a cross-process flock, fail-closed (engine
error aborts before netplane; netplane error keeps the kill-switch up).
Reconcile (SIGHUP), honest Teardown (SIGTERM), commit-confirm
Snapshot/Confirm/ArmRollback/Rollback, ACTIVE_FLAG gating. flock.go
no-op default + flock_unix.go syscall.Flock override (cross-platform).
- shater/cmd/shaterd: run/migrate/reconcile/apply/confirm/rollback/
status/nodes/stats + sub|ruleset|schedule Phase-2b no-op stubs.
Pidfile single-owner guard; reconcile cold-start no-op (fork-storm
guard); control socket at /var/run/shaterd.ctl for reply-bearing verbs.
Verified on the OpenWrt x86_64 musl VM with the D9 router tag set
(static ET_EXEC, no PT_INTERP): daemon stays inert while globals.enabled=0
(no inet-shater table), applies on SIGHUP (table + fwmark 0x2000->shater +
loop-guard mark 0xff + tproxy divert appear, status all true, engine
hash set), and does an honest teardown on SIGTERM (table/rule/active-flag/
pidfile all gone, process exits 0).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Generate(m)/GenerateWithWarnings(m) build a full option.Options from the
neutral model: outbounds (share-link + AWG-endpoint via shater/parse),
urltest/selector groups, tproxy/mixed inbounds, route rules with the
kill-switch Final gate (closed->block, open->direct), and DNS. Every
inbound/egress carries the netplane loop-guard RoutingMark.
Verified on the OpenWrt x86_64 musl VM: engine.New().Apply (box.New +
Start) accepts AND starts all 7 cases — ss+tproxy+killswitch-closed,
AmneziaWG endpoint, 2-node urltest group, killswitch-open->direct, all
reachable share-link protocols, white-box hy2/tuic/shadowtls mapping, and
bad-node-skipped. RoutingMark validates only on Linux, so the engine
suite is //go:build linux and runs on the VM; Windows/macOS get build+vet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
DPI-bypass stays a per-ruleset egress choice, never a global toggle.
Evaluated zapret (NFQUEUE packet plane) vs ByeDPI (local SOCKS desync
proxy); chose ByeDPI because it *is* an egress and composes with our
routing model with zero conflict against the verified inet-shater TPROXY
plane. zapret explicitly rejected. Native tls_fragment/spoof (already
compiled in) stay as free complementary egress presets.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Defines the in-process daemon model: 'shaterd run' owns the box, CLI verbs signal
it (SIGHUP=reconcile, SIGTERM=teardown); shater/apply orchestrates generate+engine+
netplane under flock; openwrt/shater-core supervises shaterd. MVP scope for the gate.
Engine drives sing-box in-process (D11): Apply builds+validates via box.New,
swaps atomically (start new, close old), gates on a config hash (no churn on
unchanged apply), keeps the old instance running if a new config fails validation,
and supports Rollback to last-good. Instance()/Hash() expose state for stats later.
Integration test verifies apply/no-op/swap/invalid-keeps-old/rollback/close.
No go.mod changes.
Full map of the v0.1 xrayctl/shater-core internals + the sing-box option surface,
what ports verbatim (nft/routing plane, UCI model, parsers, subsystems) vs. what is
rewritten (generator, DNS, box lifecycle, stats), the v0.2 package layout, and the
wave plan. Authoritative reference for all Phase 2 agents.
Records the Phase 2 architecture: embed the engine in-process (apply = atomic
instance swap), rewrite only the generator (xray JSON -> sing-box options),
port parsers + nft/routing near-verbatim. Ship one binary shaterd.
AmneziaWG 2.0 proven E2E vs live Cloudflare WARP (warp=off->on through tunnel);
embedding via box.New verified on VM; router musl build ~9-11MB UPX. Next: Phase 2.
shater/cmd/shater-proto: drives the engine via the library API (box.New /
Start / Close) from our own Go main — the pattern the control-plane reuses.
Builds an option.Options in code (mixed inbound + direct outbound), proves the
data path with an in-process socks5 probe, and validates a tproxy+shadowsocks
variant through box.New. Verified musl-static on the x86_64 OpenWrt VM (egress
via the embedded engine), arm64 build proof, go vet clean. No go.mod changes.
Key API notes captured for Phase 2:
- ctx = include.Context(service.ContextWith(bg, deprecated.NewStderrManager(...)))
- option.Inbound/Outbound.Options MUST be a POINTER to the concrete struct.
- box.New constructs+validates every adapter; a single non-special outbound is
auto-selected as default route.
Orchestrator-laid foundation per CLAUDE.md: lightweight single-bundle SPA to be
embedded in the forked binary and served by the daemon on its own port.
- tokens.css: Faceplate tokens ported verbatim from docs-shater/DESIGN.md
(light/dark, prefers-color-scheme default + data-theme override both ways),
mono instrument voice, tabular-nums, focus-visible, reduced-motion floor.
- Placeholder App shell (component library <Faceplate>/<Module>/<Toggle>/<Led>/
<SegMeter>/<QueryLog> + pages land in Phase 3, delegated).
Builds clean: 145K dist (46KB gzip JS), tsc --noEmit passes.
Phase 1 VM findings: canonical LX_TAGS links glibc (naive/cronet/purego dlopen)
and won't run on musl OpenWrt; router build drops with_naive_outbound,with_purego
for a fully-static binary. Ship UPX-lzma (~9-11MB from ~40MB raw).
Upstream sing-box-lx already ships a docs/ mkdocs site; keep our project docs
separate and unambiguous in docs-shater/ (parallels upstream's docs-lx/).
Updated all references in README.md, CLAUDE.md, CONTEXT.md, ARCHITECTURE.md.
Foundation pivot. The complete, working, VM-verified xray-based project is
preserved on the `v0.1` branch; `main` is reset to a docs-first scaffold for
v0.2, which will be built as a FORK of sing-box-lx with our control-plane,
DNS filter, stats and admin panel embedded in the one binary.
- Preserve everything on branch v0.1 (pushed).
- Remove the v0.1 implementation + old design docs from main (recoverable from
v0.1); keep LICENSE, .gitignore, .gitattributes, dist/shater-feed.pub (feed
signing key 5ac4b177689cb8e0 carries over).
- License -> GPL-3.0 (sing-box is GPL-3.0).
- Add full project context so it survives compaction:
docs/CONTEXT.md (start here), DECISIONS.md, ARCHITECTURE.md, ROADMAP.md,
FEATURES.md, and a new README.
Engine/UI decisions (see docs/DECISIONS.md): fork sing-box-lx (AmneziaWG 2.0 +
broad protocols, GPL-3.0, library-first) and embed the whole product for tight
integration; keep the fork maintainable via an additive overlay (shater/, panel/,
openwrt/) rebased on upstream tags. UI = thin LuCI launcher + a separate admin
panel served by the daemon, entered via a short-lived token minted in the
authenticated LuCI session. Do NOT write a proxy engine from scratch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
Promote v1.14.0-lx.3-rc.2 to a stable (Latest) release. Functionally
identical — no runtime change since the rc, only this changelog entry.
Payload: DNS command-multiplex (rc.1) + AWG re-graft onto wireguard-go
v0.0.5 with upstream merge (rc.2). Device-verified by owner.
Merges 14 upstream commits including L3-forwarding support (which bumped
wireguard-go v0.0.3->v0.0.5, already re-grafted in the prior commit),
snell protocol, bridge outbound, flow-tracking/sniff improvements, and
DNS/dialer fixes.
lx conflict resolutions:
- protocol/wireguard/endpoint.go: took upstream's new flow API
(PreMatchFlow/PortAddresses/PortMTU/AttachReturn/DetachReturn/JudgeFlow),
dropped our old PrepareConnection/NewDirectRouteConnection. SPEC 020
idle-suspend wake guard (resumeOnDial) moved to WritePackets — the single
point every L3-forwarded packet transits, incl. established flows that
bypass DialContext.
- adapter/outbound.go: kept lx IdleSuspendable/ReachabilityInvalidator,
restored 'time' import dropped by auto-merge.
- go.mod/go.sum + test/: took upstream dependency bumps (tailscale, sing,
sing-tun); wireguard-go stays v0.0.5 with local submodule replace.
Green: full sing-box CLI with LX_TAGS (Go 1.24.7), libbox, wireguard/
adapter/dns/daemon packages, transport+protocol/wireguard tests, AWG
config validation.
SPEC.md rewritten to current (multiplex) architecture, no chronology.
HISTORY.md captures v1 standalone class-error, the field bug, rejected paths.
New project rule (README + CONSTITUTION 3.2): SPEC.md = current state first,
chronology/rationale of architecture changes go to HISTORY.md.
DNS stream now runs on the shared c.ctx via dispatchCommands (CommandDNS),
mirroring handleConnectionsStream: auto-reconnects with Connect(), dies with
the client, no per-stream Close()/OnError. Removes DnsQueryHandler,
DnsQuerySubscription and the standalone SubscribeDNSQueries client method.
DnsQuery/DnsAnswer/dnsQueryFromGRPC unchanged.
Add DNS as a first-class multiplexed command, uniform with CommandConnections:
- command.go: CommandDNS constant (next in iota)
- command_client.go: case CommandDNS in dispatchCommands; DNSIncludeAnswers
option field (like StatusInterval); WriteDNSQuery in CommandClientHandler
All three upstream touch-points wrapped in // lx:begin dns / // lx:end dns
(CONSTITUTION 3.3). SPEC 018 v2.
Runtime detour/selector rings crash the core (fatal stack overflow via
unbounded DialContext recursion); static rings are already rejected at start
by lintOutbound. Worked out the full event model (E1-E5) and topologyMu race
linearization, then adversarially verified it (7-agent workflow): deadlock and
false-positive attacks HOLD, but TOCTOU BREAKS — even a correct core guard is
not airtight without also covering the endpoint manager, Manager.Remove,
history side-channels, and the pointer-vs-tag graph divergence after a runtime
Create. Owner decision (2026-07-06): protection lives at the UI level (LxBox
validates before SelectOutbound); core stays a minimal delta to upstream.
No core code changed. SPEC is a design record + Roadmap row (status DEFERRED).
Split the lx go vet step into two passes so every lx-owned package keeps
the full analyzer set; only daemon/ and experimental/libbox/ (upstream
TriggerDebugCrash/TriggerGoPanic) drop the unsafeptr check.
Producer run populated musl-toolchain-cache with 4 arch assets; restore path
validated locally (asset name, gh download, tar layout under naiveproxy/src).
Status -> C, Roadmap updated.
snapshot.debian.org intermittently 503s during the musl sysroot build and
blocks releases (v1.14.0-lx.2-rc.1 failed twice on it). actions/cache also
misses across tag builds (ref-scoping). Add a producer workflow that uploads
the built toolchain to a musl-toolchain-cache release, and a restore step in
lx-release.yml that pulls it on cache-miss before falling back to
snapshot.debian.org. Both workflows are lx-owned; zero upstream diff.
Full audit of the LX delta (10 axes, adversarial verification): 32 findings,
27 confirmed, 24 fixed on branch lx-spec022-audit-fixes, 3 skipped by design
(#12/#17/#18). Records #19 resolution (SPEC 013 test kept — upstream ships none).
Exact-string match on quic-go's formatted CRYPTO_ERROR silently disabled the
Cloudflare Access hint on any error-text reformat. Match the inner
"tls: access denied" alert as a substring instead.
s3 pads only cookie-reply messages (paddings.cookie); s4 pads every transport
data packet (paddings.transport). Folding s3 into the MTU budget dropped MTU and
warned spuriously for an atypical s3>s4 config. Fix calc, comment, warning, docs.
The 'if c.reader == nil' fast path read reader without synchronising against the
RoundTrip goroutine's write (a data race -race flagged). Always receive on
created first; on an already-closed channel that is effectively free.
calculateIPv4Checksum summed a fixed 20 bytes, producing a wrong checksum after
TTL decrement when the header carried options (IHL>5). Take the real IHL*4 span;
validate IHL against the buffer before the read. Adds an IHL=6 test.
Pool() keyed URL-test history by the raw slot tag; history is stored under
RealTag(detour), so a nested-group member always reported Delay=0 in GetPool.
Resolve the slot tag and read under RealTag, matching seedPool/rebuildPool.
questionCache returned a fresh cached response without logging/emitting; only
the stale (optimistic) branch emitted. Mirror the Exchange path: fresh->cached,
stale->optimistic. Gated by HasSubscribers, so zero cost with no profiler.
SuspendAmneziaWG left idleAsleep untouched, so an endpoint idle-suspended
BEFORE the guard fired could be resurrected by the next dial (resumeOnDial
keys only on idleAsleep) — reintroducing the AmneziaWG-over-WireGuard kernel
hang the guard exists to prevent. Now clears idleAsleep under resumeMu so
the guard is ordered against a concurrent wake.
sendConnect's ReadFrame loop blocked forever on a peer that completes
TCP+TLS but never returns the CONNECT HEADERS, wedging the outbound under
o.runMu (Close hangs too). A ctx watcher now trips tlsConn's deadline on
timeout/cancel and is joined before the long-lived readLoop starts.
The main README's 'Features & status' table and Feature-configuration section
listed XHTTP/AWG/observability/round_robin but not the SPEC 021 MASQUE
(CONNECT-IP / WARP) outbound. Add it in both README.md and README.ru.md:
intro line, a Features table row (device-verified on Wi-Fi + LTE, h3/h2), and
a short config example with the network=transport footgun and the h2 fallback
for UDP:443-filtered networks. Links to docs-lx/lx-config §4 and SPECS/021.
The user-facing config guide covered XHTTP/AWG/urltest but not the SPEC 021
MASQUE (CONNECT-IP / WARP) outbound — only the internal SPECS/021 had it.
Add a full section (§4, renumbering Observability→§5, Validate→§6): feature
table row, field table, h3/h2 example, the network=transport footgun, dns-block
requirement, h3-vs-h2 guidance (UDP:443 filtering, cold-start), device-verified
status, and a link to SPECS/021/CONFIG.md. RU mirror kept in sync.
The 1.14 merge (6b63cee4) brought daemon/managed_service.go and
experimental/libbox/debug.go into vet scope; both crash Go ON PURPOSE via
*(*int)(unsafe.Pointer(uintptr(0)))=0 (TriggerDebugCrash/TriggerGoPanic),
which vet's unsafeptr analyzer flags. They are upstream files we don't edit,
and no lx-owned file uses unsafe at all, so disable just that analyzer.
Fixes the red lint job on every push since the merge.
A failed dial is an actionable error and should be visible where the success
(tunnel established, INFO) is — the LxBox core-log forwarder only surfaces
INFO+, so a DEBUG failure was invisible on-device. WARN makes established/
failed a symmetric, forwardable pair. The detailed dial phases (establishing,
udp-socket-up) stay DEBUG.
Promotion of the rc.1..rc.22 series to a non-prerelease tag (publishes as
Latest). Functionally rc.22 + the linux-mips-softfloat asset (#6). Section
header matches the tag exactly so lx-release.yml extracts it into the notes.
Rides out transient snapshot.debian.org 503s in the linux-musl toolchain
download/keyring steps (rc.22 failure cause). No code change; publish still
requires all builds.
rc.22 linux-musl jobs failed on 'get-clang.sh: 503 No healthy backends' from
snapshot.debian.org (transient mirror outage). get-clang.sh retries internally
but back-to-back, so a whole outage window fails all attempts. Wrap the two
Debian-fetching steps with an external retry + growing backoff:
- Download Chromium musl toolchain: 5 attempts, 30→60→120→240s
- Regenerate Debian keyring: 4 attempts, 20→40→80s
Conservative: publish still needs all builds (a real musl breakage still blocks
the release, only transient mirror flakes are ridden out). No code change.
Requested in #6 (Atheros AR9344). Chromium/cronet has no big-endian MIPS
toolchain, so the musl+naive path is impossible for this target; add it to
the plain cross-compile matrix instead as a pure-Go CGO_ENABLED=0 build —
statically linked (runs on musl/OpenWrt as-is), with with_naive_outbound
and with_purego dropped (purego has no mips port either). Everything else
matches the desktop tag set.
Verified locally: GOOS=linux GOARCH=mips GOMIPS=softfloat build with the
reduced tag set compiles clean; `file` reports ELF 32-bit MSB MIPS32,
statically linked.
Adds establish/handshake/success debug logs to the MASQUE dial path so a
stuck tunnel is diagnosable from /logs/core alone (motivated by the LxBox
§130 device case: h3 hung in QUIC handshake because inbound UDP:443 was
Log the tunnel-establish phases so a stuck dial is diagnosable from
/logs/core alone, without a goroutine dump:
- "establishing <h3|h2> tunnel to <server> (sni=...)" on start
- "udp socket up, starting QUIC handshake" (h3) — pinpoints whether a
hang is the socket or the handshake (inbound UDP:443 filtered → our
ClientHello left but no ServerHello came back)
- "tunnel established" on success, "tunnel failed: <err>" on failure
Motivated by a live device case (LxBox §130): h3 hung in the QUIC
handshake because inbound UDP:443 was filtered by the network while AWG
(UDP on a non-443 port) worked — invisible in logs before this, required
a pprof dump to locate. h2 (TCP:443) is the fix there.
Refs: SPEC 021
Complete masque outbound config reference verified against code:
full JSONC template (all params), minimal config, per-field table
(masque-specific + inherited DialerOptions), profile matrix, value
formats (duration/keys/ip), start-time validation, and common footguns
(network=transport not tcp/udp, dns block required, exit-IP changes on
reconnect, keepalive vs idle_timeout).
Rework the tunnel lifecycle around a *session (device + ipConn + closer +
ctx + activity counter), guarded by runMu with a generation guard.
- C1 (was HIGH): a dropped or suspended tunnel is now rebuilt on the next
dial. Previously 'running' latched true, so after the tunnel died every
DialContext short-circuited and dialed into a dead stack — permanent
blackhole. teardownSession clears o.sess so ensureSession rebuilds.
- C2: teardownSession closes ipConn, which unblocks the paired pump parked in
a blocking read (context cancellation alone can't interrupt it); no leaked
goroutine, no zombie half-open tunnel. Idempotent via sync.Once.
- B1: idleWatcher suspends the whole tunnel (gVisor netstack, pumps, QUIC
keepalive) after idle_timeout of no traffic; the next dial rebuilds it —
near-zero resident cost when idle.
- B4: idle_timeout (default 5m) and keep_alive_period (default 30s) are config
options; negative disables. A5: fail-fast mtu<=16000 on h2.
- D2: drop the dead congestion_control option field.
lifecycle_test.go covers the generation guard, idempotent teardown and
close-guard under -race.
Refs: SPEC 021 audit B1/B4/C1/C2/A5/D2
- A4/A5/B5: cap peer-declared capsule payloadLen at maxCapsulePayload (64KiB)
to prevent int-overflow/OOM on a hostile length; shrink recvCh 64->8 and the
receive window 1GiB->8MiB so real HTTP/2 flow-control provides backpressure
instead of unbounded RAM.
- B2: reuse a per-conn scratch for the outgoing capsule frame (tx pump is the
sole writer; writeData flushes before returning) instead of allocating per
packet.
Refs: SPEC 021 audit A4/A5/B2/B5
- A1: a single malformed/empty inbound datagram (unparseable context-ID
varint) no longer tears the tunnel down — drop-and-continue like the
sibling context-ID!=0 / bad-payload cases.
- A3: snapshot the IP header before composeDatagram mutates TTL/checksum, so
the ICMP 'packet too big' reply quotes the original datagram (RFC 1191/792).
- A2: drop the redundant second status check in ConnectTunnelH3 (2xx already
validated in dialCONNECTIP); removes the unused responseInfo carrier.
- B3: reuse a per-Conn scratch for the outgoing datagram (contextID + packet)
instead of allocating per packet — safe, quic-go copies the slice before
SendDatagram returns and the tx pump is the sole writer.
Refs: SPEC 021 audit A1/A2/A3/B3
h2 (network: h2) now works on live Cloudflare WARP (warp=on, http/2).
The high-level HTTP/2 clients can't drive WARP's CONNECT-IP: stdlib
http.Client.Do(CONNECT) uses classic tunnel semantics (400), and
x/net/http2's RoundTrip refuses because WARP never advertises
SETTINGS_ENABLE_CONNECT_PROTOCOL ("extended connect not supported by
peer") — the same RFC-noncompliance it shows on h3.
Drive the h2 connection manually with x/net/http2's public Framer + hpack
(both already deps): own client preface, SETTINGS, WINDOW_UPDATE, one
HEADERS frame, DATA frames carrying capsule DATAGRAM frames. This skips
the peer-settings gate. WARP h2 is a *plain* CONNECT (:method+:authority)
keyed off the cf-connect-proto header, NOT an extended CONNECT with
:protocol (that got PROTOCOL_ERROR). No http fork, no new dependency.
Also resolve domains before L3 dial was already in; this commit adds the
h2 framer, capsule-reassembly unit tests (across DATA-frame boundaries),
and updates SPEC/TEST_PLAN — risk #1 now closed.
Refs: SPEC 021 TEST_PLAN.md
The gVisor userspace stack operates at L3 and panicked ("As4 called on IP
zero value") when handed a domain destination. Resolve via DNSRouter before
dialing (as the WireGuard endpoint does): DialContext/ListenPacket now do
Lookup + N.DialSerial/ListenSerial for domain destinations, and reject
invalid non-domain destinations.
Live-tested against Cloudflare WARP with real registration key material:
- h3 (CONNECT-IP/QUIC): WORKS — cdn-cgi/trace returns warp=on, Cloudflare
edge IP, clean connection teardown, tunnel reuse.
- h2 (CONNECT-IP/HTTP2): WARP responds 400 — stdlib net/http CONNECT
semantics differ from WARP's expected extended-CONNECT authority/headers
(SPEC risk #1, materialized). Deferred to phase 2. Documented in TEST_PLAN.
Refs: SPEC 021 TEST_PLAN.md
The "GRO off + batch 8" idea (a global alternative to Down/Up) was measured on-device
and REJECTED, for three independent reasons (SPEC.md §14):
1. Wrong holder — the main android RAM holder is device.pool.messageBuffers
(PreallocatedBuffersPerPool=4096 × ~64KB ≈ 100MB), which does NOT depend on
BatchSize; the batch-sized bufsArrs held only ~14MB. Shrinking batch wouldn't
have touched the ~100MB.
2. Not deliverable — the LX_WG_NO_GRO env switch never reaches Go's os.Getenv on
Android (wrap.<pkg> prop shows in /proc/environ but not in the runtime's env
snapshot), forcing a hardcode.
3. Fragile — hardcoded batch=8 crashed at start (SIGABRT): device.BatchSize()=
max(bind,tun) clamped back to 128 via the TUN offload while msgsPool was 8, so
Send sliced out of range. Coherent only by also gating TUN offload across three
submodule layers.
Down/Up (rc.19) stays the only viable mechanism. Brings the experiment folder
(protocol + device heap snapshots + RESULT) into lx-1.14 for the record; the
experiment CODE stays on the lx-1.14-nogro-* branches, not merged.
Sync SPEC 002 with the code (commit c0bbb1c5): GET on a non-packet-up node no
longer hard-errors — it falls back to POST + WARN so one bad subscription node
doesn't fail the whole config. Updated the mode-gate wording in SPEC.md §verif,
PARAM_MAP.md (full rationale + the old error text it replaces), URL_PARSING.md
table, and IMPLEMENTATION_REPORT.md. header/cookie uplink outside packet-up
stays a hard error (no safe default).
A subscription node sometimes ships uplink_http_method=GET on a non-packet-up
node (auto/stream-up/stream-one). GET can only carry the uplink in packet-up
(other modes put the body in the request, which GET has none of), so the strict
check rejected it — failing the ENTIRE config over one bad outbound in a large
subscription (observed: initialize outbound[361] ... can be GET only in
packet-up mode → whole tunnel won't start).
Fall back to POST (the safe default that works in every mode) and log a WARN
instead of erroring, so the rest of the config still loads. POST is what the
node should have used; the fallback just makes one malformed remote node
self-healing rather than fatal. Kept strict for packet-up (GET honoured there).
Uses log.StdLogger() for the warning (the client transport layer gets no
logger in its constructor signature; threading one through would touch upstream
signatures). Test: TestUplinkGetFallsBackToPostOutsidePacketUp; removed the GET
case from TestValidationRejections. Verified on a real binary — check on a
stream-one+GET config now warns and passes (exit 0) instead of FATAL.
The base-version step (c4fd73cd) adds an `upstream` remote (SagerNet/sing-box)
so git-describe can see the v1.14.0-alpha.* tags. But `gh release create` without
--repo resolves the target repo from the remotes and picked `upstream` →
HTTP 403 "Resource not accessible by integration" against
api.github.com/repos/SagerNet/sing-box/releases (the token has no rights there).
This is why rc.19's builds all succeeded but publish failed, while rc.18 (before
the upstream remote existed) published fine.
Pin --repo "${{ github.repository }}" so publish always targets this fork
regardless of what remotes the earlier steps added.
rc.19 gates idle-suspend behind with_lx_idle_suspend (mobile-only) and records the
on-device Android verification: suspending 8 idle+unreachable WG endpoints freed
134MB of bufsArrs live heap (223.9→89.9MB, recv-workers 18→2), matching the
~8.4MB/worker model — ~10x the desktop delta, on the platform the feature targets.
Adds ANDROID_RESEARCH/live-baseline/ — a full pprof snapshot of a real production
config with the feature OFF (263MB bufsArrs, 56% CPU on GC at idle) and its ON
"after" counterpart, closing the RESEARCH.md device gap end-to-end (buffer pool
and GC cost measured together, not inferred). Credentials scrubbed.
Idle-suspend frees the recv-worker bufsArrs, which are ~8MB each only where
BatchSize=128 (Android/Linux) — on desktop BatchSize is small and the feature
saves almost nothing. Make that platform scope explicit in the build instead of
running the tick everywhere.
The idle-suspend tick now compiles only with the new `with_lx_idle_suspend` tag,
baked into the mobile AAR (build_libbox sharedTags) but NOT the desktop LX_TAGS.
Without the tag, a config that sets route.lx_idle_suspend fails fast at start
("rebuild with -tags with_lx_idle_suspend (mobile-only feature)") rather than a
silent no-op. The gate is a single function: reachability_lx.go carries the tick
under the tag, idle_suspend_stub_lx.go is the no-tag stub that errors, and
reachability_common_lx.go keeps InvalidateReachability (needed by the group
interface in every build). The dial hot path (resumeOnDial/stampActivity) and the
upstream group files are untouched — without the tick, idleAsleep is never set, so
resumeOnDial always takes its fast path.
Adds stub unit tests (option set → error, unset → no-op). Both build variants and
the full route/wireguard/group suites are green; gofmt/vet clean; desktop lx-check
passes without the tag. Docs (lx-config.md + ru, SPEC.md §3/§10) describe the tag.
The base-version derivation used `git describe --match v1.14.0-alpha.*` as the
primary source, but actions/checkout only fetches THIS repo's tags — the alpha
tags are SagerNet/sing-box (upstream) tags, absent in the CI clone. So `git
describe` found nothing and silently fell to the subject-grep fallback, which
resolves alpha.36 (alpha.37 was merged in a commit whose subject omits the
number). That's why rc.17/rc.18 notes shipped "base alpha.36" while a local
clone with upstream tags gets 37.
Fetch just the upstream v1.14.0-alpha.* tags before git describe so the primary
graph-based path works in CI. Both hand-fixed on the published releases; this
makes the next tag correct automatically.
Full write-up of the Android device run (CPH2411, Android 15, rc.18) in a
dedicated ANDROID_RESEARCH/ subfolder: README (report), METHOD (reproducible
procedure), RESULTS (per-scenario + heap A/B), and artifacts/ (raw evidence:
lx idle log lines, goroutine dumps, pprof heap .pb + top renders). No access
credentials anywhere.
Headline, now measured on the target platform: PopulatePools.func3 (the
bufsArrs holder from RESEARCH.md) inuse_space 223.93 -> 89.89 MB (-134 MB /
-60%), recv-workers 18 -> 2, on suspending 8 of 9 WG endpoints. = 16 workers x
~8.4 MB (BatchSize=128), matching the source model, ~10x the desktop RSS delta.
RESEARCH.md status + SPEC.md §12/§13 updated: the Android heap A/B gap is
closed (only the battery A/B remains deferred).
Device run on CPH2411 (Android 15, rc.18) via the LxBox app Debug API.
9 WG endpoints (1 real WARP reachable + 8 synthetic unreachable),
lx_idle_suspend=30s. All behaviors confirmed on-device: suspend fires,
reachable final stays up, wake-by-dial, no-flap, kill-switch.
Headline: PopulatePools.func3 (the bufsArrs holder from RESEARCH.md)
inuse_space 223.93 to 89.89 MB (-134 MB / -60%), recv-workers 18 to 2.
= 16 freed workers x ~8.4 MB (BatchSize=128), matching the model, ~10x
the desktop RSS delta. Closes the Android device-verification gap.
Device-verified idle-suspend: idle AND unreachable WG/AWG endpoints go Down to
free their recv-worker bufsArrs (the Android GC-heat holder), waking on next dial.
Opt-in via route.lx_idle_suspend; off by default. Notes cover the reachability
walk, the GRO reason for Down-over-smaller-batch, the concurrency fixes, and the
2026-07-01 live-run results (recv-workers 16→0, RSS −31%).
Selectively brings idle + unreachable WireGuard/AmneziaWG endpoints Down,
freeing their recv-worker bufsArrs (the measured GC-scan heat holder on
Android) and stopping their per-peer timers (battery), then wakes them lazily
on the next dial. Off by default (lx_idle_suspend absent/0 = zero overhead).
Includes the fix for the shipped tick iterating the wrong manager (it never
reached any endpoint — the feature was inert on a live box), the full
reachability walk (final/rule/selector Now/urltest pool/detour, event-driven
cached), and the rewritten as-built spec (SPEC.md) + research doc (RESEARCH.md)
+ test plan.
Live-verified: suspend/wake/probe-wake/re-sleep/no-flap/kill-switch across
selector, urltest pool, nested groups, AWG-guard, and the real production
config; resource A/B recv-workers 16->0, RSS -31% on desktop. 29 unit tests,
adversarially checked.
The idle-suspend feature is implemented, the tick bug is fixed, and every
reachability node type plus suspend/wake/probe/no-flap/kill-switch and the
resource A/B (recv-workers 16→0, RSS -31%) are live-verified. Reflect that in
the docs and give the folder clean roles:
- SPEC.md (was SPEC_idle_suspend_lever.md): rewritten from scratch in Russian
as the as-built implementation spec — Down/Up model, reachability walk +
event-driven cache, endpoint-side suspend/wake, the tick bug and its fix
(§11), full test coverage (§12, 29 units named), and what is deliberately
deferred (§13: Tier B netstack teardown, keys-safe BindUpdate path, on-device
battery/heap measurement). Old Tier-A "light sleep" design (never shipped)
removed.
- RESEARCH.md (was SPEC.md): the diagnostic root-cause doc keeps its unique
on-device heap A/B proof (holder = recv-worker bufsArrs) — renamed so its
role (research, not implementation spec) is unambiguous.
- TEST_PLAN_idle_suspend.md: all pass criteria checked, §RESULTS + edge-case
matrix + wake-latency series (cold ~50ms / warm ~36ms, +14-21ms ≈ 1 handshake;
far-server caveat) + production-config run.
Cross-references and section numbers updated across all three files.
The shipped idle-suspend tick (c55cf11e) iterated r.outbound.Outbounds(),
which never lists WG/AWG endpoints — they live in the endpoint manager.
outbound.Manager.Outbounds() returns only m.outbounds; the endpoint
fallback exists for Outbound(tag) lookups, not the iteration. So the tick
never reached a single IdleSuspendable and the feature was inert on a live
box (0 suspends over minutes idle), despite green unit tests that exercised
the walk and the per-endpoint decision only in isolation.
Fix: Router pulls adapter.EndpointManager from ctx (service.FromContext, no
box.go change — it is already registered there) and the tick body moves into
suspendIdleEndpoints(), which scans both r.endpoint.Endpoints() (where the
IdleSuspendables actually are) and r.outbound.Outbounds() (kept for a future
non-endpoint IdleSuspendable). Nil-guarded for the stub case.
Tests: new route/idle_tick_endpoints_lx_test.go drives the tick through a
stub endpoint manager — fails pre-fix (wg-1=0 wg-2=0, tick blind to
endpoints), passes after. Adds reachability walk tests for the production
topology this fix enables (nested selector→urltest pool, dual-path dedup,
dormant nested subtree) and the AWG-guard idle invariant. All adversarially
checked. See SPECS/020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md §11.
Code recon of the wireguard-go submodule disproved SPEC.md's original PRIMARY
lever (shrink StdNetBind.BatchSize() 128→8). It cannot be done without breaking
GRO receive:
- GRO-rx is ENABLED on android (UDP_GRO set with no android gate,
controlfns_linux.go:90-104; SPEC 010 gated only GSO-tx, not GRO-rx) → rxOffload=true.
- The GRO path splits one coalesced packet into up to 64 datagrams; readAt =
len(msgs) - IdealBatchSize/udpSegmentMaxDatagrams (bind_std.go:269) HARDCODES
IdealBatchSize=128, and getMessages() allocs a 128-slot array. Shrinking bufsArrs
to 8 either desyncs bufs(8) vs array(128) → OOB panic, or overflows the split
("splitting coalesced packet resulted in overflow", bind_std.go:565). GRO can't
be disabled (needed for download throughput, §010).
So the old claim "packet loss excluded, array just shorter" was wrong for the GRO
path. Lever 1 (and lever 2, which inherits the same idle-socket GRO problem) are
rejected. PRIMARY becomes lever 3 — Down idle+unreachable devices: BindClose ends
the recv-workers and frees bufsArrs whole, while the active node keeps batch=128 so
its GRO is intact. The "most expensive fallback" is in fact the only viable lever.
Updated: status line, the lever-candidates section (struck lever 1, promoted lever 3),
the fix-logic section (renamed + rewritten around Down), verification, and residual
risks (the fast-channel risk is gone; the new risk is handshake-on-wake). Implemented
on lx-spec020-idle-suspend; see SPEC_idle_suspend_lever.md §13 + TEST_PLAN.
Add TEST_PLAN_idle_suspend.md — build/config/commands/pass-criteria to
device-verify the shipped idle-suspend on a real run: suspend fires for
idle+unreachable WG/AWG endpoints, reachable ones never suspend, wake-on-dial,
the bufsArrs memory drop (pprof heap), and no flapping. Uses the user's
WARP/AWG + plain-WG nodes. Link it from SPEC_idle_suspend_lever.md §13.
Reachability is now recomputed ONLY when the active routing tree changes, not
every idle tick. Per user direction: events decide WHO is reachable; the timer
only checks WHEN (last-activity comparison).
- adapter.ReachabilityInvalidator: narrow interface (not folded into the large
adapter.Router), registered into ctx in box.go, pulled by groups via
service.FromContext — no route<-group import.
- Router: reachMu/reachCache/reachDirty. InvalidateReachability() is a lock-free
atomic store (safe under any group lock — no lock-order cycle). reachableOutbounds()
recomputes the walk OUTSIDE the cache lock (the walk calls into groups that hold
their own locks), clears dirty BEFORE the walk so a concurrent event re-dirties
for next tick rather than being lost, publishes under RWMutex. Starts dirty so
the first tick (and every reload = fresh Router) computes.
- 4 invalidation sources: selector switch (selector.go after selected.Store),
legacy urltest auto-switch (urltest.go performUpdateCheck), and a balancer
onChange hook fired from setSlots — one hook covers all pool-rebuild call sites.
- idle tick: now one cached-map lookup + atomic idle compare per endpoint, no walk.
Design independently verified against source (no data race, no import cycle, no
deadlock — walk runs outside the lock). Builds + go vet + race-build clean.
Record the as-built decision so it is not re-derived:
- §13.1 what shipped (c55cf11e): idle-suspend via Down/Up, not light sleep
(source-verified holder is recv-worker bufsArrs; timersStop does not free it).
- §13.2 bind-swap investigation (the promised comment): BindUpdate resizes the
bind WITHOUT zeroing keys; key-zeroing lives only in Down/peer.Stop. So a
keys-safe wake is possible only while still Up — the shipped Down path can't.
- §13.3 the three reduced-bind paths (B=Down shipped / A=BindUpdate keys-safe /
Hybrid) with the GRO-off + max(bind,tun) gotchas.
- §13.4 recommendation: ship B, escalate to A/Hybrid only if device INFO logs
show handshake flapping hurts. Reduced-bind urltest wake deferred (low value
on path B since keys are already zeroed).
Selectively bring Down any WG/AWG endpoint that is idle past a threshold AND
unreachable from the active routing tree — freeing its recv-worker bufsArrs
(the dominant per-endpoint GC-scan holder), cutting the multi-WG heat. The next
dial through the endpoint wakes it (device.Up); wake pays a fresh handshake.
- option: route.lx_idle_suspend (Duration, 0/absent = off, kill-switch).
- route/reachability_lx.go: ReachableOutbounds walk — seeds = final + rule
outbounds, descend via selector Now(), urltest active pool (ActiveTags), and
static detour deps. Fresh walk per tick (no gen-cache: graph is tiny, tick is
~XX/2; a cache would need upstream-body invalidation hooks — not worth it yet).
- protocol/wireguard/endpoint.go: lastActivity/IdleSince, SuspendIfIdle (Down on
live->asleep CAS), resumeOnDial (stamp + lazy Up on dial). idleAsleep is kept
distinct from started so a guard-suspended endpoint is never idle-woken.
- transport/wireguard/endpoint.go: Resume() = device.Up() alongside Suspend().
- adapter: IdleSuspendable interface so the router tick iterates endpoints
without importing protocol/wireguard.
- route/router.go: idle tick (period max(XX/2, 5s)) started in PostStart,
stopped in Close.
- INFO log on each state transition only (edge-triggered): suspend / wake.
- group: URLTest.ActiveTags() exposes the whole active pool to the walk.
Builds clean, go vet clean. Reduced-bind urltest wake + bind-swap/keys
investigation land next.
Complete RU translation of docs-lx/lx-config.md, mirroring its structure:
the §0 exhaustive "every field at a glance" example, all per-section field
tables (XHTTP v1+v2, AmneziaWG 2.0, id/ip/ib masquerade, urltest balancer),
examples and the build section. JSONC code is preserved; only prose and
inline comments are translated. Intra-doc #anchors are re-pointed to the
Russian heading slugs (all 6 verified to resolve); the §0 example validates
as JSON.
Cross-link both ways (en ↔ ru) and re-point README.ru.md's four lx-config
links to the Russian version. File name follows the README.ru.md convention
(.ru.md, not -ru.md).
Add a §0 kitchen-sink config carrying ALL 52 lx-added fields in one place —
XHTTP transport (26), AmneziaWG 2.0 endpoint incl. id/ip/ib (21), urltest
round_robin balancer (5) — each with its default and allowed values inline,
and mutually-exclusive / server-ignored fields flagged. Sourced by reading
option/*.go directly (not the prior doc), so it is complete.
This also surfaced that §1 documented only 7 of the 26 XHTTP fields (the v1
set); fill in the 19 missing v2 fields (session/seq placement, uplink-data
placement, X-Padding obfs family, packet-up tuning, accepted-but-ignored)
as grouped tables. Fix three code-vs-doc disagreements the extraction found:
- `mode: auto` resolves to stream-one on Reality (not always packet-up);
- s3/s4 are AWG 2.0 junk-size params (not "cookie-reply/transport" junk);
- h1-h4 unset spelling includes "" as well as 0; x_padding_bytes framing.
The default wire shape is unchanged; all v2 fields are opt-in. §0 example
validated as JSON. Russian translation (lx-config.ru.md) to follow.
docs/ is an upstream-owned tree (it arrives wholesale from SagerNet on
every rebase). Our three downstream docs lived inside it — lx-config.md,
lx-changelog.md, lx-release-runbook.md — mixing fork files into the
upstream surface against CONSTITUTION principle #1 (thin layer / minimal
diff). Move them to a dedicated root-level docs-lx/ so the boundary
between our docs and upstream's is explicit.
- git mv preserves history.
- Updated every reference (docs/lx-* -> docs-lx/lx-*): README.md/.ru.md,
SPECS/{003,004,005,009,020,README}, transport/wireguard/endpoint.go
comments, lx-ci.yml, and lx-release.yml (the release-notes extractor +
fallback URL now read docs-lx/lx-changelog.md).
- Fixed the now-relative links inside the moved files that pointed at
upstream docs/ siblings: lx-config.md -> ../docs/configuration/outbound/
urltest.md; lx-changelog.md -> ../docs/changelog.md (x2).
Verified: all relative + external links resolve, both workflows are valid
YAML, the release-notes awk path is docs-lx/, go vet clean on the touched
package. No release feature — folds into the next tag naturally.
Secondary design doc alongside the authoritative SPEC.md, NOT a replacement.
SPEC.md has on-device proof (heap A/B) that the GC-scan holder is bufsArrs of
recv-workers and its primary lever is shrinking StdNetBind.BatchSize() — that
stays authoritative.
This companion works out SPEC.md's 'lever 3' (suspend inactive devices) in
detail: a light variant (per-peer timersStop, keypairs/socket kept live, cheap
wake without handshake) plus a reachability walk (final + rules + active
selector/pool choices, generation-cached) that decides which devices are idle
AND unreachable from the active routing tree.
Banner up top flags where this doc's source-reading diverged from SPEC.md's
measurements (it guessed gvisor netstack; SPEC.md measured bufsArrs) and notes
light-suspend does NOT free bufsArrs — only Down/BatchSize does. Use as the
fallback design for lever 3, not a competing primary.
Spec only — no code changes.
The subject-grep base-detection missed alpha.37: it was merged in a commit
titled "Merge upstream/testing (bump version, fix linux ping)" with no
"alpha.37" in the subject (upstream tagged it after we merged), so the grep
found only alpha.36 and rc.17 notes shipped a stale base.
Make `git describe --match v1.14.0-alpha.*` the primary source — it reads
HEAD's ancestry in the commit graph, independent of merge-message wording —
and keep the subject-grep as the fallback for a fork checkout without
upstream tags. Verified locally: now resolves v1.14.0-alpha.37.
SPEC 014 dropped with_clash_api because LxBox (Android) drives the core
over the native libbox CommandClient, making the Clash REST server dead
weight in the AAR. But the drop landed in the shared Makefile.lx LX_TAGS,
which also feeds every desktop/CLI release build (mac/windows/linux-musl
via `make -s lx-print-tags`). A CLI binary has no CommandClient channel —
it is managed by external dashboards (yacd/MetaCubeXD) over the Clash REST
API — so every desktop release since rc.1 shipped with no way to manage
the core; a config with experimental.clash_api failed fast. CI stayed
green (lx-ci BASE_TAGS kept the tag), so it was invisible in CI.
Restore with_clash_api to the desktop LX_TAGS; leave build_libbox (AAR)
unchanged. The two tag sets now diverge by design: desktop = with Clash
API, AAR = without.
Verified: desktop binary builds with with_clash_api in Tags; `check`
accepts an experimental.clash_api config; the Clash REST server comes up
live (endpoints answer 401 security-middleware, not the stub's fail-fast).
Docs: Makefile.lx comment, SPEC 014 (§2/§3.1 scoped to AAR + new §3.4),
lx-release.yml tag comment + notes line, changelog rc.17.
The release-notes template hardcoded "base v1.14.0-alpha.35"; it went stale and
had to be hand-edited on rc.14, rc.15 and rc.16 (each was actually on alpha.36).
Resolve the base dynamically in the "Resolve tag" step: take the highest alpha.NN
named in any "Merge upstream" commit subject (robust on a fork without upstream
tags fetched), falling back to git describe against upstream alpha tags, then a
generic v1.14.x label. The notes line now interpolates steps.ver.outputs.base.
§8: validated transport JSON with all 14 new fields at non-default values,
the equivalent flat-camelCase vless:// URL, and a defaults table for the
toUri() omitempty logic. Fixture verified with sing-box check; mirrors
lx-test/config/xhttp_obfs_full.json.
Relocate docs/lx-xhttp-url-parsing.md -> SPECS/002-XHTTP_CLIENT_TRANSPORT/URL_PARSING.md
so all XHTTP docs live together with the spec. Fix internal/back links.
Scanned igareck/vpn-configs-for-russia, extracted+deduped 10 unique XHTTP
nodes, ran each through our with_xhttp binary. 4 alive — all downloaded 1MB,
traffic egressed via the server IP:
- 2x plain -> packet-up
- 2x reality -> stream-one (hu99.bearbeer.digital, bez3.stream-room.com)
The two reality nodes resolve auto->stream-one and work live, closing the
open stream-one live-verification TODO from task 011 (previously synthetic-only).
Other 6 nodes dead for server-side reasons (504, HTTP/1.1-not-H2, reset,
TLS hang) — our transport errored cleanly in every case.
Remaining live TODO: obfs/placement modes (no public node is configured for them).
Our XHTTP client is HTTP/2 only (http2.Transport); Xray supports H1/H2/H3.
h3-only nodes won't connect — flag for the link parser. Out of SPEC 002 scope
(separate future 'XHTTP over HTTP/3' task).
scMaxConcurrentPosts is a removed Xray knob (grep + GitHub code search
total:0 in current XTLS/Xray-core and sing-box-extended). Current Xray
serializes to one upload POST body in flight at a time, which our sequential
packet-up Write already matches, so the field is accepted for config/link
symmetry but ignored by the client.
- option: V2RayXHTTPOptions.ScMaxConcurrentPosts (json sc_max_concurrent_posts)
- PARAM_MAP: document as legacy/ignore tier with the real concurrency mechanism
(bounded pipe + WroteRequest serialization, server-side seq reorder)
- url-parsing doc: scMaxConcurrentPosts -> accept-but-ignore
- xhttp_obfs_full.json: include the field so check covers it
Verified: build/gofmt/vet clean, 16 unit tests pass, sing-box check passes
on all 3 xhttp configs incl the field.
Self-contained reference for the link parser: maps every vless://...type=xhttp
URL param (flat query + extra={...} JSON) to sing-box transport snake_case fields.
Covers TLS/Reality mapping, the extra-JSON number→"min-max" coercion, mode=auto
pass-through, path-with-query-tail, and ignored fields (scMaxConcurrentPosts,
server-only). Examples validated with sing-box check.
Implement all 12 client-relevant Xray/sing-box-extended XHTTP params on the
existing lean-native client (no Xray vendoring):
- session/seq placement (path|query|header|cookie) + keys
- uplink-data placement (body|auto|header|cookie, chunked base64) + key + chunk size
- uplink_http_method (upper-cased; GET only in packet-up)
- X-Padding obfs mode: placement (cookie|header|query|queryInHeader) + key/header +
method repeat-x | tokenish (HPACK-Huffman-tuned via golang.org/x/net/http2/hpack)
- packet-up tuning: sc_max_each_post_bytes (split), sc_min_posts_interval_ms (throttle)
4 server-only fields (server_max_header_bytes/no_sse_header/sc_max_buffered_posts/
sc_stream_up_server_secs) accepted but ignored by the client.
New files: transport/v2rayxhttp/{meta.go,xpadding.go}, xhttp_test.go.
Range fields use the "min-max" string form (no badoption.Range in sing).
Default (non-obfs) wire shape kept byte-identical to the live-verified v1
(x_padding='0' in Referer, session/seq on path, payload in body).
Verified: 16/16 unit tests, sing-box check on 3 configs incl full obfs,
go vet/gofmt/build (tagged+untagged) clean, negative (no with_xhttp) rejects.
Adversarial wire-protocol review against PARAM_MAP found no bugs.
Live test of a non-default mode against an Xray server remains an open TODO.
The rc.15 domain fix was confirmed on a real device: with the default
sticky_hash ["process","domain"] and no dest_ip workaround, browser traffic
spreads across the pool (on-device per-domain uniformity ~0.27 -> 0.95+).
Update README + lx-config.md status from "not yet device-verified" to
device-verified.
PRIMARY lever spelled out: bufsArrs size = bind.BatchSize() (128 on StdNetBind/android);
shrink it at one point (conn/bind_std.go:322, android branch -> 8/16). BindUpdate
(device.go:558) and getMessages() (bind_std.go:260) follow automatically, both recv
goroutines (v4+v6) covered, no packet loss (array stays full, just shorter), MaxSegmentSize
untouched (GRO intact). Effect 8MB->~0.5MB/recv = 176MB->~11MB at 11 devices. Only risk:
gigabit channel may lose throughput -> fall back to dynamic batch.
On-device throughput A/B (static arm64 curl, download via tunnel):
- baseline batch=128 = 10.7 MB/s median (~86 Mbps), stable 9.1-11.2.
- CPU under load: Syscall6 36%, scanobject 6%, crypto ~3%. The bottleneck is the
WARP channel + syscall overhead, NOT batch processing. GRO/batch only matters at
hundreds-of-Mbps/gigabit, so shrinking batch does NOT cost throughput on a typical
mobile/WARP channel (where the heat is reported).
Re-ranked the levers: PRIMARY is now the global smaller StdNetBind.BatchSize()
(128->8-16) — one point, no activity detection, cuts bufsArrs 8MB->~1MB/recv (176MB->
~11-22MB at 11 devices). Dynamic-batch and Down-idle drop to secondary. Caveat: verify
on a fast Wi-Fi/gigabit channel before release (batch may matter there). Also confirmed
heap scales linearly with live device count (11->269MB, 4->104MB, 1->0). No code changed.
The lx feature docs had drifted: README (en/ru) and docs/lx-config.md still said
"currently XHTTP + AWG2" and covered only SPEC 002/003/009 — the observability
layer (SPEC 014-018) and round_robin load balancing (SPEC 019) were undocumented
in the lx overview, and urltest.md still described the pre-rc.15 domain behaviour.
- docs/lx-config.md: new "## 3. round_robin load balancing" (mode/balancer,
pool/pool_tolerance/sticky_hash, ["none"] sentinel + badjson-[] caveat, slot-hash
binding, example, status) and "## 4. Observability (CommandClient extensions)"
(URLTestOutbound/GetRules/GetGroups/GetOutbounds/GetPool/SubscribeDNSQueries +
Connection.detourList, all behind with_lx_command); Validate&build -> ## 5.
- README.md / README.ru.md: broaden the stale "XHTTP + AWG2" framing; add feature
rows for observability and round_robin with honest status.
- docs/configuration/outbound/urltest.md: reconcile sticky_hash "domain" with the
rc.15 fix — domain reads metadata.Domain (survives domain->IP resolve), so it
works for normal sniffed domain traffic, not only literal-IP destinations; the
warning is reframed (domain works; dest_ip is an alternative).
Docs-only; no code change.
On-device A/B (Debug API /diag/pprof) settles it with high confidence:
- RoutineReceiveIncoming holds 180MB (61% cum); peek = 100% via sync.Pool.Get (in
worker hands, not the pool, not the channels).
- 11 live wireguard endpoints, 22 RoutineReceiveIncoming ALL on StdNetBind (batch=128),
0 on ClientBind. 22 x bufsArrs[128] x 64KB = 176MB ~= 180MB.
- batch=128 because WARP/AWG endpoints use a WireGuardListener dialer => StdNetBind
(endpoint.go:200-202), whose BatchSize()=128 on android.
- A/B: switching the active node WARP->home does NOT free memory (buffers do not sleep);
config 11 ep -> 1 ep gives 269MB -> 0 (= Iliya's workaround, reproduced via profile).
Both prior diagnoses were wrong: batch=1 (no, StdNetBind=128) and drain-on-Suspend of
channels/sync.Pool (misses; bufsArrs of live recv-workers holds it). MaxSegmentSize
2200->65535 is the volume trigger (x30 bytes), not the holder (downLocked/pools/channels/
batch identical 1.13<->1.14).
Fix = shrink batch for INACTIVE devices (naive lazy-bufsArrs impossible: StdNetBind
getMessages() is a fixed 128). Levers + a required download-throughput measurement
documented; pending lever choice.
On the client path BatchSize()=1 (client_bind/stackDevice/systemDevice all return 1),
so the old bufsArrs=128 => 8MB/device claim is wrong; bufsArrs is ~64KB. The 224MB
pprof attributes to PopulatePools is the sync.Pool.New alloc SITE, not the holder.
Real holders: (A) device.pool messageBuffers sync.Pool local+victim cache, (B) the 3
buffered device channels. scanobject 52% comes from the pointer-dense element/container
wrappers (4 of 5 WaitPools are scan-type), not the noscan [65535]byte arrays.
Fix rewritten to drain-on-Suspend: park RoutineReadFromTUN via a stackDevice.suspended
seam in Read (the lockless pool writer surviving Down), drain the 3 device channels
after Down, then ONE runtime.GC()+FreeOSMemory() per selector transition. PopulatePools
swap rejected (WaitPool.count underflow -> cond.Wait deadlock). slim-batch and the
route-graph refcount are dropped (not needed for heat). Added in-repo verification
(device/suspenddrain_test.go + HeapInuse bench) since on-device A/B is impossible.
Android 100% CPU / heat on configs with many WG/AWG endpoints in a
selector. Diagnosed via on-device pprof (§207): scan-bound GC over a
224MB live heap = wireguard-go Device.PopulatePools buffers, held by
~10 idle WG devices. Suspend()=Down() marks the device idle but does
NOT release pools/workers (only Close() does). A/B on device: dropping
spare WG endpoints removes the heat.
SPEC 020: SLIM idle devices (shrink maxBatchSize 128->1-4, do not Close
— keepalives + shared-node safety) gated by a route-reachability
refcount (rules + final + Now, not just selector).
SPEC 010: note our GRO split-brain patch is now upstream-native on
v0.0.3 (commit 24ea133); MaxSegmentSize=65535 must stay (GRO fuel) —
heat is fixed by device count/slimming, never by shrinking the buffer.
Device verification of round_robin on a real 51-node pool surfaced three bugs,
all fixed here. Listed by impact.
1. sticky key 'domain' was always empty -> all traffic collapsed to one node.
The router resolves a domain destination to an IP and overwrites
metadata.Destination before a group's DialContext runs, so destination.Fqdn
is empty when the balancer builds the key. stickyComponent("domain") read
that empty Fqdn, so a single process's key was process+NUL for every site
-> one fixed slot. On device this measured 28/1/1 across a 3-node pool
(uniformity 0.27). Fix: read metadata.Domain (survives the resolve), fall
back to destination.Fqdn only for a direct dial. After: spread 0.95+.
2. living pool nodes could change slot index during a health-check, moving
sticky keys. balancePoolFirstLive compacted with a filtering append (a
transiently-dead slot shifted every later live node left); planTolerantPool
did delete(inPool, occupant) (an evicted-but-living node re-entered a later
slot, cascading); manual URLTest rebuild ran the tolerant planner even at
pool_tolerance==0. All now replace-in-slot (fixed-length copy(current), only
dead/empty slots rewritten by index; dedicated planFirstLivePool for the
tolerance==0 rebuild).
3. stickiness could not be disabled via sticky_hash: [] -- the config decoder
(badjson.UnmarshallExcludedContext) re-marshals the struct and collapses an
empty array to nil, indistinguishable from omitted, so the default always
applied. Disabling now uses the explicit sentinel sticky_hash: ["none"].
Tests: domain-from-metadata + fallback, replace-in-slot survivor/cascade/
first-live regressions (fail against pre-fix code), ["none"] disable + []
defaults + none-mixed error. All green under -race; gofmt clean.
v2 superseded v1; keeping both as separate files (SPEC.md + SPEC_V2.md + the v1
TEST_REPORT) was just confusing. Delete the v1 SPEC and its TEST_REPORT (they remain in
git history) and rename SPEC_V2.md → SPEC.md as the one canonical doc. Drop the "v2"
suffix and stale "design not started" status from the header.
Desktop smoke-test of the rc.13 binary surfaced this: a Go int with omitempty can't tell
`pool: 0` from an omitted field, so `pool: 0` hit the `< 1` validation and rejected a
config that should have defaulted. Now pool 0/omitted → default 3; only a negative pool
errors. Added TestBalancerZeroPoolIsDefault; renamed the negative-pool test. SPEC_V2,
urltest.md, changelog rc.14 updated.
Verified on the rc.13 desktop binary: round_robin pool fill (pool_tolerance:0 tests only
pool-many nodes, >0 tests all), config fail-fast (balancer+least_test, unknown sticky_hash,
unknown mode, negative pool), and live routing through the group.
Reworks urltest round_robin to scale to large node lists. v1 rotated over ALL live nodes,
which meant URL-testing every node each interval (unworkable at 1000 nodes). v2:
- Fixed-size pool of slots (balancer.pool, default 3). Slot indices never move; a
replacement takes the exact slot it evicts. round_robin rotates only within the pool.
- Lazy health-check: pool_tolerance=0 tests no more nodes than needed to keep the pool
full of live nodes, then stops; pool_tolerance>0 tests all and keeps the fastest with a
per-slot eviction threshold. Dead pool node keeps its slot until a live replacement is
found (pool never empties). A dial error never changes the pool — only the health-check.
- sticky = slot-hash (slot[hash(key)%pool], FNV-64a). Binds to a fixed slot index, so a
living node keeps ALL its keys when other slots churn: strict zero reconnects, zero
per-key state. Default sticky_hash ["process","domain"]; explicit [] disables.
- Removes v1 jumphash (broke on mid-list eviction), ttl_map, and least_connection (dropped
from the roadmap — round_robin is statistically even).
- GetPool RPC: CommandClient.GetPool(tag) -> []PoolSlot{slot,tag,delay} so clients can show
the N nodes actually in rotation. delay clamped 0->1 for live nodes; non-round_robin
group -> empty. Additive proto/daemon/libbox, behind with_lx_command.
Config moved under a `balancer` object (breaking for the rc.11/12 round_robin shape; no
prod configs, tests only). least_test (default) is byte-for-byte unchanged.
Tests: newBalancer validation/defaults, rotation distribution, slot-hash stable +
living-node-keeps-keys-across-other-slot-churn, empty-key fixed slot, planTolerantPool
top-N / keep-in-tolerance / evict-beyond / dead-slot-replace. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean. Not yet device-verified.
Clarify the Now() cold-start tradeoff: variant B (write the fallback node straight
into selectedOutbound*) would eliminate the micro-gap entirely — Now() and DialContext
would read one field, so they can't diverge — at the cost of touching upstream's
selection logic (stub in the choice field + one extra Interrupt() on the first real
switch, which is no worse than any later latency switch). Variant A (Now() stays a
reader) was chosen purely for minimal upstream intrusion; its only cost is a negligible
micro-gap from two separate Select() calls racing on the first seconds. Documents the
path to B if the feature outgrows upstream's selectedOutbound* later.
Doc-only; rc.12 already shipped variant A, no retag.
Before the first URL-test fills the delay history, urltest's selectedOutbound* is
nil but traffic already flows via the Select() fallback (first usable outbound).
Now() returned "" in that window, so the UI showed no server while connections were
live. Now() now falls through to Select(tcp)/Select(udp) and reports the exact node
the next DialContext will pick — same source of truth as the dial path, not a guess.
Only least_test (default) affected; round_robin/ttlmap already report the last-picked
tag (lastSelected) and are untouched. Added TestSelectColdStartFallback /
TestSelectColdStartNoOutbounds. SPEC + changelog rc.12. go build (+with_lx_command),
go test -race ./protocol/group/, gofmt all clean.
The TEST_REPORT landed after the tag was cut, so the as-tagged notes still said
"not device-verified" and base alpha.35. Feature was live-verified on 5 vless nodes;
base is alpha.36 after the pre-rc merge. Published GitHub release notes edited to match.
Live run on 5 vless nodes (3 instances, one per mode): round_robin rotates strictly
across the live set and skips dead nodes; both sticky strategies pin deterministically;
bad config is rejected at start; -race clean on units and live. Feature is now
device-verified, not just isolated.
The run surfaced a config caveat (not a bug, by design): dest_ip is empty until the
destination is resolved, so a sticky key of only source_ip/dest_ip/dest_port collapses
to "" for domain traffic and pins everything to one node. Documented in urltest.md —
use `domain` in `hash` for domain-based traffic.
The lx-ci gofmt-lint step only checks files matching the lx-owned glob
(_xhttp|_awg|_lx.go|_command_lx); the new SPEC 019 files fell outside it. Rename
to the _lx.go convention so CI gofmt-checks them, and fix the changelog reference.
No code change.
Pre-rc.11 sync. Upstream changes: darwin local DNS refactored to a raw
mDNSResponder call, iOS deb upload fix, version bump. No overlap with lx files
(protocol/group, option, constant untouched).
Add a `mode` to the urltest group so it can distribute traffic instead of only
picking the lowest-delay node, with optional per-flow stickiness.
- mode: least_test (default, unchanged) | round_robin (rotate across live nodes)
| least_connection (reserved, phase 2 — rejected at config time).
- round_robin selects once per connection over the tag-sorted live set (nodes with
a fresh URL-test result supporting the network); UDP/QUIC sessions stay on one
node; first usable outbound is the fallback when nothing is live. The legacy
selectedOutbound* cache path is untouched — balancing is a separate branch in
DialContext/ListenPacket.
- sticky {mode, timeout, cap, hash}: binds one flow to one node. hash components
process|domain|source_ip|dest_ip|dest_port concatenate in order; absent -> "",
all-empty key -> one fixed node (keyless flows never rotate). mode jumphash
(default, stateless consistent hash — ~1/n remap on node-set change) or ttlmap
(key->node table, lazy + ticker eviction, 2000 LRU cap, 10m TTL, dead-node re-pin).
Reuses the existing urltest health ticker/history as the single liveness source;
no new probing. Now() reports the last-picked tag in balanced modes.
Tests (go test -race, 15 cases): distribution, dead-node skip, all-dead fallback,
jumphash stability + empty-key fixed node, ttlmap stick/expire/cap/dead-repick,
key building, validation. The race detector caught a real bug in the sticky
sweeper (read t.ticker unlocked while close() nilled it) — fixed by passing the
channels into the goroutine, mirroring URLTestGroup.loopCheck.
Also folds the SPEC 016 connections-map mutex (ebf9cc07) into the rc.11 changelog
section, which had not yet shipped in a release.
Connections is the client-side CommandConnections accumulator. With 2+
subscribers (LxBox screenClient + profilerClient) one goroutine writes
connectionMap in ApplyEvents while another ranges it in Iterator →
"concurrent map iteration and map write" fatal error → SIGABRT of the
whole process (reproduced in ~20s under traffic, CPH2411/Android15).
Add access sync.Mutex; lock every public method touching
connectionMap/input/filtered: ApplyEvents, FilterState, SortBy*, Iterator.
- FilterState split into public (locks) + private filterState (no lock);
ApplyEvents calls the private one under its already-held lock
(sync.Mutex is not reentrant). Field filterState -> filterStateValue to
free the name for the method.
- evictClosedConnections stays lock-free: private, only called from
ApplyEvents under lock.
- Iterator returns a COPY of filtered — the gomobile caller walks it
after the Go call returns (lock released), so it must not read the
live slice a concurrent ApplyEvents/SortBy is rewriting.
This is the UI/command channel, not the data plane — uncontended lock
~20ns. LxBox per-client accumulators (§170) stay as the consumer scheme;
the mutex is class-correctness insurance against a 3rd consumer.
Verified: go test -race TestConnectionsConcurrentAccess (writer || 3
readers, 2000 rounds) green; go build ./... and -tags with_lx_command
green; gofmt clean.
Codify the rule: before cutting any lx release/prerelease tag, check whether
upstream/testing moved ahead of our last merge and, by default, merge it in
first — then build/gofmt/lx-check, then changelog, then tag.
- docs/lx-release-runbook.md: pre-release gate checklist, drift-check commands,
the manual `git merge upstream/testing` flow (replaces SPECS/004 auto-rebase
while upstream is v1.14.*-alpha), conflict zones (.pb.go, wireguard-go submodule,
build_libbox marker, observability files), and the one-liner sequence.
- SPECS/004 SPEC.md: pointer to the runbook + note that manual merge superseded
auto-rebase on this branch.
Two nits surfaced by the lx-vs-upstream cleanliness audit (no runtime impact):
- box.go: the dnstrack registration comment said "service.FromContext" — the
§180 dead-stream signature. The actual readers use PtrFromContext (pairs with
MustRegisterPtr). Fixed the comment + noted why FromContext[*T] returns nil,
so a future debugger doesn't "fix" the readers back into §180.
- common/dnstrack/manager.go: removed the unused SourceRejected constant —
rejected resolutions are folded into SourceFailed at the emit site, so
"rejected" never reaches the wire. Replaced with a comment to prevent re-adding
an unreachable client case.
Audit verdict: code clean — no concurrency/wire/behaviour issues; dns/client.go
byte-identical to upstream, emits additive and subscriber-gated.
LxBox feedback: DnsQuery lacked which DNS server / outbound channel the query went
through. A DNS rule selects a server (matchDNS by action.Server), not an outbound;
the channel is the server's own detour, fixed at config time. Add to DnsQueryEvent:
- dnsServer/dnsServerType = transport.Tag()/Type() (transport is the Exchange param,
so available on all emit paths incl. failures);
- outbound = the server's detour tag (TransportAdapter.OutboundTag() from
DialerOptions.Detour), with a selector expanded to its live node via Now()
server-side (like Connection.Detour), empty on cached/optimistic.
Also gate event construction on HasSubscribers(): with no profiler attached the DNS
hot path builds nothing (no event/answers/outbound lookup) — previously every
resolution built an event just to be dropped for lack of a listener. The Now()
resolution therefore never touches the hot path.
Wire: additive proto fields + OutboundTag() on DNSTransport (embedded adapter
satisfies it). libbox DnsQuery.DNSServer/DNSServerType/Outbound(). Changelog rc.10.
DNS attribution was empty (0/119 on device): TUN+DNS hijack returns on a fast-path
(route.go:91/226) BEFORE matchRule, and searchProcessInfo — which fills
metadata.ProcessInfo — lives inside matchRule (:416). So fast-path DNS (most DNS on
a VPN) reached the SubscribeDNSQueries emit with nil ProcessInfo. Fix: call
r.searchProcessInfo(ctx, &metadata) before both fast-path hijacks (stream+packet);
idempotent + cached, one lookup per flow. Corrects SPEC 018 пункт 3 (the earlier
'cached attribution correct' claim checked ctx consistency, not that ProcessInfo
was populated before the resolve).
Also: DnsAnswer.rdata was the full RR string ('google.com. 29 IN A 1.2.3.4'); strip
the header prefix so clients get the bare value ('1.2.3.4' / CNAME target).
No proto/wire change. LxBox §180 needs no client change. Changelog rc.9.
SubscribeDNSQueries returned Unimplemented on device and emitted nothing: the
dnstrack.Manager is registered via MustRegisterPtr (key *dnstrack.Manager) but
read via service.FromContext[*dnstrack.Manager] (key **dnstrack.Manager), so the
lookup always found nil. Server -> Unimplemented; emit sites -> silent drop.
Fix all three readers to service.PtrFromContext[dnstrack.Manager] (the pair of
MustRegisterPtr, as trafficManager does in daemon/instance.go). Verified the
manager resolves to the exact pointer box.go registered. No proto/wire change;
rc.7 contract intact. LxBox §180 needs no client change.
Changelog rc.8.
The release-notes heredoc hardcoded 'base v1.13.13' and only AWG+XHTTP — stale
since the 1.14 migration, identical for every rc, and never reflecting what a tag
actually shipped (SPEC 014/015/017/018 were invisible). Now the 'What's new'
section is extracted from docs/lx-changelog.md for the current version (awk between
'#### vX' and the next '#### '), spliced via 'sed r' so changelog backticks/$()
stay inert (no command injection from doc prose). Base line fixed to alpha.35;
standing-features list updated with the CommandClient extensions.
Two doc-hygiene fixes from LxBox review (no logic change):
- Point the реализатор at 'Согласованная форма' as the binding contract; the
earlier 'Решение' proto sketch (Empty input, no failed/answers) is illustrative.
- Note that rcode=-1 ships as signed int32 (distinct on the wire from 65535,
verified); client must map -1 -> 'no answer' before any .toUInt().
Hijacked DNS (the norm on an Android VPN) is answered before a connection becomes
a traffic tracker, so DNS queries never reach the connections stream — the only
egress was the text log, which carries no app attribution. Add common/dnstrack
(a Subscriber[QueryEvent] mirror of trafficcontrol) emitting one event per
resolution from dns/client.go, attributed via adapter.ContextFrom(ctx).ProcessInfo
(same ctx on cache-hit and miss, so cached queries are attributed too).
Failures are first-class: timeout/loopback/rejected-cached/SERVFAIL-reject emit
failed=true + error + rcode=-1 (no response) — without this the stream is blind to
DNS failures, the primary throttling signal. CNAME chains preserved: with
includeAnswers, each event carries the full response.Answer in wire order (CNAME
hops + final A/AAAA, not filtered to IPs).
Wire: rpc SubscribeDNSQueries(SubscribeDNSQueriesRequest) returns (stream
DnsQueryEvent) + DnsAnswer; event-driven server stream (no ticker); libbox
SubscribeDNSQueries(includeAnswers, handler). Tag-less core -> Unimplemented.
Detour/Chain and other streams unchanged.
Docs: SPECS/018, lx-changelog rc.7.
chain omits the final outbound's own detour by design (upstream loop only
unwinds OutboundGroup via Now() and breaks on the first non-group), so a node
detouring through e.g. WARP never shows in the routing chain. Add Detour
[]string to TrackerMetadata, unwound from the final outbound's Dependencies()
(= its detour for a non-group outbound), descending into groups via Now()
against the same atomic snapshot, with a seen-guard against cycles.
Wire: additive 'repeated string detourList = 23' on the Connection proto
message (hand-applied to keep the generated diff minimal — no toolchain churn),
mapped in connectionToProto, surfaced on libbox Connection as Detour()
StringIterator. Chain / Clash-API unchanged.
Docs: SPECS/017, lx-changelog rc.6.
Parent the per-node delay test to the gRPC per-call ctx instead of the
long-lived boxService.ctx, so cancelling the call aborts the in-flight
dial before C.TCPTimeout without tearing down the connection. Restores
the granular per-node cancel the Clash API had implicitly via r.Context()
(there was never a cancelDelays endpoint).
Mass-cancel is unblocked client-side on the existing gomobile binding
via a separate ping CommandClient + Disconnect() (no native-surface
change, no server batch RPC) — closes the LxBox feedback.
Docs: SPEC 015 §3.6 (cancellation), SPEC 014 (#4240 deleted upstream →
seam-removal criterion switched to upstream-code), lx-changelog rc.5.
Ignore test/cache.db.
Connections (command_types.go:115) держит connectionMap/input/filtered без
синхронизации; ApplyEvents/evictClosedConnections/FilterState/SortBy*/Iterator
зовутся из разных gRPC-горутин (по одной на подписчика handleConnectionsStream).
≥2 подписчика CommandConnections → concurrent map iteration and map write →
fatal error → SIGABRT всего процесса.
Всплыло после CommandClient-миграции (раньше connections слушал ≤1 потребитель).
Обойдено клиент-стороной в LxBox §170 (per-client accumulator), но в ядре не
починено — третий потребитель вернёт краш. Фикс: sync.Mutex вокруг состояния.
Референс: LxBox docs/spec/tasks/170.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
Serialize probe rounds in startProber to eliminate unbounded fan-out of
fire-and-forget probe goroutines (up to 100/sec per direction), and close
HTTP/3 transports via transport.Close() in addition to CloseIdleConnections.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.
Three command-protocol additions completing the Clash-API -> CommandClient
migration (SPEC 015, behind with_lx_command):
- GetGroups / GetOutbounds: unary pull-snapshots over the existing readGroups()
and the SubscribeOutbounds builder. The CommandClient is push-only; if the
SubscribeGroups stream never opened (service not STARTED at subscribe) or broke,
the client had no cheap way to re-read group state and the main screen stayed
empty (tunnel connected, groups=[]). These getters close that gap without
recreating the whole client. Both needed: SubscribeGroups covers only in-group
nodes, endpoints (WG/AWG) + standalone outbounds appear only via the flat list.
Errors via status.Error (unary read convention, like GetRules).
- len<2 fix: readGroups() silently dropped groups with < 2 items (upstream commit
5bc0dfa9), hiding single-node selectors -- a regression vs Clash, whose /proxies
returned group.All() unfiltered. readGroups() is the single source feeding both
SubscribeGroups (startup broadcast) and GetGroups, so the fix covers both.
Handlers in started_service_command_lx{,_stub}.go behind with_lx_command;
client methods in command_client_command_lx.go reuse the existing gRPC->libbox
iterators. proto seam under // lx: marker, regenerated via pinned lx-proto.
E2E tests (test/command_lx_test.go) drive the public daemon API: single-node
group survives, flat list returned, not-started rejected. Both tag/no-tag builds
green; no-tag answers Unimplemented.
Split the original SPEC 014 by NATURE of change:
- 014 CLASH_API_TO_COMMANDCLIENT_MIGRATION (dir renamed) — the migration itself:
with_clash_api drop (rc.1) + box.go Android-start fix (rc.3). No RPC tech-spec.
- 015 COMMAND_PROTOCOL_RPC_EXTENSIONS — single home of all command-RPC work:
URLTestOutbound + GetRules (DONE, rc.2) + GetGroups + GetOutbounds + the len<2
readGroups bugfix (TODO, rc.4). All §3.6 class, with_lx_command.
Docs-only; shipped rc.2 code unchanged. 015 documents the pull-vs-push gap
(GetGroups/GetOutbounds) and the upstream len<2 group-drop defect, plus an
upstream-candidacy plan (§7): pull-getters + len<2 are clean upstream defects;
RPCs currently ship in lx-form, an upstream PR would need upstream-form (future).
Records the Android start fatal fixed in rc.3 (commit 029acd11): PlatformLogWriter
no longer forces the Clash server; observability served by the native
CommandClient. Self-contained in the feature SPEC — fix + WATCH
SagerNet/sing-box#4240 + the obligation to drop the // lx: box.go seam on the
next rebase if upstream resolves it.
Upstream box.go forced needClashAPI whenever PlatformLogWriter is set (always
on Android/libbox), because the Clash server was historically the only log/
traffic observer. With with_clash_api dropped (rc.1), that made every Android
start fatal: 'clash api is not included in this build' — even with no clash_api
in the config.
Split the concern behind a // lx: seam: PlatformLogWriter now requests
observability (Observable log factory + connection/traffic tracker), served by
the native CommandClient (SubscribeLog/SubscribeConnections), NOT the Clash
server. Only an explicit experimental.clash_api block still creates the Clash
server (and still fails fast without the tag). daemon is already nil-safe to a
missing clashServer, so Clash-mode degrades gracefully. Desktop unaffected.
Verified: core starts with no clash_api config; still fail-fast with one.
Clarify where a URLTestOutbound result surfaces, since the original §3.2 wording
only named OutboundGroupItem ("for nodes in groups") and omitted SubscribeOutbounds
— the actual channel that carries endpoint (WG/AWG/Tailscale) delay.
- §3.2: add a 3-row channel-map table (synchronous RPC response = any node;
SubscribeOutbounds = all outbounds AND all endpoints; SubscribeGroups = only
OutboundGroup members). All three share urlTestObserver via urlTestHistoryStorage.
- §3.2: sync the client signature to what shipped — (*URLTestOutboundResult, error)
with int32 timeout (gomobile can't bind the draft (uint16,string,error)); note why.
- §3.7 / §5: replace the narrow "history flows to OutboundGroupItem" line with the
SubscribeOutbounds + SubscribeGroups split.
Docs-only; no code or artifact change (rc.2 binaries unchanged).
Restore over the native libbox CommandClient what upstream only exposed through
the dropped Clash API: per-node delay testing and a route+DNS rule-table snapshot.
Both RPCs are a pure bridge (CONSTITUTION §3.6) gated by the with_lx_command tag.
- daemon/started_service.proto: URLTestOutbound + GetRules RPCs and messages under
the // lx:begin/end lx_command marker; regenerated .pb.go/_grpc.pb.go.
- daemon/started_service_command_lx.go (+ _stub.go): handlers behind with_lx_command,
stub twin returns codes.Unimplemented. URLTestOutbound resolves an outbound OR an
endpoint (no OutboundGroup assert), honours link+timeout, error-in-payload Variant B
(delay==0 && error=="" is success 0ms), history Store/Delete via group.RealTag.
GetRules returns route + DNS rules split by isDNS.
- adapter/dns.go + dns/router.go: new adapter.DNSRouter.Rules() getter (route Router
already had one), read under rulesAccess; both under // lx: markers.
- experimental/libbox/command_client_command_lx.go: CommandClient.URLTestOutbound
(*URLTestOutboundResult, error) and GetRules (RuleIterator, error) — gomobile-bindable
shapes (the SPEC's bare (uint16,string,error) does not bind); Variant B preserved.
- cmd/internal/build_libbox/main.go: with_lx_command into sharedTags (AAR).
- Makefile.lx: with_lx_command in LX_TAGS; pinned lx-proto/lx-proto-install targets
(protoc-gen-go v1.36.11, protoc-gen-go-grpc v1.5.1) for reproducible regeneration.
- lx-ci.yml: vet+gofmt cover the lx files; build-check proves both builds toggle the
stub marker.
- Collateral one-time pin normalisation of managed_service/v2rayapi/v2raygrpc .pb.go
(audited in SPEC §3.5).
docs(lx-changelog): v1.14.0-lx.1-rc.2. SPEC 014 → accepted.
CONSTITUTION: three owner principles (thin layer / follow upstream / build
what we+users need) wired into §1–2 as a priority hierarchy with a
necessary-not-sufficient lock; §3.1(а) rewritten from binary 'not in upstream'
to a three-prong test (needed by us/users / absent from OUR built channel /
cheaper than the alternative, with an auditable touched-file count); new §3.6
legalizes the 'libbox command-protocol extensions' change-class — handlers
gated by with_lx_command behind the proven started_service_usbip{,_stub}.go
pattern, .proto seam under a // lx: marker, .pb.go regenerated (never
hand-edited). §3.5 version bumped 1.13.13-lx.N → 1.14.0-lx.N.
SPEC 014: two CommandClient RPCs restoring what was lost when with_clash_api
was dropped — URLTestOutbound (per-node delay for outbound OR endpoint, custom
url + timeout, synchronous {delay,error}, all errors in payload) and GetRules
(route + DNS rule table snapshot). DNS-rules need a marked getter on
adapter.DNSRouter/dns.Router (route-only doesn't). Deterministic proto
regeneration (pinned protoc in Makefile.lx) is a mandatory deliverable. No
separate history RPC, no cancel-handle, no batch — those live in the client.
LxBox is moving to manage the core over the native libbox CommandClient
(group/url-test/select/connections streams), so the Clash REST API is dead
weight on the client. Drop with_clash_api from both the Android AAR
(build_libbox sharedTags) and the desktop LX_TAGS. A config referencing
experimental.clash_api now fails fast (no silent fallback); lx configs won't.
lx-release.yml: tags with an -rc.N / -alpha.N / -beta.N suffix now publish as
GitHub pre-releases (--prerelease), so an unverified build never displaces the
stable lx release as Latest.
First build on the upstream 1.14 base. The WG-endpoint GRO fix (010) lands at
the AmneziaWG v0.0.3 submodule source (no downstream guard), but the Android
download-stall path is NOT yet re-verified on hardware -- hence the -rc.1 tag.
Step 2 of 2 of the 1.14 migration. Points the submodule at e5feca7
(AmneziaWG 2.0 obfuscation re-grafted onto sagernet/wireguard-go v0.0.3).
Verification on lx-1.14:
- full sing-box build with lx tags (with_gvisor/quic/wireguard/utls/clash_api/xhttp/awg): OK
- submodule builds clean for linux/android/windows/darwin (library packages)
- transport/wireguard, protocol/wireguard, protocol/group, route/rule tests: green
- broad test (option/route/transport/common): green; gofmt + go vet clean
- binary runs on 1.14; package_name_regex config validates; awg2_basic + awg2_ranged validate
§010 android UDP_GRO guard dropped (v0.0.3 fixes split-brain at source) — pending
on-device re-verification before any release tag.
Full 1.14 migration, step 1 of 2 (sing-box repo layer). Three conflicts
resolved, all as predicted by the feasibility analysis:
- route/rule/rule_item_package_name_regex.go (add/add): took upstream's
canonical version (slices.ContainsFunc) — our lx.15 backport collapses
back into upstream, so the file no longer diverges going forward.
- route/rule_conds.go: kept our package_name_regex in isProcess{,DNS}Rule
and took upstream's new isNeighbor{,DNS}Rule additions.
- cmd/internal/build_libbox/main.go: kept lx with_xhttp/with_awg append and
the no-tailscale block; deliberately dropped upstream's new with_usbip
(server-side USB/IP, contradicts client-trim).
go.mod auto-merged: wireguard-go require bumped to v0.0.3, lx replace block
(=> ./submodules/wireguard-go) preserved. Submodule pointer unchanged here —
the AmneziaWG graft rebase onto v0.0.3 is step 2 (next commit). This commit
does NOT build yet (submodule still on the old wireguard-go base).
Backport upstream 1.14 feature 941ce58b onto the 1.13.13 base without the
full migration. Adds the package_name_regex rule item (regex match over
ProcessInfo.AndroidPackageNames) to route, DNS and headless rules.
- new route/rule/rule_item_package_name_regex.go (verbatim upstream) + unit test
- PackageNameRegex option field in RawDefaultRule/RawDefaultDNSRule/DefaultHeadlessRule
- item registration in NewDefault{,DNS,Headless}Rule with E.Cause(err, package_name_regex)
- package_name_regex added to isProcess{,DNS,Headless}Rule conds
The commit's RuleSetVersion5 hunk is intentionally NOT ported (that is 1.14
rule-set v5, unrelated; base is RuleSetVersion4). Full 1.14 migration deferred
to v1.14.0 stable. SPEC 013 + Roadmap entry.
builds (no-tags + lx-tags), go vet, gofmt and rule tests all green.
Folder names were NNN-T-S-NAME, so the status (S) letter forced a rename on
every status change — and refs to the full name went stale each time. One was
already broken in-tree (client.go pointed at 002-F-O-… while the folder was
002-F-C), and the submodule needed a cosmetic commit once already (010 O→C).
Make the number the only stable anchor:
- Rename all 12 folders NNN-T-S-NAME → NNN-NAME (git mv, history preserved).
- Type/status now live in a table header at the top of each SPEC.md (canon),
aggregated by the Roadmap in SPECS/README.md (added missing 010, 012).
- Fix every ref to the old full name: README(.ru), docs/lx-changelog,
docs/lx-config, intra-SPECS cross-links, PROBE.md git-apply path,
TASKS.md titles, and transport/v2rayxhttp/client.go:5 (also un-stales O).
- Rewrite the convention + Workflow in SPECS/README.md and the DoD ritual in
IMPLEMENTATION_PROMPT.md ("rename folder to …-C-…" → "set status in header
+ Roadmap").
- Bump submodule wireguard-go (0c0c10b): fix comments point at the new folder
name SPECS/010-WG_ENDPOINT_GRO_SPLIT_BRAIN. Comment-only, no behavior change.
Not touched: gro-probe.patch (historical diagnostic diff artifact; §010 closed,
no longer applied).
The symptom was seen on DIFFERENT nodes including WG, so "↑/↓0 download stall" is
an umbrella over the symptom, not one bug. Correcting the overconfident closure:
- §010 GRO fix lives entirely in the wireguard-go submodule (imported only by
transport/wireguard/) and gates UDP_GRO — it physically cannot affect VLESS/
reality (TCP) nodes. It closes the WG share of the symptom only.
- In the lx.12→lx.14 window, route/conn.go (shared relay) and the sing copy path
were unchanged; the only non-WG-relevant change is §011 (xhttp stream-one), which
applies only if the node uses xhttp. For VLESS+reality-direct there is NO code
change that explains the disappearance.
- "also hangs on VLESS" was never strictly confirmed (wlan0 encrypted), so the
non-WG share has no confirmed root cause — it currently just doesn't reproduce.
Status stays C (not reproducible). Refs SPECS/012.
On the same exit node (VLESS NL 154.83.159.64:8443) and same network where the
baseline zombie download-stall was caught, the bug no longer reproduces on core
1.13.13-lx.14 — neither normally nor under the probe with LX_CONN_TRACE=0 (code =
release, env set via Android wrap.<pkg> prop, verified in /proc/<pid>/environ).
Likely cause: the §010 GRO split-brain fix landed in lx.14 (lx.12→lx.14 bumped the
wireguard-go submodule 27290b6d→6513629). §010 was literally "no-detour WG-endpoint
killed download on android" via receive-side coalescing — the same symptom class.
The original "also hangs on VLESS" note (→ "not a §010 dup") was never strictly
confirmed (wlan0 encrypted, core-log silent on direction), so §012 is most likely
a manifestation of §010 rather than a separate VLESS bug.
Honest caveats: the counter-proof (run baseline on a pre-lx.14 core on the same
node) was not done; the bug was intermittent. Status is "not reproducible", not
"root cause proven". Folder renamed 012-B-O → 012-B-C.
Probe stays as history on branch lx-conn-trace-probe (b6d8c40a, not deleted) — if
the symptom returns, activate via wrap.<pkg> LX_CONN_TRACE=… per RUN-PLAN.md, but
first make the wrapper transparent (it currently silences ReadWaiter/copyDirect on
reality-download and could mask the bug).
Closes SPECS/012.
Synchronous dual tcpdump (tun0+wlan0, single phone clock, from SYN) localized the
stall inside the kernel on the download direction (remoteConn→conn): 777B arrived
on wlan0 but never reached the app on tun0. Root cause not yet confirmed from
inside the kernel — pcap + code-reading only.
SPEC documents symptom, the synchronous-pcap proof, ruled-out causes (MSS, exit
protocol, server/edge/node, RST storm, §010 GRO), and strictness caveats. Adds
22.06 corrections that retire false leads: conn.go:262 pointed at the UDP
canceler not the copy; run_core.log was a launcher error (no macOS `timeout`),
not a kernel log; the device ran the release core without lx changes (the
LX_TCP_RESPONSE_TIMEOUT prototype was never on it).
PROBE.md describes the LX_CONN_TRACE probe (byte read/write counters per copy
direction, periodic tick + final snapshot) that splits the fork: read=0 → above
copy (proxy decrypt); read>0,write=0 → tun write stall. instrumentation.patch is
the self-contained probe diff; RUN-PLAN.md is the on-device run procedure.
Probe code lands on a separate branch (lx-conn-trace-probe) for CI builds, not on
lx — it forces the buffered copy path (loses splice) and is diagnostic-only.
Refs SPECS/012 (status O — probe written, not yet run on device).
No-detour WireGuard-endpoint killed download on android: UDP_GRO was enabled and
rxOffload read true, but the GRO receive dispatcher in bind_std.go is gated on
GOOS=="linux" (android is not "linux") → a coalesced super-packet was read as one
datagram and corrupted the WG stream. Gate UDP_GRO + rxOffload behind !android
(TX/GSO untouched; non-android linux unchanged).
Confirmed on device (CPH2411/Android-15): pre-fix probe rxoffload=true+dispatch=
single; post-fix rxoffload=false, download 0.44→20.7 Mbps, on par with a control
node on the same LTE cell. Candidate #2 (silent handover) not needed.
Bumps wireguard-go submodule pin to 6513629 (fix, no probe). Probe instrumentation
was never on lx — it lived only on the temporary gro-probe-010/*-verify branches.
Closes SPECS/010.
No reality+xhttp node available; per owner decision the fix is accepted on
synthetic evidence (line-by-line Xray contract match, issue #5635, hiddify
parity, green unit tests + check + builds). Live against a real Xray server
remains an open TODO documented in the 011 REPORT — re-open if it diverges.
- SPECS/011 → status C (folder 011-B-C); REPORT carries an honest live caveat.
- 002 REPORT: stream-one marked fixed-by-011 (was 'known bug').
- SPECS/README roadmap: add 011 row, update 002.
Branch lx-xhttp-streamone; NOT merged into lx.
stream-one was sending <path>/<sessionId>; Xray's splithttp server routes the
bidirectional stream-one handler only on an empty sessionId, so the request must
target the bare normalized path. With the sessionId present the server took the
stream-down branch and the response body carried non-VLESS bytes → VLESS
'unknown version'. Now dialStreamOne uses requestURL() (bare <path>, no trailing
slash); stream-up/packet-up keep their sessionId/seq.
mode=auto now mirrors Xray: reality → stream-one, otherwise packet-up. Reality is
detected by runtime type name (reality_detect.go) with kTLS unwrapping, avoiding a
with_xhttp→with_utls compile dependency so with_xhttp builds without with_utls.
Refs SPECS/011. Synthetic-validated; live pending.
Manual workflow_dispatch builder for any branch/tag: target ∈
{android-aar, apple-xcframework, binary, linux-musl, all}, branch = any ref.
Mirrors lx-release build jobs but uploads artifacts instead of releasing.
Lives on lx (default branch) so `gh workflow run` can find it; each job
checks out the requested branch, so lx source is never required to build it.
gh workflow run lx-build.yml -f target=android-aar -f branch=<branch>
No project code touched — CI file only.
The URL test history update hook and the Clash mode update hook were
single-slot: the API service's attached service overwrote the hook set
by the daemon, so clients stopped receiving group updates. Replace both
with multicast hook lists.
Also share a single URL test history storage via context: Clash API
looked it up under a key nobody registered and fell back to its own
empty storage, so dashboards showed no delay once an API service was
configured. Selector changes now notify through the shared storage,
covering selections made from any API surface.
Serialize probe rounds in startProber to eliminate unbounded fan-out of
fire-and-forget probe goroutines (up to 100/sec per direction), and close
HTTP/3 transports via transport.Close() in addition to CloseIdleConnections.
DNS rules referencing rule-sets that contain only ip_cidr predicates
silently stopped matching when legacy DNS mode was disabled, because the
IP-CIDR branch cannot match against an in-flight DNS query. The existing
validation intentionally let every rule_set through on the premise that
mixed sets still work via their non-IP branches, which is only true when
such a branch exists. Track whether a rule-set carries any non-IP-CIDR
predicate and reject pure-IP references the same way bare ip_cidr fields
are already rejected.
QUIC — revert to ONE Initial. The earlier i1+i2 "developing session" was
conceptually wrong: each DCID is a distinct QUIC connection, so two
Initials with different DCIDs read as two ABANDONED connections (more
anomalous to a DCID-tracking DPI, not less), and a real same-DCID
continuation is impossible (short header is device-blocked; a 1-RTT
packet before the server's reply is an invalid QUIC state). So ip=quic
now emits a single fragmented Initial; realism comes from the
browser-accurate ClientHello (ib → uTLS, device-confirmed working), not
from packet count. masqueI1I2 quic branch returns i2=""; the dead
masqueQUICSecondInitialCPS is removed; the i2-conflict guard is now
sip-only.
SIP — INVITE (i1) + matching 100 Trying (i2), one dialog (separate work):
both whole valid SIP messages sharing Via branch / From tag / Call-ID /
CSeq from a single newSIPDialog pass; pseudo user/host names. (Device
result: still times out on the WARP DPI — see memory; kept for other
providers.)
Both build tags (with_utls / no-utls) build & test green; gofmt/vet clean;
lx-build ok. Docs: SPEC §9 rewritten (multi-packet QUIC considered &
rejected), §10 scope, IMPLEMENTATION_REPORT R11 marked rejected; tests
updated (TestAwgIpcLinesQUICSingleInitial, NonSIPNoI2, SIPExplicitI2Conflict).
The Ib hint finally affects the wire: ip=quic + ib=chrome|firefox builds
the ClientHello with uTLS (github.com/metacubex/utls, the same lib Reality
uses) so the decoy carries a genuine browser JA3/JA4 instead of our
generic ClientHello.
- buildClientHello is now a dispatcher: ib=""/curl → buildGenericClientHello
(the ~294B device-proven CH, unchanged default; uTLS has no curl-QUIC fp),
ib=chrome/firefox → buildBrowserClientHello.
- quic_clienthello_utls_awg.go (with_awg && with_utls): UQUICClient in QUIC
mode with HelloChrome_120 / HelloFirefox_120, ALPN forced to h3, and the
PQ hybrid key_share (X25519MLKEM768, ~1.2KB) stripped so the CH fits one
Initial (reality_client.go pattern). TLSVersMin/Max pinned to 1.3 (QUIC
requirement). Result ~510-620B — a real late-2023 browser JA3.
- quic_clienthello_utls_stub_awg.go (with_awg && !with_utls): graceful
fallback to the generic CH when uTLS isn't built.
- The larger CH re-shapes fragmentation, but planFragmentsN cuts any length
and I1–I4 hold (verified). i2 (multi-packet) uses the same browser too.
WHY ib is optional / forward-looking: on the target DPI ip=quic already
passes on fragmentation alone (no fingerprint check), so the default ib=""
keeps the device-proven generic path; the uTLS CH is a knob against a
future JA3/JA4-classifying DPI and is not itself device-verified. Honest
caveats: JA3 matches a pre-PQ browser (no MLKEM key_share), and JA4 (sorts
+ ignores GREASE) is not fooled by it.
Test (quic_clienthello_utls_awg_test.go, with_utls): chrome/firefox yield
distinct larger ClientHellos that still decrypt + carry SNI + offset≠0
(I1); chrome has GREASE ciphers, firefox doesn't; ""/curl stay generic.
Both tag combos build & test green; gofmt/vet clean; lx-build ok;
sing-box check passes for ib=chrome. SPEC §4/§6/§7 + IMPLEMENTATION_REPORT
R12 + TASKS updated.
ip=dns no longer requires id: when absent, the QNAME is a generated
pronounceable pseudo-domain — consistent with ip=sip's pseudo-host
fallback, and it removes the hardcoded-default-beacon problem (every
default user would otherwise share one QNAME).
- pgDomainHost (pseudo_gen_awg.go): domain-only pseudo name (2-/3-level
LDH), NEVER an IP or a "sip." subdomain — a DNS query for a bare IP or
a sip-prefixed name is implausible, unlike pgHost (which sip uses and
where an IP host is fine). Per-build (baked into the <b> blob), not
per-packet: fresh between users/regenerations, removing the cross-user
signature; the QNAME is fixed within one node's packets (CPS can't do a
pronounceable variable-length name per packet).
- masque_awg.go dispatch: ip=dns with empty id → pgDomainHost(); a set id
is still LDH-validated. id is now REQUIRED only for quic (SNI).
Tests: TestMasqueI1DomainRequiredForQUICOnly (only quic errors on empty
id); TestMasqueI1DomainOptionalForNonQUIC now also checks dns-without-id
produces a valid query whose QNAME is a multi-label pseudo-domain (no IP).
Docs (SPEC/EXAMPLES/IMPLEMENTATION_REPORT/TASKS/README/lx-config): id
required only for quic. sing-box check: ip=dns without id now passes.
Bring the user-facing docs in line with the as-built 009 masquerade after
the lx.12 release:
- README.md / README.ru.md: feature table + masquerade section rewritten —
quic is the only device-proven profile on a real LTE/WARP DPI (~330 ms),
now multi-packet (i1+i2) with a randomized per-call layout; dns/stun/sip
are correct client-initiated requests but blocked as a protocol class to
the WARP edge (kept for other providers). id required for quic/dns only.
- docs/lx-config.md: quic = i1+i2 + randomized; stun = Binding Request (was
"Binding Success Response"); sip = INVITE+SDP (was "200 OK response");
profiles framed as client-initiated, not WireSock server responses.
- docs/lx-changelog.md (new): fork changelog (lx.11, lx.12). Kept separate
from upstream changelog.md so a rebase onto upstream stays conflict-free.
Docs only.
ip=quic now emits TWO independent fragmented QUIC Initials (i1 + i2), so
the decoy flow reads as a developing QUIC session (two session starts)
instead of a single opener — lowering the single-packet signature.
Device-verified: an explicit i1+i2 config brings the WARP tunnel up with
NO latency regression vs i1-only (~340ms), confirming the multi-packet
form is safe for the handshake.
- masqueQUICSecondInitialCPS (quic_initial_awg.go): a second full
fragmented Initial with its OWN fresh DCID. NOT a short-header (that
was device-blocked, commit 64ce4a47) and NOT a DCID-reuse 1-RTT (an
impossible QUIC state that reads anomalous) — two independent Initials
just look like two QUIC sessions starting, which a browser does
routinely.
- masqueI2 (masque_awg.go): dispatch — only ip=quic fills i2; dns/stun/sip
return "" (single-packet decoys).
- awgIpcLines (device_awg.go): wires masque i2 into the i2 slot; guards an
explicit user i2 alongside id/ip/ib as a conflict (mirrors the i1 guard).
Safe by construction: i1/i2 are separate UDP datagrams sent before the
independently-built MessageInitiation (send.go), so neither touches the
real handshake.
Tests (device_awg_test.go): ip=quic fills both i1 and i2 as valid
independent Initials (different DCID, both carry the SNI, first CRYPTO
offset≠0); non-quic leaves i2 empty; explicit-i2 conflict rejected.
Spec §9 updated from hypothesis to implemented + device-verified; §9.2
residual risks (retry head-of-line budget — i3..i5 kept empty), §9.3
what's verified vs deferred. IMPLEMENTATION_REPORT R11 added.
Full package green; gofmt/vet clean; lx-build ok; sing-box check passes
for ip=quic and rejects explicit-i2 conflict.
Add two analysis sections to the 009 spec and normalize naming to 009
(the feature lives here; the LxBox task 146 is only the upstream source
of requirements/device facts, not "our" number).
SPEC.md:
- §8 "Active probing — граница односторонней маскировки (гипотезы)":
H3 (high confidence) a one-sided client decoy is only as strong as
what the TARGET SERVER genuinely serves on that port; the dns/stun/sip
timeouts are consistent with three DPI models (passive
destination-reputation / protocol allowlist / active probing) — H1/H2
(active-probe wording) are medium-confidence, with the honest caveat
that ":2408 answers QUIC" is unverified (it's the WG port, not :443).
Falsifiable device tests T1–T3 (incl. T3: point ip=quic at a
non-QUIC-serving host → should time out, isolating the borrowed
responder from the QUIC bytes).
- §9 "Многопакетная QUIC-последовательность (гипотеза усиления)":
i1..i5 multi-packet design, with a line-by-line send.go proof it CANNOT
break the WARP handshake (decoys are separate UDP datagrams before the
independently-built MessageInitiation). Flags the critical regression
trap: a short-header i2 is exactly the construct commit 64ce4a47
deleted as device-blocked; DCID-reuse is likely a fingerprint, not a
win; retry amplifies head-of-line bytes. Hypothesis to device-test
(bar: i2 must be no worse than i1-only), NOT a shipping decision.
Naming: drop "§146 §N" cross-refs to the LxBox spec from kernel comments
(they don't resolve inside 009); keep the two honest external source
refs (the task file path + "источник device-фактов — LxBox-задача 146").
Also refresh the masque_awg.go header (profiles are now Initial / query /
Binding Request / INVITE, not the old short-header/response list).
Docs only; no code-logic change (build/tests/gofmt/vet green).
ip=sip emitted a `SIP/2.0 200 OK` response as the client's first,
unsolicited packet — a server-role packet in the client's slot, the same
wrong-direction anomaly the old STUN/DNS profiles had, and it was missing
the Contact/Max-Forwards a flow opener needs. Replace it with a SIP
INVITE request carrying an SDP offer — what a UA legitimately sends first
to start a call.
New sip_invite_awg.go (masqueSIPInviteCPS):
- request-line INVITE sip:<user>@<host> SIP/2.0 (method, not a status).
- Via(branch=z9hG4bK)/Max-Forwards:70/From(tag)/To(no tag yet)/Call-ID/
CSeq:N INVITE/Contact, Content-Type: application/sdp, exact
Content-Length, SDP body (v=0, m=audio, rtpmap PCMU/PCMA/telephone-event).
- Hybrid randomization (no cross-user signature): pronounceable user names
and (when id is empty) the host come from PseudoGen and are baked into
<b> at build time (unique between users); volatile tokens (branch /
From-tag / Call-ID / CSeq / SDP session-id+version) are per-packet
<rc>/<rd> of fixed width, so Content-Length stays exact.
New pseudo_gen_awg.go: pronounceable pseudo names / hosts / public IPs
(ported from the LxBox PseudoGen §127) — plausible without being a
hardcoded RFC beacon (bob@biloxi.com) or obvious garbage, and never a
private IP. crypto/rand, not seeded.
id is now OPTIONAL for sip (empty → pgHost()); required only for quic
(SNI) and dns (QNAME); stun ignores it. Removed masqueSIPResponseCPS.
HONEST STATUS: not device-tested on WARP, but expected to time out like
dns/stun — SIP to the datacenter WARP edge :2408 is the same
destination-class anomaly (SIP lives on :5060 / a SIP server). The INVITE
form fixes the direction anomaly of the old 200 OK but not the
destination one. QUIC remains the only proven WARP mechanism; sip is the
strictly-better, direction-clean form kept for other providers.
Tests: TestMasqueSIPResponseStructure → TestMasqueSIPInviteStructure +
TestMasqueSIPInviteNoID (request-line, To-without-tag, exact
Content-Length, names not hardcoded, no-id → pseudo-host). Validation
test updated (id required for quic/dns only). Docs updated.
ip=dns emitted an EDNS OPT *response* (QR=1) as the client's first,
unsolicited packet — a wrong-direction anomaly (a response is a
server-role packet), the same defect STUN had. Replace it with a client
DNS *query* (QR=0, QTYPE HTTPS/65): what a client legitimately sends
first. Only two wire changes from the old code — FLAGS 0x8180→0x0100 and
QTYPE 0x0001→0x0041 — everything else (encodeDNSName, OPT RR, 0xFDE9
cover option) reused. Renamed masqueDNSResponseCPS → masqueDNSQueryCPS;
TXID/cover stay fresh per packet (<r 2>/<r 40>).
DEVICE RESULT (honest): the DNS query also TIMED OUT on the target
LTE/WARP DPI, as the design predicted. Confirms the fundamental finding:
packet quality and direction (request vs response) are secondary — the
blocker is the (protocol + destination) pair. The DPI cuts DNS/STUN/SIP
to the WARP edge 162.159.x:2408 as a protocol class, because raw
DNS/STUN/SIP to a datacenter IP is itself anomalous (DNS lives on :53, a
resolver — not a datacenter edge). QUIC alone bypasses the destination
check: QUIC/HTTP3 legitimately goes anywhere (the whole HTTP/3 web), so
QUIC to a Cloudflare IP is expected traffic.
So QUIC remains the only proven mechanism on this provider. The DNS query
is committed as the strictly-better (direction-correct, RFC-clean) form
and kept — like stun/sip — for other providers whose DPI only checks
well-formedness, NOT protocol-to-destination. Marked not-confirmed-on-WARP
in SPEC / EXAMPLES / IMPLEMENTATION_REPORT / lx-config.
Test: TestMasqueDNSResponseStructure → TestMasqueDNSQueryStructure
(QR=0, QNAME round-trips, QTYPE HTTPS, OPT to end). Docs updated.
QUIC (the proven mechanism) hardened, and the dud STUN profile rebuilt
as the strongest possible shape for other providers — guided by device
A/B on the target LTE/WARP DPI (only QUIC passes there; STUN is blocked
as a protocol class regardless of packet quality).
QUIC — randomize the fragment layout per call + robustness knobs:
- planFragmentsN: random cut points (was the fixed etalon offsets).
- randomizedWirePlan: random out-of-order CRYPTO permutation, repaired so
the offset-0 fragment is never first; PING/PADDING woven into random
gaps; one flex PADDING run pins the payload to the length field. I1–I4
hold by construction (stress test: 300 random packets).
- quicGenParams knobs (default 6 frags / 2 PING / 1250B): fragment count,
PING count, datagram-size range — escalation without a code change if a
DPI ever starts keeping a reassembly buffer. Length field / payload are
recomputed from the chosen size.
- Removed the now-dead etalonWirePlan / etalonCutpoints / planFragments.
STUN — Binding Request instead of Success Response:
- New stun_request_awg.go: a full WebRTC connectivity check (USERNAME,
ICE-CONTROLLING, PRIORITY, SOFTWARE=libwebrtc, MESSAGE-INTEGRITY
HMAC-SHA1, FINGERPRINT CRC-32), fresh txn/ufrag/key per call.
- A response sent unsolicited as the client's first packet is a
wrong-direction anomaly; a request is what an ICE client sends first.
- Removed masqueSTUNResponseCPS (+ orphaned be32/stunSoftwareLen).
- HONEST: this did NOT pass the target DPI (Timeout, like the old
response) — that DPI blocks STUN to a datacenter Cloudflare IP as a
class. Kept as the best shape in case another provider's DPI only
checks well-formedness. QUIC stays the only proven mechanism.
Tests: TestQUICInitialRandomizedInvariants (80 samples, I1–I4 + offsets
differ), TestQUICInitialRobustnessKnobs (4/10/12 frags, variable size),
TestMasqueSTUNRequestStructure (type 0x0001, FINGERPRINT verifies,
USERNAME+MESSAGE-INTEGRITY present), TestMasqueSTUNRequestUniqueness.
Docs: SPEC.md / IMPLEMENTATION_REPORT.md updated incl. the device record.
QUIC ip=quic now generates an out-of-order fragmented Initial (146,
commit 64ce4a47), not a 1-RTT short header. Rewrite the 009 spec docs to
describe the as-built state only — no design history, no superseded
short-header rationale, no open forks.
- SPEC.md: rewritten — I1 CPS mechanism, fragmented QUIC Initial with
I1–I4 invariants, crypto, validation, file map; drops the revert note,
the S1–S4 fork, and the short-header design.
- IMPLEMENTATION_REPORT.md: rewritten as a register of decisions R1–R7,
each with rationale and code refs (file:line) + commits.
- TASKS.md: clean checklist of the current state (no amendment block,
no strikethrough); device-smoke on DPI marked passed.
- EXAMPLES.md: drop §146-amendment blocks and the dangling masque_quic_awg.go
reference; fix the wrong "ib selects a ClientHello profile" claim (ib does
not affect the bytes); drop the dead PLAN.md link.
- Delete PLAN.md and HANDOFF_PROMPT.md (process docs, obsolete post-impl).
Docs only; no code change.
The ip=quic masquerade emitted a QUIC 1-RTT short header, which was
empirically BLOCKED by a real LTE-operator DPI (device-proven A/B, LxBox
task §146). Replace it with a full out-of-order fragmented QUIC Initial
(RFC 9001): a realistic browser-shaped ClientHello (id as SNI) split
across 6 CRYPTO frames in a permuted wire order — first frame offset≠0,
offset-0 frame near the end, PING/PADDING interleaved. A line-rate DPI
grabs the first frame, assumes offset 0, parses garbage and fails open;
a real QUIC server reorders the frames normally. This reverses the old
short-header rationale (the ≥1200-byte Initial it called impossible is
exactly what RFC 9000 §14.1 mandates, and the short header lost on DPI).
- New: quic_initial_awg.go (varint encoder, fragment plan + I1–I4
invariants, RFC 9001 Initial assembly), quic_clienthello_awg.go
(realistic ~294B TLS 1.3 ClientHello), quic_crypto_awg.go (HKDF /
AES-128-GCM-XOR-nonce / header protection, mirrored byte-for-byte from
common/sniff qtls so the keys match the live sniffer).
- id is now REQUIRED for ip=quic (it becomes the ClientHello SNI):
required for quic/dns/sip, optional only for stun.
- A flex PADDING run pins the payload to the length field for any SNI
length, so a long (≤253B) valid domain no longer overflows generation.
- Deleted masque_quic_awg.go (masqueQUICShortHeaderCPS / quicFirstByte).
- CPS transport unchanged: whole encrypted Initial emitted as one <b>
blob; fresh DCID + TLS random + ephemeral x25519 baked in per call.
- Tests reverse-parse our own output (decrypt, frame-walk, reassemble,
SNI) — §5 control vectors, I1–I4, uniqueness, long-SNI regression.
Cross-checked: the live common/sniff QUIC sniffer parses our Initial
and classifies it as chromium.
- Docs: README/README.ru/lx-config + SPECS/009 updated (id required for
quic, fragmented-Initial mechanism, short-header marked superseded).
Device-smoke on the blocked LTE network is a manual gate (not CI):
tunnel up + real traffic through DPI, control node alongside.
Add a Masquerade id/ip/ib row to the feature table and an AmneziaWG sugar
subsection (quic/dns/stun/sip, id required only for dns/sip, ib quic-only) in
both English and Russian READMEs. Link SPECS/009 report + examples.
Add declarative masquerade fields id (domain) / ip (protocol) / ib (browser)
on a wireguard endpoint — WireSock-style sugar over the AmneziaWG I1 CPS string.
Profiles quic/dns/stun/sip generate a protocol-shaped decoy packet, ported in
structure from the open-source WireSock reference (amneziawg-proxy/src/
transform.rs, MIT).
Mechanism: I1 CPS only (S1-S4 padding is impossible against Cloudflare WARP,
the target this eases connecting to); the vendored wireguard-go submodule is
untouched. QUIC is a 1-RTT short header (no SNI/ClientHello/JA3), matching
WireSock — no false "byte-perfect"/"fingerprint" claims. id is required only for
dns/sip (it lands on the wire as QNAME / SIP host) and optional for quic/stun;
when set it is always LDH-validated (mirror of is_valid_sni_hostname) as a
security boundary against SIP/DNS injection. ip is mandatory whenever any of
id/ip/ib is set; id/ip/ib are mutually exclusive with an explicit i1.
Gated by with_awg (rejected with a clear error otherwise); empty id/ip/ib leave
the config byte-identical to upstream. Tests assert each profile parses back as
its protocol (not tautologies); every CPS spec was verified against the real
amneziawg-go newObfChain. sing-box check passes for all four profiles and
rejects the conflict/injection/bad-value cases. See SPECS/009.
SPEC/REPORT/TASKS updated: selector-in-the-middle is now covered by a
runtime selector-guard (suspend AWG consumers before the switch), no longer
an uncovered case. Two complementary guards: Start-guard (static chain) +
selector-guard (runtime selector switch).
Start-guard covers a static detour chain but stops at a selector (its
chosen member is runtime-resolved). This adds the runtime half: in
Selector.SelectOutbound, BEFORE committing the switch, if the new member
reaches a wireguard endpoint, walk up the reverse-dependency ledger
(OutboundManager.ConsumersOf) and SuspendAmneziaWG() every AmneziaWG
consumer of the group — device down, started=false. Suspending before
s.selected.Store closes the race: by the time the group points at the WG
member, the consumer is down and a reconnect fails with "not ready"
instead of sending a junk handshake into WireGuard.
New adapter.AmneziaWGSuspendable marker + OutboundManager.ConsumersOf let
protocol/group act without importing protocol/wireguard. Plain-WG and
non-AWG consumers are left untouched. Variant B throughout.
Refs #2
SPEC/REPORT/memory updated: lazy dialer-guard removed (unverifiable +
sync.Once stale), selector-in-the-middle is a known uncovered case pending
a reliable selector-switch hook.
The lazy DetourDialer guard (lx.8) never fired on device — the hang is in
Endpoint.Start, before any dial — and it is unverifiable in the LxBox UI
(detour targets real servers, not groups) and sync.Once-caches its verdict
so it can't catch a selector changing at runtime. Revert
common/dialer/{detour,dialer}.go to upstream and remove its test.
The Start-guard in protocol/wireguard (field-verified on lx.9) stays as the
sole guard. Selector-in-the-middle is now a known uncovered case.
Refs #2
Field feedback: the user-facing message named Android / kernel hang, but
the restriction is architectural — amneziawg over wireguard is not
supported, period. Reword both guards (Start + dialer) to that; keep the
why (Android hang) in code comments for developers.
Refs #2
lx.8 lazy-only guard didn't fire on device (hang in Start before dial,
logcat-proven). SPEC/REPORT updated: two echelons (Start-guard for direct
transitive detour chain, dialer-guard for selector-in-the-middle).
The lazy DetourDialer guard (lx.8) never fired on Android: an AWG node
whose detour reaches a wireguard endpoint hangs synchronously in
Endpoint.Start (peer-domain resolve over the detour + junk handshake),
before any dial. Proven by logcat — kernel stuck in Starting, no guard
error logged.
Add a Start-guard in protocol/wireguard.Endpoint.Start: walk the
transitive detour chain (OutboundManager + Dependencies); if it reaches a
type=wireguard endpoint, log and skip device startup (started stays false)
so the instance comes up and other outbounds keep working — variant B,
never abort start. Stops at selector/urltest groups (runtime target),
leaving that case to the lazy dialer guard, which stays as the second
echelon.
Refs #2
- 007 AWG_OVER_WIREGUARD_DETOUR_GUARD: SPEC/PLAN/TASKS/REPORT, status C
- 008 AWG_JUNK_PARAM_VALIDATION: SPEC/PLAN/TASKS/REPORT, status C
- SPECS/README roadmap rows for 007 (#2) and 008 (#3)
An AmneziaWG node with detour into any wireguard-based endpoint (plain WG
or AWG) ends up tunnelling AWG traffic inside WireGuard, which hangs the
kernel on Android. Guard it in DetourDialer.init() like the empty-direct
check: lazy error, so the instance still starts and other outbounds keep
working while this node fails every dial (variant B).
Owner-is-AWG flows in via dialer.Options.IsAmneziaWG; the target is matched
by Type()==wireguard, expanding selector/urltest groups recursively. Detour
into a non-wireguard outbound (vless, …) and WG->AWG stay allowed.
Fixes#2
amneziawg-go sizes junk packets rand(0..jmax-jmin)+jmin before each
handshake; jmin>jmax makes rand.Int's argument <=0 and panics in the
retransmit-timer goroutine. validateJunk rejects it in awgIpcLines so the
config fails at endpoint build / sing-box check instead of crashing later.
Only the crash case is guarded; jc/size inconsistency stays allowed
(harmless, keeps the diff minimal, avoids rejecting working configs).
Fixes#3
The 006 spec folder was committed under its -N- name, then renamed to -C-
on disk without removing the old paths from git, leaving a duplicate
SPEC/PLAN/TASKS under 006-F-N-. Drop it; 006-F-C- is the canonical set.
Cosmetic (SPECS docs only) — does not affect release artifacts.
Replace the hard 'exactly two features and nothing else' with a small
client-side feature set (currently XHTTP + AWG2). New features are
allowed over time if they meet the constitution criteria: not planned
upstream, isolated per §3.2-3.3 (new files, own build tag, marked
seams), full Spec Kit cycle. Mirrored in both READMEs and lx-config.md.
lx-ci linux_musl smoke green for amd64/arm64/armv7/mipsle-softfloat —
all statically linked, no libdl.so.2, naive preserved. mipsle+naive
built with musl static, no fallback needed.
Widening H1..H4 from uint32 to MagicHeader (005) changed the longest type
in the struct, so gofmt re-aligns the json tags. go vet didn't catch it;
the lx-ci gofmt check did (red since lx.6). Format-only, no behavior change.
lx-release.yml: new build_linux_musl job (amd64/arm64/armv7/mipsle) that
clones cronet-go, fetches the Chromium musl toolchain via cmd/build-naive,
and builds CGO_ENABLED=1 with with_musl (swapping with_purego) so libcronet
is linked statically — no libdl.so.2, runs on musl routers, naive kept.
Linux moves out of the desktop build job. Artifact names mirror upstream
arch suffixes (armv7, mipsle-softfloat) without the -musl suffix since
Linux ships a single (musl) variant.
lx-ci.yml: dispatch-only linux_musl smoke job runs the same pipeline
(build + verify statically-linked / no libdl) without publishing.
Keep NaiveProxy (upstream feature) by mirroring upstream build.yml's musl
path instead of dropping it. Closes the libdl.so.2 failure on AsusWRT
Merlin + adds linux-armv7. CI-only, no Go code. See issue #1.
awg2_ranged.json (fake keys) exercises ranged H1-H4 through sing-box
check; wire it into the positive and negative CI checks alongside
awg2_basic.json.
Route h1..h4 through the writeStr path as canonical spec strings
("N" or "N-M") instead of writeUint, re-validating each with the key
name in the error. Unset headers are omitted, so a plain WireGuard
endpoint still yields a byte-identical device config.
Add option.MagicHeader (string-based, comparable): a single uint32 or an
inclusive "N-M" range (AWG 2.0 ranged headers from awg2 exports).
- UnmarshalJSON accepts a JSON number (backward compatible with the prior
uint32 field) and a JSON string "N"/"N-M"; canonicalizes, 0 -> unset.
- MarshalJSON keeps type fidelity: single value -> number, range -> string.
- Spec() re-validates for options built in code (libbox/launcher) bypassing
JSON. string base keeps AmneziaWGOptions comparable so IsSet() still works.
H1..H4 change from uint32 to MagicHeader.
README claimed naive/cronet builds CGO-free on "every target"; that is now
inaccurate — the windows/386 legacy (Win7) build drops with_naive_outbound
because cronet-go has no windows/386. Corrected README EN/RU and added the
caveat to docs/lx-config.md §3.
Mirrors upstream build.yml: the windows/386 leg uses a Win7-patched Go
(.github/setup_go_for_windows7.sh — MetaCubeX/go reverts of the Win7
removals) so the binary runs on Windows 7. Drops with_naive_outbound for
this leg (cronet-go has no windows/386 build); the rest of LX_TAGS
compiles for 386. Archive: sing-box-<ver>-windows-386-legacy-windows-7.zip,
matching the launcher's singbox-launcher-win7-32 (also 386).
AmneziaWG s3/s4 prepend junk to every transport message, so a plain-WG
MTU overflows the path and data packets fail with EMSGSIZE while the
handshake still succeeds. On an AWG endpoint (max(s3,s4) > 0):
- when mtu is unset, default to the recommended 1280 (not upstream 1408)
- when mtu is set too high for a conservative 1492-byte (PPPoE) budget,
log an advisory warning: mtu <= 1492 - 28 - 32 - max(s3,s4)
Plain WireGuard is untouched. Docs: lx-config.md §2 + SPECS/003 report.
Upstream-file edit (// lx:no-tailscale marker): remove with_tailscale (+ts_omit_*)
from build_libbox sharedTags. Client fork has no tailscale endpoints and it is the
largest dependency in the APK; keeps the AAR aligned with the desktop LX_TAGS set.
- IMPLEMENTATION_REPORT: live test vs real Xray (3x-ui) — packet-up/auto pass
(handshake + DNS + HTTPS + 2 MB download); padding fix (x_padding in Referer);
stream-one has a known downlink-framing bug
- README (EN/RU), docs/lx-config.md, SPECS roadmap updated; 002 folder -O- -> -C-
Verified live against a real Xray (3x-ui) XHTTP server (VLESS + Reality):
- padding must be carried as x_padding=<zeros> inside the Referer header
(Xray default PlacementQueryInHeader), not a standalone X-Padding header —
the server validates x_padding length (default 100-1000) and replies 400 Bad
Request when it is missing/out of range.
- auto now maps to packet-up (validated working); stream-one has a known
downlink-framing bug ("unknown version") and must be selected explicitly.
packet-up/auto: handshake + DNS + HTTPS + 2 MB download all flow through the tunnel.
- README.md is now the lx README in English (GitHub renders it → an arriving
visitor immediately sees this is a thin sing-box fork with XHTTP + AmneziaWG 2.0)
- README.ru.md: Russian version; mutual language switcher in both
- drop the static README.sing-box.md copy (it would go stale) in favor of a link
to the live upstream sing-box README on GitHub
- docs/lx-config.md: config reference for XHTTP transport and AmneziaWG 2.0
endpoint (field tables + examples with placeholder keys)
- .github/workflows/lx-ci.yml: matrix over the two features
(baseline / with_xhttp / with_awg / full) + negative check that feature-off
rejects its config; vet job; cross-platform matrix {linux,darwin,windows}x
{amd64,arm64} building the full lx set with the merged AWG fork (submodules)
- link the config doc from SPECS/README
- replace github.com/sagernet/wireguard-go => ./submodules/wireguard-go
(Leadaxe/wireguard-go @27290b6: sagernet base + AmneziaWG obfuscation, 3-way merge)
- add S3/S4 padding to option.AmneziaWGOptions + device_awg.go IpcSet emitter
(AWG 2.x; server config carries s1/s2/s3/s4)
LIVE-VALIDATED against a real AmneziaWG 2.0 server: handshake initiation ->
received handshake response -> keepalive -> traffic egresses via the server.
AWG is now functional, not just config-valid. Secrets never committed.
Verified against XTLS/Xray-core splithttp source (no live server):
- sessionId now formatted as dashed UUID (was 32-char hex) to match uuid.New().String()
- documented version-dependent padding placement (current Xray: x_padding query
param in Referer; older: standalone X-Padding) for live-test reconciliation
build + check + vet green.
- mark registry refactor + xhttp constant complete (pushed earlier)
- SPEC §7: hiddify port pulls vendored common/xray/* + quic-go/http3; record
faithful-vendor (A) vs lean-native (B) decision; recommend A
- status N -> O (in progress)
Return SUCCESS with empty answers instead of an error when the
queried address family has no range configured. Reject configurations
where neither inet4_range nor inet6_range is set.
The TTL computation and assignment loops treat OPT record's Hdr.Ttl
as a regular TTL, but per RFC 6891 it encodes EDNS0 metadata
(ExtRCode|Version|Flags). This corrupts cached responses causing
systemd-resolved to reject them with EDNS version 255.
Also fix pointer aliasing: storeCache() stored raw *dns.Msg pointer
so subsequent mutations by Exchange() corrupted cached data.
- Skip OPT records in all TTL loops (Exchange + loadResponse)
- Use message.Copy() in storeCache() to isolate cache from mutations
Treat rule_set items as merged branches instead of standalone boolean
sub-items.
Evaluate each branch inside a referenced rule-set as if it were merged
into the outer rule and keep OR semantics between branches. This lets
outer grouped fields satisfy matching groups inside a branch without
introducing a standalone outer fallback or cross-branch state union.
Keep inherited grouped state outside inverted default and logical
branches. Negated rule-set branches now evaluate !(...) against their
own conditions and only reapply the outer grouped match after negation
succeeds, so configs like outer-group && !inner-condition continue to
work.
Add regression tests for same-group merged matches, cross-group and
extra-AND failures, DNS merged-branch behaviour, and inverted merged
branches. Update the route and DNS rule docs to clarify that rule-set
branches merge into the outer rule while keeping OR semantics between
branches.
Before 795d1c289, nested rule-set evaluation reused the parent rule
match cache. In practice, this meant these fields leaked across nested
evaluation:
- SourceAddressMatch
- SourcePortMatch
- DestinationAddressMatch
- DestinationPortMatch
- DidMatch
That leak had two opposite effects.
First, it made included rule-sets partially behave like the docs'
"merged" semantics. For example, if an outer route rule had:
rule_set = ["geosite-additional-!cn"]
ip_cidr = 104.26.10.0/24
and the inline rule-set matched `domain_suffix = speedtest.net`, the
inner match could set `DestinationAddressMatch = true` and the outer
rule would then pass its destination-address group check. This is why
some `rule_set + ip_cidr` combinations used to work.
But the same leak also polluted sibling rules and sibling rule-sets.
A branch could partially match one group, then fail later, and still
leave that group cache set for the next branch. This broke cases such
as gh-3485: with `rule_set = [test1, test2]`, `test1` could touch
destination-address cache before an AdGuard `@@` exclusion made the
whole branch fail, and `test2` would then run against dirty state.
795d1c289 fixed that by cloning metadata for nested rule-set/rule
evaluation and resetting the rule match cache for each branch. That
stopped sibling pollution, but it also removed the only mechanism by
which a successful nested branch could affect the parent rule's grouped
matching state.
As a result, nested rule-sets became pure boolean sub-items against the
outer rule. The previous example stopped working: the inner
`domain_suffix = speedtest.net` still matched, but the outer rule no
longer observed any destination-address-group success, so it fell
through to `final`.
This change makes the semantics explicit instead of relying on cache
side effects:
- `rule_set: ["a", "b"]` is OR
- rules inside one rule-set are OR
- each nested branch is evaluated in isolation
- failed branches contribute no grouped match state
- a successful branch contributes its grouped match state back to the
parent rule
- grouped state from different rule-sets must not be combined together
to satisfy one outer rule
In other words, rule-sets now behave as "OR branches whose successful
group matches merge into the outer rule", which matches the documented
intent without reintroducing cross-branch cache leakage.
PreMatch and full match phases each created a fresh InboundContext,
causing process search (expensive OS syscalls) to run twice per
connection. Use a freelru ShardedLRU cache with 200ms TTL to serve
the second lookup from cache.
Add fpm-based Alpine APK packaging alongside existing DEB/RPM/Pacman
packages. Alpine APKs use `linux` in the filename to distinguish from
OpenWrt APKs which use the `openwrt` prefix.
CCM: Fix 1M context detection - use prefix match for versioned
beta strings (e.g. "context-1m-2025-08-07") and include cache
tokens in the 200K threshold check per Anthropic billing docs.
OCM: Add GPT-5.4 family pricing (standard/priority/flex) with
extended context (>272K) premium pricing support. Add context
window tracking to usage combinations, mirroring CCM's pattern.
Update normalizeGPT5Model defaults to latest known models.
Support the OpenAI Responses WebSocket API (`wss://.../v1/responses`)
for bidirectional frame proxying with usage tracking.
Fix Codex CLI client config examples to use profiles and correct flags.
Update openai-go v3.24.0 → v3.26.0.
When clients (e.g. Node.js Anthropic SDK) explicitly set Accept-Encoding: gzip,
Go's http.Transport does not transparently decompress the response body, because
it only does so when it added the header itself. This causes CCM's json.Unmarshal
to receive raw gzip bytes, silently failing to parse usage data and leaving the
usage counter unchanged.
Fix: remove Accept-Encoding from the outgoing proxy request. Transport adds it
automatically and transparently decompresses response.Body before CCM reads it.
Wire compression (CCM→Anthropic) is preserved — Transport still negotiates gzip.
Only CCM→localhost path is affected; compression on loopback has no practical
benefit.
Move hardcoded build tags and ldflags from Makefile, Dockerfile, CI
workflows, and local build scripts into canonical files under release/:
- release/DEFAULT_BUILD_TAGS (Linux common archs, Darwin, Android)
- release/DEFAULT_BUILD_TAGS_WINDOWS (includes with_purego)
- release/DEFAULT_BUILD_TAGS_OTHERS (no with_naive_outbound)
- release/LDFLAGS (shared linker flags)
The cache deduplication in Client.Exchange uses a channel-based lock
per DNS question. Waiting goroutines blocked on <-cond without context
awareness, causing them to accumulate indefinitely when the owning
goroutine's transport call stalls. Add select on ctx.Done() so waiters
respect context cancellation and timeouts.
When bbolt encounters corrupted page data at runtime, it panics
instead of returning an error. Wrap all DB transactions with
recover to catch these panics, delete the corrupted database
file, and reopen a fresh one.
- Enable ECH for NaiveProxy outbound with DNS resolver integration
- Add query_server_name option to override domain for ECH HTTPS record queries
- Update cronet-go dependency and remove windows_386 support
Align dev-next-grpc with wip2 by adding UsePlatformWIFIMonitor()
to the new PlatformInterface, allowing platform clients to indicate
they handle WIFI monitoring themselves.
We mistakenly believed that `libresolv`'s `search` function worked correctly in NetworkExtension, but it seems only `getaddrinfo` does.
This commit changes the behavior of the `local` DNS server in NetworkExtension to prefer DHCP, falling back to `getaddrinfo` if DHCP servers are unavailable.
It's worth noting that `prefer_go` does not disable DHCP since it respects Dial Fields, but `getaddrinfo` does the opposite. The new behavior only applies to NetworkExtension, not to all scenarios (primarily command-line binaries) as it did previously.
In addition, this commit also improves the DHCP DNS server to use the same robust query logic as `local`.
Previously, the buffer was not reset within the response loop. If a packet
handle failed or completed, the buffer retained its state. Specifically,
if `ReadPacketFrom` returned `io.ErrShortBuffer`, the error was ignored
via `continue`, but the buffer remained full. This caused the next
read attempt to immediately fail with the same error, creating a tight
busy-wait loop that consumed 100% CPU.
Validates `buffer.Reset()` is called at the start of each iteration to
ensure a clean state for 'ReadPacketFrom'.
The cache lookup was performed before rule matching, using the caller's
strategy (usually AsIS/0) instead of the resolved strategy. This caused
cache misses when ipv4_only was configured globally but the cache lookup
expected both A and AAAA records.
Remove LookupCache and ExchangeCache from Router, as the cache checks
inside client.Lookup and client.Exchange already handle caching correctly
after rule matching with the proper strategy and transport.
The Chinese documentation incorrectly stated that the default value for the domain_strategy field in the direct outbound module is dns.strategy. The correct value should be inbound.domain_strategy, as specified in the English documentation. This commit corrects the Chinese documentation to align with the accurate behavior described in the English version.
Signed-off-by: Monica <1379531829@qq.com>
For historical reasons, sing-box's `domain_suffix` rule matches literal prefixes instead of the same as other projects.
This change modifies the behavior of `domain_suffix`: If the rule value is prefixed with `.`,
the behavior is unchanged, otherwise it matches `(domain|.+\.domain)` instead.
The `process_path` rule of sing-box is inherited from Clash,
the original code uses the local system's path format (e.g. `\Device\HarddiskVolume1\folder\program.exe`),
but when the device has multiple disks, the HarddiskVolume serial number is not stable.
This change make QueryFullProcessImageNameW output a Win32 path (such as `C:\folder\program.exe`),
which will disrupt the existing `process_path` use cases in Windows.
I read other rule_item_xxx.go files, they are all snake case. This description is showed on dashboard like yacd.
Signed-off-by: kkocdko <31189892+kkocdko@users.noreply.github.com>
This helps the daemon work better on IoT devices
like RaspberryPi.
According to systemd's documentation,
`network.target` means there has already been
a network manager started, but the network may
not be "up". On most PCs this does not matter
because the network will turn to "up" almost
immidiately. The IoT devices' network interface
may not be set up quickly enough, so they may
meet that the sing-box daemon is started before
network is ready, which results that sing-box
cannot find a working route. The workaround
of this is restarting sing-box daemon but it
absolutely is not the perfect solution.
As `network-online.target` must be triggered by
network manager after you configured it, I keep
`network.target` so there will be no change to
those who do not enabled proper trigger service
like `NetworkManager-wait-online.service`.
See also: https://systemd.io/NETWORK_ONLINE/
Refactor Authenticator interface to struct &
Update smux &
Update gVisor to 20231204.0 &
Update quic-go to v0.40.1 &
Update wireguard-go &
Add GSO support for TUN/WireGuard &
Fix router pre-start &
Fix bind forwarder to interface for systems stack
Enhanced the issue reporting templates for both English and Chinese versions by adding more structured and comprehensive guideline checkboxes. This aims to ensure contributors provide sufficient and beneficial information for reproducing and resolving issues, thereby improving the quality of reports and making issue tracking more efficient.
Remove the information on password generation for `2022-blake3-aes-128-gcm` cipher from the Server Example section in the shadowsocks.md file as it is no longer needed.
The old meaning is wrong. Correct the meaning according to the English documentation and the actual effect of the option.
Signed-off-by: 嫦悅 <lomombwlo@gmail.com>
description:Please provide the operating system version
validations:
required:true
- type:dropdown
attributes:
label:Installation type
description:Please provide the sing-box installation type
options:
- Original sing-box Command Line
- sing-box for iOS Graphical Client
- sing-box for macOS Graphical Client
- sing-box for Apple tvOS Graphical Client
- sing-box for Android Graphical Client
- Third-party graphical clients that advertise themselves as using sing-box (Windows)
- Third-party graphical clients that advertise themselves as using sing-box (Android)
- Others
validations:
required:true
- type:input
attributes:
description:Graphical client version
label:If you are using a graphical client, please provide the version of the client.
- type:textarea
attributes:
label:Version
description:If you are using the original command line program, please provide the output of the `sing-box version` command.
render:shell
- type:textarea
attributes:
label:Description
description:Please provide a detailed description of the error.
validations:
required:true
- type:textarea
attributes:
label:Reproduction
description:Please provide the steps to reproduce the error, including the configuration files and procedures that can locally (not dependent on the remote server) reproduce the error using the original command line program of sing-box.
validations:
required:true
- type:textarea
attributes:
label:Logs
description:|-
In addition, if you encounter a crash with the graphical client, please also provide crash logs.
For Apple platform clients, please check `Settings - View Service Log` for crash logs.
For the Android client, please check the `/sdcard/Android/data/io.nekohasekai.sfa/files/stderr.log` file for crash logs.
render:shell
- type:checkboxes
id:supporter
attributes:
label:Supporter
options:
- label:I am a [sponsor](https://github.com/sponsors/nekohasekai/)
- type:checkboxes
attributes:
label:Integrity requirements
description:|-
Please check all of the following options to prove that you have read and understood the requirements, otherwise this issue will be closed.
Sing-box is not a project aimed to please users who can't make any meaningful contributions and gain unethical influence. If you deceive here to deliberately waste the time of the developers, you will be permanently blocked.
options:
- label:I confirm that I have read the documentation, understand the meaning of all the configuration items I wrote, and did not pile up seemingly useful options or default values.
required:true
- label:I confirm that I have provided the server and client configuration files and process that can be reproduced locally, instead of a complicated client configuration file that has been stripped of sensitive data.
required:true
- label:I confirm that I have provided the simplest configuration that can be used to reproduce the error I reported, instead of depending on remote servers, TUN, graphical interface clients, or other closed-source software.
required:true
- label:I confirm that I have provided the complete configuration files and logs, rather than just providing parts I think are useful out of confidence in my own intelligence.
--body "Branch \`$BRANCH\` is rebased onto \`$TARGET\`, builds, and passes \`check\`. Auto-PR was blocked — enable Settings → Actions → General → \"Allow GitHub Actions to create and approve pull requests\", or open the PR by hand."
- **XHTTP** transport (\`with_xhttp\`) — Xray-compatible "splithttp", composes with Reality (use \`auto\`; \`stream-one\` has a known framing bug).
- **CommandClient extensions** (\`with_lx_command\`) — native libbox gRPC parity for the Clash API dropped from the **Android AAR**: URLTestOutbound, GetRules, GetGroups/GetOutbounds, Connection.Detour, SubscribeDNSQueries. (Desktop/CLI binaries keep \`with_clash_api\` for external dashboards.)
### Binaries
Drop-in \`sing-box\` for **darwin / windows** × {amd64, arm64}, plus a **Windows 7 (32-bit)** legacy build (\`sing-box-${{ steps.ver.outputs.version }}-windows-386-legacy-windows-7.zip\` — built with a Win7-patched Go; without naive/cronet, which has no windows/386 target).
**Linux — static musl builds for routers** (AsusWRT Merlin, OpenWrt, Keenetic): \`linux-amd64\`, \`linux-arm64\`, \`linux-armv7\`, \`linux-mipsle-softfloat\`. These are statically linked (no \`libdl.so.2\`/glibc dependency) and **keep NaïveProxy** — they run on musl routers where the previous dynamic builds failed with \`libdl.so.2: cannot open shared object file\`. See SPECS/006.
**Linux — big-endian MIPS** (OpenWrt \`mips_24kc\`, e.g. Atheros AR93xx): \`linux-mips-softfloat\` — pure-Go static build **without NaïveProxy** (Chromium/cronet has no big-endian MIPS toolchain); everything else matches the desktop tag set.
Each archive contains the \`sing-box\` binary (\`sing-box version\` reports \`${{ steps.ver.outputs.version }}\`). Verify downloads against \`SHA256SUMS\`.
### Android
\`libbox-${{ steps.ver.outputs.version }}.aar\` (+ \`libbox-legacy-…\` for SDK 21) — gomobile build of \`experimental/libbox\` with \`with_xhttp\`+\`with_awg\` enabled, for embedding in an Android app. \`Libbox.version()\` reports the lx version.
stale-issue-message:'This issue is stale because it has been open 60 days with no activity. Remove stale label or comment or this will be closed in 5 days'
`sing-box-lx` — **тонкий downstream** апстрима [SagerNet/sing-box](https://github.com/SagerNet/sing-box): upstream **плюс ровно две фичи** и ничего больше:
1.**XHTTP** — клиентский v2ray-транспорт (совместимость с Xray XHTTP).
Главная ценность проекта — **согласованность с upstream**. Любое изменение оценивается по тому, насколько легко оно переживёт ребейз на следующий тег upstream.
-`.github/workflows/lx-ci.yml` — build(lx tags) → version → `go vet` → `sing-box check` (linux/amd64; полная матрица — в 004).
-`lx-test/config/minimal.json` — валидный конфиг для `check` (mixed-in + direct-out). Положен в `lx-test/`, **не** в upstream `test/` (там отдельный Go-модуль).
-`AGENTS.md` — указатель для агентов (force-add: upstream его `.gitignore`-ит; новый файл → нулевой конфликт при ребейзе).
-`SPECS/**` — Spec Kit (CONSTITUTION, IMPLEMENTATION_PROMPT, README, задачи 001–004).
**Правок upstream-файлов: 0.**`constant/version.go`, `Makefile`, `.gitignore` — не тронуты.
## Проверки (DoD)
```
$ make -f Makefile.lx lx-version → 1.13.13-lx.1
$ make -f Makefile.lx lx-build → ./sing-box (28 MB)
Хранить в `Makefile` (переменная `LX_TAGS`) и продублировать в `SPECS/CONSTITUTION.md` при изменениях.
> **Обновлено в §004:** набор расширен до полного upstream feature-set (`release/DEFAULT_BUILD_TAGS`) + `with_purego` + наши две фичи, с обязательным `-checklinkname=0` в `LX_LDFLAGS`. Актуальный источник истины — `Makefile.lx` (`make -f Makefile.lx lx-print-tags`) и `SPECS/004`.
## 2. Изменяемые / новые файлы
| Файл | Тип | Изменения |
|------|-----|-----------|
| `Makefile` | new (или дополнение) | Цель `lx-build`: `go build -tags "$(LX_TAGS)" -ldflags "$(LX_LDFLAGS)" -o sing-box ./cmd/sing-box`; переменные `LX_TAGS`, `VERSION=…-lx.$(LX_BUILD)` |
| `lx-test/config/*.json` | new | Sample-конфиги для `sing-box check` (минимальный валидный, без фич — для 001) |
| `SPECS/001-.../IMPLEMENTATION_REPORT.md` | new | Отчёт |
> Версия: upstream хранит строку версии в `constant/version.go` (или собирается через ldflags в `cmd/sing-box`). Проверить фактический механизм и **задавать `-lx` суффикс через `-ldflags -X`**, не правя `constant/version.go` напрямую (иначе лишний `// lx:` дифф на каждый ребейз). Если upstream не поддерживает ldflags-override — тогда минимальная `// lx:` правка в `constant/version.go`.
## 3. Зона касания upstream
-В идеале **ноль** правок upstream-файлов (всё через новые файлы + ldflags).
- Допустимый минимум: одна `// lx:` строка в `constant/version.go`, если ldflags-override невозможен.
## 4. Порядок работ
1. Проверить механизм версии upstream (`constant/version.go`, `cmd/sing-box`).
2.`Makefile`с`LX_TAGS`/`LX_LDFLAGS`/`lx-build`.
3. Sample-конфиг + CI-скелет.
4. Прогнать DoD, заполнить отчёт.
## 5. Риски
- Версионный механизм upstream может не принимать ldflags-override — fallback на `// lx:` правку.
-`with_xhttp`/`with_awg` как несуществующие теги не ломают сборку (Go игнорирует неизвестные build-теги) — но файлов с этими тегами пока нет, это нормально.
Заложить скелет downstream'а`sing-box-lx`: remotes, рабочая ветка, build-теги, версия с `-lx`, конвенция маркеров `// lx:` и шаблон гейтинга. После задачи репозиторий — корректный «upstream + ноль фич», готовый принимать XHTTP (002) и AWG2 (003).
---
## 1. Проблема / контекст
`Leadaxe/sing-box-lx` — форк-зеркало upstream (родословная `SagerNet/sing-box` сохранена). Нужна повторяемая инфраструктура downstream'а, при которой каждое будущее изменение изолировано и ребейзопригодно (см. CONSTITUTION § 3).
- Ветка `lx` базируется на стабильном теге `v1.13.13`. **(сделано)**
- Default branch на GitHub = `lx`; шумные зеркальные ветки (`dependabot/*`, `dev-*`, `copilot/*`) — вне внимания (можно удалить с origin, не обязательно).
### 2.2 Build-теги
- Ввести **`with_xhttp`** и **`with_awg`** как опознаваемые теги проекта (фактический код — в 002/003). Зафиксировать **канонический набор тегов сборки lx** в одном месте (см. PLAN), переиспользуемый в DoD и CI.
- Инвариант: без `with_xhttp`/`with_awg` бинарь ведёт себя как upstream.
### 2.3 Версия
-`sing-box version` должен печатать суффикс **`-lx.N`** (напр. `1.13.13-lx.1`).
- Суффикс задаётся при сборке (ldflags), не хардкодом в исходниках upstream (минимальный дифф).
## v2 — полная клиентская поддержка параметров (2026-06-29)
**Статус:** реализация code-complete + все проверки зелёные; **дефолтный путь лайв-подтверждён на реальных нодах** (4 живых XHTTP-сервера, packet-up + stream-one/reality, скачивание 1 МБ); **лайв obfs/placement** — остаётся открытым TODO (нужен сервер с такой настройкой).
Клиентский XHTTP-транспорт, подход **lean-native** (на примитивах sing-box, минимум зависимостей) — реализован многоагентным workflow в изолированном worktree, влит в `lx` (коммиты `2d97ff56` registry/const + `d1b434fc` транспорт).
**Файлы (новые, если не указано иное):**
-`transport/v2ray/registry.go`, `// lx` в `transport/v2ray/transport.go`, константа в `constant/v2ray.go` — registry-рефактор (ранее).
-`option/v2ray_xhttp.go` — тип `V2RayXHTTPOptions` (Host, Path, Mode, Headers, padding).
-`option/v2ray_transport.go` — **единственная upstream-правка** (// lx): поле `XHTTPOptions` + xhttp-case в Marshal/Unmarshal.
-`transport/v2rayxhttp/{client,conn,register}.go` — клиент; `register.go` под `//go:build with_xhttp`.
-`include/v2rayxhttp.go` (`//go:build with_xhttp`) — blank-import для запуска `init()`.
-`lx-test/config/xhttp_reality.json` — VLESS+xhttp+reality для `check`.
-`go vet` (lx-теги) по `transport/v2rayxhttp`, `option`, `transport/v2ray` → чисто; `go build ./...` без тегов → ок; `gofmt` чисто.
- Негатив: бинарь **без**`with_xhttp` отвергает xhttp-конфиг (`unknown transport type: xhttp`). Невалидный mode → `v2ray-xhttp: unknown mode`. Все 4 mode конструируются.
## Зона касания upstream (ребейз)
Ровно **1 файл**: `option/v2ray_transport.go` (3 правки в // lx-маркерах). Реестр и весь пакет `v2rayxhttp` — новые файлы, конфликтов не дают.
## Лайв-тест (реальный Xray/3x-ui XHTTP-сервер)
Проверено против VLESS + Reality + `type=xhttp` ноды (панель 3x-ui):
- ✅ **packet-up** (и `auto` → packet-up): handshake + DNS + HTTPS (example.com 200) + скачивание 2 МБ @ ~2.1 МБ/с — трафик выходит через IP сервера.
- ❌ **stream-one**: `unknown version` — баг при чтении downlink-ответа (выбирается только явно). → **Исправлено в задаче 011** (корень: stream-one должен слать голый путь без sessionId; auto+reality → stream-one). Принято на синтетике, лайв отложен.
**Ключевой фикс (по исходникам Xray hub.go/config.go + лайв):** padding кладётся как `x_padding=<нули>` в **query внутри заголовка `Referer`** (Xray default `PlacementQueryInHeader`, key `x_padding`), а**не** отдельным `X-Padding`. Сервер валидирует длину `x_padding` (дефолт 100–1000) и без неё отвечает **400 Bad Request**. Плюс `mode=auto` переключён на **packet-up**. Коммит `5a398a5e`. Также ранее: `sessionId` → UUID-формат, path-layout `<path>/<sessionId>[/<seq>]` сверены.
## Остаточные пробелы
1.~~**stream-one** — баг framing downlink (`unknown version`)~~ → **исправлено в 011** (голый путь без sessionId; `auto`+reality → stream-one). Лайв-подтверждение — открытый TODO в 011.
2.**packet-up** без xmux/переиспользования соединений; **stream-up** не лайв-тестился.
3.`x_padding_bytes` — строка «min-max» (нет Range-типа в badoption); дефолт 100–1000.
## Дальше
- Лаунчер: маппинг `type=xhttp` (его задача 023 сейчас маппит в `httpupgrade`) → реальный xhttp-транспорт.
| `serverMaxHeaderBytes` | `http.Server{MaxHeaderBytes}` — лимит размера заголовков входящего запроса | `8192` | У client-only транспорта нет `http.Server` |
| `noSSEHeader` | Сервер не шлёт `Content-Type: text/event-stream` на stream-down GET | `false` (SSE шлётся) | Клиент не читает Content-Type — обрабатывает оба случая без кода |
| `scMaxBufferedPosts` | Ёмкость серверной очереди переупорядочивания upload-POST (packet-up) | `30` | Клиент не знает о глубине буфера сервера |
| `scStreamUpServerSecs` | Интервал (сек, Range) периодической записи `X`-padding в ответ stream-up | `{20,80}` | Клиент тихо отбрасывает эти байты (`io.Discard`) |
**Решение для реализации:** не реализуем. В конфиге — `expose-but-ignore` (принимаем поля, чтобы
server-образные конфиги не падали на парсинге), помечены как inbound-only. Альтернатива (просто
document-and-skip без полей в struct) тоже допустима; финальный выбор зафиксирован в SPEC §6.
- Встроенные типы регистрируются в `init()` (в `transport.go` или соседнем файле) — поведение для http/ws/quic/grpc/httpupgrade без изменений.
-`NewClientTransport` → `ctor, ok := clientRegistry[options.Type]`; нет — прежняя ошибка.
- XHTTP-конструктор регистрируется из пакета `v2rayxhttp` через `init()`**только** под `//go:build with_xhttp` (через проводящий файл, чтобы импорт пакета подтягивался лишь с тегом).
Конструктор XHTTP должен соответствовать сигнатуре `ClientConstructor` (см. upstream `transport.go`): `(ctx, dialer, serverAddr, options, tlsConfig) → (adapter.V2RayClientTransport, error)`. Опции достаются из `options.XHTTPOptions`.
-`mode=auto` в sing-box-портах исторически падает в `packet-up`, что ломало, напр., аплоад в Telegram ([hiddify#2082](https://github.com/hiddify/hiddify-app/issues/2082)) — задокументировать фактический выбор режима.
- Рефактор `switch`→registry должен **точно** сохранить семантику ошибок и nil-обработку (`options.Type == ""` → `nil, nil`).
Добавить **клиентский XHTTP-транспорт** (совместимость с Xray XHTTP) для VLESS/VMess/Trojan, встроив его через **registry-рефактор** диспетчера v2ray-транспортов, за build-тегом `with_xhttp`.
---
## 1. Проблема / контекст
- Upstream sing-box XHTTP не поддерживает и не планирует ([#3550](https://github.com/SagerNet/sing-box/issues/3550)). Сервера на Xray всё чаще только XHTTP (после депрекации части транспортов в Xray).
- В sing-box диспетчер v2ray-транспортов — **хардкод-`switch`** по `options.Type` в `transport/v2ray/transport.go`. Добавлять `case` на каждый ребейз — точка постоянных конфликтов.
## 2. Цель
VLESS/VMess/Trojan outbound с`transport.type = "xhttp"` поднимают рабочее соединение к XHTTP-серверу Xray, в т.ч. поверх **TLS/Reality**. Без тега `with_xhttp` тип `xhttp` отвергается с понятной ошибкой.
- Превратить выбор клиентского транспорта в **реестр**: `transport.RegisterClient(type, ClientConstructor)` + `map[string]ClientConstructor`, заполняемый при `init()`.
- Встроенные транспорты (`http`, `ws`, `quic`, `grpc`, `httpupgrade`) регистрируются как раньше (поведение идентично upstream).
-`NewClientTransport` ищет конструктор в реестре вместо `switch` (поведение для известных типов — без изменений; для неизвестных — та же ошибка `unknown transport type`).
-`option/v2ray_transport.go`: поле `XHTTPOptions XHTTPOptions` в `_V2RayTransportOptions` + тип `XHTTPOptions` (в новом файле `option/v2ray_xhttp.go`, чтобы минимизировать дифф основного файла; в `_V2RayTransportOptions` — одна `// lx:` строка).
### 3.4 TLS/Reality
-`tlsConfig` прокидывается в конструктор как у прочих транспортов → связка **XHTTP + Reality** работает без доп. кода. (XHTTP + XTLS-Vision несовместимы — ограничение протокола, не наше.)
## 4. Критерии приёмки
-`sing-box check -c` принимает VLESS + `transport.type=xhttp` + `tls.reality`.
- Реальный коннект к XHTTP-серверу Xray (ручная проверка), хотя бы `mode=stream-one` и `packet-up`.
- Сборка **без**`with_xhttp`: конфиг с`xhttp` → ошибка `unknown transport type: xhttp` (или эквивалент реестра).
-`go test ./transport/...`, `go vet ./...` зелёные.
- Ребейз-проверка: при следующем upstream-теге конфликты возможны **только** в `transport/v2ray/transport.go`, `constant/v2ray.go`, `option/v2ray_transport.go`.
## 7. Разведка порта и выбор подхода (добавлено по ходу)
**Что показал референс `hiddify/hiddify-sing-box` (`transport/v2rayxhttp/client.go`):** XHTTP в hiddify реализован НЕ поверх примитивов sing-box, а через **вендорённое поддерево Xray** под `common/xray/{buf,net,pipe,signal/done,uuid}` + зависимости `quic-go`, `http3`, `golang.org/x/net/http2`, абстракция `DialerClient`/`XmuxClient` и опции `option.V2RayXHTTPOptions{ V2RayXHTTPBaseOptions }`. Целевой интерфейс прост — `adapter.V2RayClientTransport = { DialContext(ctx) (net.Conn, error); Close() error }` — но реализация тянет много транзитивного кода и завязана на старую версию sing-box hiddify.
**Развилка подхода (зафиксировать перед кодом порта):**
- **(A) Faithful-vendor.** Перенести hiddify `common/xray/*` + пакет `v2rayxhttp` как **новые файлы** (namespaced), адаптировать импорты под v1.13.13. Плюс: максимальная совместимость с реальными XHTTP-серверами, проверенный код. Минус: больший footprint (но всё — новые файлы → **нулевая зона касания upstream**, что согласуется с CONSTITUTION). Тащит `quic-go`/`http3` (часть уже в go.mod sing-box).
- **(B) Lean-native.** Написать компактный XHTTP-клиент на примитивах sing-box (по образцу in-tree `transport/v2rayhttpupgrade`). Плюс: меньше кода, меньше зависимостей. Минус: больше оригинальной работы и риск несовпадения с Xray по краям (`mode=auto`, padding, xmux).
**Рекомендация:****(A)** — приоритет проекта №2 (корректность/совместимость) важнее объёма, а изоляция в новых файлах сохраняет ребейзопригодность. Footprint велик, но не увеличивает конфликтность ребейза.
**Обязательно для приёмки:** живой XHTTP-сервер (Xray) для end-to-end проверки — синтетического `sing-box check` недостаточно (XHTTP под активной разработкой, версии client↔server должны совпадать).
> ⚠️ В `extra` эти значения часто приходят **числом** (`"scMaxEachPostBytes":"1000000"`,
> `"scMinPostsIntervalMs":30.0`). Транспорт sing-box-lx ждёт **строку `"min-max"`** — превратить
> одиночное число `N` в строку `"N-N"` (или просто `"N"` — парсер примет и то, и то). Дробную часть
> у `30.0` отбросить → `"30"`.
### 2.5 Игнорируемые / серверные
| URL-параметр | Действие |
|--------------|----------|
| `scMaxConcurrentPosts` | **Accept-but-ignore.** Legacy-поле старого Xray (в текущем Xray/extended его нет — там 1 POST-тело за раз). Клиент sing-box-lx шлёт upload-POST последовательно (= текущий Xray). Можно влить как `sc_max_concurrent_posts` (принято, но не используется) — или опустить (см. §6). |
| `serverMaxHeaderBytes`, `noSSEHeader`, `scMaxBufferedPosts`, `scStreamUpServerSecs` | server-only. Можно влить как `server_max_header_bytes`/`no_sse_header`/`sc_max_buffered_posts`/`sc_stream_up_server_secs` (клиент их принимает, но игнорирует) — или просто опустить. |
| `fragment`, `fm`, `fragment=...` | TLS-фрагментация (Xray-специфика). **Не часть XHTTP.** Маппить в свою TLS-fragment-фичу, если есть; иначе опустить. |
| `flow` | Для XHTTP всегда пустой (vision несовместим). |
## 6. Известные ограничения клиента (что НЕ маппить)
-`scMaxConcurrentPosts` — legacy-поле (удалено из текущего Xray-core и sing-box-extended; там upload сериализован в 1 POST-тело за раз). Наш клиент тоже шлёт последовательно = текущий Xray. Поле принимается (`sc_max_concurrent_posts`), но игнорируется.
-`downloadSettings` (асимметричный download-транспорт) — не поддержан; `mode=auto`+reality+downloadSettings
у нас всё равно даст stream-one, не stream-up.
-`spx` (spiderX), Xray browser-dialer — нет аналога.
- **HTTP/3 (`alpn=h3` / QUIC).** Наш XHTTP-клиент работает поверх **HTTP/2** (`http2.Transport`). Xray
умеет H1/H2/H3. Ноды, помеченные `alpn=h3`, мы обслуживаем по H2 (если сервер допускает); если сервер
**требует строго h3** — коннект не встанет. Это архитектурное ограничение транспорта, вне SPEC 002
(отдельная будущая задача «XHTTP over HTTP/3»). Парсеру: `alpn` маппить как есть, но `h3`-only ноды
помечать как потенциально неработающие.
-`fragment` / `fm` (TLS-фрагментация Xray) — не часть XHTTP; маппить в свою TLS-fragment-фичу (если есть)
или опускать.
---
## 7. Чек-лист для интегратора
- [ ]`type=xhttp` распознаётся как XHTTP-транспорт.
- [ ]`extra` декодируется как URL-encoded JSON и вливается в transport.
- [ ] Числовые `sc*`-поля из `extra` → строка `"min-max"`.
Хронология вендоренного wireguard-go: базы графта, миграции, что менял upstream. Актуальное состояние — в [SPEC.md](SPEC.md); здесь только «как было раньше и почему переделали».
---
## Почему граф, а не прямой `replace` на amneziawg-go
Первая идея — подключить `amnezia-vpn/amneziawg-go` напрямую через `replace`. **Не работает:** amneziawg-go основан на *upstream* wireguard-go и не имеет sagernet-добавок (`Send(offset)`, `InputPacket`, `conn` reserved/control), на которых держится `transport/wireguard` sing-box. Прямой replace ломает сборку.
Решение — **3-way graft**: обфускация Amnezia накладывается поверх `sagernet/wireguard-go` (а не наоборот). Так контракт sing-box↔device остаётся sagernet'овским, обфускация аддитивна. Форк-модуль — `Leadaxe/wireguard-go-awg2-lx`.
## База графта: эволюция
| Дата | Submodule commit | Sagernet-база | wireguard-go версия | Контекст |
-`9de6dc3 Add batched InputPackets` + `2c27bbf FIx batched InputPackets` — новый батч-вход `InputPackets([]*InputPacketRef) []*InputPacketRef` (возвращает unmatched refs — для L3-forward, где нет пира → вызывающий строит ICMP-unreachable). **`InputPacket` (singular) НЕ удалён** — переписан на size-based буфер + backpressure-кап `maxQueuedInputPackets`.
-`8403cdb Rework outbound buffer management` — **`QueueOutboundElement.buffer` сменил тип `*[MaxMessageSize]byte` → `[]byte`** (size-based пул через `GetOutboundBuffer(n)`/`PutOutboundBuffer` из sing-аллокатора, вместо фиксированного `messageBuffers`-пула). Элемент-пулы `outboundElements*` перешли с`WaitPool` на `sync.Pool`. Добавлен `peer.queuedOutboundPackets atomic.Int32` (backpressure-счётчик).
-`57baac9 Add batched UDP I/O on Darwin` + `fcbb7c4 Coalesce UDP GSO segments` — новый `conn/msgx_darwin.go` (sendmsg_x/recvmsg_x), GSO-iovec coalescing в `bind_std.go`.
**Оценка риска для графа ДО работы** (по памяти) была завышена: «`buffer`-type change ломает все AWG-хуки в send.go — основная работа». **По факту оказалось иначе:**
**Итог re-graft (`git apply --3way` граф-diff'а на v0.0.5):**
- **15 из 16** граф-файлов легли **чисто**. Конфликт — **только `send.go`**, и **на одной строке**: upstream добавил `peer.queuedOutboundPackets.Add(-…)` там, где граф добавил пустую строку. Взяли upstream (backpressure нужен).
- **Почему `buffer`-type change НЕ сломал граф:** AWG-хуки уже везде работают с `elem.buffer` как со **срезом** (`buffer[:MessageTransportHeaderSize]`, сдвиг `buffer[i+padding]`), а не как с массивом-указателем. Переход `*[N]byte → []byte` для них прозрачен.
- **Почему upstream `InputPacket`/`InputPackets` встали verbatim:** граф `send.go` их **не трогает** (junk-логика графа — в `SendHandshakeInitiation`, а не в input-пути), поэтому конфликта не было — upstream-версии сохранились.
- **Почему `RoutineEncryption` сшилась без ручного weave:** при `MessageEncapsulatingTransportSize = 0` upstream-offset `buffer[METS:METS+HeaderSize]` схлопывается к графовому `buffer[:HeaderSize]`. Граф-версия (заголовок в начале, без финального encapsulating re-slice) наложилась как есть.
**Вывод:** несущий инвариант `MessageEncapsulatingTransportSize = 0` — то, что делает re-graft дешёвым: он нейтрализует единственную точку, где upstream и граф расходятся по layout буфера.
Сборка после re-graft: device/conn/tun на linux/android/windows/darwin ✅; полный sing-box CLI с LX_TAGS (Go 1.24.7) ✅; тесты `transport/wireguard` + `protocol/wireguard` зелёные ✅.
## MTU / EMSGSIZE — находка 2026-06-10
При лайв-тесте AWG2-узла рукопожатие проходило, но трафик не шёл: `sendmsg: message too long` (**EMSGSIZE**). Причина — `S3`/`S4`: junk дописывается к **каждому** transport-сообщению, и обфусцированный data-пакет перерастает path MTU (1500, DF). Handshake маленький — проходит; transport — нет. Plain WG к тому же серверу с `mtu 1420` работает (S-junk нет).
Эмпирика (тот же узел, менялся только `mtu`, `S3=S4=60`):
| mtu | результат |
|----:|-----------|
| 1420 | ❌ EMSGSIZE |
| 1380 | ✅ ~58 ms |
| 1280 | ✅ ~55 ms |
| 1200 | ✅ ~60 ms |
Результат — MTU-политика в текущем SPEC.md (auto-default 1280 + warn при превышении бюджета). Источник находки — заметка агента лаунчера (`singbox-launcher`). Это не баг ядра, а размерный оверхед S-junk.
## Безопасность
Секреты живого AWG-сервера **никогда** не попадали в репозитории — лайв-конфиг держался только в `/tmp` и затирался (`shred`). Репо-конфиг `lx-test/config/awg2_basic.json` — с фейк-ключами.
-`transport/wireguard/device_awg.go` (`//go:build with_awg`) — `awgIpcLines()` шлёт IpcSet-ключи `jc=/jmin=/jmax=/s1..s4=/h1..h4=/i1..i5=` в device; `device_stub_awg.go` без тега даёт явную ошибку при заданных AWG-полях.
**2. amneziawg-go активирован через merged-форк (главное достижение):**
- amneziawg-go основан на *upstream* wireguard-go и не имеет sagernet-добавок (`Send(offset)`, `InputPacket`, `conn` reserved/control), на которых держится `transport/wireguard`. Прямой `replace` ломает сборку.
- Ключевое упрощение: **`MessageEncapsulatingTransportSize = 0`** — нейтрализует 8-байтный headroom sagernet (sing-box-lx его не использует), и обфускация Amnezia встаёт чисто без weave-конфликтов в send-пути.
-`conn/tun/ipc` оставлены **чисто sagernet** (обфускация только в `device/`: новые `obf*.go`+`magic-header.go` + графты в `send/receive/device/uapi`).
## MTU при ненулевых S3/S4 (EMSGSIZE) — дополнение 2026-06-10
При лайв-тесте AWG2-узла рукопожатие проходило, но трафик не шёл: ядро спамило `failed to send data packets: … sendmsg: message too long` (**EMSGSIZE**). Причина — прямое следствие `s3`/`s4`: junk дописывается к **каждому transport-сообщению**, и обфусцированный data-пакет перерастает path MTU физического интерфейса (1500, DF). Handshake маленький — проходит; transport — нет. Plain WG к тому же серверу с `mtu 1420` работает (S-junk нет).
Бюджет: `mtu ≤ 1500 − 28 (UDP/IP) − 32 (WireGuard) − max(S3, S4)`. Для `S3=S4=60` → `mtu ≤ 1380`; рекомендуемый клиентский MTU AmneziaWG — **1280** (запас на PPPoE/вложенные туннели). Эмпирика (тот же узел/сервер, менялся только `mtu`):
- **auto-default**: при незаданном `mtu` на AWG-эндпоинте ставим рекомендованный **1280** вместо upstream-дефолта `1408` (который сам бы превышал бюджет и триггерил наш же warn).
- **warn**: при явно заданном `mtu` выше бюджета — предупреждение (handshake пройдёт, данные — нет). Path MTU зашит консервативно **1492** (PPPoE): `mtu ≤ 1492 − 28 − 32 − max(s3,s4)` → для `s3=s4=60` это `1372`. Эмпирический потолок выше (1380), т.к. тест шёл по реальному 1500-Ethernet; 1492 — запас под узкие пути.
- Проверено (`check`): AWG `s3=s4=60` без `mtu` → тихо (default 1280); `mtu=1420` → `WARN … consider mtu <= 1372`; plain WG без `mtu` → тихо (1408).
Подтверждение (amneziawg-go docs): рекомендуемый клиентский MTU 1280; если `Jmax` ≥ системного MTU — junk-пакет фрагментируется и теряется на узких путях. Это не баг ядра, а размерный оверхед S-junk. Источник находки — заметка агента лаунчера (`singbox-launcher`, 2026-06-10).
## Безопасность
Секреты сервера **никогда** не попадали в репозитории — лайв-конфиг держался только в `/tmp` и затёрт (`shred`). Репо `lx-test/config/awg2_basic.json` — с фейк-ключами.
## Зона касания upstream (ребейз)
sing-box-lx: `go.mod` (replace), `option/wireguard*`, `protocol/wireguard/endpoint.go`, `transport/wireguard/*` — всё `// lx`. Форк wireguard-go ребейзится отдельно на новый тег sagernet (повтор 3-way merge амнезии).
## Остаточное / дальше
- reserved-feature не применяет reserved-байты в obfuscated send (для plain-AWG не нужно — карта пуста).
- Можно перевести `replace` с submodule на pinned-pseudoversion (submodule достаточно).
- Лаунчер: AWG-поля (S1–S4, I1–I5) в визард + парсер `.conf`/awg-quick; рассматривает кламп MTU для AWG-узлов (первичная истина про оверхед `s3`/`s4` — здесь, см. раздел MTU).
AmneziaWG = WireGuard-девайс с расширенным конфигом. В sing-box девайс создаётся в `transport/wireguard` поверх `github.com/sagernet/wireguard-go`. Стратегия: **подменить модуль на `amneziawg-go`** (API-совместим с wireguard-go) и **под `with_awg`** прокидывать AWG-поля в строку конфигурации девайса; endpoint остаётся типом `wireguard`.
> Проверить: совпадает ли публичный API amneziawg-go (пакеты `device`, `conn`, `tun`) с тем, что импортирует `transport/wireguard`. Если расходится — минимальные `patches/` или адаптерный слой в новом файле.
| `lx-test/config/awg2_*.json` | **new** | Конфиги для `sing-box check` |
## 4. Зона касания upstream (для ребейза)
`go.mod`/`go.sum`, файл опций wireguard-endpoint, `protocol/wireguard/endpoint.go`, `transport/wireguard/*` (минимально). Девайс-логика и опции AWG — в **новых** файлах под тегом → основной конфликт только в `go.mod` и одной ветке endpoint.
## 5. Порядок работ
1. Submodule + `go.mod` replace; собрать обычный WG (без `with_awg`) — поведение upstream.
2. Сверить API amneziawg-go vs `transport/wireguard`; при необходимости `patches/`.
4.`device_awg.go` (формат `jc=/h1=/i1=…`) под `with_awg`; stub без тега.
5. Прокидка в endpoint; конфиги; `check`; ручной коннект к AWG2-серверу.
## 6. Риски
- **API-дрейф** amneziawg-go относительно версии wireguard-go, на которую завязан upstream (`v0.0.2-beta.1.0.20260224…`). Возможен лаг — фиксировать совместимый коммит сабмодуля, не «latest».
- **Регистр I1–I5** (uppercase) — silent ignore при ошибке; валидировать.
- Взаимодействие junk/CPS с `persistent_keepalive` и MTU — проверять на реальном сервере.
- Доменный `server` + FakeIP: может потребоваться override резолва (референс hoaxisr) — добавлять только при подтверждённой необходимости.
- **Junk** (`Jc`/`Jmin`/`Jmax`) — `Jc` случайных пакетов размером `rand(Jmin..Jmax)` перед handshake initiation.
- **Магические заголовки** (`H1–H4`) — подменяют 4-байтный тип сообщения (init/response/cookie/transport); в AWG 2.0 — диапазоны `"N-M"`, из которых значение генерируется на лету.
- **Размерный padding** (`S1/S2` — на handshake, `S3/S4` — на **каждый** transport-пакет).
- **CPS-пакеты** (`I1–I5`) — снимки реального протокола (напр. QUIC Initial, STUN), которые уходят вперемешку с handshake, имитируя посторонний трафик. `I1` — центральный (см. [SPEC 009](../009-WIRESOCK_MASQUERADE_PROFILES/SPEC.md) — декларативные masquerade-профили `ip=quic/sip/dns`, которые генерируют `I1`).
Upstream sing-box AWG не принимает ([#4045](https://github.com/SagerNet/sing-box/issues/4045), closed not-planned) — реализовано в форке.
## Архитектура (два слоя)
Обфускация живёт **в вендоренном wireguard-go** (submodule), а sing-box только пробрасывает параметры. Это ключевое разделение: контракт `transport/wireguard` ↔ device остаётся sagernet'овским, обфускация — аддитивна.
- **6 modified**: `device.go` (AWG-state: `junk`, `headers`, `paddings`, `ipackets [5]*obfChain`), `send.go` (junk + CPS + padding в handshake/transport-путях), `receive.go` (детект magic-header на входе), `cookie.go`/`noise-protocol.go`/`uapi.go` (типы сообщений через генератор, парсинг AWG-ключей в IpcSet).
**Ключевой инвариант — `MessageEncapsulatingTransportSize = 0`** ([device/noise-protocol.go](../../submodules/wireguard-go/device/noise-protocol.go)). Upstream держит 8-байтный headroom перед transport-заголовком (для `conn.Bind.Send()`-префикса). Граф его **обнуляет**: AWG-обфускация формирует префикс сама (junk/CPS уходят отдельными буферами через `SendBuffers`, а не через encapsulating-space). При `= 0` upstream-выражения вида `buffer[MessageEncapsulatingTransportSize+MessageTransportHeaderSize:]` схлопываются к графовому виду `buffer[MessageTransportHeaderSize:]` — поэтому большинство upstream-функций компонуются с графом **без ручного weave**. Это несущий инвариант re-graft (§ ниже).
**Что граф НЕ трогает:**`conn/`, `tun/` — чисто sagernet (берутся из upstream verbatim). Обфускация замкнута в `device/`.
### Слой 2 — sing-box (проброс параметров, всё `// lx`)
- **`option/wireguard_awg.go`** — `AmneziaWGOptions`: `Jc/Jmin/Jmax`, `S1–S4`, `H1–H4` (тип `MagicHeader` — строка `"N"` или диапазон `"N-M"`, JSON-совместим с прежним uint32), `I1–I5` (string, регистр сохраняется). Promoted-встроены в `WireGuardEndpointOptions`.
- **`transport/wireguard/device_awg.go`** (`//go:build with_awg`) — `awgIpcLines()` рендерит IpcSet-ключи `jc=/jmin=/jmax=/s1..s4=/h1..h4=/i1..i5=`, дописываемые к WireGuard-конфигу устройства. `device_stub_awg.go` (`//go:build !with_awg`) даёт явную ошибку при заданных AWG-полях.
- **`transport/wireguard/endpoint.go`** — MTU-политика для AWG (см. ниже).
- **`validateJunk`** — отвергает `jmin > jmax` до старта: `amneziawg-go` считает `rand(0..jmax-jmin)+jmin`, и `jmax < jmin` даёт `rand.Int`с аргументом `≤ 0` → **паника ядра**. Гардим только этот crash-кейс.
Регистрация endpoint остаётся `C.TypeWireGuard` (AWG = WG + доп. поля, отдельный тип не вводим).
## MTU-политика (следствие S3/S4)
`S3`/`S4` дописывают junk к **каждому** transport-сообщению → обфусцированный data-пакет перерастает path MTU физического интерфейса (1500, DF) → ядро спамит `sendmsg: message too long` (**EMSGSIZE**), handshake проходит, а трафик — нет.
Логика в [transport/wireguard/endpoint.go](../../transport/wireguard/endpoint.go) (gated `max(s3,s4) > 0`, plain WG нетронут):
- **auto-default**: при незаданном `mtu` на AWG-эндпоинте — рекомендованный **1280** вместо upstream-дефолта 1408.
- **warn**: при явном `mtu` выше бюджета — предупреждение (`pathMTU = 1492`, консервативно под PPPoE). Для `s3=s4=60` → `mtu ≤ 1372`.
Держать `Jmax` ниже системного MTU (иначе junk-пакет фрагментируется и теряется на узких путях). Подробности: `docs-lx/lx-config.md` §2 (MTU).
## Процедура re-graft (при бампе upstream wireguard-go)
Когда upstream `sagernet/wireguard-go` двигает версию, граф переносится на новую базу. **Не merge, а controlled 3-way apply** граф-diff'а:
1.**База**: submodule → новый sagernet-коммит.
2.**Apply graft**: `git diff <старая-база> <старый-graft> | git apply --3way`. По практике 15/16 файлов ложатся чисто; конфликтует обычно только `send.go` (плотный upstream-путь).
3.**Разрешить конфликты вручную**, порядок по риску: `cookie`→`device`→`noise-protocol`→`uapi`→`receive`→**`send.go`** (высший — junk/padding-хуки в hot-path).
4.**Сверить несущие инварианты**: `MessageEncapsulatingTransportSize = 0`; графовый `RoutineEncryption` (заголовок в начале буфера, без финального encapsulating re-slice); AWG-state поля в `device.go`.
5.**Проверки**: сборка `device/conn/tun` на linux/android/windows/**darwin** (darwin особо — там upstream добавляет платформенный batch-send), затем полный `sing-box`с LX_TAGS, `go test ./transport/wireguard/ ./protocol/wireguard/`, **device-verify** живого AWG-туннеля (junk/handshake/трафик).
История конкретных re-graft'ов (какие базы, что менял upstream) — в [HISTORY.md](HISTORY.md).
## Критерии готовности
-`sing-box check -c` принимает wireguard-endpoint c `jc/h1/i1…` под `with_awg`.
- Реальный коннект к AmneziaWG 2.0 (device-verify): `sending handshake initiation` → `received handshake response` → keepalive → трафик через сервер, с непустыми `Jc` и хотя бы одним `I1`.
- Сборка **без**`with_awg`: обычный WG как upstream; AWG-поля → явная ошибка.
- **sing-box-lx**: `go.mod` (replace + pin), `option/wireguard_awg.go` + `// lx`-поля в основной struct, `transport/wireguard/device_awg*.go`, MTU-блок в `transport/wireguard/endpoint.go`, проброс в `protocol/wireguard/endpoint.go` — всё `// lx`.
- **submodule wireguard-go**: ребейзится отдельно (см. процедуру re-graft), не входит в merge-зону основного репо кроме pin в `go.mod`.
- [SPEC 020](../020-MULTI_WG_IDLE_BUFFER_HEAT/SPEC.md) — idle-suspend WG/AWG-устройств (Down/Up); опирается на стабильный device-API той же вендоренной базы.
- on tag `v*-lx.*` → `build` (6 desktop, tar.gz/zip) + `build_android` (2 AAR) → `release`: `SHA256SUMS` + GitHub Release с notes (база `v1.13.13` + фичи + `lx-print-tags` + строка про AAR). Версия из тега, `sing-box version` → `-lx.N`.
- **`v1.13.13-lx.3` опубликован** (Latest): 6 архивов + `libbox-1.13.13-lx.3.aar` + `libbox-legacy-1.13.13-lx.3.aar` + `SHA256SUMS` — всё зелёное. Этот прогон впервые вживую подтвердил тяжёлый путь (cross ×6 с naive/cronet/purego + gomobile AAR + publish).
- **Windows 7 (32-bit)** legacy-таргет: `windows/386` собирается **пропатченным Go** (`.github/setup_go_for_windows7.sh` — реверты удаления Win7 из `MetaCubeX/go`, как в upstream `build.yml`) и **без `with_naive_outbound`** (`cronet-go` не имеет windows/386 — build constraints исключают всё). Артефакт `sing-box-<ver>-windows-386-legacy-windows-7.zip` — под лаунчер-сборку `singbox-launcher-win7-32` (она тоже 386). Остальной `LX_TAGS` (gvisor/quic/xhttp/awg/…) под 386 компилируется — проверено.
- **Никогда не force-push'ит `lx`** — только новая ветка + PR/issue на ревью.
- Демо (`workflow_dispatch tag=v1.13.13`): `Pick target` → `Up to date?` → success, остальное skipped, **0 side-effects** (ни веток, ни PR, ни issue).
## Операционные настройки репозитория (критично для CI)
`gh api repos/OWNER/REPO/actions/permissions/workflow`:
- **`default_workflow_permissions: write`** — иначе `gh release create` падает с`403 Resource not accessible by integration` (это и был корень падений первых релизных прогонов lx.2). NB: «релиз для тега уже существует» — **другая** ошибка (`already exists`), не 403.
- **`can_approve_pull_request_reviews: true`** («Allow GitHub Actions to create and approve pull requests») — иначе авто-PR ребейза ботом блокируется (есть fallback в issue).
- Оба включены 2026-06-09.
## Зона касания upstream (ребейз)
Все lx-артефакты — **новые файлы**: `.github/workflows/lx-{ci,release,rebase}.yml`, `Makefile.lx`, `lx-test/config/`. Единственная правка upstream-файла — `// lx`-блок в `cmd/internal/build_libbox/main.go` (теги AAR). При ребейзе новые файлы переносятся как есть, блок в `build_libbox` — вручную по маркеру.
## Остаточное / дальше
- Старый релиз `v1.13.13-lx.1` можно удалить (предшествует XHTTP-фиксу / полным тегам / libbox; `lx.3` его замещает).
- Лаунчер (репо `singbox-launcher`, отдельно): маппинг `type=xhttp` → реальный xhttp (его задача 023 сейчас в httpupgrade); AWG-поля (Jc/S1–S4/H1–H4/I1–I5) в визард + парсер `awg.conf`; замена бандлового `bin/sing-box` на lx-релиз.
- (опц.) XHTTP `stream-one` framing-баг (`auto`/`packet-up` работают, не блокер).
| `.github/workflows/lx-release.yml` | расширение | on tag `v*-lx.*`: cross-build desktop (через `Makefile.lx`, без дублирования тегов) + job **`build_android`** (AAR) → zip/checksums/GitHub Release |
| `lx-test/config/xhttp_reality.json`, `awg2_basic.json` | из 002/003 | Используются в CI `check` |
## 2. Версия / ldflags
`LX_LDFLAGS = -X github.com/sagernet/sing-box/constant.Version=<upstream>-lx.<N> -checklinkname=0 -s -w -buildid=`. `<N>` — счётчик lx-релизов поверх upstream-тега. **`-checklinkname=0` обязателен** для полного набора тегов (`badlinkname` → `go:linkname` в `crypto/tls` через `common/badtls`; Go 1.24 блокирует без флага). AAR версионируется отдельно — `build_libbox` берёт `git describe`, поэтому в обоих workflow перед сборкой AAR создаётся/обновляется тег.
4.`lx-rebase.yml` (сначала `workflow_dispatch`, потом cron).
5. Демо-прогон ребейз-workflow на текущем теге.
## 5. Риски
- **`-checklinkname=0`** (РЕШЕНО): без него полный набор не линкуется (`badtls`/`crypto/tls`). Локально подтверждено для linux/amd64 и windows/arm64; остальные 4 таргета верифицирует CI-матрица.
- **`with_naive_outbound` через cronet** тянет prebuilt `cronet-go/lib/<os>_<arch>` — если под какой-то таргет prebuilt отсутствует, naive там не соберётся → дропнуть naive на этой платформе (или из набора целиком). Проверяет CI-матрица.
- **AAR-сборка**: требует NDK r28 + OpenJDK 17 + gomobile (`make lib_install`); `build_libbox.checkJavaVersion()` ждёт строго `openjdk 17`. Версия AAR = `git describe`, поэтому тег должен существовать в чекауте.
- Авто-ребейз на **alpha/beta** теги нежелателен — фильтровать только стабильные (`vX.Y.Z` без суффиксов).
-`git submodule` в CI — не забыть `--init --recursive` и pin (нужно и для `with_awg`, и для AAR).
Собрать воспроизводимый конвейер сборки/CI/релизов `sing-box-lx`: кросс-платформенные бинари `sing-box`**и Android `libbox.aar`** с клиентским feature-set (полный upstream минус серверные/AI-теги) + lx-фичами (`with_xhttp`/`with_awg`), версия `-lx.N`, и **авто-ребейз на новый upstream-тег**.
> **Процедура выпуска → [docs-lx/lx-release-runbook.md](../../docs-lx/lx-release-runbook.md).**
> Главное правило: **перед любым тегом проверить дрейф upstream и обычно смержить его себе, и только
> потом резать релиз/пререлиз.** На ветке `lx-1.14` авто-ребейз на стабильный тег (ниже) заменён
> ручным `git merge upstream/testing` — пока upstream на `v1.14.*-alpha`, стабильного тега нет, а
> rc-линия `vX-lx.1-rc.N` сама является форматом поставки.
---
## 1. Проблема / контекст
Реальная стоимость downstream'а — не первичная разработка, а N ребейзов в год и регулярные сборки на 3 платформы. Нужен конвейер, который ловит «фичи поломались об новый upstream» раньше пользователя и выпускает drop-in бинарь для лаунчера.
## 2. Требования
### 2.1 Сборка
- **Desktop-бинарь `sing-box`** (drop-in для лаунчера): цель `make -f Makefile.lx lx-build`, output `sing-box`.
- **Набор `LX_TAGS`** — upstream feature-set (`release/DEFAULT_BUILD_TAGS`) **минус нерелевантные клиенту**: `with_tailscale` (нет tailscale-endpoint'ов), `with_ccm`/`with_ocm` (прокси Claude Code / OpenAI Codex — серверные AI-сервисы), `with_acme` (серверный выпуск TLS-сертов). Итог = `gvisor/quic/dhcp/wireguard/utls/clash_api/naive_outbound + badlinkname/tfogo_checklinkname0`**+ `with_purego`** (CGO-free кросс-сборка `with_naive_outbound` через prebuilt cronet) **+ `with_xhttp,with_awg`**. `Makefile.lx` — единственный источник истины (`make -f Makefile.lx lx-print-tags`).
- **`LX_LDFLAGS` обязан содержать `-checklinkname=0`** — иначе `badlinkname`/`tfogo_checklinkname0` ломают линк (`common/badtls` использует `go:linkname` в `crypto/tls`, который Go 1.24 блокирует). Зеркалит upstream `build_libbox`.
- **Android `libbox.aar`**: `make lib_install && make lib_android` (gomobile, NDK r28 + OpenJDK 17). `with_xhttp`/`with_awg` зашиты в `cmd/internal/build_libbox` (lx:-блок) → попадают в `libbox.aar` (SDK 23) и `libbox-legacy.aar` (SDK 21). Набор тегов AAR = upstream mobile-set **минус `with_tailscale`** (как desktop — самая тяжёлая либа в APK; правка обёрнута `// lx:no-tailscale`) + наши две фичи (NDK/CGO-сборка, `with_purego` не нужен).
- Версия `vX.Y.Z-lx.N` через ldflags (из 001); для AAR — через `git describe` внутри `build_libbox`.
### 2.2 CI-матрица
- **Политика триггеров (стоимость per-commit ↓).** Doc-only коммиты (`**.md`/`docs/**`/`SPECS/**`/LICENSE) **не запускают CI** (`paths-ignore`). На каждый push/PR — **только дешёвые** job'ы `lint` + `build-check`. Тяжёлые `cross` (6 таргетов) и `android` (gomobile AAR) — **только вручную, на `workflow_dispatch`** (`gh workflow run lx-ci.yml --ref lx` или кнопка Actions → Run workflow); на push их нет. Полную кросс-сборку + обе AAR на каждый релиз-тег и так гарантирует `lx-release.yml`. Серия быстрых пушей отменяет устаревшие прогоны (`concurrency: cancel-in-progress`).
- **`lint`** (push/PR): `go vet` по lx-пакетам с полными тегами + `gofmt` только по lx-файлам (`v2rayxhttp|_xhttp|_awg`, не по всему дереву upstream).
- **`build-check`** (push/PR): один нативный build `with_xhttp,with_awg` + `sing-box check` XHTTP/AWG2-конфигов (должны пройти); затем tagless baseline-бинарь → `check minimal.json` (проходит) + negative-check (XHTTP/AWG2-конфиги без тегов отвергаются).
- **`cross`** (dispatch): `{linux, darwin, windows} × {amd64, arm64}`, full `LX_TAGS`, CGO=0 — проверка, что полный набор + `with_purego` кросс-собирается везде.
- **`android`** (dispatch): `make lib_android` (NDK r28 + JDK17 + gomobile) — libbox AAR собирается с lx-фичами.
- Все job'ы — с submodule (`submodules: recursive`) и `fetch-depth: 0` (для `-lx` версии через `git describe`).
3. Успех + сборка/`check` зелёные → пуш ветки `lx-rebase/<tag>` и **PR**; конфликт → **issue**с диффом `// lx:` зон.
- Никогда не пушить силой в `lx` автоматически — только через PR с ревью.
### 2.4 Релизы
- Тег `vX.Y.Z-lx.N` → артефакты: desktop-архивы (`sing-box`× 6 платформ) **+ `libbox-<ver>.aar` и `libbox-legacy-<ver>.aar`**, общий `SHA256SUMS`.
- Release notes: upstream-база + состояние фич (`with_xhttp`/`with_awg`) + полный `LX_TAGS` desktop-бинаря (через `lx-print-tags`) + строка про AAR.
## 3. Критерии приёмки
- CI зелёный: на push/PR — дешёвые `lint` + `build-check`; полная матрица (`cross`×6 с полным `LX_TAGS` + `android` AAR) — вручную на `workflow_dispatch`. Doc-only коммиты CI не триггерят; релиз-тег собирает всё через `lx-release.yml`.
- Артефакты собираются, бинарь называется `sing-box`, `version` → `-lx.N`.
- **libbox AAR собирается в CI (job `android`) и публикуется в Release**; `Libbox.version()` → `-lx.N`; конфиг с AWG2/XHTTP не падает с «support not built».
- Авто-ребейз workflow отрабатывает на `workflow_dispatch` (демо на текущем теге → «уже актуально» или PR).
## 4. Вне скоупа
- Подпись кода/нотаризация; публикация AAR в Maven/jitpack (отдаём только GitHub Release asset).
- Интеграция AAR в приложение-потребитель (LxBox) — задача на стороне приложения.
- Полностью автоматический мёрж ребейза (всегда ревью).
## 5. Ссылки
- [Build from source — sing-box](https://sing-box.sagernet.org/installation/build-from-source/)
- [x]**`workflow_dispatch` — тяжёлое (вручную):** `cross``{linux,darwin,windows}×{amd64,arm64}` на полном `LX_TAGS` + `android` (`make lib_android`); на push не запускаются
- [x] submodule init (`submodules: recursive`) + `fetch-depth: 0` во всех job'ах
- [x] Зелёный прогон `lint`+`build-check` на push (подтверждено); `cross`×6 + `android` AAR — зелёные в релизном прогоне v1.13.13-lx.3 (cronet/naive/purego на всех 6)
## Авто-ребейз
- [x]`lx-rebase.yml`: fetch upstream tags → выбрать новейший **стабильный** (`^v[0-9]+\.[0-9]+\.[0-9]+$`) → rebase в CI
`H1`–`H4` в конфиге wireguard-endpoint теперь принимают **диапазон** AWG 2.0 (`"43613244-384550127"`) наряду с одиночным числом (`1234567890`, обратная совместимость). Spec-строка доезжает до IpcSet, vendored `submodules/wireguard-go` (который уже умел диапазоны) поднимает обфусцированный handshake. Реальный awg2-экспорт (`seliv_for_awg2.conf`) импортируется без правок.
Изменения — **только в lx-собственных файлах** (`option/wireguard_awg.go`, `transport/wireguard/device_awg.go`, оба созданы в 003). Ноль новых касаний upstream, vendored wireguard-go не тронут.
- База `string` ⇒ `AmneziaWGOptions` остаётся comparable, `IsSet()` (`o != AmneziaWGOptions{}`) работает.
-`Spec()` повторно валидирует и отдаёт канон для IpcSet — ловит опции, собранные в коде (libbox/лаунчер) мимо JSON.
- Валидация (uint32, start ≤ end) и в `UnmarshalJSON`, и в `Spec()`; имя поля в ошибке даёт contextjson (`h1: invalid magic header …`) на парсе и `E.Cause(err, "h1")` в `awgIpcLines` на сборке device.
**Эмит (`transport/wireguard/device_awg.go`):**
-`h1..h4` через `writeStr`-путь (spec-строка), unset → не эмитится.
- Plain WG (без AWG-полей) по-прежнему даёт `""` → byte-identical конфиг.
- ✅ baseline (без `with_awg`) — `awg2_ranged.json` отклонён явной ошибкой «awg support not built».
- ✅ Правки только в lx-own файлах; новые `// lx:`-зоны не добавлялись.
## Приёмка (живой сервер)
Лайв-тест против awg2-сервера из `seliv_for_awg2.conf` (ranged H1–H4):
```
peer - sending handshake initiation
peer - received handshake response ← uapi принял ranged-строки, обфусцированный handshake прошёл
```
-`curl --socks5-hostname 127.0.0.1:21080 https://api.ipify.org` → `{"ip":"64.188.69.128"}` (выходной IP = endpoint сервера): трафик идёт сквозь туннель.
- 0 ошибок `message too long` / EMSGSIZE (MTU не задан → дефолт `1280` под s3/s4 из 003 / `f806e24f`).
- IpcError при выставлении h1..h4-строк — **нет**.
Секреты в репозиторий не попадали: лайв-конфиг и лог держались в `/tmp`, затёрты после теста. `lx-test/config/awg2_ranged.json` — с фейк-ключами.
## Зона касания при следующем ребейзе
Без изменений относительно 003: `option/wireguard_awg.go` и `transport/wireguard/device_awg.go` — **lx-собственные** файлы, в upstream их нет, конфликтов на ребейзе не дают. Vendored `submodules/wireguard-go` не тронут.
## Вне скоупа
- Перепин `app/android/libbox.version` в лаунчере — задача на стороне LxBox.
- AWG inbound/server — отдельная будущая задача (как в 003).
Диапазон уже понимает нижний слой (`submodules/wireguard-go`: `device/magic-header.go` + `device/uapi.go` case `"h1".."h4"` → `newMagicHeader("N"/"N-M")`). Задача — донести spec-строку от JSON-конфига до IpcSet. Меняются **только lx-собственные файлы** (`option/wireguard_awg.go`, `transport/wireguard/device_awg.go`) — ноль новых касаний upstream, ребейз-стоимость не растёт.
Тип `option.MagicHeader` — `string` с канонизацией при парсе:
-`MarshalJSON`: одиночное значение → JSON number (type-fidelity со старым `uint32`), диапазон → JSON string.
- База string ⇒ тип comparable, `IsSet()` (`o != AmneziaWGOptions{}`) работает; zero value `""` = unset.
- Повторная валидация в `awgIpcLines` с именем ключа в ошибке (`E.Cause(err, "h1")`) — покрывает и программно собранные опции (libbox/лаунчер мимо JSON).
| `SPECS/README.md` | docs | строка 005 в roadmap |
## 3. Зона касания upstream (для ребейза)
**Ничего нового.** Оба изменяемых Go-файла — lx-собственные (созданы в 003), upstream-файлы не трогаются, `// lx:`-зоны не расширяются. Vendored `submodules/wireguard-go` не меняется.
## 4. Порядок работ
1.`option`: тип `MagicHeader` + замена полей + тесты.
3. Фикстура `awg2_ranged.json`; `go build` (с тегами и без), `go vet`, `go test`, `sing-box check` обеих AWG-фикстур.
4. Docs + roadmap.
5. Лайв/handshake-приёмка (конфиг с реальными ключами — только во временных файлах, как в 003).
6. REPORT, статус C, тег `v1.13.13-lx.6`.
## 5. Риски
- **Совместимость marshal**: код, который сериализует опции обратно в JSON (`sing-box format`, экспорт из лаунчера), должен получить number для одиночных значений — закрыто MarshalJSON-логикой + тестом round-trip.
- **`"h1": 0` / `"0"`**: прежняя семантика «0 = не задано» (omitempty по zero value uint32) сохраняется канонизацией в `""`.
- **Ошибка без имени поля** при парсе JSON — закрыто дублирующей валидацией в `awgIpcLines` с ключом и тестом текста ошибки.
Поддержать **диапазонные magic headers** AmneziaWG 2.0 (`H1`–`H4` вида `N-M`) в конфиге sing-box-lx. Только прослойка option → IpcSet; протокольный слой (vendored `Leadaxe/wireguard-go`) уже умеет диапазоны.
---
## 1. Проблема / контекст
- Лаунчер (LxBox) начал импортировать реальные awg2-экспорты. Живые конфиги содержат H-поля в новом формате AWG 2.0 — **диапазон** вместо числа:
```ini
H1 = 43613244-384550127
H2 = 826869626-2105069164
```
- Vendored merged-форк `submodules/wireguard-go` **уже умеет** диапазоны: `device/magic-header.go` — `newMagicHeader(spec)` принимает `"N"` и `"N-M"` (uint32, start ≤ end); `device/uapi.go` case `"h1".."h4"` парсит value через `newMagicHeader`.
- Но прослойка ядра не пропускает диапазон:
- `option/wireguard_awg.go`: `H1..H4 uint32` — диапазон не выразить, JSON-строка `"h1": "N-M"` не анмаршалится;
- `transport/wireguard/device_awg.go`: `writeUint("h1", o.H1)` — в IpcSet уходит только одиночное число.
## 2. Цель
Конфиг wireguard-endpoint с `"h1": "43613244-384550127"` (и одиночными `"h1": 1234567890` как раньше) парсится, проходит `sing-box check`, и диапазонная spec-строка доезжает до uapi девайса без IpcError.
## 3. Требования
### 3.1 Опции (`option/wireguard_awg.go`)
- `H1..H4` → тип «число-или-диапазон» (`MagicHeader` на базе string):
- `UnmarshalJSON`: принимает **JSON number** (обратная совместимость — существующие конфиги с `"h1": 1234567890` читаются без изменений) и **JSON string** `"N"` / `"N-M"`;
- валидация: обе части — uint32, start ≤ end; мусор → явная ошибка с именем поля;
- `MarshalJSON`: одиночное значение → number (type-fidelity как раньше), диапазон → string;
- тип comparable — `IsSet()` (`o != AmneziaWGOptions{}`) продолжает работать.
- `h1..h4` эмитятся как spec-строка (`writeStr`-путь), unset → не эмитить.
- Гарантия «plain WG даёт byte-identical конфиг» сохраняется.
### 3.3 Документация (`docs-lx/lx-config.md`)
- `h1`–`h4`: `int | "min-max"` + пример с диапазоном; обновить таблицу и маппинг awg.conf.
## 4. Критерии приёмки
- Тесты: unmarshal number / string-число / диапазон / ошибки (start > end, > uint32, мусор); ipc-строки с диапазоном (`\nh1=43613244-384550127`); существующие AWG-фикстуры не ломаются.
- `sing-box check` принимает конфиг с ranged H1–H4.
- Лайв-тест против awg2-сервера с ranged-конфигом (инфраструктура lx-test из 003), либо хотя бы handshake-проверка, что uapi принимает выставленные строки без IpcError.
- Сборка без `with_awg`: поведение upstream, AWG-поля → явная ошибка (как раньше).
- `go vet`, тесты затронутых пакетов — зелёные.
## 5. Вне скоупа
- Vendored `submodules/wireguard-go` — **не трогать**, он уже умеет диапазоны.
- Перепин `app/android/libbox.version` в лаунчере — задача на стороне LxBox.
- Диапазоны для S1–S4/Jc — в формате AWG 2.0 их нет (только H1–H4).
- [x]`transport/wireguard/device_awg.go`: `h1..h4` → spec-строка через writeStr-путь; валидация с именем ключа; unset → не эмитить
- [x]`device_awg_test.go` (`with_awg`): `\nh1=43613244-384550127`; одиночные значения как раньше; plain WG → `""`
## Проверки
- [x]`lx-test/config/awg2_ranged.json` (фейк-ключи) + `sing-box check` обеих AWG-фикстур
- [x] Сборка с тегами и без; `go vet`; `go test` затронутых пакетов
- [x] Существующие AWG-фикстуры (`awg2_basic.json`) не ломаются
## Приёмка
- [x] Лайв/handshake-тест против awg2-сервера с ranged-конфигом (uapi принял строки без IpcError, handshake + трафик прошли); секреты — только в temp-файлах, затёрты
## Документация и закрытие
- [x]`docs-lx/lx-config.md`: `h1`–`h4``int | "min-max"`, no-overlap, пример с диапазоном, маппинг awg.conf
Релизные Linux-бинари переведены на **статическую musl-сборку с сохранением NaïveProxy**, добавлены роутерные арки. Закрывает [issue #1](https://github.com/Leadaxe/sing-box-lx/issues/1): `libdl.so.2: cannot open shared object file` на AsusWRT Merlin + отсутствие `linux-armv7`.
**Go-кода нет** — задача чисто инфраструктурная (CI). Единственная Go-правка в этой ветке — hotfix gofmt-выравнивания в `option/wireguard_awg.go` (хвост 005, см. ниже).
## Диагноз (эмпирически подтверждён)
`with_naive_outbound` тянет `cronet-go`. Прежний релиз собирался в режиме `with_purego` (`CGO_ENABLED=0`): `purego` на Linux содержит `//go:cgo_import_dynamic … "libdl.so.2"` → бинарь **динамический**, требует `libdl.so.2`. На glibc ок, на musl (роутеры) — падает до старта. Кросс-сборкой проверено: `file` → `dynamically linked`, `strings|grep libdl.so.2` → 1; без naive/purego → `statically linked`, 0.
naive **сохраняем** — это upstream-фича (`release/DEFAULT_BUILD_TAGS` содержит `with_naive_outbound`; `protocol/naive/outbound.go` без `lx:`-маркеров). Поэтому не дропаем, а используем третий режим cronet-go — `with_musl` (статический `libcronet.a` + musl-toolchain), как делает upstream `build.yml`.
## Что сделано
**`.github/workflows/lx-release.yml`** — новый job `build_linux_musl` (зеркало upstream musl-pipeline):
- Linux убран из desktop-job `build` (остаются darwin/windows/win7); `release.needs += build_linux_musl`; release-notes обновлены.
**`.github/workflows/lx-ci.yml`** — dispatch-only smoke-job `linux_musl` (те же 4 арки): полный musl-pipeline + build + verify static, **без публикации** — безопасная приёмка. Помечен «keep in sync with build_linux_musl».
**Нейминг** — по upstream-схеме арочных суффиксов (`armv7` = arm+`v`+GOARM; `mipsle-softfloat` = arch+GOMIPS), но **без суффикса `-musl`**: upstream добавляет его, т.к. собирает и glibc, и musl на арку; у нас Linux — единственный (musl) вариант, и `linux-arm64`/`linux-armv7` совпадают с ожиданием скриптов потребителей.
## Приёмка
- ✅ YAML валиден (`python yaml`), `actionlint` чист (rc=0) для обоих workflow.
- ✅ CI smoke (`lx-ci` workflow_dispatch, job `linux_musl`×4, run [27407702652](https://github.com/Leadaxe/sing-box-lx/actions/runs/27407702652)) — все 4 **success**, `file` → `statically linked`, `libdl.so.2=0`:
- ✅ Боевой релиз [v1.13.13-lx.7](https://github.com/Leadaxe/sing-box-lx/releases/tag/v1.13.13-lx.7) опубликован — 4 musl-арки + desktop (darwin/win/win7) + 2 AAR + SHA256SUMS.
- ✅ **Field-verified** репортером issue #1 на AsusWRT Merlin RT-AX (`linux/arm64`): ядро устанавливается, стартует, работает; `sing-box version` → `1.13.13-lx.7`, теги включают `with_naive_outbound,with_musl`, `CGO: enabled`. Подтверждение на реальном устройстве, не только в CI.
> **Нейминг — апдейт от потребителя:** репортер подтвердил, что суффикс `-musl` для его скрипта **некритичен** (берёт архив с суффиксом или без). То есть наш выбор «без `-musl`» валиден без оглядки на чужой скрипт — это просто следствие единственного варианта на арку. `-softfloat` у mipsle при этом **обязателен** (FP-ABI, не линковка): softfloat запускается на любом MIPS-роутере, hardfloat — только на чипах с FPU.
## Побочный hotfix (хвост 005)
`option/wireguard_awg.go`: расширение `H1..H4 uint32 → MagicHeader` сменило самый длинный тип в struct, gofmt перевыровнял json-теги. `go vet` это не ловит, поэтому ушло в lx.6 с красным дешёвым CI (`lint` job). Исправлено `gofmt -w`, format-only. Урок: прогонять `gofmt -l` на lx-owned файлах перед коммитом.
## Зона касания upstream (для ребейза)
`lx-release.yml` / `lx-ci.yml` — **lx-собственные** файлы (в upstream их нет) → конфликтов на ребейзе не дают. `.github/CRONET_GO_VERSION` — upstream-файл, **только читаем**. Паттерн musl-pipeline заимствован из upstream `build.yml` как референс.
## Вне скоупа
- Экзотика (`386`/`riscv64`/`loong64`/`mips64le`) — точечно по запросу; cronet-musl под `mips64le` нет.
- DEB/RPM/Pacman/OpenWrt-пакеты — не публикуем (только `.tar.gz`).
- naive на Win7 (windows/386) — физически невозможен (нет `cronet-go/lib/windows_386`).
- Перепин `libbox.version` в лаунчере — на стороне LxBox.
Один новый job в `lx-release.yml` — `build_linux_musl` — повторяющий upstream `build.yml` musl-секцию, но с нашим `LX_TAGS`. Существующий desktop-путь не ломаем: из старого `build` job убираем строки `linux/*`, остальное (darwin/windows/win7) остаётся на `Makefile.lx`/purego.
```
build (desktop, как сейчас минус linux): darwin×2, windows×2, win7-386
build_linux_musl (NEW): linux musl-static + naive: amd64, arm64, armv7, mipsle
**Go-кода нет.**`Makefile.lx` править не обязательно (теги берём из него же; musl-логика живёт в CI, т.к. требует Chromium-toolchain env, которого в Makefile не выразить переносимо).
## 4. Зона касания upstream (для ребейза)
`lx-release.yml` — **lx-собственный** файл (создан в 004), в upstream его нет → конфликтов на ребейзе не даёт. Паттерн заимствован из upstream `build.yml`, но как референс, не как правка upstream-файла. `.github/CRONET_GO_VERSION` — upstream-файл, мы его **только читаем** (pin уже совпадает с go.mod), не меняем.
## 5. Порядок работ
1. SPEC/PLAN/TASKS (done).
2. Реализовать `build_linux_musl` + почистить linux из `build` + `release.needs` + notes.
3. Локально: валидация YAML (actionlint/python-yaml). Полную musl-сборку локально не проверить — Chromium toolchain только в CI.
5. На зелёном — verify-шаги (static/libdl) в логах; по возможности запуск armv7 под qemu-user.
6. REPORT, статус C, ответ в issue #1. Боевой релиз — тегом `v1.13.13-lx.7`.
## 6. Риски
- **Локально не верифицируемо** — отладка только через CI; закладываем несколько прогонов. Кеш toolchain критичен для скорости.
- **Время/размер**: Chromium toolchain — гигабайты; musl-бинарь крупнее (вшит libcronet, +неск. МБ). Приемлемо для релиза по тегу (не на каждый push).
- **mipsle softfloat**: проверить, что toolchain build-naive поддерживает target `linux/mipsle` + `GOMIPS=softfloat`. Если cronet-musl/mipsle не соберётся — fallback: mipsle через `DEFAULT_BUILD_TAGS_OTHERS` (без naive, `CGO_ENABLED=0`, статика) как делает upstream для арок без cronet. Зафиксировать в REPORT.
- **keyring/sysroot download** может флапать (внешняя Chromium infra) — ретраи.
Публиковать **статические musl-бинари**`sing-box` под роутерные Linux-арки, **сохраняя NaïveProxy-outbound**. Закрывает [issue #1](https://github.com/Leadaxe/sing-box-lx/issues/1): нужен `linux-armv7`, и текущий `linux-arm64` не запускается на AsusWRT Merlin (`libdl.so.2: cannot open shared object file`).
---
## 1. Проблема / контекст
- Лаунчер-аудитория ставит ядро на роутеры (AsusWRT Merlin, OpenWrt, Keenetic) — это **musl**-окружения.
- Текущие релизные `linux-amd64/arm64` собраны в режиме `with_purego` (`CGO_ENABLED=0`). `purego` через `//go:cgo_import_dynamic … "libdl.so.2"` делает бинарь **динамическим** и вешает зависимость от `libdl.so.2`. На glibc-десктопе ок, на musl — загрузчик падает до старта. Проверено эмпирически: `file` → `dynamically linked`, `strings | grep libdl.so.2` → 1 совпадение.
-`linux-armv7` в релизе **отсутствует** вовсе.
- **NaïveProxy-outbound — штатная upstream-фича** (`release/DEFAULT_BUILD_TAGS` содержит `with_naive_outbound`; `protocol/naive/outbound.go` — upstream-код без `lx:`-маркеров). По CONSTITUTION (upstream + ровно 2 фичи, из upstream ничего не выкусываем) её **нельзя** дропать ради статики.
Без `with_awg`/`with_xhttp` поведение не меняется — это чисто сборочная задача (CI), **Go-кода нет**.
## 3. Требования
### 3.1 Механизм — по подобию upstream `build.yml`
- Сборка musl-варианта повторяет upstream: clone `cronet-go` по pin `.github/CRONET_GO_VERSION` (уже совпадает с`go.mod`: `2faf34666c2c`), regenerate Debian keyring, download Chromium **musl** toolchain через `go run ./cmd/build-naive --target=linux/<arch> --libc=musl download-toolchain`, выставить env (`… env >> $GITHUB_ENV`), затем `CGO_ENABLED=1 go build`с тегом `with_musl` — `libcronet.a` линкуется статически.
- **zig не используется** — официальный путь cronet-go (Chromium toolchain) надёжнее и совпадает с upstream.
### 3.2 Теги
- Брать `LX_TAGS` (Makefile.lx, single source of truth), заменить `with_purego` → `with_musl`. `with_naive_outbound`**остаётся**. Остальные фичи (`with_xhttp,with_awg,…`) без изменений.
- **Без суффикса `-musl`**: upstream добавляет его, т.к. собирает и glibc, и musl на арку; у нас на Linux единственный вариант — musl, поэтому суффикс избыточен и сломал бы ожидание скриптов (`linux-arm64`/`linux-armv7`).
- naive присутствует (тег `with_naive_outbound` в сборке; по возможности — функциональная проверка).
- armv7/mipsle запускаются на реальном/эмулированном musl-роутере (`sing-box version` без ошибки загрузчика).
- darwin/windows/win7/android-ассеты не изменились по составу.
- **Верификация — через CI** (`workflow_dispatch`): Chromium musl-toolchain (гигабайты) недоступен локально на macOS, поэтому локальной сборки musl нет — приёмка по прогону workflow.
## 5. Вне скоупа
- Экзотические арки (`386`, `riscv64`, `loong64`, `mips64le`) — добавляются точечно строкой матрицы по запросу; cronet-musl под `mips64le` вообще нет.
- DEB/RPM/Pacman/OpenWrt-пакеты (upstream их делает) — нам не нужны, публикуем `.tar.gz`.
- naive на Win7 — физически невозможен (нет `cronet-go/lib/windows_386`).
Практический how-to по полям маскировки фичи 009. Это сахар над AmneziaWG `i1`:
вместо ручной CPS-строки `i1=<b 0x...>` пишешь домен/протокол/браузер, а движок
сам собирает пакет-приманку нужного протокола и шлёт его как `i1` перед handshake.
> Требуется сборка с `with_awg`. Без тега любой `id`/`ip`/`ib` отвергается:
> `AmneziaWG (awg) support is not included in this build, rebuild with -tags with_awg`.
---
## 1. Три поля
| Поле | Имя | Значения | Обязательно |
|------|-----|----------|-------------|
| `id` | домен | LDH-хост (`www.google.com`, `ozon.ru`, `_dmarc.example.com`) | **обязателен только для `quic`** (SNI); опционален для `dns` (QNAME или псевдо-домен), `sip` (host или псевдо-host) и `stun` (игнорируется) |
# 010 — WG-endpoint без `detour` режет download на Android (GRO split-brain)
| Поле | Значение |
|------|----------|
| Тип | B (bug) — расследование |
| Статус | **C (closed)** — корень подтверждён на железе (probe v2.1: `rxoffload=true`+`dispatch=single`), фикс верифицирован (download 0.44→20.7 Mbps, вровень с контрольной нодой), вмержен в `lx` (submodule `fb8d8d8`). Кандидат №2 не понадобился. **Обновление 1.14:** наш патч БОЛЬШЕ НЕ НУЖЕН — при миграции на v0.0.3 (re-graft submodule) фикс стал upstream-родным: коммит upstream `24ea133 «conn: harmonize GOOS checks between "linux" and "android"»` добавил `\|\| runtime.GOOS == "android"` в gейты приёмного пути `conn/bind_std.go` (строки 206/215/267/323/458). Наш §010-guard при миграции DROPPED (memory `wg-1.14-migration-is-submodule-rebase`). **Следствие:** GRO на Android теперь полностью рабочий (включается в `controlfns_linux.go:104` без android-guard + разбирается upstream-кодом) → большой `MaxSegmentSize=65535` (`device/queueconstants_android.go`) ему нужен как топливо. ⚠️ Откат `MaxSegmentSize→2200` ради экономии памяти задушит GRO-производительность download (то самое, что §010 чинил). Память от multi-WG нагрева лечить числом устройств, НЕ размером буфера. |
| Зона | ядро `sing-box-lx` + submodule `wireguard-go` (`conn/`) |
---
## Симптом
WireGuard-**endpoint** на Android без `detour`: download почти мёртв при живом
upload. Тот же конфиг с `"detour": "direct"` на endpoint'е — download нормальный.
Асимметрия (download убит, upload жив) указывает на дефект **только на приёме (RX)**.
`ClientBind` (detour-путь) не вызывает `controlFns`/`supportsUDPOffload`, читает по
одной датаграмме (`BatchSize()=1`) — offload-машинерии нет вообще, поэтому путь
иммунен. Это и объясняет «`detour: direct` лечит».
**Детерминирована только мёртвая RX-ветка разбора** (split-путь на android никогда
не зовёт `splitCoalescedMessages`). А вот **активация** offload — нет: `rxOffload`
взводится, лишь если ядро android вернуло `UDP_GRO == 1` (см. п.1), и даже при
взведённом `rxOffload` баг **проявляется**, только когда ядро коалесит в моменте —
нужен плотный входящий поток (download). На редком трафике ядро отдаёт по пакету —
склейки нет — работает. Поэтому «баг в коде» = неразбираемый GRO **при условии**, что
GRO вообще включился; первое детерминировано, второе — рантайм-зависимо и требует
repro (или прямого замера `rxOffload`, см. шаг 0).
---
## Опровергнутые гипотезы (по коду — не повторять)
| # | Гипотеза | Почему отвергнута |
|---|----------|-------------------|
| 1 | MTU / фрагментация | Симптом асимметричный (RX-only); при MTU резало бы симметрично. |
| 2 | У no-detour нет network-strategy / умного выбора интерфейса | Endpoint и direct зовут **один** конструктор `dialer.NewWithOptions → NewDefault`; `networkStrategy` гейтится только на `AutoDetectInterface`/`platformInterface`/`!disableDefaultBind` — одинаково для обоих. На Android оба привязаны к интерфейсу через `ProtectFunc == AutoDetectInterfaceFunc` (`route/network.go:340,368`). Разница не в выборе интерфейса. |
| 3 | Флаг `DirectOutbound` влияет на dialer | `DirectOutbound` — write-only поле; в `NewDefault` не передаётся и нигде не читается. |
| 4 | «Голый `ListenPacket` без стратегии» (`client_bind.go:89`) — корень | Эта ветка — multi-peer / без явного endpoint. Single-peer + валидный endpoint даёт `isConnect=true` (`endpoint.go:208`) → `DialContext`, не `ListenPacket`. А на no-detour `ClientBind` вообще не используется. |
---
## Открытый второй кандидат (если фикс GRO не лечит полностью)
**Тихий хэндовер / смена IP без смены интерфейса.** Re-bind сокета на смене сети идёт
через `onPauseUpdated` → `device.Up()` → `BindUpdate()`. Но событие `NetworkWake`
эмитится только из `notifyInterfaceUpdate`, а мобильный монитор
(`experimental/libbox/monitor.go:95-98`) дедуплицирует по **Name+Index, игнорируя
Addresses**. Смена source-IP на том же интерфейсе → событие подавляется → сокет не
переоткрывается. Не путь `StdNetBind`-vs-`ClientBind`, но самостоятельный сетевой
кандидат — держать открытым.
> Замечание против самой GRO-версии, которое надо снять repro: если баг чисто в
> GOOS-логике, он должен бить download и на стабильной сети, не только на сотовой.
> Если на стабильной сети download жив — либо GRO коалесит по-разному на разных
> интерфейсах, либо GRO не единственная причина (тогда вес смещается к кандидату №2).
---
## План проверки (эксперимент, не релиз)
Один дискриминирующий замер, изолирующий именно RX:
0.**(Бесплатно, без сборки ядра) Сначала снять факт `rxOffload`.** Залогировать
фактический возврат `supportsUDPOffload(conn)` → `(txOffload, rxOffload)` на
целевом Android. Это дискриминирует всю GRO-гипотезу до любого патча:
-`rxOffload == false` → GRO на этом ядре не активируется, корень №1 **мёртв** без
repro; вес немедленно уходит на кандидат №2 (тихий хэндовер).
-`rxOffload == true` → GRO взведён, переходим к шагу 1 (изолировать именно RX).
1. Запатчить `conn/controlfns_linux.go`: пропускать `setsockopt(UDP_GRO)` при
`runtime.GOOS == "android"` (TX/GSO **не трогать**, чтобы `txOffload` оставался
`true` и эксперимент проверял только RX). Тогда `rxOffload` читается `false` и
приём идёт обычным путём.
2. Собрать ядро и воспроизвести no-detour WG-endpoint на **любом** Android
(эмулятор/устройство — баг детерминирован в коде, оператор не нужен), снять
download под плотным потоком.
3. Желательно — пакетная проверка: реально ли `recvmsg` отдаёт >MTU датаграммы на
этом сокете (подтверждает, что GRO коалесит).
Исходы:
- download починился при живом upload → корень = GRO-на-android **подтверждён**, и
минимальный фикс найден тем же шагом;
- не починился → переходим к кандидату №2 (тихий хэндовер).
---
## Решение
Пока **нет** — это таска-расследование. Код в ядро/submodule — только после repro.
**Кандидат на фикс (когда подтверждён):** гейтить `UDP_GRO`-setsockopt
(`controlfns_linux.go`) и/или чтение `rxOffload` (`features_linux.go`) за `!android`,
чтобы `StdNetBind` на android не объявлял offload, который не умеет разбирать. Это
откатывает android на «без offload» — поведение, идентичное рабочему detour-пути.
> NB: фикс «добавить RX self-disable в linux-ветку» был бы **no-op** — эта ветка на
> android мёртвый код. Корень = split-brain GOOS, а не отсутствие fallback.
---
## Acceptance (для будущего фикса)
- [ ] No-detour WG-endpoint на Android даёт download, сопоставимый с `detour: direct`.
- [ ] Фикс не ломает offload/производительность на «настоящем» Linux (не-android) —
гейт за `!android`, а не глобальное отключение.
- [ ] Регресс: plain WG **и** AmneziaWG endpoint; single-peer (`isConnect`) и
multi-peer (`ListenPacket`) пути.
- [ ] Юнит/интеграционный тест на coalesced-receive (сейчас отсутствует — поэтому
дефект и проскочил).
---
## Источники (проверено по живому коду)
-`transport/wireguard/endpoint.go:200-215` — развилка StdNetBind vs ClientBind.
**Дата:** 2026-06-21 · **Статус:** Complete (синтетика) · **Лайв:** ⚠️ НЕ ПРОГОНЯЛСЯ — нет доступа к reality+xhttp ноде; приёмка на синтетике по решению владельца · **База:** ветка `lx` (`1.13.13-lx.13`)
> **Honest caveat.** Фикс принят на основании: построчной сверки с исходниками Xray (контракт stream-one = голый путь / пустой sessionId), независимого подтверждения [issue #5635](https://github.com/XTLS/Xray-core/issues/5635), совпадения с портом hiddify, и зелёной синтетики (юнит-тесты URL-layout + reality-детект, `check`, сборки). Это **не** заменяет лайв против реального Xray-сервера. Лайв остаётся открытым TODO — при первом доступе к reality+xhttp ноде прогнать сценарии из раздела «Дальше» и, если что-то не так, переоткрыть задачу.
## Проблема (из жалобы)
`vless + reality + xhttp`, `mode:auto`, `path:/` работает на стороннем (Xray-логика) ядре, **не работает** на нашем. Две связанные первопричины, сверены построчно с исходниками Xray-core `transport/internet/splithttp` (`main`) и подтверждены [issue #5635](https://github.com/XTLS/Xray-core/issues/5635) + референс-портом hiddify:
1.**stream-one слал `sessionId` в пути.** Xray-сервер (`hub.go`) роутит stream-one (двунаправленный) ТОЛЬКО при пустом sessionId. Наш `dialStreamOne` строил `<path>/<sessionId>` → сервер уходил в stream-down ветку → downlink не-VLESS → VLESS `unknown version`. (Зафиксировано как known bug ещё в 002.)
2.**`mode=auto` всегда → packet-up.** Xray: auto + Reality → stream-one. У нас auto лип к packet-up (т.к. stream-one был сломан).
## Что сделано
Изменения **только** в пакете `transport/v2rayxhttp` (новый код — ребейз-зона = ∅, upstream не тронут).
- ⚠️ **Лайв против реального Xray reality+xhttp сервера** — НЕ выполнен (нет доступа к ноде). По решению владельца задача принята на синтетике (статус → **C**), лайв остаётся открытым TODO (см. «Дальше»). При первом доступе к ноде — прогнать и при расхождении переоткрыть.
## Ребейз-зона
**∅.** Все изменения — в новых файлах пакета `v2rayxhttp` и его правках (новый код фичи). Upstream-файлы (`transport/v2ray/transport.go`, `constant`, `option/v2ray_transport.go`) — не тронуты.
## Остаточные риски
- **Матч Reality по имени типа** (`reality_detect.go`) хрупок к переименованию `RealityClientConfig` в sing/upstream. Митигировано юнит-тестом (двойники с теми же именами) — но тест останется зелёным при переименовании реального типа, а лайв сломается. При ребейзе сверять имя типа в `common/tls/reality_client.go`.
- **stream-up** по-прежнему не лайв-тестился (как и в 002).
## Дальше (открытый TODO — лайв)
При первом доступе к реальной Xray reality+xhttp ноде:
1.`make -f Makefile.lx lx-build`; собрать аутбаунд с параметрами рабочей подписки (server/uuid/sni/pbk/sid/path), локальный socks/mixed на `127.0.0.1:2080`.
2.`mode:stream-one` — `curl -x socks5h://127.0.0.1:2080 https://api.ipify.org` должен вернуть IP сервера (handshake+DNS+HTTPS+download).
3.`mode:auto` на той же ноде — идентичный результат (резолв в stream-one).
4. Регрессия: `mode:packet-up` на packet-up-ноде по-прежнему работает.
5. Если ок — оставить как есть (уже C). Если расхождение — переоткрыть задачу.
Слияние в `lx` и релиз `-lx.N` — на усмотрение владельца (фикс на ветке `lx-xhttp-streamone`, в `lx` не влит).
sessionID в stream-one больше нигде не используется (можно убрать его генерацию для этой ветки, но проще оставить — он безвреден, в URL не идёт). Остальное (POST, pipe-body, streamConn late-binding) — без изменений: разведка подтвердила, что late-binding корректен, баг был только в URL.
Сейчас `case modeAuto, modePacketUp:` → `dialPacketUp`.
**Правка:** выделить `modeAuto` в отдельную ветку:
```go
casemodeAuto:
ifc.realityEnabled{
returnc.dialStreamOne(ctx,sessionID)
}
returnc.dialPacketUp(ctx,sessionID)
casemodePacketUp:
returnc.dialPacketUp(ctx,sessionID)
```
`c.realityEnabled` — новое bool-поле на `Client`, проставляется один раз в `NewClient`.
### 3.4 REALITY-детект в `NewClient` — БЕЗ межтеговой зависимости (РАЗВИЛКА)
`*tls.RealityClientConfig` / `*tls.KTLSClientConfig` — под `//go:build with_utls`; `v2rayxhttp` — под `with_xhttp`. Прямой `tlsConfig.(*tls.RealityClientConfig)` введёт жёсткую связь `with_xhttp → with_utls` и сломает сборку `with_xhttp` без `with_utls` (нарушение CONSTITUTION §3.2). Варианты:
- **(A) Матч по имени типа (рекомендуется).** В `NewClient`:
```go
realityEnabled := tlsConfigIsReality(tlsConfig)
// helper: reflect.TypeOf(unwrap(tlsConfig)).String() содержит "RealityClientConfig"
```
Разворачивать KTLS-обёртку: у `*KTLSClientConfig` встроено поле `Config Config` → если имя типа = KTLS, взять inner и проверить снова. Делать через рефлексию по имени поля/типа, **без** импорта with_utls-типов. Плюс: нулевая межтеговая связь, работает и для kTLS. Минус: матч по строке имени типа — хрупковато к переименованию upstream (митигируется тестом).
- **(B) Проброс флага из вызывающего слоя.** Добавить признак reality в `option.V2RayXHTTPOptions` или в сигнатуру конструктора. Минус: правка upstream-сигнатуры `ClientConstructor`/диспетчера — расширяет ребейз-зону, противоречит «новый код в новых файлах». Отклонено.
- **(C) Эвристика по ServerName/NextProtos.** Ненадёжно (reality неотличим от обычного uTLS по этим полям). Отклонено.
**Решение: (A)** — изолированный helper в `client.go` (или соседнем lx-файле пакета), детект по имени типа с разворачиванием KTLS, покрытый юнит-тестом. Если по ходу выяснится, что рефлексия по приватному полю KTLS недоступна — fallback: матчить и `RealityClientConfig`, и `KTLSClientConfig` по суффиксу имени (для не-Linux kTLS не используется, риск низкий).
---
## 4. Порядок работ
1. Ветка `lx/xhttp` от `lx` (по git-дисциплине; сейчас HEAD на `lx-gro-probe-010`).
2. `requestURL` bare-path ветка + `dialStreamOne` голый путь (фикс 3.1 главный — проверяем stream-one лайв сразу).
4. Тест-конфиги, `sing-box check`, лайв-проверка (stream-one + auto на reality-ноде; packet-up регрессия).
5. DoD, IMPLEMENTATION_REPORT, статус (шапка SPEC.md + Roadmap) → C.
---
## 5. Риски
- **Лайв-сервер обязателен.** Синтетического `check` мало — XHTTP под активной разработкой, нужна reality-нода Xray. stream-one лайв-валидируем явно; auto — на той же ноде убеждаемся, что резолвится в stream-one.
- **Матч по имени типа (3.4-A)** хрупок к переименованию `RealityClientConfig` в upstream/sing. Митигировать юнит-тестом, который при ребейзе сразу покраснеет.
- **h2 для stream-one обязателен** — наш h2-only транспорт это обеспечивает; не регрессируем packet-up (он тоже h2 поверх reality).
- Не сломать stream-up/packet-up URL (оставляют sessionId) — правим только пустую-elem ветку `requestURL` и только `dialStreamOne`.
Починить XHTTP-режим **`stream-one`** (сломан с момента 002) и привести **`mode=auto`** к поведению Xray, чтобы конфиги `vless + reality + xhttp + mode:auto` поднимались на нашем ядре «как есть».
Build-tag: `with_xhttp`. Scope: **client-only**.
---
## 1. Проблема / контекст
Жалоба (2026-06-21): простейший `vless + reality + xhttp`, `mode:auto`, `path:/` работает на стороннем ядре (Xray-логика), **не работает на нашем**. Это типовой конфиг из панелей/подписок.
Две связанные первопричины (обе сверены построчно с исходниками Xray-core `transport/internet/splithttp` ветки `main` и подтверждены [issue #5635](https://github.com/XTLS/Xray-core/issues/5635), а также референс-портом hiddify):
### 1.1 `stream-one` шлёт `sessionId` в пути — главный баг
В Xray `stream-one` — это **один POST на голый путь без sessionId**; сервер (`hub.go`) роутит режим по наличию sessionId:
Наш `dialStreamOne` ([transport/v2rayxhttp/conn.go:21](../../transport/v2rayxhttp/conn.go)) строит URL как `c.requestURL(sessionID)` → `<path>/<sessionId>`. Сервер парсит **непустой** sessionId, уходит в stream-down ветку (ждёт парный stream-up POST, которого нет), и в response.Body летят **не VLESS-байты**. VLESS-парсер читает первый байт как версию → `unknown version` (часто `0x58='X'` из HTTP-обвязки). Зафиксировано как «known bug» ещё в [002 IMPLEMENTATION_REPORT](../002-XHTTP_CLIENT_TRANSPORT/IMPLEMENTATION_REPORT.md).
### 1.2 `mode=auto` всегда → packet-up
Xray (`dialer.go`): `auto` → packet-up по умолчанию, **но если REALITY → stream-one** (если ещё и `downloadSettings` → stream-up). Решает только наличие REALITY/downloadSettings, не h2/h3.
Наш `DialContext` ([client.go:155](../../transport/v2rayxhttp/client.go)) намеренно лепит `auto` к packet-up (т.к. stream-one был сломан). После фикса 1.1 `auto` должен резолвиться как Xray: **REALITY → stream-one**, иначе packet-up. Тогда конфиг из жалобы работает без правок пользователя.
---
## 2. Цель
`vless/vmess/trojan` outbound с`transport.type=xhttp`, `mode:stream-one` — поднимает рабочее соединение к Xray XHTTP-серверу (handshake + DNS + HTTPS + загрузка), в т.ч. поверх Reality. `mode:auto` при включённом Reality резолвится в `stream-one` (как Xray). `mode:packet-up`/`stream-up` — без регрессий.
---
## 3. Требования
### 3.1 Фикс `stream-one` (главное)
-`stream-one` шлёт запрос на **голый нормализованный путь** (`<path>`), **без**`sessionId` в URL (и нигде — ни query, ни header). Метод — `POST` (как сейчас), тело — uplink-pipe, downlink — response.Body того же запроса (late-binding уже корректен — `streamConn.created`).
-`requestURL()` при **пустом** наборе элементов обязан вернуть голый `<path>`**без** trailing-slash. Сейчас `requestURL()` с пустым elem даёт `<path>/` (ловушка `strings.Join([], "/")==""` → `c.path + "/"`), что отличается от Xray/hiddify (`<path>`).
-`stream-up` и `packet-up` URL **не трогаем** — они законно используют sessionId (`conn.go:48,52,86,236`).
### 3.2 `mode=auto` как Xray
-`auto` + **Reality включён** → `stream-one`.
-`auto` + Reality выключен → `packet-up` (текущая совместимая ветка; download-settings у нас нет — stream-up в auto не выбираем).
- REALITY-признак определяется **в конструкторе** (`NewClient`), один раз, и сохраняется на `Client` (в `DialContext` tlsConfig вне области видимости).
-`*tls.RealityClientConfig`/`*tls.KTLSClientConfig` объявлены под `//go:build with_utls`, а пакет `v2rayxhttp` — под `with_xhttp`. **Запрещён прямой type-assert** на with_utls-типы из v2rayxhttp: это введёт жёсткую зависимость `with_xhttp → with_utls` и сломает сборку с `with_xhttp` без `with_utls` (нарушение §3.2 CONSTITUTION — фича за своим тегом).
- Детект REALITY делать **без прямой ссылки на with_utls-тип**: по имени конкретного типа (`reflect`/`fmt %T`, сопоставление с суффиксом `RealityClientConfig`) либо иным способом, не вводящим импорт-связь между тегами. Способ фиксируется в PLAN.
### 3.4 Что НЕ трогаем (доказано в разведке)
- **ALPN.** Наш форсинг `["h2"]` → uTLS добавляет `http/1.1` → `["h2","http/1.1"]`. Xray `decideHTTPVersion` для списка длиной ≠1 → h2; reality всегда h2. Форсинг benign, а h2 для bidirectional stream-one **обязателен**. Оставляем как есть.
- **Padding.** `x_padding` в query внутри `Referer` совпадает с Xray (client request direction). Не трогаем.
- **Submodule, server/inbound** — вне scope.
---
## 4. Критерии приёмки
-`sing-box check -c` принимает `vless + reality + xhttp + mode:stream-one` и `mode:auto` (есть/будет тест-конфиг в `lx-test/config`).
- **Лайв:** реальный Xray XHTTP-сервер (reality-нода) — `mode:stream-one` поднимает соединение (handshake + DNS + HTTPS-страница + загрузка); `mode:auto` на той же ноде даёт идентичный результат (резолвится в stream-one). packet-up/stream-up — без регрессий.
- Сборка **с**`with_xhttp`**без**`with_utls` — компилируется (изоляция тега не нарушена).
- Сборка без `with_xhttp` = поведение upstream (xhttp отвергается).
-`go vet` (lx-теги) и `gofmt -l` по затронутым файлам — чисто. `go build ./...` без тегов — ок.
- Ребейз-зона не расширяется: правки только в новых файлах пакета `v2rayxhttp` и (возможно) комментарий в `option/v2ray_xhttp.go`. Upstream-файлы — не трогаем.
Разбор для зомби-соединения (download, ↓0 на tun0):
| Снимок | Где застряло | Трактовка |
|---|---|---|
| `read=0 write=0` | **выше** copy | proxy (reality/vless) не отдал НИ ОДНОГО расшифрованного байта, хотя pcap показал зашифрованные 777B → дефект расшифровки/фрейминга, НЕ в `connectionCopy` |
| `read>0 write=0` | **запись в tun** | байты из upstream прочитаны, но не записаны/не флашатся в gVisor tun-сокет → подтверждает гипотезу SPEC «застряло в `remoteConn→conn`» |
| `read>0 write>0` | copy шёл | смотреть на `err`/`timeout` — копирование двигалось, причина в завершении/лаге, не в самом stuck |
Ключ к развилке pcap: обёртка стоит на **расшифрованном**`remoteConn`, поэтому
`read` — это plaintext. `read=0` при наличии зашифрованного трафика в pcap снимает
с`connectionCopy` подозрение и переводит расследование выше по стеку (proxy-слой).
---
## Механика (почему обёртка перехватывает, а не обходится)
Обёртка `lxTraceConn` (`route/conn_trace_lx.go`) встраивает `net.Conn`, считает байты
в `Read`/`Write`, и **намеренно НЕ реализует**`Upstream()` / `ReaderReplaceable()` /
`WriterReplaceable()` / `SyscallConn()`. Это критично — иначе её обойдут:
висит**. Для висящего зомби это и есть прямой снимок: серия `tick#1 write=0`,
`tick#2 write=0`… показывает застревание в реальном времени, разрывать НЕ нужно.
- `lx-trace download final: read=… write=… err=…` — один раз, при отвисании
(↓0→↓2820) или разрыве (`DELETE /connections/<id>`).
Искать в `/tmp/lx_trace_run.log` строки `lx-trace download` для нужного conn (сверить
с `destinationIP`/`sourcePort` из шага 4).
6. **Сопоставить** найденную строку с pcap-потоком по времени/объёму (тот же conn).
## Чтение результата (развилка)
| Снимок | Вывод | Следующий шаг расследования |
|---|---|---|
| `read=0 write=0` | proxy (reality/vless) не отдал plaintext, хотя pcap видел зашифрованные 777B | копать **выше** `connectionCopy`: расшифровка/фрейминг vless/reality на download |
| `read>0 write=0` | байты прочитаны из upstream, не записаны в tun | подтверждение гипотезы SPEC; копать запись в gVisor tun-сокет (флаш/блокировка) |
| `read>0 write>0` | copy двигался | смотреть `err`/`timeout`; причина в завершении/лаге, не в чистом stuck |
## Период тика
Тик уже встроен. Период задаётся значением `LX_CONN_TRACE`: `1`/`true` → 5s (дефолт);
`3s`/`1s`/`500ms` → этот период. Для висящего зомби 5s достаточно; если нужен более
плотный снимок прогресса — поставить `LX_CONN_TRACE=1s`. Помнить: чем чаще тик, тем
больше строк в core-логе.
## Не забыть
- `LX_CONN_TRACE` — только на время диагностики (теряется zero-copy splice, half-close
деградирует в Close). После прогона — выключить / поставить релизное ядро обратно.
- Метод, который сработал и который НЕ менять: синхронный tun0+wlan0 ОДНОЙ командой
| Статус | **C (closed) — НЕ удалось воспроизвести на lx.14.** Симптом наблюдался на РАЗНЫХ нодах, включая WG → это зонтик над «↓0»-сталлом, не один баг. WG-долю закрыл фикс **§010 GRO** (вошёл в lx.14, доказан). Для не-WG нод (VLESS/reality) §010 не применим, отдельного код-фикса нет — там симптом сейчас не воспроизводится без подтверждённого объяснения. См. раздел «Закрытие 22.06» ниже. |
| Зона | ядро `sing-box-lx` — `route/conn.go` (`connectionCopy`, общий relay); фикс WG-доли — submodule `wireguard-go` (GRO, §010, UDP/WG-only) |
| Связь | WG-доля симптома = [010-WG_ENDPOINT_GRO_SPLIT_BRAIN](../010-WG_ENDPOINT_GRO_SPLIT_BRAIN/SPEC.md) (closed, фикс в lx.14). §010 это UDP/WG-only → НЕ объясняет не-WG (VLESS) случаи; «виснет и на VLESS» строго не подтверждено (см. «Закрытие») |
| Артефакты | [PROBE.md](PROBE.md) — зонд (env-гейт `LX_CONN_TRACE`, не прогнан в боевом режиме); [instrumentation.patch](instrumentation.patch) — диф зонда; [RUN-PLAN.md](RUN-PLAN.md) — процедура прогона (на случай повторного появления) |
---
## Симптом
Жалоба: WhatsApp/Telegram «висят» — чаты/медиа не грузятся. В клиенте (LxBox
Conns) у приложения одно TCP-соединение, оно **зомби**: `↑517/629 ↓0` — ClientHello
Добавить rule-item **`package_name_regex`** (route / DNS / headless) — матчинг имени Android-пакета по регулярному выражению. Точечный бэкпорт апстрим-фичи 1.14 на стабильную базу 1.13.13 **без** полной миграции на 1.14.
Scope: **все платформы** (фича активна там, где заполняется `ProcessInfo.AndroidPackageNames`, т.е. Android). Build-tag: нет — встроена в ядро роутинга.
> **Обновление (2026-07-02, база 1.14.0-alpha.35, аудит SPEC 022 #19):** после миграции базы на 1.14 сам impl (`route/rule/rule_item_package_name_regex.go` + поле `option.RawDefaultRule.PackageNameRegex`) стал **нативным upstream** — бэкпорт-дельта по коду больше не нужна и растворилась в базе. НО LX-тест `route/rule/rule_item_package_name_regex_test.go` **сохраняется намеренно**: upstream своего теста для этого item не поставляет (проверено на базе и на `upstream/testing`), так что это единственное покрытие фичи. Тест изолирован в своём `_test.go`, тестирует стабильный публичный API (`NewPackageNameRegexItem`/`Match`) и ребейз-конфликтов не несёт.
---
## 1. Проблема / контекст
Запрос (2026-06-23): нужен `package_name_regex` в проекте. У апстрима поле существует **только с sing-box 1.14.0** (commit [`941ce58b`](https://github.com/SagerNet/sing-box/commit/941ce58b) «Add `package_name_regex` route, DNS and headless rule item»), в ветке 1.13.x его нет. На нашей базе уже есть `package_name` (точное совпадение, map-lookup) — но не regex-вариант.
Полная миграция 1.13.13→1.14 оценена отдельным feasibility-разбором как ~1,5–2 дня работы с главным риском в ребейзе AmneziaWG-подмодуля `wireguard-go` (база 506b763 → v0.0.3, ветки diverged 52/51, ручная переинсерция §010 android-GRO fix). Сама же фича `package_name_regex` — изолированный add в `route/rule`, **не трогает** ни один awg/xhttp/selector/build-tag файл и **не гейтится** новым build-тегом. Поэтому выбран точечный бэкпорт, а полная миграция отложена до выхода **v1.14.0 stable** (её штатно подхватит существующий `lx-rebase.yml`, который по дизайну исключает alpha/beta/rc).
Апстрим-коммит `941ce58b` дополнительно содержит хунк про `C.RuleSetVersion5` в `option/rule_set.go` — это часть отдельного rule-set v5 (1.14), **не относится** к фиче и **не переносится** (на нашей базе `RuleSetVersionCurrent = RuleSetVersion4`).
---
## 2. Цель
Правило роутинга / DNS-правило / headless-правило (rule-set) с полем `package_name_regex: ["^com\\.termux.*", ...]` матчит соединение, если хотя бы одно из имён пакетов в `metadata.ProcessInfo.AndroidPackageNames` удовлетворяет хотя бы одному из выражений. Семантика и сообщения об ошибках — идентичны апстрим-1.14.
---
## 3. Требования
### 3.1 Новый rule-item
- Файл `route/rule/rule_item_package_name_regex.go` — дословно апстрим-версия из `941ce58b`: `PackageNameRegexItem`с`[]*regexp.Regexp`, конструктор `NewPackageNameRegexItem([]string) (*PackageNameRegexItem, error)` (компиляция через `regexp.Compile`, ошибка `parse expression <i>`), `Match` по `AndroidPackageNames`, человекочитаемый `String()` (усечение до 3 выражений в описании).
### 3.2 Option-поля
-`PackageNameRegex badoption.Listable[string]`с тегом `json:"package_name_regex,omitempty"` — сразу после `PackageName` в трёх структурах: `option.RawDefaultRule` ([option/rule.go](../../option/rule.go)), `option.RawDefaultDNSRule` ([option/rule_dns.go](../../option/rule_dns.go)), `option.DefaultHeadlessRule` ([option/rule_set.go](../../option/rule_set.go)). Выравнивание struct-тегов — под существующий столбец каждого файла (gofmt-чисто).
### 3.3 Регистрация item в правилах
-В`NewDefaultRule` ([route/rule/rule_default.go](../../route/rule/rule_default.go)), `NewDefaultDNSRule` ([route/rule/rule_dns.go](../../route/rule/rule_dns.go)), `NewDefaultHeadlessRule` ([route/rule/rule_headless.go](../../route/rule/rule_headless.go)) — блок `if len(options.PackageNameRegex) > 0 { ... }` сразу после `PackageName`-блока, с проброской ошибки `E.Cause(err, "package_name_regex")`. Все три конструктора уже возвращают `error` на нашей базе — сигнатуры не меняются.
### 3.4 Cond-функции
-`isProcessRule` / `isProcessDNSRule` ([route/rule_conds.go](../../route/rule_conds.go)) и `isProcessHeadlessRule` ([route/rule/rule_set.go](../../route/rule/rule_set.go)) — добавить `|| len(rule.PackageNameRegex) > 0`, чтобы правило с одним лишь `package_name_regex` корректно классифицировалось как process-rule.
### 3.5 Что НЕ трогаем
- Хунк `RuleSetVersion5` из апстрим-коммита (см. §1) — **не переносим**.
-`go build ./...` без тегов и сборка с lx-тегами (`with_gvisor with_quic with_wireguard with_utls with_clash_api with_xhttp with_awg`) — ок. ✅
-`go vet ./route/... ./option/...` и `gofmt -l` по затронутым файлам — чисто. ✅
- Юнит-тест `route/rule/rule_item_package_name_regex_test.go` зелёный: матч префикса/якоря `$`, матч одного из нескольких пакетов, no-match, nil `ProcessInfo` (без паники), ошибка на невалидном выражении. ✅
- Ребейз-зона: фича — это новый файл + точечные правки в 6 файлах роутинга/опций; коллизий с awg/xhttp/selector/CI-кластерами нет (подтверждено feasibility-разбором).
---
## 5. Вне скоупа
- Полная миграция на 1.14 (отложена до v1.14.0 stable; отдельный feasibility-отчёт).
- [Документация route/rule#package_name_regex](https://sing-box.sagernet.org/configuration/route/rule/#package_name_regex) — «Match android package name using regular expression», since 1.14.0.
- Feasibility-разбор миграции 1.13.13→1.14 (этой сессии) — обоснование точечного бэкпорта вместо полного перехода.
| Тип | F (feature) — смена канала управления ядром (client-side) |
| Статус | A (accepted) — `with_clash_api` drop из AAR в `v1.14.0-lx.1-rc.1`; box.go-фикс в `rc.3`; десктоп-регрессия исправлена в `rc.17` (§3.4) |
**Переезд управления ядром с Clash API на нативный libbox CommandClient — на Android.** LxBox перестаёт использовать Clash REST API и переходит на нативный gRPC-канал `StartedService` (поверх unix-сокета). Из **AAR-сборки** убирается `with_clash_api` — отпадает HTTP-сервер Clash и связанный attack surface. **Десктоп/CLI сохраняют `with_clash_api`** (внешние дашборды ходят по Clash REST API; нативного CommandClient-канала у CLI нет) — см. §3.4.
Этот SPEC фиксирует **сам переезд и его последствия**. Доработки command-протокола, понадобившиеся, чтобы CommandClient заменил Clash API по функциональности (per-node delay, таблица правил, pull-снапшоты групп, фикс потери групп), вынесены в отдельный **[SPEC 015 — COMMAND_PROTOCOL_RPC_EXTENSIONS](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md)**.
Графические клиенты sing-box исторически управляли ядром через два канала: Clash REST API (`experimental.clash_api`, под build-tag `with_clash_api`) и нативный libbox **CommandClient**. Clash API — это слой совместимости со сторонними дашбордами; для **своего** клиента на устройстве разработчики предполагают именно CommandClient (gRPC поверх локального unix-сокета).
LxBox переходит на CommandClient как единственный канал управления, потому что:
- это нативный, более богатый интерфейс (closed-connections история, per-event дельты соединений, ProcessInfo раздельными полями, NQ/STUN/Tailscale-инструменты) — то, что Clash REST не покрывает;
- Clash API — это лишний HTTP-сервер в процессе и открытый локальный порт (attack surface), не нужный, когда клиент ходит по нативному каналу;
- убрав `with_clash_api`, мы уменьшаем дифф и размер AAR.
---
## 2. Цель
LxBox (Android) управляет ядром **только** через CommandClient; `with_clash_api` не входит в **AAR-сборку**. Конфиг, ссылающийся на `experimental.clash_api`, fail-fast с понятной ошибкой (а не молчаливо деградирует). Функциональный паритет с Clash API по нужным UI возможностям достигается доработками CommandClient — см. [SPEC 015](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md).
> **Важно (исправлено):** дроп `with_clash_api` относится **только к Android AAR**. Десктоп/CLI-бинари (mac/windows/linux) управляются внешними дашбордами (yacd/MetaCubeXD) **именно через Clash REST API** — нативного CommandClient-канала вне gomobile/libbox у них нет. Поэтому `with_clash_api` **остаётся** в десктоп `LX_TAGS`. Изначально (rc.1) тег был ошибочно убран и из десктоп-набора тоже — см. §3.4.
- Убрать `with_clash_api` из `sharedTags` ([cmd/internal/build_libbox/main.go](../../cmd/internal/build_libbox/main.go), `// lx:`-блок). **Десктоп `LX_TAGS` (`Makefile.lx`) тег сохраняет** — см. §3.4.
- Без тега (в AAR) подключается `include/clashapi_stub.go` — конфиг с`experimental.clash_api` получает `clash api is not included in this build, rebuild with -tags with_clash_api` (fail-fast, **не** молчаливый отказ).
- lx-конфиги на Android `clash_api` не используют — управление идёт через CommandClient.
- Сделано в `v1.14.0-lx.1-rc.1` (commit `57b5b5e5`) — но изначально ошибочно срезано и с десктопа, исправлено в §3.4.
### 3.2 Доработки CommandClient → SPEC 015
Нативный CommandClient беднее Clash API по ряду возможностей, нужных UI (per-node delay-тест, таблица правил, pull-снапшоты групп/узлов, баг потери одно-узловых групп). Все эти доработки — **в [SPEC 015](../015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md)** (класс §3.6, build-tag `with_lx_command`). `URLTestOutbound` и `GetRules` уже зашиплены (rc.2); `GetGroups`/`GetOutbounds` + фикс `len<2` — target rc.4. Здесь они только упоминаются как часть полного перехода; тех-спека — в 015.
синхронный), `cc_channel.dart:177` (`ccUrlTestOutbound` через MethodChannel).
---
# Ответ ядра — вариант #2 работает БЕЗ правок биндинга
**От:** команда ядра sing-box-lx
**К:** LxBox
**Дата проверки:** против `experimental/libbox/command_client.go` + `common/urltest/urltest.go` (HEAD ветки lx-1.14)
**Итог:** ваш блокер снимается вариантом #2 — отдельный ping-`CommandClient` + его `Disconnect()`. Правка биндинга (#1/#3) НЕ нужна для устранения «зомби». Разбор ниже.
## Прямой ответ на главный вопрос
> «Рвёт ли `CommandClient.disconnect()` уже-ушедшие в dial per-call тесты?»
**Да.** Цепочка проверена по коду:
1.`Disconnect()` ([command_client.go:295](../../experimental/libbox/command_client.go)) делает ДВЕ вещи: `c.cancel()` (отменяет общий `c.ctx`) **и**`c.grpcConn.Close()`.
2. После слоя 1 серверный хэндлер привязал тест к gRPC **per-call**`ctx` (`testCtx := ctx`, [started_service_command_lx.go](../../daemon/started_service_command_lx.go)). Обрыв клиентского вызова/транспорта → gRPC-Go рантайм отменяет серверный stream-ctx этого вызова (стандартный `grpc.NewServer`, никакой обёртки, отвязывающей ctx, нет — [daemon/server.go:16](../../daemon/server.go), [command_server.go:164](../../experimental/libbox/command_server.go)).
3. Этот ctx течёт в `urltest.URLTest(testCtx, …)` → в **оба** ctx-aware этапа: `detour.DialContext(ctx, …)` (TCP/proxy connect+handshake) и `client.Do(req.WithContext(ctx))` (HTTP HEAD) — [urltest.go:99,127](../../common/urltest/urltest.go). Отмена ctx обрывает уже-ушедший dial, не дожидаясь `C.TCPTimeout`.
Итог: dial **НЕ** доживает до timeout независимо от disconnect — он падает по отмене ctx. «Зомби» закрываются.
## Но: рвите ОТДЕЛЬНЫЙ ping-client, не общий
Нюанс, который надо учесть. `Disconnect()` через `c.cancel()` отменяет **общий**`c.ctx` ([command_client.go:32-33](../../experimental/libbox/command_client.go)) — один на ВСЕ вызовы этого инстанса. Если дёрнуть `Disconnect()` на вашем основном client'е, оборвутся и Connections/Groups/Status-стримы. Поэтому:
**Держите ОТДЕЛЬНЫЙ `CommandClient`-инстанс под масс-пинг.** Каждый `NewCommandClient` ([command_client.go](../../experimental/libbox/command_client.go)) поднимает СВОЙ `c.ctx`/`c.cancel` и СВОЙ `grpcConn` (`Connect()`/`ConnectWithFD` — [:239](../../experimental/libbox/command_client.go), [:266](../../experimental/libbox/command_client.go)). Значит:
-`pingClient.disconnect()` отменяет только per-call ctx ping-тестов, рвёт только ping-conn;
- остальные стримы целы.
Это **ровно ваш fallback #2**, и он реализуем на текущем `v1.14.0-lx.1`-биндинге как есть: `urlTestOutbound` уже экспонирован, `disconnect()` уже экспонирован. Новой нативной поверхности не требуется.
## Почему epoch-гейт сам по себе не закрывал зомби (и почему теперь закроет)
Ваш epoch-гейт гасит **применение** результатов на стороне UI, но `urlTestOutbound` — синхронный блокирующий gomobile-вызов, и до слоя 1 серверный тест был привязан к `boxService.ctx` (жил, пока жив сервис) — отменить его было нечем, кроме сноса всего client'а. Теперь, при отдельном ping-client: epoch-бамп (мгновенный UI) **+** `pingClient.disconnect()` (рвёт серверные dial'ы) = и UI, и ядро реагируют. Воркер-пул после disconnect просто получит ошибки на оставшихся вызовах — их и так гасит epoch-гейт.
Практически: при «отмене» зовите `pingClient.disconnect()` и поднимайте свежий `pingClient` под следующий прогон (либо реконнект того же). Стоимость — один short-lived conn на прогон масс-пинга, дёшево.
## Про варианты #1 / #3 (per-call handle / CancelToken)
Не отвергаем, но считаем **избыточными** для вашей задачи:
-#1/#3 дают гранулярность «отменить ОДИН узел из батча, оставив остальные». Для масс-отмены («отменить весь прогон») это не нужно — disconnect ping-client'а гасит весь батч разом.
- Цена #1/#3 — новый stateful слой в gomobile-поверхности (handle-реестр / `CancelToken`-тип, биндинг, версионирование AAR). Это та самая «новая подсистема», которую SPEC 015 §3.6 просит не плодить.
- Если позже появится UX «перепинговать только этот узел с отменой» — вернёмся к #1. Пока YAGNI.
## Что нужно от вас для подтверждения
Проверьте на устройстве сценарий: запустить масс-пинг (concurrency=10) на медленных/недоступных узлах → нажать «отмена» → убедиться, что (а) серверные dial'ы рвутся в пределах ~момента, не висят до `C.TCPTimeout`; (б) Connections/Groups-стримы основного client'а не мигают/не пересоздаются. Если (а) не подтвердится — пришлите лог, копнём транспортный слой конкретного outbound (теоретически отдельный outbound мог бы игнорировать ctx в своём dial — но `urltest.URLTest` зовёт его правильно).
## Сводка ответа
- Вопрос #2 (disconnect рвёт per-call ctx тестов) — **ДА**, проверено по коду. Реализуйте #2 сами.
- Условие: **отдельный** ping-`CommandClient`-инстанс (свой `c.ctx`/`c.cancel`/conn — уже так устроено), чтобы disconnect не задел другие стримы.
- Правки биндинга (#1/#3) — не требуются; держим #1 в запасе под будущий per-node UX.
- Слой 2 разблокирован на текущем биндинге. ✅
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.