Compare commits

...
Author SHA1 Message Date
omarandClaude Opus 5 e38108a7c4 test(gate): install iproute2 in the docker lane — without ip every slot is free
test / go + panel tests (push) Successful in 15m18s
release / test gate (push) Successful in 10m57s
release / apk aarch64_cortex-a53 (push) Successful in 5m52s
release / apk x86_64 (push) Failing after 28s
release / release apk (push) Successful in 6s
This change was already in the working tree when this session started; it is
committed here because it is load-bearing and an uncommitted load-bearing file
is a trap.

netplane.L3SlotFor asks the kernel through `ip link show` and reclaims through
`ip link del`. golang:1.26 ships no iproute2, so in the docker re-exec lane
every slot read as FREE, TestIntegrationL3StaleSlotIsReclaimed stood itself
down rather than pass while proving the opposite of what it claims, and [5/7]
then failed the gate — correctly, since this environment HAS root and
/dev/net/tun and the capability guard is therefore not what skipped it.

Installing it is also what made the concurrent-namespace defect visible at all
(see 06c04c157): with no `ip` on PATH, no `ip link del` was ever issued and the
two test binaries that were destroying shater/generate's TUN looked innocent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 03:25:58 +03:00
omarandClaude Opus 5 06c04c157d fix(gate): a unit test in one package was deleting another package's TUN
`go test` runs package binaries CONCURRENTLY and every one of them shares the
host's network namespace. netplane.L3SlotFor is destructive by design — it
DELETES a candidate slot it finds occupied rather than waiting for it — and
netplane.removeL3Devices deletes both slots unconditionally. Two test binaries
reached those for real:

  shater/engine  l3slot_test.go calls l3RetargetForNext for its return value
  shater/apply   Applier.Teardown -> netplane.TeardownRouting -> removeL3Devices

Measured with an `ip` shim on PATH inside the gate container: apply.test issued
9 `ip link del shater-l3a` + 9 `ip link del shater-l3b` per run, engine.test one
per l3slot test — into the namespace where shater/generate's privileged tests
were holding a live TUN. From the other side that is

  post-start inbound/tun[l3-in]: starting TUN interface: find tun interface: Link not found
  no [shater-l3a shater-l3b] device exists after a successful Start

i.e. an intermittently red [2/7]/[4/7] in a package that did nothing wrong,
while [5/7] — which runs only `^TestIntegration`, so neither binary reaches the
slot code — passed the very same test seconds later. It only became visible when
iproute2 was installed into the gate container: without `ip` every slot read as
free and no deletion was ever issued.

Not a product defect. shaterd is one process with one engine; the running
generation's slot is excluded before anything is deleted, and nothing else on
the router calls L3SlotFor.

The kernel is faked rather than the CHOICE: making the engine's tests stub the
slot answer would delete the only place the ENGINE checks that the running
generation's slot is excluded, which is the invariant the production outage
violated. netplane.L3StubKernelForTest points the two kernel operations at an
in-memory set; engine and apply install it from TestMain (forget-proof, unlike a
per-test helper whose omission fails in a different package on some runs only).
netplane's TestL3StubKernelTakesTheSlotChoiceOffTheKernel is the control, in
both directions: stubbed, nothing reaches the exec seam; restored, the same call
does.

Mutation: with the engine TestMain reverted, the generate binary's
TestIntegrationL3* failed 8 of 8 runs beside a loop of the engine binary; with
it, 0 of 8. With L3StubKernelForTest degraded to a no-op, the control fails
naming the three escaped `ip` calls.

Also: the DoH3 ownership test's control now retries.
requireInstrumentFindsPackedQuery packed a query into a pooled buffer, released
it and demanded the scan find it — but under -race sync.Pool.Put drops one
object in four on purpose, so the control failed 18 of 60 measured runs and took
the whole -race pass down with it. Its sibling control in the same file already
retried for exactly this reason. The claim is existential ("this instrument CAN
find a released buffer"), so one success out of 32 proves it and nothing is
diluted; 0 of 60 after. What it does not buy is stated in the code: the VERDICT
is still a 3-in-4 detector under -race, which is the safe direction, and the
non-race pass runs the same test as a certainty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 03:25:46 +03:00
omarandClaude Opus 5 3654acf7fb fix(egress): tunnel was a device to the router and an unknown type to the engine
An egress type was read by two halves that never call each other. netplane's
EgressDevice accepted `tunnel`, so addEgressRouting gave it a mark, an `ip rule`,
a routing table with an unreachable floor and a prerouting mark bypass, and
`untunnelable_egress` (D26) carried ESP/AH/GRE/IGMP/SCTP out of it by kernel
routing with the engine nowhere in the path. generate's outbound switch had never
heard of `tunnel`: default arm, no outbound, so every node, group and rule bound
to the same egress was fail-closed. One name, two answers.

Refusing `tunnel` would have broken the half that works to match the half that
does not — D26's kernel egress is shipped and verified, and the generator's
refusal is already loud and fail-closed. `tunnel` is not a distinct kind either:
the data plane treats it identically to `interface` in every line that mentions
it, and the panel's own `interface` label already reads "out a specific WAN or
tunnel". So it is an ALIAS, and it is folded to `interface` ONCE, at the config
boundary (Model.NormalizeEgressTypes, called by ParseUCIExport/ReadUCI). Teaching
the generator a second string would have left two strings for the next consumer
to forget; after the fold there is one.

- model: CanonicalEgressType / EgressTypeKnown / KnownEgressTypes — a closed,
  positive registry, plus NormalizeEgressTypes on the load path. An unrecognised
  type is left as written, never defaulted: substituting `direct` for a typo
  would send traffic somewhere nobody asked for.
- model: ValidateEgresses now NAMES an unknown type at validate time. Until now
  the only notice was a generator warning raised while building an engine config,
  which said nothing about the data plane — and the two disagreed anyway.
- netplane: EgressDevice and the prerouting mgmt-bypass consult the registry
  instead of carrying their own copies of the rule. The bypass now keys off
  EgressDevice, so a device-kind egress with no interface no longer gets an
  accept for a mark addEgressRouting never installs.
- panel: the egress editor cleared Interface/Port/DPI for every type it had no
  branch for — including types it renders no field for — so opening an egress it
  labels "(unknown)", changing only the NAME and saving deleted its `interface`.
  On a `tunnel` egress that silently unbound untunnelable_egress and dropped the
  ESP/GRE carrier back to policy. A save may now only clear a field the editor
  was in a position to show.
- panel: the unknown-type hint said "This engine builds no outbound for that
  type", which was false for the one unknown type anybody had — the data plane
  was building it a routing table at that moment. It now names both halves and
  states what saving does.

Tests: TestEgressTypeMeansTheSameInBothHalves runs one table of written types
through the real boundary and then asks netplane AND generate, requiring one
verdict (external test package: generate imports netplane, so nothing inside
netplane can import generate). Mutation-checked both ways — dropping the fold
fails on `tunnel`; restoring the old EgressDevice string test reproduces the
historical split with "generate emitted outbound egress-probe = false ... want
true". Panel: egressEdit.test.ts, mutation-checked by restoring the
unconditional clear (Interface undefined, want 'wg0').

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:52:56 +03:00
omarandClaude Opus 5 3a9b3f523d fix(l3): a test that is not about the TUN must not open one
The l3_tunnel default flip (164b703a7) turned 32 ORDINARY tests in
shater/generate red — the whole CI — because every fixture with a tproxy inbound
now generates the `l3-in` TUN and engine.Apply then wants /dev/net/tun, which the
act_runner LXC guest does not have. Three PRIVILEGED tests failed too, on a host
that DOES have the device.

The proposed fix was to move the TUN inbound out of generate and have the engine
add it at apply time. Refuted, on three grounds:

- it does not fix the 32. Thirty of them fail inside engine.Apply, not box.New;
  the engine adding the inbound leaves them exactly as red, unless the l3_tunnel
  signal travels OUTSIDE option.Options — and then
- the hash gate stops seeing it. Apply's fast path is a hash of the options; a
  decision that is not in them makes toggling l3_tunnel a no-op reconcile, i.e.
  the device stays up with the option off, or never comes up with it on;
- and the `icmp "tunnel"` warning cannot move. It needs the model, and the panel
  reads it out of GenerateWithWarnings. Leaving it in a package that no longer
  makes the decision it explains is a lie generator by construction.

What the failures actually were was contention. Measured under `docker run
--cap-add NET_ADMIN --device /dev/net/tun`: run alone, all three privileged tests
PASS; run as a package, all three FAIL — and one fails by finding a `shater-l3`
device that a DNS-filter test created. There are two L3 slots and they are global
to the process. So the fix is that the engine instrument in this suite does not
open a kernel device it does not own: withoutL3Ingress, one helper, applied at
applyAndClose and at the six other call sites.

Nothing is skipped, and the ingress does not lose coverage — it gains some:

- TestL3TunnelChangesNothingButTheTunInbound (ordinary, portable) proves the
  default config MINUS the l3-in inbound is byte-identical, through the engine's
  own marshaller, to the l3_tunnel=0 config. That is what lets the 32 Starts keep
  speaking for the default config instead of merely for a config near it;
- TestL3TunInboundIsAcceptedByBoxNew (ordinary) puts the registry half of the
  privileged test on a gate that can actually run it: a slim registry that loses
  tun.RegisterInbound now fails on EVERY CI run with `type not found: tun`
  instead of only where /dev/net/tun exists. That regression changes no generated
  byte and costs a LAN-wide outage on the router;
- TestIntegrationL3StaleSlotIsReclaimed (privileged) covers what a RESTART finds:
  an engine with l3Device == "" next to a device it did not open. It must take
  the other slot, leave that one alone, and RECLAIM it on the next apply. The
  occupied slot is held by a second live engine, not planted with `ip tuntap
  add` — a planted device is PERSISTENT and therefore attachable, and the
  planted version of this test passed with netplane.L3SlotFor's reclaim loop
  deleted, i.e. proved nothing.

generate's placeholder device name is now longer than IFNAMSIZ allows. box.New
accepts it (measured), so the emitted config is still one the engine can
validate; Start refuses it and creates NO device. A caller that builds a box from
generate's output without going through engine.Apply therefore fails at once and
visibly, instead of quietly creating `shater-l3` — the one name every generation
wants, and the intermittent TUNSETIFF EBUSY that netplane/l3.go exists to refuse.

The "leaked TUN" in the sentinel's message was not a leak. Instrumented: Close
returns in ~300 µs with ZERO open /dev/net/tun fds (control: 1 fd immediately
before Close), and the device survives 3.8-4.6 s longer purely as the kernel's
deferred unregister_netdevice. On the stand (ImmortalWrt 25.12.1 r37978, kernel
6.12.94 — the router's revision) the same test takes 0.10 s, so the lag is a
nested-netns container artefact. l3GoneTimeout goes 5s -> 20s: a leak is
unbounded, so the longer budget costs one slow failure and gives up no
sensitivity.

Verification. CONTROL, the criterion that matters: without /dev/net/tun
`ok shater/generate` (was 32 failures). With `--device /dev/net/tun --cap-add
NET_ADMIN`: green, privileged tests really ran. On local_openwrt, cross-built
with the shipped tags: the WHOLE package green with every privileged test
executed, no contamination. `go build ./...`, `go vet ./shater/...` clean.

Mutation-verified, each reverted after: shortening the placeholder fails
TestL3PlaceholderCannotBecomeAKernelDevice by name; making withoutL3Ingress a
no-op brings back exactly 32 failures; gating a second config change on
l3_tunnel, and stripping nothing in the comparison, each fail
TestL3TunnelChangesNothingButTheTunInbound; removing tun.RegisterInbound fails
TestL3TunInboundIsAcceptedByBoxNew with the right hint; deleting L3SlotFor's
reclaim loop fails TestIntegrationL3StaleSlotIsReclaimed with the production
error verbatim (`TUNSETIFF: device or resource busy`); l3GoneTimeout at 1ms still
fires the leak sentinel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:49:40 +03:00
omarandClaude Opus 5 201aa7c168 fix(panel): traceroute never printed a hop — stop saying it works
Five untunnelable notes told the operator that a plain `traceroute` works,
"still follows your rules", or that the hops it prints are the tunnel's path.
Measured on the production router: it prints `* * *` and nothing else, under
every rung of the ladder — `direct` included — with the L3 ingress on or off.

There is no mechanism that could print a hop. The UDP probe is diverted by
tproxy and delivered LOCALLY to the engine's socket; local delivery is not
forwarding, so the TTL is never decremented and no router on the path is
provoked into a time-exceeded. The engine opens its own connection with a
fresh TTL, and an ICMP error raised against that has no way back to the
client's datagram. `traceroute -I` and Windows `tracert` are ICMP echo and do
work — that half of the text was true and is kept.

One shared udpTracerouteFacts now carries the symptom, the cause and the way
out, so the panel cannot fork the claim; netplane/untunnelable.go states the
same fact in the same terms.

Second correction in the same notes: the outbounds that carry an echo are not
just WireGuard/AmneziaWG. generate/route.go's l3Target is exhaustive by
adapter registration — a wireguard/AWG node AND the direct outbound behind
`direct` or an interface egress. In the commonest configuration here that is
most of the address space, and those pings answer out of the ordinary uplink
with its real address. The old text let an operator conclude either
"tunnelled" or "dropped"; it was neither.

traceroute_honesty_test.go is the ratchet: an exhaustive matrix over policy x
kill switch x L3 x egress, asserting the retired sentences never return and
that any note mentioning a trace carries the shared facts verbatim — with a
control that fails if the matrix stopped mentioning tracing at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:20:41 +03:00
omarandClaude Opus 5 78bb6a1be8 fix(netplane): a read that fails, a floor nobody checked, a flow that predates the plane
Four defects, all of the same family: something the plane relies on stops being
true and nothing says so.

1. One failed `uci -q export firewall` opened a hole AND switched off the alarm
   for it. nftZoneDevices answered nil on a read failure — the same answer as an
   empty zone — so a rule with `src: zone:lan` produced no divert line, no
   fail-closed drop and no accept_local; and uncoveredNetworkWarnings, whose job
   is to report exactly that, ran the same command, got the same nil and stayed
   silent. The read now carries its error: renderNft refuses under a closed
   kill-switch (same contract as an unusable device name) and warns under an
   open one, and the coverage check names the blindness itself.

2. RoutingPresent did not check the fail-closed floor its Apply twin installs.
   addEgressRouting/addL3Routing install three things per binding; the presence
   checks knew two. A floor that failed to install once was never retried, and
   the table fell through to `main` the first time its device went down. The
   checklist test grows clause (e) so the next mark cannot repeat it.

3. A flow established before the divert plane existed bypassed it for life:
   confirmed by conntrack while nothing diverted it, offloaded to fw4's
   flowtable, steered by netdev-ingress ahead of our prerouting hook and
   refreshed by its own packets. On the divert going from ABSENT to PRESENT —
   not on every apply — the TCP/UDP entries of flows forwarded from the divert
   devices' subnets are dropped, so they re-derive their path. Not a flush: the
   router's own addresses and LAN-to-LAN are excluded, so SSH, LuCI and the panel
   survive. Measured on the stand: 3 client flows cut, the live SSH session and
   the router's own connections untouched; `conntrack` CLI confirmed absent
   there, which is why this is ctnetlink.

4. The untunnelable text claimed Linux/macOS traceroute "still prints hops". It
   prints none, under any policy: the UDP probe is delivered locally by tproxy,
   local delivery does not decrement TTL, and no router raises time-exceeded.
   `traceroute -I` is what works.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 01:01:00 +03:00
omarandClaude Opus 5 ea3a4c518e test(generate): the L3 ingress is the default now — say so in the fixtures, not in 51 rewrites
The l3_tunnel default flip (164b703a7) turned 51 tests in shater/generate red.
Two premises had changed, and each is repaired where it broke rather than at the
assertion:

- ~43 fixtures build an engine-topology model with no inbounds at all and assert
  "this config produces no diagnostics". On the seeded-ON default such a model
  earns an honest `icmp "tunnel"` warning: the L3 ingress is fed only by the
  tproxy divert plane, and a model with no tproxy inbound raises none. The
  warning is TRUE of those fixtures — they are not routers. So they now say they
  run neither router-wide plane (nonDNSGlobals became plainGlobals, and gained
  the same treatment for l3_tunnel that D24 gave dns_intercept), and every
  "no warnings" assertion keeps its original strength instead of being loosened
  to "no warnings except this one".

- 8 assertions counted len(opts.Inbounds). The subject of every one of them is
  how many TPROXY LISTENERS survive a guard, and a total that also counts a
  synthetic inbound answers a different question — one whose right number
  changes whenever an unrelated global flips. They count tproxy listeners now,
  and while there they gained the assertion the count was standing in for: that
  the SURVIVOR of the clash guard is the first-declared listener, and that two
  distinct ports keep the ports their nft diverts aim at.

TestL3TunnelOffEmitsNoTunInbound had lost its meaning rather than its fixture.
It read the default and asserted "off", so after the flip it was pinning
DefaultGlobals, not l3_tunnel. It now sets the opt-out explicitly and says why
the opt-out has to keep working, and TestL3TunnelOnByDefaultEmitsTunInbound
pins the other direction — that a model which never mentions l3_tunnel gets the
ingress — which nothing in this package did.

TestSniffIsNotAnInboundField asserted "exactly 1 inbound" purely so it could
index ins[0]. It checks every emitted listener now and counts what it checked,
so the guarantee that assertion was really providing (the loop ran) survives
without a count that any future synthetic inbound breaks for no reason.

The warning text is rewritten. "l3_tunnel is on but no tproxy inbound is
enabled" accused the reader of a choice they no longer made: since the flip it
is the default, and a message that reads as "you turned this on" sends them
hunting for a switch they never touched. It now says what is not happening, that
the ingress is on by default, and names BOTH exits — a tproxy inbound restores
it, `option l3_tunnel '0'` says the router does not want it — because which one
is right is a fact about their router the generator cannot know.

model/dnsintercept_test.go had the blindness its l3 twin documented: a plain
strings.Contains is satisfied by `#option dns_intercept '1'`, and the parse half
cannot tell either, because a commented option falls back to the seed, which
since D24 is also true. A config shipping the option commented out would have
passed both halves while giving a fresh install no visible option to flip. The
check is line-wise and comment-aware now, and its "config unreadable" branch is
a Fatal instead of a Skip — a guard that skips itself is how one ends up
reporting ok while guarding nothing.

Mutation-verified, each reverted after: seeding L3Tunnel=false fails the
default test by name; removing the l3_tunnel guard fails the opt-out test;
stripping either exit from the warning fails TestL3TunnelWithoutTproxySkipped;
setting a legacy SniffEnabled on the tproxy listener fails the sniff test;
disabling the listen-clash guard fails TestDuplicateTproxyPortSkipped; freezing
the tproxy port at the default fails TestMultiLanDistinctTproxyPortsBothKept;
commenting out the shipped dns_intercept fails the shipped-config test (and the
parse half stayed silent, which is the blindness).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:54:48 +03:00
omarandClaude Opus 5 0f69880150 test(gate): a skipped test is a test that did not run — name it, or fail
Three holes, one shape: work that reads as coverage and is not.

1. shater/apply's TestApplyInstallsHoldWhenEngineFailsToStart — the only
   end-to-end test between "the engine died" and "the LAN forwards to the
   WAN in the clear" — asserted nothing. It broke the engine by pointing a
   rule-set at /nonexistent/nope.srs and stood itself down with t.Skip when
   that failed to break anything; it stopped breaking anything once
   LocalRuleSet.reloadFile began treating an unreadable file as empty.
   Measured in golang:1.26: the skip fired unconditionally and the package
   still printed `ok shater/apply`.

   It now injects the failure at the engineApply seam — the branch under
   test is applyLocked's, and a particular cause that stops causing retires
   the test silently — and COUNTS the seam calls, so applyLocked ceasing to
   go through it fails by name instead of quietly asserting something else.
   Everything else stays real: the model, generate, the kill-switch
   decision, netplane.RenderHoldNft, the latch, Status. New companion
   TestEngineApplyReallyFailsWithoutStarting is the control that the real
   engine.Apply can fail with the engine left stopped, so the simulated
   state is one this fork can be in.

   Mutation-checked both ways: drop the holdLocked call from applyLocked and
   the test fails with "0 holding planes were installed, want 1"; bypass the
   seam and it fails with "the engine-swap seam ran 0 times, want exactly 1".

2. warnings_test.go had two of the same genre. The len(genWarnings)==0
   t.Skip is now a t.Fatal — an unloadable blocklist must always warn, and a
   generate that stops saying so is the W7 regression, not a reason to stand
   down. TestStatusWarningsAlwaysNonNil pins readConfig itself: its
   "zero warnings" assertion was true on a build host only because the
   config read failed SILENTLY, so once that failure started publishing a
   critical warning the same line meant two different things in two
   environments.

3. The gate could not see any of it. It now runs the suites with -v and
   matches every `--- SKIP` against SKIP_DECLARED; an undeclared skip fails
   BY NAME, a declared one prints its reason on every run. check_skips
   proves its own instrument first (no `=== RUN` line => the check was
   reading a blank page), and it also reports on a suite that failed
   elsewhere, so a red tree cannot become a hiding place. -v costs no test
   time (38/25/24 s plain vs 38/24/24 s, warm) — only output, which is
   filtered on a green run.

Also closes the same hole one language over: [6/7] requires every non-Go
test file in the tree to be claimed by a named runner, and [7/7] runs the
ones this gate owns with a verdict by name. openwrt/luci-app-shater/tests/
status-readout.test.js — 24 assertions over the one screen an operator
reaches while the LAN is cut off — was executed by nothing at all, and
[1/7] could not report it because `go list` is its instrument. The non-Go
suites run on the HOST before the docker re-exec, so the local loop really
executes them rather than printing "did not run" every time; where there is
no node at all they are named and the notice replaces the closing banner.

Controls, all run and reverted: a planted t.Skip is caught and named; a
planted failing .test.js is caught and named; an unclaimed test file is
caught and named.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:30:39 +03:00
omarandClaude Opus 5 164b703a7d feat(l3): ping travels the tunnel by default, and every LAN zone can reach it
l3_tunnel was opt-in, and "off" had no honest win left in it. Off, a LAN ping
is decided by `untunnelable` alone and every rung is a drop (block) or a
disclosure (icmp/direct send the echo out of the WAN with the client's real
address). "Ping works" was never the state where ping was tunnelled — it was
the state where ping was leaking. On, an L3-capable outbound carries the echo
and one that is not drops it honestly: adapter.JudgeFlow returns ActionDrop for
an ICMP flow whose outbound is not a tun.Port, so no reply is forged. The price
is a standing TUN + gVisor netstack, ~2 MB RSS, and it is stated where the
option is.

The switch stays. It is a real answer on a 32/64 MB device and when bisecting
whether the L3 ingress is what broke a box — but it is now a WARNED answer:
ValidateGlobals says what the off state does to ping and names the policy that
takes over. Two combinations also changed meaning and are now reported:
untunnelable=icmp is no longer "block plus working ping" (the prerouting L3
mark claims every ICMP packet before the forward chain the echo accept lives
in, and a LAN host's ICMP errors are marked in with them and dropped in the
TUN), and the existing =direct report gains a sibling rather than standing
alone.

The fw4 seeding was the second half of the same problem. The divert set spans
every LAN inbound and every iface:/zone: rule source, but 30_shater-core seeded
a forwarding into shater_l3 for `lan` only — so on a multi-zone router ICMP
from the other zones is marked, routed, accepted by `inet shater`, and dropped
by fw4's zone policy with nothing in any log. Every zone gets a forwarding now,
guarded by a scan of the actual src/dest pairs so a re-run adds nothing. Every
zone including an uplink, because guessing which zones hold clients is wrong
somewhere and a superfluous entry authorises nothing: accept_to_shater_l3 is
`oifname "shater-l3*" accept`, and the only thing that routes a packet into
that device is our own fwmark rule.

scripts/testbed-lao.sh builds the second LAN zone this needs to be visible at
all. It is not installed by the package — that is the whole opt-in mechanism.

Verified on local_openwrt (ImmortalWrt 25.12.1 r37978): three runs of the
seeder leave exactly one forwarding per zone (lan/wan/lao) and no existing
section altered; deleting the lao forwarding removes `jump accept_to_shater_l3`
from chain forward_lao and re-seeding restores it; with the idempotency guard
disabled two runs produce nine forwardings instead of three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:26:00 +03:00
omarandClaude Opus 5 de6fa8ebf4 fix(doh3): Close is not an ownership handoff — stop pooling the query buffer
Review found the hole and it is real. My previous fix gave the pooled buffer to
the transport and released it when the transport closed the body, on the grounds
that "http3.Transport closes the request body on every path, hence the Once".
That sentence is true about how many times the body is closed and says nothing
about when — the failure mode this project keeps writing down.

Verified against the pinned quic-go: on every error path RoundTripOpt
(http3/transport.go:167-173) closes the body the moment doRequest returns, and
doRequest (http3/client.go:338-341) waits only on the request-CANCELLATION
watchdog — close(reqDone); <-done — never on the goroutine writing the body.
Nothing in quic-go joins that goroutine. So Close is not a handoff point, and
the sync.Once stopped a double Release while doing nothing about a read after
one.

One correction to the review's severity, since it changes what we tell people:
on the failure path the bytes do not reach the resolver. Every ReadResponse
error branch (http3/stream.go:325, :336, :343, :363) calls str.CancelWrite
BEFORE RoundTripOpt closes the body, so what the writer reads out of the
recycled buffer is thrown at a cancelled stream. The disclosure primitive is the
success path only; the failure path is a read of somebody else's memory, which
is undefined behaviour and a -race finding, and not shippable either.

Fixed by not sharing at all: Pack() into memory the body owns. The alternative —
a lock around Read and Close — would also be correct and was rejected because it
keeps a released-but-referenced object alive, and that is now twice in one day
that an assumption about quic-go's internal lifetimes has been wrong.

The cost is negative, measured rather than assumed: Pack is 87 ns/op at 64 B and
1 alloc against 108 ns/op at 64 B and 1 alloc for the pooled version, because
buf.NewSize allocates the Buffer struct itself — the same 64 bytes — and then
adds Get/Put on top. The pool was never saving an allocation here.

The failure path cannot be caught on the wire, so the new test pins the cause:
a query tagged with a random needle, an exchange that fails (server never
answers; context already cancelled), then the pool drained on the goroutine
RoundTripOpt ran on, demanding the needle is not there. Mutations run without
-race: restoring pooledRequestBody fails both subtests 5/5, and blunting the
scan trips its control. -race is a separate pass, green at -count=3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:12:15 +03:00
omarandClaude Opus 5 b91fba1295 fix(l3): a covered last fragment must complete; say what the timeout really does
Two review findings on the fragment reassembler.

1. A whole datagram could vanish. entry.total was assigned before addRange
   was asked, so a last fragment (MF=0) whose range was already covered by
   MF=1 fragments answered fragInsertDuplicate and returned nil — while the
   entry was already complete(). Nothing re-examined it, because every later
   fragment is a duplicate too, so it died at its deadline with all its bytes
   present. A duplicate now falls through to the completion check: it
   contributes no bytes (held bytes still win) but it does contribute the
   total length. This is what the documented first-wins policy always
   implied; the code just did not do it.

   The sender needed is non-conforming, so the old behaviour was safe rather
   than exploitable — but it contradicted the comment three screens up, and
   that comment is the next reader's only defence.

   Also closed positively: a last fragment declaring an end BELOW the bytes
   already held now poisons the datagram instead of quietly never completing.

2. The 5 s timeout was not a memory ceiling and the comment said it was.
   sweep ran only when a NEW key was created, so once fragmented traffic
   stopped, up to fragMaxEntries entries stayed resident indefinitely.

   Both halves are fixed, and the honest one is the comment. sweep now runs
   on EVERY fragment — an O(64) scan on a path that is already the rare one —
   which releases residue as soon as any fragment arrives instead of waiting
   for an unrelated new datagram. That still does not cover total silence, so
   fragTimeout now documents the guarantee the code actually keeps: bounded
   by fragMaxEntries/fragMaxTotalBytes at all times, released on the next
   fragment, NOT "freed within 5 s".

   No timer, deliberately: it would need a goroutine with a lifecycle tied to
   something returnDeviceWrapper has no teardown hook for, and a goroutine
   that must be stopped and might not be is a failure this project has
   already paid for — to reclaim at most ~1.1 MiB that only exists after
   fragmented traffic has already happened. What bounds growth is the byte
   and entry ceiling; this timeout's job is correctness, and for that a
   check driven by the arriving fragment is exact.

   The now-unreachable per-key deadline check is removed rather than left as
   dead defence in depth.

16 mutations, all red. M15 (duplicate returns early again) reds only the
buggy case while the control and the poison case stay green, so the test is
shown able to see both an assembled datagram and a lost one. M17 (sweep back
inside the new-key branch) reds the new test while both old timeout subtests
stay green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:08:25 +03:00
omarandClaude Opus 5 df078c3205 fix(model): the write rollback may not swallow its own failure
The rollback added for the "failed import commits the deletion" defect went
through migrate.go's staged(), which drops the revert's error on the floor
(`_ = u.Revert("shater")`). That is defensible where staged() lives — a
migration that cannot revert leaves a half-migrated config, wrong but visible —
and it is not defensible here, because the delta this path stages STARTS WITH A
DELETE OF THE WHOLE PACKAGE. A revert that silently does not take leaves that
delete in /tmp/.uci, the caller is told only "import failed" and believes
nothing happened, and the next `uci commit shater` from any process publishes
an EMPTY /etc/config/shater. The guard reintroduced the exact loss it was
added to prevent.

writeUCIWith now uses its own revertStagedWrite, which reports both failures.
migrate.go's staged() is untouched: changing its signature to suit this caller
would rewrite a contract three migration paths depend on, for a hazard those
paths do not have.

The wrapped error names the CONSEQUENCE and the one command that clears it
("a staged DELETE ... will publish it ... run `uci revert shater` NOW"), not
just the fact — "revert failed" tells an operator nothing about what it costs.
ErrStagedWriteStuck makes it machine-detectable, so a caller can tell "your
change did not happen" from "your change did not happen and this router is one
unrelated `uci commit` away from an empty config".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:06:52 +03:00
omarandClaude Opus 5 d0471b2418 build(shater-core): ship the keep.d entry, or the node inventory dies at the next flash
files/ is not installed wholesale — every path in Package/shater-core/install is
explicit — so the keep.d file added alongside it would never have reached a
router. sysupgrade's "keep settings" walks /lib/upgrade/keep.d/*, and without
this entry /etc/shater/subs does not survive a flash: the restored box has its
rules and its groups and no nodes for them to point at, and the only repair is
`sub update`, which needs the internet the tunnel was going to provide.

/etc/config/shater needs no entry — it is a package conffile and sysupgrade
already keeps it that way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:02:14 +03:00
omarandClaude Opus 5 3314927bef fix(panel): stop the readouts claiming things the daemon never said
Nine places where the panel asserted more than it could know. Each was
checked against the daemon before being changed, and the two that a test
can reach are pinned by tests proven with a mutation.

MULTICAST IPTV WAS AN INSTRUCTION, AND IT WAS WRONG. The `direct` rung
said "Ping, multicast IPTV, and connecting to a VPN ... all work", so
someone who wanted IPTV read it and moved to the most open setting on the
ladder — the one that also lets a client's ESP/GRE past the proxy — and
still had no IPTV. The stream is UDP; every rule the policy emits carries
`meta l4proto != { tcp, udp }`, and the fail-closed forward chain accepts
only the RFC1918/link-local daddr sets, with no 224.0.0.0/4 among them.
The daemon says so itself in the note drawn a few pixels below. IPTV is
now stated once, and it says it does not work.

THE `block` COST LINE WAS UNCONDITIONAL, and three settings contradict
it: an open kill-switch (no drops are emitted at all), Globals.L3Tunnel
(ICMP is marked into the engine's TUN before the forward chain) and
Globals.UntunnelableEgress (ESP/AH/GRE/SCTP are routed out a named
device). The last two were not in the panel's `Globals` type, so the page
could not have been honest about them even in principle; they were added
rather than papered over with a vaguer sentence, and the copy is now
derived from all three.

THE KILL-SWITCH WAS READ WITH `=== 'closed'`. The daemon decides with
!EqualFold(TrimSpace(v), "open") and `Status.kill_switch` is the raw UCI
string, so `'Closed'`, `' closed '` and `''` — all of which BLOCK on the
router — drew OPEN, amber, "Nothing is meant to be blocked", and through
protectionState downgraded a plane-less router from crit to amber. One
normaliser now, `planeState.killSwitchClosed`, used by all five callers
that had their own spelling of it.

AN UNREADABLE CONFIG IS NOT "TURNED OFF". `enabled`, `kill_switch` and
`panel_port` are sourced from the config and are placeholders when it
could not be read (new `config_readable`). That happens on a full
/overlay or an interrupted `uci commit` — exactly when the fail-closed
plane has the LAN cut off on purpose — and the daemon publishes
plane:"hold" with enabled:false. Checking `!enabled` first rendered
"Turned off", amber, no alarm, and pointed at a Settings page backed by
the same unreadable file. The check now comes first, carries the daemon's
"do not turn anything off to fix it", and the kill-switch readout refuses
to name a policy it could not read instead of printing ARMED from "".

Also: the holding plane promises "no client TRAFFIC reaches the WAN", not
"nothing" — DNS to the router still goes to the ISP in the clear, by
design, so the daemon can recover; the stats backend is bbolt, not SQLite,
and reclaims space by rebuilding the file, not by a VACUUM that does not
exist (and skips it when the disk cannot fit the copy); the lock screen
sent people to System → shater when the menu entry is admin/services/shater,
which is the one instruction the product gives to someone who has just
lost access; and the panel port is configured, not confirmed — a failed
listen is only a log line.

RULESET.FORMAT WAS DESTROYED BY RENAMING A LIST. The edit form rebuilt
the object from its own controls and has no control for `Format`, so the
value could only be restored over SSH. It decides how a `file` list is
parsed and stops a `url` .srs being read as text; without it the list
matches nothing, the rule stops firing, and the traffic falls silently
through to the next rule. Carried now for the two sources the generator
consults it for. The same class of loss is made loud elsewhere: the two
other rebuild sites return `Complete<T>`, so adding a field to `Inbound`
or `DNSRule` fails the build in the function that has to decide.

Egress.Target is deleted: it is not in the Go model, so the "which egress
points at this node" branches could never fire, and had anything ever put
a string on it PUT would have rejected the whole write under
DisallowUnknownFields.

One layout fix on the way past: at 390px the policy plate's grid column
was sized by the select's longest option, so the sentence beside it was
clipped mid-word — which is how a line about what leaks loses its second
half.

Verified: npm run build + tsc clean; 57 tests pass; mutation-checked by
restoring the old comparison, the old check order and the old rebuild in
turn, each time watching the matching tests fail with the exact inverted
reading; browser-checked at 390 and 1280 against the mock, which now
reproduces `?ks=Closed` and `?cfg=unreadable` verbatim instead of
normalising them out of existence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:01:08 +03:00
omarandClaude Opus 5 6801146240 feat(core): back up the product state, and let the watchdog see a crash loop
Two things the box could not survive, both silent.

BACKUPS CARRIED NOTHING. No shater package put a single entry in
/lib/upgrade/keep.d, so "keep settings" and LuCI Backup took /etc/config/shater
(a conffile) and nothing else. Everything the product knows besides UCI lives in
/etc/shater: the entire node inventory (subs/*.json, hundreds of nodes on the
live router), the boot-armor arm token, the compiled blocklists. Restored onto a
new router the config looked complete and had no nodes to route to — and the
repair, `sub update`, needs the internet the tunnel was supposed to provide.

keep.d/shater-core keeps subs/, boot.nft, lists/ and alert-state.json, and names
what it refuses and why: stats.db is history bounded only by stats_disk_limit_mb
(0 = unlimited) and the archive is built in RAM; cache.db is sing-box's cache and
a stale one is worse than none; shaterd.log is a log carrying the query history
of the box it came from.

THE WATCHDOG COULD NOT SEE A CRASH LOOP. /etc/init.d/shater respawns every 5s,
forever; shater-cron escalated only after five consecutive ticks where `pidof`
found nothing. A daemon dying seconds into startup is back before the next
60s sample, so the counter reset every time — while the fail-closed plane held
the LAN shut and the panel, served by that daemon, never came up.

The tick's sleep is now spent sampling the daemon's identity (via its pidfile,
not `pidof`, which also matches the CLI verbs this loop runs) every 5s. A tick in
which 3 different daemons lived is churn; two such ticks in a row is the verdict.
A legitimate bounce replaces the daemon once and is announced twice over
(RESTART_FLAG up, ACTIVE_FLAG down), either of which discards the tick.

The action is the one the operator already chose: kill_switch=open stops the
stack, exactly as the dead-daemon path does; kill_switch=closed — and an absent
or unrecognised value, which is the documented default — reports at daemon.crit
and leaves the decision to the person, naming the command that opens the LAN.

Also drops the ruleset loop from shater_run_due. `shaterd ruleset update` has
never existed; it exited 0, so the loop stamped every url rule-set as freshly
updated and fired a reconcile for work that never happened. Now that it exits
non-zero the same loop would emit ~288 syslog lines a day per rule-set instead.
The comment says who does own the refresh, and where the gap that is left is.

Verified: sh -n and busybox `ash -n`; the pure detector driven with synthetic
sample streams under busybox ash (13 cases); shater_sample_pid against a real
/proc with a live process named shaterd as the positive control; and the whole
chain end to end against a real 2s-lifetime crash loop. Each threshold and each
veto is pinned by a mutation that makes the gate fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:58:14 +03:00
omarandClaude Opus 5 2f8c692c39 fix(l3): the TUN is a reclaimable slot — one fixed name made every apply fatal
On the production router every configuration change with l3_tunnel=1 killed the
engine and held the LAN down, three times in a row:

  19:10:33  reconcile failed: start inbound/tun[l3-in]: open tun: TUNSETIFF: device or resource busy
  19:14:02  start instance failed and could not restore previous config; engine stopped
  19:14:38  reconcile failed: TUNSETIFF: device or resource busy

A new generation had to open the device the outgoing one still held. That alone
is a failed apply; what made it an outage is that the recovery path rebuilds the
PREVIOUS config, which named the same device — so the rescue failed for exactly
the reason it was needed. A recovery path must not depend on the resource whose
contention it is recovering from.

The device is now one of two slots, chosen by the ENGINE at box-build time, on a
copy of the options taken AFTER the hash — so the stored config stays canonical
and a no-op reconcile is still a no-op. It cannot be chosen in generate: generate
runs every minute and its output is what Apply hashes, so an alternating name
there would rebuild the engine once a minute forever.

Rotation alone was NOT enough, and that was measured, not reasoned: the two-slot
build survived five applies of five kinds and then failed on 4 of 10 back-to-back
changes with the original outage in full, because a retired generation keeps its
TUN until its budgeted Close finishes. So an occupied non-current slot is now
DELETED rather than waited for — the running generation's slot is excluded first
and never touched, every other slot belongs to a box that is carrying nothing.
No bounded wait: waiting on an asynchronous kernel teardown is the race this
design removes.

The firewall never learns which slot is live — our accepts and the fw4 zone match
`shater-l3*`, verified to validate AND load on ImmortalWrt 25.12.1 / nftables
1.1.6, so the ruleset is byte-identical across a swap. Routers seeded by a
pre-slot build are migrated in place, or fw4 would silently resume dropping the
forward.

A2: turning the feature off left the device, the ip rule and table 8200 behind —
addL3Routing returned early instead of tearing down, and nothing else owns that
device. The disabled branch and TeardownRouting now remove all three.

Two smaller lies found while proving this, both measured: `ip -6 route flush`
does not take a non-unicast route, so the fail-closed floor survived and the next
add answered `File exists` — reported as a CRITICAL "this table has no floor,
traffic can leave over the plain WAN" on every apply, about a floor that was
right there; and teardown left it behind. Fixed both.

Verified on local_openwrt (ImmortalWrt 25.12.1, kernel 6.12.94 — the router's
revision) before and after, with binaries built from the same tree: the pre-fix
binary reproduces the outage and the leftovers; the fixed one survives all five
apply kinds and 12 back-to-back changes and leaves nothing behind. Ten reverted
mutations, each shown failing. See D28.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:55:44 +03:00
omar c788425cad feat(stats): the connection log now says which rule sent it there
The tracker has carried the matched route rule and the outbound chain since
upstream (common/trafficcontrol/tracker.go Rule/Chain); nothing in shater/ ever
read them, so "why did this connection go out that exit" was unanswerable from
the log and cost hours per report.

ConnLogEntry gains RuleKind/Rule/Chain. Rule is the engine rule text, not the
model rule name: nothing survives generation that ties an emitted option.Rule
back to the /etc/config/shater rule it came from, and a guessed name would be
worse than none. RuleKind keeps the two empty cases apart — "default" is a
recorded fact (nothing matched, took route.Final), "" means not recorded at all,
which is what an old persisted row decodes to.

Both fields are interned, so the ring pays 56 B/row of headers instead of a
private copy of text that is identical across every connection one rule matched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
@
2026-07-26 23:47:49 +03:00
omarandClaude Opus 5 0144282f5e fix(apply): a finding that is still true may not erase itself
Three ways this package published calm over a router that was not doing what
its config said. All three are the inverted failure: not an error raised when
things are fine, but silence when they are not.

1. Critical policy-routing findings were erased by the next no-op reconcile.
   applyDataPlaneLocked set routeWarnings only on the full path; applyLocked
   published the set unconditionally, so a minute later the fast path replaced
   it with one that no longer contained the finding. Neither surviving finding
   ("this egress CANNOT REACH ANYTHING outside its own subnet", "table could
   not be given a fail-closed floor") makes RoutingPresent false, so nothing
   brought it back: zero findings, plane full, green, over an egress carrying
   nothing. The comment on the gate claimed the previous set stood; it did not.

   planeOutcome now distinguishes "nothing was found" from "nothing was
   checked" (routeMeasured, written only by measuredRouting), and applyLocked
   carries the last MEASUREMENT forward across the fast path. A re-measurement
   still retires a finding, so this is not a latch.

2. An unreadable configuration was published as enabled=false. The panel tests
   !enabled before plane and renders "Turned off", amber, no alarm, "turn it on
   in Settings" — over a LAN the boot armor had cut off, pointing at a settings
   page backed by the same unreadable file. Status now carries config_readable
   and config_error, plus a critical finding in section "config".

3. The reason the engine failed to start existed nowhere. holdLocked logged it
   and called no publisher, and Warnings carries the last SUCCESSFUL apply — so
   plane="hold" with an empty findings list was a normal state of the product.
   The cause is recorded and published at read time while the engine is down,
   so it self-clears when the engine comes up; the boot-time arm is a warning,
   a real failure is critical.

Each fix is mutation-checked, and the route-warning test carries its control:
it sees a live finding, sees it survive the fast path, and sees a re-measured
clean state retire it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:45:28 +03:00
omarandClaude Opus 5 b642e5d8fe fix(egress): an interface egress with no interface was bound to br-lan
generate/outbound.go resolved the bind device with netplane.IfaceDevice,
whose empty-name fallback is "br-lan" — correct for an INBOUND with no
network, a black hole for an egress. netplane.EgressDevice returns "" for
the same egress on purpose (it calls br-lan "catastrophic here"), so
addEgressRouting installed no `ip rule` and no routing table for that
egress's mark, and the prerouting marking and the forward-chain accept
skipped it too.

The outbound was therefore emitted with SO_BINDTODEVICE=br-lan and a
routing mark nothing routed: every node, group and rule bound to that
egress dialled public addresses out of the LAN bridge. Not a leak — the
bind pins the socket to the LAN — but a total, silent black hole, with the
panel showing a configured, applied egress and no findings at all. The
`if dev == "" { dev = eg.Interface }` line that stood there read as a
guard against exactly this and could never execute: IfaceDevice never
returns "".

- generate now calls netplane.EgressDevice — the data plane's own
  resolution — so a bind can no longer name a device the routing was never
  installed for, and ` eth1 ` binds what the netplane routes. A device-less
  egress emits NO outbound and is reported; every reference to it then
  resolves through egressDetourOrBlock to tagBlock, so the traffic is
  blocked rather than sent out over the plain WAN.
- model.ValidateEgresses reports the same egress on the config channel
  (netplane's own skip is silent), built on model.EgressHasDevice — the
  model-side twin of EgressDevice, which ValidateUntunnelableEgress now
  shares so the two model resolutions cannot drift either.
- TestEgressDeviceResolutionParity runs one table through
  netplane.EgressDevice and model.EgressHasDevice and requires one verdict,
  the same treatment TestUntunnelableEgressResolutionLockstep gave the
  earlier validator/data-plane divergence.

Also: the UntunnelableEgress comment claimed "the panel says which, at
apply time, from whether the device is point-to-point". It does not. The
operator-facing text states both possibilities and declines to claim
either, there is no UI for the option, and isPointToPoint is consulted
only to warn that a gateway-less device can reach nothing. Said so, so the
next implementer does not read a described feature as a built one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:44:35 +03:00
omarandClaude Opus 5 35e4900769 fix(shaterd): three success reports for work that was not done
`shaterd status` fabricated a status when the daemon was unreachable and
exited 0. The stub is the same struct, printed by the same marshaller, so the
only thing that distinguished it was `plane` being "" — a value a live
Applier.Status() cannot emit. luci-app-shater was forced to key its "daemon
down" verdict off exactly that side effect, and filling `plane` in the stub for
any reason would have silently turned "dead" into "fine" on that page.

Both branches now carry an explicit "daemon_answered" boolean, and the offline
branch exits 1. The field is ADDITIVE and spliced in, not re-marshalled: every
existing key keeps its name, value and position (including plane:"" — still
emitted deliberately so dashboard.js keeps working until it moves onto the new
field), and a newer daemon's unknown fields are relayed untouched.

model.writeUCIWith committed the staged package DELETION when the import that
was supposed to refill it failed: /etc/config/shater came out empty, the caller
saw only "WriteUCI: import: ...", the next ReadUCI reported Enabled=false and
the next reconcile tore the plane down. Both error paths now revert through
migrate.go's staged() instead — the same idiom, for the same reason.

`shaterd ruleset update` printed a note and exited 0. shater-cron runs it with
output discarded and, on a zero exit, stamps the ruleset as freshly updated and
sets changed=1, so every source=url ruleset was permanently "just updated" by a
verb that fetched nothing. notImpl now exits 1 (not 2 — a caller must be able to
tell an unimplemented verb from an unknown one).

pidfilePath becomes a var so the daemon-answered / daemon-absent split is
testable without writing to the real /var/run, mirroring ctlPath.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:39:17 +03:00
omarandClaude Opus 5 d3e33294c1 fix(l3): reassemble return-path IP fragments — classifyReturn drops them
sing-tun's forwardReturn.classifyReturn refuses to judge a fragment
(flow_parse.go sets `fragment` for IPv4 MF/offset and for an IPv6
fragment extension header; flow_dispatch.go:703 answers returnPass), so
a fragmented answer coming back through a WireGuard/AmneziaWG endpoint
falls through to the endpoint's own tun stack instead of the l3 return
path, and the LAN client never sees it.

Measured on the live router: `ping -c3 -s 1400` through an AWG tunnel
with MTU 1280 is 100% loss while the WAN capture shows 3 x (1312 + 208)
in both directions — the far host answers, the peer fragments the answer
to fit the tunnel, the fragments die in classifyReturn. `-s 56` is 3/3
and PMTUD with DF works end to end, so only the fragmented return is
broken.

sing-tun is pinned upstream with no `replace`, but the fix does not need
to live there: every decrypted packet passes returnDeviceWrapper.Write
before it is offered to ReturnPackets. Reassemble there and
classifyReturn gets a whole datagram.

Hard ceilings, because this runs on a 128-256 MB router: 64 concurrent
datagrams, 1 MiB of held bytes, 64 disjoint ranges per datagram, 65535
bytes per datagram, 5 s to complete (timer starts at the first fragment
and is never refreshed). Over any ceiling evicts oldest-first.

Overlap policy: a range contained in one already held is a duplicate and
is ignored (first-wins, deterministic) because benign networks do
retransmit; any PARTIAL overlap poisons the datagram until its deadline.
No conforming fragmenter emits one, and every historical hole in this
area comes from a reassembler that tried to resolve the conflict.

The MTU of shater-l3 is untouched (65535 on purpose) and sing-tun is
untouched.

14 mutations run against the tests; each turns at least one test red,
including the two that first survived (a stale-head reuse the sweep was
covering for, and a fast-path copy).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:38:07 +03:00
omarandClaude Opus 5 32aac89139 docs: stop the docs promising a safety net that ships disarmed
Every install recipe walked the reader through `shaterd apply` + `shaterd
confirm` as if commit-confirm were armed. It is not: DefaultGlobals() never
seeds ConfirmTimeout, the shipped config carries confirm_timeout '0', and
ArmRollback returns at once on a non-positive timeout. A reader following the
README believed an apply that cut their SSH would undo itself. It would not.
README/README.en/INSTALL now arm it in the recipe and say what 0 means; the
apply-flow diagram gained the edge it always took on a stock box.

The boot armor was documented nowhere at all (`grep -rli armor --include=*.md`
returned zero) while shipping enabled and blocking LAN->WAN on every boot.
INSTALL 4 now says what it is, why SSH/LuCI stay up on purpose, every condition
under which it refuses to arm, and how to switch it off.

Also removed or corrected, each checked against the code, not inherited:

* MASQUE/CONNECT-IP is advertised in both READMEs and absent from parse,
  generate and model -- registry names it among the types deliberately left
  unregistered. Dropped, with the fork-vs-product distinction spelled out.
  The inverse too: Hysteria2/TUIC/XHTTP were tagged [T1] while shipped under
  with_quic/with_xhttp; ShadowTLS is generate+registry only, no parser.
* `direct (flow-offload on)` -- no offload/flowtable/flow_offloading anywhere
  in openwrt/, shater/ or panel/src. The product does not do this.
* shater-core deps were two releases stale in two places, one of which vouched
  for a config.buildinfo check that never covered kmod-tun. Ruling narrowed to
  what was actually checked.
* PORTING's "Full schema" -- the shipped config points at it -- was missing
  l3_tunnel and untunnelable_egress (UCI is their only path; the panel does not
  show them) and the blocklist/allowlist/device/alert sections, while listing a
  `config preset` that ReadUCI has no branch for.
* ARCHITECTURE had no L3 ingress and no kernel egress at all, though both are
  [MVP] and one creates an fw4 zone in the user's firewall config. New 3a.
* nftset-for-routing in the DNS diagram: that is the v0.1 mechanism, gone in v0.2.
* CONTEXT described a pre-Phase-1 repo and a 24.10.3 testbed. The testbed is
  ImmortalWrt 25.12.1 r37978-cd0a06bfd3fd (read off the box), which is not a
  detail: .apk does not install on 24.10 at all.
* The gate existed and no .md mentioned it. README/README.en/CONTEXT now do.
* release.yml's header still described publishing as either/or after the rolling
  pointer became unconditional. Comment only.
* Shipped /etc/config/shater: schema_version '1' against CurrentSchemaVersion=2;
  a pointer to a dns_filter line that was not in the globals block (added, '0');
  and `option sniff '1'` on the inbound -- an option the model deliberately does
  not have, which the first panel save would have silently washed out.
* lx-changelog pointed at a D25 heading that does not exist.
* ROADMAP 2b and 5 were done and unmarked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:35:06 +03:00
omarandClaude Opus 5 9e6dda22b3 fix(http3,doh3): stop releasing what is still being read
Two suspicions, both put to a test rather than to a reading. Both were real, and
neither was the leak the suspicion named — both are objects released while still
in use.

roundTripHTTP3Race ran both racers on one cancellable context and cancelled it
before returning the WINNER. quic-go and net/http reset a request's stream when
its context dies, so the caller got a response whose body stopped mid-read:
H3_REQUEST_CANCELLED (local) (read 2687 of 65536 bytes). That path is taken
whenever there is no cached HTTP/3 connection and the request is replayable —
the first request to every host, and every one after an idle close. Each racer
now has a context of its own; losers are cancelled where everything used to be,
and the winner's cancel travels with its body.

DoH3's Exchange packed the query into a POOLED buffer and released it the moment
RoundTrip returned. But http3 writes the request body on a goroutine of its own
and returns as soon as the response HEADERS arrive — the body is still being
read. With the window held open the query on the wire diverges from the query we
packed at exactly offset 8192, quic-go's copy-buffer size: everything past that
was the next pool user's memory, sent to the resolver. Not a slowdown — a data
race and a small memory-disclosure primitive. The buffer now goes back when the
transport closes the body, which http3 does on every path, and can do twice.

Both files diverge from upstream again, hours after 0a6689b29 made them
byte-identical on purpose. Upstream carries the second defect in
dns/transport/https.go too; that file is outside this audit and is named in D27
so the next person finds it instead of rediscovering it.

sing-quic moves v0.6.2-0.20260525051024 -> v0.6.4-0.20260709034545. quic.go is
byte-identical across the two, so this neither duplicates nor retires the
packet-conn ownership fix — quic-go still does not own the socket. What it does
carry is the other half of the family we took only half of: clientConn.Close in
tuic/, hysteria/ and hysteria2/ now sets a past write deadline, word for word
the fix v2rayquic already had. We ship tuic and hysteria2. Cost, measured:
+256 KiB exactly on the stripped aarch64 binary and six indirect modules for a
realm port-mapping path nothing we generate can reach.

Tests are mutation-checked: reverting each fix makes them fail, with the text
quoted above. The DoH3 test carries its own control — it first proves the pool
does hand a released buffer back and that poisoning it lands, because a clean
result from an instrument that cannot produce a dirty one proves nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:26:03 +03:00
omarandClaude Opus 5 fde4bed571 fix(luci): stop calling the daemon dead when only the engine is
`running` changed meaning on 2026-07-26 (a8970b8ac): it was a hardcoded
true and is now the ENGINE's liveness (apply.go `Running: engineUp`).
dashboard.js was last touched on 15 July and stayed in the old epoch, so
a dead engine made the page report "Daemon (shaterd): not running" in
red, advise "start the Shater service first" — the service was running —
and DISABLE the button to the panel, which is the one place the config
can be fixed. The holding plane keeps management reachable on purpose
(netplane/nft.go: "The operator can always get in to fix the config");
LuCI was the only thing taking that guarantee away.

Daemon liveness is now derived from the wire, not from `running`. "The
ubus call returned" is not enough either: `shaterd status` EXITS 0 WITH
A FABRICATED STATUS when the daemon is unreachable (cmdStatus offline
stub), and that stub is the apply.Status zero value plus a UCI read — so
it carries enabled/table/kill_switch but leaves `plane` at "", a value
no live daemon emits. A known plane word is the positive proof a daemon
answered; an explicit empty one is proof none did. Everything else —
{} from a failed call, {"error":...} from the plugin (also what a live
but WEDGED daemon produces), a pre-`plane` daemon — is unknown, and
unknown is an unlit lamp, never green. The launcher button is never
disabled again: a mint that fails already reports itself.

"Interception: active" is gone. apply.go says of `active`, verbatim:
"Never render it as 'we are proxying'" — it is the run latch that gates
hotplug and cron, it stays raised while the engine is down and the LAN
is blocked, and this page painted it green next to two more green lamps
in exactly that state. It is now "Service latch", and its lamp reports
only whether the latch agrees with globals.enabled. The row that was
missing is `plane`: full / hold (LAN->WAN BLOCKED) / none. `traffic` is
shown too, because plane=full is not "tunnelled" — a `default -> direct`
router has a full plane and no tunnel at all.

The rpcd plugin's status docstring listed five fields of fourteen and
had done since before half of them existed; it now describes the real
shape and the two fields that are easy to misread.

tests/status-readout.test.js runs the derivation against six recorded
status shapes with no browser and no router. Mutation-checked: reverting
to `st.running` fails 14 assertions including the operator-visible
"not responding - start the Shater service" over a live daemon;
restoring the "Interception: active" row fails 9; putting
openBtn.disabled back fails 1 by name; opening the closed plane list
fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:22:56 +03:00
omarandClaude Opus 5 cb26936ebf fix(wgdedup): merge identical WireGuard copies instead of blocking one
A rule pointing at node:awgout, which was already the first hop of the
default-route chain, took the house off the internet for two minutes.
The pass saw one private key materialised twice, kept the copy that
sorted first alphabetically, and fail-closed everything that routed
through the other one — which happened to be the default route for all
traffic.

The mechanism was right and the framing was wrong. The physical limit is
one DEVICE per key, not one mention per key. Two copies that build the
same device — same key, same peers, same address/MTU/AWG parameters and
the same dialer — are one device written down twice, and there is nothing
for them to fight over. Those are now MERGED: one survives and every
reference to the others is rewritten to it, silently. That makes the
shape the owner wanted expressible: one chain using awgout as an
intermediate hop and another using it as a terminal, both entering over
the same egress, coexisting on one device.

Identity is the marshalled options blob rather than a hand-picked field
list, so a field added to WireGuardEndpointOptions or DialerOptions later
reads as "different" instead of being silently merged.

Only a real incompatibility — different detour, different peers,
different device parameters — is still two devices, and then:

  - the survivor is chosen by WEIGHT, not by tag order: reachability from
    route.Final (the default route) dominates, breadth of use breaks
    ties, tag order only settles a true tie;
  - the warning names the consequence. "Everything that routed through X
    is fail-closed" is equally true of a stray test rule and of the whole
    house's default route, and that is what the operator read it as. It
    now says which of the three it is, measured on the finished config:
    the default route is dead, or it survives via another path, or it
    never touched the lost copy.

A merge must not rename away the subscription fetch detour: that
reference lives in the model and is resolved against the running box, so
this pass cannot rewrite it. Such tags win the survivor slot outright,
which costs nothing since every copy in a class is the same device.

Tests: identical copies coexist on one device; a real incompatibility
keeps the default-route copy even when it sorts last and says so; the
warning does not announce an outage when the default route survives
through a group, and does announce one when it dead-ends behind a
surviving exit; no duplication at all is a no-op. All seven mutations
(merge off, weight off, member-dedup off, pin off, detour-following off,
consequence collapsed, plus a positive control) fail the suite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:16:12 +03:00
omarandClaude Opus 5 8db29b6267 fix(apply): say when shaterd apply armed no safety net
`shaterd apply` exists for one reason: snapshot the last-good, apply, and arm
an automatic rollback so a change that costs you access to the router undoes
itself. It answered `{"changed":false}` and not one word about that.

On the live router (2026-07-26) that was a trap. The operator edited UCI, ran
`uci commit`, the `config.change` reload trigger had already restarted the
daemon, and the fresh daemon applied the new config on startup. By the time
`apply` ran there was nothing left to apply — and the last-good it snapshotted
as the ROLLBACK TARGET was the newly applied config itself. The watcher was
armed onto the very configuration it was meant to protect against: firing it
would have restored exactly what was already loaded. No safety net, no word
said, house offline.

The verb now answers the question it exists to answer, in a closed vocabulary:

  rollback_armed  true ONLY when a window was armed AND its target differs
                  from what is running. An armed watcher pointing at the
                  running config is not a net and is not reported as one.
  reason          applied | already-applied | nothing-to-apply | disabled |
                  commit-confirm-off | config-unreadable | apply-failed
  message         the same thing in the operator's words, never empty.

The two "nothing moved" cases are told apart where they CAN be: an
/etc/config/shater mtime later than this daemon's start, with the running
config already matching it, can only mean a reconcile beat this command to it
(reason=already-applied). Where they cannot — the `uci commit` reload trigger
is stop+start, so it moves the daemon's start past the edit — the text says
so instead of reading as success: no net, harmless if you changed nothing,
unprotected if you did, and shaterd cannot tell which.

Two silent holes surface as a side effect, both previously reported as plain
success: `confirm_timeout=0` (the SHIPPED DEFAULT in
openwrt/shater-core/files/etc/config/shater) makes ArmRollback a no-op, and a
failed post-apply ReadUCI skips the arming entirely.

Arming behaviour is byte-for-byte unchanged — this only makes its absence
visible. A real safeguard for the already-applied case is separate work.

Tests are mutation-verified three ways: reverting classifyApply to the old
{changed,error} fails 11 tests; blinding the mtime discriminator fails exactly
the discriminating one (and falls back to the honest ambiguous text); making
sameConfig always report "different" fails every invariant that forbids
claiming a net over an identical target.

NOT verified on hardware: local_openwrt was held by another agent, so the
control-socket round trip and the real mtime/daemon-start comparison have not
been exercised on a router.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:12:38 +03:00
omarandClaude Opus 5 564033cd10 fix(apply): chain: as a subscription fetch detour resolved to a name nothing answers to
`fetch_detour=chain:<X>` never worked. engine.ViaToTag maps "chain:X" to the
bare tag "X", but the generator materialises a chain as one wrapper per hop —
chain-<X>-h1..chain-<X>-hN — and routes into the LAST one. The lookup missed and
the update failed with "unknown outbound tag".

It failed CLOSED, so the feed was never pulled over the plain WAN by this path.
But the miss had a sharp edge: when a node or group happened to share the
chain's name, the lookup HIT it, and the subscription was fetched through a
completely different outbound with nothing said.

Applier.HTTPClient now resolves chain: before the engine sees it, against the
tags the RUNNING box actually holds (outbounds unioned with endpoints — a WG hop
is an endpoint and Outbounds() does not list those), mirroring the generator:
the highest-indexed chain-<X>-h<i> wrapper is the entry, and a chain that
flattens to one hop IS that hop. Every other via form is passed through
untouched.

The case the generator cannot serve is named rather than papered over: chains
are built lazily, only for a chain some enabled rule/egress/DNS detour targets,
and a fetch detour is not one of those references — so a chain nothing else
points at has no outbounds at all. That, and every other miss, is an explicit
refusal wrapping engine.ErrOutboundUnknown (the panel already maps it to 400).
Never a fall back to direct: that would put the feed and the owner's real
address on the plain WAN, which is the thing fetch_via=proxy is set to avoid.

Tests are mutation-checked. Pre-fix behaviour resolves "work"/"solo" and kills
every chain case; first-hop-instead-of-last, member-copies-count-as-hops,
dropped pass-through, and a silent direct fallback each kill their own test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:09:35 +03:00
omarandClaude Opus 5 add90b5b2f fix(panel): the DNS footnote was a grid item nobody placed
`.dns-filter-note` under the endpoint-resolver readout is a DIRECT child of
`.dns-filter-card`, so it is a grid item. With no explicit span it auto-placed
into column 1 — the toggle's `auto` track — and sized that track to its own
max-content: 237px at 390px, 322px at 1280px. That left the `1fr` copy column
with 0px, so "Network-wide ad & tracker blocking" laid out one word per line
and spilled 2px past the viewport, scrolling the whole page sideways on a
phone. On desktop the same cause parked the 52px toggle in a 322px column,
270px away from the copy it labels.

Measured at 390px: documentElement.scrollWidth 377 vs clientWidth 375. With
`grid-column: 1 / -1` on the footnote: 375/375, and the track list goes from
`237px 0px` to `52px 185px`. Cancelling just that one declaration in the live
DOM puts 377/375 and `237px 0px` straight back, so nothing else contributes.

Verified with playwright over 320/360/375/390/414/430/480/560/640/720/768/
1024/1280/1440: zero horizontal overflow at every width, with every rule
editor open, all three master toggles flipped, every source tab, and every
resolver type. No `overflow-x: hidden` anywhere — the page does not scroll
sideways because nothing overflows, not because the symptom is hidden.
Focus rings and prefers-reduced-motion re-checked and unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 23:06:30 +03:00
omarandClaude Opus 5 a571bd0e1a docs(claude): model is the executor's call, skills are mandatory, standards that earned their place
The old file pinned every subagent to fable — which broke the moment that
quota ran out mid-session — and spent half its length on panel scaffolding
that has been done for weeks. It said nothing about the test gate, the
testbed, or the hardware router, so none of that reached a subagent unless
it was retyped by hand into the brief.

What is new is not advice, it is the list of things whose absence cost a
day each: a test must be mutation-checked or it is decoration; an
instrument with no control proves nothing; a subagent must be told it may
refute the orchestrator, because the best results this project has had
arrived exactly that way; a formally-true sentence that reads as "it works"
is still a lie.

Skills are now a table mapping this project's areas to the skills that
cover them, with the rule that they are invoked BEFORE the work rather
than after something failed to run, and that every brief must name them —
a subagent cannot see this conversation and will not guess they exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 22:48:05 +03:00
omarandClaude Opus 5 1267d20fb8 docs: drop the L3 handoff note — it is merged, and it said to
test / go + panel tests (push) Successful in 8m33s
release / test gate (push) Successful in 8m8s
release / apk aarch64_cortex-a53 (push) Successful in 6m33s
release / apk x86_64 (push) Successful in 3m45s
release / release apk (push) Successful in 8s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 20:00:46 +03:00
omarandClaude Opus 5 35f697ed08 docs(openwrt): say why mtu_fix is inert instead of claiming an MTU we no longer set
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:52:39 +03:00
omarandClaude Opus 5 d0fb6befb1 fix(l3): the l3-in MTU is not a tunnel budget — 1420 was a forgery generator
shater-l3 was created at 1420, the WireGuard payload budget, copied one
layer too far out. It bought nothing: what actually goes into the tunnel
is sized by sing-tun's forwardToPort against Port.PortMTU(), which
already fragments to the outbound MTU without DF and answers a
well-formed `fragmentation needed` quoting it with DF. All 1420 did was
make the KERNEL split every packet above 1392 bytes of payload on its
way into the device -- and a fragment is the one thing sing-tun will not
judge. Dispatch returns on parsed.fragment before calling JudgeFlow, the
fragments reach the gVisor stack, it reassembles them, and the ICMP
forwarder's installFlow demands an unspecified port address that a
WireGuard endpoint never has. So it declined and answered the echo
itself. `ping -s 1392` honest, `ping -s 1393` a lie, and only for the
outbounds the feature exists for.

65535 rather than merely "large": no IP datagram can exceed it, so the
kernel cannot fragment at this device for any packet ever. Anything
smaller leaves a band open and re-opens the class. It is also sing-box's
own default TUN MTU on Linux.

Memory was measured, not argued. Three paired runs of the integration
test under -test.memprofilerate=1 allocate 5.41/5.48/5.47 MB at 65535
against 5.76/5.46/5.70 MB at 1420, and a -diff_base profile puts every
difference in netlink interface enumeration. Nothing in the read path
scales with the MTU: gVisor reads through fdbased.BufConfig, which
sing-tun pins to one 65535-byte view regardless. I predicted a ~1.8 MB
saving from GSO switching off above 49152 and was wrong -- protocol/tun
turns GSO back on at StartStateStart whenever a FlowOutbound exists, so
the GRO scaffolding is there at both values. The corrected reasoning is
in the constant's comment so the next reader does not redo the mistake.

The integration test now reads the MTU back off the real kernel device,
which is the assertion the value exists for: a kernel that clamped it
would restore the forgery without changing a generated byte.

D25's KNOWN HOLE block is replaced with what is genuinely left. Chiefly:
a big non-DF ping does not start WORKING, it starts failing HONESTLY --
classifyReturn declines fragments on the way back too, so the packet
really leaves, the far host really answers, and the reply is not NAT'd
home. And a client that fragments on the wire itself is still uncovered;
that is the nft carve-out's job, with a warning that conntrack defrag
may reassemble in prerouting and leave such a rule unable to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:50:08 +03:00
omarandClaude Opus 5 81c96019b5 fix(panel): let a routing rule say ICMP, instead of calling one broken
The Proto picker was a closed list of the two transports and the ten
sniffed L7 labels, and anything else drew "<value> — never matches".
The engine now routes ICMP by rule (Rule.Proto accepts icmp, icmpv4,
icmpv6), so a working ping rule was rendered as a dead one and could not
be created here at all — the operator had to hand-edit /etc/config/shater
and then watch the panel call the result broken.

Adds a third group, "Layer 3". All three spellings are offered: they are
not synonyms — icmpv4/icmpv6 pin the rule's ip_version — so hiding the
narrowing would both strand a capability outside the UI and silently
widen such a rule the first time someone edited it here.

The doc comment no longer claims the list IS generate/route.go's
sniffedProtocols; only the middle group is. ICMP goes to the emitted
rule's `network`, never to `protocol`, which is the whole reason it never
matched as a sniffed label.

An unknown value is still kept and offered as written, but the
never-matches flag is now judged on the lower-cased value, the way the
engine judges it — a hand-written `ICMP` is a live rule, not an inert one.

Verified: npm run build clean (tsc --noEmit + vite build); an icmp rule
added through the panel renders as a plain "PROTO icmp" chip; no
horizontal overflow at 360px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:26:07 +03:00
omarandClaude Opus 5 baed8ff8f2 fix(model): fwmark_base 0x7f routes the engine's own traffic into its own TUN
The panel offers fwmark_base and table_base as free hex fields under
"Advanced" and nothing has ever checked them. What makes that more than a
footgun is that the derived values are invisible from the number typed: the
L3 mark is base+0x80, so 0x7f lands it exactly on 0xff — the loop-guard mark
the engine stamps on its OWN traffic — and `ip rule fwmark 0xff lookup 8200`
then captures everything the engine sends and routes it into the engine's
TUN. The router loses the internet the moment l3_tunnel is switched on, for
a reason nothing on screen connects to a collapsed section. fwmark_base 0xff
had produced the same failure since long before the L3 offset existed.

table_base is worse and got the same treatment: its derived values can land
on the kernel's own table ids, and teardown does `ip route flush table <n>`.
It is count-sensitive (egress #i uses base+0x10+i), so the check takes the
egresses rather than living in ValidateGlobals.

Written as "derive every value this layout produces, then look for
duplicates and reserved ids" rather than as a blacklist, so a future offset
is covered by construction. The layout constants are duplicated from
netplane (the import only runs one way) and pinned by netplane's
TestMarkLayoutConstantsLockstep.

Warn-only, like every check in this file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 f80fb4dd1b fix(netplane): give every mark-driven table a floor, and check the L3 pair
Two halves of the same omission.

1. A fwmark lookup that finds an empty table does not fail — it falls
   through to main. Every mark-driven table now gets an `unreachable
   default` at the maximum metric: it loses to any real default route while
   one exists, it has no device so the kernel never garbage-collects it, and
   it turns "lookup failed, try main" into "lookup succeeded: unreachable".
   The fallthrough stops depending on somebody reading a warning at the
   moment an interface goes down. Deliberately not gated on the kill-switch:
   that switch decides whether traffic may escape the tunnel, while an egress
   binding is a statement about WHICH UPLINK, and silently substituting a
   different one is not what "fail open" was meant to permit.

   RoutingPresent's "does this table have a default route" test is tightened
   in the same breath, or the floor would answer it and turn the safety net
   into a blindfold.

2. RoutingPresent had never heard of addL3Routing. This is the same defect
   its own comment describes as already caught twice ("a presence check must
   cover everything its Apply counterpart installs"), committed a third time
   — and its trigger needs no interface to go down: editing a node URI
   restarts the engine, the kernel destroys shater-l3 and takes `default dev
   shater-l3 table 8200` with it, the rendered nft text is unchanged, so the
   fast-path skipped ApplyRouting forever and LAN ping stayed dead until
   someone restarted the daemon.

TestRoutingPresentSeesL3Table, TestEgressTableGetsFailClosedFloor and
TestEveryStampedMarkIsRoutedAndVerified all fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 b71b793681 fix(netplane): a mark says where a packet was sent, not where it went
The forward chain let untunnelable-egress traffic past the kill-switch on
the strength of its fwmark alone. `ip rule fwmark X lookup N` does not
deliver the packet to table N, it delivers the LOOKUP there — and a lookup
that finds nothing falls through to main. So when the egress interface goes
down and the kernel garbage-collects its default route, every non-TCP/UDP
packet from the LAN is still stamped, still accepted here (above the
fail-closed drop), and leaves out the plain WAN with the router's real
address. Nothing we render changes, so no apply runs and nothing notices.

Ordinary egress traffic never had this hole: the engine binds those sockets
to the device, and a dead device fails the socket. The untunnelable-egress
path is made of nothing but a mark, so the accept now carries the second
opinion instead — `meta mark X oifname "dev"`, strictly narrower than either
half, true only when the routing did what the mark asked. The comment being
replaced argued correctly that oifname ALONE would be too loose, then drew
from that the conclusion that oifname should be dropped rather than added.

Same conjunction in the holding plane, where it is theory (that plane stamps
nothing) but where a bare mark accept has no business sitting.

Also folds the egress device resolution into one EgressDevice(), because the
binding and model.ValidateUntunnelableEgress had already drifted: the
validator trimmed the interface name and the binding did not, so `option
interface '   '` gave a panel saying "the option is ignored" over a data
plane that was marking packets for a table nobody built.

TestUntunnelableEgressAcceptIsBoundToItsDevice and
TestUntunnelableEgressResolutionLockstep fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:32 +03:00
omarandClaude Opus 5 61c87ad1d9 fix(l3): guard the ICMP honest-drop at PreMatch, not inside the walk
The drop that keeps a ping from reading as tunnelled lived in
preMatchFlow, overriding the pre-declared continueResult. That covered
every exit of THAT function and none of the walk above it: the
prepareMatchMetadata error return (which arrived later, with the shared
metadata refactor), the sniff bail-outs, and the default: arm of the
rule-action switch all returned PreMatchContinue on their own.
adapter.JudgeFlow maps Continue to tun.ActionAccept, and sing-tun answers
Accept by rewriting Echo into EchoReply itself -- the exact forgery this
delta exists to remove. Narrow paths, but paths.

PreMatch is now a funnel over the renamed preMatch walk, so the guard
sits on the single return value and cannot be outgrown by a new exit.
PreMatchBypass joins the drop: sing-tun implements ActionBypass on the
nfqueue plane only, so on the TUN path it lands in the same default: arm
as Accept and forges too.

Every ICMP case has an explicit TCP/UDP twin; the JudgeFlow mapping
table is pinned outright, including the one fix that must NOT be made
there -- refusing ActionFlow for a port whose address is not unspecified
would drop every ping through WireGuard/AWG, because the forward
dispatcher and the ICMP forwarder share that function with identical
arguments and only the latter needs an unspecified address.

That leaves a real hole open, now named in D25 rather than papered over:
a FRAGMENTED echo to a WireGuard/AWG outbound is still answered by the
router. The dispatcher returns before asking for a verdict at all when
the packet is a fragment, and the reassembled packet reaches the ICMP
forwarder, whose installFlow demands the unspecified address a WireGuard
endpoint never has. The two fixes that would close it both live outside
pre-match and are written down; the Consequence paragraph is scoped
until one lands.

The stack comment in generate/inbound.go repeated the "only gvisor
really forwards ICMP" argument that D25 itself retracts -- both stacks
run the same ForwardDispatcher first. Brought in line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:21:29 +03:00
omar 4dee508e12 fix(route): let a rule say "icmp", and say when saying it is a lie
`icmp` fell through ruleMatchers' proto switch into RawDefaultRule.Protocol —
the SNIFFED-L7 field, compared against what the sniffers labelled a connection.
Nothing ever labels a flow "icmp" (PreMatch skips the sniff action for an ICMP
flow outright), so the rule was structurally valid and permanently dead. That
made the whole L3 ingress unusable on a real config: with no way to write "ICMP
goes here", every ping fell to the catch-all, which resolves to the chain's last
hop — a group of VLESS nodes that cannot carry layer 3 at all.

icmp is a NETWORK. NetworkItem.Match is a map lookup over metadata.Network, and
adapter.JudgeFlow sets that to N.NetworkICMP for BOTH ICMPv4 and ICMPv6 (one
case covers both protocol numbers), so there is exactly one network value and it
covers both families. `icmpv4`/`icmpv6` narrow that same network with an
ip_version item instead of inventing a second one: metadata.IPVersion comes from
the destination address, and an ICMPv6 packet always has an IPv6 destination —
no false positives, no false negatives.

An ICMP rule that cannot fire is not a dead setting: ICMP has no fall-through,
so route.preMatchFlow DROPS it. Four ways to get that silently are now reported:
l3_tunnel off (nothing enters the engine at all), icmpv6 with ipv6 off (neither
the nft mark nor the TUN address exists), a port matcher next to it (JudgeFlow
zeroes both ports), and a target that cannot carry layer 3 — decidable from the
model, because the capability is fixed by the outbound TYPE: only wireguard/AWG
endpoints and the direct outbound behind direct/interface egresses declare
N.NetworkICMP. A mixed group gets its own text (the answer follows group.Now()),
`block` gets none (dropping the ping IS the policy), and an unresolved target
gets none either (ruleKillFallback already said the louder thing).

Wording stays clear of shater/apply's criticalMarkers on purpose: a failed ping
is fail-CLOSED, and a cosmetic alarm is how the real one stops being read.
2026-07-26 19:18:06 +03:00
omarandClaude Opus 5 76da5134ef test(gate): the two tests that need a kernel may not skip in silence
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.

Two halves.

Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.

Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.

The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 18:50:46 +03:00
omar 4c630c9a13 docs: handoff note for the L3 branch
Transient, to be deleted when omp/work merges. Everything meant to outlive the
merge is already in D25/D26 and the lx changelog; this file is the part that is
only useful while the branch is still a branch — the verification commands, the
testbed recipe, what was proven on hardware and what was not, and the six files
that will conflict on rebase.
2026-07-26 18:31:59 +03:00
omar d8dbefcd07 docs: record the AWG site-to-site path as declined, not impossible
D26's "no port-like selector" line disposes of NAT-based forwarding and nothing
else, and read alone it says "impossible" — which is false and would be
re-derived at the cost of another research pass. The endpoint is protocol-blind
in both directions, so ESP could ride it untouched with the client's own source
address and no NAT whatsoever. That was declined for two reasons worth naming:
lx-owned code in the forward hot path, and a server-side AllowedIPs prerequisite
that turns a router option into a deployment contract.
2026-07-26 18:31:59 +03:00
omar 974208fc05 docs: record the kernel egress, and retract the reason D25 gave for the ceiling
D26 writes down where the engine's boundary actually is, because the intuitive
answer is wrong and someone will look for it again: the WG/AWG forward path
never consults gVisor in either direction, so the limit is sing-tun's
ForwardDispatcher — its parser and its port-shaped NAT — and the kernel egress
was chosen because it clears that limit without a line of new hot-path code, not
because userspace "cannot". Tailscale documents the same boundary for their
userspace mode and is quoted as corroboration, with the caveat that ours sits at
the dispatcher rather than the stack.

D25 said two things that do not survive checking, and both are corrected in
place rather than left for the next reader to trip over. It blamed the netstack
for the ICMP-echo ceiling; that was the dispatcher. And it called `stack: gvisor`
mandatory because the system stack fakes ping — the system stack runs the very
same dispatcher first and only forges an echo for packets the dispatcher
declined, so gvisor is a deliberate choice (already linked via with_wireguard,
and the combination the integration test exercises), not a necessity.

The operator note says what the option buys and refuses to call an egress a
tunnel on its own say-so: with a WireGuard device it is one, with a second WAN
the destination sees that uplink's address. It also says what the option does
not fix — multicast IPTV stays broken — and that IPsec through NAT-T is ordinary
UDP that never needed any of this.
2026-07-26 18:31:59 +03:00
omar 2eb71e8244 feat(netplane,model): hand the protocols the engine will not dispatch to the kernel
ESP, AH, GRE, IGMP and SCTP cannot enter the engine, and the reason is not the
one that looks obvious. A WireGuard or AmneziaWG endpoint forwards straight past
its gVisor stack — WritePackets reads the IP version and the destination address
and hands the raw bytes to the device, and the return path offers every
decrypted packet back before the stack sees it. WireGuard would carry ESP today
if anything handed it one. What refuses is sing-tun's ForwardDispatcher: its
parser recognises TCP, UDP and ICMP echo, and its NAT wants a port-shaped
selector that ESP, AH and GRE do not have. The retracted rationale is corrected
where it was written down, not quietly dropped.

So these protocols go to the kernel instead. untunnelable_egress names an
interface or tunnel egress; prerouting stamps that egress's OWN mark on
everything that is not TCP or UDP, and addEgressRouting has already bound that
mark to a table whose default route leaves via the device. Every protocol works
because nothing in the path has to understand any of them. No new mark, no new
table, no new code in the hot path.

Whether that is a tunnel depends on the device, and nothing here claims
otherwise: a WireGuard interface is one, a second WAN is a different uplink
whose real address the far end sees.

The wide `!= { tcp, udp }` filter is safe here and stays banned for the L3
ingress, for the same reason stated in both places: there the receiver is a
dispatcher that knows four protocols, here it is the kernel. ICMP is claimed by
the L3 ingress first when both are on. The local plane keeps its exclusions —
router-addressed traffic, private destinations, ICMPv6 ND/RA — and with IPv6 off
the marking is scoped to v4, because addEgressRouting installs no v6 rule then
and a marked v6 packet would fall into the main table.

An interface egress with an empty `interface` no longer resolves: IfaceDevice
defaults to br-lan, so it passed the binding while addEgressRouting skipped it —
mark set, no rule, straight past a closed kill switch and out the default WAN.
2026-07-26 18:31:59 +03:00
omar 668cccbf24 test(generate): the L3 device name is a singleton, so wait for the kernel to take it back
Both gated tests stand an engine up on shater-l3. Run together, the second met
`TUNSETIFF: device or resource busy` and failed for a reason that had nothing to
do with what it asserts — the first had closed its box and yielded while
unregister_netdevice was still catching up. Each passed alone, which is the
shape of a fixture bug that gets rediscovered rather than fixed.

The poll that already guarded the first test is now a shared helper both call.
It stays a poll rather than a sleep for the reason it always was: the removal is
usually immediate and a fixed wait would be either flaky or slow.
2026-07-26 18:31:59 +03:00
omar 4ea4585402 test(generate): pin that ping through an interface egress is real, and byedpi's is not
An interface egress is a direct outbound carrying BindInterface and a routing
mark, and direct builds its ICMP port from the very same dialer control — so
ping routed at that egress leaves through that device, marked, like every other
packet bound to it. Nothing said so. Both halves of that sentence are one
`common.Cast[*dialer.DefaultDialer]` away from being false: if the dialer ever
stops being a DefaultDialer, icmpPort is nil, PreMatchFlow declines, and ping
through the egress degrades to a drop without a single generated byte changing.
The gated test asserts the live outbound, not the config, because that is where
the cast happens.

The failure the codegen half guards is worse than a broken ping: losing
BindInterface or the mark does not stop the echo, it sends it out the main table
over the plain WAN with the real address, which is the one thing an egress
exists to prevent.

byedpi is a SOCKS outbound and cannot be a tun.Port, so ICMP aimed at it is
dropped. That is the honest end of l3-honest-drop and it is pinned too, because
the alternative the TUN stack offers is a forged reply.
2026-07-26 18:31:59 +03:00
omar dc6d102473 docs: put a number on the second netstack, and say what it does not bound
Measured on a throwaway harness in a container: peak RSS of a process that
brought the engine up went from ~26 MB to ~28 MB with l3_tunnel on, three
paired runs. It is x86_64, idle, with an empty ICMP NAT table, so it stays
listed as unverified for the router — an indicative figure is more useful than
silence only if it says loudly what it is not.
2026-07-26 18:31:58 +03:00
omar 683afc0a47 docs: record how ping got through the tunnel, and where it stops
D25 writes down the reasoning that is expensive to reconstruct: why a TUN rather
than TPROXY, why the interface is its own with auto_route off, why gvisor is
mandatory rather than preferred, and why the ceiling is ICMP echo — a boundary
in sing-tun's flow parser and gVisor's protocol set, not an unfinished edge of
ours. It also records what carries layer 3 and what does not, that masque could
and does not, and the two things still unproven: the live-router path end to
end, and what a second gVisor NIC costs in memory on the hardware.

D17 gains one line: its claim that TPROXY cannot carry ICMP is still true, and
is no longer the end of the story.
2026-07-26 18:31:58 +03:00
omar 2c3e20512e feat(openwrt): let fw4 know the L3 tunnel device before it exists
Both nft tables run and a drop in either one wins, so our forward accept for
shater-l3 decides nothing on its own: fw4 sees a device in no zone and drops the
forward, and the feature fails with exactly the symptom it was built to fix —
ping does not work, and nothing says why.

The zone names the device directly rather than a network. fw4 resolves a zone's
networks through netifd, and a proto-none interface for a device the daemon
creates is never up and contributes nothing, so list network would compile to an
empty device set. list device compiles to a plain iifname/oifname match that is
valid before the TUN exists and starts matching the moment shaterd creates it,
with no firewall reload at enable time.

It is seeded unconditionally, not gated on l3_tunnel: uci-defaults run once, and
a zone naming an absent device is inert. Gating it would mean the option could
be switched on and never take effect. The sections are named so a re-run is a
no-op instead of a second zone, and kmod-tun joins DEPENDS because /dev/net/tun
is not on a stock image.
2026-07-26 18:31:58 +03:00
omar 51b2f04672 feat(netplane,generate): carry LAN ping through the tunnel, on a TUN of its own
Kernel TPROXY needs a socket to hand a packet to, so it moves TCP and UDP and
nothing else. Everything else reached the forward chain and met the untunnelable
policy, whose best answer was "let it out with your real address" and whose
default was "drop it" — so on a stock install ping simply did not work, and the
setting that fixed it did so by leaking.

The engine has been able to do better for a while: sing-tun's ForwardDispatcher
does real ICMP forwarding with NAT on the echo id, and a WireGuard or AmneziaWG
endpoint is a tun.Port that carries the packet for real. What was missing was a
way in, because nothing on the router could hand it an IP packet.

l3_tunnel (opt-in, off by default) adds one: the generator emits an "l3-in" TUN
inbound and prerouting fwmarks LAN ICMP into it. The interface is its own and
auto_route is off, so the main routing table is never touched and the fwmark
plus addL3Routing's ip rule are the only entrance — the TPROXY plane is byte for
byte what it was. gvisor is not a preference: the system stack forges echo
replies locally, which is the very thing this is meant to end.

Only icmp and ipv6-icmp are ever marked, and only after the local plane is out
of the way — the router itself, private destinations, and ICMPv6 ND/RA, which
mean nothing off-link and take v6 down if one neighbour probe is tunnelled.
ESP, AH, GRE, IGMP and SCTP are deliberately left alone: sing-tun's parser and
gVisor's stack know no such protocol, so marking them would black-hole the
traffic while looking like a feature. They stay with the untunnelable policy,
which also keeps its say over what happens if the ip rule fails to install.

Ping and Windows tracert now cross the tunnel; IPv6 traceroute shows only the
destination, because the return path recognises TimeExceeded for v4 alone.
2026-07-26 18:31:58 +03:00
omar f190c8251e feat(lx): stop answering ping on behalf of a tunnel that never saw it
PreMatchContinue is not "fall back to the ordinary route" the way it is for TCP
and UDP. An ICMP flow has no ordinary route: the TUN stack takes the packet back
and answers the echo itself, swapping the addresses and writing a reply
(sing-tun stack_gvisor_icmp.go). So a ping routed to any outbound that cannot
carry layer 3 — every proxy protocol; only adapter.FlowOutbound can — came back
successful, and the operator read a working tunnel off a packet that was never
sent.

That is worse than the packet loss it replaced. Loss is a fault the operator can
see and chase; a forged reply is a fault that reports itself as health, and it
reports it on the one tool anyone reaches for first.

preMatchFlow now overrides continueResult once, at the top, for
N.NetworkICMP. One hunk covers every exit that used to fall through — no such
outbound, a group whose selection is gone, an outbound whose Network() omits
icmp, an outbound that is not a FlowOutbound — and keeps the diff to three lines
against a function upstream will keep editing. JudgeFlow carries the same
verdict in its !isPort branch, because FlowOutbound and tun.Port are separate
interfaces and drift between them must not reopen the forgery.

TCP and UDP are untouched, and the test pins that as hard as it pins the drop.
2026-07-26 18:31:58 +03:00
omarandClaude Opus 5 1945404eaa fix(armor): a reboot is not someone switching the product off
test / go + panel tests (push) Successful in 5m24s
release / test gate (push) Successful in 5m24s
release / apk aarch64_cortex-a53 (push) Successful in 3m9s
release / apk x86_64 (push) Successful in 3m9s
release / release apk (push) Successful in 8s
The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:

    case $action in restart|reload) keep;; *) DISARM;; esac

`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:

    13:28  /etc/shater/boot.nft present
    ----   reboot
    18s    at_S22: NO_TABLE  armor_file=NO_FILE

It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.

Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.

Also closed, found while proving the above:

  * Every restart left the LAN in the clear for 80-90ms. The exit path was
    `Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
    transactions with no `inet shater` between them, leaving fw4's
    `lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
    & Apply. TeardownExiting arms first under the apply lock and skips the
    delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
    REPLACES the table, so the kernel never observes its absence.
    35k-sample instrument: 7 and 6 no-table hits before, 0 across three
    runs after.

  * SaveBootArmor fsynced the payload but not the directory, so a power cut
    could lose the rename that publishes it -- a boot with no armor and no
    error anywhere.

`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.

Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 17:33:53 +03:00
omarandClaude Opus 5 6476722372 fix(panel): stop shipping a fabricated router in the binary
test / go + panel tests (push) Successful in 5m26s
release / test gate (push) Successful in 5m28s
release / apk aarch64_cortex-a53 (push) Successful in 6m7s
release / apk x86_64 (push) Successful in 3m5s
release / release apk (push) Successful in 7s
mock.ts was a static import and the mock switch was read from the query string at
runtime, so the bundle that ships inside the daemon carried a complete fictional
router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive,
without a single request to the daemon. The only tell was a line in the footer.
That is worse than any wrong number — there is no data at all and nothing says
so. It is out of the production bundle now, which is 21 kB smaller for it.

Unknown state stopped reading as good news in two more places. The kill-switch
tile treated an absent plane as armed, because the check was "not none" and
undefined satisfies it — the contract in the API types says the opposite. And the
apply page announced "daemon auto-rolled back" from its own timer, while the
daemon, seeing the state generation move, disarms and says it is NOT rolling back
in the log only.

Alerts moved to Settings. They are about the kill switch, apply failures, new
devices and subscription expiry, and they lived at the bottom of the DNS page,
while Settings mentioned them in prose with nothing to click.

Findings truncation is visible now: the notice that says how many were suppressed
arrives as info, and the attention list keeps only critical and warning, so past
fifty findings the operator saw forty-nine and no hint of the rest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:39 +03:00
omarandClaude Opus 5 4078334d85 fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.

They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.

Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.

The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.

And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.

The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:21 +03:00
omarandClaude Opus 5 0a6689b29e fix(quic,v2ray): close the sockets quic-go was never going to close
DialEarly with a packet conn the caller made sets a flag that means quic-go does
not own it: closing the transport only stops reading from the socket. Neither DNS
transport closed it. On the QUIC one it was closed on a failed handshake and
never on success, so every redial — idle timeout, retry error, engine reload —
left a UDP socket for the life of the process. On the HTTP/3 one the library
drives its own reconnects, so the leak compounds without anything in our code
looking wrong.

That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on
every reconnect without closing the previous one. Both are now owned by a watcher
tied to the connection's own context, so the socket lives exactly as long as the
connection does.

This matters more than it did last week: the shipped resolvers are DoH, and DNS
is intercepted by default now, so the whole network's query stream rides this
path on a router with 512 MB.

The same upstream commit fixes both halves. We had taken the v2ray half and not
the DNS one — the third time this session a paired fix arrived half-applied, and
the first of those cost a day of debugging. These two files are now byte-identical
to upstream so a rebase cannot reopen it.

Also from that family: websocket and httpupgrade leaked their conn on failed
handshakes, and a QUIC stream's Close did not release a blocked write.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:59 +03:00
omarandClaude Opus 5 ef22167b1a fix(apply): report the hold when the plane was armed by someone else
Booting with the armor loaded, or restarting through the handoff, left the status
saying the LAN was not being held while it was being dropped. Transient after a
successful apply, but permanent on the unreadable-config path — and there the
apply-failure alert words itself "traffic is NOT being blocked" at the exact
moment it is. That sends the operator to fix something that is not broken, past
the protection that is holding.

The table cannot be identified from here — netplane exposes no read-back and nft
does not keep comments — but identifying it is the wrong question. Holding does
not claim the holding plane is the object in the kernel; it claims the engine is
down and forwarded traffic is being dropped. A leftover full ruleset does that
too: with no engine socket the tproxy statement breaks its own rule before the
accept, so the packet reaches the forward chain unmarked and meets the primary
drop. What decides it is whether the last applied config was enabled and
fail-closed, which is exactly what the boot armor's presence already means.

So it is derived at read time rather than latched. A latch set from an inference
would have to be remembered in order to be cleared, which is the trap the active
flag already taught us. ArmHold also stops deferring to a table it cannot
inspect and installs its own render instead — the honest answer to "do not claim
a foreign table blindly" is to make it ours, and a fresh render beats a snapshot
that predates an interface rename.

Also closes the last of the detour transports: the subscription fetch took a
client and dropped it, and the exits that leak are the error ones, retried by
cron forever against a broken feed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:39 +03:00
omarandClaude Opus 5 cbda0fee0a fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.

The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.

The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.

An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.

A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:16 +03:00
178 changed files with 28803 additions and 1642 deletions
+9 -5
View File
@@ -35,11 +35,15 @@
# invalidates every deployed router's trust.
#
# AUTO-RELEASE
# push a tag `vX.Y.Z` -> versioned per-arch releases `apk-vX.Y.Z-<arch>`.
# workflow_dispatch -> rolling per-arch `apk-latest-<arch>` (always-fresh
# feed). Publish uses the Gitea API via curl (ci/gitea-release.sh) — no
# external action needed. NOTE: the apk release tags deliberately do NOT start
# with `v` so publishing them cannot re-trigger this workflow's `v*` filter.
# The rolling per-arch `apk-latest-<arch>` is published on EVERY run — tag runs
# included — and then read back over the API to assert it really serves the
# version just built. A tag push `vX.Y.Z` publishes the pinnable per-arch
# `apk-vX.Y.Z-<arch>` IN ADDITION. It is not an either/or: it used to be, and
# the rolling pointer then froze at 0.2.0 while v0.2.9/v0.2.10 shipped (see the
# long comment above the `release-apk` job). Publish uses the Gitea API via curl
# (ci/gitea-release.sh) — no external action needed. NOTE: the apk release tags
# deliberately do NOT start with `v` so publishing them cannot re-trigger this
# workflow's `v*` filter.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
+141 -47
View File
@@ -6,61 +6,155 @@
## Правила делегирования
1. ЛЮБАЯ реализация (код, тесты, конфиги, рефакторинг, отладка) выполняется
субагентами через инструмент Agent с `model: "fable"`. Сам ты правишь файлы
только в одном случае: тривиальная правка в 1–2 строки, где постановка
задачи дороже самой правки.
субагентами через инструмент Agent. Сам ты правишь файлы только в одном
случае: тривиальная правка в 1–2 строки, где постановка задачи дороже самой
правки.
2. Перед делегированием ты сам исследуешь код настолько, чтобы написать
точное ТЗ. В каждом задании субагенту обязательно указывай:
- контекст: что это за проект и над чем идёт работа;
- конкретные файлы и функции, которые нужно менять (пути, а не «найди сам»);
2. **Модель выбирает исполнитель задачи, а не привычка.** `fable` — быстрый и
дешёвый, годится для механической работы с ясным контрактом. `opus` — для
всего, где нужно рассуждение: поиск причины, аудит, дизайн, работа в чужом
коде. Если у `fable` кончилась квота — молча переходи на `opus`, это не повод
останавливать работу. Не спрашивай владельца, какую модель брать.
3. Перед делегированием ты сам исследуешь код настолько, чтобы написать точное
ТЗ. В каждом задании субагенту обязательно указывай:
- контекст: что за проект и над чем идёт работа;
- конкретные файлы и функции (пути, а не «найди сам»);
- контракт: сигнатуры, форматы данных, инварианты, что менять НЕЛЬЗЯ;
- definition of done: как проверить, что задача выполнена
(какие команды/тесты прогнать и какой ожидается результат);
- что вернуть в финальном ответе: список изменённых файлов, результаты
проверок, найденные проблемы и принятые решения.
- definition of done: какие команды прогнать и какой ждать результат;
- что вернуть: изменённые файлы, результаты проверок, найденные проблемы,
принятые решения.
3. Скиллы: при постановке задачи посмотри список доступных скиллов и ЯВНО
перечисли в ТЗ, какие скиллы субагент обязан вызвать через инструмент Skill
до начала работы (например: «сначала вызови Skill "openwrt-procd-services"
и следуй ему»). Субагент не видит наш диалог и сам не догадается — пиши
названия скиллов прямо в текст задания.
4. **Скиллы использовать по максимуму — и тебе, и агентам.** Это не
формальность: в них лежит выстраданное знание по ровно тем предметным
областям, в которых мы работаем, и игнорировать их — значит переоткрывать
чужие грабли. См. раздел «Скиллы» ниже.
4. Независимые задачи запускай ПАРАЛЛЕЛЬНО — несколько вызовов Agent в одном
сообщении, каждый с `model: "fable"`. Зависимые — последовательно, передавая
в следующее ТЗ результаты предыдущего.
5. Независимые задачи запускай ПАРАЛЛЕЛЬНО — несколько вызовов Agent в одном
сообщении. Зависимые — последовательно, передавая результаты предыдущего.
**Делишь файлы между параллельными агентами явно** и пишешь каждому, кто ещё
работает в дереве и что трогать нельзя. Запрещай им `git stash`,
`git checkout <файл>`, `git reset` — в этом проекте агент уже сносил правки
соседа через `git stash push`.
5. Приёмка: результат каждого субагента ты проверяешь сам (читаешь diff
ключевых мест, гоняешь проверки из definition of done). Если результат
не принят — не переделывай сам, а верни задачу: доработку заказывай тому же
агенту через SendMessage (у него сохранён контекст), а не новым спавном.
6. Приёмка: результат каждого субагента ты проверяешь сам — читаешь diff
ключевых мест, гоняешь проверки из definition of done. Не принимай отчёт на
слово: сегодня отчёт «тесты зелёные» дважды сопровождался тестом, который
ничего не прибивал. Если результат не принят — не переделывай сам, а верни
задачу тому же агенту через SendMessage (у него сохранён контекст).
6. Финальный отчёт пользователю: что сделано, кем (сколько агентов),
что проверено, что осталось.
7. Финальный отчёт владельцу: что сделано, сколько агентов, что проверено,
**что осталось непроверенным и почему** — последнее так же важно.
## Инженерные стандарты
Это не пожелания. Каждый пункт здесь появился после того, как его отсутствие
стоило рабочего дня.
- **Тест обязан быть проверен мутацией.** Откатить фикс → показать, что тест
падает, и с каким текстом → вернуть фикс. Тест, не падающий на сломанном коде,
не тест, а украшение.
- **Прибор без контроля не доказывает ничего.** Отрицательный результат чего-то
стоит, только если показано, что этот же прибор умеет дать положительный.
«Утечки не нашли» прибором, который не мог её увидеть, — это не результат.
- **Опровержение ценнее согласия.** В каждом ТЗ прямо разрешай субагенту
сказать «твоя версия неверна» и требуй доказательства, а не вежливости.
Лучшие результаты этого проекта приходили именно так.
- **Не обещать непроверенного.** Комментарий, предупреждение и текст в панели —
это утверждения о поведении. Если поведение не проверено, так и писать.
Формально верная фраза, которая читается как «работает», — тоже ложь.
- **Умолчание падает в восстановимую сторону.** Открытый `default:` в разборе
вариантов — источник целого класса дефектов: неучтённое значение уходит туда,
где дороже всего ошибиться. Списки делать положительными и закрытыми.
- **Проверка присутствия обязана покрывать всё, что ставит её Apply-двойник.**
Иначе идемпотентный быстрый путь становится ловушкой: «всё на месте» при
отсутствующем маршруте.
- **Никакого молчаливого скипа.** Тест, который не выполнился, обязан быть
назван поимённо в выводе гейта. Однажды CI гонял два теста из 116 файлов, и
все считали, что покрыто.
## Скиллы
**Правило: если задача касается области, по которой есть скилл, — скилл
вызывается ДО начала работы, а не после того, как что-то не заработало.**
Это относится и к тебе, и к каждому субагенту.
Субагент не видит наш диалог и сам не догадается, что скиллы существуют.
Поэтому **в каждом ТЗ перечисляй поимённо**, какие скиллы он обязан вызвать
через инструмент Skill: «сначала вызови Skill "openwrt-nftables" и Skill
"openwrt-networking", следуй им». Требуй в отчёте сказать, что именно из скилла
он применил, — так видно, вызвал он его или упомянул.
Соответствие областей этого проекта и скиллов:
| Трогаешь | Обязательные скиллы |
|---|---|
| `/etc/config/*`, `uci`, uci-defaults, парсер модели | `openwrt-uci` |
| nftables, fw4, зоны, метки, tproxy, kill-switch | `openwrt-nftables` |
| интерфейсы, мосты, VLAN, policy routing, `ip rule`, sysctl, dnsmasq | `openwrt-networking` |
| init-скрипты, procd, respawn, service triggers, boot armor | `openwrt-procd-services` |
| перехват трафика целиком (tproxy + маршрутизация + DNS) | `openwrt-transparent-proxy` |
| сборка пакетов, SDK, фид, CI, подпись, `apk`/`opkg` | `openwrt-package-build-ci`, `openwrt-native-packages` |
| LuCI-приложение, ubus/rpcd, ucode | `openwrt-luci-plugin`, `openwrt-ubus-rpcd`, `openwrt-ucode` |
| панель (React/TS) | `react-expert`, `frontend-design:frontend-design` |
| Go: конкурентность, каналы, профилирование, идиоматика | `fullstack-dev-skills:golang-pro` |
| TypeScript | `fullstack-dev-skills:typescript-pro` |
| стратегия тестирования, покрытие, тестовые данные | `fullstack-dev-skills:test-master` |
| поиск причины по логам и трассам | `fullstack-dev-skills:debugging-wizard` |
| проверка в браузере, скриншоты | `fullstack-dev-skills:playwright-expert` |
| ревью | `review`, `fullstack-dev-skills:code-reviewer` |
| безопасность | `security-review`, `fullstack-dev-skills:security-reviewer` |
| графики и визуализация данных | `dataviz` |
Список неполный — **смотри доступные скиллы под задачу**, а не только в эту
таблицу. Если скилл выглядит смежным, дешевле вызвать его и не воспользоваться,
чем не вызвать и потом отлаживать то, что там уже описано.
## Проверки
- **Гейт:** `bash scripts/run-tests.sh` — Linux в Docker, боевой набор тегов,
`-race`, и шаг, требующий вердикта по имени для привилегированных тестов.
Зелёный гейт — необходимое условие, но не достаточное: он не видит стыков с
ядром, procd и nftables.
- **Стенд:** сервер `local_openwrt` в ssh-manager — ImmortalWrt 25.12.1 той же
ревизии, что боевой роутер. Сюда — всё, что касается init-скриптов, nft,
policy routing, TUN.
- **Боевой роутер:** `mini_router` (BPI-R3), через него идёт весь домашний
трафик. Перед изменением конфигурации — резервная копия. Проверять приборно,
а не по логу: лог может печатать одно и то же в честном и в ложном случае.
## Релиз и деплой
- Тег → CI (Gitea Actions) → apk-фид → установка на роутер.
- **Обновлять только поимённо**, никогда не `apk upgrade` целиком:
`apk upgrade shaterd shater-core luci-app-shater byedpi`.
- **Не трогать кеш CI-раннера** — сборка растянется на часы.
- Число тегов на порцию работы — на твоё усмотрение, если владелец не сказал
иначе.
## Фронтенд (admin panel)
Дизайн-направление ЗАФИКСИРОВАНО: **Faceplate** (панель сетевого железа).
Полная спека, токены, компоненты и ссылка на живой эталон — в
[`docs-shater/DESIGN.md`](docs-shater/DESIGN.md). Эталон:
https://claude.ai/code/artifact/9f7c07e8-d8ac-4ae1-b113-5b25d0ba5dd2
Спека, токены и компоненты — в [`docs-shater/DESIGN.md`](docs-shater/DESIGN.md).
Эталон: https://claude.ai/code/artifact/9f7c07e8-d8ac-4ae1-b113-5b25d0ba5dd2
- **Стек:** Vite + React + TypeScript, лёгкий (SPA встраивается в бинарь —
без тяжёлых зависимостей). Расположение: папка `panel/` в корне.
- **Порядок работ:**
1. Сам (оркестратор) скаффолдишь `panel/`, переносишь токены из
`docs-shater/DESIGN.md` в `panel/src/tokens.css` один-в-один и задаёшь каркас
компонентов. Это фундамент — делай аккуратно сам или отдай ОДНОМУ агенту.
2. Дизайн-систему в компоненты: `<Faceplate> <Module> <Toggle> <Led>
<SegMeter> <QueryLog>` + кнопки — строго по эталону.
3. Страницы раздаёшь ПАРАЛЛЕЛЬНО Opus-агентам (`model: "opus"`), по одной на
агента: Overview, Nodes/Subscriptions, Routing rules, DNS/Blocklists,
Devices, Apply/Rollback.
- **В КАЖДОМ ТЗ агенту обязательно:** ссылка на `docs-shater/DESIGN.md` и на эталон;
требование сначала вызвать Skill `react-expert` и Skill
`frontend-design:frontend-design` и следовать им; список готовых компонентов,
которые он ДОЛЖЕН переиспользовать (не изобретать заново); какие токены и
семантические цвета применять; DoD — страница совпадает с языком эталона,
адаптив + фокус + reduced-motion соблюдены.
- **Не отходить от Faceplate.** Любой новый экран наследует ту же визуальную
систему. Оранжевый — только акцент; семантика good/warn/crit — отдельно.
- **Стек:** Vite + React + TypeScript в `panel/`. SPA встраивается в бинарь —
тяжёлые зависимости недопустимы.
- **Панель целиком на английском.** Ни одного символа кириллицы в `panel/src`.
- **В КАЖДОМ ТЗ на панель:** ссылка на `DESIGN.md` и на эталон; требование
сначала вызвать Skill `react-expert` и Skill
`frontend-design:frontend-design`; список существующих компонентов, которые
надо ПЕРЕИСПОЛЬЗОВАТЬ (`<Faceplate> <Module> <Toggle> <Led> <SegMeter>
<QueryLog>` и кнопки), а не изобретать заново; какие токены и семантические
цвета применять; DoD — совпадение с языком эталона, адаптив, фокус,
`prefers-reduced-motion`.
- Оранжевый — только акцент; семантика good/warn/crit — отдельно.
- **Панель не должна врать про состояние.** Значение, которое движок примет,
не может рисоваться как «never matches»; настройка, которой управляет другая
подсистема, не может описываться так, будто управляет ею.
+30 -5
View File
@@ -24,7 +24,8 @@ The engine is a **fork of [sing-box](https://github.com/SagerNet/sing-box) via
[sing-box-lx](https://github.com/Leadaxe/sing-box-lx)**, compiled into a single Go
binary `shaterd` together with the control plane, DNS filter, stats aggregator and
the web panel itself. Broad protocol set: VLESS/VMess/Trojan/Shadowsocks,
Reality/XTLS, WireGuard, **AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
Reality/XTLS, WireGuard, **AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP — exactly what
`shater/parse` can read and `shater/registry` registers in the engine.
A thin **LuCI launcher** (mini-dashboard + "Open panel" button) hands the browser a
single-use token into the standalone SPA the daemon serves on its own port
@@ -39,8 +40,8 @@ single-use token into the standalone SPA the daemon serves on its own port
selector / chain / direct / block; node groups with balancer/observatory;
multi-hop chains; per-rule egress.
- **Fail-closed kill-switch** (dead group → block, never a silent direct leak); own
`inet shater` nft table; atomic apply with `nft -c` validation and commit-confirm
auto-rollback.
`inet shater` nft table; atomic apply with `nft -c` validation. Commit-confirm
auto-rollback exists but **ships OFF** (`confirm_timeout=0`) — arm it yourself.
- **DNS filtering & blocklists** with flexible sources (inline / file / url /
geosite), compiled `.srs` matcher; Block-DoH/DoT to stop filter bypass.
- Subscriptions (Clash / sing-box / Xray-JSON) and manual nodes; node health board.
@@ -68,8 +69,23 @@ keeps the router current. Point the repo line at `apk-vX.Y.Z-<arch>` instead to
pin a build; that file then has to be edited by hand for every upgrade.
shater ships **inert** (globals off) so install never breaks connectivity. After
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
then `shaterd apply` and `shaterd confirm`.
configuring nodes/rules:
```sh
uci set shater.globals.enabled=1
uci set shater.globals.confirm_timeout=120 # commit-confirm ships OFF — arm it
uci commit shater
shaterd apply && shaterd confirm
```
Without that middle line `shaterd apply` arms no auto-rollback (and says so), so an
apply that costs you SSH/LuCI access has to be undone by hand.
Once an enabled, fail-closed config has been applied, `/etc/init.d/shater-armor`
loads a saved fail-closed plane at **boot**, before the daemon exists: LAN→WAN
forwarding is blocked until `shaterd` applies, while SSH/LuCI/the panel stay
reachable on purpose (the chain hooks `forward` only). What arms it, what refuses
to arm, and how to switch it off — `INSTALL.md` §4.
## Build from source
@@ -78,6 +94,15 @@ then `shaterd apply` and `shaterd confirm`.
into `openwrt/shaterd/files/`. Details in
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
`bash scripts/run-tests.sh` is the test gate: the whole suite under the **shipped**
build tags (`scripts/router-tags.sh`), on linux (it re-execs in Docker from a
non-linux host), with `-race`, plus three machine checks against a silent skip —
the tag set may only add test files, every package with tests must report `ok` by
name, and every `TestIntegration*` must produce a verdict by name.
`scripts/check-router-tags.sh` separately proves no feature declared in
`FEATURES.md` lost a build tag it needs. A green gate is necessary but not
sufficient: it does not see the kernel, procd or nftables seams.
## Repository layout
| Path | What |
+56 -8
View File
@@ -27,7 +27,8 @@ BananaWRT** (Banana Pi BPI-R3, BPI-R4 и совместимые). Он проз
Go-бинарь `shaterd` вместе с control-plane, DNS-фильтром, агрегатором статистики и
самой веб-панелью. За счёт sing-box поддерживается широкий и актуальный набор
протоколов: VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, WireGuard,
**AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
**AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP — ровно то, что умеет разобрать
`shater/parse` и что регистрирует `shater/registry` в движке.
Интеграция в OpenWrt — тонкий **LuCI-лаунчер**: мини-дашборд и кнопка «Открыть
панель», которая по одноразовому токену передаёт браузер в полноценную SPA-панель,
@@ -50,8 +51,10 @@ Go-бинарь `shaterd` вместе с control-plane, DNS-фильтром,
- **Fail-closed kill-switch**: мёртвая группа → block, а не тихая утечка мимо
прокси; собственная nft-таблица `inet shater` и свои марки/таблицы, fw4 не
трогаем.
- Атомарный apply с валидацией движком и `nft -c`, **commit-confirm** с
авто-откатом к последней рабочей конфигурации.
- Атомарный apply с валидацией движком и `nft -c`. **Commit-confirm** с
авто-откатом к последней рабочей конфигурации есть, но **на стоковой установке
выключен**: `confirm_timeout` поставляется нулём, и apply не вооружает ничего,
пока вы не зададите окно (см. «Включение»).
- Идемпотентный reconcile из hotplug/boot под flock; management-bypass
(SSH/LuCI/LAN) всегда в обход.
@@ -121,8 +124,11 @@ flowchart TB
Путь трафика: LAN-клиент → `nft tproxy` (mark → tproxy-порт) → tproxy-inbound
sing-box (сниффинг SNI/Host/QUIC) → маршрут по правилу → outbound/selector/chain
(проксировано) · direct (flow-offload) · block. Подробные диаграммы (auth-handoff,
data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md).
(проксировано) · direct (обычный маршрут, без туннеля) · block. TPROXY несёт
только TCP и UDP; ICMP и остальные протоколы — через отдельные опциональные
механизмы (`l3_tunnel`, `untunnelable_egress`, ARCHITECTURE §3a). Подробные
диаграммы (auth-handoff, data-plane, DNS-flow, apply-flow) — в
[`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md).
---
@@ -195,15 +201,32 @@ shater ставится **инертным** (globals выключены), чт
```sh
uci set shater.globals.enabled=1
# Предохранитель: commit-confirm поставляется ВЫКЛЮЧЕННЫМ (confirm_timeout=0),
# и без этой строки apply ничем не подстрахован. 120 с — окно на проверку связи.
uci set shater.globals.confirm_timeout=120
uci commit shater
shaterd apply # apply + вооружить commit-confirm на живом демоне
shaterd confirm # подтвердить (отменяет авто-откат)
shaterd apply # применить и вооружить авто-откат на 120 с
shaterd confirm # подтвердить в пределах окна (отменяет авто-откат)
```
`shaterd apply` печатает, вооружил ли он что-нибудь, и почему нет: при
`confirm_timeout=0` он прямо говорит, что автоматического отката НЕТ. Оставить
ноль — сознательный выбор: тогда apply, отрезавший вам SSH/LuCI, придётся
откатывать руками.
`/etc/init.d/shater enable && /etc/init.d/shater start` поднимает демона под procd.
Кнопка «Открыть панель» в LuCI чеканит одноразовый токен и передаёт браузер в
панель (`:8088` по умолчанию).
После первого же применённого включённого fail-closed конфига появляется
**загрузочная защита**: `/etc/init.d/shater-armor` (START=21) грузит сохранённый
fail-closed план ещё до старта демона, закрывая те секунды между поднятием LAN и
первым apply, когда роутер форвардил трафик в WAN открытым. Форвардинг LAN→WAN
заблокирован, пока `shaterd` не применит конфиг; SSH, LuCI и панель при этом
доступны **намеренно** — цепочка вешается только на `forward`. Чем защита
вооружается, когда отказывается вооружаться и как её снять —
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) §4.
---
## Сборка из исходников
@@ -226,6 +249,28 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
(набор build-тегов, почему `shaterd` — prebuilt-пакет, порядок CI) — в
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
### Проверка
```sh
bash scripts/run-tests.sh # полный гейт
bash scripts/run-tests.sh --no-race # без -race, для локального цикла
```
Гейт гоняет весь набор **под теми же build-тегами, с которыми собирается
роутерный бинарь** (`scripts/router-tags.sh`), на Linux (с не-Linux хоста — сам
перезапускается в Docker), с `-race`, и содержит три машинные проверки против
молчаливого скипа: набор тегов может только ДОБАВЛЯТЬ тест-файлы; каждый пакет с
тестами обязан отчитаться `ok` поимённо; каждый `TestIntegration*` обязан выдать
вердикт по имени. Причина такая: до 2026-07 релизный тракт не гонял почти ничего
— 115 тест-файлов из 116 под `shater/**` в CI не исполнялись ни разу.
Отдельно `scripts/check-router-tags.sh` проверяет, что ни одна заявленная в
`FEATURES.md` фича не потеряла нужный ей build-тег.
Зелёный гейт — необходимое, но не достаточное условие: он не видит стыков с
ядром, procd и nftables. Это проверяется на стенде (см.
[`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md)).
---
## Структура репозитория
@@ -274,7 +319,10 @@ CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает вс
shater вкомпилирует **форк движка sing-box-lx** — тонкий downstream апстрима
[SagerNet/sing-box](https://github.com/SagerNet/sing-box), добавляющий набор
клиентских фич (XHTTP, AmneziaWG 2.0, MASQUE, расширения наблюдаемости) за
build-тегами и живущий **ребейзом на каждый upstream-тег, а не merge**. Форк
build-тегами и живущий **ребейзом на каждый upstream-тег, а не merge**. Это набор
самого форка, а не shater: MASQUE/CONNECT-IP мы намеренно **не регистрируем** —
`shater/generate` его не порождает, а отказ от него и остального незадействованного
зоопарка экономит ~6 МБ бинаря и столько же RAM на роутере (`shater/registry`). Форк
разрабатывается по Spec Kit; неизменяемые принципы — в
[`SPECS/CONSTITUTION.md`](SPECS/CONSTITUTION.md), справочник фич движка — в
[`docs-lx/lx-config.ru.md`](docs-lx/lx-config.ru.md).
+145
View File
@@ -0,0 +1,145 @@
// lx:begin l3-honest-drop
package adapter
import (
"net/netip"
"testing"
"github.com/sagernet/sing-tun"
"github.com/sagernet/sing-tun/gtcpip/header"
"github.com/stretchr/testify/require"
)
// judgeFlowRouter answers PreMatch with a canned verdict; JudgeFlow reads
// nothing else off the Router.
type judgeFlowRouter struct {
Router
result PreMatchResult
}
func (r *judgeFlowRouter) PreMatch(InboundContext, []byte) PreMatchResult { return r.result }
// judgeFlowPort is the tun.Port half of a FlowOutbound. inet4 is what
// PortAddresses reports for IPv4 — the one field the two ICMP consumers in
// sing-tun disagree about (see the comment on
// TestJudgeFlowICMPToBoundPortStaysAFlow).
type judgeFlowPort struct {
Outbound
inet4 netip.Addr
}
func (o *judgeFlowPort) Tag() string { return "wg-out" }
func (o *judgeFlowPort) Type() string { return "wireguard" }
func (o *judgeFlowPort) PortAddresses() (netip.Addr, netip.Addr) {
return o.inet4, netip.Addr{}
}
func (o *judgeFlowPort) PortMTU() uint32 { return 1420 }
func (o *judgeFlowPort) AttachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) DetachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) WritePackets(packets [][]byte) error { return nil }
// judgeFlowNonPort is a FlowOutbound-shaped result that is NOT a tun.Port — the
// interface drift the second line of defense in JudgeFlow exists for.
type judgeFlowNonPort struct {
Outbound
}
func (o *judgeFlowNonPort) Tag() string { return "drifted" }
func (o *judgeFlowNonPort) Type() string { return "drifted" }
func judgeFlow(t *testing.T, protocol uint8, result PreMatchResult) tun.FlowVerdict {
t.Helper()
return JudgeFlow(
&judgeFlowRouter{result: result},
"l3-in", "tun", protocol,
netip.MustParseAddrPort("192.168.1.2:1234"),
netip.MustParseAddrPort("1.1.1.1:1234"),
nil,
)
}
const (
judgeFlowICMP = uint8(header.ICMPv4ProtocolNumber)
judgeFlowTCP = uint8(header.TCPProtocolNumber)
)
// TestJudgeFlowICMPToBoundPortStaysAFlow is the guard on the ONE fix that must
// not be made here.
//
// sing-tun has two ICMP consumers with different requirements on the port:
//
// - ForwardDispatcher.createFlow (flow_dispatch.go) needs only a VALID port
// address — it NATs the echo identifier and rewrites the source to that
// address. This is the path every unfragmented LAN ping takes, and it is
// what makes ping-through-WireGuard/AWG work at all.
// - ICMPForwarder.installFlow (stack_gvisor_icmp.go) additionally requires the
// address to be UNSPECIFIED, because it writes the packet to the port
// unmodified. A WireGuard endpoint reports its concrete interface address
// (transport/wireguard/port.go), so installFlow declines and HandlePacket
// falls through to forging the echo reply.
//
// The tempting fix — "for ICMP, refuse ActionFlow when PortAddresses() is not
// unspecified, so the verdict becomes a drop and the forgery is unreachable" —
// is applied HERE, in the one function both consumers share, with byte-identical
// arguments from either. It would therefore kill the working path too: every
// ping through WireGuard/AWG, fragmented or not, would drop, and l3_tunnel would
// carry nothing but `direct`. Keep this test failing loudly if anyone tries.
func TestJudgeFlowICMPToBoundPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.MustParseAddr("10.2.0.2")}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action,
"ICMP to a WireGuard/AWG endpoint must stay a flow: the forward dispatcher NATs it by echo identifier and this is the whole point of l3_tunnel")
require.Same(t, tun.Port(port), verdict.Port)
}
// The `direct` shape: an unspecified port address. Both consumers accept it.
func TestJudgeFlowICMPToUnspecifiedPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.IPv4Unspecified()}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action)
require.Same(t, tun.Port(port), verdict.Port)
}
// PreMatchDrop is the honest verdict and must arrive as ActionDrop: it is the
// only value (besides Reject) that stops ICMPForwarder.HandlePacket before the
// Echo -> EchoReply rewrite.
func TestJudgeFlowICMPDropReachesTheStackAsDrop(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchDrop})
require.Equal(t, tun.ActionDrop, verdict.Action)
}
// The second line of defense: a PreMatchFlow whose outbound is not a tun.Port
// must not degrade ICMP to ActionAccept, because Accept is the forged reply.
func TestJudgeFlowICMPNonPortOutboundDrops(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionDrop, verdict.Action,
"FlowOutbound and tun.Port are distinct interfaces; a drift between them must not silently re-enable the echo forger")
}
func TestJudgeFlowTCPNonPortOutboundAccepts(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionAccept, verdict.Action,
"for TCP, falling back to Accept is upstream behaviour and must stay untouched")
}
// TCP keeps every mapping it had, including the Continue -> Accept default that
// is a forgery only for ICMP.
func TestJudgeFlowTCPContinueStaysAccept(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchContinue})
require.Equal(t, tun.ActionAccept, verdict.Action)
}
func TestJudgeFlowTCPBypassStaysBypass(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchBypass})
require.Equal(t, tun.ActionBypass, verdict.Action)
}
// lx:end l3-honest-drop
+11
View File
@@ -75,7 +75,18 @@ func JudgeFlow(router Router, inbound string, inboundType string, network uint8,
case PreMatchFlow:
port, isPort := result.Outbound.(tun.Port)
if !isPort {
// lx:begin l3-honest-drop
// Second line of defense behind route.(*Router).preMatchFlow: a
// PreMatchFlow result already implies the outbound is an
// adapter.FlowOutbound, but FlowOutbound and tun.Port are distinct
// interfaces, and a drift between them must not degrade ICMP to
// ActionAccept — the TUN stack would then forge the echo reply
// itself instead of admitting the tunnel cannot carry the packet.
if networkName == N.NetworkICMP {
return tun.FlowVerdict{Action: tun.ActionDrop}
}
return tun.FlowVerdict{Action: tun.ActionAccept}
// lx:end l3-honest-drop
}
verdict := tun.FlowVerdict{Action: tun.ActionFlow, Port: port, UDPTimeout: result.UDPTimeout, NewTracker: result.NewTracker}
if result.Destination.IsValid() {
+223
View File
@@ -0,0 +1,223 @@
//go:build with_quic
package httpclient
import (
"context"
stdTLS "crypto/tls"
"io"
"net"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
)
// raceProbePayload is large enough that it cannot ride along in the response
// headers: the caller has to read the body off the QUIC stream AFTER
// roundTripHTTP3Race has returned. That is the whole point of the test.
const raceProbePayload = 64 * 1024
var _ N.Dialer = (*plainDialer)(nil)
type plainDialer struct{}
func (d *plainDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
return (&net.Dialer{}).DialContext(ctx, network, destination.String())
}
func (d *plainDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
// splitDialer sends the HTTP/3 racer and the HTTP/2 racer to two different
// listeners, so a test can decide which one of them wins without having to bind
// a TCP and a UDP socket on the same port number.
type splitDialer struct {
udp M.Socksaddr
tcp M.Socksaddr
}
func (d *splitDialer) DialContext(ctx context.Context, network string, _ M.Socksaddr) (net.Conn, error) {
destination := d.tcp
if network == N.NetworkUDP {
destination = d.udp
}
return (&net.Dialer{}).DialContext(ctx, network, destination.String())
}
func (d *splitDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func startH3Server(t *testing.T, handler http.Handler) M.Socksaddr {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
listener, err := quic.ListenAddrEarly("127.0.0.1:0", &stdTLS.Config{
Certificates: []stdTLS.Certificate{*certificate},
NextProtos: []string{http3.NextProtoH3},
MinVersion: stdTLS.VersionTLS13,
}, nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: handler}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
return M.ParseSocksaddr(listener.Addr().String())
}
func newRaceProbeTransport(t *testing.T, serverAddr M.Socksaddr) (*http3FallbackTransport, string) {
return newRaceProbeTransportWithDialer(t, &plainDialer{}, serverAddr)
}
func newRaceProbeTransportWithDialer(t *testing.T, dialer N.Dialer, serverAddr M.Socksaddr) (*http3FallbackTransport, string) {
t.Helper()
baseTLSConfig, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
h2Fallback, err := newHTTP2FallbackTransport(dialer, baseTLSConfig, option.HTTP2Options{})
if err != nil {
t.Fatal(err)
}
inner, err := newHTTP3FallbackTransport(dialer, baseTLSConfig, h2Fallback, option.QUICOptions{}, 300*time.Millisecond)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { inner.Close() })
return inner.(*http3FallbackTransport), "https://" + serverAddr.String() + "/probe"
}
// TestHTTP3RaceWinnerBodyStaysReadable pins that the response handed back by the
// HTTP/3 race is a LIVE response: its body must still be readable after
// roundTripHTTP3Race returns. Cancelling the context the winner was issued on
// resets its QUIC stream, so a "successful" round trip would hand the caller a
// response it can never read.
func TestHTTP3RaceWinnerBodyStaysReadable(t *testing.T) {
payload := make([]byte, raceProbePayload)
for i := range payload {
payload[i] = byte(i)
}
serverAddr := startH3Server(t, http.HandlerFunc(func(writer http.ResponseWriter, request *http.Request) {
writer.Header().Set("Content-Type", "application/octet-stream")
writer.Write(payload)
}))
transport, url := newRaceProbeTransport(t, serverAddr)
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
request, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
t.Fatal(err)
}
// No cached HTTP/3 connection yet and a bodyless GET is replayable, so this
// takes the racing path.
response, err := transport.RoundTrip(request)
if err != nil {
t.Fatal("round trip: ", err)
}
defer response.Body.Close()
if response.ProtoMajor != 3 {
t.Fatalf("expected the HTTP/3 racer to win, got HTTP/%d.%d", response.ProtoMajor, response.ProtoMinor)
}
body, err := io.ReadAll(response.Body)
if err != nil {
t.Fatalf("the race winner's body died with the race: %v (read %d of %d bytes)", err, len(body), len(payload))
}
if len(body) != len(payload) {
t.Fatalf("short body: got %d bytes, want %d", len(body), len(payload))
}
}
// TestHTTP3RaceFallbackWinnerBodyStaysReadableAndH3LoserIsCancelled covers the
// other half of the race: the HTTP/2 fallback wins, so its body must survive the
// race, and the HTTP/3 racer that lost must be torn down instead of being left
// to run to completion on the caller's behalf.
func TestHTTP3RaceFallbackWinnerBodyStaysReadableAndH3LoserIsCancelled(t *testing.T) {
payload := make([]byte, raceProbePayload)
for i := range payload {
payload[i] = byte(i)
}
h3Started := make(chan struct{}, 1)
h3Cancelled := make(chan struct{}, 1)
// The HTTP/3 handler never answers, so the fallback wins on the timer.
h3Addr := startH3Server(t, http.HandlerFunc(func(_ http.ResponseWriter, request *http.Request) {
select {
case h3Started <- struct{}{}:
default:
}
<-request.Context().Done()
select {
case h3Cancelled <- struct{}{}:
default:
}
}))
h2Server := httptest.NewUnstartedServer(http.HandlerFunc(func(writer http.ResponseWriter, _ *http.Request) {
writer.Header().Set("Content-Type", "application/octet-stream")
writer.Write(payload)
}))
h2Server.EnableHTTP2 = true
h2Server.StartTLS()
t.Cleanup(h2Server.Close)
transport, _ := newRaceProbeTransportWithDialer(t, &splitDialer{
udp: h3Addr,
tcp: M.ParseSocksaddr(h2Server.Listener.Addr().String()),
}, h3Addr)
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
request, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://localhost:443/probe", nil)
if err != nil {
t.Fatal(err)
}
response, err := transport.RoundTrip(request)
if err != nil {
t.Fatal("round trip: ", err)
}
if response.ProtoMajor != 2 {
t.Fatalf("expected the HTTP/2 fallback to win, got HTTP/%d.%d", response.ProtoMajor, response.ProtoMinor)
}
body, err := io.ReadAll(response.Body)
if err != nil {
t.Fatalf("the fallback winner's body died with the race: %v (read %d of %d bytes)", err, len(body), len(payload))
}
response.Body.Close()
if len(body) != len(payload) {
t.Fatalf("short body: got %d bytes, want %d", len(body), len(payload))
}
select {
case <-h3Started:
case <-time.After(5 * time.Second):
t.Fatal("the HTTP/3 racer never reached the server, the test proves nothing about cancelling it")
}
select {
case <-h3Cancelled:
case <-time.After(5 * time.Second):
t.Fatal("the losing HTTP/3 request was left running after the fallback won")
}
}
+67 -19
View File
@@ -6,6 +6,7 @@ import (
"context"
stdTLS "crypto/tls"
"errors"
"io"
"net/http"
"sync"
"time"
@@ -168,32 +169,65 @@ func (t *http3FallbackTransport) roundTripHTTP3(request *http.Request) (*http.Re
return t.roundTripHTTP3Race(request, authority)
}
// cancelOnBodyClose releases a racer's context when the caller is done with the
// response it won. The race cannot release it on the way out: the body is read
// after RoundTrip returns, and the context the request was issued on is what
// keeps its stream alive.
type cancelOnBodyClose struct {
io.ReadCloser
cancel context.CancelFunc
cancelOnce sync.Once
}
func (b *cancelOnBodyClose) Close() error {
err := b.ReadCloser.Close()
b.cancelOnce.Do(b.cancel)
return err
}
func withCancelOnBodyClose(response *http.Response, cancel context.CancelFunc) *http.Response {
if response == nil || response.Body == nil {
cancel()
return response
}
response.Body = &cancelOnBodyClose{ReadCloser: response.Body, cancel: cancel}
return response
}
func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, authority string) (*http.Response, error) {
ctx, cancel := context.WithCancel(request.Context())
defer cancel()
type result struct {
response *http.Response
err error
h3 bool
}
results := make(chan result, 2)
startRoundTrip := func(request *http.Request, useH3 bool) {
request = request.WithContext(ctx)
var (
response *http.Response
err error
)
if useH3 {
response, err = t.h3Transport.RoundTrip(request)
} else {
response, err = t.h2FallbackRoundTrip(request)
}
results <- result{response: response, err: err, h3: useH3}
// Each racer runs on a context of its own. A context shared by both cannot be
// cancelled when one of them wins: quic-go and net/http reset the winner's
// stream on cancellation, so the caller would be handed a response whose body
// stops mid-read with H3_REQUEST_CANCELLED. Only losers are cancelled here;
// the winner's cancel travels with its body and fires on Close.
startRoundTrip := func(useH3 bool) context.CancelFunc {
ctx, cancel := context.WithCancel(request.Context())
raceRequest := cloneRequestForRetry(request).WithContext(ctx)
go func() {
var (
response *http.Response
err error
)
if useH3 {
response, err = t.h3Transport.RoundTrip(raceRequest)
} else {
response, err = t.h2FallbackRoundTrip(raceRequest)
}
results <- result{response: response, err: err, h3: useH3}
}()
return cancel
}
goroutines := 1
received := 0
var fallbackCancel context.CancelFunc
h3Cancel := startRoundTrip(true)
drainRemaining := func() {
cancel()
for range goroutines - received {
go func() {
loser := <-results
@@ -203,7 +237,6 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
}()
}
}
go startRoundTrip(cloneRequestForRetry(request), true)
timer := time.NewTimer(t.fallbackDelay)
defer timer.Stop()
var (
@@ -215,20 +248,28 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
case <-timer.C:
if goroutines == 1 {
goroutines++
go startRoundTrip(cloneRequestForRetry(request), false)
fallbackCancel = startRoundTrip(false)
}
case raceResult := <-results:
received++
if raceResult.err == nil {
winnerCancel := fallbackCancel
if raceResult.h3 {
t.clearH3Broken(authority)
winnerCancel = h3Cancel
if fallbackCancel != nil {
fallbackCancel()
}
} else {
h3Cancel()
}
drainRemaining()
return raceResult.response, nil
return withCancelOnBodyClose(raceResult.response, winnerCancel), nil
}
if raceResult.h3 {
t.markH3Broken(authority)
h3Err = raceResult.err
h3Cancel()
if goroutines == 1 {
goroutines++
if !timer.Stop() {
@@ -237,14 +278,21 @@ func (t *http3FallbackTransport) roundTripHTTP3Race(request *http.Request, autho
default:
}
}
go startRoundTrip(cloneRequestForRetry(request), false)
fallbackCancel = startRoundTrip(false)
}
} else {
fallbackErr = raceResult.err
if fallbackCancel != nil {
fallbackCancel()
}
}
if received < goroutines {
continue
}
h3Cancel()
if fallbackCancel != nil {
fallbackCancel()
}
drainRemaining()
switch {
case h3Err != nil && fallbackErr != nil:
+141
View File
@@ -0,0 +1,141 @@
// lx:begin health-board
package urltest
import (
"strconv"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/adapter"
)
// captureEvictions swaps the eviction notice sink for the duration of a test and
// returns a func that reads back everything reported.
func captureEvictions(t *testing.T) func() []string {
t.Helper()
var (
mu sync.Mutex
msgs []string
)
orig := boardEvictionLog
boardEvictionLog = func(m string) {
mu.Lock()
msgs = append(msgs, m)
mu.Unlock()
}
t.Cleanup(func() { boardEvictionLog = orig })
return func() []string {
mu.Lock()
defer mu.Unlock()
return append([]string(nil), msgs...)
}
}
// TestBoardHoldsAGenerationWithoutEvicting is the "what it holds" half of the
// bound. A live generation on this box is ~1200 tags (≈380 nodes plus their
// per-group egress copies and chain hops); the board must carry that — and a
// second generation's worth of overlap during a subscription rename — with no
// eviction at all, or the ceiling would be silently degrading real health data.
func TestBoardHoldsAGenerationWithoutEvicting(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
const generation = 1200
for gen := 0; gen < 2; gen++ {
for i := 0; i < generation; i++ {
s.StoreURLTestHistory("gen"+strconv.Itoa(gen)+"-node-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 20})
}
}
if got := s.Evicted(); got != 0 {
t.Fatalf("two full generations (%d tags) evicted %d entries; the board must hold them",
2*generation, got)
}
if msgs := read(); len(msgs) != 0 {
t.Fatalf("unexpected eviction notices: %v", msgs)
}
// Everything is still readable.
if s.LoadURLTestHistory("gen0-node-0") == nil {
t.Fatalf("the first tag of the first generation was lost without an eviction")
}
}
// TestBoardEvictsOldestAndSaysSo is the "what happens when it overflows" half.
// Overflow must (a) actually bound the map, (b) drop the LEAST RECENTLY MEASURED
// tags — on this box, exactly the ones no config names any more — and (c) be
// audible: a silent eviction is a health board quietly forgetting nodes it is
// still being asked about.
func TestBoardEvictsOldestAndSaysSo(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
base := time.Now().Add(-24 * time.Hour)
// Stale generation first: measured a day ago, nothing since.
const stale = 1500
for i := 0; i < stale; i++ {
s.StoreURLTestHistory("stale-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: base.Add(time.Duration(i) * time.Millisecond), Delay: 30})
}
if s.Evicted() != 0 {
t.Fatalf("evicted before the ceiling was reached")
}
// Now push past the ceiling with fresh measurements.
for i := 0; i <= maxBoardEntries; i++ {
s.StoreURLTestHistory("fresh-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 15})
}
if got := s.Evicted(); got == 0 {
t.Fatalf("board grew past %d entries without evicting anything — it is still unbounded", maxBoardEntries)
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("board holds %d entries, above the %d ceiling", size, maxBoardEntries)
}
// The day-old generation is what went, not the fresh one.
if s.LoadURLTestHistory("stale-0") != nil {
t.Fatalf("the oldest observation survived while newer ones were dropped")
}
if s.LoadURLTestHistory("fresh-"+strconv.Itoa(maxBoardEntries)) == nil {
t.Fatalf("the newest measurement was evicted")
}
msgs := read()
if len(msgs) == 0 {
t.Fatalf("entries were evicted with no notice — eviction must never be silent")
}
m := msgs[0]
for _, want := range []string{"health board full", "evicted", "re-probed"} {
if !strings.Contains(m, want) {
t.Fatalf("eviction notice %q does not say %q", m, want)
}
}
}
// TestBoardEvictionThroughMarkFailed pins the OTHER write path. MarkFailed is how
// a dead node is recorded, and a flood of dead renamed nodes is exactly the shape
// of the leak — so it has to prune too, not just the success path.
func TestBoardEvictionThroughMarkFailed(t *testing.T) {
captureEvictions(t)
s := NewHistoryStorage()
for i := 0; i <= maxBoardEntries; i++ {
s.MarkFailed("dead-" + strconv.Itoa(i))
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("MarkFailed grew the board to %d, above the %d ceiling", size, maxBoardEntries)
}
if s.Evicted() == 0 {
t.Fatalf("MarkFailed never prunes — the failure path is still unbounded")
}
}
// lx:end health-board
+118
View File
@@ -10,11 +10,128 @@
package urltest
import (
"sort"
"strconv"
"sync"
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/log"
)
// --- board capacity ---------------------------------------------------------
//
// The board is the one structure in the daemon whose key space is chosen by
// somebody else. Its keys are outbound TAGS, and on this box a tag is a node
// NAME straight out of the subscription — plus the derived per-group egress
// copies ("group-<g>-m<i>-<node>") and per-chain hop copies the probe planner
// creates for the same nodes. Providers rename their nodes freely, so a daily
// subscription refresh introduces a whole new generation of keys, while the
// store itself is pinned to the ENGINE's context (shater/engine.New) and so
// outlives every generation and every Apply — by design, so health survives a
// config change.
//
// Nothing ever removed a key. DeleteURLTestHistory exists but no shater path
// calls it (only daemon/ and clashapi/, which this fork does not run), so the
// map was strictly append-only for the life of the process — and the process is
// expected to live for months.
//
// The arithmetic: ~380 nodes, and a config with a couple of egress-bound groups
// plus a handful of chains puts a LIVE generation at roughly 380 base tags +
// 2x380 group copies + ~100 chain copies ≈ 1200 keys. One new generation per day
// is ~440k keys a year, at ~200 B per entry (map bucket + a tag string that is
// routinely 30-50 B with flag emoji, + a 56 B URLTestHistory) ≈ 88 MB of a
// 512 MB box — spent entirely on nodes that no longer exist.
const (
// maxBoardEntries is the hard ceiling. 4096 is ~3.4 live generations, so the
// board comfortably holds the current config plus the overlap while a
// subscription refresh swaps names, and still costs under a megabyte. A tighter
// bound would start evicting tags the running config actually uses; a looser one
// would stop being a bound in any useful sense.
maxBoardEntries = 4096
// keepBoardEntries is the prune target: drop a quarter at a time so the
// O(n log n) selection is amortised over ~1024 inserts instead of running on
// every probe once the board is full.
keepBoardEntries = 3072
)
// boardEvictionLog reports an eviction. A package var so tests can capture it;
// production leaves it writing to the process log, which under procd is the same
// syslog/logsink stream every other daemon line lands in.
//
// Eviction is NEVER silent. It is not free either: an evicted tag reverts to
// "untested" and its next probe re-measures it, so a board that evicts entries
// belonging to the LIVE config is a board whose ceiling is too low — and the only
// way anyone finds that out is this line.
var boardEvictionLog = func(msg string) { boardLogger().Warn(msg) }
// pruneLocked drops the least-recently-OBSERVED entries when the board exceeds
// maxBoardEntries. "Least recently observed" is max(LastOK, LastFail): the entry
// nothing has measured for the longest is, on this box, precisely a tag that no
// longer exists in any config — a renamed node, a removed group copy, a retired
// chain hop. Caller holds access.
func (s *HistoryStorage) pruneLocked() {
if len(s.delayHistory) <= maxBoardEntries {
return
}
type kv struct {
tag string
seen time.Time
}
all := make([]kv, 0, len(s.delayHistory))
for tag, h := range s.delayHistory {
seen := h.LastOK
if h.LastFail.After(seen) {
seen = h.LastFail
}
all = append(all, kv{tag, seen})
}
sort.Slice(all, func(i, j int) bool { return all[i].seen.Before(all[j].seen) })
drop := len(all) - keepBoardEntries
var oldest time.Time
for i := 0; i < drop; i++ {
if i == 0 {
oldest = all[i].seen
}
delete(s.delayHistory, all[i].tag)
}
s.evicted += uint64(drop)
msg := "urltest: health board full (" + strconv.Itoa(maxBoardEntries) + " tags) — evicted " +
strconv.Itoa(drop) + " least-recently-measured entries (" + strconv.FormatUint(s.evicted, 10) +
" total since start); they revert to untested and will be re-probed"
if !oldest.IsZero() {
msg += "; oldest observation was " + time.Since(oldest).Truncate(time.Second).String() + " ago"
}
boardEvictionLog(msg)
}
// Evicted reports how many entries the capacity bound has dropped since the store
// was created. Nonzero means the board reached maxBoardEntries at least once.
func (s *HistoryStorage) Evicted() uint64 {
if s == nil {
return 0
}
s.access.RLock()
defer s.access.RUnlock()
return s.evicted
}
// boardLogger is the process-wide fallback logger for eviction notices. The store
// is built from a plain constructor with no logger in sight (box.New, the daemon,
// shater/engine all call NewHistoryStorage()), so rather than change that
// signature everywhere the notice goes to the standard logger — which on the
// router is the daemon's own stderr, i.e. the same sink logsink owns.
var (
boardLogOnce sync.Once
boardLog log.ContextLogger
)
func boardLogger() log.ContextLogger {
boardLogOnce.Do(func() { boardLog = log.StdLogger() })
return boardLog
}
// HealthVerdict classifies a stored history entry at read time.
type HealthVerdict int
@@ -54,6 +171,7 @@ func (s *HistoryStorage) MarkFailed(tag string) {
updated.Delay = previous.Delay
}
s.delayHistory[tag] = updated
s.pruneLocked()
s.notifyUpdated()
s.access.Unlock()
}
+9
View File
@@ -21,6 +21,10 @@ type HistoryStorage struct {
access sync.RWMutex
delayHistory map[string]*adapter.URLTestHistory
updateHooks []*observable.Subscriber[struct{}]
// evicted counts entries dropped by the capacity bound (board_lx.go). The map
// is keyed by outbound tags chosen by a subscription provider, so it needs a
// ceiling; see the comment on maxBoardEntries.
evicted uint64
}
func NewHistoryStorage() *HistoryStorage {
@@ -71,6 +75,11 @@ func (s *HistoryStorage) StoreURLTestHistory(tag string, history *adapter.URLTes
}
// lx:end health-board
s.delayHistory[tag] = history
// lx:begin health-board — the map is keyed by provider-chosen tags and the
// store outlives every engine generation, so it must bound itself here: no
// shater path ever calls DeleteURLTestHistory. See maxBoardEntries.
s.pruneLocked()
// lx:end health-board
s.notifyUpdated()
s.access.Unlock()
}
+90 -3
View File
@@ -10,6 +10,7 @@ import (
"net/url"
"strconv"
"sync"
"sync/atomic"
"time"
"github.com/sagernet/sing-box/adapter"
@@ -171,6 +172,73 @@ func (t *HTTPSTransport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
return response, nil
}
// requestBuffer owns the pooled buffer that backs one DoH query.
//
// Both transports behind HTTPSTransportWrapper write the request body on a
// goroutine of their own and return from RoundTrip as soon as the response
// HEADERS arrive: net/http's write loop is still copying out of the body a
// bufferful at a time (4 KiB of write buffer, or io.Copy's 32 KiB once it hands
// the body to the connection), and http2's writeRequestBody has read only the
// first max-frame-size bytes of it. Returning the buffer to the pool at that
// point handed live memory to the next caller while the query was still going
// out — everything past that first copy left the router as whatever that caller
// had written there. A data race, and a memory-disclosure primitive aimed at the
// resolver. Measured, not reasoned: with the write parked mid-query the bytes on
// the wire diverge from the bytes we packed at exactly one copy buffer in.
//
// Ownership is counted rather than handed over once, because a retry holds two
// bodies at a time and the two transports order that differently:
// http.Transport.rewindBody CLOSES the old body before asking GetBody for a
// new one, while http2's shouldRetryRequest asks GetBody first and closes the
// old body on a goroutine. exchange keeps a count of its own until RoundTrip
// returns — the only window in which either can call GetBody — so neither
// ordering can free the buffer under the other. If a transport ever fails to
// close a body, the count never reaches zero and the buffer is simply not
// reused: garbage, not corruption.
type requestBuffer struct {
buffer *buf.Buffer
raw []byte
refs atomic.Int32
}
func newRequestBuffer(buffer *buf.Buffer, raw []byte) *requestBuffer {
holder := &requestBuffer{buffer: buffer, raw: raw}
holder.refs.Store(1)
return holder
}
// body hands out a reader over the packed query as one more owner. It refuses
// once the buffer is back in the pool, so a late caller gets an error instead
// of a reader over memory that now belongs to somebody else.
func (b *requestBuffer) body() (*pooledRequestBody, bool) {
for {
refs := b.refs.Load()
if refs < 1 {
return nil, false
}
if b.refs.CompareAndSwap(refs, refs+1) {
return &pooledRequestBody{Reader: bytes.NewReader(b.raw), owner: b}, true
}
}
}
func (b *requestBuffer) release() {
if b.refs.Add(-1) == 0 {
b.buffer.Release()
}
}
type pooledRequestBody struct {
*bytes.Reader
owner *requestBuffer
closeOne sync.Once
}
func (b *pooledRequestBody) Close() error {
b.closeOne.Do(b.owner.release)
return nil
}
func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg, error) {
exMessage := *message
exMessage.Id = 0
@@ -181,11 +249,31 @@ func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
requestBuffer.Release()
return nil, err
}
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), bytes.NewReader(rawMessage))
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
// Drops the count exchange holds once RoundTrip is done with the request;
// the bodies handed to the transport keep their own until it closes them.
defer queryBuffer.release()
requestBody, _ := queryBuffer.body() // cannot fail: the count above is ours
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), requestBody)
if err != nil {
requestBuffer.Release()
requestBody.Close()
return nil, err
}
// http.NewRequestWithContext infers both only for the body types it knows,
// and pooledRequestBody is not one of them. Upstream got them for free from
// *bytes.Reader; GetBody is what lets a POST be replayed when a pooled
// connection turns out to have been closed under us. Being unknown to
// net/http also costs one packet on the HTTP/1.1 leg: isKnownInMemoryReader
// no longer recognises the body, so the request headers are flushed before
// the query instead of travelling with it.
request.ContentLength = int64(len(rawMessage))
request.GetBody = func() (io.ReadCloser, error) {
retryBody, ok := queryBuffer.body()
if !ok {
return nil, E.New("DoH request buffer already released")
}
return retryBody, nil
}
request.Header = t.headers.Clone()
request.Header.Set("Content-Type", MimeType)
request.Header.Set("Accept", MimeType)
@@ -193,7 +281,6 @@ func (t *HTTPSTransport) exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
currentTransport := t.transport
t.transportAccess.Unlock()
response, err := currentTransport.RoundTrip(request)
requestBuffer.Release()
if err != nil {
return nil, err
}
@@ -0,0 +1,541 @@
package transport
import (
"bytes"
"context"
"errors"
"io"
"net"
"net/http"
"net/http/httptest"
"net/url"
"os"
"strconv"
"sync"
"sync/atomic"
"testing"
"time"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
mDNS "github.com/miekg/dns"
"golang.org/x/net/http2"
)
// The request body of a DoH query is backed by a POOLED buffer. Neither
// transport behind HTTPSTransportWrapper is done with that body when RoundTrip
// returns: net/http hands the request to a write loop of its own and returns as
// soon as the response HEADERS have been read, and golang.org/x/net/http2 writes
// the body on the goroutine that runs writeRequest while roundTrip waits on
// respHeaderRecv. Returning the buffer to the pool at that point hands live
// memory to the next caller while the query is still being written to the wire,
// and what goes out is whatever that next caller put there.
//
// Both tests below force a window that is normally microseconds wide to stay
// open, and drain the pool while it is open:
//
// - HTTP/1.1: the client connection stops accepting writes past the request
// headers, so net/http's write loop is parked having copied only the first
// io.Copy buffer (32 KiB) of the query.
// - HTTP/2: the server pins a 1 KiB stream receive window and does not read
// the body, so writeRequestBody is parked in awaitFlowControl having copied
// only the first max-frame-size bytes of the query.
//
// In both, the server sends the response HEADERS first and withholds the
// response BODY until the pool has been drained, so Exchange has returned from
// RoundTrip — and released the buffer, on the broken build — while the query is
// still going out.
//
// Both queries are padded past the transport's copy buffer on purpose. Below it
// the transport lifts the whole query out of the pooled buffer in a single Read
// that RACES the release rather than provably following it, and a test built on
// that race would be a coin toss. The ownership defect is the same at every
// size; only its deterministic proof needs the padding.
const (
// Past io.Copy's 32 KiB buffer, which is the granularity net/http moves a
// request body at (persistConnWriter.ReadFrom -> io.Copy), and still inside
// buf.MaxPooledBufferSize so the buffer really comes from the pool.
httpsH1PaddedQuerySize = 40000
// Past http2's max frame size, which is how much of the body
// writeRequestBody lifts into its scratch buffer per round.
httpsH2PaddedQuerySize = 20000
// Pinned on the HTTP/2 server so the client cannot write the whole body
// before the response headers come back.
httpsPinnedStreamWindow = 1024
// Pinned too: Go's HTTP/2 server advertises a 1 MiB max frame size by
// default, and the client sizes its body-copy buffer from that — with the
// default it would slurp a 20 KB query in one Read and the divergence would
// be hidden by the copy size rather than absent. 16384 is the protocol
// minimum and what real resolvers advertise.
httpsPinnedMaxFrameSize = 16384
// How many times the HTTP/2 scenario is repeated; see the test.
httpsH2Rounds = 8
// How long to wait after the response headers before draining the pool, so
// that Exchange has certainly returned from RoundTrip.
httpsReleaseSettleDelay = 200 * time.Millisecond
)
// httpsPaddedQuery returns a query and the exact bytes HTTPSTransport.exchange
// packs for it.
func httpsPaddedQuery(t *testing.T, padding int) (*mDNS.Msg, []byte) {
t.Helper()
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: make([]byte, padding)})
message.Extra = append(message.Extra, opt)
onWire := *message
onWire.Id = 0
onWire.Compress = true
expected, err := onWire.Pack()
if err != nil {
t.Fatal(err)
}
return message, expected
}
func httpsTestReply(t *testing.T) []byte {
t.Helper()
query := new(mDNS.Msg)
query.SetQuestion("example.com.", mDNS.TypeA)
response := new(mDNS.Msg)
response.SetReply(query)
raw, err := response.Pack()
if err != nil {
t.Fatal(err)
}
return raw
}
// httpsPoisonPool takes buffers of one size class out of the pool and fills them
// with a pattern no DNS message contains. They are returned, not released: the
// caller holds them so nothing can hand them back while the check runs.
func httpsPoisonPool(size int, count int) []*buf.Buffer {
poison := make([]*buf.Buffer, 0, count)
for range count {
buffer := buf.NewSize(size)
poison = append(poison, buffer)
free := buffer.FreeBytes()
for i := range free {
free[i] = 0xEE
}
}
return poison
}
func httpsReleaseAll(buffers []*buf.Buffer) {
for _, buffer := range buffers {
buffer.Release()
}
}
// httpsRequirePoisonReachesReleasedBuffer is the CONTROL for the tests below. A
// clean result there means nothing unless this instrument is shown to be able to
// produce a dirty one: it must be true that a buffer released while its bytes
// are still referenced comes back out of the pool and gets overwritten. If that
// stops holding — a different allocator, a pool that zeroes, a size class that
// is not pooled at all — the tests below would go green on broken code.
//
// Retried, because under -race sync.Pool.Put drops one object in four on
// purpose. That same dice roll is why the checks below are 3-in-4 detectors
// under -race and certainties without it; it can only make a broken build look
// clean, never a clean build look broken.
func httpsRequirePoisonReachesReleasedBuffer(t *testing.T, size int, pattern []byte) {
t.Helper()
for range 32 {
control := buf.NewSize(size)
free := control.FreeBytes()
if len(free) < len(pattern) {
t.Fatalf("control failed: a %d-byte buffer came back %d bytes long", size, len(free))
}
copy(free, pattern)
alias := free[:len(pattern)]
control.Release()
held := httpsPoisonPool(size, 8)
poisoned := !bytes.Equal(alias, pattern)
httpsReleaseAll(held)
if poisoned {
return
}
}
t.Fatal("control failed: poisoning the pool never touched a released buffer, so a clean result below would prove nothing")
}
// httpsTestDialer hands HTTPSTransportWrapper a connection to a local test
// server, optionally wrapped.
type httpsTestDialer struct {
target string
wrap func(net.Conn) net.Conn
access sync.Mutex
conns []net.Conn
}
func (d *httpsTestDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, "tcp", d.target)
if err != nil {
return nil, err
}
var wrapped net.Conn = conn
if d.wrap != nil {
wrapped = d.wrap(conn)
}
d.access.Lock()
d.conns = append(d.conns, conn)
d.access.Unlock()
return wrapped, nil
}
func (d *httpsTestDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, os.ErrInvalid
}
func (d *httpsTestDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
// httpsGatedConn stops accepting writes once limit bytes have gone out, until
// the gate is opened. HTTP/1.1 has no flow-control knob to park the writer with,
// so the connection provides one.
type httpsGatedConn struct {
net.Conn
limit int64
written atomic.Int64
gate chan struct{}
}
func (c *httpsGatedConn) Write(p []byte) (int, error) {
if c.written.Load()+int64(len(p)) > c.limit {
select {
case <-c.gate:
case <-time.After(30 * time.Second):
return 0, errors.New("gated conn: nobody opened the gate")
}
}
n, err := c.Conn.Write(p)
c.written.Add(int64(n))
return n, err
}
// httpsSlowServer is the handler both tests share: response HEADERS first, then
// nothing until the pool has been drained, then the request body, then the
// response body.
type httpsSlowServer struct {
reply []byte
served atomic.Int32
warmups int32
headersSent chan struct{}
bodyGate chan struct{}
received chan []byte
readErr chan error
}
func newHTTPSSlowServer(reply []byte) *httpsSlowServer {
return &httpsSlowServer{
reply: reply,
headersSent: make(chan struct{}, 1),
bodyGate: make(chan struct{}),
received: make(chan []byte, 1),
readErr: make(chan error, 1),
}
}
func (s *httpsSlowServer) ServeHTTP(writer http.ResponseWriter, request *http.Request) {
if s.served.Add(1) <= s.warmups {
// Warm-up: answer normally, so the connection is established and the
// client has applied the server's SETTINGS before the query that
// matters goes out.
io.Copy(io.Discard, request.Body)
writer.Header().Set("Content-Type", MimeType)
writer.Header().Set("Content-Length", strconv.Itoa(len(s.reply)))
writer.Write(s.reply)
return
}
// Without this, net/http's HTTP/1.1 server drains up to 256 KB of the
// request body before it will write response headers, precisely so that a
// half-duplex client cannot deadlock. That would consume the query before
// the client is anywhere near done sending it, and there would be nothing
// left in flight to catch. Full duplex is how a resolver that answers from
// cache before reading the whole query behaves; HTTP/2 is full duplex
// already and returns an error here, which is fine.
http.NewResponseController(writer).EnableFullDuplex()
writer.Header().Set("Content-Type", MimeType)
// Content-Length matters: without it Exchange falls into io.ReadAll and
// waits for the end of the response, which this handler is about to
// withhold on purpose.
writer.Header().Set("Content-Length", strconv.Itoa(len(s.reply)))
writer.WriteHeader(http.StatusOK)
writer.(http.Flusher).Flush()
s.headersSent <- struct{}{}
// A real resolver would be reading the query by now. Withholding it is what
// keeps the client parked mid-body while the pool is drained.
<-s.bodyGate
body, err := io.ReadAll(request.Body)
s.readErr <- err
s.received <- body
writer.Write(s.reply)
}
// drainPoolOnceHeadersAreOut waits for the response headers, gives Exchange time
// to return from RoundTrip, drains the size class the query buffer came from —
// on this goroutine, so a buffer released on the way out lands in our hands and
// not somewhere harmless — and only then lets the server read the query.
func (s *httpsSlowServer) drainPoolOnceHeadersAreOut(bufferSize int) <-chan []*buf.Buffer {
poisoned := make(chan []*buf.Buffer, 1)
go func() {
<-s.headersSent
time.Sleep(httpsReleaseSettleDelay)
poisoned <- httpsPoisonPool(bufferSize, 32)
close(s.bodyGate)
}()
return poisoned
}
func (s *httpsSlowServer) requireQueryOnWire(t *testing.T, expected []byte) {
t.Helper()
var sent []byte
select {
case sent = <-s.received:
case <-time.After(30 * time.Second):
t.Fatal("the server never received the request body")
}
if err := <-s.readErr; err != nil {
t.Fatal("reading the request body: ", err)
}
if bytes.Equal(sent, expected) {
return
}
firstDiff := -1
for i := 0; i < len(sent) && i < len(expected); i++ {
if sent[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the query on the wire is not the query we packed: %d of %d bytes received, first difference at offset %d — "+
"the pooled request buffer was reused while the transport was still reading it", len(sent), len(expected), firstDiff)
}
// TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP1 proves that the query an
// HTTP/1.1 resolver receives is the query we asked to send, even when the pool
// is drained the instant the response headers arrive.
func TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP1(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH1PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
handler := newHTTPSSlowServer(httpsTestReply(t))
server := httptest.NewServer(handler)
t.Cleanup(server.Close)
dialer := &httpsTestDialer{
target: server.Listener.Addr().String(),
wrap: func(conn net.Conn) net.Conn {
// One 4 KiB flush of net/http's write buffer gets through, which is
// what carries the request headers to the server, and the write loop
// parks on the next one — still holding the query.
return &httpsGatedConn{Conn: conn, limit: 4096, gate: handler.bodyGate}
},
}
t.Cleanup(dialer.closeAll)
// Scheme http puts HTTPSTransportWrapper on its HTTP/1.1 leg, the one it
// also falls back to whenever a resolver does not negotiate h2.
destination := &url.URL{Scheme: "http", Host: "doh.invalid", Path: "/dns-query"}
dnsTransport := &HTTPSTransport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTPS, "test-doh-h1", nil),
logger: logger.NOP(),
dialer: dialer,
destination: destination,
headers: http.Header{},
transport: NewHTTPSTransportWrapper(dialer, M.ParseSocksaddr(server.Listener.Addr().String()), destination),
}
t.Cleanup(func() { dnsTransport.Close() })
poisoned := handler.drainPoolOnceHeadersAreOut(bufferSize)
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
defer httpsReleaseAll(<-poisoned)
handler.requireQueryOnWire(t, expected)
}
// TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP2 does the same over h2,
// the leg every resolver that speaks HTTP/2 lands on.
func TestHTTPSExchangeRequestBufferOutlivesRoundTripHTTP2(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH2PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
// Repeated because a buffer released on the goroutine running Exchange
// usually lands in that P's private sync.Pool slot, which the goroutine
// draining the pool cannot steal: one round catches a broken build about
// half the time, eight catch it better than 99 times in 100. Every round
// must come back clean.
for round := range httpsH2Rounds {
if !t.Run(strconv.Itoa(round), func(t *testing.T) {
httpsH2Round(t, message, expected, bufferSize)
}) {
return
}
}
}
func httpsH2Round(t *testing.T, message *mDNS.Msg, expected []byte, bufferSize int) {
handler := newHTTPSSlowServer(httpsTestReply(t))
// x/net/http2 may put the first request on the wire before it has applied
// the server's SETTINGS, and would then overrun the 1 KiB window this test
// pins and be reset with FLOW_CONTROL_ERROR. One small query first settles
// that: reading its response proves the SETTINGS frame ahead of it was
// processed.
handler.warmups = 1
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { listener.Close() })
h2server := &http2.Server{
MaxUploadBufferPerStream: httpsPinnedStreamWindow,
MaxReadFrameSize: httpsPinnedMaxFrameSize,
}
go func() {
for {
conn, acceptErr := listener.Accept()
if acceptErr != nil {
return
}
go h2server.ServeConn(conn, &http2.ServeConnOpts{Handler: handler})
}
}()
dialer := &httpsTestDialer{target: listener.Addr().String()}
t.Cleanup(dialer.closeAll)
// Scheme https keeps HTTPSTransportWrapper on its h2 leg. The dialer hands
// back a plain connection, which x/net/http2 speaks prior-knowledge h2 over;
// TLS adds nothing this test is about.
destination := &url.URL{Scheme: "https", Host: "doh.invalid", Path: "/dns-query"}
dnsTransport := &HTTPSTransport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTPS, "test-doh-h2", nil),
logger: logger.NOP(),
dialer: dialer,
destination: destination,
headers: http.Header{},
transport: NewHTTPSTransportWrapper(dialer, M.ParseSocksaddr(listener.Addr().String()), destination),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
warmup := new(mDNS.Msg)
warmup.SetQuestion("warmup.invalid.", mDNS.TypeA)
if _, err = dnsTransport.Exchange(ctx, warmup); err != nil {
t.Fatal("warm-up exchange: ", err)
}
poisoned := handler.drainPoolOnceHeadersAreOut(bufferSize)
if _, err = dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
defer httpsReleaseAll(<-poisoned)
handler.requireQueryOnWire(t, expected)
}
// TestHTTPSRequestBufferSurvivesRewind covers the second owner a retry creates.
// net/http rewinds a dead connection's request by CLOSING the body it has and
// then asking GetBody for another one (rewindBody), while x/net/http2 asks
// GetBody first and closes the old body on a goroutine (shouldRetryRequest,
// closeReqBodyLocked). Either ordering frees the buffer under the retry if the
// first Close is what returns it to the pool, and the retry then sends whatever
// the next pool user wrote — the same disclosure, one attempt later.
func TestHTTPSRequestBufferSurvivesRewind(t *testing.T) {
message, expected := httpsPaddedQuery(t, httpsH2PaddedQuerySize)
bufferSize := 1 + message.Len()
httpsRequirePoisonReachesReleasedBuffer(t, bufferSize, expected)
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
requestBuffer := buf.NewSize(bufferSize)
rawMessage, err := exMessage.PackBuffer(requestBuffer.FreeBytes())
if err != nil {
t.Fatal(err)
}
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
defer queryBuffer.release()
first, ok := queryBuffer.body()
if !ok {
t.Fatal("the first body was refused while exchange still holds the buffer")
}
// The transport got some of the query out before the connection turned out
// to be dead, then closed the body.
if _, err = io.CopyN(io.Discard, first, 128); err != nil {
t.Fatal(err)
}
first.Close()
// GetBody, as the retry would call it.
second, ok := queryBuffer.body()
if !ok {
t.Fatal("GetBody was refused after the first body was closed: the retry has no query left to send")
}
poison := httpsPoisonPool(bufferSize, 32)
defer httpsReleaseAll(poison)
retried, err := io.ReadAll(second)
if err != nil {
t.Fatal(err)
}
if !bytes.Equal(retried, expected) {
firstDiff := -1
for i := 0; i < len(retried) && i < len(expected); i++ {
if retried[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the retried query is not the query we packed: %d of %d bytes, first difference at offset %d — "+
"closing the first body returned the buffer to the pool while the retry still needed it", len(retried), len(expected), firstDiff)
}
second.Close()
}
// TestHTTPSRequestBufferRefusesBodyAfterRelease pins the recoverable end of the
// contract: once the buffer really is back in the pool, GetBody must hand out an
// error rather than a reader over memory that now belongs to somebody else.
func TestHTTPSRequestBufferRefusesBodyAfterRelease(t *testing.T) {
requestBuffer := buf.NewSize(64)
rawMessage := requestBuffer.FreeBytes()[:8]
queryBuffer := newRequestBuffer(requestBuffer, rawMessage)
body, ok := queryBuffer.body()
if !ok {
t.Fatal("the first body was refused while the caller still holds the buffer")
}
body.Close()
body.Close() // http3 and net/http both manage to close a body twice
queryBuffer.release()
if _, ok = queryBuffer.body(); ok {
t.Fatal("a body was handed out over a buffer that is already back in the pool")
}
}
+29 -5
View File
@@ -126,6 +126,12 @@ func (t *HTTP3Transport) newTransport() *http3.Transport {
conn.Close()
return nil, dialErr
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-quicConn.Context().Done()
conn.Close()
}()
return quicConn, nil
},
TLSClientConfig: t.tlsConfig,
@@ -156,15 +162,34 @@ func (t *HTTP3Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
requestBuffer := buf.NewSize(1 + message.Len())
rawMessage, err := exMessage.PackBuffer(requestBuffer.FreeBytes())
// NOT a pooled buffer, deliberately — the request body must own memory this
// transport can never hand back.
//
// quic-go writes the request body on a goroutine of its own (http3's
// doRequest spawns it and goes on to block in ReadResponse), and NOTHING ever
// joins that goroutine. On the success path sendRequestBody closes the body
// when it is finished, but on every error path RoundTripOpt closes it as soon
// as doRequest returns — and doRequest waits only on the request-cancellation
// watchdog, not on the writer. So there is no moment at which this code can
// know the body is no longer being read, and therefore no moment at which it
// may return a pooled buffer. Releasing on Close looks like an ownership
// handoff and is not one.
//
// Owning it costs nothing here, measured rather than assumed: for a typical
// query (a 36-byte name, A record) Pack is 87 ns/op at 64 B and 1 alloc,
// against 108 ns/op at 64 B and 1 alloc for packing into a pooled buffer. The
// pool never avoided an allocation on this path — buf.NewSize allocates the
// Buffer struct itself, the same 64 bytes the message needs — it only added
// Get/Put on top. This path is hot in queries, not in bytes.
//
// The response buffer below stays pooled: it is read to completion and
// unpacked before Exchange returns, and nothing outlives it.
rawMessage, err := exMessage.Pack()
if err != nil {
requestBuffer.Release()
return nil, err
}
request, err := http.NewRequestWithContext(ctx, http.MethodPost, t.destination.String(), bytes.NewReader(rawMessage))
if err != nil {
requestBuffer.Release()
return nil, err
}
request.Header = t.headers.Clone()
@@ -174,7 +199,6 @@ func (t *HTTP3Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS
currentTransport := t.transport
t.transportAccess.Unlock()
response, err := currentTransport.RoundTrip(request)
requestBuffer.Release()
if err != nil {
return nil, err
}
@@ -0,0 +1,426 @@
package quic
import (
"bytes"
"context"
"crypto/rand"
"crypto/tls"
"io"
"net"
"net/http"
"net/url"
"strconv"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
mDNS "github.com/miekg/dns"
)
// The request body of a DoH3 query used to be backed by a POOLED buffer. quic-go
// sends that body on a goroutine of its own which outlives RoundTrip (http3's
// doRequest spawns it and returns as soon as the response HEADERS arrive), and
// NOTHING joins that goroutine, so there is no moment at which the transport may
// hand the buffer back.
//
// Two tests, because the two paths are observable in different ways.
//
// - On the SUCCESS path the body keeps flowing, so the damage is visible on the
// wire: TestHTTP3ExchangeRequestBufferOutlivesRoundTrip pins a 2 KB server
// stream window and answers before reading the body, so the client is still
// writing when Exchange returns, and compares what the server received.
//
// - On the FAILURE and CANCELLATION paths the damage is not visible on the wire
// at all: every ReadResponse error in quic-go calls str.CancelWrite BEFORE
// RoundTripOpt closes the body, so whatever the writer reads afterwards is
// thrown at a dead stream. What is left is a read of memory that belongs to
// somebody else. TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory therefore
// pins the CAUSE instead of the symptom: the bytes of a query must never end
// up in a buffer this transport can return to the pool.
const (
// Big enough to need more than one 8 KiB read out of the request body
// (http3's bodyCopyBufferSize), small enough to still come from the pool
// (buf.MaxPooledBufferSize).
paddedQuerySize = 20000
// Pinned on the server so the client cannot write the whole body before the
// response comes back.
pinnedStreamWindow = 2048
// Padding for the marked query of the ownership test. Only has to be
// distinctive and pooled, not large.
markedQueryPadding = 4096
markedQueryNeedle = 64
// How deep to drain a size class when looking for the needle.
poolScanDepth = 64
// How many times a CONTROL may repeat before it gives up.
//
// Both controls in this file assert the same thing — a buffer released while
// its bytes are still referenced comes back out of the pool — and under
// `-race` that is a DICE ROLL, not a certainty: sync.Pool.Put drops one
// object in four on purpose (runtime_randn(4) == 0, sync/pool.go). Measured
// in golang:1.26 with `go test -race -count=60`: the single-attempt control
// failed 18 times out of 60, i.e. the gate's -race pass had a ~30% chance of
// going red on a tree with nothing wrong with it.
//
// A retry is the honest repair rather than a papering-over, because the
// control's claim is EXISTENTIAL — "this instrument is able to find a
// released, still-referenced buffer" — and one success proves it. It is not
// an average over attempts, so nothing is diluted by taking more than one.
// 32 attempts leave a (1/4)^32 chance of a false alarm.
//
// What this does NOT do, said plainly: it does not make the VERDICT below
// certain under -race. The same 1-in-4 drop means a scan that comes back
// clean has a 1-in-4 chance of being clean because the pool threw the
// evidence away. That direction is the safe one — it can only let a broken
// build look clean, never make a clean build look broken — and the -race
// pass is not the only one that runs this test: [2/7] of scripts/run-tests.sh
// runs the same file WITHOUT -race, where both the control and the verdict
// are certainties.
controlAttempts = 32
)
func paddedQuery(t *testing.T) (*mDNS.Msg, []byte) {
t.Helper()
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: make([]byte, paddedQuerySize)})
message.Extra = append(message.Extra, opt)
// Exactly what HTTP3Transport.Exchange puts on the wire.
onWire := *message
onWire.Id = 0
onWire.Compress = true
expected, err := onWire.Pack()
if err != nil {
t.Fatal(err)
}
return message, expected
}
// poisonPool takes buffers of one size class out of the pool and fills them with
// a pattern no DNS message contains. The buffers are returned, not released: the
// caller holds them so nothing can hand them back while the check runs.
func poisonPool(size int, count int) []*buf.Buffer {
poison := make([]*buf.Buffer, 0, count)
for range count {
buffer := buf.NewSize(size)
poison = append(poison, buffer)
free := buffer.FreeBytes()
for i := range free {
free[i] = 0xEE
}
}
return poison
}
func releaseAll(buffers []*buf.Buffer) {
for _, buffer := range buffers {
buffer.Release()
}
}
// requirePoisonReachesReleasedBuffer is the CONTROL for the test below. A clean
// result there means nothing unless this instrument is shown to be able to
// produce a dirty one: it must be true that a buffer released while its bytes
// are still referenced comes back out of the pool and gets overwritten. If this
// stops holding — a different allocator, a pool that zeroes, a size class that
// is not pooled at all — the test below would go green on broken code.
//
// Retried, because under -race sync.Pool.Put drops one object in four on
// purpose. That same dice roll is why the check below is a 3-in-4 detector under
// -race and a certainty without it; it can only make a broken build look clean,
// never a clean build look broken. See controlAttempts.
func requirePoisonReachesReleasedBuffer(t *testing.T, size int, pattern []byte) {
t.Helper()
for range controlAttempts {
control := buf.NewSize(size)
free := control.FreeBytes()
if len(free) < len(pattern) {
t.Fatalf("control failed: a %d-byte buffer came back %d bytes long", size, len(free))
}
copy(free, pattern)
alias := free[:len(pattern)]
control.Release()
held := poisonPool(size, 8)
poisoned := !bytes.Equal(alias, pattern)
releaseAll(held)
if poisoned {
return
}
}
t.Fatal("control failed: poisoning the pool never touched a released buffer, so a clean result below would prove nothing")
}
// TestHTTP3ExchangeRequestBufferOutlivesRoundTrip proves that the query the
// server receives is the query we asked to send, even when the pool is drained
// the instant Exchange returns.
func TestHTTP3ExchangeRequestBufferOutlivesRoundTrip(t *testing.T) {
message, expected := paddedQuery(t)
bufferSize := 1 + message.Len()
requirePoisonReachesReleasedBuffer(t, bufferSize, expected)
drainGate := make(chan struct{})
received := make(chan []byte, 1)
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
// Answer BEFORE reading the request body. A real resolver would not, but
// any peer, middlebox or loss pattern that delays the body has the same
// effect, and this makes the window deterministic.
response := new(mDNS.Msg)
response.SetReply(testQuery())
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
// Content-Length matters here: without it Exchange falls into io.ReadAll
// and waits for the stream FIN, which this handler is about to withhold.
writer.Header().Set("Content-Length", strconv.Itoa(len(rawResponse)))
writer.Write(rawResponse)
writer.(http.Flusher).Flush()
<-drainGate
body, _ := io.ReadAll(request.Body)
received <- body
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), &quic.Config{
InitialStreamReceiveWindow: pinnedStreamWindow,
MaxStreamReceiveWindow: pinnedStreamWindow,
InitialConnectionReceiveWindow: 1 << 16,
MaxConnectionReceiveWindow: 1 << 16,
})
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3-buffer", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: M.ParseSocksaddr(listener.Addr().String()),
tlsConfig: &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
},
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, message); err != nil {
t.Fatal("exchange: ", err)
}
// Exchange has returned, the body is still in flight. Drain the size class it
// came from, on this very goroutine, so a buffer released on the way out lands
// in our hands and not somewhere harmless. The buffers are held until after
// the comparison below.
poison := poisonPool(bufferSize, 32)
defer releaseAll(poison)
close(drainGate)
var sent []byte
select {
case sent = <-received:
case <-time.After(20 * time.Second):
t.Fatal("the server never received the request body")
}
if !bytes.Equal(sent, expected) {
firstDiff := -1
for i := 0; i < len(sent) && i < len(expected); i++ {
if sent[i] != expected[i] {
firstDiff = i
break
}
}
t.Fatalf("the query on the wire is not the query we packed: %d of %d bytes received, first difference at offset %d — "+
"the pooled request buffer was reused while quic-go was still reading it", len(sent), len(expected), firstDiff)
}
}
// markedQuery builds a query whose EDNS0 padding carries a random tag, so the
// packed bytes contain a needle that can be searched for in pool memory and
// cannot collide with anything else.
func markedQuery(t *testing.T) (*mDNS.Msg, []byte) {
t.Helper()
padding := make([]byte, markedQueryPadding)
if _, err := rand.Read(padding); err != nil {
t.Fatal(err)
}
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
opt := new(mDNS.OPT)
opt.Hdr.Name = "."
opt.Hdr.Rrtype = mDNS.TypeOPT
opt.Option = append(opt.Option, &mDNS.EDNS0_PADDING{Padding: padding})
message.Extra = append(message.Extra, opt)
return message, padding[:markedQueryNeedle]
}
// poolHoldsNeedle drains one size class of the buffer pool and reports whether
// any buffer in it still carries the needle. It must run on the goroutine that
// released the buffer: sync.Pool keeps a per-P private slot that no other P can
// steal from, and on the paths this test covers the release happens inline in
// RoundTripOpt, on the caller's own goroutine.
func poolHoldsNeedle(size int, needle []byte, count int) bool {
held := make([]*buf.Buffer, 0, count)
defer func() { releaseAll(held) }()
var found bool
for range count {
buffer := buf.NewSize(size)
held = append(held, buffer)
if bytes.Contains(buffer.FreeBytes(), needle) {
found = true
}
}
return found
}
// requireInstrumentFindsPackedQuery is the CONTROL. It does exactly what the old
// Exchange did — pack a query into a pooled buffer and release it — and demands
// that the scan below FINDS the needle. Without it, "the pool does not hold the
// query" would also be the verdict for a scan that can never find anything.
//
// Retried for the same reason its sibling control above is, and it was NOT
// before: under -race sync.Pool.Put drops one object in four, so a single
// attempt made this control — and with it the whole -race pass of the gate —
// fail on 18 of 60 measured runs with nothing wrong in the tree. A fresh
// needle is packed on each attempt, so a later one cannot be answered by an
// earlier one's bytes. See controlAttempts for what the retry does and does not
// buy.
func requireInstrumentFindsPackedQuery(t *testing.T) {
t.Helper()
for range controlAttempts {
message, needle := markedQuery(t)
size := 1 + message.Len()
exMessage := *message
exMessage.Id = 0
exMessage.Compress = true
buffer := buf.NewSize(size)
if _, err := exMessage.PackBuffer(buffer.FreeBytes()); err != nil {
t.Fatal(err)
}
buffer.Release()
if poolHoldsNeedle(size, needle, poolScanDepth) {
return
}
}
t.Fatalf("control failed: %d times in a row, a query packed into a pooled buffer and released was NOT "+
"found by the scan, so a clean verdict below would prove nothing. Under -race sync.Pool.Put drops "+
"one object in four, which is what the retries absorb; this many consecutive misses is something "+
"else — a pool that zeroes on Put, a size class that stopped being pooled, or buf.Buffer no longer "+
"handing its array back at all", controlAttempts)
}
// TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory pins the ownership rule the
// failure paths depend on.
//
// quic-go's http3.Transport closes the request body on every error path
// (transport.go RoundTripOpt) the moment doRequest returns, and doRequest waits
// only on the request-cancellation watchdog — never on the goroutine writing the
// body. So releasing the buffer when the body is closed is not an ownership
// handoff, and the only safe arrangement is for the query never to live in pool
// memory at all.
//
// This test encodes THAT design. A future guarded-pool design (a lock around
// Read and Close, refusing reads after release) would also be correct and would
// fail this test on purpose — it would have to replace it, and say so.
func TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory(t *testing.T) {
requireInstrumentFindsPackedQuery(t)
// A UDP socket nobody answers on: the handshake runs to the context deadline
// instead of being refused, which is the shape a router sees when the tunnel
// carrying its resolver drops.
blackhole, err := net.ListenUDP("udp", &net.UDPAddr{IP: net.IPv4(127, 0, 0, 1)})
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { blackhole.Close() })
for _, testCase := range []struct {
name string
ctx func(t *testing.T) (context.Context, context.CancelFunc)
}{
{
// RoundTripOpt closes the body after the handshake gives up.
name: "server never answers",
ctx: func(t *testing.T) (context.Context, context.CancelFunc) {
return context.WithTimeout(context.Background(), 500*time.Millisecond)
},
},
{
// The cancellation watchdog fires, then RoundTripOpt closes the body.
name: "context already cancelled",
ctx: func(t *testing.T) (context.Context, context.CancelFunc) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
return ctx, func() {}
},
},
} {
t.Run(testCase.name, func(t *testing.T) {
message, needle := markedQuery(t)
size := 1 + message.Len()
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3-ownership", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: M.ParseSocksaddr(blackhole.LocalAddr().String()),
tlsConfig: &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
},
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := testCase.ctx(t)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, message); err == nil {
t.Fatal("expected the exchange to fail; this test is about the failure path")
}
// Same goroutine that ran RoundTripOpt, so the per-P private slot a
// release would have landed in is the one being drained.
if poolHoldsNeedle(size, needle, poolScanDepth) {
t.Fatal("the bytes of the query came back out of the buffer pool: the request body was packed into pooled " +
"memory and released while quic-go's body writer could still be reading it")
}
})
}
}
+351
View File
@@ -0,0 +1,351 @@
package quic
import (
"context"
"crypto/tls"
"net"
"net/http"
"net/url"
"sync"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
mDNS "github.com/miekg/dns"
)
var _ N.Dialer = (*trackingDialer)(nil)
// These tests pin down who owns the UDP socket handed to quic-go.
//
// quic-go's Dial/DialEarly take a net.PacketConn but do NOT take ownership of
// it: quic.setupTransport() builds a Transport with createdConn=false, and
// Transport.Close() then only calls conn.SetReadDeadline(time.Now()) instead of
// conn.Close(). So every QUIC connection torn down here — idle timeout, a
// retryable error, an engine reload calling Reset() — used to strand the UDP
// socket that carried it for the rest of the process's life. On a router that
// resolves through DoQ/DoH3 for months that is an unbounded fd leak.
//
// Both tests reconnect once and assert the socket from the FIRST connection is
// actually closed. Without the `<-conn.Context().Done() -> rawConn.Close()`
// watchdogs in quic.go / http3.go they fail on that assertion.
type trackedConn struct {
net.Conn
closeOnce sync.Once
closed chan struct{}
}
func (c *trackedConn) Close() error {
c.closeOnce.Do(func() { close(c.closed) })
return c.Conn.Close()
}
// trackingDialer hands out real UDP sockets and remembers every one of them.
type trackingDialer struct {
access sync.Mutex
conns []*trackedConn
}
func (d *trackingDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn, closed: make(chan struct{})}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *trackingDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func (d *trackingDialer) count() int {
d.access.Lock()
defer d.access.Unlock()
return len(d.conns)
}
func (d *trackingDialer) at(index int) *trackedConn {
d.access.Lock()
defer d.access.Unlock()
return d.conns[index]
}
func (d *trackingDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
func requireClosed(t *testing.T, conn *trackedConn, what string) {
t.Helper()
select {
case <-conn.closed:
case <-time.After(5 * time.Second):
t.Fatalf("%s: the UDP socket of the retired QUIC connection was never closed — quic-go does not own it, we must", what)
}
}
func requireDialed(t *testing.T, dialer *trackingDialer, want int) {
t.Helper()
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if dialer.count() >= want {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Fatalf("expected at least %d dial(s), got %d", want, dialer.count())
}
func testServerTLSConfig(t *testing.T, nextProtos []string) *tls.Config {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
return &tls.Config{
Certificates: []tls.Certificate{*certificate},
NextProtos: nextProtos,
MinVersion: tls.VersionTLS13,
}
}
func testClientTLSConfig(t *testing.T, nextProtos []string) sbTLS.Config {
t.Helper()
config, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
config.SetNextProtos(nextProtos)
return config
}
// startDoQServer serves a minimal DoQ responder and returns its address.
func startDoQServer(t *testing.T) M.Socksaddr {
t.Helper()
listener, err := quic.ListenAddr("127.0.0.1:0", testServerTLSConfig(t, []string{"doq"}), nil)
if err != nil {
t.Fatal(err)
}
ctx, cancel := context.WithCancel(context.Background())
t.Cleanup(func() {
cancel()
listener.Close()
})
go func() {
for {
conn, acceptErr := listener.Accept(ctx)
if acceptErr != nil {
return
}
go func(conn *quic.Conn) {
for {
stream, streamErr := conn.AcceptStream(ctx)
if streamErr != nil {
return
}
go func(stream *quic.Stream) {
defer stream.Close()
request, readErr := transport.ReadMessage(stream)
if readErr != nil {
return
}
response := new(mDNS.Msg)
response.SetReply(request)
transport.WriteMessage(stream, 0, response)
}(stream)
}
}(conn)
}
}()
return M.ParseSocksaddr(listener.Addr().String())
}
func testQuery() *mDNS.Msg {
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
return message
}
func TestQUICTransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
serverAddr := startDoQServer(t)
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeQUIC, "test-doq", nil),
dialer: dialer,
serverAddr: serverAddr,
tlsConfig: testClientTLSConfig(t, []string{"doq"}),
connection: transport.NewConnPool(transport.ConnPoolOptions[*quic.Conn]{
Mode: transport.ConnPoolSingle,
IsAlive: func(conn *quic.Conn) bool {
return conn != nil && !common.Done(conn.Context())
},
Close: func(conn *quic.Conn, _ error) {
conn.CloseWithError(0, "")
},
}),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
// Retire the connection the way a retryable error or an engine reload does.
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
// The reconnect must still work, on a fresh socket.
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err := dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func TestHTTP3TransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
message, err := readRequestMessage(request)
if err != nil {
writer.WriteHeader(http.StatusBadRequest)
return
}
response := new(mDNS.Msg)
response.SetReply(message)
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
writer.Write(rawResponse)
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
serverAddr := M.ParseSocksaddr(listener.Addr().String())
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
stdConfig := &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
}
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: serverAddr,
tlsConfig: stdConfig,
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err = dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func readRequestMessage(request *http.Request) (*mDNS.Msg, error) {
defer request.Body.Close()
rawMessage := make([]byte, 4096)
n, err := readFull(request.Body, rawMessage)
if err != nil {
return nil, err
}
var message mDNS.Msg
err = message.Unpack(rawMessage[:n])
if err != nil {
return nil, err
}
return &message, nil
}
func readFull(reader interface{ Read([]byte) (int, error) }, buffer []byte) (int, error) {
var total int
for total < len(buffer) {
n, err := reader.Read(buffer[total:])
total += n
if err != nil {
if total > 0 {
return total, nil
}
return total, err
}
}
return total, nil
}
+12
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"os"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/sing-box/adapter"
@@ -117,6 +118,12 @@ func (t *Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg,
rawConn.Close()
return nil, E.Cause(err, "establish QUIC connection")
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-earlyConnection.Context().Done()
rawConn.Close()
}()
return earlyConnection, nil
})
if err != nil {
@@ -144,6 +151,11 @@ func (t *Transport) exchange(ctx context.Context, message *mDNS.Msg, conn *quic.
return nil, E.Cause(err, "open stream")
}
defer stream.CancelRead(0)
stopWatch := context.AfterFunc(ctx, func() {
stream.CancelRead(0)
_ = stream.SetWriteDeadline(time.Now())
})
defer stopWatch()
err = transport.WriteMessage(stream, 0, message)
if err != nil {
stream.Close()
+44
View File
@@ -12,6 +12,50 @@ as GitHub **pre-releases** and never become "Latest".
#### Unreleased (shater)
**`l3-honest-drop` — ICMP routed to an L4-only outbound is dropped, not
forged** — ships with `shaterd` (part of the shater L3 ingress,
`docs-shater/DECISIONS.md` D25), not as an lx release tag; recorded here because
it edits two upstream files. Without it the TUN stack answers an unroutable echo
ITSELF — sing-tun's `ICMPForwarder.HandlePacket` rewrites Echo→EchoReply
whenever the flow judgment comes back Accept (`stack_gvisor_icmp.go`) — so a
ping routed to vless/vmess/… would read as a working tunnel while the packet
never left the router.
* **`route/route.go` (`PreMatch`)** — the pre-match walk was renamed to
`preMatch` and the exported `PreMatch` became a thin FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for
`N.NetworkICMP`. An earlier version overrode `continueResult` inside
`preMatchFlow` instead; that covered only the exits reaching that function and
left three of the walk's own exits forging — the `prepareMatchMetadata` error
return, the sniff bail-outs, and the `default:` arm of the rule-action switch
(every action pre-match has no arm for: `hijack-dns`, `direct`, …). A guard on
the single return value cannot be outgrown by a new exit. `PreMatchBypass` is
folded in because sing-tun implements `ActionBypass` on the nfqueue plane only
— on the TUN path it lands in the same `default:` arm as Accept, i.e. forges.
* **`adapter/router.go` (`JudgeFlow`, the `!isPort` branch)** — ICMP returns
`ActionDrop` where it fell through to `ActionAccept`. Second line of defense:
`adapter.FlowOutbound` and `tun.Port` are distinct interfaces, and a drift
between them must not quietly re-enable the forged reply.
* **TCP/UDP behaviour is unchanged** — `PreMatchContinue` still means "take the
ordinary connection route" for both, `PreMatchBypass` still means bypass, and
the `!isPort` fallthrough still returns `ActionAccept` for them; pinned by
`route/prematch_icmp_lx_test.go` and `adapter/judgeflow_icmp_lx_test.go`
(both inside the marker), each ICMP case having an explicit TCP/UDP twin.
* **NOT covered: a FRAGMENTED echo to a WireGuard/AWG outbound is still
forged** — sing-tun's `ForwardDispatcher.Dispatch` returns before asking for a
verdict at all when `parsed.fragment`, and the reassembled packet reaches
`ICMPForwarder.HandlePacket`, whose `installFlow` demands an UNSPECIFIED port
address that a WireGuard endpoint never has. Fixing it inside `JudgeFlow`
is NOT possible — both consumers call it with identical arguments and the
working path needs the concrete address. Full chain, the two viable fixes and
the trap are in `docs-shater/DECISIONS.md` D25, under "What is still NOT
covered, said plainly", item 2.
* **Rebase cost: two small marked blocks** (`lx:begin/end l3-honest-drop`, a
wrapper function in `route/route.go` and one branch body in
`adapter/router.go`) plus the two self-contained test files — carried across
an upstream rebase by eye. Note that `PreMatch`'s own body now lives in
`preMatch`, so an upstream change to the walk applies to that function.
**Fork-layer + control-plane rework of proxy health** — ships with `shaterd`
(the shater router daemon), not as an lx release tag; recorded here because the
load-bearing half lives in fork zones (`common/urltest`, `protocol/group`).
+63 -8
View File
@@ -62,16 +62,61 @@ flowchart LR
C["LAN client"] -->|"nft tproxy, mark → tproxy port"| IN["sing-box tproxy inbound (sniff SNI/Host/QUIC)"]
IN --> R{"route: rule match — src / dst / list / geo / client"}
R -->|"proxied"| OUT["outbound / selector (balancer, chain)"]
R -->|"direct"| DIR["direct (flow-offload on)"]
R -->|"direct"| DIR["direct (out the normal route, untunnelled)"]
R -->|"blocked"| BLK["block"]
OUT --> NET["exit — VLESS/Reality/AmneziaWG2/Hysteria2/…"]
```
Reliability (ported from v0.1): own nft table `inet shater` + own marks/tables
(never touch fw4); atomic validate→stage→swap; commit-confirm rollback;
idempotent reconcile under flock; management-bypass always; fail-closed
(never touch fw4); atomic validate→stage→swap; commit-confirm rollback (opt-in —
see §5); idempotent reconcile under flock; management-bypass always; fail-closed
kill-switch (dead group → block, not a silent direct leak).
Only TCP and UDP reach that path — TPROXY carries nothing else. What happens to
the rest is §3a.
### 3a. L3 ingress and kernel egress — what TPROXY cannot carry
Two opt-in globals cover the protocols the tproxy plane leaves on the floor.
Both are off in a stock config, and both are configured through UCI only (the
panel does not expose them).
**`globals.l3_tunnel` — LAN ICMP through the tunnel.** The generator adds a
synthetic `tun` inbound tagged `l3-in` (gVisor stack, `auto_route` **off**, MTU
65535, `shater/generate/inbound.go`), so ICMP is routed by the engine's own rules
instead of being dropped or answered by a forged local reply. The device is not
one fixed name: the generator emits a stable placeholder (so a no-op reconcile
still hashes identical and does not rebuild the engine once a minute), and
`shater/engine/l3slot.go` substitutes one of the two slots `shater-l3a` /
`shater-l3b` (`netplane/l3.go`) just before `box.New` — a new generation must
never reopen the name the outgoing one still holds
(`TUNSETIFF: device or resource busy` took the whole LAN down once). The routing half is scoped and lives entirely outside
the main table: our nft prerouting chain stamps LAN `icmp`/`ipv6-icmp` with
`L3Mark` (`fwmark_base + 0x80`), and `netplane.addL3Routing` binds that mark to
`L3Table` (`table_base + 8`), whose only content is a default route out the live
slot. Because the daemon creates the device at runtime, netifd never learns about
it and fw4 would reject the forward on its own account — so `30_shater-core`
seeds a **`shater_l3` zone in the user's `/etc/config/firewall`**, matching
`list device 'shater-l3*'` (a string match that is valid before the TUN exists
and covers both slots). Ceiling: ICMP echo only, and only for L3-capable
egresses; see `DECISIONS.md` D25 for what is still not covered.
**`globals.untunnelable_egress` — everything else, carried by the kernel.** It
names an existing interface/tunnel egress. Whatever the L3 block above did not
claim — ESP/AH, GRE, IGMP, SCTP, and ICMP too when `l3_tunnel` is off — is
stamped in prerouting with **that egress's own mark** (`netplane/nft.go`,
`UntunnelableEgressBinding`) and accepted; the `fwmark → table` pair
`addEgressRouting` already installed for the egress then routes it out the
egress's device. No new mark, no new table, and the engine never sees a byte —
which is why any IP protocol works here while the L3 TUN is narrow. Order is
load-bearing: this sweep runs **after** the L3 marking (first match wins) and
**after** the local-plane accepts, so LAN-to-LAN, router-addressed traffic and
IPv6 neighbour discovery never leave through an uplink. With `ipv6=0` the mark
is scoped to `nfproto ipv4`, because `addEgressRouting` installs the `-6`
rule/table pair only when IPv6 is on and marked v6 without it would fall through
to the main table past the kill-switch. `globals.untunnelable` (block | icmp |
direct) stays in charge of whatever neither mechanism carries.
## 4. DNS + filtering + stats
```mermaid
@@ -79,7 +124,7 @@ flowchart LR
C["client :53"] -->|"hijack"| DNS["sing-box DNS (in-process)"]
DNS --> FILT{"shater filter: blocklists + allowlist + per-device policy"}
FILT -->|"blocked"| NX["NXDOMAIN / 0.0.0.0"]
FILT -->|"allowed"| RES["resolvers (DoH/DoT/plain/FakeIP) + nftset for routing"]
FILT -->|"allowed"| RES["resolvers (DoH/DoT/plain/local/FakeIP), per-rule detour"]
DNS -->|"query events (engine observability)"| AGG["shater stats aggregator"]
AGG --> PANEL["panel: top domains · per-device · allowed/blocked · timeline"]
```
@@ -87,9 +132,10 @@ flowchart LR
Because the engine's DNS runs **in our process**, every query (domain, client,
verdict, latency) is available to the stats aggregator without log-scraping —
this is the payoff of embedding. Blocklist matching uses an efficient compiled
matcher, not dnsmasq megalists (see `DECISIONS.md` D5). Per-device blocking =
engine route/DNS rule keyed by client, or nftset(device) × nftset(blocked-domain)
→ drop.
matcher, not dnsmasq megalists (see `DECISIONS.md` D5). Per-device blocking is an
engine route/DNS rule keyed by client. Routing decisions come from in-engine
rule-sets: the v0.1 mechanism where dnsmasq populated nft sets does not exist in
v0.2 (`generate/dns.go`).
## 5. Config & apply flow
@@ -100,11 +146,20 @@ stateDiagram-v2
Render --> Validate: engine config check + nft -c
Validate --> KeepOld: fail
Validate --> Apply: ok (atomic swap: engine reload + nft/route reconcile)
Apply --> ConfirmWindow
Apply --> Committed: confirm_timeout = 0 (SHIPPED DEFAULT — nothing armed)
Apply --> ConfirmWindow: confirm_timeout > 0
ConfirmWindow --> Committed: confirmed
ConfirmWindow --> Rollback: timeout
Rollback --> LastGood
```
**The confirm window is opt-in and ships closed.** `model.DefaultGlobals()` leaves
`ConfirmTimeout` at zero, the shipped `/etc/config/shater` says
`option confirm_timeout '0'`, and `apply.ArmRollback` returns immediately on a
non-positive timeout — so on a stock install every apply takes the left edge above
and there is no net under it. `shaterd apply` reports which edge it took
(`reason: commit-confirm-off` vs an armed window). Set
`globals.confirm_timeout` to arm it.
## 6. Roadmap tiers
See `ROADMAP.md` for the phased plan and `FEATURES.md` for the full feature list.
+29 -16
View File
@@ -51,6 +51,9 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
MASQUE/WARP, and gRPC observability (DNS queries / rules / outbounds). Upstream
sing-box brings VLESS/VMess/Trojan/Shadowsocks/WireGuard/Reality + Hysteria2/
TUIC. It is library-first (`libbox`) and **GPL-3.0** (compatible with us).
That list is what the FORK can build, not what shater ships: `shater/registry`
registers only what `shater/generate` can emit, and MASQUE is one of the types
deliberately left out (~6 MB of binary and resident RAM). See `FEATURES.md`.
- We **fork it** (not just depend on it) so we can embed literally everything —
control-plane, admin panel, DNS filter — and integrate tightly with the
engine internals (DNS, routing, stats). This is a deliberate, decided
@@ -84,13 +87,13 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
## Repository model
- **`shater` `main` = our fork of sing-box-lx.** After Phase 1 it contains the
full sing-box-lx tree PLUS our additive overlay (`shater/`, `panel/`,
`openwrt/`, `docs-shater/`). Upstream is tracked via a git remote and merged by tag.
- **`shater` `main` = our fork of sing-box-lx.** It contains the full sing-box-lx
tree PLUS our additive overlay (`shater/`, `panel/`, `openwrt/`, `docs-shater/`,
`scripts/`, `ci/`). Upstream is tracked via a git remote and merged by tag.
Phase 1 merged the engine in on 2026-07-14 (`v1.14.0-lx.3`); `main` has not been
a docs-only seed since.
- **`shater` branch `v0.1`** = the standalone xray-based version (frozen, ported
from).
- Until Phase 1 merges the engine in, `main` is the docs-first overlay seed you
are reading now (LICENSE, README, `docs-shater/`, the feed signing key).
## What to port from v0.1 (don't rewrite these ideas)
@@ -119,20 +122,30 @@ filter/stats engine wired into sing-box's DNS.
v0.2 fork; branch `v0.1` = the working xray-based version.
- **Upstream to track:** `https://github.com/Leadaxe/sing-box-lx` (which tracks
`https://github.com/SagerNet/sing-box`).
- **CI:** Gitea Actions (act_runner + Docker). v0.1's workflow was removed from
`main`; new CI is added when the v0.2 build exists.
- **CI:** Gitea Actions (act_runner + Docker), `.gitea/workflows/release.yml` —
builds the four packages through the ImmortalWrt 25.12.1 SDK and publishes the
signed per-arch apk repo. The opkg/`.ipk` lane was deleted, not disabled (D22).
- **Test gate:** `bash scripts/run-tests.sh` — the whole suite under the SHIPPED
build tags, on linux (in Docker from a non-linux host), with `-race`, and with
three anti-silent-skip checks. Not optional reading before touching `shater/`.
- **Feed signing:** EC (prime256v1) key for the apk index; secret in the repo
secret `KEY_APK`; public key `dist/shater-apk.pem`, installed on routers as
`/etc/apk/keys/shater-apk.pem`. Never regenerate it (D22).
- **Test VM:** OpenWrt 24.10.3 x86_64 in Docker (`docker ps --filter
name=openwrt-vm`). SSH via the ssh-manager MCP server `local_openwrt`
(localhost:2222, root/openwrt). LuCI at `http://127.0.0.1:8080` (root/openwrt),
drivable with the Playwright MCP.
- **Test VM:** **ImmortalWrt 25.12.1** (`r37978-cd0a06bfd3fd`) x86_64 in Docker
(`docker ps --filter name=openwrt-vm`), apk-tools 3.0.5 — deliberately the same
revision as `mini_router`, and required: the only package format we publish is
`.apk`, which does not install on 24.10 at all. SSH via the ssh-manager MCP
server `local_openwrt` (localhost:2222, root/openwrt). LuCI at
`http://127.0.0.1:8080` (root/openwrt), drivable with the Playwright MCP.
- **Routers:** `mini_router` (BPi-R3 Mini, ImmortalWrt 25.12.1) carries the real
home traffic; `main_router` (BPi-R4, OpenWrt 25.12.0). Both `aarch64_cortex-a53`,
both apk-tools 3.0.5 — see the table in D22.
## Current status
Repo reset done: v0.1 preserved on its branch; `main` cleaned to this docs-first
scaffold. Next is Phase 1 in `ROADMAP.md` — fork sing-box-lx into `main`
(add upstream remote, merge a pinned tag), stand up the embedding prototype
(prove AmneziaWG 2.0, measure binary size with feature-trim + `-s -w` + UPX)
before building the control plane and panel.
**v0.2 is feature-complete and running on real hardware.** ROADMAP Phases 0–8 are
done and VM-verified; the product ships as a signed apk feed and is installed on
`mini_router`. Read `ROADMAP.md` for what each phase delivered, `FEATURES.md` for
the honest MVP/T1/T2 state of each feature (including what is declared but not
shipped), and `DECISIONS.md` for why. Work since Phase 8 has been correctness and
honesty passes rather than new phases.
+646
View File
@@ -319,6 +319,10 @@ Three values, not two, because the leaks differ in *kind*: an ICMP echo is ephem
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
*(Refined 2026-07-26 by D25: still true of TPROXY — but ICMP echo now has an
opt-in data plane of its own, the dedicated L3 TUN, so the policy no longer
speaks alone for ping; it keeps sole charge of ESP/GRE/IGMP and of the degraded
paths.)*
**Fail-open degradations must be visible in the panel, not only in `logread`.** The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
@@ -772,3 +776,645 @@ server, which restores exactly the pre-D24 behaviour and clears the notice. Unti
that lands, an operator can get the same result by setting `endpoint_resolver` to a
direct resolver. Note the hazard is **not** created by D24 — any config with two
resolvers has it today; the default merely makes it universal.
## D25 — L3 ingress: LAN ICMP rides a dedicated TUN through the tunnel, not a policy verdict
Decided 2026-07-26. D17 made everything TPROXY cannot divert an explicit policy
(`Globals.Untunnelable` = block | icmp | direct) — and its premise still holds:
kernel TPROXY delivers a packet by handing it to a listening SOCKET, and sockets
exist for TCP and UDP only, so an ICMP echo has nothing to be handed to. But a
policy can only choose between losing the packet and leaking it with the
client's real source address; neither ever puts a ping THROUGH the tunnel. This
decision adds the data plane D17 could not have: **`globals.l3_tunnel` (opt-in,
default off; `model.Globals.L3Tunnel`) opens a second, dedicated ingress — a TUN
device — and LAN ICMP enters the engine as raw IP packets**, where the ordinary
route rules pick an outbound exactly as for any flow. The policy is refined, not
repealed: it keeps sole charge of the protocols the engine cannot ingest at all,
and of the degraded paths (both below).
**The whole mechanism is one mark, one rule, one device — the TPROXY plane is
untouched.** The nft prerouting chain stamps `L3Mark` (= fwmark_base + 0x80,
`netplane/nft.go` `l3MarkOffset`) on LAN `ip protocol icmp` / `meta l4proto
ipv6-icmp` ONLY, and only after every local plane was already accepted
(fib-local, RFC1918/link-local/multicast daddr sets) and — for v6 — after a
unicast ND/NA carve-out, because one tunnelled neighbour probe is enough to take
the LAN's v6 plane down (`renderNft`, the L3 block). `addL3Routing`
(`netplane/apply.go`) binds that mark to a table (= table_base + 0x08) whose
only content is `default dev shater-l3`; del-then-add idempotent, and a failed
rule or route is a NAMED operator warning, never an apply abort. `generate`
emits the synthetic `l3-in` TUN inbound bound to exactly `netplane.L3Device`,
MTU 65535 (the largest IP datagram there can be, so the KERNEL can never
fragment on the way in — see "the device MTU is not a tunnel budget" below),
point-to-point /30 + /126 addresses from private space,
the v6 one only when `globals.ipv6` is on — and only next to a tproxy inbound:
the ingress rides the same LAN divert plane, and without one the TUN would sit
dark while the config claims ICMP is tunnelled, so it is skipped with a warning
(`generate/inbound.go`, `appendL3TunInbound`). `shater/registry` registers the
`tun` inbound type; that costs no new build tag and no meaningful size because
`with_wireguard` already requires `with_gvisor` (D23, `scripts/router-tags.sh`).
**`auto_route: false` is load-bearing, not a default we happened to keep.**
sing-box's auto_route rewrites the router's MAIN routing table — it would drag
everything the router itself sends (WAN traffic, DNS, the tunnel's own underlay)
into this TUN. The fwmark rule + dedicated table above is deliberately the ONLY
entrance, and disabling the feature can never strand a stale default route in
main (`generate/inbound.go`; pinned by `TestL3TunnelEmitsTunInbound`).
**`stack: "gvisor"` is a deliberate choice, and the tempting reason for it is
wrong.** It is TRUE that sing-tun's system stack answers an ICMP echo LOCALLY —
`processIPv4ICMP` rewrites Echo→EchoReply in place and swaps the addresses
(sing-tun `stack_system.go:648`; the v6 twin sits right under it). It is FALSE
that this makes the system stack unusable here: `dispatchIPv4`
(`stack_system.go:355-372`) hands the packet to the SAME `ForwardDispatcher`
first and only falls through to that forger for packets addressed to the TUN
itself, exactly as the gVisor filter does (`stack_gvisor_filter.go:52-113`).
Both stacks would forward. gvisor is chosen because it is already linked —
`with_wireguard` requires `with_gvisor` (D23), so it costs no tag and no new
code path — and because it is the combination the integration test actually
exercises. Do not re-derive this as "the system stack fakes ping": it fakes ping
only where the dispatcher declined the packet.
**The ceiling is ICMP echo, and it is upstream's dispatcher — NOT the netstack.**
This distinction matters because the netstack answer is the intuitive one and it
is wrong. On the forward path a WireGuard/AWG endpoint never consults gVisor at
all: `Endpoint.WritePackets` (`transport/wireguard/port.go:21-58`) reads the IP
version and the destination address and hands the raw bytes to
`wgDevice.InputPackets` — the protocol byte is never examined — and
`returnDeviceWrapper.Write` (`:127-157`) offers every decrypted packet to
`returnPath.ReturnPackets` before the stack sees it. WireGuard would carry ESP
today if anything handed it one. What refuses is `ForwardDispatcher`: its parser
sets `hasFlow` for TCP, UDP and ICMP echo alone (`flow_parse.go`,
`parseTransport`, the echo identifier serving as the pseudo-port), and
`createFlow` NATs through a port-shaped selector (`flow_dispatch.go:325`,
`allocateSelector`) that ESP, AH and GRE do not have. So ESP/AH/GRE/IGMP/SCTP
cannot enter the engine in ANY configuration and REMAIN on the D17 policy —
or on the kernel egress of D26, which sidesteps the dispatcher entirely. The nft
plane encodes the same boundary on purpose: it marks `icmp`/`ipv6-icmp` only,
never `l4proto != { tcp, udp }`, because a marked ESP packet would enter the
device and vanish — a black hole wearing a tunnel's name — instead of receiving
the policy's honest verdict (`netplane/nft.go`, the prerouting L3 comment).
**What works and what does not, read off the upstream source.** ping v4/v6 —
yes. Windows `tracert` — yes: the gVisor return path recognises
`ICMPv4TimeExceeded` and `ICMPv4DstUnreachable` alongside EchoReply and NATs
them back to the LAN client (`stack_gvisor_icmp.go:341+`, `returnPacket`). IPv6
traceroute — intermediate hops stay invisible: the v6 branch of the same
function accepts EchoReply only, so just the final destination answers. Several
LAN clients behind the one tunnel address are already solved upstream:
`ForwardDispatcher` NATs by echo identifier and rewrites the source to the
outbound's port address (`flow_dispatch.go:325+`, `createFlow`; `icmpFlowKey`) —
we wrote no NAT of our own.
**Which outbounds can carry it.** The contract is `adapter.FlowOutbound`
(= `Outbound` + `tun.Port` + `PreMatchFlow`, `adapter/outbound.go`). In-tree
implementors: the WireGuard/AWG endpoint (`protocol/wireguard`), `direct`
(`protocol/direct`), `bridge` (`protocol/bridge`), `tailscale`
(`protocol/tailscale`). Of those, the shaterd registry can construct only
WireGuard/AWG and direct (`shater/registry/registry.go` — bridge and tailscale
are not registered). Every proxy protocol — vless/vmess/trojan/shadowsocks/
hysteria2/tuic/socks/http/shadowtls — is L4-only and cannot. Recorded as a known
gap: `masque` is L3 by nature (CONNECT-IP; it builds a userspace gVisor stack
per tunnel, `protocol/masque/outbound.go`) but implements no `tun.Port` and is
not in the shater registry, so today it cannot carry the ingress. Wiring it up
is possible future work, not a promise.
**ICMP to an L4-only outbound is DROPPED, and that took patching upstream files
(the `lx:l3-honest-drop` delta — see `docs-lx/lx-changelog.md`).** In the gVisor
stack the fallthrough verdict is a forgery: `ICMPForwarder.HandlePacket` answers
the echo ITSELF (Echo→EchoReply + address swap) whenever the flow judgment comes
back Accept (`stack_gvisor_icmp.go:120`), and upstream maps "no flow route" to
exactly that Accept — so a ping routed to vless would read as tunnelled while
the packet died on the router. Two small marked hunks make the truth observable:
`route/route.go` wraps the whole pre-match walk — the walk itself became
`preMatch`, and the exported `PreMatch` is now a FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for `N.NetworkICMP` —
and `adapter/router.go` (`JudgeFlow`, the `!isPort` branch) returns `ActionDrop`
for ICMP where it fell through to `ActionAccept` — the second line of defense,
because `FlowOutbound` and `tun.Port` are distinct interfaces and a drift
between them must not quietly re-enable the forger. TCP/UDP verdicts are
byte-identical; `route/prematch_icmp_lx_test.go` and
`adapter/judgeflow_icmp_lx_test.go` pin both directions. The operator-facing
text says the same out loud (`shater/apply/warnings.go`): proxy-routed addresses
"cannot be pinged at all — deliberately".
> **Why a funnel and not an override inside the walk.** The first version of
> this delta overrode the pre-declared `continueResult` inside `preMatchFlow`
> and claimed to cover "every exit point of the function at once". It covered
> every exit of THAT function; the walk above it has exits of its own that never
> reach it — the `prepareMatchMetadata` error return (which arrived later, with
> the shared-metadata refactor, upstream `b911fb078`), the sniff bail-outs, and
> the `default:` arm of the rule-action switch, which catches every action
> pre-match has no arm for (`hijack-dns`, `direct`, and whatever upstream adds
> next). Each of those returned `PreMatchContinue`, i.e. `tun.ActionAccept`,
> i.e. the forged reply. A guard on the single return value cannot be outgrown
> by a new exit. `PreMatchBypass` joined the drop for the same reason: sing-tun
> implements `ActionBypass` on the nfqueue plane only — the name appears nowhere
> in `flow_dispatch.go` or `stack_gvisor_icmp.go` — so on the TUN path it lands
> in the same `default:` arm as Accept and forges too. There is no honest bypass
> for a packet that is already inside the engine's TUN.
**The device MTU is NOT a tunnel budget, and pretending it was manufactured
forged replies.** `l3-in` is created with MTU **65535**, not the tunnel's 1420,
and the maximum is the whole argument. This MTU governs exactly one thing:
whether the KERNEL splits a packet on its way INTO the device. What the engine
then puts into the tunnel is sized separately and correctly, against the
OUTBOUND's MTU — `ForwardDispatcher.forwardToPort` (`flow_dispatch.go:445-481`)
measures every forwarded packet against `Port.PortMTU()` and either fragments to
it (no DF, `fragmentIPv4Packet`) or answers a well-formed `fragmentation needed`
quoting it (DF, `buildFragmentationNeeded`, source = the far host, so PMTU
discovery works end to end). That machinery was always there; it was simply
never handed a whole packet.
At 1420 it wasn't. Anything above 1392 bytes of payload was fragmented by the
kernel at this device, and a fragment is the one thing sing-tun will not judge:
`Dispatch` (`flow_dispatch.go:176-177`) returns on `parsed.fragment` BEFORE
calling `JudgeFlow` at all. The fragments fell through to the gVisor stack —
promiscuous and spoofing (`stack_gvisor.go:219-223`) — which reassembled them
and handed the echo to `ICMPForwarder.HandlePacket` (`stack_gvisor_icmp.go:105+`),
whose `installFlow` (`:233-244`) writes to the port UNMODIFIED and therefore
demands a port address that is valid **and UNSPECIFIED**. `direct` qualifies
(`IPv4Unspecified()`); a WireGuard/AWG endpoint reports its concrete interface
address (`transport/wireguard/port.go:13`) and does not. So it declined, and
`HandlePacket` fell past the switch and FORGED the reply: `SetType(EchoReply)` +
address swap. Net effect on the operator's bench: `ping -s 1392` honest,
`ping -s 1393` a lie told by the router — and the lie was, of course, only for
the outbounds this feature exists for. (Upstream applies the very same
unspecified test and answers it honestly in the cloudflared ICMP handler,
`protocol/cloudflare/inbound.go:163-167`: it drops. Only the TUN path forges.)
65535 rather than "big enough": no IP datagram can exceed it, so the kernel
CANNOT fragment at this device, for any packet, ever. Any smaller value leaves
a band of sizes open and re-opens the class. It is also sing-box's own default
TUN MTU on Linux. Pinned by `TestL3TunnelMTULeavesNothingForTheKernelToFragment`
and `TestL3TunnelMTUIsNotATunnelBudget` (`generate/l3mtu_test.go`), and — the
assertion that matters — by the integration test reading the MTU back off the
real kernel device, since a kernel that clamped it would restore the forgery
without changing a generated byte.
Memory was MEASURED, not reasoned about: three paired runs of
`TestIntegrationL3TunInboundStarts` under `-test.memprofilerate=1` (exact
accounting, not sampled) allocate 5.41 / 5.48 / 5.47 MB at 65535 against
5.76 / 5.46 / 5.70 MB at 1420, and a `-diff_base` profile attributes every
difference to netlink interface enumeration. Nothing in the read path scales
with the MTU: gVisor reads through `fdbased.BufConfig`, which sing-tun's `init`
pins to a single 65535-byte view regardless of MTU, and `fdbased` keeps `mtu`
only to return it from `MTU()`. Two adjacent facts, recorded because both are
easy to derive wrongly: (a) `protocol/tun` computes
`enableGSO = stack == gvisor && mtu < 49152`, so this MTU turns GSO off there —
and then `StartStateStart` turns it back ON unconditionally because an
`adapter.FlowOutbound` exists in the config, so the ~1.98 MB of TCP/UDP GRO
scaffolding is present at BOTH MTUs and is priced by the flow-capable outbound,
not by this number; (b) the `mtu_fix` on the `shater_l3` fw4 zone is now inert —
only ICMP is ever marked into the device — and its uci-defaults comment still
says "the tunnel MTU is 1420".
**What is still NOT covered, said plainly.**
1. **A big ping does not start WORKING — it starts FAILING HONESTLY.** Upstream's
ICMP NAT is unfragmented-only in BOTH directions: `classifyReturn`
(`flow_dispatch.go:703-710`) returns `returnPass` on `parsed.fragment` exactly
as the forward path does. So a non-DF `ping -s 2000` now genuinely leaves the
router (fragmented to the tunnel MTU by `forwardToPort`), the far host really
answers, and the reply — fragmented by the peer to fit the tunnel — is not
NAT'd back to the LAN client. The operator sees a timeout. That is the
feature's promise ("travels or fails honestly"), not a capability claim.
Carrying oversized ICMP end to end would need reassembly upstream does not
have; it is not planned.
2. **A client that puts fragments on the wire ITSELF.** The device MTU cannot
un-fragment what already arrived fragmented, so such packets still reach the
gVisor stack, still get reassembled there, and still receive a forged reply
when the outbound is WireGuard/AWG. This is the residue the planned
`ip frag-off & 0x3fff != 0` prerouting carve-out (`netplane/nft.go`) is for.
**Whoever writes that rule must first check whether it can ever match:** fw4's
ruleset uses conntrack, conntrack pulls in `nf_defrag_ipv4`/`nf_defrag_ipv6`,
and defrag REASSEMBLES in PREROUTING before our marking rules run. Where
defrag is active the case does not arise (the MTU covers it) and the rule is
dead; where it is not, the rule is the only cover. Verify on the bench with
`nft list ruleset | grep -c ct` and a fragment counter, do not assume.
3. **The DF path changed hands and is untested on hardware.** It used to be the
kernel that answered `fragmentation needed` (from the router's LAN address,
MTU 1420); it is now the engine (from the far host's address, quoting
`Port.PortMTU()`). Both are correct PMTUD; only the first has ever run on a
real router.
**fw4 has to be told about the device, and `list device` is the only spelling
that works.** nftables runs EVERY table on every packet and a drop in any one of
them wins — an accept in `inet shater` cannot override fw4, and fw4 WILL reject
this forward: netifd never learns about a device the daemon creates at runtime,
so `shater-l3` belongs to no zone and falls into fw4's zone-less defaults. Hence
a real fw4 zone `shater_l3` + a lan→shater_l3 forwarding, seeded idempotently
(NAMED sections) and unconditionally in uci-defaults
(`openwrt/shater-core/files/etc/uci-defaults/30_shater-core`, `seed_l3_zone`),
with `mtu_fix` set. That `mtu_fix` is now inert and should be read as such: it
clamps forwarded TCP MSS to the route MTU, the device MTU is 65535, and nothing
but ICMP is ever marked into this device — the uci-defaults comment still says
"the tunnel MTU is 1420" and is stale. The device is attached
via `list device`, deliberately NOT `list network`: fw4 resolves a zone's
networks through netifd, which yields an EMPTY device set for a runtime-created
TUN (a proto-none stub would have to be brought UP to contribute an l3_device,
and nothing ever brings it up), while `list device` compiles to a plain
iifname/oifname string match — valid before the TUN exists, matching from the
moment shaterd creates it. `kmod-tun` joined DEPENDS so a slimmed image cannot
lose `/dev/net/tun` (`openwrt/shater-core/Makefile`). Our own forward chain
accepts both TUN legs ahead of the fail-closed drops — accepts that speak for
OUR table only (`netplane/nft.go`, forward chain step 4).
**What the policy still owns, and the one combination that now warns.** With the
ingress on, the mark is stamped in prerouting and the ROUTING decision carries
echo into the TUN before the forward chain — where the policy's verdicts live —
is ever consulted; that holds under every `untunnelable` value. The policy
therefore governs exactly two things: the never-markable protocols above, and
the fallback when the L3 rule/route did not come up (engine down, partial apply)
— `block` turns that failure into an honest loss, `direct` into a silent leak
with the real address. That is why `l3_tunnel` + `untunnelable=direct` draws a
validation warning naming the safe choice (`model/validate.go`), why every
rule/route failure surfaces as a named panel warning rather than an abort
(`addL3Routing`), and why the D17 HOLDING plane never marks: the TUN is created
BY the engine, and the holding plane exists precisely because the engine is not
running — marking would dead-end ping in a device that does not exist
(`netplane/nft.go`, hold comment).
- **Rejected: `auto_route` / letting the engine own the routing.** It rewrites
the main table and intercepts the router's own WAN/DNS/underlay traffic; the
blast radius of a toggle meant for LAN ping would be the whole router.
- **Rejected: marking all `l4proto != { tcp, udp }` into the TUN.** ESP/AH/GRE/
IGMP/SCTP cannot be parsed into flows upstream; they would vanish inside the
device. A drop with a name (the policy's) beats a silent black hole.
- **Rejected: keeping upstream's accept-and-forge for unroutable ICMP.** A ping
that "works" without leaving the router is the inverted lie this project keeps
deleting (D17's fiction purge, D23's dead WireGuard, D24's obedient-client
leak).
**Proven, and not proven, said plainly.** The cold start is PROVEN, not assumed:
`TestIntegrationL3TunInboundStarts`
(`shater/generate/l3_integration_linux_test.go`, run as root with NET_ADMIN and
`/dev/net/tun`, PASS) drives an `l3_tunnel=1` config through the SLIM registry
(`registry.Context`, not upstream's `include.Context`) under the shipped router
tag set: `box.New` + `Start` accept it, the kernel really ends up with the
`shater-l3` device at the contract MTU 65535 — the assertion the value exists
for, since a kernel that clamped it would silently restore the forged-reply
band — and Close removes it; precisely
the "built with X, verified with Y" gap class D23 exists for (a lost
`tun.RegisterInbound` or a trimmed `with_gvisor` changes no generated byte and
would otherwise surface only on the operator's router). Each layer contract is
pinned besides (`generate` `TestL3Tunnel*`, `netplane` `TestL3Ingress*`,
`route/prematch_icmp_lx_test.go`). Exactly two things remain UNVERIFIED:
(a) the end-to-end path on live hardware — LAN client → prerouting mark →
ip rule → TUN → WireGuard peer → reply back to the client — has not been
exercised on a real router; (b) the steady-state memory cost of the second
gVisor netstack (the `l3-in` TUN beside the WireGuard endpoint's own) is
unmeasured on the target hardware. An indicative figure exists and is only
that: on x86_64 in a container, idle and carrying no flows, peak RSS of a
process that brought the same engine up went from ~26.0-26.8 MB without
`l3_tunnel` to ~28.3-28.7 MB with it over three paired runs — about +2.2 MB.
That was measured on a throwaway harness, not on aarch64, not under load, and
with an empty ICMP NAT table, so it bounds nothing on the router. Neither
item is folded into any claim above.
Consequence: a ping from the LAN either genuinely travels through the tunnel
(WireGuard/AWG, direct) or fails honestly, at every size the router itself can
put into the device — and a router that never opts in renders the pre-feature
plane byte-for-byte (`TestL3IngressOptIn` pins the off-state render). Read
"fails honestly" strictly: above the tunnel MTU a non-DF ping now leaves the
router for real and then times out, because upstream's ICMP NAT does not carry
fragments back either. The one qualifier left is item 2 above — a client that
puts fragments on the wire ITSELF, on a router where conntrack defrag is not
reassembling them first. This paragraph has been overclaimed twice already;
extend it only against a bench result, never against a reading.
## D26 — What the engine cannot carry, the kernel carries: `untunnelable_egress`
Decided 2026-07-26. D25 ended with ESP/AH/GRE/IGMP/SCTP still owned by the D17
policy — that is, with a choice between dropping them and leaking them out the
WAN, never a data plane. This decision gives them one, and deliberately NOT
ours: **`globals.untunnelable_egress` (default empty;
`model.Globals.UntunnelableEgress`) names an existing egress of type
interface/tunnel, and LAN traffic that is neither TCP nor UDP is stamped in
prerouting with that egress's own mark, so the KERNEL routes it out that
egress's device with the kernel's own NAT.** No proxy, no engine, no userspace
stack ever touches the packet — which is exactly why every protocol works.
**Where the engine's boundary actually is — recorded so nobody digs for it
twice.** It is NOT the gVisor stack, and it is not WireGuard: on the forward
path the WG/AWG endpoint never consults gVisor at all. `Endpoint.WritePackets`
(`transport/wireguard/port.go:21-58`) takes the raw IP packet bytes, reads
exactly the IP version and the destination address, and hands
`device.InputPacketRef`s to `wgDevice.InputPackets` — the protocol byte is
never read; on the way back (`port.go:127-157`) `returnDeviceWrapper.Write`
offers every decrypted packet to `returnPath.ReturnPackets` first and only the
unconsumed remainder falls through to the gVisor device. gVisor serves
`DialContext`/`ListenPacket` — traffic the ENGINE originates — while forwarded
traffic bypasses the stack in both directions, indifferent to protocol. The
real ceiling sits one step earlier, in sing-tun's `ForwardDispatcher`:
`parseTransport` (`flow_parse.go:106-153`) sets `hasFlow` for exactly TCP, UDP,
ICMPv4 Echo/EchoReply and ICMPv6 EchoRequest/EchoReply — a packet of any other
protocol is never dispatched as a flow — and `createFlow`
(`flow_dispatch.go:325`) builds its NAT through
`allocateSelector(packet.protocol, …, packet.source.Port())` (line 355), which
needs a port-like selector that ESP/AH/GRE simply do not have (SCTP has ports,
but the parser above never grants it a flow either). Tailscale documents the
same frontier for its own userspace mode — "Any IP protocol other than TCP or
UDP (such as SCTP) is not supported in userspace mode… All IP protocols are
supported" in kernel mode
(https://tailscale.com/docs/reference/kernel-vs-userspace-routers) — useful as
external corroboration of where userspace data planes generally end, though OUR
boundary is the dispatcher, not the stack. The kernel egress was therefore
chosen not because userspace "cannot" in principle, but because the kernel
delivers all protocols with zero new code on the hot path.
**The mechanism already existed; the feature is one binding and one marking
step.** `addEgressRouting` (`netplane/apply.go`) has always installed, for
every interface/tunnel egress, an `ip rule fwmark <EgressMark> lookup
<EgressTable>` plus a `default dev <device>` route in that table — per-rule
egress selection rides on it. The only missing piece was that nothing ever
marked non-TCP/UDP traffic: `untunnelable=direct` merely ACCEPTED it in the
forward chain, so it left over the main table, i.e. the WAN.
`UntunnelableEgressBinding` (`netplane/nft.go`) resolves the option to the
egress's index, its OWN mark and its OWN device — deliberately no third
mark/table pair to keep coherent — and the prerouting chain stamps that mark on
the untunnelable protocols. A name that does not resolve to an interface/tunnel
egress with a device renders nothing and is reported: the D17 policy stays in
sole charge, which is the fail-closed reading of a typo.
**Why `l4proto != { tcp, udp }` is safe here when D25 banned it.** D25 rejected
the broad filter because the receiving side was the `ForwardDispatcher`, which
classifies nothing beyond TCP/UDP/ICMP echo — a marked ESP packet would enter
the TUN and vanish, a black hole wearing a tunnel's name. Here the receiving
side is the kernel, which forwards ANY IP protocol and NATs what it has
machinery for: SCTP carries ports and NATs like TCP/UDP; GRE is NATed only
through the PPTP helper keyed on the call-id — the kernel's own comment calls
GRE "generally not very suited for NAT, as it has no protocol-specific part as
port numbers" (`net/netfilter/nf_conntrack_proto_gre.c`); ESP/AH pass as plain
routed IP. Nothing on this path can silently swallow a protocol it does not
understand, which was the entire objection.
**Order against D25: the L3 ingress claims ICMP first.** With `l3_tunnel` on,
LAN ICMP is marked into the engine's TUN before the egress carrier is consulted
— the engine path routes ping by the operator's rules, which a kernel egress
cannot do — and only the remaining protocols go to the egress. With `l3_tunnel`
off, ICMP goes to the egress with everything else. In both shapes marked
traffic is settled by ROUTING before the forward chain speaks, so the D17
policy now governs exactly the failure case — the rule or route that did not
come up — the same division D25 already established for the L3 mark.
**What the feature refuses to promise — and the operator text refuses with it
(`shater/apply/warnings.go`, the egress-carrier note).** (a) It is not a tunnel
per se: the option accepts any interface/tunnel egress, and on the target
routers a WireGuard device is the exception (`kmod-wireguard` is usually
absent) while a second WAN is routine. Through a WireGuard egress this
genuinely is a tunnel; through a second WAN it is simply another uplink, and
the destination sees that uplink's real address. No text, comment or doc line
may call it a tunnel unconditionally. (b) It does not revive IPTV: IGMP is
LAN-side multicast group management, WireGuard is L3 point-to-point and carries
no multicast, and multicast never crossed this router under any setting —
routing IGMP out an egress restores nothing, and no wording may hint otherwise.
(c) IPsec through NAT-T never needed it: RFC 3948 encapsulates ESP in UDP/4500,
so a modern IPsec client behind NAT is ordinary UDP that already follows the
routing rules; the raw-ESP case this feature carries is the no-NAT-T remainder.
- **Rejected: teaching the engine these protocols.** Extending `parseTransport`
and the selector NAT upstream would be new hot-path code in an
actively-maintained adversarial area, for protocols the kernel already
forwards for free — and for ESP/AH/GRE there is no port-like selector to NAT
by in the first place.
- **Rejected 2026-07-26: carrying them through the userspace AWG endpoint
site-to-site, with no NAT at all.** This is the alternative the "no port-like
selector" line above does NOT dispose of, and it is written down because the
obvious reading of that line — "impossible" — is wrong and would be
re-derived. The endpoint is already protocol-blind in both directions
(`transport/wireguard/port.go:21-58`, `:127-157`), so an ESP packet could be
forwarded UNTOUCHED, keeping the LAN client's own source address, and the
reply would come back addressed to that client and need only be written to the
TUN. No selector, no NAT, every protocol. It needs two things we declined to
take on: lx-owned code in the forward hot path, bypassing `ForwardDispatcher`
on both legs — precisely the surface CONSTITUTION §2 exists to keep small on
an actively-maintained upstream — and a SERVER-side prerequisite (our LAN
prefix in the peer's `AllowedIPs`, plus a route back), which turns a router
option into a deployment contract. The kernel egress above buys the same
protocols with zero hot-path code, so this stays a design on file, not a gap.
- **Rejected: a dedicated mark/table pair for the carrier.** `addEgressRouting`
already binds `EgressMark`/`EgressTable` to the device; a third pair would be
a second copy of the same route that could drift from the first.
**Not verified, said plainly.** The end-to-end path — LAN client → prerouting
mark → ip rule → egress device → far end and back — has not been exercised with
real ESP or GRE on live hardware. Nothing above claims it has.
Consequence: raw IPsec, PPTP/GRE, SCTP — and ICMP when the L3 ingress is off —
leave through an egress the operator explicitly named, under kernel routing and
kernel NAT, instead of being dropped or silently leaking out the WAN; and with
the option empty (the default) the plane renders byte-for-byte as before, with
the D17 policy in sole charge.
## D27 — The `sing-quic` pin moves forward; two use-after-release defects in the QUIC/HTTP-3 client path
Decided 2026-07-26, after an audit of `common/httpclient`, `dns/transport/quic`
and `transport/v2rayquic`. Two suspicions were put to a test rather than to a
reading. Both turned out to be real, and neither was the resource leak the
suspicion named — both are objects released while still in use.
**1. The HTTP/3 race handed back a response nobody could read
(`common/httpclient/http3_transport.go`).** `roundTripHTTP3Race` ran both racers
on one `context.WithCancel` child and its `drainRemaining()` called `cancel()`
before returning the WINNER. quic-go and net/http both reset a request's stream
when its context is cancelled, so the caller received a `*http.Response` whose
body died mid-read. Measured, not inferred:
`H3_REQUEST_CANCELLED (local) (read 2687 of 65536 bytes)`. The path is taken
whenever there is no cached HTTP/3 connection and the request is replayable —
that is, the FIRST request to every host, plus every request after an idle
close. Anything configured with `"version": 3` and no
`disable_version_fallback` was affected: subscription fetches, remote rule-set
downloads, URLTest probes.
Fixed by giving each racer a context of its own. Losers are cancelled where the
old code cancelled everything; the winner's `cancel` travels with its body and
fires on `Close`. Pinned by
`common/httpclient/http3_race_lx_test.go` — one test per winner, and the
loser-is-torn-down assertion so the fix cannot be "stop cancelling" either.
**How defect 2 is pinned, and why it takes two tests.** The success path is
visible on the wire, so `TestHTTP3ExchangeRequestBufferOutlivesRoundTrip`
compares the query the server received with the query we packed. The failure
path is NOT visible on the wire — see the `CancelWrite` note below — so
`TestHTTP3ExchangeNeverPacksQueriesIntoPooledMemory` pins the CAUSE instead of
the symptom: it tags a query with a random needle, runs an exchange that fails
(a server that never answers; a context already cancelled) and then drains the
buffer pool on the same goroutine `RoundTripOpt` ran on, demanding the needle is
not in it. Both carry a control that must produce a DIRTY result first — the
poison must be shown to reach a released buffer, the scan must be shown to find a
query that really was packed into pool memory — because a clean verdict from an
instrument that cannot produce a dirty one is not evidence. The second test
encodes the chosen design; a future guarded-pool implementation would fail it on
purpose and would have to replace it, saying so.
**2. DoH3 sent the DNS server whatever the next caller put in a recycled buffer
(`dns/transport/quic/http3.go`).** `Exchange` packed the query into a POOLED
`buf.Buffer`, handed `bytes.NewReader` over it to the request, and called
`requestBuffer.Release()` the moment `RoundTrip` returned. But http3's
`doRequest` writes the request body on a goroutine of its own and returns as
soon as the response HEADERS arrive — the body is still being read. The test
holds that window open (a 2 KB server stream window, an answer written before
the body is read) and shows the query on the wire diverging from the query we
packed **at exactly offset 8192** — `bodyCopyBufferSize`, the amount quic-go had
already copied out before the buffer went back to the pool. Everything past that
was the next pool user's memory, sent to the resolver. That is a data race and a
small memory-disclosure primitive, not a slowdown.
**The first fix for this was wrong, and the way it was wrong is the point.** It
transferred ownership to the transport: the body became a `pooledRequestBody`
whose `Close` released the buffer, justified as "`http3.Transport` closes the
request body on every path, hence the `sync.Once`". That sentence is true about
HOW MANY TIMES the body is closed and says nothing about WHEN — the exact shape
of dishonesty this document keeps having to name. Review caught it before
release. On the failure path `RoundTripOpt` (`http3/transport.go:167-173`) closes
the body the moment `doRequest` returns, and `doRequest`
(`http3/client.go:338-341`) waits only on the request-CANCELLATION watchdog —
`close(reqDone); <-done` — never on the goroutine writing the body. Nothing in
quic-go ever joins that goroutine. So `Close` is not a handoff point, and the
`sync.Once` prevented a double `Release` while doing nothing about a read after
one.
**One correction to the review's severity, for the record:** on the failure path
the damage stops at the data race. Every `ReadResponse` error branch
(`http3/stream.go:325`, `:336`, `:343`, `:363`) calls `str.CancelWrite` BEFORE
`RoundTripOpt` closes the body, so the bytes the writer reads out of the recycled
buffer are thrown at an already-cancelled stream and never reach the resolver.
The memory-disclosure primitive is the SUCCESS path only. The failure path is
"merely" a read of memory owned by somebody else — still undefined behaviour,
still a `-race` finding, still not shippable.
**Fixed instead by not sharing at all:** the query is packed with `Pack()` into
memory the request body owns outright, and no pooled buffer is involved. The
alternative on the table — a lock around `Read` and `Close` so reads after
release return an error — would also be correct, and was rejected because it
keeps a released-but-referenced object alive and leaves a live invariant for the
next person to break, which is now twice in one day that an assumption about
quic-go's internal lifetimes has been wrong.
The cost turned out to be negative, measured rather than assumed: `Pack` runs
87 ns/op at 64 B and 1 alloc against 108 ns/op at 64 B and 1 alloc for the pooled
version, because `buf.NewSize` allocates the `Buffer` struct itself — the same 64
bytes — and then adds `Get`/`Put` on top. **The pool was never saving an
allocation on this path.** The response buffer stays pooled: it is read and
unpacked before `Exchange` returns and nothing outlives it.
Both files thereby DIVERGE from upstream again, six hours after `0a6689b29`
made them byte-identical on purpose. That was the right call then and this is
the right call now; upstream carries defect 2 in `dns/transport/https.go` as
well (same shape, HTTP/1.1 and HTTP/2 write bodies asynchronously too) and that
file was left alone — it is outside the audit's scope, and it is written down
here so the next person finds it instead of rediscovering it.
**3. The pin moved: `sing-quic` v0.6.2-0.20260525051024 -> v0.6.4-0.20260709034545.**
`quic.go` — `Dial`/`DialEarly`/`CreateTransport` — is byte-identical across the
two, so the packet-conn ownership fix in `0a6689b29` is NOT duplicated by the
bump and is not made redundant by it: quic-go still does not own the socket, and
we still close it. What the newer module does carry is the OTHER half of the
same family, and we had taken only our half:
`clientConn.Close()` in `tuic/`, `hysteria/` and `hysteria2/` now sets a past
write deadline after `Stream.Close()`, word for word the fix
`transport/v2rayquic/stream.go` already had — quic-go's `Stream.Close` does not
release a write blocked on flow control. We ship tuic and hysteria2, so on the
old pin every such close could park a goroutine for the life of the process.
It also brings a QUIC-connection-death watchdog and a handshake deadline to the
hysteria clients.
**The cost, measured, not estimated:** six new indirect modules (`libp2p/go-nat`
and its UPnP/NAT-PMP/gopacket tail) for hysteria2's realm port mapping, and
**+256 KiB exactly** on the stripped aarch64 `shaterd` (26 542 242 ->
26 804 386 bytes, +0.99%), before UPX. The port-mapping path is unreachable from
anything `shater/generate` emits — `realm` is only built when the JSON names it,
and it never does — so the growth is dead weight, but it is small dead weight
next to a goroutine leak on the two QUIC protocols we actually ship. `upstream/lx`
is already on this pin, so keeping the old one would mean fighting every rebase.
**Not verified here:** the tuic/hysteria2 close fix is read from the module diff,
not exercised — those packages are outside this audit's file set and testing them
needs a live tuic/hysteria2 server. `go test ./common/... ./dns/... ./transport/...
./protocol/...` is green under the shipped tag set apart from
`common/windivert`'s `TestIntegration*`, which want Windows SCM access and fail
on any developer machine, bump or no bump.
## D28 — The L3 TUN is a reclaimable SLOT, not a name: a fixed device made every apply fatal
Decided 2026-07-26. D25 gave the L3 ingress one fixed device, `shater-l3`. On the
production router that single name made **every** configuration change with
`l3_tunnel=1` an outage, three times in a row:
```
19:10:33 reconcile failed: start inbound/tun[l3-in]: open tun: TUNSETIFF: device or resource busy
19:14:02 start instance failed and could not restore previous config; engine stopped
19:14:38 reconcile failed: TUNSETIFF: device or resource busy
```
followed by `plane: hold` — the fail-closed ruleset — i.e. the whole house
offline until somebody ran `/etc/init.d/shater restart` by hand.
**The mechanism is a collision between two GENERATIONS of the engine, and the
fatal part is where the collision lands.** An apply builds a fresh box and starts
it; with one name that box must open the device the outgoing box still holds.
That alone is a failed apply. What turned it into an outage is the recovery path:
`closeOldThenStart` answers a failed start by rebuilding the PREVIOUS config —
and that config names the same device, so **the rescue failed for exactly the
reason the rescue was needed**. A recovery path must never depend on the resource
whose contention it is recovering from; that sentence, not the device name, is
the decision here.
**Two slots (`shater-l3a` / `shater-l3b`), chosen by the ENGINE at box-build
time.** Not by `generate`, and that is load-bearing rather than incidental:
`generate` runs on every reconcile including the once-a-minute no-ops, and its
output is what `Apply` hashes to decide whether anything changed. A device name
that alternated there would change the hash every minute and rebuild the whole
engine forever; a name that tracked "whichever slot exists" would hand the BUSY
one to every real change. Only the engine knows it is building a new generation.
So `generate` emits `netplane.L3DeviceBase` as a **placeholder that is never
created**, and `engine.newBox` substitutes a slot on a COPY of the options —
after the hash, so the stored config stays canonical (`engine/l3slot.go`).
**Rotation alone is NOT the fix, and that was measured, not reasoned.** The
two-slot build survived five applies of five different kinds and then failed on
4 of 10 back-to-back changes with the original outage in full. A retired
generation does not hand its device back when its replacement is adopted:
`Box.Close` walks its subsystems under a budget and the TUN dies with the last
fd. One apply gives it time; two inside that window do not. Rotation widens the
race by one generation — a better outage, not the absence of one.
**So an occupied non-current slot is DELETED, not waited for** (`L3SlotFor`).
The running generation's slot is excluded first and never touched; every other
slot belongs to a retired generation that is not in the L3 routing table and is
carrying nothing, so taking its device away is safe and, if anything, helps the
close already in flight. There is deliberately **no bounded wait**: waiting on an
asynchronous kernel teardown is the race this design removes, and adding it back
as a "safety net" would only make the failure intermittent.
**The firewall never learns which slot is live.** Our forward accepts and the fw4
zone match by PREFIX — `iifname "shater-l3*"` / `oifname "shater-l3*"`, and
`list device 'shater-l3*'` in uci-defaults. Verified on ImmortalWrt 25.12.1
(kernel 6.12.94, nftables 1.1.6) that both forms validate AND load, and that fw4
compiles the wildcard into exactly those matches. So the rendered ruleset is
byte-identical across a swap: no nft reload, and no window in which the accept
names a device that is already gone. Routers seeded by a pre-slot build are
migrated in place (`migrate_l3_zone_wildcard`); without it an upgrade would
silently go back to fw4 dropping the forward.
**A2 — turning the feature off used to leave the plane installed.** `addL3Routing`
returned early when `l3_tunnel=0`, so after switching it off the router still had
the device, `ip rule fwmark 0x2080 lookup 8200` and table 8200. Nobody else
removes them: the engine's new config simply has no TUN inbound. The surviving
device is also the commonest way back into the EBUSY above. The disabled branch
now removes rule, table and device, symmetrically with `removeEgressRouting`, and
`TeardownRouting` sweeps every name including the legacy `shater-l3`.
**Two smaller lies found while proving the above, both measured on the stand and
both fixed here.** `ip -6 route flush table N` does not remove a non-unicast
route while the v4 flush does, so (a) the floor survived the flush and the next
`add` answered `File exists`, which was reported as a CRITICAL "this table has NO
fail-closed floor, traffic can leave over the plain WAN" — on every single apply,
about a table whose floor was sitting right there; and (b) teardown left that
floor behind. An already-present floor is now success, and teardown deletes it
explicitly. (b) cannot misroute anything — no rule points at the table — but a
teardown whose result is not "the table prints nothing" is one nobody can verify.
**Verified on the stand** (`local_openwrt`, ImmortalWrt 25.12.1, kernel 6.12.94 —
the router's revision), before-and-after with binaries built from the same tree:
the pre-fix binary reproduces the production failure on the "edit a node URI and
apply" round (engine stopped, `plane: hold`) and leaves device + rule behind on
`l3_tunnel=0`; the fixed binary survives all five apply kinds, 12 back-to-back
changes, and leaves nothing at all — no device, no rule on either family, both
tables printing empty.
**Known fragility, deliberately NOT addressed here.** Our `ip rule` priorities
are whatever the kernel hands out (the L3 rule lands at 32764, counting down from
32765), so the ORDER of our rules between reboots depends on insertion sequence.
It is not a live defect — the marks are disjoint, each rule catches its own, and
all of them end up above `main` — but it is luck, not design. Moving to explicit
`pref` values touches every existing rule and needs a migration for rules already
installed on implicit numbers; that is its own piece of work, not a rider on this
one.
+39 -3
View File
@@ -6,12 +6,44 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Proxy engine & protocols (from the sing-box fork)
- **[MVP]** VLESS, VMess, Trojan, Shadowsocks, WireGuard, Reality/XTLS.
- **[MVP]** **AmneziaWG 2.0** (I1–I5 CPS decoy packets) — a driving requirement.
- **[T1]** Hysteria2, TUIC, ShadowTLS, XHTTP, MASQUE/CONNECT-IP (Cloudflare WARP).
- **[MVP]** Hysteria2, TUIC (`hysteria2://`/`hy2://`/`tuic://`, `shater/parse`),
XHTTP transport — all shipped: the router tag set carries `with_quic` and
`with_xhttp` and `shater/registry` registers them (`scripts/router-tags.sh`,
`buildtags.Features`).
- **[T1]** ShadowTLS — half-built: `shater/generate` emits it and `shater/registry`
registers it, but no parser produces one (there is no `shadowtls://` share link
and no subscription path), so a config cannot reach it today.
- **NOT SHIPPED** MASQUE/CONNECT-IP (Cloudflare WARP). `masque` appears nowhere in
`shater/parse`, `shater/generate` or `shater/model`, and `shater/registry` names
it among the upstream types it deliberately does not register (~6 MB of binary
and resident RAM). The engine fork can build it; this product does not.
- **[MVP]** Transports: TCP/WS/gRPC/HTTPUpgrade/H2/QUIC as upstream provides.
## Transparent proxying & routing
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** **L3 ingress for ICMP** (`globals.l3_tunnel`, opt-in, default off):
LAN ping travels THROUGH the tunnel instead of being dropped or answered by a
forged local reply. The engine opens a dedicated TUN (`shater-l3`, gVisor
stack, `auto_route` off); nft marks LAN icmp/icmpv6 only and a scoped
`ip rule` routes it in — the TPROXY plane and the main routing table stay
untouched (D25). Carried only by L3-capable egresses (WireGuard/AmneziaWG,
direct); ICMP routed to vless/vmess/… is honestly dropped, never faked.
Ceiling is upstream sing-tun's: ICMP echo only — Windows tracert works, IPv6
traceroute shows just the destination; ESP/AH/GRE/IGMP stay with the
`untunnelable` policy (D17) unless `untunnelable_egress` carries them (D26).
- **[MVP]** **Kernel egress for untunnelable protocols**
(`globals.untunnelable_egress`, opt-in, default empty): names an existing
interface/tunnel egress, and IPsec (ESP/AH), PPTP/GRE, SCTP — everything that
is neither TCP nor UDP, plus ICMP when the L3 ingress is off — is routed out
that egress's device by the KERNEL with kernel NAT, reusing the egress's own
fwmark/table from `addEgressRouting`; the proxy never sees a byte, which is
why every protocol works (D26). What that buys depends on the device: a
WireGuard interface really is a tunnel, a second WAN is just another uplink
whose real address the destination sees. It does not revive multicast IPTV,
and UDP-based VPNs (WireGuard, OpenVPN-UDP, IPsec NAT-T) never needed it —
they follow the routing rules as before. The `untunnelable` policy (D17)
keeps only the failure case: a route that did not come up.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
@@ -87,8 +119,12 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Reliability ("железно")
- **[MVP]** Fail-closed kill-switch (dead group → block, never silent direct leak);
IPv6 dropped when disabled.
- **[MVP]** Atomic apply with engine + `nft -c` validation; commit-confirm
auto-rollback to last-good.
- **[MVP]** Atomic apply with engine + `nft -c` validation. Commit-confirm
auto-rollback to last-good is built and works, but it is **opt-in and ships
OFF**: `DefaultGlobals()` leaves `ConfirmTimeout` at 0, the shipped
`/etc/config/shater` says `confirm_timeout '0'`, and `apply.ArmRollback` returns
at once on a non-positive timeout. Until an operator sets a window, an apply on
a stock box has no net under it — and `shaterd apply` says so.
- **[MVP]** Idempotent reconcile from hotplug/boot under flock; restart engine only
on real config change; management-bypass (SSH/LuCI/LAN) always exempt.
- **[MVP]** Own nft table `inet shater` + own marks/tables; never touch fw4.
+79 -8
View File
@@ -88,7 +88,7 @@ Four OpenWrt packages live under `openwrt/`:
| Package | Arch | What it ships |
|--------------------|-----------|---------------|
| `shaterd` | per-arch | **Prebuilt** static `shaterd` binary → `/usr/bin/shaterd` (this is the ship artifact from step 1). |
| `shater-core` | all | procd init (supervises `shaterd run`), cron, hotplug, sysctl, inert default UCI. `DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full`. |
| `shater-core` | all | procd init (supervises `shaterd run`), the boot armor (§4), cron, hotplug, sysctl, inert default UCI. `DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle`. |
| `luci-app-shater` | all | Thin LuCI launcher: mini dashboard + token-handoff "Open panel" button. `DEPENDS:=+shater-core +rpcd`. |
| `byedpi` | per-arch | *Optional* ByeDPI (`ciadpi`) local desync SOCKS proxy for a `type='byedpi'` egress. |
@@ -176,11 +176,27 @@ install. Configure nodes/rules (via the LuCI panel or `uci`), then enable and ap
```sh
uci set shater.globals.enabled=1
# The safety net is NOT on by default — see below. 120 s is a window wide enough
# to re-open SSH/LuCI and decide whether the new config is any good.
uci set shater.globals.confirm_timeout=120
uci commit shater
shaterd apply # apply + arm commit-confirm on the running daemon
shaterd confirm # confirm (cancels the auto-rollback)
shaterd apply # apply + arm the auto-rollback for 120 s
shaterd confirm # confirm inside that window (cancels the auto-rollback)
```
> **Commit-confirm ships OFF.** `model.DefaultGlobals()` does not seed
> `ConfirmTimeout`, the shipped `/etc/config/shater` carries
> `option confirm_timeout '0'`, and `apply.ArmRollback` returns immediately on a
> non-positive timeout — so on a stock box `shaterd apply` arms **nothing** and an
> apply that costs you SSH/LuCI access simply stays. The daemon says so rather
> than implying otherwise: the `commit-confirm-off` outcome of `shaterd apply`
> prints *"globals.confirm_timeout is 0, so commit-confirm is switched OFF: this
> apply armed NO automatic rollback"*, and the panel's Overview reads
> `confirm: no auto-rollback`. Set a window (UCI as above, or Settings in the
> panel) if you want the net. Non-obvious detail: the option is written back only
> when non-zero, so an explicit `0` disappears from `/etc/config/shater` on the
> first write — absent and `0` mean the same thing.
`/etc/init.d/shater enable && /etc/init.d/shater start` brings up the procd-supervised
daemon (`shaterd run`), which owns the engine, the `inet shater` data plane, policy
routing, in-process DNS, and the admin panel (default `:8088`). The LuCI app's
@@ -218,6 +234,55 @@ Your `0` is kept: `/etc/config/shater` is a conffile (upgrades never replace it)
the daemon always writes the option back explicitly, so it is never re-enabled by a
default.
### The boot-time fail-closed armor
`shater-core` installs a **third** init script, `/etc/init.d/shater-armor`, and
`30_shater-core` enables it at install time. It exists because `/etc/init.d/shater`
is `START=99`: by then fw4 (19) has loaded `lan -> wan ACCEPT` and netifd (20) has
brought the LAN bridge up, so between link-up and the daemon's first apply the
router forwards LAN traffic to the WAN in the clear — on router hardware with a
UPX-packed binary that is the seconds in which Wi-Fi associates and every client
reconnects. `kill_switch=closed` covered none of it, because the protection lived
inside a process that had not started.
**How it works.** On every apply the daemon persists a copy of its fail-closed
*holding plane* — the same ruleset it installs when the engine is down — to
`/etc/shater/boot.nft`. `shater-armor` runs at `START=21` (after fw4 and netifd),
validates that file with `nft -c` and loads it. When the daemon comes up it
replaces the table atomically, so there is never a moment with no table. Its
`stop()` is deliberately a no-op.
**LAN forwarding is blocked until the daemon applies — management access is not.**
The chain hooks `forward` only, so SSH, LuCI and the admin panel (all `input` hook,
to the router's own addresses) stay reachable **on purpose**: a kill switch you
cannot switch off is a brick. If you see the syslog line
```
fail-closed plane armed from /etc/shater/boot.nft: LAN->WAN forwarding is BLOCKED
until shaterd applies. SSH, LuCI and the admin panel stay reachable.
```
that is the mechanism working, not a fault.
**When it refuses to arm** — each is a state check made at boot, never a record of
something that happened on the way down:
| Condition | Behaviour |
|---|---|
| `/etc/shater/boot.nft` absent | Nothing to do, silent. The file exists only while the last applied config was **both** `enabled=1` **and** `kill_switch=closed`; either being off removes it at the next apply, and an operator-typed `/etc/init.d/shater stop` removes it there and then. Powering off does **not** — and neither does the `stop` a package upgrade issues while the service stays enabled, so being replaced cannot leave the next boot unprotected. |
| the file is empty, or fails `nft -c` | Refuses, logs an error — the LAN is unprotected until `shaterd` starts. |
| `/usr/bin/shaterd` missing, or no `S??shater` symlink in `/etc/rc.d` | Refuses: nothing would ever come along to replace the block with a working data plane. This is what makes an uninstalled or disabled product safe regardless of what the file says. |
| UCI is readable **and** says `globals.enabled` is not `1` | Removes `boot.nft` and does not arm. An **unreadable** UCI is not a refusal — that case is exactly why the armor is a file rather than a query. |
| `nft` not installed | Refuses, logs an error. |
**Turning it off.** The durable off-states are the two the script itself asks
about — `uci set shater.globals.enabled=0 && uci commit shater && shaterd apply`
(the next apply removes `boot.nft`), or `/etc/init.d/shater disable`. A bare
`/etc/init.d/shater stop` typed at the shell also removes the file, but it is not
durable: `S99shater` is still linked, so procd starts the daemon again on the next
boot. To remove just the armor and keep the stack: `/etc/init.d/shater-armor
disable`.
## 5. The signed apk repo (the normal install path)
OpenWrt/ImmortalWrt **25.12** packages with Alpine's **apk**: `.apk` files, a
@@ -321,8 +386,14 @@ The mtk-vendor channel (base: `SuperKali/immortalwrt-mt798x-rebase`, branch
`downloads.immortalwrt.org/releases/25.12-SNAPSHOT` — so packages built with the
vanilla ImmortalWrt 25.12 filogic SDK install cleanly; no SuperKali-special SDK
is needed. We ship **no kmods** (shaterd is a static Go binary, byedpi plain C),
so the vendor 6.6 kernel is irrelevant to our packages; the kmod *dependencies*
of shater-core (`kmod-nft-tproxy`, `kmod-nft-socket`, plus `ip-full`) are
already **baked into the BananaWRT mtk-vendor image** (verified in its
`config.buildinfo`). On a self-built 25.12 image, make sure those kmods come
from the image's own kernel build.
so the vendor 6.6 kernel is irrelevant to our packages.
What was actually checked in the BananaWRT mtk-vendor `config.buildinfo` is
`kmod-nft-tproxy`, `kmod-nft-socket` and `ip-full` — those three are baked into
the image. `shater-core` also depends on `kmod-tun`, `nftables-json` and
`ca-bundle` (added later; see the annotated `DEPENDS` in
`openwrt/shater-core/Makefile`), and **those were not part of that check**. They
are ordinarily present on a stock image — apk will pull whatever is missing from
the distfeeds — but if you install offline or from a slimmed image, verify them
yourself. On a self-built 25.12 image, make sure the kmods come from the image's
own kernel build.
+24 -5
View File
@@ -273,10 +273,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
| `block_doh` | `0` | NXDOMAIN the known public DoH hostnames + the Firefox canary and reject `:443` to their IPs, so clients fall back to `:53` (which the engine catches) |
| `group_health` | `1` | OUR background group probing (the observatory). Does not touch sing-box's own urltest inside a group |
| `untunnelable` | `block` | policy for what TPROXY cannot carry (ICMP/IGMP/ESP/AH/GRE/SCTP): `block` \| `icmp` (echo out, rest dropped) \| `direct` (all out, bypassing the tunnel) |
| `l3_tunnel` | `0` | **opt-in**, UCI-only (the panel does not expose it). Opens the synthetic `l3-in` TUN so LAN ICMP is routed by the engine instead of dropped/forged; nft marks LAN `icmp`/`ipv6-icmp` with `fwmark_base+0x80` and a scoped `ip rule` sends it to table `table_base+8`. Absent option ⇒ OFF; only an explicit `1` opens it. See D25 and `ARCHITECTURE.md` §3a |
| `untunnelable_egress` | unset | **opt-in**, UCI-only. Names a `config egress`; everything the L3 block did not claim (ESP/AH, GRE, IGMP, SCTP, and ICMP when `l3_tunnel=0`) is stamped with that egress's OWN mark and routed out its device by the kernel — no new mark, no new table, engine not in the path. Empty ⇒ `untunnelable` above stays in sole charge (D26) |
| `geo_provider` | unset = auto | `sagernet` \| `loyalsoldier` \| `metacubex` \| `custom`; auto = country codes from SagerNet, everything else from Loyalsoldier |
| `geosite_url` / `geoip_url` | unset | `{category}` templates, honoured only when `geo_provider=custom` |
| `geosite_index_url` / `geoip_index_url` | unset | git-trees URLs used to SUGGEST categories in the panel; empty = no suggestions |
| `stats_backend` | `memory` | `off` (no aggregation at all) \| `memory` (RAM, lost on restart) \| `sqlite` (aggregates in RAM + query/connection log on disk) |
| `stats_backend` | `memory` | `off` (no aggregation at all) \| `memory` (RAM, lost on restart) \| `sqlite` (aggregates in RAM + query/connection log on disk). The value NAME is historical: the on-disk store is **bbolt**, not SQLite, since the migration — a leftover sqlite-era `stats.db` is detected by its file magic and replaced (`shater/stats/boltring.go`) |
| `stats_ring_size` / `stats_timeline_minutes` / `stats_max_domains` | `200` / `60` / `5000` | live-log length, sparkline minutes, domain-map cap. **`0` = UNLIMITED** (grows with traffic), which is why these three are always emitted |
| `stats_disk_limit_mb` | `64` | on-disk cap of `stats.db`; only meaningful for `stats_backend=sqlite`; `0` = unlimited |
| `stats_retention_disabled` | `0` | master switch that turns OFF all trimming/pruning — every aggregate then grows unbounded |
@@ -285,7 +287,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
Deleted options still parse (unknown keys are ignored) and drain out on the next
render: `dns_mode` (D17 — fake-IP is a resolver TYPE), `sweep_interval` (D19).
- `config inbound`: name, enabled, type, network, tproxy_port(12345), listen, port, auth, user, pass, target_addr, target_port, target_network, tcp, udp, sniff.
- `config inbound`: name, enabled, type, network, tproxy_port(12345), listen, port, auth, user, pass, target_addr, target_port, target_network, tcp, udp.
**No `sniff`.** Since sing-box 1.11 sniffing is a leading route ACTION rule with no
inbound matcher, so every inbound is sniffed always; the flag was read by nothing but
its own UCI round-trip. Re-adding it would be a regression, not a restored feature —
the hijack-dns rule matches the SNIFFED `dns` protocol, so a per-inbound toggle is a
DNS-leak switch wearing a performance label (`model.go`, `Inbound`).
- `config subscription`: name, enabled, url, update_interval, fetch_via(direct|proxy), ua, hwid, device_os, ver_os, device_model, list header, format, list include/exclude/filter_proto/filter_country, dedup, expire_alert_days.
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
@@ -296,8 +303,19 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
- `config blocklist`: name, enabled, source(inline|file|url|geosite), url, path, list category, list entry, response(nxdomain), update_interval.
- `config allowlist`: the same minus `response` (an allowlist has no verdict to render); it overrides every blocklist.
- `config device`: name, mac, ip, enabled, list block, list allow.
- `config alert`: name, enabled, type(telegram), token, chat_id, url, list event, via, fallback.
- **`config preset` is NOT a section type.** `ReadUCI`'s type switch has no `preset`
branch, so such a section is parsed by nothing and reaches no part of the model.
It survives only because `30_shater-core` still seeds three of them
(`block_ads`, `ru_bypass`, `private`) with the comment "so the LuCI Rules page
renders their toggles" — and v0.2's LuCI app is a thin launcher with no Rules
page. Those `uci set` calls should be dropped from the uci-defaults script; until
they are, three inert sections appear in every fresh `/etc/config/shater`.
### Subscriptions & HAPP fetch
Schemes: `vless:// vmess:// trojan:// ss:// wireguard:// wg://`. Body formats (`DetectSubFormat`): clash-YAML, xray-JSON, singbox-JSON, base64/plain link list. All converge to URIs re-parsed by `ParseShareLink`. HAPP fetch: UA default `Happ/3.13.0`; headers `x-hwid` (auto UUIDv4/sub), `x-device-os`, `x-ver-os`, `x-device-model`, custom. `fetch_via=proxy` dials local socks. Quota/expiry from `Subscription-Userinfo` (`upload;download;total;expire`). Reconcile by `Fingerprint` (sha256 of proto|addr|port|id|net|sec|sni|path) → new/keep/stale (drop after 3 stale refreshes).
@@ -305,10 +323,11 @@ Schemes: `vless:// vmess:// trojan:// ss:// wireguard:// wg://`. Body formats (`
---
## PART B — v0.1 packaging (`shater-core/`, branch `v0.1`)
Pure scripts+config, `PKGARCH:=all`. v0.1 DEPENDS: `+xrayctl +xray-core +dnsmasq-full +kmod-nft-tproxy +kmod-nft-socket +ip-full`. → **v0.2 deps: `+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full`** (engine does DNS in-process, so dnsmasq-full may be droppable — confirm the :53 listener is our engine). `/etc/config/shater` is a conffile.
Pure scripts+config, `PKGARCH:=all`. v0.1 DEPENDS: `+xrayctl +xray-core +dnsmasq-full +kmod-nft-tproxy +kmod-nft-socket +ip-full`. → **v0.2 deps (authoritative: `openwrt/shater-core/Makefile`, which annotates each one): `+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle`** — `dnsmasq-full` is gone (the engine owns the `:53` hijack listener); `kmod-tun` is `/dev/net/tun` for the L3 ingress, `nftables-json` is the `nft -j` output `netplane/stats.go` parses, `ca-bundle` is the cert store a `CGO_ENABLED=0` binary has no host fallback for. `/etc/config/shater` is a conffile.
- **init.d/shater** (procd, START=99/STOP=10): v0.1 supervised `xray run -c /etc/xray/run.json`; → v0.2 supervises `shaterd`. `respawn 3600 5 0` (infinite). **No `procd_set_param file` watch** (would bounce tunnel on commit). Inert unless `globals.enabled=1`. `ACTIVE_FLAG=/var/run/shater.active` gates hotplug/cron. `stop` clears flag + tears down nft table + reserved routing tables. `reload_service`→start/stop. trigger `procd_add_reload_trigger "shater"`.
- **init.d/shater-cron** (START=96): supervised `loop`; per-item due-check, runs sub/ruleset update + reconcile + schedule due; watchdog: engine dead 5 ticks ⇒ kill_switch=open stops stack (fail-open), closed logs crit.
- **uci-defaults/30_shater-core**: seed `rt_tables` (8192 shater), enable both inits, seed preset packs (disabled), run migrate, apply sysctl.
- **init.d/shater-armor** (START=21/STOP=89, v0.2-only — no v0.1 counterpart): the fail-closed plane BEFORE the daemon exists. `/etc/init.d/shater` is START=99, so from netifd's `ifup` until the daemon's first apply the router forwarded LAN→WAN in the clear. The daemon persists its holding plane to `/etc/shater/boot.nft` on every apply; this loads it after fw4 (19) and netifd (20), `nft -c`-validated. Four state checks refuse to arm (no/empty/invalid file, missing `shaterd`, no `S??shater` rc-link, readable UCI saying `enabled≠1`) — asked ON THE WAY UP, deliberately not recorded on the way down. Hooks `forward` only, so SSH/LuCI/panel stay reachable. `stop()` is a NO-OP. Operator-facing writeup: `INSTALL.md` §4.
- **uci-defaults/30_shater-core**: seed `rt_tables` (8192 shater), `mkdir /etc/shater`, seed the `shater_l3` fw4 zone + `lan→shater_l3` forwarding (named sections, `list device 'shater-l3*'`) and migrate a legacy exact-name entry to the wildcard, run `shaterd migrate`, apply sysctl, then a DETACHED bring-up (enable+restart `shater`/`shater-cron`, enable `shater-armor`, conditional `firewall reload`) — detached because an inline init call inside an apk/opkg transaction deadlocks on procd's flock. Also still seeds three `config preset` sections, which nothing parses (see the schema note above); those calls should go.
- **hotplug.d/iface/99-shater**: ifup/ifdown → debounced (2s) `reconcile` (netifd wipes ip rules on reload). Guarded by enabled + ACTIVE_FLAG.
- **sysctl.d/99-shater.conf**: `ip_forward=1`, `rp_filter=0` (all+default), `lo.route_localnet=1`, `lo.accept_local=1`, `all.src_valid_mark=1`, `ipv6.all.forwarding=1`.
+2 -2
View File
@@ -49,7 +49,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
the gate: fail-closed forward drop (4f618140), engine apply-swap close-first
fallback (9b6b9406), DNS hijack-dns per D14 (86194ce6).
## Phase 2b — DPI-bypass egress = ByeDPI (D13)
## Phase 2b — DPI-bypass egress = ByeDPI (D13) ✅ DONE
- The one external desync tool is **ByeDPI (ciadpi)** — chosen over zapret because
it *is* a SOCKS egress (fits shater's "routing picks the egress" model with zero
packet-plane conflict); zapret is explicitly rejected (see D13).
@@ -86,7 +86,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
well-known lists (StevenBlack/OISD/AdGuard).
- **Gate:** ad/tracker domains blocked network-wide; big list loads fast; RAM sane.
## Phase 5 — Statistics (per-domain / client / device)
## Phase 5 — Statistics (per-domain / client / device) ✅ DONE
- Stats aggregator consuming the engine's DNS/routing/stats observability + nft
counters: top domains, allowed vs blocked, per-device breakdown, timelines,
per-node/per-rule traffic, live query log with one-click block.
+7 -1
View File
@@ -46,7 +46,7 @@ require (
github.com/sagernet/sing v0.8.12-0.20260702081104-2ded2af32d3d
github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3
github.com/sagernet/sing-mux v0.3.5
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc
github.com/sagernet/sing-shadowsocks v0.2.8
github.com/sagernet/sing-shadowsocks2 v0.2.1
github.com/sagernet/sing-shadowtls v0.2.1
@@ -104,14 +104,20 @@ require (
github.com/google/btree v1.1.3 // indirect
github.com/google/go-cmp v0.7.0 // indirect
github.com/google/go-querystring v1.1.0 // indirect
github.com/google/gopacket v1.1.19 // indirect
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806 // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/hashicorp/yamux v0.1.2 // indirect
github.com/hdevalence/ed25519consensus v0.2.0 // indirect
github.com/huin/goupnp v1.2.0 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/jackpal/go-nat-pmp v1.0.2 // indirect
github.com/klauspost/compress v1.18.0 // indirect
github.com/klauspost/cpuid/v2 v2.3.0 // indirect
github.com/koron/go-ssdp v0.0.4 // indirect
github.com/kr/fs v0.1.0 // indirect
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138 // indirect
github.com/libp2p/go-netroute v0.2.1 // indirect
github.com/mdlayher/socket v0.5.1 // indirect
github.com/mitchellh/go-ps v1.0.0 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
+26 -2
View File
@@ -101,6 +101,8 @@ github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/google/go-querystring v1.1.0 h1:AnCroh3fv4ZBgVIf1Iwtovgjaw/GiKJo8M8yD/fhyJ8=
github.com/google/go-querystring v1.1.0/go.mod h1:Kcdr2DB4koayq7X8pmAG4sNG59So17icRSOU623lUBU=
github.com/google/gopacket v1.1.19 h1:ves8RnFZPGiFnTS0uPQStjwru6uO6h+nlr9j6fL7kF8=
github.com/google/gopacket v1.1.19/go.mod h1:iJ8V8n6KS+z2U1A8pUwu8bW5SyEMkXJB8Yo/Vo+TKTo=
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806 h1:wG8RYIyctLhdFk6Vl1yPGtSRtwGpVkWyZww1OCil2MI=
github.com/google/nftables v0.2.1-0.20240414091927-5e242ec57806/go.mod h1:Beg6V6zZ3oEn0JuiUQ4wqwuyqqzasOltcoXPtgLbFp4=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
@@ -109,10 +111,14 @@ github.com/hashicorp/yamux v0.1.2 h1:XtB8kyFOyHXYVFnwT5C3+Bdo8gArse7j2AQ0DA0Uey8
github.com/hashicorp/yamux v0.1.2/go.mod h1:C+zze2n6e/7wshOZep2A70/aQU6QBRWJO/G6FT1wIns=
github.com/hdevalence/ed25519consensus v0.2.0 h1:37ICyZqdyj0lAZ8P4D1d1id3HqbbG1N3iBb1Tb4rdcU=
github.com/hdevalence/ed25519consensus v0.2.0/go.mod h1:w3BHWjwJbFU29IRHL1Iqkw3sus+7FctEyM4RqDxYNzo=
github.com/huin/goupnp v1.2.0 h1:uOKW26NG1hsSSbXIZ1IR7XP9Gjd1U8pnLaCMgntmkmY=
github.com/huin/goupnp v1.2.0/go.mod h1:gnGPsThkYa7bFi/KWmEysQRf48l2dvR5bxr2OFckNX8=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
github.com/insomniacslk/dhcp v0.0.0-20260220084031-5adc3eb26f91 h1:u9i04mGE3iliBh0EFuWaKsmcwrLacqGmq1G3XoaM7gY=
github.com/insomniacslk/dhcp v0.0.0-20260220084031-5adc3eb26f91/go.mod h1:qfvBmyDNp+/liLEYWRvqny/PEz9hGe2Dz833eXILSmo=
github.com/jackpal/go-nat-pmp v1.0.2 h1:KzKSgb7qkJvOUTqYl9/Hg/me3pWgBmERKrTGD7BdWus=
github.com/jackpal/go-nat-pmp v1.0.2/go.mod h1:QPH045xvCAeXUZOxsnwmrtiCoxIr9eob+4orBN1SBKc=
github.com/jessevdk/go-flags v1.4.0/go.mod h1:4FA24M0QyGHXBuZZK/XkWh8h0e1EYbRYJSGM75WSRxI=
github.com/jsimonetti/rtnetlink v1.4.0 h1:Z1BF0fRgcETPEa0Kt0MRk3yV5+kF1FWTni6KUFKrq2I=
github.com/jsimonetti/rtnetlink v1.4.0/go.mod h1:5W1jDvWdnthFJ7fxYX1GMK07BUpI4oskfOqvPteYS6E=
@@ -122,6 +128,8 @@ github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zt
github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/klauspost/cpuid/v2 v2.3.0 h1:S4CRMLnYUhGeDFDqkGriYKdfoFlDnMtqTiI/sFzhA9Y=
github.com/klauspost/cpuid/v2 v2.3.0/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/koron/go-ssdp v0.0.4 h1:1IDwrghSKYM7yLf7XCzbByg2sJ/JcNOZRXS2jczTwz0=
github.com/koron/go-ssdp v0.0.4/go.mod h1:oDXq+E5IL5q0U8uSBcoAXzTzInwy5lEgC91HoKtbmZk=
github.com/kr/fs v0.1.0 h1:Jskdu9ieNAYnjxsi0LbQp1ulIKZV1LAFgK1tWhpZgl8=
github.com/kr/fs v0.1.0/go.mod h1:FFnZGqtBN9Gxj7eW1uZ42v5BccTP0vu6NEaFoC2HwRg=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
@@ -138,6 +146,10 @@ github.com/libdns/cloudflare v0.2.2 h1:XWHv+C1dDcApqazlh08Q6pjytYLgR2a+Y3xrXFu0v
github.com/libdns/cloudflare v0.2.2/go.mod h1:w9uTmRCDlAoafAsTPnn2nJ0XHK/eaUMh86DUk8BWi60=
github.com/libdns/libdns v1.1.1 h1:wPrHrXILoSHKWJKGd0EiAVmiJbFShguILTg9leS/P/U=
github.com/libdns/libdns v1.1.1/go.mod h1:4Bj9+5CQiNMVGf87wjX4CY3HQJypUHRuLvlsfsZqLWQ=
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138 h1:YohuNPT/1k3VcThCQlBZ43PCPWPfMRS1zcxWBF2SLK8=
github.com/libp2p/go-nat v1.0.1-0.20250821073202-01afc089f138/go.mod h1:TXQg5tfSy+bUjnhT5728j5j/MBj7keIYqqZ1+8k/ui8=
github.com/libp2p/go-netroute v0.2.1 h1:V8kVrpD8GK0Riv15/7VN6RbUQ3URNZVosw7H2v9tksU=
github.com/libp2p/go-netroute v0.2.1/go.mod h1:hraioZr0fhBjG0ZRXJJ6Zj2IVEVNx6tDTFQfSmcq7mQ=
github.com/logrusorgru/aurora v2.0.3+incompatible h1:tOpm7WcpBTn4fjmVfgpQq0EfczGlG91VSDkswnjF5A8=
github.com/logrusorgru/aurora v2.0.3+incompatible/go.mod h1:7rIyQOR62GCctdiQpZ/zOJlFyk6y+94wXzv6RNZgaR4=
github.com/mdlayher/netlink v1.9.0 h1:G8+GLq2x3v4D4MVIqDdNUhTUC7TKiCy/6MDkmItfKco=
@@ -264,8 +276,8 @@ github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3 h1:3y6
github.com/sagernet/sing-cloudflared v0.1.3-0.20260706062323-d9787e794aa3/go.mod h1:XEqEDYRCAYLaoPjZ1ifVWJg5iWAJHL2gOAXe/PM28Cg=
github.com/sagernet/sing-mux v0.3.5 h1:RHnhVEc+SFqkrK4xMygYjDwwLhzp2Bj3lztSukONfhI=
github.com/sagernet/sing-mux v0.3.5/go.mod h1:QvlKMyNBNrQoyX4x+gq028uPbLM2XeRpWtDsWBJbFSk=
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7 h1:hFLPJ21uNZSbRnzhOKz4Zv0b4F93mpDorWyN93BeRcM=
github.com/sagernet/sing-quic v0.6.2-0.20260525051024-9467ede27fb7/go.mod h1:+oqD54aHel4ALKkp1hVXWCgLU/EjLojvm6AUzDfvj0I=
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc h1:zdc0fj4JdAdgAmQIoh7ZF+B/wPTEF2X75lYDqTmvlaw=
github.com/sagernet/sing-quic v0.6.4-0.20260709034545-e23afe1172dc/go.mod h1:9k+dzGsWMttUGldBzq3dU792YHXzW6NgfbOGltnXq+0=
github.com/sagernet/sing-shadowsocks v0.2.8 h1:PURj5PRoAkqeHh2ZW205RWzN9E9RtKCVCzByXruQWfE=
github.com/sagernet/sing-shadowsocks v0.2.8/go.mod h1:lo7TWEMDcN5/h5B8S0ew+r78ZODn6SwVaFhvB6H+PTI=
github.com/sagernet/sing-shadowsocks2 v0.2.1 h1:dWV9OXCeFPuYGHb6IRqlSptVnSzOelnqqs2gQ2/Qioo=
@@ -360,6 +372,8 @@ go4.org/mem v0.0.0-20240501181205-ae6ca9944745 h1:Tl++JLUCe4sxGu8cTpDzRLd3tN7US4
go4.org/mem v0.0.0-20240501181205-ae6ca9944745/go.mod h1:reUoABIJ9ikfM5sgtSF3Wushcza7+WeD01VB9Lirh3g=
go4.org/netipx v0.0.0-20231129151722-fdeea329fbba h1:0b9z3AuHCjxk0x/opv64kcgZLBseWJUpBw5I82+2U4M=
go4.org/netipx v0.0.0-20231129151722-fdeea329fbba/go.mod h1:PLyyIXexvUFg3Owu6p/WfdlivPbZJsZdgWZlrGope/Y=
golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w=
golang.org/x/crypto v0.0.0-20191011191535-87dc89f01550/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/crypto v0.0.0-20210513164829-c07d793c2f9a/go.mod h1:P+XmwS30IXTQdn5tA2iutPOUgjI07+tq3H3K9MVA1s8=
golang.org/x/crypto v0.48.0 h1:/VRzVqiRSggnhY7gNRxPauEQ5Drw9haKdM0jqfcCFts=
golang.org/x/crypto v0.48.0/go.mod h1:r0kV5h3qnFPlQnBSrULhlsRfryS2pmewsg+XfMgkVos=
@@ -367,17 +381,24 @@ golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93 h1:fQsdNF2N+/YewlRZiricy4P1i
golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93/go.mod h1:EPRbTFwzwjXj9NpYyyrvenVh9Y+GFeEvMNh7Xuz7xgU=
golang.org/x/image v0.27.0 h1:C8gA4oWU/tKkdCfYT6T2u4faJu3MeNS5O8UPWlPF61w=
golang.org/x/image v0.27.0/go.mod h1:xbdrClrAUway1MUTEZDq9mz/UpRwYAkFFNUslZtcB+g=
golang.org/x/lint v0.0.0-20200302205851-738671d3881b/go.mod h1:3xt1FjdF8hUf6vQPIChWIBhFzV8gjjsPE/fR3IyQdNY=
golang.org/x/mod v0.1.1-0.20191105210325-c90efee705ee/go.mod h1:QqPTAvyqsEbceGzBzNggFXnrqF1CaUcvgkdR5Ot7KZg=
golang.org/x/mod v0.33.0 h1:tHFzIWbBifEmbwtGz65eaWyGiGZatSrT9prnU8DbVL8=
golang.org/x/mod v0.33.0/go.mod h1:swjeQEj+6r7fODbD2cqrnje9PnziFuw4bmLbBZFrQ5w=
golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s=
golang.org/x/net v0.0.0-20210226172049-e18ecbb05110/go.mod h1:m0MpNAwzfU5UDzcl9v0D8zg8gWTRqZa9RBIspLL5mdg=
golang.org/x/net v0.0.0-20210525063256-abc453219eb5/go.mod h1:9nx3DQGgdP8bBQD5qxJ1jj9UTztislL4KSBs9R2vV5Y=
golang.org/x/net v0.50.0 h1:ucWh9eiCGyDR3vtzso0WMQinm2Dnt8cFMuQa9K33J60=
golang.org/x/net v0.50.0/go.mod h1:UgoSli3F/pBgdJBHCTc+tp3gmrU4XswgGRgtnwWTfyM=
golang.org/x/oauth2 v0.34.0 h1:hqK/t4AKgbqWkdkcAeI8XLmbK+4m4G5YeQRrmiotGlw=
golang.org/x/oauth2 v0.34.0/go.mod h1:lzm5WQJQwKZ3nwavOZ3IS5Aulzxi68dUSgRHujetwEA=
golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20210220032951-036812b2e83c/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.19.0 h1:vV+1eWNmZ5geRlYjzm2adRgW2/mcpevXNg50YZtPCE4=
golang.org/x/sync v0.19.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200217220822-9197077df867/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20200728102440-3e129f6d46b1/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -389,6 +410,7 @@ golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.40.0 h1:36e4zGLqU4yhjlmxEaagx2KuYbJq3EwY8K943ZsHcvg=
golang.org/x/term v0.40.0/go.mod h1:w2P8uVp06p2iyKKuvXIm7N/y0UCRt3UfJTfZ7oOpglM=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.34.0 h1:oL/Qq0Kdaqxa1KbNeMKwQq0reLCCaFtqu2eNuSeNHbk=
@@ -396,8 +418,10 @@ golang.org/x/text v0.34.0/go.mod h1:homfLqTYRFyVYemLBFl5GgL/DWEiH5wcsQ5gSh1yziA=
golang.org/x/time v0.11.0 h1:/bpjEDfN9tkoN/ryeYHnv5hcMlc8ncjMcM4XBk5NWV0=
golang.org/x/time v0.11.0/go.mod h1:CDIdPxbZBQxdj6cxyCIdrNogrJKMJ7pr37NYpMcMDSg=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.0.0-20200130002326-2f3ba24bd6e7/go.mod h1:TB2adYChydJhpapKDTa4BR/hXlZSLoq2Wpct/0txZ28=
golang.org/x/tools v0.42.0 h1:uNgphsn75Tdz5Ji2q36v/nsFSfR/9BRFvqhGBaJGd5k=
golang.org/x/tools v0.42.0/go.mod h1:Ma6lCIwGZvHK6XtgbswSoWroEkhugApmsXyrUmBhfr0=
golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
golang.org/x/xerrors v0.0.0-20200804184101-5ec99f83aff1 h1:go1bK/D/BFZV2I8cIQd1NKEZ+0owSTG1fDTci4IqFcE=
golang.org/x/xerrors v0.0.0-20200804184101-5ec99f83aff1/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -26,15 +26,32 @@ var callMintToken = rpc.declare({
expect: { '': {} }
});
// led renders a small status dot: state is 'good' | 'warn' | 'bad'.
// LED palette. 'unknown' is an UNLIT socket — never amber and never green.
// Amber is this page's "degraded", and there is nothing to be degraded about
// when no reading has arrived; green on a missing reading is how the panel used
// to claim health it had not measured (see panel/src/planeState.ts, which says
// the same thing and is the wording this page is kept in step with).
var LED_COLORS = {
good: '#37b24d',
warn: '#f59f00',
bad: '#e03131',
unknown: '#6b6b6b'
};
// led renders a small status dot: state is 'good' | 'warn' | 'bad' | 'unknown'.
// The dot is decorative — every row states its condition in words beside it — so
// it is hidden from assistive tech rather than being the only carrier of meaning.
function led(state) {
var color = state === 'good' ? '#37b24d'
: state === 'warn' ? '#f59f00'
: '#e03131';
// Closed positive list. An unrecognised state resolves to UNKNOWN, never to
// green: an open default here is exactly how a state nobody thought about
// ends up painted healthy.
var color = Object.prototype.hasOwnProperty.call(LED_COLORS, state)
? LED_COLORS[state] : LED_COLORS.unknown;
var glow = (color === LED_COLORS.unknown) ? '' : ';box-shadow:0 0 5px ' + color;
return E('span', {
'aria-hidden': 'true',
'style': 'display:inline-block;width:.72em;height:.72em;border-radius:50%;' +
'margin-right:.6em;vertical-align:-.05em;background:' + color +
';box-shadow:0 0 5px ' + color
'margin-right:.6em;vertical-align:-.05em;background:' + color + glow
});
}
@@ -49,63 +66,251 @@ function row(state, label, value) {
]);
}
// statusRows maps the shaterd status object to LED rows. An empty object (the
// ubus call failed / daemon down) degrades every row to a "down" reading.
function statusRows(st) {
st = st || {};
var down = (st.running !== true);
// ---------------------------------------------------------------------------
// Pure state derivation — no DOM below this line until statusRows().
//
// statusReadout() maps a shaterd status object to a list of
// { state, label, value } descriptors. It is deliberately free of E()/DOM so it
// can be run against recorded fixtures offline; tests/status-readout.test.js
// does exactly that for the three cases this page has to tell apart.
// ---------------------------------------------------------------------------
// PLANES is the closed set of values a LIVE Applier.Status() can put in `plane`
// (shater/apply/apply.go: "full" | "hold" | "none"). It is also how this page
// tells a live daemon from a dead one — see daemonState().
var PLANES = { full: true, hold: true, none: true };
// daemonState — is the shaterd PROCESS answering?
//
// 'up' — a live Applier produced this status.
// 'down' — proven not: `shaterd status` printed its OFFLINE STUB.
// 'unknown' — no usable answer, or an answer from a daemon older than `plane`.
//
// `running` MUST NOT be used for this. It changed meaning on 2026-07-26
// (a8970b8ac): it used to be a hardcoded true, and is now the ENGINE's liveness
// (apply.go `Running: engineUp`). A daemon that is perfectly alive with a dead
// engine reports running=false — and this page used to answer that with a red
// "Daemon: not running", the advice "start the Shater service first", and a
// DISABLED button to the one place the config can be fixed. The holding plane
// keeps management reachable on purpose (shater/netplane/nft.go); LuCI was the
// only thing taking that guarantee away.
//
// Nor is "the ubus call returned" sufficient, which is the trap here: the rpcd
// plugin shells out to `shaterd status`, and that command EXITS 0 WITH A
// FABRICATED STATUS when the daemon is unreachable (cmd/shaterd/main.go,
// cmdStatus offline stub). The stub is the apply.Status zero value plus a UCI
// read, so it carries enabled/table/kill_switch/panel_port but leaves `plane` at
// "" — a value no live daemon ever emits, because Status() always assigns one of
// the three words. So a known plane word is the one positive proof on the wire
// that a daemon answered, and an explicit empty one is positive proof that none
// did.
//
// Everything else is unknown and is painted as unknown: {} from a failed ubus
// call, {"error":...} from the plugin (which is ALSO what a live-but-wedged
// daemon produces — cmdStatus prints nothing and exits 1 on a control-socket
// timeout, so "wedged" must not be reported as "dead"), and a status from a
// daemon predating the `plane` field.
function daemonState(st) {
if (!st || typeof st !== 'object')
return 'unknown';
if (typeof st.plane === 'string' && PLANES[st.plane] === true)
return 'up';
if (st.plane === '')
return 'down';
return 'unknown';
}
// engineState — is a sing-box instance actually started?
//
// The engine lives INSIDE the shaterd process, so a dead daemon is a dead engine
// and this page may say so without guessing. With the daemon up, `engine_running`
// is the self-documenting field and `running` carries the same fact by
// construction; either may prove a NEGATIVE, and a negative always wins. Neither
// asserting anything leaves 'unknown'.
function engineState(st) {
var d = daemonState(st);
if (d === 'down')
return 'down';
if (d === 'unknown')
return 'unknown';
if (st.running === false || st.engine_running === false)
return 'down';
if (st.running === true || st.engine_running === true)
return 'up';
return 'unknown';
}
function mk(state, label, value) {
return { state: state, label: label, value: value };
}
function statusReadout(st) {
st = (st && typeof st === 'object') ? st : {};
var dstate = daemonState(st);
var estate = engineState(st);
var traffic = (st.traffic && typeof st.traffic === 'object') ? st.traffic : {};
var rows = [];
// Daemon process itself.
rows.push(row(
down ? 'bad' : 'good',
_('Daemon (shaterd)'),
down ? _('not running') : _('running')
));
// --- The shaterd process itself. ------------------------------------------
// Its own liveness is not a field; it is whether a live daemon answered.
if (dstate === 'up')
rows.push(mk('good', _('Daemon (shaterd)'), _('responding')));
else if (dstate === 'down')
rows.push(mk('bad', _('Daemon (shaterd)'),
_('not responding — start the Shater service')));
else
rows.push(mk('unknown', _('Daemon (shaterd)'),
_('no usable answer — state unknown')));
// Desired state: globals.enabled in UCI.
rows.push(row(
st.enabled ? 'good' : 'warn',
_('Service enabled'),
st.enabled ? _('enabled') : _('inert (disabled)')
));
// --- The engine (sing-box) inside it. -------------------------------------
if (estate === 'up')
rows.push(mk('good', _('Engine (sing-box)'), _('running')));
else if (estate === 'down' && dstate === 'down')
rows.push(mk('bad', _('Engine (sing-box)'),
_('stopped — it runs inside shaterd, which is not answering')));
else if (estate === 'down')
rows.push(mk('bad', _('Engine (sing-box)'),
_('stopped — the daemon is up but no instance is running')));
else
rows.push(mk('unknown', _('Engine (sing-box)'), _('not reported')));
// Interception raised (ACTIVE_FLAG present after a successful enabled apply).
rows.push(row(
st.active ? 'good' : (st.enabled ? 'warn' : 'bad'),
_('Interception'),
st.active ? _('active') : _('inactive')
));
// --- Desired state: globals.enabled in UCI. -------------------------------
if (st.enabled === true)
rows.push(mk('good', _('Service enabled'), _('enabled')));
else if (st.enabled === false)
rows.push(mk('warn', _('Service enabled'), _('inert (disabled)')));
else
rows.push(mk('unknown', _('Service enabled'), _('not reported')));
// Data plane: the `inet shater` nft table is loaded.
rows.push(row(
st.table ? 'good' : (st.enabled ? 'warn' : 'bad'),
_('Data plane'),
st.table ? _('nft table inet shater loaded') : _('not loaded')
));
// --- The ACTIVE_FLAG latch. -----------------------------------------------
// NOT a health signal, and this row must never read as one. apply.go states
// the contract: it is the "the service is meant to be running" latch that
// gates hotplug and cron; it is raised by a successful enabled apply and
// cleared only by teardown, so it STAYS UP while the engine is down and the
// fail-closed holding plane is blocking the LAN — deliberately, because
// clearing it would switch off the very cron reconcile that brings the engine
// back. This page used to render it as "Interception: active", in green, over
// a dead engine and a blocked LAN. The lamp now reports only whether the
// latch AGREES with globals.enabled.
if (typeof st.active !== 'boolean')
rows.push(mk('unknown', _('Service latch'), _('not reported')));
else if (st.active)
rows.push(mk(st.enabled === true ? 'good' : 'warn', _('Service latch'),
_('raised — the service is meant to be running')));
else
rows.push(mk(st.enabled === false ? 'good' : 'warn', _('Service latch'),
_('cleared — the service is torn down')));
// Kill-switch: fail-closed ("closed") is the safe posture; "open" leaks
// LAN→WAN if the engine goes down. Unknown (older daemon) degrades to warn.
var ks = st.kill_switch;
rows.push(row(
ks === 'closed' ? 'good' : 'warn',
_('Kill-switch'),
ks === 'closed' ? _('closed (fail-closed)')
: ks === 'open' ? _('open (leaky)')
: _('unknown')
));
// --- What is loaded in the kernel right now. ------------------------------
// The row this page was missing. `plane` distinguishes the working ruleset
// from the FAIL-CLOSED HOLDING PLANE, which `table` cannot: `table` is a bare
// existence check, so a held LAN and a working one look identical through it.
switch (st.plane) {
case 'full':
// Deliberately mechanical wording. "full" means the table, the policy
// routing and the engine are all in place — it does NOT mean traffic is
// tunnelled. That claim belongs to the traffic verdict below.
rows.push(mk('good', _('Traffic plane'),
_('full — ruleset, routing and engine are all installed')));
break;
case 'hold':
rows.push(mk('bad', _('Traffic plane'),
_('hold — the engine is down and LAN→WAN forwarding is BLOCKED')));
break;
case 'none':
rows.push(mk(st.kill_switch === 'closed' ? 'bad' : 'warn', _('Traffic plane'),
st.kill_switch === 'closed'
? _('none — nothing is installed; traffic reaches the WAN unprotected')
: _('none — no data plane is installed')));
break;
default:
rows.push(mk('unknown', _('Traffic plane'),
dstate === 'down'
? _('not reported — no daemon answered')
: _('not reported by this daemon')));
break;
}
// Running engine config hash ("" when the engine is not started).
rows.push(row(
st.hash ? 'good' : 'warn',
_('Config hash'),
st.hash ? st.hash : '—'
));
// --- Where the traffic goes under the running config. ---------------------
// Separate from the plane on purpose: a router with one `default -> direct`
// rule has a fully installed plane and sends every packet out the plain WAN
// with its real address.
switch (traffic.verdict) {
case 'tunnel':
rows.push(mk('good', _('Traffic verdict'), _('tunnel — unmatched traffic is proxied')));
break;
case 'split':
rows.push(mk('warn', _('Traffic verdict'),
_('split — the default leaves directly; only matched rules are tunnelled')));
break;
case 'direct':
rows.push(mk('warn', _('Traffic verdict'),
_('direct — nothing is tunnelled; traffic leaves over the plain WAN')));
break;
case 'blocked':
rows.push(mk('warn', _('Traffic verdict'), _('blocked — unmatched traffic is dropped')));
break;
default:
rows.push(mk('unknown', _('Traffic verdict'), _('not reported')));
break;
}
// --- The nft table, as a bare presence check. -----------------------------
// Kept because the offline stub still reads it straight from the kernel, so
// it is the one plane fact available when no daemon answers. It says nothing
// about WHICH ruleset is loaded — that is the Traffic plane row.
if (st.table === true)
rows.push(mk('good', _('nft table'), _('inet shater is loaded')));
else if (st.table === false)
rows.push(mk(st.enabled === false ? 'warn' : 'bad', _('nft table'), _('not loaded')));
else
rows.push(mk('unknown', _('nft table'), _('not reported')));
// --- Kill-switch: the configured policy, and whether it is in force. ------
// "closed" with no plane installed is a setting that is not in effect, which
// is worse news than "open" and must not share its amber lamp.
if (st.kill_switch === 'closed' && st.plane === 'none')
rows.push(mk('bad', _('Kill-switch'),
_('closed, but NOT in effect — no data plane is installed')));
else if (st.kill_switch === 'closed' && dstate === 'up')
rows.push(mk('good', _('Kill-switch'), _('closed (fail-closed)')));
else if (st.kill_switch === 'closed')
rows.push(mk('warn', _('Kill-switch'),
_('configured closed; whether it is installed is not known')));
else if (st.kill_switch === 'open')
rows.push(mk('warn', _('Kill-switch'), _('open (leaky)')));
else
rows.push(mk('unknown', _('Kill-switch'), _('not reported')));
// --- Running engine config hash ("" when the engine is not started). ------
if (typeof st.hash === 'string' && st.hash !== '')
rows.push(mk('good', _('Config hash'), st.hash));
else if (estate === 'down')
rows.push(mk('unknown', _('Config hash'), _('none — the engine is not started')));
else
rows.push(mk('unknown', _('Config hash'), _('not reported')));
return rows;
}
// statusRows turns the readout into LED table rows.
function statusRows(st) {
return statusReadout(st).map(function(r) {
return row(r.state, r.label, r.value);
});
}
// panelHint describes the button's target and what is known about it. It never
// promises the panel is up — only where the launcher will point.
function panelTitle(dstate) {
if (dstate === 'up')
return _('Mint a session token and open the admin panel');
if (dstate === 'down')
return _('shaterd is not answering, so this will probably fail — but the panel is served by the daemon, not by the engine, so it is worth trying: any failure is reported here.');
return _('The daemon state is not known. Try it — a failure is reported here rather than hidden.');
}
// handleOpenPanel mints a single-use token and hands it to the panel via the
// ARCHITECTURE §2 browser bridge: GET http://<router>:<port>/?t=<token>. The panel
// validates+consumes the token and drops a session cookie.
@@ -151,13 +356,17 @@ return view.extend({
handleSave: null,
handleReset: null,
// Exposed so the offline fixture harness (tests/status-readout.test.js) can
// exercise the state derivation without a browser, a router, or a DOM.
statusReadout: statusReadout,
daemonState: daemonState,
engineState: engineState,
load: function() {
return L.resolveDefault(callStatus(), {});
},
render: function(st) {
var self = this;
var table = E('table', { 'class': 'table' }, statusRows(st));
var openBtn = E('button', {
@@ -174,25 +383,29 @@ return view.extend({
'style': 'margin-left:1em;color:#888;font-size:90%'
}, hintText());
// Reflect daemon reachability on the button up front, then keep the whole
// dashboard live. Also track the panel port reported in status so the
// launcher redirect and hint follow globals.panel_port.
// Track the panel port reported in status so the launcher redirect and the
// hint follow globals.panel_port.
//
// THE BUTTON IS NEVER DISABLED. It used to be locked whenever
// `running !== true`, which after a8970b8ac means "the engine is down" —
// precisely the situation the panel exists to get you out of, and one in
// which the daemon and its web server are still up and still minting
// tokens (cmd/shaterd/main.go starts the panel server independently of the
// engine). Locking it on a guess is the failure; a mint that fails already
// reports itself through ui.addNotification, which is the recoverable
// direction for an unknown state.
function reflect(state) {
state = state || {};
state = (state && typeof state === 'object') ? state : {};
panelPort = state.panel_port || DEFAULT_PANEL_PORT;
hint.textContent = hintText();
var down = (state.running !== true);
openBtn.disabled = down;
openBtn.title = down
? _('shaterd is not running — start the Shater service first')
: _('Mint a session token and open the admin panel');
openBtn.title = panelTitle(daemonState(state));
}
reflect(st || {});
reflect(st);
poll.add(function() {
return L.resolveDefault(callStatus(), {}).then(function(s) {
dom.content(table, statusRows(s));
reflect(s || {});
reflect(s);
});
}, 5);
@@ -209,7 +422,7 @@ return view.extend({
E('div', { 'class': 'cbi-section' }, [
E('h3', {}, _('Admin panel')),
E('p', { 'class': 'cbi-value-description' },
_('The rich admin panel is served by shaterd on its own port. LuCI mints a short-lived, single-use token for your browser — the panel has no separate login.')),
_('The rich admin panel is served by shaterd on its own port — by the daemon, not by the engine, so it stays reachable while the engine is down. LuCI mints a short-lived, single-use token for your browser; the panel has no separate login.')),
E('div', {}, [ openBtn, hint ])
])
]);
@@ -4,9 +4,27 @@
# Registers the ubus object "shater" (object name == this file's name) with two
# read-side methods the thin LuCI launcher calls over ubus:
#
# status -> passthrough of `shaterd status` ({running,enabled,active,table,hash})
# status -> passthrough of `shaterd status`
# mint_token -> passthrough of `shaterd mint-token` ({"token":"..."} | {"error":"..."})
#
# The status object is whatever apply.Status marshals (shater/apply/apply.go is the
# only definition; this script never parses or reshapes it). As of 2026-07-26 that is:
#
# running, engine_running, enabled, active, table, plane, traffic, hash,
# kill_switch, panel_port, can_rollback, warnings, started_unix, uptime_seconds
#
# Two of those are load-bearing for the caller and easy to misread:
#
# running / engine_running — the ENGINE's liveness, not this daemon's. `running`
# was a hardcoded true until a8970b8ac (2026-07-26) and is now `engineUp`, so
# a healthy daemon with a dead engine reports running=false. The daemon's own
# liveness is not a field at all.
# plane — "full" | "hold" | "none" from a LIVE daemon. `shaterd status` also has an
# OFFLINE STUB path: when the daemon is unreachable it still exits 0 and prints
# a status built from the apply.Status zero value plus a UCI read, which leaves
# plane at "". So this method returning an object is NOT evidence that a daemon
# answered; a known plane word is. dashboard.js relies on exactly that.
#
# Why shell out to shaterd instead of talking to /var/run/shaterd.ctl directly:
# a reliable AF_UNIX client is NOT guaranteed on stock OpenWrt (busybox `nc` is
# usually built without `-U`; socat/ucode-socket aren't in the base image). shaterd
@@ -0,0 +1,236 @@
#!/usr/bin/env node
/*
* Offline harness for the dashboard's state derivation.
*
* Run: node openwrt/luci-app-shater/tests/status-readout.test.js
*
* Why this exists: the LuCI page is the ONE screen an operator reaches when the
* engine is down and the fail-closed holding plane is blocking the LAN. What it
* says there is a claim about the router's behaviour, and until now nothing
* checked those claims. There is no browser and no router in this loop — the view
* exposes statusReadout/daemonState/engineState as plain functions, and this file
* feeds them recorded status objects.
*
* The fixtures are not invented. Each is what the wire actually carries:
*
* ENGINE_UP — apply.Status() from a live daemon with a started engine.
* ENGINE_DOWN — apply.Status() from a live daemon whose engine died; the
* holding plane is installed and the LAN is blocked.
* DAEMON_DOWN — the OFFLINE STUB `shaterd status` prints when the daemon is
* unreachable (cmd/shaterd/main.go cmdStatus): apply.Status zero
* value + a UCI read, marshalled by Status.JSON(), so every field
* is present and `plane` is "".
* NO_ANSWER — {} , what L.resolveDefault hands render() when the ubus call
* fails outright.
* PLUGIN_ERROR — {"error":...} from the rpcd plugin, which is ALSO what a
* live-but-wedged daemon produces.
* LEGACY — a daemon predating plane/engine_running (packages do not update
* atomically).
*
* Mutation check: revert dashboard.js to reading `st.running` for daemon
* liveness and this file fails on ENGINE_DOWN with the exact text the operator
* would have been shown.
*/
'use strict';
var fs = require('fs');
var path = require('path');
// --- Load the view module with LuCI's globals stubbed. ----------------------
// The view file is a module body LuCI wraps in a function, so it ends in a
// top-level `return` and cannot be require()d. Wrapping it in new Function is the
// same thing LuCI's loader does. The 'require x' lines are bare string literals
// and evaluate to nothing.
// DASHBOARD_JS points the harness at a copy of the view. It exists so the
// mutation check is repeatable: copy dashboard.js, reintroduce the defect in the
// copy, run this file against it, and watch the named assertions fail. A test
// that cannot be shown to fail on the broken code is decoration.
var SRC = process.env.DASHBOARD_JS || path.join(__dirname, '..', 'htdocs',
'luci-static', 'resources', 'view', 'shater', 'dashboard.js');
function loadView() {
var src = fs.readFileSync(SRC, 'utf8');
var factory = new Function('view', 'dom', 'poll', 'rpc', 'ui', 'E', '_', 'L',
'window', src);
return factory(
{ extend: function(o) { return o; } }, // view
{ content: function() {} }, // dom
{ add: function() {} }, // poll
{ declare: function() { return function() {}; } }, // rpc
{ createHandlerFn: function() { return function() {}; }, addNotification: function() {} },
function() { return {}; }, // E
function(s) { return s; }, // _ (identity)
{ resolveDefault: function(p, d) { return Promise.resolve(d); } },
{ location: { hostname: 'router' }, open: function() { return null; } }
);
}
var page = loadView();
// --- Fixtures ---------------------------------------------------------------
var ENGINE_UP = {
running: true, engine_running: true, enabled: true, active: true, table: true,
plane: 'full', traffic: { verdict: 'tunnel', default: 'proxy', tunnel_rules: 3 },
hash: 'a1b2c3d4', kill_switch: 'closed', panel_port: 8088, can_rollback: true,
warnings: [], started_unix: 1753500000, uptime_seconds: 3600
};
var ENGINE_DOWN = {
running: false, engine_running: false, enabled: true, active: true, table: true,
plane: 'hold', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: true,
warnings: [], started_unix: 1753500000, uptime_seconds: 3600
};
// Exactly what Status.JSON() emits for the cmdStatus offline stub.
var DAEMON_DOWN = {
running: false, engine_running: false, enabled: true, active: true, table: true,
plane: '', traffic: { verdict: '', default: '', tunnel_rules: 0 },
hash: '', kill_switch: 'closed', panel_port: 8088, can_rollback: false,
warnings: null, started_unix: 0, uptime_seconds: 0
};
var NO_ANSWER = {};
var PLUGIN_ERROR = { error: 'shaterd unavailable' };
var LEGACY = {
running: true, enabled: true, active: true, table: true, hash: 'deadbeef',
kill_switch: 'closed', panel_port: 8088
};
// --- Assertions -------------------------------------------------------------
var failures = [];
function check(name, cond, detail) {
if (cond) return;
failures.push(name + (detail ? ': ' + detail : ''));
}
// readout indexes the rows by label. A row that is NOT emitted must fail by name
// rather than by throwing on `undefined.state`: a harness that dies mid-run stops
// reporting the assertions after it, which is the silent-skip failure this
// project has been bitten by. Missing rows come back as a loud sentinel instead.
var MISSING = { state: '<row absent>', value: '<row absent>', missing: true };
function readout(st) {
var out = {};
page.statusReadout(st).forEach(function(r) { out[r.label] = r; });
return new Proxy(out, {
get: function(t, k) {
if (typeof k !== 'string' || k in t) return t[k];
return MISSING;
},
has: function(t, k) { return k in t; }
});
}
function show(title, st) {
process.stdout.write('\n=== ' + title + ' ===\n');
process.stdout.write(' daemon=' + page.daemonState(st) +
' engine=' + page.engineState(st) + '\n');
page.statusReadout(st).forEach(function(r) {
process.stdout.write(' [' + r.state.padEnd(7) + '] ' +
r.label.padEnd(18) + ' ' + r.value + '\n');
});
}
function lamps(st) {
return page.statusReadout(st).map(function(r) { return r.state; });
}
// 1. Live daemon, engine up.
show('A. daemon alive, engine running', ENGINE_UP);
check('A/daemon', page.daemonState(ENGINE_UP) === 'up');
check('A/engine', page.engineState(ENGINE_UP) === 'up');
check('A/no-red', lamps(ENGINE_UP).indexOf('bad') === -1,
'a fully healthy router must show no red lamp');
check('A/no-unknown', lamps(ENGINE_UP).indexOf('unknown') === -1,
'every field is present, so nothing may read as unknown');
// 2. THE DEFECT. Live daemon, dead engine, LAN held.
show('B. daemon alive, engine DOWN, holding plane', ENGINE_DOWN);
check('B/daemon-up', page.daemonState(ENGINE_DOWN) === 'up',
'the daemon is answering; calling it dead is the bug being fixed');
check('B/engine-down', page.engineState(ENGINE_DOWN) === 'down');
var b = readout(ENGINE_DOWN);
check('B/daemon-row-green', b['Daemon (shaterd)'].state === 'good',
'got ' + b['Daemon (shaterd)'].state + ' / ' + b['Daemon (shaterd)'].value);
check('B/daemon-row-no-start-advice',
b['Daemon (shaterd)'].value.indexOf('start') === -1,
'must not tell the operator to start a service that is already running');
check('B/plane-red', b['Traffic plane'].state === 'bad');
check('B/plane-says-blocked', /BLOCKED/.test(b['Traffic plane'].value));
check('B/latch-not-called-interception',
!b['Service latch'].missing && b['Interception'].missing === true,
'`active` is the run latch, not a "we are proxying" signal — apply.go: ' +
'"Never render it as \'we are proxying\'"');
check('B/latch-value-is-a-latch', /meant to be running/.test(b['Service latch'].value),
'the latch row must state the latch, not claim traffic is being proxied');
check('B/latch-not-active-word', !/^active$/.test(b['Service latch'].value));
check('B/verdict-unknown', b['Traffic verdict'].state === 'unknown',
'no verdict was published; it must not be painted as tunnel');
check('B/some-red', lamps(ENGINE_DOWN).indexOf('bad') !== -1,
'a blocked LAN must not be an all-green screen');
// The button is a property of render(), not of the pure readout, so it is
// guarded at the source level: nothing may ever set `disabled` on the launcher.
// Locking the way into the panel while the engine is down is the defect this
// whole file exists for, and it must not come back by a different route.
check('B/button-never-disabled',
!/openBtn\s*\.\s*disabled/.test(fs.readFileSync(SRC, 'utf8')),
'dashboard.js assigns openBtn.disabled — the launcher must never be locked');
// 3. Daemon not answering at all — the offline stub.
show('C. daemon NOT running (offline stub)', DAEMON_DOWN);
check('C/daemon-down', page.daemonState(DAEMON_DOWN) === 'down',
'plane:"" is the stub signature; got ' + page.daemonState(DAEMON_DOWN));
check('C/engine-down', page.engineState(DAEMON_DOWN) === 'down');
var c = readout(DAEMON_DOWN);
check('C/daemon-row-red', c['Daemon (shaterd)'].state === 'bad');
check('C/plane-unknown', c['Traffic plane'].state === 'unknown',
'the stub reports no plane; got ' + c['Traffic plane'].value);
check('C/kill-switch-not-green', c['Kill-switch'].state !== 'good',
'"closed" from a dead daemon proves nothing is installed to enforce it');
check('C/distinct-from-B',
c['Daemon (shaterd)'].value !== b['Daemon (shaterd)'].value,
'engine-down and daemon-down must not render identically');
// 4/5/6. Degenerate answers must degrade to unknown, never to healthy.
[['D. ubus call failed ({})', NO_ANSWER],
['E. rpcd plugin error / wedged daemon', PLUGIN_ERROR],
['F. legacy daemon (no plane, no engine_running)', LEGACY]].forEach(function(p) {
show(p[0], p[1]);
var st = p[1];
check(p[0] + '/daemon-unknown', page.daemonState(st) === 'unknown');
check(p[0] + '/engine-unknown', page.engineState(st) === 'unknown');
var r = readout(st);
check(p[0] + '/plane-unknown', r['Traffic plane'].state === 'unknown');
check(p[0] + '/daemon-row-unknown', r['Daemon (shaterd)'].state === 'unknown');
check(p[0] + '/no-false-green-plane', r['Traffic plane'].state !== 'good');
});
// The legacy fixture additionally must not crash and must not lose the fields it
// DOES carry — a non-atomic package update must degrade, not black out.
var f = readout(LEGACY);
check('F/enabled-still-read', f['Service enabled'].state === 'good');
check('F/hash-still-read', f['Config hash'].value === 'deadbeef');
check('F/table-still-read', f['nft table'].state === 'good');
// A plane word this build does not know about must land in unknown, not in the
// last-listed branch. (Closed positive list, recoverable default.)
var FUTURE = Object.assign({}, ENGINE_UP, { plane: 'partial' });
check('G/unknown-plane-word', page.daemonState(FUTURE) === 'unknown',
'an unrecognised plane value must not be read as a live daemon');
check('G/unknown-plane-row', readout(FUTURE)['Traffic plane'].state === 'unknown');
// --- Report -----------------------------------------------------------------
process.stdout.write('\n');
if (failures.length) {
process.stdout.write('FAIL (' + failures.length + ')\n');
failures.forEach(function(f) { process.stdout.write(' - ' + f + '\n'); });
process.exit(1);
}
process.stdout.write('OK — all cases distinguished\n');
+19 -1
View File
@@ -40,6 +40,10 @@ define Package/shater-core
# shaterd : the daemon our init supervises (`shaterd run`)
# kmod-nft-tproxy : kernel TPROXY (shaterd emits the `inet shater` rules)
# kmod-nft-socket : socket match used by the tproxy divert chain
# kmod-tun : /dev/net/tun — the daemon opens the `shater-l3` TUN
# for L3 ingress (globals.l3_tunnel); usually built-in
# on stock images, but a slimmed image without it would
# make the option fail with a cryptic open() error.
# ip-full : `ip rule`/`ip route`/rt_tables for policy routing
# nftables-json : shaterd shells out to `nft`, and netplane/stats.go
# parses `nft -j list ...` — the JSON output only exists
@@ -49,7 +53,7 @@ define Package/shater-core
# ca-bundle : the daemon is CGO_ENABLED=0, so crypto/x509 has no
# host cert fallback — without /etc/ssl/certs every
# HTTPS subscription / .srs ruleset fetch fails.
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full +nftables-json +ca-bundle
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle
PKGARCH:=all
endef
@@ -80,6 +84,10 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/init.d
$(INSTALL_BIN) ./files/etc/init.d/shater $(1)/etc/init.d/shater
$(INSTALL_BIN) ./files/etc/init.d/shater-cron $(1)/etc/init.d/shater-cron
# START=21 one-shot that loads the persisted fail-closed plane before fw4's
# `lan -> wan ACCEPT` can be the only thing on the box (the main init is
# START=99, i.e. seconds of plaintext forwarding on every boot).
$(INSTALL_BIN) ./files/etc/init.d/shater-armor $(1)/etc/init.d/shater-armor
$(INSTALL_DIR) $(1)/etc/hotplug.d/iface
$(INSTALL_BIN) ./files/etc/hotplug.d/iface/99-shater $(1)/etc/hotplug.d/iface/99-shater
@@ -92,6 +100,16 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/uci-defaults
$(INSTALL_BIN) ./files/etc/uci-defaults/30_shater-core $(1)/etc/uci-defaults/30_shater-core
# sysupgrade's "keep settings" walks /lib/upgrade/keep.d/*, and without this the
# node inventory in /etc/shater/subs does NOT survive a flash: the restored box
# has its rules and its groups and no nodes for them to point at, and the only
# repair is `sub update`, which needs the internet the tunnel was going to
# provide. Package metadata, not user config, so INSTALL_DATA and not
# INSTALL_CONF. (/etc/config/shater needs no entry — it is a conffile and
# sysupgrade already keeps it that way.)
$(INSTALL_DIR) $(1)/lib/upgrade/keep.d
$(INSTALL_DATA) ./files/lib/upgrade/keep.d/shater-core $(1)/lib/upgrade/keep.d/shater-core
endef
$(eval $(call BuildPackage,shater-core))
+45 -5
View File
@@ -49,17 +49,58 @@ config globals 'globals'
# Queries aimed at an EXTERNAL resolver are dropped with the rest of the LAN's
# forwarded traffic.
option dns_intercept '1'
# Carry LAN ping through the tunnel. ON by default, and the alternative is
# why: without it a ping is decided by `untunnelable` below, whose rungs are
# "drop it" (block, the default) or "let it out of the WAN interface with the
# client's real IP on it" (icmp/direct). There was no setting in which ping
# both worked and stayed inside the tunnel. With this on, the engine opens a
# TUN, LAN ICMP is routed into it, and an outbound that speaks layer 3
# (WireGuard/AmneziaWG, or a direct route) carries the echo for real. An
# outbound that does not (vless/trojan/shadowsocks) makes the ping DROP —
# honestly: no reply is forged, ping reports loss. So a ping that used to
# "work" through such a node was a ping that was leaking.
#
# It costs a permanent TUN device plus the gVisor netstack behind it, about
# 2 MB of RSS for as long as the daemon runs.
#
# Set to '0' to opt out — worth it on a 32/64 MB router, or to bisect whether
# the L3 ingress is what broke something. `shaterd apply` will tell you what
# the off state costs. Your explicit value is never overwritten: this file is
# a conffile and the daemon always writes the option back as '1'/'0'.
#
# NOTE the interaction: with this ON, `untunnelable` no longer governs ping at
# all (the L3 route decision happens before the firewall chain its verdicts
# live in). It still governs ESP/AH/GRE/IGMP/SCTP, which no tunnel of ours can
# carry. `untunnelable 'icmp'` in particular stops meaning "block plus working
# ping" and is reported as such.
option l3_tunnel '1'
option ipv6 '1'
# Reserved fwmark base and routing-table base (do not overlap fw4/other apps).
option fwmark_base '0x2000'
option table_base '0x2000'
# Seconds to auto-rollback an unconfirmed apply (0 = commit-confirm off).
# Seconds to auto-rollback an unconfirmed apply. SHIPPED AS 0, i.e.
# commit-confirm is OFF: `shaterd apply` arms nothing, and an apply that costs
# you SSH/LuCI access stays until you undo it by hand. Set a window (e.g.
# '120') to arm it, and run `shaterd confirm` inside that window to keep the
# new config. Note the option is written back only when NON-zero, so an
# explicit '0' disappears from this file on the first write by the daemon or
# the panel — absent and 0 are the same thing.
option confirm_timeout '0'
option schema_version '1'
# Master enable of the DNS blocklist/allowlist filter (D15). OFF by default;
# it needs at least one `config resolver` to have a DNS plane to filter with.
# See the "DNS filter" section at the end of this file.
option dns_filter '0'
option schema_version '2'
# LAN interception inbound. `network` is a UCI interface name; shaterd resolves
# it to its device (e.g. 'lan' -> br-lan) for the nft TPROXY plane. Enable
# globals above and adjust `network` to the interface(s) you want proxied.
#
# There is no per-inbound `sniff` option: since sing-box 1.11 sniffing is a
# leading route ACTION rule with no inbound matcher, so EVERY inbound is sniffed,
# always. Do not add one back — the hijack-dns rule matches the SNIFFED `dns`
# protocol, so a per-inbound sniff toggle would be a DNS-leak switch (D14, and
# the long argument at shater/model/model.go Inbound).
config inbound
option name 'lan'
option enabled '1'
@@ -68,7 +109,6 @@ config inbound
option tproxy_port '12345'
option tcp '1'
option udp '1'
option sniff '1'
# --- Commented examples (copy, uncomment, adjust, then enable globals) -------
#
@@ -153,8 +193,8 @@ config inbound
#
# --- DNS filter (D15) -------------------------------------------------------
# Network-wide domain blocking, built on sing-box rule-sets + reject DNS rules.
# Turn it ON by setting `option dns_filter '1'` in `config globals` above (it is
# OFF by default). Filtering needs at least one `config resolver` (the in-engine
# Turn it ON by flipping `option dns_filter` to '1' in `config globals` above (it
# is shipped '0'). Filtering needs at least one `config resolver` (the in-engine
# DNS plane). A blocklist answers matched domains with NXDOMAIN; an allowlist
# always OVERRIDES the blocklists (allowlisted domains resolve normally).
#
+259 -1
View File
@@ -33,6 +33,30 @@
# be running. `start` raises ACTIVE_FLAG, `stop` clears it; hotplug/cron
# reconcile ONLY while the flag is up, so an admin `stop` STICKS — no
# background actor may resurrect interception behind a stopped daemon.
# * BEING REPLACED IS NOT BEING SWITCHED OFF. `restart` and `reload` (which is
# stop+start, i.e. every LuCI Save & Apply) both run through `stop`, and the
# daemon's SIGTERM teardown removes the fail-closed table unconditionally — it
# does not consult kill_switch at all. Between that teardown and the
# successor's first apply the init GUARANTEES a gap: it waits for the old
# process to exit (shater_wait_stopped), then runs `shaterd migrate`, then
# starts a daemon that still has to build an engine. So a restart is announced
# with RESTART_FLAG, which tells the outgoing daemon to leave the fail-closed
# holding plane STANDING — apply.TeardownExiting swaps it in with one nft
# transaction and then skips the delete, so the table is never absent, not even
# for the 80-90 ms the old arm-after-teardown order measured. A real `stop`
# raises no flag and therefore still means what it says.
# (A package UPGRADE does not come through here at all on apk v3: shater-core's
# script table is post-install / pre-deinstall / post-upgrade, with no
# pre-upgrade, so default_prerm — and its `stop` — runs only on REMOVAL.)
# * The FAIL-CLOSED PLANE MUST ALSO EXIST BEFORE THIS SCRIPT DOES. START=99 is
# after fw4 (19) and netifd (20), so at every boot the LAN forwards to the WAN
# in the clear for as long as it takes procd to decompress the daemon off
# flash and get an engine up. /etc/init.d/shater-armor (START=21) loads
# BOOT_ARMOR — a copy of the holding plane the daemon persists on every apply
# — to close that window. This script owns the DISARM half, and it owns it
# with a CLOSED LIST: an operator's `stop`, or a removal, and nothing else.
# Powering the box down must not — `shutdown` reaches stop_service too, and it
# is not a person switching the product off (see shater_stop_disarms).
# * The engine must never be permanently abandoned while interception stands:
# respawn retries are infinite (procd never gives up); a sustained-dead
# daemon is additionally escalated by the shater-cron watchdog.
@@ -50,11 +74,166 @@ ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Raised around a restart/reload, read by the OUTGOING `shaterd run` at SIGTERM:
# present => "you are being replaced, leave the fail-closed plane standing";
# absent => "you are being switched off, take everything down". tmpfs, so a
# power cut can never make the next boot look like a restart.
RESTART_FLAG=/var/run/shater.restarting
# The persisted fail-closed holding plane. Written by the daemon on every apply,
# loaded by /etc/init.d/shater-armor at boot. Its PRESENCE is the arm token, so
# removing it here is how a deliberate stop stops the next boot from blocking.
BOOT_ARMOR=/etc/shater/boot.nft
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# WHICH ACTION rc.common was invoked with, frozen at source time.
#
# rc.common does, in this order:
# initscript=$1; action=${2:-help}; shift 2; ...; . "$initscript"; $action "$@"
# so `action` is ALREADY assigned when this file is sourced, and every action then
# runs as a function in THAT SAME shell. MEASURED on the target (ImmortalWrt
# 25.12.1 r37978) with a throwaway probe init script, not read off documentation:
#
# /etc/init.d/X restart -> stop_service action=[restart], start_service [restart]
# /etc/init.d/X stop -> stop_service action=[stop]
# /etc/init.d/X reload -> reload_service action=[reload]
# `reboot` -> stop_service action=[SHUTDOWN] <-- see below
# the boot after it -> start_service action=[boot]
#
# A previous probe reported this variable EMPTY and the emptiness was written up as
# the defect. It was the probe: `sh -x /etc/init.d/shater restart` bypasses the
# `#!/bin/sh /etc/rc.common` shebang, so rc.common never runs, never assigns
# `action`, and the variable reads empty no matter what this file does.
#
# Frozen into our own variable because `action` is a short, generic name that other
# framework helpers also use as a local; a snapshot taken before any function runs
# cannot be shadowed later.
SHATER_RC_ACTION="$action"
# --- what an action MEANS --------------------------------------------------
#
# THE BUG THESE TWO PREDICATES REPLACE (v0.2.17, measured on the live router).
# The old stop_service was `case $action in restart|reload) keep;; *) DISARM;; esac`
# — an open default that swept up every action nobody had enumerated. `reboot` is
# one of them: procd runs the K-links with the action `shutdown`, so the shutdown
# path deleted the arm token on the way down and the next boot had nothing to load.
# The mechanism destroyed itself at exactly the moment it exists for. Instrument
# reading from the router, one minute apart across a reboot:
#
# 13:28 /etc/shater/boot.nft present
# ---- reboot (stop_service action=[shutdown] -> old `*` branch -> rm)
# 18s at_S22: NO_TABLE armor_file=NO_FILE
#
# So both lists below are POSITIVE and CLOSED. An action nobody thought about —
# `shutdown` above all, but also whatever a future procd invents — falls through
# both and changes nothing. The default now fails in the recoverable direction: at
# worst a boot arms when it need not have, which costs the second before the daemon
# applies and is still gated by shater-armor's own four state refusals. The old
# default failed in the direction of the plaintext window the feature was built to
# close.
#
# They are predicates rather than an inline `case` so the test gate can execute the
# real thing: it sources THIS FILE in /bin/sh and calls them with every action procd
# actually uses (shater/cmd/shaterd/initscript_test.go). A comment claiming
# `shutdown` is handled is what shipped last time.
# True only for the ONE action that means "the operator switched the product off".
# Deliberately not `shutdown`: powering a router down is not turning a feature off.
#
# NOT sufficient on its own — see shater_stop_disarms. `stop` is also how the
# package manager's plumbing reaches us, and a package manager is not a person.
shater_action_disarms() {
case "$1" in
stop) return 0 ;;
*) return 1 ;;
esac
}
# Is a package manager in the middle of a transaction RIGHT NOW?
#
# This is a state, read at the moment the decision is made, exactly like
# shater-armor's four refusals — not a record of an event. The same question is
# already asked (for the same reason: prerm/postinst plumbing is not a user
# action) by the detached bring-up in /etc/uci-defaults/30_shater-core.
shater_pkg_transaction() {
pidof apk >/dev/null 2>&1 && return 0
pidof opkg >/dev/null 2>&1 && return 0
return 1
}
# Is the main service still enabled at boot? Same glob, and for the same reason,
# as shater-armor's own check: `/etc/init.d/shater enabled` would source procd.sh
# and take a blocking flock, which is not something to do from inside a package
# manager's transaction.
shater_rc_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
# THE ACTUAL DISARM DECISION.
# $1 = action
# $2 = 1 when a package transaction is in flight
# $3 = 1 when the service is still enabled in rc.d
# All three are passed in rather than read inside, so the gate can drive every
# combination without a package manager or an /etc/rc.d.
#
# WHY IT IS NOT JUST THE ACTION. base-files' default_prerm runs, in this order:
#
# if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
# "$i" stop
#
# so a package manager reaches stop_service wearing the operator's clothes. Two
# different intentions arrive as the same action, and the difference between them
# is readable at the moment of the decision:
#
# REMOVAL — prerm has ALREADY run `disable`, so S99shater is gone. The product
# is going away; the armor goes with it. (It is belt-and-braces even
# so: shater-armor refuses to arm without that symlink, and the whole
# init script is about to be deleted anyway.)
# REPLACED — the service is still enabled, so something intends to bring it
# back. That is not an operator switching anything off, and deleting
# the armor here would leave the next boot unprotected. "The next
# apply will rewrite it" is not an answer: the armor exists precisely
# to cover a reboot, and a reboot between an update and the first
# apply is how this product is deployed.
#
# MEASURED, because the paragraph above is about a path I got wrong once already.
# On THIS target (apk-tools 3.0.5, ImmortalWrt 25.12.1) shater-core's script table
# is post-install / pre-deinstall / post-upgrade, with NO pre-upgrade — so an apk
# UPGRADE never executes default_prerm and never calls `stop` at all. Verified with
# a real `apk fix --reinstall shater-core` while sampling the armor file: 245 625
# samples, zero disappearances, even with this guard mutated off. The upgrade half
# of this predicate is therefore defence-in-depth for a shape that is one
# `pre-upgrade` script (or a returning opkg lane) away, NOT a fix for an observed
# failure. The removal half is live today.
shater_stop_disarms() {
shater_action_disarms "$1" || return 1
# No package manager involved => a person typed it. The escape hatch must work.
[ "$2" = "1" ] || return 0
# A package transaction that has NOT disabled the service is replacing it.
[ "$3" = "1" ] && return 1
return 0
}
# True when a successor is coming, so the outgoing daemon should leave the
# fail-closed holding plane standing instead of removing it.
#
# `shutdown` is deliberately NOT a handoff either: nothing is coming, and the
# kernel that would hold the plane is going away with it. Leaving the flag down
# there also keeps the marker's meaning exact — it says "you are being replaced",
# and at shutdown nothing is.
shater_action_handoff() {
case "$1" in
restart|reload) return 0 ;;
*) return 1 ;;
esac
}
# --- helpers ---------------------------------------------------------------
# True only when the stack is explicitly enabled in UCI.
@@ -73,6 +252,25 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Announce/withdraw "this daemon is being replaced, not switched off". Read by
# `shaterd run` when it receives SIGTERM.
shater_mark_restart() {
mkdir -p "$(dirname "$RESTART_FLAG")" 2>/dev/null
: > "$RESTART_FLAG"
}
shater_clear_restart() { rm -f "$RESTART_FLAG"; }
# Remove the persisted boot armor, so the LAN is NOT blocked at the next boot
# before the daemon starts. Called from exactly two places, both of which are a
# statement about the PRODUCT rather than about this process: an operator typing
# `stop`, and a daemon binary that is no longer on the box. In neither case is
# anything going to come along and replace the armor with a real data plane, and a
# kill switch with nothing behind it is just a brick.
#
# NOT called on the shutdown path. That is the whole fix — see
# shater_action_disarms.
shater_disarm_boot() { rm -f "$BOOT_ARMOR"; }
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
@@ -136,9 +334,20 @@ start_service() {
# Guard: never claim to run without the daemon binary. A half-removed/failed
# shaterd upgrade must degrade to "plugin off", not to a box that thinks
# interception is live with nothing behind it.
#
# "Plugin off" now has to include DISARMING. With the boot armor in play, a
# missing binary is the one case where the fail-closed plane could stand
# forever with nothing able to replace it: the armor loads at START=21, the
# daemon never starts, and every later boot repeats it. The product being gone
# is not a security event — it is an uninstall — so the plane comes down and
# the LAN returns to plain routing, loudly.
if [ ! -x "$PROG" ]; then
shater_clear_restart
shater_disarm_boot
rm -f "$ACTIVE_FLAG"
nft delete table inet shater 2>/dev/null
_slog -p daemon.err \
"shaterd binary missing/not executable at $PROG — refusing to start (LAN stays on plain routing)"
"shaterd binary missing/not executable at $PROG — refusing to start; the fail-closed plane and its boot armor have been REMOVED (LAN back to plain routing, unprotected). Reinstall shaterd."
return 0
fi
@@ -150,6 +359,11 @@ start_service() {
# running, which is the boot case.
shater_wait_stopped
# The predecessor is gone and has already consumed the flag (it reads it in its
# SIGTERM handler). Withdraw it now, so a LATER `stop` is unambiguous even if
# this start fails further down.
shater_clear_restart
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
#
@@ -214,6 +428,45 @@ start_service() {
}
stop_service() {
# Say WHY we are stopping before procd sends the signal, because the daemon
# cannot tell from the signal alone and the answer changes what it leaves in
# the kernel. Two INDEPENDENT questions, and the old code conflated them into
# one two-armed `case` whose else-branch answered both wrongly for `shutdown`:
#
# 1. IS A SUCCESSOR COMING (this process only)? restart / reload.
# Raise RESTART_FLAG so the outgoing daemon replaces its data plane with
# the fail-closed HOLDING plane instead of removing it. The gap until the
# successor applies is not a moment: this script waits out the old
# process, runs `shaterd migrate`, then starts a daemon that must build an
# engine — all of it, before this flag existed, with `lan -> wan ACCEPT`
# and nothing else.
#
# 2. IS THE PRODUCT BEING SWITCHED OFF (across boots)? `stop` — and only
# `stop`, and only when a PERSON is behind it (shater_stop_disarms; the
# package manager reaches us through `stop` too). Then the boot armor goes
# with it, so the next boot does not quietly reinstate what the operator
# just switched off — the same rule ACTIVE_FLAG has always enforced for
# hotplug/cron.
#
# `shutdown` answers NO to both, which is the defect this replaced: a reboot is
# not a successor and it is certainly not an operator switching the product off.
# It is the boot the armor exists for. An upgrade answers NO to the second for
# the same kind of reason.
if shater_action_handoff "$SHATER_RC_ACTION"; then
shater_mark_restart
else
shater_clear_restart
fi
local in_pkg=0 rc_en=0
shater_pkg_transaction && in_pkg=1
shater_rc_enabled && rc_en=1
if shater_stop_disarms "$SHATER_RC_ACTION" "$in_pkg" "$rc_en"; then
shater_disarm_boot
elif [ "$in_pkg" = "1" ] && shater_action_disarms "$SHATER_RC_ACTION"; then
_slog -p daemon.info \
"stop came from a package transaction that left the service enabled — keeping the boot armor, so being replaced cannot leave the next boot unprotected"
fi
# Drop the live-flag FIRST so a concurrent hotplug/cron tick cannot rebuild
# what we are about to tear down. procd then sends SIGTERM to `shaterd run`,
# which runs its OWN honest teardown (engine.Close + netplane restore) — we
@@ -233,6 +486,11 @@ reload_service() {
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
#
# Marked EXPLICITLY as well as via SHATER_RC_ACTION: this is the path a routine
# Save & Apply takes, so it is the one that must not depend on reading an
# rc.common variable correctly. Belt and braces, one line.
shater_mark_restart
stop
start
}
@@ -0,0 +1,162 @@
#!/bin/sh /etc/rc.common
# /etc/init.d/shater-armor — the fail-closed plane, before the daemon exists.
#
# WHAT THIS CLOSES
#
# /etc/init.d/shater is START=99. By then fw4 (START=19) has long since loaded
# `lan -> wan ACCEPT` and netifd (START=20) has brought the LAN bridge up, so the
# router forwards LAN traffic to the WAN in the clear from the moment the link
# comes up until `shaterd run` has been decompressed off flash, has waited out any
# predecessor, has migrated UCI, has read the config and has installed its first
# table. On router-class hardware with a UPX-packed binary that is seconds — and
# they are exactly the seconds in which Wi-Fi finishes associating and every
# client on the network reconnects and starts talking. `kill_switch=closed` was
# configured the whole time and covered none of it.
#
# There was nothing in the package that could cover it either: no /etc/nftables.d
# include, no `nft -f` in uci-defaults. Protection existed only inside a Go
# process that had not started yet.
#
# HOW
#
# The daemon persists a copy of its fail-closed HOLDING plane (the same ruleset it
# installs when the engine is down: one forward chain, LAN-to-LAN and router
# traffic accepted, everything else from the diverted devices dropped) to
# $ARMOR on every apply. This script loads it early. When the daemon comes up it
# replaces the table atomically — the ruleset begins with `delete table` and adds
# its own in one netlink transaction — so there is never a moment with no table.
#
# `iifname` matches by NAME at packet time, not by ifindex at load time, so
# loading this before netifd has created br-lan is fine: the rules simply start
# matching when the device appears. That is why START can sit here rather than
# racing netifd.
#
# START=21: after fw4 (19) and netifd (20), because fw4's own start tears its
# table down and rebuilds it and we do not want to be in the middle of that, and
# because there is nothing to protect before the LAN device is being created. The
# residual exposure is the fraction of a second between netifd's `ifup` and this
# script, against seconds-to-a-minute before.
#
# THE ESCAPE HATCHES (a kill switch that cannot be switched off is a brick)
#
# These are STATE checks, evaluated here, at the moment of arming — not a record
# of something that happened on the way down. That distinction is the whole
# lesson of v0.2.17: the arm token was deleted by an EVENT on the shutdown path
# ("this looks like a stop"), and since `reboot` also runs the K-links, the
# mechanism reliably erased itself on the one transition it was built for. An
# event on the way down cannot be trusted to describe the world on the way up; a
# question asked on the way up can be.
#
# * $ARMOR only exists while the daemon's last applied config was BOTH enabled
# and fail-closed. `globals.enabled=0` and `kill_switch=open` each remove it
# at the next apply, and an operator typing `/etc/init.d/shater stop` removes
# it there and then. Powering the box off does NOT.
# * We refuse to arm when the main service is disabled in rc.d, or when the
# daemon binary is gone — in either case nothing would ever come along to
# replace the armor with a real data plane. These two are what makes a
# genuinely uninstalled/disabled product safe REGARDLESS of what the file
# says, which is why they are checked here rather than trusted to have been
# acted on earlier.
# * We refuse to arm when UCI can be read AND says the stack is disabled. A
# config that cannot be read is NOT a refusal: that case is precisely why the
# armor is a file rather than a query.
# * The chain hooks `forward` only, so SSH, LuCI and the admin panel (all input
# hook, to the router's own addresses) stay reachable. The operator can always
# get in and undo this.
#
# Note what a bare `/etc/init.d/shater stop` does NOT mean: it does not survive a
# reboot, because S99shater is still linked and procd starts the daemon again. So
# "stopped" is not a durable off-state and this script must not be designed as if
# it were — the durable ones are `disable` (no S??shater) and `globals.enabled=0`,
# and those are the two refusals above.
#
# busybox ash only — no bashisms.
START=21 # after firewall (19) and network (20), long before shater (99)
STOP=89
ARMOR=/etc/shater/boot.nft
PROG=/usr/bin/shaterd
# Syslog line that honors globals.log_syslog, like the other two inits. An
# unreadable UCI leaves the option empty => ON, which is what we want here: the
# one boot where the config cannot be read is the boot worth logging.
_slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater-armor "$@"
}
# Is the MAIN service enabled at boot? Answered by looking for its rc.d symlink
# rather than by running `/etc/init.d/shater enabled`: that is a USE_PROCD script,
# so every action of it sources procd.sh, which takes a blocking flock — and this
# runs at START=21, in the middle of boot, for a question a glob answers exactly
# as well. The START number is not hardcoded; any S<NN>shater counts.
shater_service_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
start() {
# No saved plane => the stack has never applied an enabled, fail-closed config
# (or it was explicitly switched off). Nothing to do, and nothing to say.
[ -f "$ARMOR" ] || return 0
[ -s "$ARMOR" ] || {
_slog -p daemon.err "$ARMOR is empty — NOT arming; the LAN is unprotected until shaterd starts"
return 0
}
# Never arm something nothing can disarm.
[ -x "$PROG" ] || {
_slog -p daemon.err \
"$PROG is missing — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
shater_service_enabled || {
_slog -p daemon.warn \
"the shater service is disabled in rc.d — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
# A READABLE config that says "off" wins over the saved plane (it means the
# daemon was stopped before it could disarm). An UNREADABLE config does not:
# that is the case this whole mechanism exists for.
en=$(uci -q get shater.globals.enabled 2>/dev/null)
if [ -n "$en" ] && [ "$en" != "1" ]; then
rm -f "$ARMOR"
_slog -p daemon.info "globals.enabled=$en — boot armor removed, not arming"
return 0
fi
command -v nft >/dev/null 2>&1 || {
_slog -p daemon.err "nft is not installed — cannot arm; the LAN is unprotected until shaterd starts"
return 0
}
# Validate before loading: a truncated/incompatible snapshot must not leave a
# half-built table behind on the one boot it is needed.
if ! nft -c -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.err \
"$ARMOR did not validate (nft -c) — NOT arming; the LAN is unprotected until shaterd starts"
return 0
fi
if nft -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.warn \
"fail-closed plane armed from $ARMOR: LAN->WAN forwarding is BLOCKED until shaterd applies. SSH, LuCI and the admin panel stay reachable."
else
_slog -p daemon.err \
"could not load $ARMOR — the LAN is unprotected until shaterd starts"
fi
return 0
}
stop() {
# Deliberately a NO-OP. By the time anything stops this service the daemon owns
# `inet shater`, and deleting the table here would dismantle a LIVE data plane
# on the strength of a service that only ever ran for one second at boot. The
# disarm paths that matter live where the decision is actually made:
# /etc/init.d/shater stop (operator switched it off) and the daemon itself
# (globals.enabled=0 / kill_switch=open).
return 0
}
+270 -28
View File
@@ -3,11 +3,12 @@
# plus the data-plane watchdog (v0.2).
#
# A tiny procd-supervised loop that, once per tick, checks every enabled
# subscription and url-ruleset against its per-item `update_interval` and runs
# subscription against its per-item `update_interval` and runs
# shaterd sub update <name> (subscriptions)
# shaterd ruleset update <name> (url rulesets)
# when the item is due, then a single `shaterd reconcile` if anything changed
# (the daemon's config-hash gate rebuilds the engine only on a real change).
# `config ruleset` items are NOT touched here — see shater_run_due for who owns
# their refresh and where the remaining gap is.
#
# RELIABILITY CONTRACT (same "железно" posture as /etc/init.d/shater):
# * The loop body is fully INERT unless globals.enabled=1 AND the main shater
@@ -27,6 +28,15 @@
# the main service is STOPPED (tears interception down — fail-open, LAN
# returns to plain routing); with kill_switch=closed the rules stay
# (blocked-by-design) and we log loudly.
# * CRASH-LOOP WATCHDOG: the dead-daemon counter above cannot see the failure
# it matters most for. /etc/init.d/shater sets `respawn 3600 5 0`, so a
# daemon that dies a few seconds into startup is back within 5s and a single
# `pidof` per 60s tick nearly always finds a process — the counter resets and
# never reaches WATCHDOG_TICKS, while the fail-closed plane keeps the LAN shut
# and the panel (served BY the daemon) never comes up. So the tick's sleep is
# spent SAMPLING the daemon's identity instead of sleeping blind, and a tick in
# which several different daemons lived is counted as churn. See
# shater_churn_scan / shater_churn_verdict / shater_churn_action.
# * The loop never self-exits (procd would respawn-churn an exiting body);
# it idles on its guards instead. busybox ash only — no bashisms.
@@ -42,14 +52,34 @@ INIT_SCRIPT=/etc/init.d/shater-cron
SHATER_INIT=/etc/init.d/shater
SHATERD=/usr/bin/shaterd
ACTIVE_FLAG=/var/run/shater.active
# Raised by /etc/init.d/shater around a restart/reload and cleared by the
# successor's start_service. Read here ONLY as a "a person/package asked for this
# bounce" veto on the crash-loop verdict — never as a liveness signal.
RESTART_FLAG=/var/run/shater.restarting
# Written by `shaterd run` itself (main.go writePidfile) before it builds anything,
# and removed by that same process on a clean exit. It is the only handle that
# names THE daemon: `pidof shaterd` also matches the short-lived CLI verbs this
# very loop runs (`sub update`, `reconcile`, `schedule due`).
PIDFILE=/var/run/shaterd.pid
STAMP_DIR=/var/run/shater/cron
TICK=60 # seconds between due-checks
RETRY_SECS=300 # backoff before retrying a FAILED fetch
WATCHDOG_TICKS=5 # consecutive dead-daemon ticks before escalating
DEFAULT_SUB_INTERVAL=6h
DEFAULT_RS_INTERVAL=24h
DEFAULT_BL_INTERVAL=24h # url blocklist refresh interval (D16)
# --- crash-loop watchdog tuning --------------------------------------------
#
# Every number here is chosen against ONE question: what can a legitimate restart
# produce? A legitimate bounce (`restart`, LuCI Save & Apply -> reload, a package
# transaction) replaces the daemon EXACTLY ONCE, and it is announced twice over —
# /etc/init.d/shater raises RESTART_FLAG in stop_service and clears ACTIVE_FLAG for
# the duration. A crash loop is announced by nothing and repeats without bound.
LOOP_POLL=5 # seconds between identity samples inside one tick
LOOP_MIN_GENS=3 # distinct daemons in ONE tick that count as churn
LOOP_WINDOWS=2 # consecutive churn ticks before we call it a loop
LOOP_REPORT_TICKS=30 # do not repeat the report more often than this
# --- helpers ---------------------------------------------------------------
shater_enabled() {
@@ -133,7 +163,7 @@ shater_stamp_retry() {
# Walk anonymous `config subscription` / `config ruleset` sections by index and
# run any that are due. Echoes non-empty on stdout if at least one item updated.
shater_run_due() {
local i name en ivl secs stamp src changed=""
local i name en ivl secs stamp changed=""
# Subscriptions.
i=0
@@ -159,28 +189,36 @@ shater_run_due() {
i=$(( i + 1 ))
done
# Rulesets (only url sources auto-update; others have nothing to fetch).
i=0
while uci -q get "shater.@ruleset[$i]" >/dev/null 2>&1; do
name=$(uci -q get "shater.@ruleset[$i].name")
src=$(uci -q get "shater.@ruleset[$i].source")
if [ -n "$name" ] && [ "$src" = "url" ]; then
ivl=$(uci -q get "shater.@ruleset[$i].update_interval")
secs=$(shater_ivl_secs "$ivl" "$DEFAULT_RS_INTERVAL")
stamp="$STAMP_DIR/rs.$(shater_safe_name "$name")"
if shater_due "$stamp" "$secs"; then
if "$SHATERD" ruleset update "$name" >/dev/null 2>&1; then
shater_stamp "$stamp"
changed=1
else
_slog -p daemon.warn \
"ruleset update '$name' failed; retrying in ${RETRY_SECS}s"
shater_stamp_retry "$stamp" "$secs"
fi
fi
fi
i=$(( i + 1 ))
done
# RULE-SETS ARE NOT UPDATED FROM HERE, AND NEVER WERE.
#
# There used to be a second loop that ran `shaterd ruleset update <name>` for
# every `config ruleset` with source=url. That verb has never existed: it
# printed a note and exited 0, so this loop stamped the item as freshly updated
# and raised `changed`, which cost a reconcile per item per interval and told
# the operator the list was current when not one byte had been fetched. The verb
# now exits non-zero (shater/cmd/shaterd/main.go, notImpl), which turns the same
# loop into one failed attempt and one syslog line every RETRY_SECS — ~288 lines
# a day, per rule-set, about work that has no owner here. Noise in the log hides
# real problems as effectively as a lie about success does.
#
# WHO REFRESHES A url RULE-SET NOW, so the next reader does not think this was
# forgotten. `source=url` splits into two shapes in shater/generate/ruleset.go:
#
# * the URL serves an engine-native .srs/.json -> it stays a REMOTE rule-set
# and sing-box owns fetch/cache/refresh through RemoteRuleSet.UpdateInterval
# on the running box. This cron loop never had anything to contribute.
# * the URL serves a plain-text list -> it is compiled locally into
# /etc/shater/lists/<tag>.srs, and that artifact is refreshed by the
# GENERATOR, "when missing or older than update_interval" — i.e. only when
# something else already caused a generate. Nothing schedules one, so this
# shape has NO periodic refresh at all today. That is a real gap, and it is
# stated here rather than papered over with a call to a verb that does
# nothing: closing it needs a daemon-side timer (or a real `ruleset update`),
# not a shell loop, because only the daemon can force a rebuild past the
# config-hash gate.
#
# `config blocklist` url items are a different mechanism and DO refresh — see
# shater_run_due_blocklists below.
[ -n "$changed" ] && echo 1
}
@@ -256,6 +294,171 @@ shater_watchdog() {
echo "$dead"
}
# --- crash-loop watchdog ----------------------------------------------------
#
# THE HOLE. shater_watchdog above answers "is a daemon there?" once every TICK
# seconds. /etc/init.d/shater sets `respawn 3600 5 0`, so a daemon that dies a few
# seconds into startup is back 5s later and that one sample nearly always finds a
# process: the dead-counter resets, never reaches WATCHDOG_TICKS, and the single
# failure the watchdog exists for — a new binary or a bad config that cannot get
# an engine up while the fail-closed plane holds the LAN shut — is the one it can
# never see. The daemon also serves the panel, so in that state the operator has
# neither internet nor a way to look at the box.
#
# THE SIGNAL. Not "is it there" but "is it the SAME one". The tick's sleep is
# spent taking an identity sample every LOOP_POLL seconds instead of sleeping
# blind, and a tick in which LOOP_MIN_GENS different daemons lived is a churn
# tick. LOOP_WINDOWS consecutive churn ticks is the verdict.
#
# WHY A LEGITIMATE RESTART CANNOT REACH IT. Four independent reasons, in order of
# how much they are relied on:
#
# 1. A bounce replaces the daemon ONCE. One restart scores 2 generations in the
# tick it happens in and 1 in every tick after, so it cannot even produce a
# single churn tick at LOOP_MIN_GENS=3, let alone LOOP_WINDOWS of them in a
# row. Reaching the verdict takes >= 4 replacements inside 2 consecutive
# minutes, >= 2 in each.
# 2. Bounces are ANNOUNCED. /etc/init.d/shater raises RESTART_FLAG in
# stop_service and clears ACTIVE_FLAG for the whole stop->start, and either
# one seen in any sample of a tick discards that tick outright.
# 3. The panel's Apply does not restart anything: it writes UCI and applies over
# the daemon's control socket, in process. Only `restart`, a LuCI Save &
# Apply (config.change -> reload) and a package transaction bounce the
# daemon, and a human cannot produce those at four a minute.
# 4. The sample names THE daemon via its pidfile, not `pidof shaterd` — the
# short-lived CLI verbs this very loop runs share that process name.
#
# WHAT IT DOES NOT COVER, stated rather than implied: a daemon that dies
# INSTANTLY (well under a second) is almost never caught alive by a 5s sample, so
# it scores few generations and this detector stays quiet. That case is exactly
# the one the existing dead-tick counter does see — its `pidof` misses too, tick
# after tick — so the two cover opposite ends and are deliberately left as two
# independent instruments rather than merged into one clever number.
# One identity sample: echoes the pid of the live `shaterd run`, or "-" for none.
#
# Through the PIDFILE, which `shaterd run` writes before it builds anything and
# removes on a clean exit, because that is the only handle that names THE daemon:
# `pidof shaterd` also matches `shaterd sub update` / `reconcile` / `schedule due`.
# /proc/<pid>/comm is checked so a stale pidfile whose pid has been reused by an
# unrelated process cannot read as a live daemon. No forks: `read` is a builtin.
shater_sample_pid() {
local pid="" comm=""
[ -r "$PIDFILE" ] && read -r pid 2>/dev/null < "$PIDFILE"
case "$pid" in
''|*[!0-9]*) echo -; return ;;
esac
[ -r "/proc/$pid/comm" ] && read -r comm 2>/dev/null < "/proc/$pid/comm"
[ "$comm" = "shaterd" ] || { echo -; return; }
echo "$pid"
}
# shater_churn_scan <sample>... -> "<generations> <absent-samples>"
#
# A GENERATION is one distinct daemon lifetime observed during the tick: a live
# pid that differs from the last live pid seen. A daemon that simply keeps running
# therefore scores exactly 1 generation and 0 absent samples for as long as it
# runs — the signal is flat unless something is actually being replaced.
#
# A GAP (samples with no daemon at all, e.g. procd's 5s respawn hole) is counted
# but does NOT by itself open a new generation: only a different pid does. An
# earlier draft reset the comparison across a gap so that "same pid seen again
# after a gap" would score two. That case cannot occur — a respawn always gets a
# fresh pid — so it was unfalsifiable code, and resetting also meant a momentarily
# unreadable pidfile could inflate the count. Not resetting is both simpler and
# the safer direction.
#
# Pure: no I/O, no globals, every input on the command line. That is what lets the
# gate drive it with synthetic sample streams instead of a live router.
shater_churn_scan() {
local gens=0 absent=0 last="" s
for s in "$@"; do
if [ "$s" = "-" ]; then
absent=$(( absent + 1 ))
continue
fi
[ "$s" = "$last" ] || gens=$(( gens + 1 ))
last="$s"
done
echo "$gens $absent"
}
# shater_churn_verdict <gens> <samples> <announced> <churn-so-far>
# -> the new consecutive-churn-tick count
#
# Also pure. `announced`=1 means a sample during the tick saw RESTART_FLAG up or
# ACTIVE_FLAG down, i.e. /etc/init.d/shater said out loud that it was bouncing the
# daemon: that tick proves nothing and resets the run. A tick with no samples at
# all (the first pass through the loop) likewise scores 0 rather than guessing.
shater_churn_verdict() {
local gens="$1" n="$2" announced="$3" churn="$4"
[ "$announced" = "1" ] && { echo 0; return; }
[ "$n" -gt 0 ] || { echo 0; return; }
if [ "$gens" -ge "$LOOP_MIN_GENS" ]; then
echo $(( churn + 1 ))
return
fi
echo 0
}
# shater_churn_action <churn-ticks> <kill_switch> -> none | log | stop
#
# WHAT TO DO, and why it is not our call to make twice. A crash loop leaves the
# box in the same state a dead daemon does — no engine, fail-closed plane standing
# — so the answer is the one the operator already gave with kill_switch, not a new
# policy invented here:
#
# open The operator asked for connectivity over interception. Stop the stack,
# exactly as shater_watchdog does for a sustained-dead daemon: the plane
# comes down and the LAN returns to plain routing. It also disarms the
# boot armor, so the NEXT boot is clean too instead of repeating the loop
# behind a closed LAN. Nothing else can end the loop: procd's retries are
# infinite by design.
# closed The operator asked for blocked-over-leaking. Blocked is what they get,
# and opening their LAN from a background loop would be the opposite of
# what the knob says. Report it loudly and let the person decide; the
# message names the one command that opens it.
#
# The list is POSITIVE and CLOSED, and the fall-through goes to `log`: an absent
# or unrecognised kill_switch is the model's documented default ("closed", see
# shater/model/model.go DefaultGlobals), and `log` is the recoverable side — it
# changes nothing and can be acted on, where a wrong `stop` silently drops a
# household onto the unproxied WAN.
#
# NOTE (not changed here, deliberately): shater_watchdog above answers the same
# question with `if closed ... else stop`, so for an ABSENT kill_switch it fails
# open — the opposite of the documented default. It is left alone because that
# behaviour predates this file's crash-loop work; it is reported upward instead.
shater_churn_action() {
local churn="$1" ks="$2"
[ "$churn" -ge "$LOOP_WINDOWS" ] || { echo none; return; }
case "$ks" in
open) echo stop ;;
closed) echo log ;;
*) echo log ;;
esac
}
# Sleep out one tick in LOOP_POLL slices, sampling the daemon's identity as we go.
# Publishes CHURN_SAMPLES / CHURN_N / CHURN_ANNOUNCED for the next pass of loop().
# Deliberately NOT a subshell (globals must survive), and it always returns 0 so a
# false `[ -f ]` at the end cannot look like a failure.
shater_tick_sample() {
local slept=0
CHURN_SAMPLES=""
CHURN_N=0
CHURN_ANNOUNCED=0
while [ "$slept" -lt "$TICK" ]; do
sleep "$LOOP_POLL"
slept=$(( slept + LOOP_POLL ))
CHURN_SAMPLES="$CHURN_SAMPLES $(shater_sample_pid)"
CHURN_N=$(( CHURN_N + 1 ))
[ -f "$RESTART_FLAG" ] && CHURN_ANNOUNCED=1
[ -f "$ACTIVE_FLAG" ] || CHURN_ANNOUNCED=1
done
return 0
}
# loop: the foreground body supervised by procd. Never exits on its own — it
# idles while disabled/inactive so procd is not respawn-churned by a
# self-exiting body when the stack is off.
@@ -274,7 +477,12 @@ loop() {
# the flock immediately and keeps children (sleep/shaterd) from inheriting
# it. A no-op where fd 1000 is not open (older procd.sh without procd_lock).
exec 1000>&-
local changed dead=0
local changed dead=0 churn=0 quiet=0 scan gens absent ks act
# No tick has been sampled yet on the first pass; shater_churn_verdict scores
# an empty tick as 0 rather than guessing.
CHURN_SAMPLES=""
CHURN_N=0
CHURN_ANNOUNCED=0
mkdir -p "$STAMP_DIR"
while :; do
if shater_enabled && shater_active; then
@@ -299,10 +507,44 @@ loop() {
"$SHATERD" schedule due >/dev/null 2>&1
fi
dead=$(shater_watchdog "$dead")
# Crash-loop verdict on the tick that has just elapsed. Unquoted on
# purpose: CHURN_SAMPLES is a whitespace-separated token list and word
# splitting is how it becomes arguments.
scan=$(shater_churn_scan $CHURN_SAMPLES)
gens=${scan%% *}
absent=${scan##* }
churn=$(shater_churn_verdict "$gens" "$CHURN_N" "$CHURN_ANNOUNCED" "$churn")
ks=$(uci -q get shater.globals.kill_switch)
act=$(shater_churn_action "$churn" "$ks")
case "$act" in
stop)
_slog -p daemon.crit \
"shaterd is CRASH-LOOPING: $gens distinct daemons in the last ${TICK}s (absent in $absent of $CHURN_N samples), $churn such windows in a row — it is being respawned faster than it can bring an engine up. kill_switch=open, so shater is being STOPPED: interception comes down and the LAN returns to plain, UNPROXIED routing. Find the reason with 'logread -e shaterd', then '/etc/init.d/shater start'."
"$SHATER_INIT" stop
churn=0
quiet="$LOOP_REPORT_TICKS"
;;
log)
# Rate-limited: a standing condition, not an event. Never
# silent for good, though — an operator who looks at the log an
# hour later must still find it being said.
if [ "$quiet" -le 0 ]; then
_slog -p daemon.crit \
"shaterd is CRASH-LOOPING: $gens distinct daemons in the last ${TICK}s (absent in $absent of $CHURN_N samples), $churn such windows in a row — it is being respawned faster than it can bring an engine up. kill_switch=${ks:-closed} keeps the fail-closed plane standing, so the LAN stays blocked and the admin panel is down with the daemon that serves it. Nothing is decided for you: find the reason with 'logread -e shaterd', or open the LAN with '/etc/init.d/shater stop'."
quiet="$LOOP_REPORT_TICKS"
fi
churn=0
;;
esac
[ "$quiet" -gt 0 ] && quiet=$(( quiet - 1 ))
else
dead=0
churn=0
quiet=0
fi
sleep "$TICK"
# Sleeps out the tick, sampling the daemon's identity while it does.
shater_tick_sample
done
}
@@ -59,6 +59,199 @@ if uci -q get shater.globals >/dev/null 2>&1 || [ -f /etc/config/shater ]; then
uci -q commit shater
fi
# Introduce the daemon-created `shater-l3*` TUN to fw4 (L3 ingress, D-L3). The
# daemon policy-routes LAN ICMP into that device from OUR nft table
# `inet shater`, but nftables runs EVERY table on every packet and a drop in
# any one of them wins — an accept in `inet shater` cannot override fw4. And
# fw4 WILL drop this forward: netifd knows nothing about a device the daemon
# creates at runtime, so it belongs to no zone and falls into fw4's zone-less
# defaults (REJECT). The device has to be declared to fw4 itself; it cannot be
# fixed from our own table.
#
# Seeded UNCONDITIONALLY (not gated on globals.l3_tunnel): uci-defaults run
# once, so gating on the option would require re-running this script when the
# option is flipped later — which never happens. An idle zone is harmless: its
# device match is a plain iifname/oifname STRING compare that simply never hits
# while the TUN does not exist.
#
# Idempotency: `config zone`/`config forwarding` are normally ANONYMOUS
# sections, and a naive `uci add firewall zone` would append a duplicate on
# every re-run (uci-defaults re-run on package upgrade/reinstall). The zone is
# NAMED instead, guarded by an existence check — a re-run re-finds the section
# and touches nothing. The forwardings are named too where the name is free, but
# their guard is a scan of the actual src/dest pairs, which is stronger; see
# seed_l3_forwarding below.
seed_l3_zone() {
# No fw4 on this image (bare nftables build) => nothing drops the forward
# on fw4's behalf and there is nothing to punch through.
[ -f /etc/config/firewall ] || return 0
if ! uci -q get firewall.shater_l3 >/dev/null; then
uci set firewall.shater_l3=zone
uci set firewall.shater_l3.name='shater_l3'
uci set firewall.shater_l3.input='REJECT'
uci set firewall.shater_l3.output='ACCEPT'
uci set firewall.shater_l3.forward='REJECT'
uci set firewall.shater_l3.masq='0'
# INERT TODAY, kept for the day it is not. mtu_fix clamps forwarded TCP
# MSS to the route MTU — but the L3 TUN is 65535 (deliberately: at any
# smaller value the kernel fragments into the device, and the flow
# dispatcher refuses to judge a fragment and lets the stack forge the
# echo reply — see l3MTU in shater/generate/inbound.go), so the clamp has
# nothing to clamp to. And only ICMP is ever marked into this device, so
# no TCP rides here to be clamped in the first place. It earns its keep
# the moment either of those changes; removing it would make that day
# silent.
uci set firewall.shater_l3.mtu_fix='1'
# `list device`, deliberately NOT the usual `list network`: fw4
# resolves a zone's networks through netifd, and netifd never learns
# about a device the daemon creates at runtime — a stub interface
# (proto none) would need to be brought UP to contribute an l3_device,
# and nothing ever brings it up, so `list network` resolves to an
# EMPTY device set and fw4 keeps dropping the forward. `list device`
# instead compiles to an iifname/oifname STRING match, valid before
# the TUN exists and matching from the moment shaterd creates it —
# no netifd involvement and no firewall reload at enable time. Do not
# "normalize" this to `list network` in a refactor; it breaks silently.
#
# The WILDCARD is load-bearing too. The daemon no longer opens one fixed
# device: it alternates between `shater-l3a` and `shater-l3b` so that a
# new engine generation never has to reopen the name the previous one is
# still holding (that collision — TUNSETIFF: device or resource busy —
# took the whole LAN down on the production router, because the recovery
# path rebuilt the same config and hit the same busy name). fw4 compiles
# `shater-l3*` to `iifname "shater-l3*"` / `oifname "shater-l3*"`,
# verified on ImmortalWrt 25.12.1 with nftables 1.1.6, so ONE zone covers
# every slot and no firewall reload is needed when the slot changes.
uci add_list firewall.shater_l3.device='shater-l3*'
fi
seed_l3_forwardings
uci -q commit firewall
}
# EVERY zone gets a forwarding into shater_l3, not just `lan`.
#
# The bug this closes is silent by construction. The daemon's divert set is built
# from every enabled `config inbound`'s network PLUS every device a rule names
# through an `iface:`/`zone:` source (shater/netplane/nft.go, nftDivertRefs) — so
# on a router with several LAN zones, ICMP from ALL of them is marked and routed
# into the TUN by our table. Our table then accepts it and fw4 drops it anyway,
# because the forward is judged in `forward_<source zone>` and only `lan` had a
# jump to `accept_to_shater_l3`. Result: ping through the tunnel works from one
# subnet and not from the next, with nothing in any log to say why — fw4's drop
# is the zone's policy verdict, not a rule with a name. The owner's production
# router has a single LAN zone, which is exactly why this went unnoticed; his
# second router has three.
#
# Every zone, including an uplink zone, and that is deliberate rather than lazy:
#
# - The alternative is guessing which zones hold clients, and every available
# signal is wrong somewhere. `masq='1'` marks the WAN on a stock config and
# also marks a double-NAT LAN. The name `wan*` is a convention, not a rule.
# A guess that is wrong reintroduces exactly the silent breakage above, while
# a superfluous entry costs a line of ruleset.
# - A forwarding into shater_l3 permits nothing on its own. It authorises the
# forward of packets ROUTED INTO the TUN, and the only thing that routes a
# packet there is our own fwmark rule, which matches solely on the divert
# device set. A packet arriving on the WAN is not marked and never reaches
# this decision; if an operator ever puts a WAN device in the divert set,
# they meant to and this is the entry that makes it work.
# - The reverse direction is NOT opened: no `src shater_l3` forwarding exists,
# so nothing comes out of the TUN into a zone by way of these sections. The
# engine's own replies return on the conntrack `established,related accept`
# at the top of fw4's forward chain.
#
# LIMIT, stated because it is not obvious: this is a SNAPSHOT. uci-defaults run
# at first boot and on package install/upgrade, so a zone created AFTER the last
# shater-core install has no forwarding until the next one. Re-running this
# script (or reinstalling the package) re-seeds. The durable fix belongs in the
# daemon, which recomputes the divert set on every apply and already knows which
# zones are in it; it is deliberately not attempted from here.
seed_l3_forwardings() {
uci -q show firewall 2>/dev/null |
sed -n "s/^firewall\.\([^.=]*\)=zone\$/\1/p" |
while read -r sid; do
zone=$(uci -q get "firewall.$sid.name")
# Unnamed zone: fw4 cannot reference it from a forwarding either.
[ -n "$zone" ] || continue
# Our own zone: a forwarding from shater_l3 to itself is meaningless.
[ "$zone" = "shater_l3" ] && continue
seed_l3_forwarding "$zone"
done
}
# One `config forwarding` <zone> -> shater_l3, created only if no such forwarding
# exists yet.
#
# The guard scans the ACTUAL src/dest pairs rather than trusting a section id,
# which covers all three ways one can already be there: the legacy named section
# `shater_l3_fwd` seeded by earlier releases (src=lan), the per-zone names this
# function writes, and an anonymous one an operator added by hand. Without that,
# a re-run — uci-defaults re-run on every package upgrade — would append a
# duplicate for `lan` on every upgrade.
seed_l3_forwarding() {
local zone="$1" sid found
found=$(uci -q show firewall 2>/dev/null |
sed -n "s/^firewall\.\([^.=]*\)=forwarding\$/\1/p" |
while read -r f; do
[ "$(uci -q get "firewall.$f.dest")" = "shater_l3" ] || continue
[ "$(uci -q get "firewall.$f.src")" = "$zone" ] || continue
echo yes
break
done)
[ -n "$found" ] && return 0
# Section ids are [a-zA-Z0-9_] only, while a zone name may legally carry a
# hyphen — sanitise, and keep the legacy id for `lan` so an existing install
# is recognised as already seeded rather than gaining a second section.
if [ "$zone" = "lan" ]; then
sid="shater_l3_fwd"
else
sid="shater_l3_fwd_$(printf '%s' "$zone" | sed 's/[^a-zA-Z0-9_]/_/g')"
fi
# The id may still be taken — by a section for a DIFFERENT zone whose name
# sanitises to the same thing, or by something else entirely. Fall back to an
# anonymous section rather than overwrite: the src/dest scan above is what
# makes this idempotent, the name is only there to be readable.
if uci -q get "firewall.$sid" >/dev/null; then
sid=$(uci add firewall forwarding) || return 0
else
uci set "firewall.$sid=forwarding"
fi
uci set "firewall.$sid.src=$zone"
uci set "firewall.$sid.dest=shater_l3"
}
seed_l3_zone
# Upgrade path for routers seeded by a pre-slot build.
#
# The block above only runs when the zone does NOT exist, which is exactly right
# for idempotency and exactly wrong here: an already-installed router has the
# zone with the OLD exact device `shater-l3`, that name matches no slot, and fw4
# would go back to dropping the forward — i.e. LAN ping through the tunnel dies
# silently on upgrade while everything reports healthy. Rewrite it in place.
#
# Narrow on purpose: only the literal legacy entry is replaced, and only when the
# wildcard is not already listed, so an operator who added devices of their own
# keeps them and a re-run changes nothing (uci-defaults re-run on every package
# upgrade). No `fw4 reload` here — uci-defaults run before the firewall starts on
# boot, and on a package upgrade the daemon's next apply is what needs the zone,
# not this script.
migrate_l3_zone_wildcard() {
[ -f /etc/config/firewall ] || return 0
uci -q get firewall.shater_l3 >/dev/null || return 0
devs=$(uci -q get firewall.shater_l3.device) || return 0
case " $devs " in
*" shater-l3* "*) return 0 ;; # already migrated
*" shater-l3 "*) ;; # legacy exact name present
*) return 0 ;;
esac
uci -q del_list firewall.shater_l3.device='shater-l3'
uci add_list firewall.shater_l3.device='shater-l3*'
uci -q commit firewall
}
migrate_l3_zone_wildcard
# Bring the UCI schema forward on upgrade (idempotent; refuses a newer schema).
[ -x /usr/bin/shaterd ] && /usr/bin/shaterd migrate >/dev/null 2>&1
@@ -110,8 +303,27 @@ SHATER_BRINGUP='
done
[ -x /etc/init.d/shater ] && /etc/init.d/shater enable
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron enable
# The boot-time fail-closed armor. `enable` only — it is a one-shot that loads
# the persisted holding plane at START=21, and running it NOW would install a
# block on a live box moments before the daemon replaces it anyway. It has to
# be enabled here regardless of whether the stack is on: the file it loads only
# exists while the daemon wants it to, so an enabled-but-unarmed service is a
# no-op, and enabling it later would mean the first boot after an upgrade is
# the one boot still exposed.
[ -x /etc/init.d/shater-armor ] && /etc/init.d/shater-armor enable
[ -x /etc/init.d/shater ] && /etc/init.d/shater restart
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron restart
# Fold the seeded shater_l3 zone into the LIVE ruleset — matters on a live
# opkg/apk install only, where firewall started long before our commit and
# nothing else would re-read it until the next reboot. Gated on the fw4
# table actually being loaded: at FIRST boot this job can run before the
# S19 firewall start, and an early reload would install a ruleset built
# from a half-initialized netifd AND make the later start a no-op (fw4
# start skips when its table already exists). No table => the pending S19
# start reads the committed config by itself, no reload needed.
if nft list tables 2>/dev/null | grep -q "inet fw4"; then
[ -x /etc/init.d/firewall ] && /etc/init.d/firewall reload
fi
exit 0
'
SHATER_TMO=""
@@ -0,0 +1,80 @@
# /lib/upgrade/keep.d/shater-core — what sysupgrade and LuCI "Backup" must carry
# out of /etc/shater.
#
# HOW THIS FILE IS READ. /sbin/sysupgrade (base-files, list_static_conffiles):
#
# find $(sed -ne '/^[[:space:]]*$/d; /^#/d; p' \
# /etc/sysupgrade.conf /lib/upgrade/keep.d/* 2>/dev/null) \
# \( -type f -o -type l \) $filter 2>/dev/null
#
# so blank lines and lines starting with '#' are stripped, and every other line is
# a path handed to `find`: a directory is recursed, a path that does not exist is
# silently skipped (hence a trailing '/' for the two directories, and no need to
# guard for a fresh install that has neither). The result is tarred and, on a real
# sysupgrade, HELD IN RAM across the flash — which is why this is a per-file
# decision and not simply "/etc/shater/".
#
# WHY IT EXISTS. Everything the product knows besides /etc/config/shater lives in
# /etc/shater, and nothing shipped a keep.d entry for it. A "keep settings"
# sysupgrade, or a LuCI backup restored onto a new router, therefore produced a
# box whose config looked complete and whose node inventory was EMPTY — silently.
#
# /etc/config/shater is NOT listed here: it is declared in
# Package/shater-core/conffiles, and sysupgrade backs CHANGED conffiles up on its
# own (list_changed_conffiles). Listing it again would work, but it would claim
# ownership of a mechanism that already covers it.
# THE NODE INVENTORY. Subscription-fetched nodes deliberately live OUTSIDE UCI
# (shater/model/subcache.go) — one JSON file per subscription. Without them the
# restored box has groups and rules that reference nodes which do not exist, so no
# tunnel comes up, and the only repair is `sub update`, which needs the internet
# the tunnel was supposed to be providing. Indented JSON: a few hundred KiB even
# for a several-hundred-node subscription.
/etc/shater/subs/
# THE BOOT-ARMOR ARM TOKEN. Its PRESENCE is what lets /etc/init.d/shater-armor
# (START=21) load the fail-closed plane before fw4's `lan -> wan ACCEPT` is the
# only rule on the box. Without it the first boot after a restore forwards LAN to
# WAN in the clear until the daemon has built an engine. One small nft script.
/etc/shater/boot.nft
# COMPILED LIST ARTIFACTS (.srs). Losing these fails SILENTLY in the worst
# direction: a missing LOCAL rule-set is left out of the generated config and the
# engine starts perfectly happily with the filtering simply gone
# (shater/generate/ruleset.go, compiledListRuleSet). "It will re-download itself"
# is NOT true for them either — a compiled url list is rebuilt only by the next
# generate, and nothing schedules one (see the note in /etc/init.d/shater-cron
# about `ruleset update`). Cheap to keep: compiled .srs is 3-6% of the source
# text (~80 KiB for a 150k-domain list), under a 4 MiB soft cap.
/etc/shater/lists/
# ALERT DE-DUPLICATION STATE. A few hundred bytes mapping subscription -> when its
# expiry warning last fired. Without it every subscription already announced
# announces itself again on the restored box — the exact re-alert storm the file
# was created to prevent (shater/alert/expiry.go).
/etc/shater/alert-state.json
# DELIBERATELY NOT KEPT. Each of these is history or cache, and the backup is
# built in RAM:
#
# /etc/shater/stats.db Traffic/query HISTORY, not configuration. Bounded
# only by globals.stats_disk_limit_mb, whose default is
# 64 MB and whose 0 means UNLIMITED — one file able to
# outweigh everything else here by two orders of
# magnitude, and the only entry whose loss costs the
# operator nothing but a chart.
# /etc/shater/cache.db sing-box's own cache (8 MiB cap, deleted above it).
# Rebuilt on demand by design, and a stale rule-set
# cache carried onto a different box is worse than no
# cache at all.
# /etc/shater/shaterd.log A log (capped by globals.log_max_kb). A restored box
# wants its own log, and this one carries the DNS query
# history of the box it came from — which is not
# something to move into an archive a person then puts
# somewhere else.
#
# ON SECRECY, since this archive routinely ends up in cloud storage: subs/*.json
# carries every node's credentials (UUID/password/keys). That is not a NEW
# exposure — /etc/config/shater already carries the subscription URLs and every
# manual node's credentials, and it is already in the backup as a conffile — but a
# shater backup is a secret-bearing file and should be treated as one.
+66
View File
@@ -262,6 +262,65 @@
}
}
/* ---- fixture band (dev builds only; see App.tsx MockBanner) ----
Deliberately outside the crit/amber vocabulary: nothing is wrong with the
router, there is no router. The hazard hatch is the service-sticker language a
piece of network hardware already uses for "this unit is not in service". */
.mock-band {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
margin-top: calc(var(--u, 8px) * 2);
padding: 10px 14px;
border: 1px dashed var(--faint);
border-radius: 9px;
background: repeating-linear-gradient(
-45deg,
var(--sink),
var(--sink) 9px,
var(--panel) 9px,
var(--panel) 18px
);
}
.mock-band-tag {
flex-shrink: 0;
align-self: flex-start;
padding: 3px 7px;
border: 1px solid var(--faint);
border-radius: 4px;
background: var(--raised);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 700;
letter-spacing: 0.14em;
color: var(--dim);
}
.mock-band-copy {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 2px;
}
.mock-band-headline {
font-family: var(--font-mono);
font-size: 12.5px;
font-weight: 700;
letter-spacing: 0.02em;
color: var(--ink);
}
.mock-band-detail {
font-size: 12.5px;
line-height: 1.5;
color: var(--dim);
max-width: 76ch;
}
.mock-band-detail code {
font-family: var(--font-mono);
font-size: 11.5px;
color: var(--ink);
}
/* ---- commit-confirm band (every page except Apply, which has the full panel) ----
Same plate as the protection banner so the two read as one family; the seconds
are the loud element because they are the only thing that is running out. */
@@ -397,6 +456,13 @@
.finding--warning {
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
}
/* The daemon's "the list is capped" disclosure. Dashed, because the row is about
what ISN'T here — it must not read as one more finding to work through. */
.finding--truncated {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
background: var(--panel);
}
.finding-copy {
flex: 1;
min-width: 0;
+39 -2
View File
@@ -8,6 +8,7 @@ import { usePendingConfirm } from './pendingConfirm'
import { bootstrapSession } from './session'
import { ROUTES, navigate, useRoute } from './router'
import { engineState, protectionState } from './planeState'
import { truncationNote } from './findings'
import type { Route } from './router'
import { Overview, Placeholder, Nodes, Routing, Apply, DNS, Devices, Targets, Settings, Profiles, Insights, Networks } from './pages'
@@ -100,6 +101,7 @@ export function App() {
footer={<StatusBar status={status} />}
>
<Nav route={route} />
<MockBanner />
<PlaneBanner status={status} route={route} />
<ConfirmBand route={route} onChanged={() => void refreshStatus()} />
<Page route={route} status={status} onStatusChange={() => void refreshStatus()} />
@@ -170,6 +172,31 @@ function ConfirmBand({ route, onChanged }: { route: Route; onChanged: () => void
)
}
/**
* Says, on every page, that nothing on screen came from a router.
*
* Only a DEV build can ever render this — the fixtures are not in a production
* bundle (api.ts initMockBackend), so an operator cannot reach this state at all.
* It is here for the person who CAN: a footer line reading "DEMO DATA" is easy to
* work past for an afternoon and then screenshot into a bug report, and every
* number above it is invented.
*/
function MockBanner() {
if (!MOCK) return null
return (
<div className="mock-band" role="status">
<span className="mock-band-tag">FIXTURES</span>
<div className="mock-band-copy">
<span className="mock-band-headline">No router is being read</span>
<span className="mock-band-detail">
Every reading on this page is invented by <code>src/mock.ts</code> for offline
development. Drop <code>?mock</code> from the address to talk to a daemon.
</span>
</div>
</div>
)
}
/**
* The protection state, pinned under the nav on every page EXCEPT Overview
* (which shows the same state as its own headline readout — see planeState.ts).
@@ -195,12 +222,15 @@ function PlaneBanner({ status, route }: { status: Status | null; route: Route })
const criticals = (status.warnings ?? []).filter((w) => w.severity === 'critical').length
const state = protectionState(status)
// The published list is capped at 50, so with a note attached the count is a
// floor. Say "at least" rather than quoting a total the daemon didn't send.
const atLeast = truncationNote(status.warnings) ? 'At least ' : ''
// Wording comes from the shared source of truth so the banner and Overview can
// never describe the same router differently.
const headline = state.alarm
? state.headline
: `${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
: `${atLeast}${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
const detail = state.alarm
? state.detail
: 'Something you configured isn’t in effect. Review the findings before relying on it.'
@@ -321,8 +351,15 @@ function UnauthPlate() {
<Faceplate ariaLabel="shater — not authenticated" header={<FaceplateHeader wordmark="SHATER" subline="v0.2 · openwrt appliance" />}>
<div className="plate-msg">
<Module name="Session" value="LOCKED" led={{ variant: 'amber' }}>
{/* THE ONLY RECOVERY INSTRUCTION THE PRODUCT GIVES, so it has to point at
the real menu entry. It said "System → shater"; the page is registered
at `admin/services/shater` (luci-app-shater/root/usr/share/luci/menu.d/
luci-app-shater.json, title "Shater"), which LuCI renders under
SERVICES. Anyone reading this line has just lost access to the panel
and is looking for the one door back — sending them to the wrong menu
costs far more than its size. */}
<p className="placeholder-note">
No active session. Open the panel from the LuCI menu (System → shater →{' '}
No active session. Open the panel from the LuCI menu (Services → Shater →{' '}
<strong>Open panel</strong>) to hand off a fresh access token.
</p>
</Module>
+228 -42
View File
@@ -13,8 +13,12 @@
// serves in-memory fixtures instead of hitting the network, so `npm run dev`
// and screenshot runs render without a live backend. A real backend in dev is
// reachable instead via the Vite proxy in vite.config.ts (no flag ⇒ real fetch).
//
// THE FIXTURES ARE A DEV-BUILD-ONLY ARTEFACT — see initMockBackend below. They
// used to be a plain static import, decided at RUNTIME off `location.search`, so
// the invented router shipped inside the binary that goes on real hardware and a
// link ending in `?dev` painted a healthy appliance without making one request.
import * as mock from './mock'
import { armPendingConfirm, clearPendingConfirm, noteConfirmTimeout } from './pendingConfirm'
// --- error type -------------------------------------------------------------
@@ -108,13 +112,55 @@ export interface StatusWarning {
/** GET /api/status — live daemon + data-plane state. */
export interface Status {
running: boolean
/**
* `globals.enabled` in UCI — MEANINGLESS unless `config_readable` is true. Test
* that first; see it for why "off" and "cannot tell" must never share a branch.
*/
enabled: boolean
active: boolean
table: boolean
hash: string
version: string
// The RAW `option kill_switch` string, echoed straight off model.Globals
// (apply.go: `s.KillSwitch = m.Globals.KillSwitch`) — NOT normalised. So it can
// be "Closed", " closed ", or "" as well as the two documented spellings, and
// the daemon reads it as closed unless it case-insensitively equals "open"
// (apply.killSwitchClosed). Never compare it with `===`; use
// planeState.killSwitchClosed, which is that same rule.
kill_switch?: string // "closed" (fail-closed) | "open"
panel_port?: number // configured admin-panel port (default 8088)
// The CONFIGURED admin-panel port (default 8088) — NOT a port anything has
// confirmed is being listened on. The daemon echoes the config value, and a
// failed listen is only `logger.Warn("panel server unavailable (daemon
// continues)")`, so this field reads exactly the same whether the panel is up or
// was never bound. Do not render it as "the panel is at :N": say configured.
panel_port?: number
/**
* COULD THE CONFIGURATION BE READ AT ALL when this status was taken?
*
* It QUALIFIES the only three fields sourced from the config — `enabled`,
* `kill_switch`, `panel_port`. When it is false those three are zero values and
* mean NOTHING: not "switched off", not "kill switch unset", not "port 0". Test
* it before reading any of them.
*
* The phrasing is positive on purpose, and the panel must keep it that way: a
* client that predates the field sees it missing, reads `false`, and lands on the
* ALARMING side. Reading it as `!== false` would invert that and hand the
* reassuring branch to every daemon too old to answer.
*
* Why it matters more than it looks: the failure is a full /overlay, or a
* `uci commit` caught half-written — precisely when the fail-closed plane has the
* whole LAN cut off on purpose. The daemon then publishes `plane:"hold"` WITH
* `enabled:false`, and a panel that checks `!enabled` first renders "Turned off",
* amber, no alarm, and points at a Settings page backed by the same unreadable
* file. It tells the owner they did this to themselves while the house has no
* internet. See planeState.protectionState, where the order of those two checks
* is the whole fix.
*/
config_readable?: boolean
/** Why the config read failed, verbatim, or absent/"" when it did not. A
* diagnostic, not the operator-facing sentence — that one is published as a
* critical warning in section "config", name "unreadable". */
config_error?: string
can_rollback?: boolean // a rollback would revert something (armed snapshot or engine last-good)
// Is the sing-box engine process actually up? Absent on older daemons.
engine_running?: boolean
@@ -579,12 +625,46 @@ export interface Stats {
// --- Model shapes (PascalCase keys; slices may be null) ---------------------
/**
* `T` with every key REQUIRED to be written down — `undefined` still allowed as a
* VALUE, so nothing changes on the wire (`JSON.stringify` omits undefined, and the
* result stays assignable to `T`).
*
* WHAT IT IS FOR. Several editors REBUILD a model object from their form controls
* instead of extending the one they were given, because rebuilding is what stops a
* stale field surviving a change of shape (an inbound switched from `socks` to
* `tproxy` must not keep its old `Listen`). The cost is that the rebuild is only
* correct for as long as somebody remembers to touch it: add a sixteenth field to
* `Inbound` and every edit silently drops it, with no error anywhere. That already
* happened once, to `Ruleset.Format` (see ruleset.ts), and the value could only be
* put back over SSH.
*
* Annotating the rebuilt literal `Complete<T>` turns the next occurrence into a
* BUILD failure, in the function that has to decide, naming the field it forgot.
* Writing `undefined` for a field this shape has no use for is then a statement
* rather than an omission.
*
* Only the OPTIONAL keys get `| undefined`. A bare `{[K in keyof T]-?: T[K] |
* undefined}` would widen the required ones too — `Name: string | undefined` —
* which both weakens them and stops the result being assignable back to `T`; an
* intersection with `T` does not fix it either, since intersecting an optional
* `string` with `string | undefined` collapses back to `string`. So the two halves
* are split explicitly.
*/
type OptionalKeys<T> = { [K in keyof T]-?: object extends Pick<T, K> ? K : never }[keyof T]
export type Complete<T> = Pick<T, Exclude<keyof T, OptionalKeys<T>>> & {
[K in OptionalKeys<T>]-?: T[K] | undefined
}
export interface Globals {
Enabled: boolean
// debug|info|warning|error|none. `none` really is silent — it is emitted as the
// engine's own log-disable switch, not as a quieter level. Anything unrecognised
// falls back to warning (an unknown level fails engine start outright).
LogLevel: string
// The saved policy, verbatim. Same caveat as Status.kill_switch: the daemon
// normalises with EqualFold+TrimSpace and defaults to CLOSED, so read it through
// planeState.killSwitchClosed rather than comparing the string.
KillSwitch: string // "closed" | "open"
// There is deliberately no DNSMode. It was removed from the Go model (see
// model.go's package comment): "nftset" named the v0.1 dnsmasq architecture that
@@ -617,19 +697,60 @@ export interface Globals {
* no representation in them. So this traffic can only be dropped or let out
* directly; there is no third physical option, and the setting picks WHICH.
*
* block — (default) drop it all. No ping/traceroute out, no multicast IPTV,
* no client IPsec/PPTP passthrough. Nothing leaks.
* block — (default) drop it all. No ping/traceroute out, no client
* IPsec/PPTP passthrough. Nothing leaks.
* icmp — let ICMP/ICMPv6 echo out directly. Ping and traceroute work; the
* host being pinged sees the real WAN IP. IPTV/VPN passthrough stay
* blocked.
* direct — let all of it out directly. Ping, IPTV and IPsec/PPTP passthrough
* work, and all of it bypasses the tunnel with the real IP.
* host being pinged sees the real WAN IP. The rest gets out only
* toward addresses the routing rules already send direct.
* direct — let all of it out directly. Ping and IPsec/PPTP passthrough work,
* and all of it bypasses the tunnel with the real IP.
*
* MULTICAST IPTV IS NOT ONE OF THE THINGS THIS DECIDES, on any of the three.
* The stream is UDP, every line this policy emits carries `meta l4proto !=
* { tcp, udp }` (netplane/untunnelable.go), and the fail-closed forward chain's
* surviving accepts cover the RFC1918/link-local daddr sets only — 224.0.0.0/4
* is not among them (netplane/nft.go). The daemon says so itself in the notes it
* publishes for this section. Listing IPTV as something `direct` restores is the
* one lie this comment previously told.
*
* Absent, empty, or unrecognised ⇒ `block` (the daemon normalises to the safe
* side). Note the policy is inert while KillSwitch is "open", because then
* nothing is being blocked in the first place.
* side). Three other settings override it, and the panel must read them before
* describing it: an OPEN KillSwitch (the forward chain then has no drops at all,
* so nothing is blocked whatever this says), L3Tunnel (takes ICMP echo into the
* tunnel in prerouting, before the forward chain is consulted), and
* UntunnelableEgress (routes ESP/AH/GRE/IGMP/SCTP — and ICMP too, when L3Tunnel
* is off — out a named interface). See netplane/nft.go's prerouting chain.
*/
Untunnelable?: string // block|icmp|direct
/**
* The L3 ingress (model.Globals.L3Tunnel, UCI `l3_tunnel`). The engine opens a
* TUN device and prerouting policy-routes ICMP ECHO AND NOTHING ELSE into it,
* where the engine's own route rules pick the outbound — so ping and Windows
* tracert travel THROUGH the tunnel toward every address the rules send to an
* outbound that can carry plain IP (WireGuard/AmneziaWG), and are not answered
* at all for addresses routed to a stream-only outbound. Raw ESP/AH/GRE cannot
* enter it: sing-tun's dispatcher NATs through a port-shaped selector they do
* not have.
*
* Load-bearing for the panel because it happens BEFORE the forward chain, so it
* silently rewrites the ICMP half of every Untunnelable promise. Absent ⇒ false
* (opt-in), so read it as `=== true`.
*/
L3Tunnel?: boolean
/**
* The name of an interface/tunnel egress that carries what the engine will not
* dispatch — ESP, AH, GRE, IGMP, SCTP, plus ICMP when L3Tunnel is off
* (model.Globals.UntunnelableEgress, UCI `untunnelable_egress`). The kernel
* routes those packets out that device with its own NAT; the forward chain,
* where Untunnelable's verdicts live, never decides them.
*
* It is NOT necessarily a tunnel — the option takes any interface/tunnel egress,
* and a second WAN is just another uplink whose real address the far end sees.
* The daemon's own note for this section says which, from whether the device is
* point-to-point, so the panel does not guess. "" (the default) ⇒ Untunnelable
* is in sole charge.
*/
UntunnelableEgress?: string
DNSFilter?: boolean // master enable for the in-engine blocklist filter
DNSIntercept?: boolean // force ALL LAN plaintext DNS (:53) through the engine, incl. router-addressed queries
BlockDoH?: boolean // block known public DoH resolvers (by host + IP:443 + Firefox canary) so clients fall back to plaintext :53
@@ -659,13 +780,20 @@ export interface Globals {
StatsTimelineMinutes?: number // trailing per-minute sparkline buckets; 0 = unlimited
StatsMaxDomains?: number // network-wide domain-map size before prune; 0 = unlimited
StatsRetentionDisabled?: boolean // master switch: disable trimming for all aggregates
// SQLite-only: hard cap on the on-disk stats.db size, in MB. 0 = unlimited (bounded
// only by the device); a positive N caps the DB at N MB (oldest rows pruned + VACUUM).
// Applies only when StatsBackend === "sqlite"; ignored for off/memory.
// Disk-backend only: hard cap on the on-disk stats.db size, in MB. 0 = unlimited
// (bounded only by the device); a positive N aims the DB at N MB — oldest rows are
// deleted, then the file is REBUILT to give the pages back (stats/boltring.go:
// bbolt.Compact into a temp file + atomic swap). Not sqlite and not VACUUM: the
// store is bbolt, and the rebuild is SKIPPED when the filesystem cannot fit the
// transient second copy, which leaves the DB over its cap until space frees up.
// Applies only when StatsBackend === "sqlite" (a historical value name — see
// StatsBackend below); ignored for off/memory.
StatsDiskLimitMB?: number // stats.db disk cap in MB; 0 = unlimited
// Logging/stats backend selector. "off" collects nothing (all stats lists empty);
// "memory" keeps aggregates in RAM (lost on restart); "sqlite" persists the query
// and connection logs to an on-disk, disk-bounded store that survives a restart
// "memory" keeps aggregates in RAM (lost on restart); "sqlite" — a HISTORICAL value
// name, kept because it is on disk in every shipped config; the store behind it is
// bbolt (stats/boltring.go), pure Go and already linked into the binary — persists
// the query and connection logs to an on-disk, disk-bounded store that survives a restart
// (the DNS/nft aggregates stay in RAM; if the DB can't be opened it falls back to
// memory and the snapshot honestly reports "memory"). Default "memory". The
// EFFECTIVE running backend is echoed on the /api/stats snapshot.
@@ -955,11 +1083,22 @@ export interface Inbound {
* address. Route to a group/node/chain for the former, and to the `block` TARGET
* for the latter.
*/
/*
* THERE IS DELIBERATELY NO `Target`. It was declared here as "legacy field of the
* removed `proxy` type; read by nothing", and it is not in the Go model at all —
* so it could only ever be `undefined`, which made every panel branch asking
* "which egress points at this node/group?" (Nodes.tsx, Targets.tsx) permanently
* unreachable: dead code that read as coverage. The other end was worse than
* inert. PUT /api/config decodes with DisallowUnknownFields, so the day anything
* had put a string on it — a rename pass, a migration, a hand-written fixture —
* `JSON.stringify` would have started emitting it and the daemon would have
* rejected the WHOLE write with `json: unknown field "Target"` (verified against
* the running daemon), losing an unrelated edit somewhere else on the page.
*/
export interface Egress {
Name: string
Type: string // interface|direct|byedpi
Interface?: string // type=interface: the UCI interface name
Target?: string // legacy field of the removed `proxy` type; read by nothing
Port?: number // type=byedpi ONLY: the local ciadpi listen port (default 1080)
DPI?: string // type=interface|direct: off|fragment|record|spoof (byedpi desyncs itself)
}
@@ -1120,12 +1259,59 @@ export interface Model {
// --- transport --------------------------------------------------------------
/** True when the URL asks for the offline fixture backend (?mock or ?dev). */
export const MOCK: boolean = (() => {
// --- the offline fixture backend (dev builds only) ---------------------------
/**
* True when the in-memory fixtures are serving this session instead of the
* daemon. ALWAYS false in a production build — see {@link initMockBackend}.
*
* A live binding, not a constant: it is decided once during boot, before the
* first render, and every importer sees the same value for the whole session.
*/
export let MOCK = false
/** The loaded fixture module. `null` unless a dev build was asked for `?mock`. */
let fixtures: typeof import('./mock') | null = null
/**
* Load the fixture backend, if this build has one and the URL asks for it.
* Call ONCE from the entry point and await it before the first render — the
* pages read {@link MOCK} while they render, so flipping it afterwards would
* leave a half-mocked screen.
*
* Two gates, and the order matters. `import.meta.env.DEV` is folded to a literal
* `false` by Vite at build time, so in a production build the whole body is
* unreachable, `import('./mock')` is tree-shaken out of the module graph, and the
* fixtures are not in the emitted bundle AT ALL — not lazily, not behind a flag.
* `vite.config.ts` fails the build if that ever stops being true.
*
* This is deliberately stronger than "hide the mock behind a query flag". The
* flag was the bug: `?dev` on a production URL rendered an invented healthy
* router — 119 of 122 nodes alive, "Protected" — with no request made and one
* line of small print in the footer to say so. A person cannot audit a bundle;
* the only honest guarantee is that the invented data is not in it.
*/
export async function initMockBackend(): Promise<boolean> {
if (import.meta.env.DEV && mockRequested()) {
fixtures = await import('./mock')
MOCK = true
}
return MOCK
}
/** Does the URL ask for the offline fixture backend (`?mock` or `?dev`)? */
function mockRequested(): boolean {
if (typeof location === 'undefined') return false
const q = new URLSearchParams(location.search)
return q.has('mock') || q.has('dev')
})()
}
/** The fixture backend, for the `MOCK ? … : …` branches below. Throws rather
* than inventing data if it is ever reached without having been loaded. */
function mock(): NonNullable<typeof fixtures> {
if (!fixtures) throw new Error('mock backend not loaded — call initMockBackend() first')
return fixtures
}
/** A decoded response plus the raw Headers, for endpoints whose contract puts
* pagination metadata outside the JSON body (see the stats log endpoints). */
@@ -1174,11 +1360,11 @@ async function req<T>(path: string, init?: RequestInit): Promise<T> {
// --- endpoints --------------------------------------------------------------
export function getStatus(): Promise<Status> {
return MOCK ? mock.getStatus() : req<Status>('api/status')
return MOCK ? mock().getStatus() : req<Status>('api/status')
}
export async function getConfig(): Promise<Model> {
const m = await (MOCK ? mock.getConfig() : req<Model>('api/config'))
const m = await (MOCK ? mock().getConfig() : req<Model>('api/config'))
// Every page reads the config, and the commit-confirm window's length is the
// only thing needed to arm a countdown — so it is captured here once instead of
// being threaded through eight pages. See pendingConfirm.ts.
@@ -1188,7 +1374,7 @@ export async function getConfig(): Promise<Model> {
export function putConfig(m: Model): Promise<{ ok: boolean; applied: boolean }> {
return MOCK
? mock.putConfig(m)
? mock().putConfig(m)
: req('api/config', { method: 'PUT', body: JSON.stringify(m) })
}
@@ -1203,25 +1389,25 @@ export function putConfig(m: Model): Promise<{ ok: boolean; applied: boolean }>
* state. See pendingConfirm.ts.
*/
export async function apply(): Promise<ApplyResult> {
const r = await (MOCK ? mock.apply() : req<ApplyResult>('api/apply', { method: 'POST' }))
const r = await (MOCK ? mock().apply() : req<ApplyResult>('api/apply', { method: 'POST' }))
if (!r.error && r.changed) armPendingConfirm()
return r
}
export async function confirm(): Promise<ApplyResult> {
const r = await (MOCK ? mock.confirm() : req<ApplyResult>('api/confirm', { method: 'POST' }))
const r = await (MOCK ? mock().confirm() : req<ApplyResult>('api/confirm', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export async function rollback(): Promise<ApplyResult> {
const r = await (MOCK ? mock.rollback() : req<ApplyResult>('api/rollback', { method: 'POST' }))
const r = await (MOCK ? mock().rollback() : req<ApplyResult>('api/rollback', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export function getStats(): Promise<Stats> {
return MOCK ? mock.getStats() : req<Stats>('api/stats')
return MOCK ? mock().getStats() : req<Stats>('api/stats')
}
// --- daemon log download ------------------------------------------------------
@@ -1273,7 +1459,7 @@ function saveBlob(blob: Blob, filename: string): void {
*/
export async function downloadLog(range: LogRange): Promise<void> {
if (MOCK) {
saveBlob(new Blob([mock.getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
saveBlob(new Blob([mock().getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
return
}
let res: Response
@@ -1352,14 +1538,14 @@ function logPage<T>(env: { body: T[] | null; headers: Headers }): StatsLogPage<T
/** GET /api/stats/log — one page of the DNS query log with its cursor metadata. */
export function getStatsLogPage(q: StatsLogQuery = {}): Promise<StatsLogPage<QueryLogEntry>> {
return MOCK
? mock.getStatsLogPage(q)
? mock().getStatsLogPage(q)
: reqFull<QueryLogEntry[] | null>(`api/stats/log${statsLogQS(q)}`).then(logPage)
}
/** GET /api/stats/conns — one page of the connection log with its cursor metadata. */
export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<ConnLogEntry>> {
return MOCK
? mock.getStatsConnsPage(q)
? mock().getStatsConnsPage(q)
: reqFull<ConnLogEntry[] | null>(`api/stats/conns${statsLogQS(q)}`).then(logPage)
}
@@ -1367,13 +1553,13 @@ export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<C
* Rows only; callers that tail the stream want {@link getStatsLogPage} instead. */
export function getStatsLog(q: number | StatsLogQuery = {}): Promise<QueryLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
return MOCK ? mock().getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
}
/** GET /api/stats/conns — the live connection-event log (device→dest), newest first. */
export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
return MOCK ? mock().getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
@@ -1428,12 +1614,12 @@ export interface RulesReachability {
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
return MOCK ? mock().getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
return MOCK ? mock().getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
}
/**
@@ -1444,7 +1630,7 @@ export function getRulesetStatus(): Promise<RulesetStatus[]> {
*/
export function updateRuleset(tag: string): Promise<RulesetStatus | { ok: boolean }> {
return MOCK
? mock.updateRuleset(tag)
? mock().updateRuleset(tag)
: req('api/ruleset/update', { method: 'POST', body: JSON.stringify({ tag }) })
}
@@ -1472,7 +1658,7 @@ export interface RulesetCheck {
*/
export function checkRulesetCategory(source: string, category: string): Promise<RulesetCheck> {
return MOCK
? mock.checkRulesetCategory(source, category)
? mock().checkRulesetCategory(source, category)
: req<RulesetCheck>('api/ruleset/check', {
method: 'POST',
body: JSON.stringify({ source, category }),
@@ -1501,18 +1687,18 @@ export interface RulesetCategories {
*/
export function getRulesetCategories(source: string): Promise<RulesetCategories> {
return MOCK
? mock.getRulesetCategories(source)
? mock().getRulesetCategories(source)
: req<RulesetCategories>(`api/ruleset/categories?source=${encodeURIComponent(source)}`)
}
/** GET /api/devices — discovered LAN clients merged with per-device config. */
export function getDevices(): Promise<DiscoveredDevice[]> {
return MOCK ? mock.getDevices() : req<DiscoveredDevice[]>('api/devices')
return MOCK ? mock().getDevices() : req<DiscoveredDevice[]>('api/devices')
}
/** GET /api/interfaces — the router's UCI network interfaces for the egress picker. */
export function getInterfaces(): Promise<Interface[]> {
return MOCK ? mock.getInterfaces() : req<Interface[]>('api/interfaces')
return MOCK ? mock().getInterfaces() : req<Interface[]>('api/interfaces')
}
/** POST /api/session — exchange a single-use handoff token for a session cookie. */
@@ -1548,7 +1734,7 @@ export function importWg(conf: string): Promise<{ uri: string; name: string }> {
*/
export function updateSubscription(name: string): Promise<{ added: number }> {
return MOCK
? mock.updateSubscription(name)
? mock().updateSubscription(name)
: req('api/subscription/update', { method: 'POST', body: JSON.stringify({ name }) })
}
@@ -1568,7 +1754,7 @@ export function updateSubscription(name: string): Promise<{ added: number }> {
export function getGroupsHealth(
opts: { group?: string; members?: boolean } = {},
): Promise<GroupsHealth> {
if (MOCK) return mock.getGroupsHealth(opts)
if (MOCK) return mock().getGroupsHealth(opts)
const p = new URLSearchParams()
if (opts.group) p.set('group', opts.group)
if (opts.members) p.set('members', '1')
@@ -1683,11 +1869,11 @@ export interface GroupTestStart {
*/
export function postGroupsTest(name = ''): Promise<GroupTestStart> {
return MOCK
? mock.postGroupsTest(name)
? mock().postGroupsTest(name)
: req<GroupTestStart>('api/groups/test', { method: 'POST', body: JSON.stringify({ name }) })
}
/** GET /api/groups/test — progress + results of the current/last group test. */
export function getGroupsTest(): Promise<GroupTestStatus> {
return MOCK ? mock.getGroupsTest() : req<GroupTestStatus>('api/groups/test')
return MOCK ? mock().getGroupsTest() : req<GroupTestStatus>('api/groups/test')
}
+116
View File
@@ -0,0 +1,116 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import type { Egress } from './api.ts'
import {
DPI_TYPES,
EGRESS_TYPES,
UNKNOWN_EGRESS_TYPE_HINT,
isKnownEgressType,
nextEgress,
} from './egressEdit.ts'
// The egress editor's save merge. The defect these pin: the submit handler
// cleared Interface, Port and DPI for every type it did not have a branch for —
// including types it renders no field for at all — so opening an egress the
// panel calls "(unknown)" and changing only its NAME deleted its interface. On a
// `tunnel` egress that was live damage: the data plane routed it for real, and
// `untunnelable_egress` resolves through exactly that field, so the ESP/AH/GRE/
// IGMP/SCTP carrier silently stopped existing and those protocols fell back to
// the untunnelable policy.
test('an unknown type keeps the fields the editor never showed', () => {
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0', DPI: 'fragment' }
// The form as the editor would hold it for an unknown type: no Interface
// input is rendered, no port input, no DPI select. Only the name was touched.
const out = nextEgress(initial, {
name: 'vpn-renamed',
type: 'wireguard',
iface: 'wg0',
port: '',
dpi: 'off',
})
assert.equal(out.Name, 'vpn-renamed')
assert.equal(out.Type, 'wireguard')
assert.equal(
out.Interface,
'wg0',
'renaming an egress whose type this panel does not know must not delete its interface',
)
assert.equal(out.DPI, 'fragment', 'nor any other field the form declined to display')
})
test('an unknown type with a blank form state still keeps what was stored', () => {
// The stricter version: the editor's `iface` state is seeded from `initial`,
// so a test that passes the same value back could pass on a broken merge too.
// Blank the form and the stored value must still survive.
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0', Port: 9050 }
const out = nextEgress(initial, { name: 'vpn', type: 'wireguard', iface: '', port: '', dpi: '' })
assert.equal(out.Interface, 'wg0')
assert.equal(out.Port, 9050)
})
test('the inputs are not mutated — the caller keeps a usable `initial`', () => {
const initial: Egress = { Name: 'vpn', Type: 'wireguard', Interface: 'wg0' }
nextEgress(initial, { name: 'other', type: 'interface', iface: 'wan2', port: '', dpi: 'off' })
assert.deepEqual(initial, { Name: 'vpn', Type: 'wireguard', Interface: 'wg0' })
})
test('a known type still clears the fields it does not use', () => {
// The other half of the contract: for a type the editor DOES render, stale
// settings from the previous type must go, or the config keeps a value the new
// type ignores and the panel shows a setting that does nothing.
const initial: Egress = { Name: 'e', Type: 'interface', Interface: 'wan2', DPI: 'fragment' }
const out = nextEgress(initial, { name: 'e', type: 'byedpi', iface: 'wan2', port: '1081', dpi: 'fragment' })
assert.equal(out.Type, 'byedpi')
assert.equal(out.Interface, undefined, 'a byedpi egress has no interface')
assert.equal(out.Port, 1081)
assert.equal(out.DPI, undefined, 'byedpi desyncs itself; the native preset is not applied')
})
test('an interface egress carries its interface and DPI, and no port', () => {
const out = nextEgress(undefined, {
name: ' wan-direct ',
type: 'interface',
iface: ' wan2 ',
port: '1080',
dpi: 'record',
})
assert.equal(out.Name, 'wan-direct', 'the name is trimmed')
assert.equal(out.Interface, 'wan2', 'the interface is trimmed')
assert.equal(out.Port, undefined, 'only a byedpi egress dials a port')
assert.equal(out.DPI, 'record')
})
test('a byedpi egress with no port falls back to the ciadpi default', () => {
const out = nextEgress(undefined, { name: 'b', type: 'byedpi', iface: '', port: ' ', dpi: 'off' })
assert.equal(out.Port, 1080)
})
test('the type list is the closed set the daemon builds outbounds for', () => {
// model.KnownEgressTypes. `tunnel` must NOT be here: the daemon folds it to
// `interface` on read (model.NormalizeEgressTypes), so the panel receives the
// canonical spelling and a second entry would put the split back into the UI.
assert.deepEqual(
EGRESS_TYPES.map((t) => t.id),
['interface', 'direct', 'byedpi'],
)
assert.equal(isKnownEgressType('tunnel'), false)
assert.equal(isKnownEgressType('interface'), true)
assert.deepEqual([...DPI_TYPES].sort(), ['direct', 'interface'])
})
test('the unknown-type hint describes what actually happens, both halves of it', () => {
// It has to name the ROUTING as well as the outbound — the old text said only
// "this engine builds no outbound", which was false for the one unknown type
// anybody had, because the data plane was building that egress a real routing
// table at the same time. And it has to promise what nextEgress now keeps.
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /routing rule or table/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /no outbound/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /blocked/)
assert.match(UNKNOWN_EGRESS_TYPE_HINT, /leaves this egress’s other settings/)
assert.doesNotMatch(
UNKNOWN_EGRESS_TYPE_HINT,
/This engine builds no outbound for that type/,
'the superseded sentence claimed the engine was the only half involved',
)
})
+142
View File
@@ -0,0 +1,142 @@
import type { Egress } from './api'
/**
* The egress editor's data half — the closed type list and the one function that
* decides which fields a save writes.
*
* It lives outside `pages/Targets.tsx` because it is the part that must be
* TESTED, and the panel's test runner is `node --test src/*.test.ts`: plain
* modules only, no JSX, no DOM. Extracting it is not tidiness — the bug below
* shipped precisely because "which fields does Save write?" was three lines
* buried in a submit handler that nothing could call.
*/
/**
* The three egress kinds that produce a real way out, in the order the editor
* offers them. This is the panel's copy of `model.KnownEgressTypes` and must
* stay equal to it: the daemon builds no outbound for anything else, and the
* router installs no mark, no `ip rule` and no routing table for it either, so
* every node, group and rule bound to such an egress is blocked.
*
* `tunnel` is deliberately NOT here. It is an accepted spelling in
* `/etc/config/shater`, but the daemon folds it to `interface` on read
* (model.NormalizeEgressTypes), so an egress written that way arrives at this
* panel already saying `interface` — with its Interface field rendered, its
* blurb correct and no "(unknown)" label. Adding a fourth entry here would put
* the second spelling back into a UI that has to agree with two backend halves.
*
* `proxy` and `block` were removed: neither ever created an outbound, so
* everything bound to them fell through to the plain WAN with the real address.
* Send traffic through a proxy by routing it at a group/node/chain, and drop it
* with the `block` target on a rule.
*/
export const EGRESS_TYPES: ReadonlyArray<{ id: string; label: string; blurb: string }> = [
{
id: 'interface',
label: 'Interface — out a specific WAN or tunnel',
blurb: 'Binds to one device (wan, wg0, …) so this traffic leaves over that uplink.',
},
{
id: 'direct',
label: 'Direct — straight out, with an optional DPI preset',
blurb: 'Uses the normal route. Its point is the DPI preset below, applied to what you route here.',
},
{
id: 'byedpi',
label: 'ByeDPI — through the local ciadpi desync proxy',
blurb: 'Hands traffic to ciadpi on 127.0.0.1, which desyncs it and goes out direct.',
},
]
/** Lookup by id, or undefined when the stored type is not one this panel knows. */
export function egressTypeInfo(type: string): (typeof EGRESS_TYPES)[number] | undefined {
return EGRESS_TYPES.find((t) => t.id === type)
}
/** Whether this panel has a definition — and therefore its own fields — for the type. */
export function isKnownEgressType(type: string): boolean {
return egressTypeInfo(type) !== undefined
}
/**
* Types whose native DPI-bypass preset applies. NOT byedpi: the desync happens
* inside the ciadpi process, and the engine's tls_* flags are never stamped on
* top of it — so the control is hidden there rather than accepted and dropped.
*/
export const DPI_TYPES: ReadonlySet<string> = new Set(['interface', 'direct'])
/**
* What the editor shows under the Type select when the stored type is not one of
* the three. It has to describe what the router actually does with such an
* egress, and what THIS FORM does to it on save — both halves were wrong before.
*
* It used to read: "This engine builds no outbound for that type, so everything
* routed here is blocked. Pick one above." Two problems. It said "this engine",
* as if only the engine were involved, at a time when the data plane happily
* built an `ip rule`, a routing table and a mark bypass for a `tunnel` egress and
* `untunnelable_egress` carried live ESP/GRE out of it — so the sentence was
* flatly false for the one unknown type anybody had. And it stayed silent about
* the thing this form was doing to the egress: saving it wiped `Interface`,
* because the field is only rendered for `type === 'interface'` and the submit
* handler cleared every field it did not render. That is fixed in nextEgress
* below, and the text now says so, because a promise about saving is only worth
* making next to the code that keeps it.
*/
export const UNKNOWN_EGRESS_TYPE_HINT =
'The router does not recognise this type: it builds no outbound for it and installs no ' +
'routing rule or table, so everything routed here is blocked — never sent out over the ' +
'plain WAN. Pick a type above to fix it; saving leaves this egress’s other settings ' +
'exactly as they are until you do.'
/** The editor's form state, as strings straight out of the inputs. */
export interface EgressForm {
name: string
type: string
iface: string
port: string
dpi: string
}
/**
* Merge the form back onto the egress being edited.
*
* # The rule, and why it is the rule
*
* A save may only CLEAR a field the editor was in a position to show. For the
* three known types the editor renders every field that type uses, so clearing
* the others is right: switching `interface` → `byedpi` must drop the stale
* interface name, or the config keeps a setting the new type ignores.
*
* For a type this panel has no definition for, the editor renders NONE of those
* fields — and used to clear all three anyway:
*
* base.Interface = type === 'interface' ? iface.trim() : undefined
*
* So opening an egress the panel calls "(unknown)", changing nothing but its
* name, and pressing Save silently deleted its `interface`. That was not
* hypothetical damage. `tunnel` was such a type, the data plane routed it for
* real, and `untunnelable_egress` pointing at it is resolved by
* netplane.UntunnelableEgressBinding through exactly that field: with the
* interface gone the binding fails, the ESP/AH/GRE/IGMP/SCTP protection quietly
* stops existing, and those protocols fall back to the untunnelable policy —
* from one rename, with no message anywhere.
*
* The daemon no longer hands this panel a `tunnel` (it is folded to `interface`
* on read), so that particular type is gone. The rule stays, because the next
* type the backend gains before the panel learns it would repeat the whole
* thing: an editor must not delete what it declines to display.
*
* `initial` is never mutated — the caller keeps a usable object if the save
* fails.
*/
export function nextEgress(initial: Egress | undefined, form: EgressForm): Egress {
const type = form.type
const base: Egress = { ...(initial ?? ({} as Egress)), Name: form.name.trim(), Type: type }
// There is no `Target` to clear — the field is not in the Go model, so GET
// never delivers one and the spread above cannot produce one.
if (!isKnownEgressType(type)) return base
base.Interface = type === 'interface' ? form.iface.trim() : undefined
base.Port = type === 'byedpi' ? Number(form.port.trim()) || 1080 : undefined
base.DPI = DPI_TYPES.has(type) ? form.dpi : undefined
return base
}
+127
View File
@@ -0,0 +1,127 @@
// findings.ts — which apply-time finding is shown where.
//
// Run with `npm test` (node's built-in test runner + native TypeScript
// stripping; no test dependency is added to the SPA, which ships inside the
// daemon binary).
//
// Two defects are pinned here.
//
// 1. THE TRUNCATION NOTE WAS UNREACHABLE. The daemon caps Status.warnings at 50
// and overwrites the last slot with an `info` note counting what it dropped.
// Overview filtered `info` away wholesale, and the settings-page route keys on
// a section (`generate`) that no page owns — so the single line telling the
// operator "you are not seeing all of it" reached no screen at all.
//
// 2. FINDINGS ABOUT AN ENTITY NEVER REACHED THAT ENTITY'S PAGE. The generator
// drops a node it cannot build and names it; the Nodes page rendered that node
// as an ordinary row with a green toggle, because it never read the findings.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
attentionFindings,
entityFindings,
findingsByName,
sectionNotes,
truncationNote,
worstSeverity,
} from './findings.ts'
import type { StatusWarning } from './api.ts'
const crit = (section: string, name: string, message = 'broken'): StatusWarning => ({
severity: 'critical',
section,
name,
message,
})
const warn = (section: string, name: string, message = 'degraded'): StatusWarning => ({
severity: 'warning',
section,
name,
message,
})
const info = (section: string, name: string, message: string): StatusWarning => ({
severity: 'info',
section,
name,
message,
})
/** Verbatim from apply/warnings.go finalizeWarnings. */
const SUPPRESSED = info(
'generate',
'',
'7 further warning(s) suppressed; run `logread -e shater` for the full list',
)
// --- the truncation note ----------------------------------------------------
test('the truncation note is found, whatever else is in the list', () => {
const note = truncationNote([crit('rule', 'a'), warn('node', 'b'), SUPPRESSED])
assert.notEqual(note, null)
assert.match(note!.message, /7 further warning/)
})
test('a whole list has no truncation note', () => {
assert.equal(truncationNote([crit('rule', 'a'), warn('node', 'b')]), null)
assert.equal(truncationNote([]), null)
assert.equal(truncationNote(undefined), null)
})
test('an ordinary info note is not mistaken for the truncation note', () => {
const notes = [info('untunnelable', 'block', 'Ping and traceroute do not work…')]
assert.equal(truncationNote(notes), null)
})
test('the truncation note is kept out of the settings-page notes it would pollute', () => {
const all = [info('generate', '', 'cache: moved to /overlay'), SUPPRESSED]
const notes = sectionNotes(all, 'generate')
assert.equal(notes.length, 1)
assert.match(notes[0].message, /cache:/)
})
test('the attention list still carries only critical and warning', () => {
const all = [crit('rule', 'a'), warn('node', 'b'), info('untunnelable', 'block', 'x'), SUPPRESSED]
const attention = attentionFindings(all)
assert.equal(attention.length, 2)
assert.ok(attention.every((w) => w.severity !== 'info'))
})
// --- per-entity findings ----------------------------------------------------
test('a page takes only the sections it owns', () => {
const all = [
crit('node', 'tokyo-01', 'parse share-link: bad scheme (skipped)'),
warn('subscription', 'qomar', 'fetch failed'),
crit('rule', 'default', 'never applies'),
info('generate', '', 'cache: x'),
]
const mine = entityFindings(all, ['node', 'subscription'])
assert.deepEqual(
mine.map((w) => w.name),
['tokyo-01', 'qomar'],
)
})
test('entity findings never include info notes', () => {
const all = [info('node', 'tokyo-01', 'just a note'), SUPPRESSED]
assert.equal(entityFindings(all, ['node', 'generate']).length, 0)
})
test('findings index by name, and global (unnamed) ones are left out', () => {
const all = [
crit('node', 'tokyo-01', 'first'),
warn('node', 'tokyo-01', 'second'),
crit('node', '', 'global to the section'),
]
const byName = findingsByName(entityFindings(all, ['node']))
assert.equal(byName.size, 1)
assert.equal(byName.get('tokyo-01')!.length, 2)
})
test('one lamp per row takes the loudest severity', () => {
assert.equal(worstSeverity([warn('node', 'a'), crit('node', 'a')]), 'critical')
assert.equal(worstSeverity([warn('node', 'a')]), 'warning')
assert.equal(worstSeverity([]), null)
})
+91 -2
View File
@@ -6,7 +6,8 @@
//
// critical / warning — something needs attention: a protection promise is
// broken, or something you configured isn't in effect. These belong on
// Overview, where the operator looks first.
// Overview, where the operator looks first — and, when they name an entity,
// ALSO on the page that owns that entity (see `entityFindings`).
//
// info — a statement ABOUT the configuration, not a problem. It never clears,
// because nothing is wrong: it is simply describing a choice that was made.
@@ -16,9 +17,49 @@
// page that never goes away and never asks for anything trains people to skim
// the list — which is exactly how a real critical finding gets missed. Anything
// standing in the findings list should be something you could act on.
//
// The one exception is carved out below: the daemon's own note that it dropped
// findings to fit the cap. It is `info` by severity and unactionable by nature,
// and it is the single most important line in the list, because it is the list
// telling you it is not the whole list.
import type { StatusWarning } from './api'
/**
* The daemon's truncation disclosure, verbatim from apply/warnings.go
* finalizeWarnings:
*
* "%d further warning(s) suppressed; run `logread -e shater` for the full list"
*
* Matched on the stable clause rather than the whole sentence so a reworded tail
* still registers. If this ever stops matching, the failure mode is a list that
* silently claims to be complete — which is why `truncationNote` is tested.
*/
const SUPPRESSED_RE = /further warning\(s\) suppressed/
/**
* The daemon's "this list is incomplete" note, or null when the list is whole.
*
* Status.warnings is capped at 50, sorted critical-first, and the last slot is
* REPLACED by an `info` note counting what was dropped. That note therefore
* arrives on the one channel the panel filtered away wholesale: `info` never
* reached Overview, and the settings-page route (`sectionNotes`) keys on
* section `generate`, which no page owns. So the single line saying "there are
* findings you are not being shown" was the only one guaranteed to be invisible.
*
* Callers must render this WITH the attention list, not instead of it.
*/
export function truncationNote(warnings: StatusWarning[] | undefined): StatusWarning | null {
return (
(warnings ?? []).find((w) => w.severity === 'info' && SUPPRESSED_RE.test(w.message)) ?? null
)
}
/** Is this the truncation disclosure rather than an ordinary note? */
function isTruncationNote(w: StatusWarning): boolean {
return w.severity === 'info' && SUPPRESSED_RE.test(w.message)
}
/** Findings that need attention — the Overview list. Info notes are excluded. */
export function attentionFindings(warnings: StatusWarning[] | undefined): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'critical' || w.severity === 'warning')
@@ -29,10 +70,58 @@ export function attentionFindings(warnings: StatusWarning[] | undefined): Status
* (e.g. `untunnelable` → the Networks page's "Other traffic" section). Only info:
* a critical/warning is an attention item and stays on Overview, so it can't be
* quietly buried on a settings page instead.
*
* The truncation note is excluded: it is about the LIST, not about any section,
* and it has its own home beside the list ({@link truncationNote}).
*/
export function sectionNotes(
warnings: StatusWarning[] | undefined,
section: string,
): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'info' && w.section === section)
return (warnings ?? []).filter(
(w) => w.severity === 'info' && w.section === section && !isTruncationNote(w),
)
}
/**
* The attention findings about entities ONE page owns — for that page to show
* beside the entities themselves.
*
* Overview is where you look when you already suspect something; a page like
* Nodes is where you look when you don't. The generator drops a node it cannot
* build — an unparseable share link, a WireGuard key materialised twice — and
* says so by name ("node \"x\": parse share-link: … (skipped)"), yet that node
* kept rendering as an ordinary row with a green toggle, because the page never
* read the findings at all. The switch says on; the engine has no such outbound.
*
* This does NOT move anything off Overview: the same finding appears in both
* places, which is correct — one list is "what is wrong with this router", the
* other is "what is wrong with this node".
*/
export function entityFindings(
warnings: StatusWarning[] | undefined,
sections: readonly string[],
): StatusWarning[] {
const want = new Set(sections)
return attentionFindings(warnings).filter((w) => want.has(w.section))
}
/** Index attention findings by entity name, for badging a row directly. Entries
* with an empty `name` are global to their section and are left out. */
export function findingsByName(findings: StatusWarning[]): Map<string, StatusWarning[]> {
const out = new Map<string, StatusWarning[]>()
for (const f of findings) {
if (!f.name) continue
const list = out.get(f.name)
if (list) list.push(f)
else out.set(f.name, [f])
}
return out
}
/** The loudest severity in a set — for a row badge that has room for one lamp. */
export function worstSeverity(findings: StatusWarning[]): 'critical' | 'warning' | null {
if (findings.some((f) => f.severity === 'critical')) return 'critical'
if (findings.length > 0) return 'warning'
return null
}
+25 -9
View File
@@ -2,17 +2,33 @@ import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './tokens.css'
import { App } from './App'
import { initMockBackend } from './api'
import { ConfirmProvider } from './components'
const rootEl = document.getElementById('root')
if (!rootEl) throw new Error('#root not found')
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
// Settle the fixture question BEFORE the first render: pages read `MOCK` while
// they render, so a backend that arrives afterwards would paint half a screen
// from the daemon and half from fixtures. In a production build this resolves
// immediately and to `false` — the fixtures are not in the bundle to load (see
// api.ts initMockBackend and the assertNoMockFixtures plugin in vite.config.ts).
function mount() {
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl!).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
}
// A fixture module that fails to load is a broken dev checkout, not a reason to
// hand the operator a blank plate — mount anyway and let the shell report that it
// cannot reach a daemon, which by then is the truth.
void initMockBackend().then(mount, (e) => {
console.error('mock backend failed to load; continuing against the real API', e)
mount()
})
+149 -8
View File
@@ -6,8 +6,15 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import { killSwitchClosed } from './planeState'
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, Traffic } from './api'
/** One URL knob, safe to read before `location` exists (SSR-less builds/tests). */
function mockParam(name: string): string | null {
if (typeof location === 'undefined') return null
return new URLSearchParams(location.search).get(name)
}
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
let hash = 'sha256:9f7c07e8d8ac4ae1'
@@ -33,7 +40,13 @@ const CONFIG: Model = {
ActiveProfile: 'mobile-uplink',
// Policy for traffic TPROXY physically can't carry (non-TCP/UDP). Override
// from the URL — ?mock&untun=icmp / &untun=direct — to see all three states.
Untunnelable: new URLSearchParams(typeof location === 'undefined' ? '' : location.search).get('untun') ?? 'block',
Untunnelable: mockParam('untun') ?? 'block',
// The two settings that decide part of the untunnelable traffic BEFORE the
// policy above is consulted, so the Networks copy has to change shape for
// them: ?mock&l3=1 (ping rides the tunnel) and ?mock&uegress=wg0 (the kernel
// routes ESP/GRE/SCTP out that interface). Both off by default, as shipped.
L3Tunnel: mockParam('l3') === '1',
UntunnelableEgress: mockParam('uegress') ?? '',
DNSIntercept: true, // force ALL LAN plaintext DNS (:53) through the engine
BlockDoH: false, // block known public DoH resolvers so clients fall back to plaintext :53
GroupHealth: true, // observatory: background probing of used groups/chains + Targets health stats (default on)
@@ -42,7 +55,7 @@ const CONFIG: Model = {
StatsMaxDomains: 5000, // fixed cap — shows the "limit" rendering (5000)
StatsRetentionDisabled: false,
StatsBackend: 'memory', // logging backend: off | memory | sqlite
StatsDiskLimitMB: 64, // SQLite-only disk cap (MB); shows once backend=sqlite (0 ⇒ Unlimited)
StatsDiskLimitMB: 64, // disk-backend-only cap (MB); shows once backend=sqlite (0 ⇒ Unlimited)
// Daemon operational log (shaterd's own log): both destinations on, file in
// tmpfs (the default), 2 MB cap — the defaults a fresh install ships with.
LogToSyslog: true,
@@ -500,14 +513,28 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
// the field case the readout used to call "Protected" (one
// rule, `default → direct`); `unknown` is a daemon too old to
// report. Default: tunnel.
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
// ?mock&plane=unreported → a daemon that sends NO `plane` field. The panel then
// knows nothing about what is installed, which is the state
// the Kill-switch module used to render as a green "ARMED"
// (`undefined !== 'none'` is true).
function mockPlane(): {
plane: 'full' | 'hold' | 'none' | undefined
engine: boolean
killSwitch: string
} {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
const killSwitch = params.get('ks') === 'open' ? 'open' : 'closed'
// Passed through VERBATIM, because that is what the daemon does: apply.go sets
// `s.KillSwitch = m.Globals.KillSwitch` with no normalisation, so `?ks=Closed`,
// `?ks=%20closed%20` and `?ks=` are all reachable readings of a router that
// BLOCKS. The mock used to fold everything that wasn't "open" to "closed",
// which made the panel's own `=== 'closed'` bug unreproducible here.
const killSwitch = params.get('ks') ?? 'closed'
const p = params.get('plane')
if (p === 'hold') return { plane: 'hold', engine: false, killSwitch: 'closed' }
if (p === 'none') return { plane: 'none', engine: false, killSwitch }
if (p === 'open') return { plane: 'none', engine: false, killSwitch: 'open' }
if (p === 'unreported') return { plane: undefined, engine: true, killSwitch }
return { plane: 'full', engine: true, killSwitch }
}
@@ -515,7 +542,7 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
// meaningful with the plane installed: with the engine down there is no running
// config to judge, and the daemon reports the unknown/zero value — so do the same
// here rather than leaving a stale "tunnel" behind a dead engine.
function mockTraffic(plane: 'full' | 'hold' | 'none'): Traffic | undefined {
function mockTraffic(plane: 'full' | 'hold' | 'none' | undefined): Traffic | undefined {
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
switch (params.get('traffic')) {
@@ -560,6 +587,29 @@ const MOCK_WARNINGS: StatusWarning[] = [
name: 'fakeip-pool',
message: 'fake-IP resolver cannot be used as a fallback; the failover chain was not built',
},
// Two findings the generator attributes to a NODE by name — the class that the
// Nodes page never showed, leaving a node the engine threw away rendered as an
// ordinary row with a green toggle. Both name real fixture nodes so the row
// badge, the collapsed-bucket "N flagged" count and the per-row strip all fire.
{
severity: 'warning',
section: 'node',
name: 'fi-trojan',
message: 'parse share-link: unsupported scheme "trojan+ws" (skipped)',
},
{
severity: 'warning',
section: 'node',
name: 'home-wg',
message:
'this WireGuard node is materialised twice in the engine config — as "home-wg" and as "group-stealth-m1-home-wg" — and traffic can reach both. A WireGuard peer keeps ONE session per public key, so two devices built from one private key evict each other continuously and NEITHER tunnel passes traffic. Only "home-wg" is kept; everything that routed through "group-stealth-m1-home-wg" is fail-closed (blocked) instead of leaving over the plain WAN',
},
{
severity: 'warning',
section: 'subscription',
name: 'backup',
message: 'fetch failed: dial tcp 203.0.113.9:443: i/o timeout — serving the nodes cached earlier',
},
{
severity: 'info',
section: 'generate',
@@ -568,6 +618,34 @@ const MOCK_WARNINGS: StatusWarning[] = [
},
]
/**
* The daemon's truncation disclosure, exactly as apply/warnings.go writes it when
* the published set overflows the 50-entry cap. Served under `?mock&trunc` so the
* "this list is incomplete" rendering is exercisable — it used to be dropped
* wholesale by the panel's `info` filter and reached no screen at all.
*/
/**
* The daemon's own critical finding when it cannot read the configuration
* (apply.go, section "config" / name "unreadable"). Copied close to verbatim: the
* sentence about NOT switching anything off is the load-bearing one — the instinct
* in front of a dead LAN is to turn things off, and that is the single action that
* makes this worse.
*/
const CONFIG_UNREADABLE_WARNING: StatusWarning = {
severity: 'critical',
section: 'config',
name: 'unreadable',
message:
"the router's configuration could NOT be read (uci show shater: exit status 1), so this status cannot say whether shater is switched on, whether the kill switch is closed, or which port this panel is served on — enabled, kill_switch and panel_port are placeholders here, not readings. If traffic is being blocked, that is the fail-closed plane doing its job and NOT the service being switched off: do not turn anything off to fix it. The usual causes are a full /overlay and a `uci commit` interrupted part-way; free space, check /etc/config/shater, then restart shaterd.",
}
const MOCK_TRUNCATION: StatusWarning = {
severity: 'info',
section: 'generate',
name: '',
message: '7 further warning(s) suppressed; run `logread -e shater` for the full list',
}
/**
* The standing `untunnelable` note the daemon reports. It is INFO, never a
* problem: it states a correct, chosen configuration. Two shapes, mirroring the
@@ -575,7 +653,36 @@ const MOCK_WARNINGS: StatusWarning[] = [
* is inert entirely while the kill-switch is open.
*/
function untunnelableNote(mode: string, killSwitch: string): StatusWarning[] {
if (killSwitch === 'open') {
const g = CONFIG.Globals as { L3Tunnel?: boolean; UntunnelableEgress?: string }
const egress = (g.UntunnelableEgress ?? '').trim()
// The daemon's own precedence: the egress carrier owns the whole story, then
// the L3 ingress, then the kill-switch, then the policy (apply/warnings.go).
if (egress) {
return [
{
severity: 'info',
section: 'untunnelable',
name: egress,
message:
(g.L3Tunnel
? 'ping and Windows tracert travel THROUGH the tunnel; everything else the tunnel cannot carry — IPsec (ESP/AH), PPTP/GRE, SCTP — now leaves'
: 'ping, Windows tracert, IPsec (ESP/AH), PPTP/GRE, SCTP and every other protocol that is neither TCP nor UDP now leave') +
` through egress "${egress}": the kernel routes them out that interface with that interface's own NAT, and none of it follows your routing rules. Multicast IPTV does not pass this router under any setting, and carrying IGMP out an egress cannot change that.`,
},
]
}
if (g.L3Tunnel) {
return [
{
severity: 'info',
section: 'untunnelable',
name: '',
message:
'ping and Windows tracert work and travel THROUGH the tunnel, toward every address your rules send to an outbound that can carry plain IP (WireGuard/AmneziaWG); addresses your rules send anywhere else cannot be pinged at all, deliberately. Raw VPN passthrough (IPsec ESP/AH, PPTP/GRE) cannot enter the tunnel and stays with the untunnelable policy. Multicast IPTV does not pass this router on any setting; the L3 ingress does not change that.',
},
]
}
if (!killSwitchClosed(killSwitch)) {
return [
{
severity: 'info',
@@ -614,9 +721,12 @@ function mockWarnings(killSwitch: string): StatusWarning[] {
const params = new URLSearchParams(q)
const mode = (CONFIG.Globals as { Untunnelable?: string }).Untunnelable ?? 'block'
const notes = untunnelableNote(mode, killSwitch)
// `?trunc` adds the daemon's "the published list is capped" disclosure, which
// it appends IN PLACE OF the last entry it had room for.
const trunc = params.has('trunc') ? [{ ...MOCK_TRUNCATION }] : []
// A degraded plane always comes with the findings that explain it.
if (params.has('warn') || params.get('plane')) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes]
if (params.has('warn') || params.get('plane') || trunc.length > 0) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes, ...trunc]
}
return notes
}
@@ -625,6 +735,34 @@ export async function getStatus(): Promise<Status> {
await wait(120)
const enabled = (CONFIG.Globals as { Enabled: boolean }).Enabled
const { plane, engine, killSwitch } = mockPlane()
// ?mock&cfg=unreadable — the daemon could not READ the configuration (full
// /overlay, or a `uci commit` caught half-written). It is not a hypothetical: it
// is the situation the fail-closed plane exists for, so it ships with the plane
// HOLDING and the whole LAN cut off deliberately — while `enabled`, `kill_switch`
// and `panel_port` are placeholders that mean nothing. Reproducing it here is how
// the "Turned off" misreading stays fixed: the panel must alarm, not reassure.
if (mockParam('cfg') === 'unreadable') {
return {
running: true,
enabled: false, // a placeholder, NOT "the owner switched it off"
active: false,
table: true,
hash,
version: '1.11.0-shater',
kill_switch: '', // placeholder likewise
panel_port: 0, // placeholder likewise
config_readable: false,
config_error: 'uci show shater: exit status 1',
can_rollback: armed || hasLastGood,
engine_running: false,
plane: 'hold',
traffic: mockTraffic('hold'),
warnings: [CONFIG_UNREADABLE_WARNING, ...mockWarnings(killSwitch)],
started_unix: MOCK_STARTED_UNIX,
uptime_seconds: Math.floor(Date.now() / 1000) - MOCK_STARTED_UNIX,
byedpi_installed: true,
}
}
return {
running: true,
enabled,
@@ -634,6 +772,9 @@ export async function getStatus(): Promise<Status> {
version: '1.11.0-shater',
kill_switch: killSwitch,
panel_port: 8088,
// The daemon read the config fine in every other mock state. Sent explicitly
// rather than left off: absent means "no reading", which is a different claim.
config_readable: true,
can_rollback: armed || hasLastGood,
engine_running: engine,
plane,
+443
View File
@@ -0,0 +1,443 @@
/* Alerts section (rendered on Settings) — inherits the Faceplate tokens and the
* shared page chrome from App.css (.toast, .mono). Every rule below is a
* one-to-one copy of the DNS.css rule the markup used before the section moved
* here, renamed `dns-*` → `alr-*` so nothing collides. Orange stays an accent. */
/* ---- section shell (matches the Settings group plates one-to-one) ---- */
.alr-section {
margin-top: calc(var(--u, 8px) * 3.5);
}
.alr-sec-hd {
display: flex;
align-items: baseline;
gap: 12px;
padding-bottom: 10px;
border-bottom: 1px solid var(--groove);
}
.alr-sec-title {
margin: 0;
font-family: var(--font-mono);
font-size: 13px;
font-weight: 700;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--dim);
}
.alr-sec-count {
font-size: 11px;
letter-spacing: 0.06em;
color: var(--faint);
}
.alr-sec-note {
margin: 10px 2px 0;
font-family: var(--font-sans);
font-size: 12.5px;
line-height: 1.55;
color: var(--dim);
max-width: 56ch;
}
/* ---- add form ---- */
.alr-add {
display: flex;
flex-direction: column;
gap: 10px;
margin-top: calc(var(--u, 8px) * 2);
}
.alr-add-top {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px;
}
.alr-input {
min-width: 0;
padding: 9px 12px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 12.5px;
letter-spacing: 0.02em;
box-shadow: 0 1px 2px var(--shadow) inset;
transition: border-color 0.15s, box-shadow 0.15s;
}
.alr-input::placeholder {
color: var(--faint);
}
.alr-input:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-input:disabled {
opacity: 0.55;
}
.alr-input--name {
flex: 0 1 14rem;
}
/* segmented type picker */
.alr-seg {
display: inline-flex;
border: 1px solid var(--groove);
border-radius: 7px;
overflow: hidden;
background: var(--sink);
}
.alr-seg-btn {
padding: 8px 14px;
border: 0;
background: transparent;
color: var(--dim);
font-family: var(--font-mono);
font-size: 11px;
letter-spacing: 0.08em;
text-transform: uppercase;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.alr-seg-btn + .alr-seg-btn {
border-left: 1px solid var(--groove);
}
.alr-seg-btn.on {
background: var(--accent);
color: #fff;
}
.alr-seg-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
.alr-resp {
display: inline-flex;
align-items: center;
gap: 8px;
}
.alr-resp-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-select {
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 11.5px;
letter-spacing: 0.04em;
cursor: pointer;
}
.alr-select:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-add-actions {
display: flex;
align-items: center;
justify-content: flex-end;
gap: 14px;
flex-wrap: wrap;
}
.alr-field-err {
flex: 1;
min-width: 0;
margin: 0;
font-family: var(--font-mono);
font-size: 11.5px;
line-height: 1.5;
color: var(--crit);
}
/* ---- rows ---- */
.alr-rows {
list-style: none;
margin: calc(var(--u, 8px) * 2) 0 0;
padding: 0;
display: flex;
flex-direction: column;
gap: 8px;
}
.alr-row {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
padding: 12px 14px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(
180deg,
var(--raised),
color-mix(in srgb, var(--raised) 82%, var(--panel))
);
box-shadow: 0 1px 0 var(--edge) inset;
}
.alr-row-main {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 4px;
}
.alr-row-l1 {
display: flex;
align-items: center;
gap: 8px;
flex-wrap: wrap;
}
.alr-row-name {
font-family: var(--font-mono);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
color: var(--ink);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 24ch;
}
.alr-row-l2 {
display: flex;
align-items: center;
gap: 10px;
flex-wrap: wrap;
font-size: 11.5px;
letter-spacing: 0.02em;
}
.alr-row-detail {
color: var(--dim);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 40ch;
}
/* badge — groove-bordered, not orange (accent stays reserved) */
.alr-badge {
display: inline-block;
padding: 2px 7px;
border: 1px solid var(--groove);
border-radius: 5px;
background: color-mix(in srgb, var(--sink) 60%, transparent);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 600;
letter-spacing: 0.1em;
text-transform: uppercase;
color: var(--dim);
white-space: nowrap;
}
.alr-badge--accent {
border-color: color-mix(in srgb, var(--accent) 55%, var(--groove));
color: var(--accent);
}
.alr-masked {
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: 0.08em;
color: var(--faint);
text-transform: uppercase;
cursor: help;
}
.alr-del {
flex: none;
padding: 6px 12px;
font-size: 10.5px;
}
/* ---- empty plate ---- */
.alr-empty {
margin-top: calc(var(--u, 8px) * 2);
padding: calc(var(--u, 8px) * 3);
border: 1px dashed var(--groove);
border-radius: 9px;
background: color-mix(in srgb, var(--raised) 55%, transparent);
text-align: center;
}
.alr-empty-title {
display: block;
font-size: 13px;
font-weight: 700;
letter-spacing: 0.06em;
color: var(--dim);
}
.alr-empty-body {
margin: 8px auto 0;
max-width: 48ch;
font-family: var(--font-sans);
font-size: 13px;
line-height: 1.55;
color: var(--dim);
}
/* ---- loading skeleton ---- */
.alr-skel {
height: 62px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(90deg, var(--raised), var(--sink), var(--raised));
background-size: 200% 100%;
animation: alr-skel-shift 1.4s ease-in-out infinite;
}
@keyframes alr-skel-shift {
from {
background-position: 200% 0;
}
to {
background-position: -200% 0;
}
}
/* the per-row delivery picker sits inline in the row */
.alr-detour {
flex: none;
display: flex;
flex-direction: column;
gap: 5px;
min-width: 0;
}
.alr-detour-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-detour-select {
max-width: 22rem;
}
/* current delivery-path readout on the row */
.alr-path {
color: var(--faint);
white-space: nowrap;
}
.alr-path[data-active='on'] {
color: var(--dim);
}
.alr-path-name {
color: var(--led-on);
font-weight: 600;
}
.alr-path[data-missing='y'] .alr-path-name {
color: var(--amber);
}
.alr-path-flag {
color: var(--amber);
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.alr-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.alr-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.alr-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-row delivery controls sit inline in the row (like .alr-detour) */
.alr-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.alr-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.alr-note--row {
margin-top: 2px;
}
/* alert event checkboxes */
.alr-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.alr-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.alr-event input {
accent-color: var(--fp-accent, currentColor);
}
/* ---- responsive ---- */
@media (max-width: 640px) {
.alr-row {
flex-wrap: wrap;
}
.alr-row-main {
flex-basis: calc(100% - 90px);
}
.alr-del {
margin-left: auto;
}
.alr-input--name {
flex-basis: 100%;
}
.alr-detour {
flex-basis: 100%;
order: 3;
flex-wrap: wrap;
}
/* A <select> won't shrink below its widest option unless it's allowed to:
without min-width:0 the long detour labels push the page into a horizontal
scroll at 390px. Let them fill the row and clip instead. */
.alr-detour-select,
.alr-resp .alr-select {
max-width: 100%;
width: 100%;
min-width: 0;
}
.alr-resp {
display: flex;
flex-wrap: wrap;
max-width: 100%;
}
.alr-ctl {
flex-basis: 100%;
order: 3;
}
}
@media (prefers-reduced-motion: reduce) {
.alr-skel {
animation: none;
}
.alr-input,
.alr-seg-btn {
transition: none;
}
}
+690
View File
@@ -0,0 +1,690 @@
import './Alerts.css'
import { useCallback, useMemo, useState } from 'react'
import { Button, Toggle, useConfirm } from '../components'
import type { Alert, Model } from '../api'
// The Alerts section — out-of-band notifications (Telegram bot / webhook) for
// kill-switch trips, apply failures, new devices and subscription expiry. It
// lived at the bottom of the DNS page, which is the last place an operator
// looking for "tell me when the tunnel dies" would think to look; it now renders
// as a group on Settings. The component owns no I/O: every mutation goes through
// the `onSave` prop so Settings keeps a single dirty banner and a single toast.
//
// NOTE on duplication: the detour helpers below (DetourCatalog, canonDetour,
// detourValues, describeDetour, DetourSelect) plus asArray / uniqueName /
// maskUrl / EmptyPlate are deliberate copies of the ones in DNS.tsx. DNS keeps
// its own for resolvers and DNS rules; extracting a shared module would couple
// two pages that otherwise share nothing, and that refactor is out of scope
// here. If a third consumer ever appears, promote them then.
// ---- local Model extension --------------------------------------------------
/** The Model with the Alerts slice surfaced (index-signature passthrough). */
type AlertsModel = Model & { Alerts?: Alert[] | null }
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers (copies of DNS.tsx — see the header note) ----------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
const HTTP_RE = /^https?:\/\//i
/** A remote URL often carries a token in its query/path — show host only. */
function maskUrl(url: string): { host: string; masked: boolean } {
try {
const u = new URL(url)
return { host: u.host, masked: u.search !== '' || u.pathname.replace(/\/+$/, '') !== '' }
} catch {
return { host: url || '—', masked: false }
}
}
function uniqueName(base: string, taken: Set<string>): string {
const seed = base.trim() || 'alert'
if (!taken.has(seed)) return seed
let i = 2
while (taken.has(`${seed}-${i}`)) i++
return `${seed}-${i}`
}
/** The live targets an alert's delivery can be pinned to (the picker). */
interface DetourCatalog {
groups: string[]
chains: string[]
egresses: { name: string; type: string }[]
nodes: string[]
}
/**
* Normalise a stored `Via` to a picker option value. Empty/`direct` ⇒
* `direct`; already-prefixed values (`group:`/`chain:`/`egress:`/`node:`) pass
* through; a bare legacy name is resolved against the catalog so a still-valid
* setup isn't mislabelled; anything unresolved is kept verbatim (shown stale).
*/
function canonDetour(raw: string | undefined, cat: DetourCatalog): string {
const d = (raw ?? '').trim()
if (!d || d.toLowerCase() === 'direct') return 'direct'
if (/^(node|group|chain|egress):/i.test(d)) return d
if (cat.egresses.some((e) => e.name === d)) return `egress:${d}`
if (cat.groups.includes(d)) return `group:${d}`
if (cat.chains.includes(d)) return `chain:${d}`
if (cat.nodes.includes(d)) return `node:${d}`
return d
}
/** Every valid option value for a catalog, including `direct`. */
function detourValues(cat: DetourCatalog): Set<string> {
const s = new Set<string>(['direct'])
for (const g of cat.groups) s.add(`group:${g}`)
for (const c of cat.chains) s.add(`chain:${c}`)
for (const e of cat.egresses) s.add(`egress:${e.name}`)
for (const n of cat.nodes) s.add(`node:${n}`)
return s
}
/** Describe a canonical detour value for the row readout. */
function describeDetour(
canon: string,
cat: DetourCatalog,
valid: Set<string>,
): { direct: boolean; prefix: string; name: string; missing: boolean } {
if (canon === 'direct') return { direct: true, prefix: '', name: '', missing: false }
const i = canon.indexOf(':')
const kind = i === -1 ? '' : canon.slice(0, i)
const name = i === -1 ? canon : canon.slice(i + 1)
const missing = !valid.has(canon)
let prefix = 'via'
if (kind === 'group') prefix = 'via group'
else if (kind === 'chain') prefix = 'via chain'
else if (kind === 'node') prefix = 'via node'
else if (kind === 'egress') {
const eg = cat.egresses.find((e) => e.name === name)
prefix = eg?.type === 'interface' ? 'via interface' : 'via egress'
}
return { direct: false, prefix, name, missing }
}
// ---- section ----------------------------------------------------------------
export function AlertsSection({
config,
busy,
loading,
onSave,
}: {
/** Full desired-state model; null until it has loaded. */
config: Model | null
/** A save/apply is in flight — controls lock. */
busy: boolean
/** The config is still loading — show a skeleton row. */
loading: boolean
/** Persist the whole next model; resolves true on success (Settings' `save`). */
onSave: (next: Model, okMsg: string) => Promise<boolean>
}): JSX.Element {
const confirm = useConfirm()
const model = config as AlertsModel | null
const alerts = useMemo<Alert[]>(() => asArray(model?.Alerts), [model])
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Built straight from the Model with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({
groups: asArray(config?.Groups).map((g) => g.Name),
chains: [],
egresses: asArray(config?.Egresses).map((e) => ({ name: e.Name, type: e.Type })),
nodes: asArray(config?.Nodes).map((n) => n.Name),
}),
[config],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
// ---- mutations — all writes go through onSave -----------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!model) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return onSave({ ...model, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[model, alerts, onSave],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void onSave({ ...model, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[model, alerts, onSave],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!model) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void onSave({ ...model, Alerts: next }, `Deleted ${target.Name}`)
},
[model, alerts, onSave, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!model) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void onSave(
{ ...model, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[model, alerts, onSave],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void onSave(
{ ...model, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[model, alerts, onSave],
)
return (
<div className="alr-section" aria-label="Alerts">
<header className="alr-sec-hd">
<h2 className="alr-sec-title">Alerts</h2>
<span className="alr-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="alr-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="alr-rows" aria-hidden="true">
<li className="alr-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="alr-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
)
}
// ---- alert add form + row ----------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="alr-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="alr-add-top">
<input
className="alr-input alr-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="alr-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="alr-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="alr-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="alr-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="alr-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="alr-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="alr-delivery">
<label className="alr-resp">
<span className="alr-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="alr-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="alr-add-actions">
{err && (
<p className="alr-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="alr-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="alr-row-main">
<div className="alr-row-l1">
<span className="alr-row-name">{alert.Name}</span>
<span className="alr-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="alr-badge alr-badge--accent">
{e}
</span>
))}
</div>
<div className="alr-row-l2 mono">
<span className="alr-row-detail">{detail.text}</span>
{detail.masked && (
<span className="alr-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="alr-path">direct</span>
) : (
<span className="alr-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="alr-path-name">{route.name}</strong>
{route.missing && <span className="alr-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="alr-note alr-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="alr-ctl">
<label className="alr-detour">
<span className="alr-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="alr-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
/** The live delivery picker: option list built from the Model's targets. */
function DetourSelect({
value,
catalog,
valid,
busy,
disabled,
ariaLabel,
onChange,
directLabel = 'Direct (no proxy)',
}: {
value: string // canonical value
catalog: DetourCatalog
valid: Set<string>
busy: boolean
disabled: boolean
ariaLabel: string
onChange: (v: string) => void
directLabel?: string
}) {
const missing = value !== 'direct' && !valid.has(value)
return (
<select
className="alr-select alr-detour-select"
value={value}
onChange={(e) => onChange(e.target.value)}
disabled={busy || disabled}
aria-label={ariaLabel}
>
<option value="direct">{directLabel}</option>
{catalog.groups.length > 0 && (
<optgroup label="Groups">
{catalog.groups.map((g) => (
<option key={g} value={`group:${g}`}>
Group {g} (balancer)
</option>
))}
</optgroup>
)}
{catalog.chains.length > 0 && (
<optgroup label="Chains">
{catalog.chains.map((c) => (
<option key={c} value={`chain:${c}`}>
Chain {c}
</option>
))}
</optgroup>
)}
{catalog.egresses.length > 0 && (
<optgroup label="Interfaces / egresses">
{catalog.egresses.map((e) => (
<option key={e.name} value={`egress:${e.name}`}>
Interface/egress {e.name}
{e.type ? ` (${e.type})` : ''}
</option>
))}
</optgroup>
)}
{catalog.nodes.length > 0 && (
<optgroup label="Nodes">
{catalog.nodes.map((n) => (
<option key={n} value={`node:${n}`}>
Node {n}
</option>
))}
</optgroup>
)}
{missing && <option value={value}>{value} (missing)</option>}
</select>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="alr-empty">
<span className="alr-empty-title mono">{title}</span>
<p className="alr-empty-body">{body}</p>
</div>
)
}
+97 -37
View File
@@ -11,7 +11,7 @@ import {
ApiError,
} from '../api'
import type { Globals, Status } from '../api'
import { engineReadout } from '../planeState'
import { engineReadout, killSwitchReadout } from '../planeState'
import { onPendingConfirmExpire, usePendingConfirm } from '../pendingConfirm'
// Short, readable config hash — drops the "sha256:" prefix like the footer does.
@@ -103,30 +103,64 @@ export default function Apply() {
void loadConfig()
}, [loadConfig])
// The window running out is the daemon reverting on its own — observe it and
// say so. The countdown itself ticks inside usePendingConfirm; this only reacts
// to the end of it, and the store makes sure that fires exactly once even with
// the app-wide band mounted alongside.
// The window running out does NOT mean the daemon rolled back.
//
// apply.ArmRollback captures the data-plane generation when it arms, and on
// expiry it compares. If anything re-applied the plane in between — another
// panel apply, SIGHUP, a hotplug or the once-a-minute cron reconcile, the WAN
// profile auto-switch — it disarms and KEEPS the running config, logging "NOT
// rolling back" and nothing else. That is the common case on a production
// router, and this page used to print "daemon auto-rolled back to last-good
// config" for it: a confident report of an event that did not happen, with a
// hash pair underneath that quietly said "unchanged".
//
// The panel cannot see which branch ran — the daemon says so only in its log.
// So it reports the one thing it CAN observe, the live config hash, and waits
// for the revert to land before reading it (a rollback is a full re-apply and
// does not complete the instant the timer fires).
const liveHashRef = useRef('')
liveHashRef.current = status?.hash ?? ''
useEffect(
() =>
onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Auto-rolled back')
void (async () => {
const after = (await refreshStatus())?.hash ?? ''
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed — daemon auto-rolled back to last-good config.',
before,
after,
})
})()
}),
[flash, refreshStatus],
)
useEffect(() => {
let cancelled = false
const off = onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Confirm window elapsed')
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed. Reading what the daemon did…',
before,
after: before,
})
void (async () => {
let after = before
for (let i = 0; i < 4 && !cancelled; i++) {
await new Promise((r) => window.setTimeout(r, 1500))
if (cancelled) return
after = (await refreshStatus())?.hash ?? after
if (after !== before) break
}
if (cancelled) return
setResult({
kind: 'expire',
tone: 'warn',
text:
after !== before
? 'Confirm window elapsed and the live config changed — the daemon reverted to its last-good config.'
: 'Confirm window elapsed and the live config has not changed, so this config is still running. ' +
'The daemon only reverts if nothing else re-applied the data plane while the window was open; ' +
'otherwise it stands down and keeps what is live. Which one happened is in the daemon log — ' +
'download it from Settings, or run `logread -e shater`.',
before,
after,
})
})()
})
return () => {
cancelled = true
off()
}
}, [flash, refreshStatus])
const confirmWindow = globals?.ConfirmTimeout ?? 0
@@ -240,7 +274,25 @@ export default function Apply() {
// The LIVE kill-switch wins over the saved one, exactly as on Overview: this row
// is a status readout, and the config on disk can already differ from what is
// installed. Falls back to the config only while /api/status is unread.
const killArmed = (status?.kill_switch ?? globals?.KillSwitch ?? 'closed') === 'closed'
// Whether that setting is actually installed — same three-plus-unknown reading
// as Overview, so the two pages cannot disagree about the same router. There is
// no separate `killArmed` here any more: it compared the raw string (so "Closed"
// read as fail-OPEN) and, being a boolean, could not express "the configuration
// could not be read". Both facts come off this one readout now.
const kill = killSwitchReadout(status, globals?.KillSwitch)
const killWord =
kill.state === 'open'
? 'open'
: kill.state === 'armed'
? 'fail-closed'
: kill.state === 'inert'
? 'closed · not in effect'
: // The unknown branch splits: "closed · not reported" asserts the policy
// and doubts only the install, which is wrong when the policy itself is
// a placeholder from a configuration nothing could read.
kill.setting === 'not known'
? 'not known'
: 'closed · not reported'
// Every engine mark on this page comes from ONE reading, and that reading is
// able to say "stopped" — see planeState.engineState for why `status.running`
// could not. This page is where someone lands when the network is down; three
@@ -251,8 +303,14 @@ export default function Apply() {
// that is a leak (crit); under an open one it is the documented choice (amber).
// It used to go amber whenever `running` was true — i.e. always — and unlit
// otherwise, so the one state worth shouting about had no colour of its own.
const dataVariant: LedVariant = status?.table ? 'on' : !status ? 'off' : killArmed ? 'crit' : 'amber'
const configVariant: LedVariant = status?.enabled ? 'on' : 'amber'
const dataVariant: LedVariant =
status?.table ? 'on' : !status ? 'off' : kill.state !== 'open' ? 'crit' : 'amber'
// `enabled` is sourced from the configuration, so it means nothing when that
// could not be read (Status.config_readable): unlit, not amber, and the pip
// beside it says so rather than printing "disabled".
const configUnreadable = status?.config_readable === false
const configVariant: LedVariant = configUnreadable ? 'off' : status?.enabled ? 'on' : 'amber'
const configWord = configUnreadable ? 'unreadable' : status?.enabled ? 'enabled' : 'disabled'
const liveHash = short(status?.hash ?? '')
const pct = armed ? Math.max(0, Math.round((armed.remaining / armed.pending.total) * 100)) : 0
@@ -275,18 +333,14 @@ export default function Apply() {
<StatusPip
label="Config"
variant={configVariant}
value={status?.enabled ? 'enabled' : 'disabled'}
value={configWord}
/>
<StatusPip
label="Data plane"
variant={dataVariant}
value={status?.table ? 'nft installed' : 'no table'}
/>
<StatusPip
label="Kill-switch"
variant={killArmed ? 'on' : 'amber'}
value={killArmed ? 'fail-closed' : 'open'}
/>
<StatusPip label="Kill-switch" variant={kill.variant} value={killWord} />
</div>
{statusError && (
@@ -320,7 +374,7 @@ export default function Apply() {
v: status?.table ? 'nft installed' : 'no table',
hot: dataVariant === 'crit',
},
{ k: 'kill-switch', v: killArmed ? 'fail-closed' : 'open', hot: !killArmed },
{ k: 'kill-switch', v: killWord, hot: kill.variant === 'crit' || kill.settingHot },
]}
/>
<Module
@@ -334,7 +388,10 @@ export default function Apply() {
led={{ variant: engineVariant }}
rows={[
{ k: 'state', v: engine.word, hot: engineVariant === 'crit' },
{ k: 'config', v: status?.enabled ? 'enabled' : 'disabled' },
// Same reading as the pip above — `status.enabled` is a placeholder
// when the configuration could not be read, and "disabled" is the one
// word that must not be printed for it.
{ k: 'config', v: configWord, hot: configUnreadable },
{ k: 'schema', v: globals ? `v${globals.SchemaVersion}` : '—' },
]}
/>
@@ -374,8 +431,9 @@ export default function Apply() {
is what "live but not kept" means — so the readout survives a
reload instead of depending on what this tab remembers. */}
Applied config <span className="mono">{liveHash}</span> is live but not yet kept.
Confirm to keep it — otherwise the daemon rolls back to the last-good config when
the timer hits zero.
Confirm to keep it. At zero the daemon rolls back to the last-good config — unless
something else re-applies the data plane first, in which case it stands down and
keeps whatever is live.
</p>
<div className="cc-bar" aria-hidden="true">
<span className="cc-bar-fill" style={{ width: `${pct}%` }} />
@@ -492,7 +550,9 @@ function labelFor(kind: ActionKind): string {
case 'rollback':
return 'Rollback'
case 'expire':
return 'Auto-rollback'
// NOT "Auto-rollback": on expiry the daemon either reverts or stands down,
// and this page cannot tell which. Name the event it did observe.
return 'Window elapsed'
}
}
+10 -67
View File
@@ -104,6 +104,16 @@
font-size: 11.5px;
color: var(--faint);
}
/* Nested inside .dns-filter-copy the note is an ordinary paragraph, but the
endpoint-resolver footnote sits as a DIRECT child of the card — which makes it
a grid item. Without a span it auto-placed into the toggle's `auto` column and
sized that column to its own max-content (322px on desktop, 237px at 390px),
which starved the `1fr` copy column down to 0px: the heading then laid out one
word per line and spilled 2px past the viewport, scrolling the whole page
sideways. It is a full-width footnote under the readout — say so. */
.dns-filter-card > .dns-filter-note {
grid-column: 1 / -1;
}
.dns-readout {
display: flex;
flex-direction: column;
@@ -648,73 +658,6 @@
}
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.dns-alert-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.dns-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.dns-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-alert-row delivery controls sit inline in the row (like .dns-detour) */
.dns-alert-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.dns-alert-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.dns-alert-note--row {
margin-top: 2px;
}
@media (max-width: 640px) {
.dns-alert-ctl {
flex-basis: 100%;
order: 3;
}
}
/* alert event checkboxes */
.dns-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.dns-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.dns-event input {
accent-color: var(--fp-accent, currentColor);
}
@media (prefers-reduced-motion: reduce) {
.dns-skel {
animation: none;
+10 -482
View File
@@ -9,7 +9,7 @@ import {
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { Alert, DNSRule, Model, Resolver, RulesetStatus } from '../api'
import type { Complete, DNSRule, Model, Resolver, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
// The DNS / Blocklists page is a thin editor over the desired-state Model —
@@ -60,28 +60,9 @@ type GlobalsX = Model['Globals'] & { DNSFilter?: boolean }
type DNSModel = Model & {
Blocklists?: Blocklist[] | null
Allowlists?: Allowlist[] | null
Alerts?: Alert[] | null
DNSRules?: DNSRule[] | null
}
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers ---------------------------------------------------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
@@ -310,7 +291,6 @@ export default function DNS() {
const blocklists = useMemo(() => asArray(config?.Blocklists), [config])
const allowlists = useMemo(() => asArray(config?.Allowlists), [config])
const resolvers = useMemo<Resolver[]>(() => asArray(config?.Resolvers), [config])
const alerts = useMemo<Alert[]>(() => asArray(config?.Alerts), [config])
// Ascending Order — the engine evaluates DNS rules first-match, so the list is
// shown and edited in the order it actually runs.
const dnsRules = useMemo<DNSRule[]>(
@@ -719,78 +699,6 @@ export default function DNS() {
[config, dnsRules, save, confirm],
)
// ---- alert mutations ------------------------------------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!config) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return save({ ...config, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[config, alerts, save],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void save({ ...config, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[config, alerts, save],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!config) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void save({ ...config, Alerts: next }, `Deleted ${target.Name}`)
},
[config, alerts, save, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!config) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void save(
{ ...config, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[config, alerts, save],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void save(
{ ...config, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[config, alerts, save],
)
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Reuse the resolver detour catalog with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({ ...detourCatalog, chains: [] }),
[detourCatalog],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
const loading = config === null && loadError === null
return (
@@ -1211,57 +1119,6 @@ export default function DNS() {
)}
</div>
{/* ---- 5. ALERTS ---- */}
<div className="dns-section" aria-label="Alerts">
<header className="dns-sec-hd">
<h2 className="dns-sec-title">Alerts</h2>
<span className="dns-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="dns-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="dns-rows" aria-hidden="true">
<li className="dns-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="dns-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
{toast && (
<div className="toast" role="status">
{toast}
@@ -1271,342 +1128,6 @@ export default function DNS() {
)
}
// ---- alert add form + row --------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="dns-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="dns-add-top">
<input
className="dns-input dns-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="dns-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="dns-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="dns-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="dns-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="dns-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="dns-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="dns-alert-delivery">
<label className="dns-resp">
<span className="dns-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="dns-alert-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="dns-add-actions">
{err && (
<p className="dns-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="dns-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="dns-row-main">
<div className="dns-row-l1">
<span className="dns-row-name">{alert.Name}</span>
<span className="dns-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="dns-badge dns-badge--accent">
{e}
</span>
))}
</div>
<div className="dns-row-l2 mono">
<span className="dns-row-detail">{detail.text}</span>
{detail.masked && (
<span className="dns-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="dns-path">direct</span>
) : (
<span className="dns-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="dns-path-name">{route.name}</strong>
{route.missing && <span className="dns-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="dns-alert-note dns-alert-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="dns-alert-ctl">
<label className="dns-detour">
<span className="dns-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="dns-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
// ---- add form --------------------------------------------------------------
interface AddDraft {
@@ -2423,8 +1944,15 @@ function ruleToDraft(r: DNSRule): DNSRuleDraft {
}
}
/** Build the DNSRule to persist. Empty matcher lists are omitted, not sent as []. */
function draftToRule(d: DNSRuleDraft): DNSRule {
/**
* Build the DNSRule to persist. Empty matcher lists are omitted, not sent as [].
*
* The return type is `Complete<DNSRule>` for the same reason as Networks.fromDraft:
* this REBUILDS the rule from the draft rather than extending the one it was given,
* so a field added to `DNSRule` would otherwise be dropped on every edit with
* nothing to notice it. Completing the type makes that a build failure here.
*/
function draftToRule(d: DNSRuleDraft): Complete<DNSRule> {
const domains = parseRuleDomains(d.Domains)
const order = Number.parseInt(d.Order, 10)
return {
+16
View File
@@ -495,6 +495,22 @@
border-left: 0;
border-top: 1px solid var(--groove);
}
/* Collapsing to one column was not enough on a phone. A grid column is sized by
its widest item's MIN-CONTENT, and a <select> reports the width of its longest
option ("Allow everything — most compatible", in the mono face) — so the
column stayed ~20px wider than the plate and the COPY beside it was clipped
mid-word at the right edge, which is how a sentence about what leaks loses its
second half. The select is allowed to shrink and ellipsise its own label
instead; the chosen option is still fully readable once opened, and no text
that states a consequence is cut. */
.nw-policy-ctl {
min-width: 0;
}
.nw-policy-ctl .fp-select {
max-width: 100%;
min-width: 0;
text-overflow: ellipsis;
}
}
/* A daemon info note about the current policy — neutral by design: it states a
+296 -34
View File
@@ -2,9 +2,10 @@ import './Networks.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Inbound, Interface, Model, Status } from '../api'
import type { Complete, Inbound, Interface, Model, Status } from '../api'
import { isLanNetwork, isWanNetwork, useInterfaces } from '../srcOptions'
import { sectionNotes } from '../findings'
import { killSwitchClosed } from '../planeState'
// The Networks page is the INGRESS editor — a thin editor over Model.Inbounds,
// following the same save-then-Apply contract as Nodes/DNS/Routing: every edit
@@ -98,6 +99,36 @@ const DEFAULT_TPROXY_PORT = 12345
*
* `icmp` still carries that destination-dependence for its NON-ping half, so its
* cost line says so rather than claiming "nothing else gets out".
*
* WHY THIS COPY IS NOW A FUNCTION AND NOT A TABLE. Two audit findings, one cause:
* a constant string cannot be true about a router whose behaviour three OTHER
* settings can override.
*
* 1. MULTICAST IPTV WAS PROMISED, AND NEVER WORKS. `direct` said "Ping, multicast
* IPTV, and connecting to a VPN ... all work" — an INSTRUCTION, and the worst
* kind of wrong: someone who wants IPTV reads it, moves to the most open rung
* on the ladder (which also permits a client's ESP/GRE straight past the
* proxy), and still has no IPTV. The stream is UDP; every rule this policy
* emits carries `meta l4proto != { tcp, udp }` so UDP never reaches one, and
* the fail-closed forward chain accepts only the RFC1918/link-local daddr
* sets — 224.0.0.0/4 is not there, and unconditional drops follow. The daemon
* says exactly this in the note rendered a few pixels below on this same page.
* IPTV is now stated ONCE, as its own line, and it says it does not work.
* 2. THE `block` COST LINE WAS UNCONDITIONAL. Three settings contradict it:
* - an OPEN kill-switch — the forward chain emits no drops at all;
* - Globals.L3Tunnel — prerouting marks ICMP echo into the engine's TUN
* BEFORE the forward chain, so ping keeps working, through the tunnel;
* - Globals.UntunnelableEgress — ESP/AH/GRE/IGMP/SCTP are marked and routed
* out a named interface, so the forward chain never rules on them.
* Neither of the last two existed in the panel's `Globals` type, so the page
* could not have told the truth about them even in principle; they were added
* (api.ts) rather than papered over with a vaguer sentence.
*
* The copy therefore describes only what the POLICY still decides, and a separate
* line names whatever another setting has taken off it. Detail beyond that belongs
* to the daemon's own note for this section (`policyNotes`), which is computed
* from the running plane and rendered right underneath — this copy's job is to not
* contradict it.
*/
type Untunnelable = 'block' | 'icmp' | 'direct'
@@ -122,31 +153,197 @@ interface PolicyCopy {
works: string
cost: string | null
tone: 'good' | 'warn'
/** What some OTHER setting decides instead of this one. `null` ⇒ nothing; this policy owns it all. */
claimed: string | null
}
const UNTUNNELABLE_COPY: Record<Untunnelable, PolicyCopy> = {
block: {
// Scoped to "this traffic" on purpose. The old line — "Nothing leaves except
// through the tunnel" — was doubly loose: it was false (see the note above),
// and even read charitably it collides with directly-routed TCP, which does
// leave outside the tunnel by design.
works:
'None of this traffic leaves the router — it’s dropped, whatever your routing rules say. It’s the only setting whose promise doesn’t depend on how the rules are written.',
cost: 'Ping and traceroute stop working from your devices. So do IPsec and PPTP VPN connections made from a device on your network, multicast IPTV, and SCTP. VPNs that run over UDP — WireGuard, OpenVPN-UDP, and IPsec through NAT (IKEv2/NAT-T) — are unaffected: they go through the tunnel like everything else.',
tone: 'good',
},
icmp: {
works: 'Ping and traceroute work everywhere, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. IPsec, PPTP and IPTV also get out — but only toward addresses your routing rules already send direct, so a VPN app on a device can still open its own connection beside this one if its server is one of those.',
tone: 'warn',
},
direct: {
works: 'Ping, multicast IPTV, and connecting to a VPN from a device on your network all work.',
cost: 'All of it goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered.',
tone: 'warn',
},
/**
* How much of this traffic the policy still decides.
*
* Only three combinations are reachable, which is why this is an enum and not two
* booleans: UntunnelableEgress claims EVERY untunnelable protocol (the kernel
* routes them out its device before the forward chain runs), so once it is set
* there is nothing left for L3Tunnel to change about the policy's scope.
*
* all — neither override is on. The policy decides everything.
* exceptPing — L3Tunnel only. Ping rides the tunnel; ESP/AH/GRE/SCTP are the
* policy's.
* none — UntunnelableEgress is set. Routing settles all of it first; the
* policy answers only for the case where that route fails to come up.
*/
type PolicyScope = 'all' | 'exceptPing' | 'none'
interface PolicyContext {
/** Globals.L3Tunnel. */
l3: boolean
/** Globals.UntunnelableEgress, trimmed. */
egress: string
/** The LIVE kill-switch, normalised the daemon's way. Open ⇒ the chain has no drops. */
killSwitchOpen: boolean
}
function policyScope(ctx: PolicyContext): PolicyScope {
if (ctx.egress) return 'none'
return ctx.l3 ? 'exceptPing' : 'all'
}
/** The sentence that stops an operator "fixing" a UDP VPN that was never broken. */
const UDP_VPNS_FINE =
'VPNs that run over UDP — WireGuard, OpenVPN-UDP, and IPsec through NAT (IKEv2/NAT-T) — are unaffected either way: they go through the tunnel like everything else.'
/** Which setting took this traffic off the policy, and what it does with it. */
function claimedCopy(ctx: PolicyContext): string | null {
// Each of these says only WHY the copy above has the shape it has — which other
// setting took the traffic, and therefore why the familiar promise is missing.
// What that setting then DOES with it is the daemon's note, published for this
// same section and rendered immediately below from the RUNNING plane. Saying it
// twice would make the shorter, staler one look like a second opinion.
if (ctx.egress && ctx.l3) {
return `Two other settings decide this before the one above is asked: ping goes through the tunnel (l3_tunnel), and everything else the tunnel can’t carry is routed out “${ctx.egress}” (untunnelable_egress).`
}
if (ctx.egress) {
return `Another setting decides this before the one above is asked: everything the tunnel can’t carry is routed out “${ctx.egress}” (untunnelable_egress).`
}
if (ctx.l3) {
// Deliberately shorter than the two above: when only the L3 ingress is on, the
// daemon publishes its own note for this section directly underneath and says
// the rest (which addresses can be pinged, and why the others cannot). This
// line exists to explain the SHAPE of the copy above it — why ping is missing
// from a policy that used to decide it — not to restate the daemon.
return 'Ping and Windows tracert are taken through the tunnel before the setting above is asked (l3_tunnel), so it no longer decides them.'
}
return null
}
/**
* The copy for the policy as it is actually behaving right now.
*
* Read it as: an open kill-switch beats everything (no drops are emitted at all,
* so no rung promises anything), then the scope decides how much of the ladder's
* usual story is still this setting's to tell.
*/
function policyCopy(policy: Untunnelable, ctx: PolicyContext): PolicyCopy {
const scope = policyScope(ctx)
const claimed = claimedCopy(ctx)
// Fail-open: the forward chain emits no drops, so every rung is inert. Saying
// what IS happening beats repeating a promise nothing is keeping.
if (ctx.killSwitchOpen) {
if (scope === 'none') {
return {
works:
'Nothing is being dropped, and nothing is left for this setting to decide: the kill-switch is open, and another setting has already taken this traffic.',
cost: 'If that route ever fails to come up, the traffic leaves through your normal connection with your real IP address, quietly, instead of failing.',
tone: 'warn',
claimed,
}
}
return {
works:
scope === 'exceptPing'
? 'Nothing is being dropped: with the kill-switch open the forward chain has no drops at all, so a device’s own IPsec or PPTP connection works too.'
: 'Nothing is being dropped: with the kill-switch open the forward chain has no drops at all, so ping, traceroute and a device’s own IPsec or PPTP connection all work.',
// No "set the kill-switch to fail-closed" here: the moot note below owns
// that instruction, and printing it twice in one section is how the second
// copy stops being read.
cost: `It reaches the internet with your real IP address, around the tunnel. ${UDP_VPNS_FINE}`,
tone: 'warn',
claimed,
}
}
switch (policy) {
case 'block':
if (scope === 'none') {
return {
works:
'Where this setting still applies, the packet is dropped rather than let out — so a route that fails to come up fails honestly instead of leaking.',
cost: null,
tone: 'good',
claimed,
}
}
return {
// Scoped to "this traffic" on purpose. The old line — "Nothing leaves
// except through the tunnel" — was doubly loose: it was false (see the
// note above), and even read charitably it collides with directly-routed
// TCP, which does leave outside the tunnel by design.
works:
scope === 'exceptPing'
? 'Everything this setting still decides is dropped, whatever your routing rules say. If the ping route ever fails to come up, ping fails outright rather than leaking.'
: 'None of this traffic leaves the router — it’s dropped, whatever your routing rules say. It’s the only setting whose promise doesn’t depend on how the rules are written.',
cost:
scope === 'exceptPing'
? `IPsec and PPTP VPN connections made from a device on your network stop working, and so does SCTP. ${UDP_VPNS_FINE}`
: `Ping and traceroute stop working from your devices. So do IPsec and PPTP VPN connections made from a device on your network, and SCTP. ${UDP_VPNS_FINE}`,
tone: 'good',
claimed,
}
case 'icmp':
if (scope === 'none') {
return {
works:
'Where this setting still applies, it lets ping out directly and drops the rest.',
cost: 'So if a route ever fails to come up, ping quietly leaves with your real IP address instead of failing.',
tone: 'warn',
claimed,
}
}
if (scope === 'exceptPing') {
return {
works:
'Ping already travels through the tunnel, so this rung’s exception for it only matters if that route fails to come up.',
cost: 'IPsec, PPTP and SCTP get out toward addresses your routing rules already send direct, with your real IP address — so a VPN app on a device can still open its own connection beside this one if its server is one of those. And if the ping route fails, ping leaves with your real address rather than failing.',
tone: 'warn',
claimed,
}
}
return {
works:
'Ping and traceroute work everywhere, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. IPsec and PPTP also get out — but only toward addresses your routing rules already send direct, so a VPN app on a device can still open its own connection beside this one if its server is one of those.',
tone: 'warn',
claimed,
}
case 'direct':
if (scope === 'none') {
return {
works:
'Nothing is left for this setting to decide: another setting has already taken this traffic.',
cost: 'If that route ever fails to come up, this setting lets the traffic leave through your normal connection with your real IP address, quietly, instead of failing.',
tone: 'warn',
claimed,
}
}
return {
works:
scope === 'exceptPing'
? 'A device on your network can make its own IPsec or PPTP VPN connection. Ping already travels through the tunnel.'
: 'Ping and traceroute work, and a device on your network can make its own IPsec or PPTP VPN connection.',
cost:
scope === 'exceptPing'
? 'That traffic goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered. If the ping route ever fails to come up, ping does the same instead of failing.'
: 'All of it goes out with your real IP, around the tunnel. A VPN app left running on a device keeps its own connection open beside this one — traffic through it isn’t proxied or filtered.',
tone: 'warn',
claimed,
}
}
}
/**
* Multicast IPTV, said once and said straight.
*
* It is stated unconditionally because it is unconditionally true — the stream is
* UDP and no rule this policy emits can match UDP, on any of the three rungs — and
* it is stated at all because the page used to promise the opposite under `direct`
* and under `icmp`. Someone whose IPTV is broken arrives here looking for the
* setting that fixes it; the useful thing to tell them is that there isn't one.
*/
const IPTV_LINE =
'Multicast IPTV is not one of these things: it doesn’t pass this router on any of the three settings, and “Allow everything” won’t bring it back.'
/** The addr:port an inbound binds — the generator's clash key (listenKey). */
function listenKey(in_: Inbound): string {
if (effectiveType(in_) === 'tproxy') {
@@ -348,11 +545,24 @@ export default function Networks({ status }: { status?: Status | null }) {
// Untunnelable-traffic policy. Normalised the same way the daemon does, so an
// absent/unknown UCI value reads as `block` here too rather than as blank.
const untunnelable = normUntunnelable(config?.Globals?.Untunnelable)
const untunnelableCopy = UNTUNNELABLE_COPY[untunnelable]
// Prefer the LIVE kill-switch off /api/status; fall back to the saved config
// when the shell hasn't got a status yet.
const killSwitchOpen =
(status?.kill_switch ?? config?.Globals?.KillSwitch ?? 'closed').toLowerCase() === 'open'
// when the shell hasn't got a status yet. The comparison is the daemon's own
// (planeState.killSwitchClosed) — this page normalised and planeState.ts did
// not, so the same router read differently on two pages.
const killSwitchOpen = !killSwitchClosed(status?.kill_switch ?? config?.Globals?.KillSwitch)
// The two settings that decide part of this traffic BEFORE the policy is
// consulted. Both are read from the saved config, like `untunnelable` itself:
// /api/status reports neither, and this section describes the setting the
// operator is editing. The daemon's own note below is the live counterpart.
const untunnelableCopy = useMemo(
() =>
policyCopy(untunnelable, {
l3: config?.Globals?.L3Tunnel === true,
egress: (config?.Globals?.UntunnelableEgress ?? '').trim(),
killSwitchOpen,
}),
[untunnelable, config?.Globals?.L3Tunnel, config?.Globals?.UntunnelableEgress, killSwitchOpen],
)
// The daemon's info notes about this policy — shown beside the control they
// describe. The fail-open case has its own dedicated line below, so drop that
// one here to avoid saying the same thing twice.
@@ -519,8 +729,8 @@ export default function Networks({ status }: { status?: Status | null }) {
</header>
<p className="nw-sec-note">
The tunnel carries the traffic almost everything uses — web, video, games, email. A few
things can’t go through it no matter what: ping, and the protocols that carry IPTV or a VPN
connection. Choose what happens to those.
things can’t go through it no matter what: ping, and the protocols a device uses to make
its own VPN connection. Choose what happens to those.
</p>
<div className="nw-policy">
@@ -549,13 +759,28 @@ export default function Networks({ status }: { status?: Status | null }) {
</div>
</div>
{/* Fail-open makes the whole policy moot — say so instead of letting the
page imply something is being blocked when nothing is. */}
{/* What another setting decides instead of this one. Unlit lamp, like the
daemon's notes below: it reports a configuration, not a fault. */}
{untunnelableCopy.claimed && (
<p className="nw-sec-note nw-policy-note" role="status">
<Led variant="off" />
<span>{untunnelableCopy.claimed}</span>
</p>
)}
{/* Said once, on every setting, because it is true on every setting. */}
<p className="nw-sec-note nw-policy-note">
<Led variant="off" />
<span>{IPTV_LINE}</span>
</p>
{/* Fail-open makes the whole policy moot. The copy above now says what IS
happening; this line stays because it is the one that says what to DO. */}
{killSwitchOpen && (
<p className="nw-sec-note nw-policy-moot" role="status">
<Led variant="amber" /> This setting isn’t doing anything right now: the kill-switch is
set to fail-open, so traffic keeps flowing directly whenever the tunnel is down. Set it
to fail-closed in Settings for this choice to take effect.
<Led variant="amber" /> The kill-switch is set to fail-open, so traffic keeps flowing
directly whenever the tunnel is down. Set it to fail-closed in Settings for this choice
to take effect.
</p>
)}
@@ -758,7 +983,22 @@ function toDraft(in_: Inbound): Draft {
* and a dokodemo listener binds whatever `TargetNetwork` says. Writing anything
* else would make the daemon warn about a flag no one can see.
*/
function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound {
/**
* Build the Inbound this draft describes — REBUILT per type, never extended.
*
* Rebuilding is the point: an inbound switched from `socks` to `tproxy` binds
* 0.0.0.0:TproxyPort, so a surviving `Listen`/`Auth` from its previous life would
* be a setting the panel shows nobody and the generator ignores. `base` is
* accepted and deliberately unused for that reason.
*
* The return type is `Complete<Inbound>` so the rebuild cannot go stale in
* silence: adding a field to `Inbound` fails the BUILD in all three branches
* below until each says what it wants done with it. That is the only mechanism
* here that makes the omission loud — `Ruleset.Format` is what a missing one
* costs (see ruleset.ts). `undefined` is a positive statement of "this shape has
* no use for it", and JSON.stringify drops it, so the wire form is unchanged.
*/
function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Complete<Inbound> {
void base
const common = { Name: d.Name.trim(), Enabled: enabled }
const port = Number.parseInt(d.Port, 10)
@@ -771,6 +1011,16 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
TproxyPort: Number.parseInt(d.TproxyPort, 10) || DEFAULT_TPROXY_PORT,
TCP: d.TCP,
UDP: d.UDP,
// Binds 0.0.0.0:TproxyPort and cannot authenticate or rewrite a
// destination, so none of the listener fields mean anything here.
Listen: undefined,
Port: undefined,
Auth: undefined,
User: undefined,
Pass: undefined,
TargetAddr: undefined,
TargetPort: undefined,
TargetNetwork: undefined,
}
case 'socks':
case 'http':
@@ -784,6 +1034,12 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
Pass: d.Auth === 'password' ? d.Pass : '',
TCP: true,
UDP: true,
// A local listener diverts no network and has no fixed target.
Network: undefined,
TproxyPort: undefined,
TargetAddr: undefined,
TargetPort: undefined,
TargetNetwork: undefined,
}
case 'dokodemo':
return {
@@ -796,6 +1052,12 @@ function fromDraft(d: Draft, base: Partial<Inbound>, enabled: boolean): Inbound
TargetNetwork: d.TargetNetwork,
TCP: true,
UDP: true,
// Diverts no network, and forwards everything on without authenticating.
Network: undefined,
TproxyPort: undefined,
Auth: undefined,
User: undefined,
Pass: undefined,
}
}
}
+60
View File
@@ -196,6 +196,52 @@
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
color: var(--amber);
}
/* "not built" — the saved switch says on and the engine has no such outbound. */
.badge--crit {
border-color: color-mix(in srgb, var(--crit) 55%, var(--groove));
color: var(--crit);
}
/* ---- last-apply findings, attached to the row they are about ----
Sits under the row's own two lines, inside the row plate, so a node the
generator threw away cannot read as an ordinary enabled node. Severity carries
the colour; the accent stays reserved for controls. */
.row-findings {
margin: 6px 0 0;
padding: 0;
list-style: none;
display: flex;
flex-direction: column;
gap: 5px;
}
.row-finding {
display: flex;
align-items: flex-start;
gap: 8px;
padding: 7px 9px;
border: 1px solid color-mix(in srgb, var(--amber) 40%, var(--groove));
border-radius: 6px;
background: color-mix(in srgb, var(--sink) 35%, transparent);
}
.row-finding--critical {
border-color: color-mix(in srgb, var(--crit) 45%, var(--groove));
}
.row-finding-msg {
flex: 1;
min-width: 0;
font-size: 12px;
line-height: 1.5;
color: var(--ink);
max-width: 82ch;
overflow-wrap: anywhere;
}
/* Findings that belong to no single row (see Nodes.tsx globalFindings). */
.node-findings {
margin-bottom: calc(var(--u, 8px) * 2);
}
.node-findings .row-findings {
margin-top: 0;
}
/* masked-credential marker */
.masked {
@@ -641,6 +687,20 @@ select.fp-input {
letter-spacing: 0.06em;
color: var(--faint);
}
/* A collapsed bucket has to carry its own bad news: a 300-node subscription is
closed by default, and the per-row findings inside it are otherwise unreachable
without knowing to look. */
.group-flagged {
flex: none;
display: inline-flex;
align-items: center;
gap: 6px;
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--amber);
}
.group-rows {
margin-top: 8px;
}
+138 -6
View File
@@ -6,12 +6,14 @@ import { Button, Led, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
getStatus,
putConfig,
importWg,
updateSubscription,
ApiError,
} from '../api'
import type { Model, Node as NodeCfg, Subscription } from '../api'
import type { Model, Node as NodeCfg, StatusWarning, Subscription } from '../api'
import { entityFindings, findingsByName } from '../findings'
import { fmtBytes, fmtDate, fmtUntil } from '../format'
// The whole page is a thin editor over the desired-state Model: every mutation
@@ -269,9 +271,10 @@ function findNodeReferences(m: Model, name: string): NodeRefSite[] {
if (isNodeRef(s.FetchDetour, name))
out.push({ kind: 'subscription', label: `subscription “${s.Name}” fetch` })
}
for (const e of asArray(m.Egresses)) {
if (isNodeRef(e.Target, name)) out.push({ kind: 'egress', label: `egress “${e.Name}” target` })
}
// An egress does not reference a node. The loop that used to sit here read
// `e.Target`, a field the Go model has never had — so it was `undefined` on
// every egress and the branch could not fire. It read as coverage for a
// reference site that does not exist, which is worse than the gap it hid.
return out
}
@@ -334,7 +337,9 @@ function renameNodeReferences(m: Model, from: string, to: string): Model {
if (m.Alerts) next.Alerts = m.Alerts.map((a) => ({ ...a, Via: pfx(a.Via) }))
if (m.Subscriptions)
next.Subscriptions = m.Subscriptions.map((s) => ({ ...s, FetchDetour: pfx(s.FetchDetour) }))
if (m.Egresses) next.Egresses = m.Egresses.map((e) => ({ ...e, Target: pfx(e.Target) }))
// No Egresses pass: an egress holds no node reference to rewrite (it has no
// Target field), and rewriting one would have WRITTEN the key onto every egress
// — which PUT rejects wholesale under DisallowUnknownFields.
return next
}
@@ -519,6 +524,43 @@ export default function Nodes() {
void loadConfig()
}, [loadConfig])
// ---- what the last apply said about these nodes ---------------------------
//
// The generator drops a node it cannot build and names it: an unparseable
// share link (generate/outbound.go), a name colliding with a reserved tag, a
// WireGuard private key materialised twice (generate/wgdedup.go). Until now
// this page never read /api/status, so a node the engine had thrown away
// rendered as an ordinary row with a green toggle — the switch said on and
// there was no such outbound anywhere in the running config.
//
// Findings are attached to the ROWS, not summarised at the top: a 300-node
// subscription makes a list of names useless, and the row is where the false
// reassurance was.
const [findings, setFindings] = useState<StatusWarning[]>([])
const loadFindings = useCallback(async () => {
try {
const s = await getStatus()
setFindings(entityFindings(s.warnings, ['node', 'subscription']))
} catch {
// Status is a supplement here, not the page. Keep the last set rather than
// clearing it — a dropped poll is not the same as "the problem is fixed".
}
}, [])
useEffect(() => {
void loadFindings()
}, [loadFindings])
const nodeFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'node')),
[findings],
)
const subFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'subscription')),
[findings],
)
// Findings about nodes/subscriptions in general, which belong to no single row.
const globalFindings = useMemo(() => findings.filter((w) => !w.name), [findings])
// ---- toast + persistent apply banner --------------------------------------
const [toast, setToast] = useState<string | null>(null)
const toastTimer = useRef<number | undefined>(undefined)
@@ -573,8 +615,11 @@ export default function Nodes() {
flash(`Apply failed — ${errText(e)}`)
} finally {
setApplying(false)
// An apply is exactly what rewrites the findings — including clearing the
// ones the operator just fixed.
void loadFindings()
}
}, [flash, loadConfig])
}, [flash, loadConfig, loadFindings])
// ---- node mutations -------------------------------------------------------
const nodes = useMemo(() => asArray(config?.Nodes), [config])
@@ -994,6 +1039,14 @@ export default function Nodes() {
</div>
)}
{/* Findings about nodes in general — no single row owns them, so they sit
above the lists rather than being dropped for having no name. */}
{globalFindings.length > 0 && (
<div className="node-findings">
<RowFindings findings={globalFindings} />
</div>
)}
{/* ---- NODES ---- */}
<div className="node-section" aria-label="Nodes">
<header className="sec-hd">
@@ -1149,6 +1202,7 @@ export default function Nodes() {
<NodeGroup
key={g.key || '__manual__'}
group={g}
findings={nodeFindings}
open={isGroupOpen(g)}
busy={busy}
egressNames={egressNames}
@@ -1238,6 +1292,7 @@ export default function Nodes() {
busy={busy}
catalog={detourCatalog}
valid={detourValid}
findings={subFindings.get(s.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => toggleSub(i, on)}
onDelete={() => removeSub(i)}
onEdit={(patch) => editSub(i, patch)}
@@ -1264,6 +1319,7 @@ function NodeGroup({
open,
busy,
egressNames,
findings,
onToggle,
onToggleNode,
onRemoveNode,
@@ -1274,6 +1330,8 @@ function NodeGroup({
open: boolean
busy: boolean
egressNames: string[]
/** Last-apply findings per node name (findings.ts findingsByName). */
findings: Map<string, StatusWarning[]>
onToggle: () => void
onToggleNode: (idx: number, on: boolean) => void
onRemoveNode: (idx: number) => void
@@ -1283,6 +1341,13 @@ function NodeGroup({
const panelId = `node-group-${group.key || 'manual'}`
// The same inventory count as the section header, scoped to this bucket.
const count = useMemo(() => fmtEnabled(group.items.map((i) => i.node)), [group.items])
// How many nodes in this bucket the last apply had something to say about —
// shown on the COLLAPSED header, because a subscription of 300 nodes is
// collapsed by default and the row badge below would never be seen otherwise.
const flagged = useMemo(
() => group.items.filter(({ node }) => findings.has(node.Name)).length,
[group.items, findings],
)
return (
<section className={`node-group${open ? ' node-group--open' : ''}`}>
<h3 className="group-hd-wrap">
@@ -1296,6 +1361,12 @@ function NodeGroup({
<span className="group-caret" aria-hidden="true" />
<span className="group-name">{group.label}</span>
<span className="group-count mono">{count}</span>
{flagged > 0 && (
<span className="group-flagged" title="Findings from the last apply">
<Led variant="amber" />
{flagged} flagged
</span>
)}
</button>
</h3>
{open && group.key !== '' && (
@@ -1312,6 +1383,7 @@ function NodeGroup({
node={node}
busy={busy}
egressNames={egressNames}
findings={findings.get(node.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => onToggleNode(idx, on)}
onDelete={() => onRemoveNode(idx)}
onRename={(name, onError) => onRenameNode(idx, name, onError)}
@@ -1324,10 +1396,29 @@ function NodeGroup({
)
}
/** One shared empty array, so a clean row doesn't get a fresh identity per render. */
const EMPTY_FINDINGS: StatusWarning[] = []
/**
* Did the generator say it left this entity OUT of the engine config?
*
* The producers all end the sentence with the same word — "(skipped)" for an
* unparseable share link or a bad WireGuard endpoint (generate/outbound.go),
* "skipped" for a name colliding with a reserved tag — and wgdedup says only one
* of the duplicates "is kept". Read the daemon's word rather than inventing a
* verdict: a finding that does NOT say this may well be about a node that is
* running perfectly, and badging it "not built" would be a new lie in place of
* the old one.
*/
function skipped(findings: StatusWarning[]): boolean {
return findings.some((f) => /\bskipped\b|\bis kept\b/i.test(f.message))
}
function NodeRow({
node,
busy,
egressNames,
findings,
onToggle,
onDelete,
onRename,
@@ -1336,6 +1427,8 @@ function NodeRow({
node: NodeCfg
busy: boolean
egressNames: string[]
/** What the last apply said about THIS node; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onRename: (name: string, onError: (msg: string) => void) => Promise<boolean>
@@ -1483,6 +1576,16 @@ function NodeRow({
)}
<span className="badge">{proto}</span>
{node.Stale && <span className="badge badge--warn">stale</span>}
{/* The toggle above is the SAVED state. When the last apply couldn't
build this node the engine has no such outbound, and the two
disagree — so the row says which, rather than leaving a green
switch to imply the node is carrying traffic. The word is the
daemon's own where it used one. */}
{findings.length > 0 && (
<span className={`badge badge--${skipped(findings) ? 'crit' : 'warn'}`}>
{skipped(findings) ? 'not built' : 'flagged'}
</span>
)}
</div>
{renameErr && (
<p className="row-err" role="alert">
@@ -1506,6 +1609,7 @@ function NodeRow({
</span>
)}
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
{canPin && (
@@ -1625,6 +1729,7 @@ function SubRow({
busy,
catalog,
valid,
findings,
onToggle,
onDelete,
onEdit,
@@ -1635,6 +1740,8 @@ function SubRow({
busy: boolean
catalog: DetourCatalog
valid: Set<string>
/** What the last apply said about THIS subscription; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onEdit: (patch: Subscription) => Promise<boolean>
@@ -1659,6 +1766,7 @@ function SubRow({
<span className="row-name">{sub.Name}</span>
{sub.Format && sub.Format !== 'auto' && <span className="badge">{sub.Format}</span>}
{sub.FetchVia === 'proxy' && <span className="badge">via proxy</span>}
{findings.length > 0 && <span className="badge badge--warn">flagged</span>}
</div>
<div className="row-line2 mono">
<span className="row-host">{host}</span>
@@ -1671,6 +1779,7 @@ function SubRow({
every {interval} · {count} node{count === 1 ? '' : 's'}
</span>
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
<Button
@@ -2142,6 +2251,29 @@ function HeaderRows({
)
}
/**
* What the last apply said about THIS row, under the row it is about.
*
* Deliberately inside the row rather than in a list at the top of the page: the
* failure being fixed is a node that looks fine, and a name in a summary three
* screens up does not fix that. The wording is the daemon's own — these messages
* already name the entity and say what was done about it ("(skipped)", "only X
* is kept"), so paraphrasing them here would only invent a second vocabulary.
*/
function RowFindings({ findings }: { findings: StatusWarning[] }) {
if (findings.length === 0) return null
return (
<ul className="row-findings" aria-label="Findings from the last apply">
{findings.map((f, i) => (
<li key={i} className={`row-finding row-finding--${f.severity}`}>
<Led variant={f.severity === 'critical' ? 'crit' : 'amber'} />
<span className="row-finding-msg">{f.message}</span>
</li>
))}
</ul>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="empty-plate">
+61 -20
View File
@@ -14,8 +14,8 @@ import type { Model, Stats, Status, StatusWarning } from '../api'
import { confirmTimeout } from '../pendingConfirm'
import { navigate } from '../router'
import type { Route } from '../router'
import { attentionFindings } from '../findings'
import { engineReadout, protectionState } from '../planeState'
import { attentionFindings, truncationNote } from '../findings'
import { engineReadout, killSwitchReadout, protectionState } from '../planeState'
// null-safe length for a Go slice that may arrive as null.
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
@@ -226,11 +226,11 @@ export function Overview({
// ---- derived display state ----
const g = config?.Globals
// The LIVE kill-switch from /api/status wins over the saved config: this is a
// status readout, so it must describe what is actually installed. Reading the
// config here let the strip claim "fail-closed" while an apply-time finding
// said the running plane was fail-open — two truths on one screen.
const killArmed = (status?.kill_switch ?? g?.KillSwitch ?? 'closed') === 'closed'
// There is deliberately no local `killArmed` any more. The page asked the same
// question twice — once here and once inside killSwitchReadout — and the local
// copy was the poorer of the two: it compared the raw string (so "Closed" read
// as fail-OPEN) and it was a boolean, which cannot say "the configuration could
// not be read". Both answers now come from the readout below.
// Offer rollback only when the daemon has something to revert to (armed
// commit-confirm snapshot or an engine last-good); otherwise hide the control.
const canRollback = status?.can_rollback ?? false
@@ -309,15 +309,19 @@ export function Overview({
const engineVariant: LedVariant = engine.variant
const protection = protectionState(status)
// Configured fail-closed AND actually enforcing it. `none` means nothing is
// installed, so the setting is inert no matter what it says.
const killInEffect = killArmed && status?.plane !== 'none'
// Configured fail-closed, actually enforcing it, or not known — three answers,
// and the third is not folded into the first. See planeState.killSwitchReadout.
const kill = killSwitchReadout(status, g?.KillSwitch)
// Findings that need attention. `info` notes are statements about the config,
// not problems, so they live beside the setting they describe (see findings.ts)
// — keeping this list to things someone could actually act on.
const warnings = attentionFindings(status?.warnings)
const criticalCount = warnings.filter((w) => w.severity === 'critical').length
// The daemon caps the published list at 50 and says so in an `info` note — the
// one channel this page filters away. Carried separately so the list can admit
// it is not the whole list. See findings.ts truncationNote.
const truncated = truncationNote(status?.warnings)
return (
<section className="page" aria-label="Overview">
@@ -340,7 +344,7 @@ export function Overview({
</p>
)}
<Findings warnings={warnings} criticalCount={criticalCount} />
<Findings warnings={warnings} criticalCount={criticalCount} truncated={truncated} />
<div className="grid">
{/* Groups, not nodes: a group is where a dial path is defined, so it is the
@@ -447,17 +451,21 @@ export function Overview({
/>
{/* A kill-switch set to fail-closed is only ARMED if something is actually
installed to enforce it. With no plane it is configured but inert, and
saying "ARMED" there would be a false reassurance next to a readout
that says nothing is protected. */}
installed to enforce it, and "we haven't been told" is neither. With no
plane it is configured but inert; with no reading the lamp stays unlit
rather than joining the healthy branch by default. */}
<Module
name="Kill-switch"
value={killInEffect ? 'ARMED' : killArmed ? 'NOT IN EFFECT' : 'OPEN'}
led={{ variant: killInEffect ? 'on' : killArmed ? 'crit' : 'amber' }}
value={kill.value}
led={{ variant: kill.variant }}
rows={[
{ k: 'setting', v: killArmed ? 'fail-closed' : 'fail-open', hot: !killArmed },
...(killArmed && !killInEffect
? [{ k: 'blocking now', v: 'no — nothing installed', hot: true }]
// From the readout, not from a second local comparison: the old
// `killArmed ? 'fail-closed' : 'fail-open'` had no third answer, so an
// unreadable configuration printed a confident "fail-closed" beneath a
// lamp that said NOT REPORTED. See planeState.killSwitchReadout.
{ k: 'setting', v: kill.setting, hot: kill.settingHot },
...(kill.blockingNow
? [{ k: 'blocking now', v: kill.blockingNow, hot: kill.hot }]
: [{ k: 'ipv6', v: g?.IPv6 ? 'covered' : 'off' }]),
{ k: 'confirm', v: g?.ConfirmTimeout ? `${g.ConfirmTimeout}s window` : 'no auto-rollback' },
]}
@@ -536,11 +544,23 @@ const SECTION_ROUTE: Record<string, Route> = {
rule: 'routing',
ruleset: 'routing',
blocklist: 'dns',
allowlist: 'dns',
resolver: 'dns',
dns_rule: 'dns',
device: 'devices',
chain: 'targets',
group: 'targets',
// A node the generator dropped (unparseable share link, duplicate WireGuard
// key, name colliding with a reserved tag) is reported under `node` — and had
// nowhere to jump to, so the one page that could show it a green toggle was
// also the one page the finding could not reach.
node: 'nodes',
subscription: 'nodes',
egress: 'targets',
inbound: 'networks',
interface: 'networks',
profile: 'profiles',
alert: 'settings',
// The standing note about non-TCP/UDP traffic — its control lives on Networks.
untunnelable: 'networks',
}
@@ -558,11 +578,14 @@ const SECTION_ROUTE: Record<string, Route> = {
function Findings({
warnings,
criticalCount,
truncated,
}: {
warnings: StatusWarning[]
criticalCount: number
/** The daemon's "N further suppressed" note, when the list was capped. */
truncated: StatusWarning | null
}) {
if (warnings.length === 0) return null
if (warnings.length === 0 && !truncated) return null
const rank = { critical: 0, warning: 1, info: 2 } as const
const sorted = [...warnings].sort((a, b) => rank[a.severity] - rank[b.severity])
@@ -572,6 +595,10 @@ function Findings({
<header className="findings-hd">
<h2 className="findings-title">Last apply</h2>
<span className="findings-count mono">
{/* "at least" whenever the list was capped: the counts below it are a
floor, not a total, and the cap drops the least severe FIRST — so
on a config with fifty criticals the thing it drops is a critical. */}
{truncated ? 'at least ' : ''}
{criticalCount > 0
? `${criticalCount} critical · ${warnings.length} total`
: `${warnings.length} note${warnings.length === 1 ? '' : 's'}`}
@@ -610,6 +637,20 @@ function Findings({
</li>
)
})}
{/* The list saying it is not the whole list. Last, because it is about
everything above it — and never filtered out with the other `info`
notes, which is where it used to disappear. */}
{truncated && (
<li className="finding finding--truncated">
<Led variant="amber" />
<div className="finding-copy">
<span className="finding-where mono">list truncated</span>
<span className="finding-msg">
Some findings are missing from this list. {truncated.message}
</span>
</div>
</li>
)}
</ul>
</section>
)
+67 -15
View File
@@ -13,6 +13,8 @@ import {
} from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
import { killSwitchClosed } from '../planeState'
import { carryRulesetFormat } from '../ruleset'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -60,16 +62,31 @@ type RuleForce = {
}
/**
* Everything `Proto` can match, and nothing else. The engine understands two
* transports and exactly ten application protocols its sniffers can name
* (generate/route.go sniffedProtocols); a value outside this set builds a rule
* that is perfectly valid and can never fire — so its traffic quietly falls
* Everything `Proto` can match, and nothing else. A value outside this set builds
* a rule that is perfectly valid and can never fire — so its traffic quietly falls
* through to whatever rule sits below it. That is why this is a closed list and
* not a text box.
*
* Split into two groups because they answer different questions: the transport is
* known the moment a packet arrives, while an app protocol is only known once the
* first bytes have been read and labelled.
* Three groups, because the engine reads them through three different matchers
* (generate/route.go, ruleMatchers) and they answer different questions:
*
* - Transport — the L4 network. Known the moment a packet arrives.
* - Detected protocol — the L7 label a sniffer puts on a connection once its
* first bytes have been read. This group, and ONLY this group, is the engine's
* `sniffedProtocols` set; anything else routed into that matcher is inert.
* - Layer 3 — ICMP. Not a sniffed label: it lands in the emitted rule's
* `network`, never in `protocol` (the sniffers are skipped outright for an
* ICMP flow, so they never report "icmp"). All three spellings are the SAME
* one network; `icmpv4`/`icmpv6` additionally pin `ip_version`, which the
* engine derives from the destination address.
*
* ICMP carries caveats the picker deliberately does not try to enforce, because
* the daemon reports each one against the whole config on apply: it reaches the
* engine only while globals l3_tunnel is on, it has no ports (a port matcher
* beside it can never be satisfied), `icmpv6` also needs globals ipv6 on, and it
* is DROPPED rather than falling through when routed at a target that cannot
* carry layer 3 — i.e. every proxy protocol. Only wireguard/AmneziaWG nodes and
* direct/interface egresses can carry a ping.
*/
const PROTO_TRANSPORT: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'tcp', label: 'TCP' },
@@ -87,10 +104,27 @@ const PROTO_APP: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'rdp', label: 'RDP' },
{ id: 'ntp', label: 'NTP' },
]
const PROTO_VALUES = new Set([...PROTO_TRANSPORT, ...PROTO_APP].map((p) => p.id))
/**
* The family-qualified spellings are offered next to plain `icmp` rather than
* hidden behind it: the engine treats them as first-class and the difference is
* observable (an `ip_version` item on the same rule), so hiding them would leave a
* capability reachable only by hand-editing /etc/config/shater — and would mean
* that anyone who edited such a rule here lost the narrowing on the next save.
*/
const PROTO_L3: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'icmp', label: 'ICMP (ping)' },
{ id: 'icmpv4', label: 'ICMP — IPv4 only' },
{ id: 'icmpv6', label: 'ICMP — IPv6 only' },
]
const PROTO_VALUES = new Set(
[...PROTO_TRANSPORT, ...PROTO_APP, ...PROTO_L3].map((p) => p.id),
)
/** The Proto picker's option list — shared by the inline add row and the editor. */
function ProtoOptions({ value }: { value: string }) {
// The engine lower-cases `Proto` before matching it, so a hand-written `ICMP`
// is a working rule; judge it the same way and flag only what really is inert.
const matches = PROTO_VALUES.has(value.trim().toLowerCase())
return (
<>
<option value="">any</option>
@@ -108,10 +142,18 @@ function ProtoOptions({ value }: { value: string }) {
</option>
))}
</optgroup>
{/* A stored value the engine can't detect is kept and flagged, never
silently rewritten — the rule it belongs to is live right now. */}
<optgroup label="Layer 3">
{PROTO_L3.map((p) => (
<option key={p.id} value={p.id}>
{p.label}
</option>
))}
</optgroup>
{/* A stored value none of the groups spells verbatim is kept and offered as
written, never silently rewritten — the rule it belongs to is live right
now. It is flagged only when the engine cannot match it either. */}
{value !== '' && !PROTO_VALUES.has(value) && (
<option value={value}>{value} — never matches</option>
<option value={value}>{matches ? value : `${value} — never matches`}</option>
)}
</>
)
@@ -732,7 +774,10 @@ export default function Routing() {
[rules],
)
const killSwitch = (config?.Globals?.KillSwitch ?? 'closed') === 'open' ? 'open' : 'closed'
// Normalised by the daemon's own rule — `=== 'open'` read "OPEN" as fail-CLOSED,
// so this dialog would have described a blocking router that isn't blocking.
// See planeState.killSwitchClosed.
const killSwitch = killSwitchClosed(config?.Globals?.KillSwitch) ? 'closed' : 'open'
const onToggle = useCallback(
async (name: string) => {
@@ -2030,9 +2075,16 @@ function RulesetForm({
return
}
// Type is meaningless for geosite/geoip (the remote .srs is self-describing).
const rs: Ruleset = isGeoSource(source)
? { Name: nm, Source: source }
: { Name: nm, Type: type, Source: source }
//
// `Format` is CARRIED, not rebuilt: this form has no control for it, so every
// save that dropped it destroyed a value only SSH could put back — and
// renaming a list came through here. See ruleset.ts for what it costs.
const rs: Ruleset = carryRulesetFormat(
isGeoSource(source)
? { Name: nm, Source: source }
: { Name: nm, Type: type, Source: source },
initial,
)
if (source === 'inline') {
const entries = parseLines(text)
if (entries.length === 0) {
+62 -18
View File
@@ -2,8 +2,10 @@ import './Settings.css'
import { useCallback, useEffect, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { AlertsSection } from './Alerts'
import { apply as apiApply, downloadLog, getConfig, putConfig, ApiError } from '../api'
import type { Globals, LogRange, Model } from '../api'
import { killSwitchClosed } from '../planeState'
// The Settings page is a thin editor over the desired-state Model's Globals —
// same save→apply split as DNS.tsx: every edit rewrites model.Globals in-place,
@@ -116,10 +118,18 @@ const LOG_LEVELS: ReadonlyArray<{ value: string; label: string }> = [
// Logging/stats backend. "off" collects nothing; "memory" keeps aggregates in RAM
// (lost on restart); "sqlite" persists logs to /etc/shater/stats.db so they survive
// a restart, bounded by the retention row caps + the disk-limit knob below.
//
// THE VALUE `sqlite` IS A HISTORICAL NAME AND THE LABELS NO LONGER REPEAT IT. There
// is no SQLite in the daemon: the store is bbolt (stats/boltring.go) — pure Go, no
// CGO, already linked into the binary via experimental/cachefile — and it was chosen
// precisely to be rid of "the stop-the-world window the sqlite VACUUM used to
// impose", in that file's own words. The wire value has to stay (it is in every
// shipped config, and the daemon still matches on it); what the operator READS
// should describe where the logs go, which is the disk.
const STATS_BACKENDS: ReadonlyArray<{ value: string; label: string }> = [
{ value: 'off', label: 'Off — no logging' },
{ value: 'memory', label: 'Memory (RAM)' },
{ value: 'sqlite', label: 'SQLite · persistent' },
{ value: 'sqlite', label: 'Disk · survives a restart' },
]
// ---- page ------------------------------------------------------------------
@@ -218,14 +228,16 @@ export default function Settings() {
const ringUnlimited = (globals?.StatsRingSize ?? 0) === 0
const timelineUnlimited = (globals?.StatsTimelineMinutes ?? 0) === 0
const domainsUnlimited = (globals?.StatsMaxDomains ?? 0) === 0
// SQLite disk cap: 0 = unlimited (stats.db grows with the disk).
// Disk cap: 0 = unlimited (stats.db grows with the disk).
const diskUnlimited = (globals?.StatsDiskLimitMB ?? 0) === 0
// Logging backend. Default "memory" when the field is absent (older config). When
// "off", nothing is collected, so the retention sizes below don't apply — dim them.
const statsBackend = globals?.StatsBackend || 'memory'
const loggingOff = statsBackend === 'off'
const loggingSqlite = statsBackend === 'sqlite'
// The wire value is still `sqlite` (historical — see STATS_BACKENDS); the store
// is bbolt on disk, so everything the operator reads calls it the disk backend.
const loggingDisk = statsBackend === 'sqlite'
// Retention controls are meaningless with logging off; disable them there.
const retentionDisabledCtl = busy || !ready || loggingOff
@@ -257,7 +269,11 @@ export default function Settings() {
// the Targets page. Absent ⇒ enabled (older config), so read it as `!== false`.
const groupHealthOn = globals?.GroupHealth !== false
const killSwitch = globals?.KillSwitch === 'open' ? 'open' : 'closed'
// Normalised the daemon's way (planeState.killSwitchClosed), not by string
// equality: `kill_switch 'OPEN'` is fail-OPEN on the router, and `=== 'open'`
// read it as closed — the panel would have drawn the protective setting over a
// router that has none.
const killSwitch = killSwitchClosed(globals?.KillSwitch) ? 'closed' : 'open'
/**
* The master switch, which is the most destructive control in the panel and was
@@ -300,10 +316,16 @@ export default function Settings() {
},
[confirm, killSwitch, setGlobal],
)
// "so nothing leaks unproxied" claimed more than the holding plane promises.
// netplane/nft.go states its own contract as "No client TRAFFIC reaches the WAN"
// and names the exception in the same paragraph: clients still reach the router's
// resolver and dnsmasq forwards those lookups to the ISP in the clear. It cannot
// be closed — blocking it would also cut the daemon's own name resolution, and
// with it any chance of recovering unattended.
const killNote =
killSwitch === 'open'
? 'Fail-open — if the engine stops, traffic falls back to the direct WAN. Stays online, but unprotected.'
: 'Fail-closed — if the engine stops, LAN→WAN is blocked so nothing leaks unproxied.'
: 'Fail-closed — if the engine stops, LAN→WAN is blocked so no traffic from your devices reaches the internet. DNS is the exception: lookups sent to the router still go out to your provider in the clear, which is what lets the router recover on its own.'
const loading = config === null && loadError === null
@@ -374,7 +396,16 @@ export default function Settings() {
<Field
label="Panel port"
note="Admin-panel port. 0 uses the default 8088. A change needs a restart to rebind."
// "needs a restart to rebind" read as a promise that the rebind
// happens. It is not one the panel can make: cmd/shaterd/main.go
// treats a failed panel listen as `logger.Warn("panel server
// unavailable (daemon continues)")` and carries on — the daemon keeps
// routing traffic and the panel simply is not there. Nothing reports
// it in the UI either, because the UI is what went missing, and
// `status.panel_port` keeps naming the CONFIGURED port regardless
// (which is also what LuCI builds its "Open panel" button from). So
// the note names the failure and where the answer actually is.
note="Admin-panel port. 0 uses the default 8088. A change takes effect on restart — and if the new port is already taken the panel does not come back at all: the daemon keeps running and only says so in its log."
>
<InlineEdit<number>
value={globals?.PanelPort ?? 0}
@@ -406,7 +437,7 @@ export default function Settings() {
<Field
label="Log level"
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level."
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level — set up where they go in the Alerts section below."
>
<Select
value={globals?.LogLevel || 'warning'}
@@ -620,6 +651,13 @@ export default function Settings() {
</Field>
</Group>
{/* ---- ALERTS ---- */}
{/* Extracted from the DNS page — out-of-band notifications belong with
the appliance-wide knobs, next to the log level whose note points
here. Renders its own section header (same plate as a Group); all
writes go through `save`, so the dirty banner and toast stay one. */}
<AlertsSection config={config} busy={busy} loading={loading} onSave={save} />
{/* ---- STATISTICS & LOGGING ---- */}
<Group
title="Statistics &amp; logging"
@@ -627,7 +665,7 @@ export default function Settings() {
>
<Field
label="Logging backend"
note="Off: collect nothing. Memory: fast, lost on restart, RAM-bounded. SQLite: survives restart, disk-bounded."
note="Off: collect nothing. Memory: fast, lost on restart, RAM-bounded. Disk: survives a restart, disk-bounded."
>
<Select
value={statsBackend}
@@ -642,7 +680,7 @@ export default function Settings() {
v === 'off'
? 'Logging off — collecting nothing'
: v === 'sqlite'
? 'Logging backend → SQLite (persistent)'
? 'Logging backend → disk (survives a restart)'
: 'Logging backend → memory',
)
}
@@ -655,13 +693,13 @@ export default function Settings() {
Insights page shows an off state. The retention limits below apply once logging
is turned back on.
</p>
) : loggingSqlite ? (
) : loggingDisk ? (
<p className="set-group-note">
Logs persist to <span className="mono">/etc/shater/stats.db</span> and survive a
restart. The <strong>entries</strong> limit below caps rows kept per log table; the{' '}
<strong>disk limit</strong> caps the whole <span className="mono">stats.db</span>{' '}
file (oldest rows are pruned to stay under it). Set any size to <strong>0</strong>{' '}
for <strong>Unlimited</strong>.
<strong>disk limit</strong> aims the whole <span className="mono">stats.db</span>{' '}
file at a size (oldest rows are deleted and the file rebuilt to stay near it). Set
any size to <strong>0</strong> for <strong>Unlimited</strong>.
</p>
) : (
<p className="set-group-note">
@@ -698,10 +736,16 @@ export default function Settings() {
</p>
)}
{loggingSqlite && (
{loggingDisk && (
<Field
label="SQLite disk limit (MB) (0 = unlimited)"
note="Hard cap on the on-disk stats.db file. A positive number is the ceiling — oldest rows are pruned and the DB vacuumed to stay under it; 0 lets it grow with the disk."
label="Disk limit (MB) (0 = unlimited)"
// Was: "oldest rows are pruned and the DB vacuumed". There is no
// SQLite and no VACUUM here — the store is bbolt, and reclaiming
// space means rebuilding the file (bbolt.Compact + atomic swap).
// The rebuild is SKIPPED when the filesystem cannot fit the
// transient second copy, so "ceiling" was a promise too: the DB
// then sits over the cap until space frees up. Both are said.
note="Target size for the on-disk stats.db file. Above it, the oldest rows are deleted and the file is rebuilt to give the space back — the rebuild needs room for a temporary second copy, so on a full disk the file stays over the limit until space frees up. 0 lets it grow with the disk."
>
<InlineEdit<number>
value={globals?.StatsDiskLimitMB ?? 0}
@@ -711,7 +755,7 @@ export default function Settings() {
inputMode="numeric"
width="9rem"
placeholder="Unlimited"
ariaLabel="SQLite disk limit in MB (0 = unlimited)"
ariaLabel="Stats database disk limit in MB (0 = unlimited)"
busy={busy}
disabled={retentionDisabledCtl}
onCommit={(v) =>
@@ -720,7 +764,7 @@ export default function Settings() {
/>
</Field>
)}
{loggingSqlite && diskUnlimited && (
{loggingDisk && diskUnlimited && (
<p className="set-warn" role="status">
Unlimited — stats.db grows with disk; set a cap (MB) to bound it.
</p>
+19 -46
View File
@@ -29,6 +29,13 @@ import type {
Node,
} from '../api'
import { fmtClock, fmtDuration } from '../format'
import {
DPI_TYPES,
EGRESS_TYPES,
UNKNOWN_EGRESS_TYPE_HINT,
egressTypeInfo,
nextEgress,
} from '../egressEdit'
// The Targets page is a thin editor over the desired-state Model — the same
// shape as DNS.tsx and Nodes.tsx. It manages the three things a routing rule
@@ -112,9 +119,9 @@ function findReferences(m: Model, kind: RefKind, name: string): RefSite[] {
if (asArray(c.Hops).some((h) => isPrefixed(h, kind, name)))
out.push({ label: `chain “${c.Name}” hop` })
}
for (const e of asArray(m.Egresses)) {
if (isPrefixed(e.Target, kind, name)) out.push({ label: `egress “${e.Name}” target` })
}
// No Egresses loop: an egress carries no target. The one that stood here read
// `e.Target`, absent from the Go model, so it was `undefined` on every egress
// and never once matched — a dead branch shaped like a covered case.
for (const r of asArray(m.Resolvers)) {
if (isPrefixed(r.Detour, kind, name)) out.push({ label: `resolver “${r.Name}” DNS path` })
}
@@ -152,7 +159,8 @@ function renameReferences(m: Model, kind: RefKind, from: string, to: string): Mo
if (m.Rules) next.Rules = m.Rules.map((r) => ({ ...r, Target: pfx(r.Target), Egress: br(r.Egress) }))
if (m.Chains)
next.Chains = m.Chains.map((c) => ({ ...c, Hops: c.Hops ? c.Hops.map((h) => pfx(h) ?? h) : c.Hops }))
if (m.Egresses) next.Egresses = m.Egresses.map((e) => ({ ...e, Target: pfx(e.Target) }))
// No Egresses pass — see targetRefs above: an egress holds no target to rewrite,
// and writing the key on would make PUT reject the whole rename with a 400.
if (m.Resolvers) next.Resolvers = m.Resolvers.map((r) => ({ ...r, Detour: pfx(r.Detour) }))
if (m.Alerts) next.Alerts = m.Alerts.map((a) => ({ ...a, Via: pfx(a.Via) }))
if (m.Subscriptions)
@@ -244,29 +252,6 @@ function normStrategy(raw: string | undefined): string {
const PROTOS = ['vless', 'vmess', 'trojan', 'ss'] as const
/**
* The three egress kinds that produce a real way out. `Proxy` and `Block` were
* removed: neither ever created an outbound, so everything bound to them fell
* through to the plain WAN with the real address. Send traffic through a proxy by
* routing it at a group/node/chain, and drop it with the `block` target on a rule.
*/
const EGRESS_TYPES: ReadonlyArray<{ id: string; label: string; blurb: string }> = [
{
id: 'interface',
label: 'Interface — out a specific WAN or tunnel',
blurb: 'Binds to one device (wan, wg0, …) so this traffic leaves over that uplink.',
},
{
id: 'direct',
label: 'Direct — straight out, with an optional DPI preset',
blurb: 'Uses the normal route. Its point is the DPI preset below, applied to what you route here.',
},
{
id: 'byedpi',
label: 'ByeDPI — through the local ciadpi desync proxy',
blurb: 'Hands traffic to ciadpi on 127.0.0.1, which desyncs it and goes out direct.',
},
]
const EGRESS_TYPE_LABEL: Record<string, string> = Object.fromEntries(
EGRESS_TYPES.map((t) => [t.id, t.label.split(' — ')[0]]),
)
@@ -278,12 +263,6 @@ const DPI_PRESETS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'spoof', label: 'Spoof' },
]
/**
* Types whose native DPI-bypass preset applies. NOT byedpi: the desync happens
* inside the ciadpi process, and the engine's tls_* flags are never stamped on top
* of it — so the control is hidden there rather than accepted and dropped.
*/
const DPI_TYPES = new Set(['interface', 'direct'])
// ---- group membership health ------------------------------------------------
//
@@ -2839,7 +2818,7 @@ function EgressEditor({
const [port, setPort] = useState(initial?.Port != null ? String(initial.Port) : '')
const [dpi, setDpi] = useState(initial?.DPI || 'off')
const [err, setErr] = useState<string | null>(null)
const typeInfo = EGRESS_TYPES.find((t) => t.id === type)
const typeInfo = egressTypeInfo(type)
// This egress was byedpi when the editor opened — its own type stays legal
// even with the package gone, so saved config can always round-trip.
const wasByedpi = initial?.Type === 'byedpi'
@@ -2856,14 +2835,11 @@ function EgressEditor({
if (type === 'byedpi' && byedpiLocked)
return setErr('Install the byedpi package to add a ByeDPI egress.')
setErr(null)
const base: Egress = { ...(initial ?? ({} as Egress)), Name: nm, Type: type }
// Only carry the fields the chosen type uses; clear the rest. `Target` belonged
// to the removed `proxy` type and is cleared unconditionally.
base.Interface = type === 'interface' ? iface.trim() : undefined
base.Target = undefined
base.Port = type === 'byedpi' ? Number(port.trim()) || 1080 : undefined
base.DPI = DPI_TYPES.has(type) ? dpi : undefined
await onSave(base)
// Carry the fields the chosen type uses and clear the rest — but ONLY for a
// type this editor renders those fields for. An unknown type keeps every
// stored setting untouched, because this form showed the operator none of
// them and must not delete what it declined to display. See nextEgress.
await onSave(nextEgress(initial, { name: nm, type, iface, port, dpi }))
}
return (
@@ -2912,10 +2888,7 @@ function EgressEditor({
})}
{!typeInfo && <option value={type}>{type || '—'} (unknown)</option>}
</select>
<p className="tg-fhint">
{typeInfo?.blurb ??
'This engine builds no outbound for that type, so everything routed here is blocked. Pick one above.'}
</p>
<p className="tg-fhint">{typeInfo?.blurb ?? UNKNOWN_EGRESS_TYPE_HINT}</p>
{byedpiLocked && (
<p className="tg-fhint">
Install the <code>byedpi</code> package to enable the ByeDPI egress.
+241 -1
View File
@@ -15,7 +15,13 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { engineReadout, engineState, protectionState } from './planeState.ts'
import {
engineReadout,
engineState,
killSwitchClosed,
killSwitchReadout,
protectionState,
} from './planeState.ts'
import type { Status, Traffic } from './api.ts'
/** A healthy, fully-installed router; `traffic` is what each case varies. */
@@ -189,3 +195,237 @@ test('no status at all is unknown, not down', () => {
assert.equal(engineReadout(null).variant, 'off')
assert.equal(engineReadout(null).word, 'checking…')
})
// --- killSwitchReadout: "I don't know" is not "it's armed" -------------------
//
// The Overview module read `killArmed && status?.plane !== 'none'`, and
// `undefined !== 'none'` is true — so a daemon that never reported `plane`, and
// the seconds before the first status arrives, both lit a green lamp over the
// word ARMED. These pin the fourth answer that expression could not express.
test('a daemon that does not report `plane` reads as not reported, never ARMED', () => {
const { plane, ...noPlane } = status()
void plane
const k = killSwitchReadout(noPlane as Status)
assert.equal(k.state, 'unknown')
assert.notEqual(k.value, 'ARMED')
assert.equal(k.variant, 'off')
assert.notEqual(k.variant, 'on')
assert.equal(k.blockingNow, 'not known')
})
test('no status at all is unknown too, and says there is no reading', () => {
const k = killSwitchReadout(null)
assert.equal(k.state, 'unknown')
assert.equal(k.variant, 'off')
assert.equal(k.blockingNow, 'no reading yet')
})
test('fail-closed with a plane installed is armed', () => {
for (const plane of ['full', 'hold'] as const) {
const k = killSwitchReadout(status({ plane }))
assert.equal(k.state, 'armed')
assert.equal(k.value, 'ARMED')
assert.equal(k.variant, 'on')
assert.equal(k.blockingNow, null)
}
})
test('fail-closed with no plane is configured but blocking nothing', () => {
const k = killSwitchReadout(status({ plane: 'none', table: false }))
assert.equal(k.state, 'inert')
assert.equal(k.value, 'NOT IN EFFECT')
assert.equal(k.variant, 'crit')
assert.equal(k.hot, true)
})
test('fail-open is the operator’s choice — amber, and never a plane question', () => {
for (const plane of ['full', 'none', undefined] as const) {
const k = killSwitchReadout(status({ kill_switch: 'open', plane }))
assert.equal(k.state, 'open')
assert.equal(k.value, 'OPEN')
assert.equal(k.variant, 'amber')
}
})
test('the live kill_switch wins over the saved one; the saved one only fills a gap', () => {
const { kill_switch, ...noKill } = status()
void kill_switch
// Live says open, config says closed → live wins.
assert.equal(killSwitchReadout(status({ kill_switch: 'open' }), 'closed').state, 'open')
// Nothing live → fall back to the saved policy.
assert.equal(killSwitchReadout(noKill as Status, 'open').state, 'open')
assert.equal(killSwitchReadout(noKill as Status, 'closed').state, 'armed')
})
// --- killSwitchClosed: the daemon's spelling, not the panel's ---------------
//
// `Status.kill_switch` is `option kill_switch` echoed verbatim (apply.go:
// `s.KillSwitch = m.Globals.KillSwitch`), and the daemon decides with
//
// !strings.EqualFold(strings.TrimSpace(g.KillSwitch), "open")
//
// so case, surrounding whitespace and an empty value all read as CLOSED on the
// router. The panel compared `=== 'closed'`, which read every one of them as
// OPEN: an amber lamp over the word OPEN and "Nothing is meant to be blocked" on
// a router that blocks — and, through protectionState's `failClosed`, a
// plane-less router downgraded from crit to amber on the same misreading.
//
// These are the inputs a hand-edited /etc/config/shater actually produces.
test('closed is closed however it is spelled — case, padding, and empty', () => {
for (const raw of ['closed', 'Closed', 'CLOSED', ' closed ', '\tclosed\n', '', ' ']) {
assert.equal(killSwitchClosed(raw), true, `killSwitchClosed(${JSON.stringify(raw)})`)
}
// Absent is the same question with no answer, and it falls closed too.
assert.equal(killSwitchClosed(undefined), true)
assert.equal(killSwitchClosed(null), true)
// Anything unrecognised is NOT taken as permission to stop blocking.
assert.equal(killSwitchClosed('nonsense'), true)
})
test('only "open" is open — but every spelling of it is', () => {
for (const raw of ['open', 'Open', 'OPEN', ' open ', 'oPeN']) {
assert.equal(killSwitchClosed(raw), false, `killSwitchClosed(${JSON.stringify(raw)})`)
}
})
test('a router blocking under "Closed" is never drawn as OPEN', () => {
for (const raw of ['Closed', ' closed ', '']) {
const k = killSwitchReadout(status({ kill_switch: raw }))
assert.equal(k.state, 'armed', `state for ${JSON.stringify(raw)}`)
assert.equal(k.value, 'ARMED')
assert.equal(k.variant, 'on')
assert.notEqual(k.value, 'OPEN')
}
})
test('the saved policy is normalised too, not just the live one', () => {
const { kill_switch, ...noKill } = status()
void kill_switch
assert.equal(killSwitchReadout(noKill as Status, ' Closed ').state, 'armed')
assert.equal(killSwitchReadout(noKill as Status, 'OPEN').state, 'open')
})
test('"Closed" with no plane is the CRIT state, not the amber fail-open one', () => {
// The inverted-lamp case with teeth: the daemon is blocking-by-policy and has
// no plane installed, which is "not protected AND nothing is stopping it".
// Reading "Closed" as fail-open downgraded that from crit to amber and told
// the operator it was their own configured choice.
for (const raw of ['Closed', ' closed ', '']) {
const p = protectionState(status({ kill_switch: raw, plane: 'none' }))
assert.equal(p.variant, 'crit', `variant for ${JSON.stringify(raw)}`)
assert.equal(p.headline, 'Not protected — traffic is going out directly')
assert.equal(p.alarm, true)
}
// The genuinely fail-open router still gets the amber, deliberate-choice text.
const open = protectionState(status({ kill_switch: 'OPEN', plane: 'none' }))
assert.equal(open.variant, 'amber')
assert.equal(open.headline, 'Not protected — running direct')
})
// --- the holding plane does not promise "nothing" ---------------------------
//
// netplane/nft.go's RenderHoldNft states its own contract as "No client TRAFFIC
// reaches the WAN" and names the exception in the same paragraph: clients still
// reach the router's resolver and dnsmasq forwards those lookups to the ISP in
// the clear, deliberately, because blocking it would also cut the daemon's own
// name resolution. The banner used to round that up to "Nothing is being let
// out".
// --- an unreadable config is not "switched off" -----------------------------
//
// `enabled`, `kill_switch` and `panel_port` all come from the configuration, so
// when the daemon could not READ it they are zero values that mean nothing
// (Status.config_readable). The failure happens on a full /overlay or a `uci
// commit` caught half-written — which is precisely when the fail-closed plane has
// the LAN cut off on purpose — and the daemon then publishes `plane:"hold"` WITH
// `enabled:false`. Checking `!enabled` first turned that into "Turned off", amber,
// no alarm, "turn the service on in Settings": the owner of a house with no
// internet told they did it themselves, and pointed at a page reading the same
// unreadable file. The ORDER of the two checks is the fix, so these pin the order.
/** The proven field state: hold plane, placeholders, and no config reading. */
function unreadable(over: Partial<Status> = {}): Status {
return status({
config_readable: false,
config_error: 'uci show shater: exit status 1',
enabled: false,
kill_switch: '',
plane: 'hold',
engine_running: false,
active: false,
...over,
})
}
test('an unreadable config is a crit alarm, never "Turned off"', () => {
const p = protectionState(unreadable())
assert.equal(p.variant, 'crit')
assert.equal(p.alarm, true)
assert.notEqual(p.headline, 'Turned off')
assert.match(p.headline, /configuration/i)
})
test('it says not to switch anything off — the instinct that makes it worse', () => {
const d = protectionState(unreadable()).detail
assert.match(d, /don’t turn anything off|do not turn anything off/i)
assert.match(d, /fail-closed plane doing its job/i)
assert.doesNotMatch(d, /Turn the service on in Settings/)
})
test('it wins over `enabled` whatever that placeholder happens to say', () => {
// `enabled` is a zero value here; neither of its readings may reach a branch.
for (const enabled of [false, true]) {
const p = protectionState(unreadable({ enabled }))
assert.equal(p.variant, 'crit', `enabled=${enabled}`)
assert.notEqual(p.headline, 'Turned off')
}
})
test('an OLDER daemon that never sends the field is not alarmed at forever', () => {
// Absent ⇒ "no reading", which is NOT "the read failed". Such a daemon's
// `enabled` means what it says, so the ordinary branches must still run —
// reading the field as `!== true` would have alarmed on every one of them.
const { config_readable, ...noField } = unreadable({ enabled: false })
void config_readable
assert.equal(protectionState(noField as Status).headline, 'Turned off')
// And a healthy older daemon still reaches its ordinary green readout.
const running = withTraffic({ verdict: 'tunnel', tunnel_rules: 2 })
assert.equal(protectionState(running).headline, 'Protected')
assert.equal(protectionState(running).variant, 'on')
})
test('the kill-switch readout refuses to name a policy it could not read', () => {
// kill_switch is "" here — which normalises to CLOSED, and that is exactly the
// trap: an unreadable config printed ARMED, green, "setting: fail-closed" from a
// placeholder. The `setting` word carries the third answer so no caller has to
// re-derive a boolean that cannot hold it.
const k = killSwitchReadout(unreadable())
assert.equal(k.state, 'unknown')
assert.notEqual(k.value, 'ARMED')
assert.equal(k.variant, 'off')
assert.equal(k.setting, 'not known')
assert.equal(k.settingHot, true)
})
test('a readable config still names its policy, both ways', () => {
const closed = killSwitchReadout(status({ config_readable: true }))
assert.equal(closed.setting, 'fail-closed')
assert.equal(closed.settingHot, false)
const open = killSwitchReadout(status({ config_readable: true, kill_switch: 'open' }))
assert.equal(open.setting, 'fail-open')
assert.equal(open.settingHot, true)
})
test('a readable config still reaches the calm "Turned off"', () => {
const p = protectionState(status({ config_readable: true, enabled: false }))
assert.equal(p.headline, 'Turned off')
assert.equal(p.alarm, false)
})
test('the hold banner does not claim nothing leaves, and names DNS', () => {
const d = protectionState(status({ plane: 'hold', engine_running: false })).detail
assert.match(d, /DNS/)
assert.doesNotMatch(d, /^Nothing is being let out/)
})
+199 -2
View File
@@ -71,6 +71,165 @@ export function engineReadout(status: Status | null): { variant: LedVariant; wor
}
}
// ---------------------------------------------------------------------------
// Is the kill-switch actually blocking anything?
// ---------------------------------------------------------------------------
/**
* Four answers, and "unknown" is one of them.
*
* armed — configured fail-closed AND a data plane is installed to enforce it.
* inert — configured fail-closed, but there is no plane. Nothing is blocking.
* unknown — the daemon has not said how much plane is installed, so whether the
* setting is in force is not known. NEVER paint this green.
* open — configured fail-open. Nothing is meant to be blocked.
*/
export type KillSwitchState = 'armed' | 'inert' | 'unknown' | 'open'
/**
* IS THE KILL-SWITCH CLOSED? The daemon's rule, exactly:
*
* killSwitchClosed = !strings.EqualFold(strings.TrimSpace(g.KillSwitch), "open")
* (apply/apply.go)
*
* Three things it does that `=== 'closed'` did not, each of which had inverted a
* lamp. `Status.kill_switch` is the RAW UCI string, echoed with no normalisation
* (`s.KillSwitch = m.Globals.KillSwitch`), so all three inputs are reachable from
* a hand-edited /etc/config/shater:
*
* - CASE. `kill_switch 'Closed'` blocks on the router. `'Closed' === 'closed'`
* is false, so the panel drew OPEN, amber, "Nothing is meant to be blocked" —
* and `plane: 'none'` was downgraded from crit to amber on the same reading.
* - WHITESPACE. `' closed '` likewise.
* - THE DEFAULT SIDE. An empty value blocks on the router. `??` only falls
* through null/undefined, so `''` reached the comparison, failed it, and read
* as OPEN — the default landing on the side that cannot be recovered from by
* looking at the page.
*
* Everything in the panel that asks this question must ask it here. Two pages
* having their own spelling of the same comparison is how they came to disagree.
*/
export function killSwitchClosed(raw: string | null | undefined): boolean {
return (raw ?? '').trim().toLowerCase() !== 'open'
}
export interface KillSwitchReadout {
state: KillSwitchState
/** The word the module puts in its readout. */
value: string
variant: LedVariant
/** Is it blocking right now — the row under the readout. `null` ⇒ nothing to add. */
blockingNow: string | null
/** True when `blockingNow` is bad news and should be drawn hot. */
hot: boolean
/**
* The CONFIGURED policy, as a word for the "setting" row: 'fail-closed',
* 'fail-open', or 'not known'.
*
* It lives here rather than being re-derived at each call site because it was
* re-derived at each call site: both pages kept their own `killArmed ?
* 'fail-closed' : 'fail-open'` beside this readout, which cannot express the third
* answer — so a router whose configuration could not be READ printed a confident
* "fail-closed" under a lamp that already said NOT REPORTED.
*/
setting: string
/** True when `setting` should be drawn hot (fail-open, or not known). */
settingHot: boolean
}
/**
* THE UNKNOWN BRANCH IS THE WHOLE POINT. This used to be
*
* killArmed && status?.plane !== 'none'
*
* and `undefined !== 'none'` is true — so a daemon that had not reported `plane`
* at all, and a panel that had not yet received its first status, both landed in
* the "ARMED" branch under a green lamp. Every other unknown in this file is an
* unlit socket for exactly this reason (see engineReadout): the kill-switch is
* the last thing standing between the LAN and the plain WAN, and "I don't know
* whether it is installed" must never be dressed as "it is".
*
* `configured` is the SAVED policy from /api/config, used only while
* /api/status has not reported one. The live value wins wherever it exists, as
* everywhere else in the panel: this is a status readout, and the config on disk
* can already differ from what is installed.
*
* AN UNREADABLE CONFIGURATION IS ITS OWN ANSWER, and it comes first. `kill_switch`
* is one of the three fields sourced from the config, so when `config_readable` is
* false it is a zero value — and the empty string normalises to "closed", which is
* how a router nobody could read printed ARMED, green, "setting: fail-closed". The
* lamp said NOT REPORTED two lines further down in an earlier draft of this fix and
* the row still said fail-closed, which is the same defect twice.
*/
export function killSwitchReadout(
status: Status | null,
configured?: string,
): KillSwitchReadout {
if (status?.config_readable === false) {
return {
state: 'unknown',
value: 'NOT REPORTED',
variant: 'off',
// The plane may well be blocking (a hold plane is installed in exactly this
// situation) — what is unknown is the SETTING, so that is what this says.
blockingNow: 'setting can’t be read',
hot: false,
setting: 'not known',
settingHot: true,
}
}
const closed = killSwitchClosed(status?.kill_switch ?? configured)
const setting = closed
? { setting: 'fail-closed', settingHot: false }
: { setting: 'fail-open', settingHot: true }
if (!closed) {
return {
state: 'open',
value: 'OPEN',
variant: 'amber',
blockingNow: null,
hot: false,
...setting,
}
}
switch (status?.plane) {
case 'full':
case 'hold':
// Something is installed, so the fail-closed guard is really in the path.
return {
state: 'armed',
value: 'ARMED',
variant: 'on',
blockingNow: null,
hot: false,
...setting,
}
case 'none':
return {
state: 'inert',
value: 'NOT IN EFFECT',
variant: 'crit',
blockingNow: 'no — nothing installed',
hot: true,
...setting,
}
default:
return {
state: 'unknown',
value: 'NOT REPORTED',
variant: 'off',
// Terse on purpose: this is a two-column readout row, and the long form
// wrapped onto three lines beside a one-word key.
blockingNow: status ? 'not known' : 'no reading yet',
hot: false,
...setting,
}
}
}
/**
* `plane` + `engine_running` express the state more precisely than the three
* booleans the old status strip exposed (engine active / config enabled / nft
@@ -110,7 +269,36 @@ export function protectionState(status: Status | null): ProtectionState {
}
}
const failClosed = (status.kill_switch ?? 'closed') === 'closed'
// COULD THE CONFIG BE READ? This has to come FIRST, before `enabled` and before
// `kill_switch`, because both of those are sourced from the configuration and are
// zero values when it could not be read (Status.config_readable).
//
// The order was the bug, and it is not a cosmetic one. The daemon cannot read the
// config when /overlay is full or a `uci commit` was interrupted part-way — which
// is exactly when the fail-closed plane has the LAN cut off on purpose. It then
// publishes `plane:"hold"` together with `enabled:false`, and checking `!enabled`
// first turned that into "Turned off", amber, alarm:false, "turn the service on in
// Settings" — telling the owner of a house with no internet that they switched it
// off themselves, and sending them to a page backed by the same unreadable file.
//
// Read as `=== false`, never `!== true`: the field is positively phrased so that
// its ABSENCE (an older daemon) reads as "no reading" — but "no reading" is not
// the same as "the read failed", and only the latter earns this crit. An older
// daemon has an `enabled` that means what it says, so it must fall through to the
// branches below rather than be alarmed at forever.
if (status.config_readable === false) {
return {
variant: 'crit',
headline: 'The router’s configuration can’t be read',
// Carries the daemon's own load-bearing sentence: the instinct here is to
// switch things off, and that is the one action that makes it worse.
detail:
'This page can’t say whether shater is on, or whether the kill-switch is closed. If traffic is being blocked, that is the fail-closed plane doing its job — not the service being off, so don’t turn anything off to fix it. The usual causes are a full disk or an interrupted save; see the findings below.',
alarm: true,
}
}
const failClosed = killSwitchClosed(status.kill_switch)
// The service being switched off is a deliberate state, not a fault.
if (!status.enabled) {
@@ -126,11 +314,20 @@ export function protectionState(status: Status | null): ProtectionState {
case 'full':
return fullPlaneState(status.traffic)
case 'hold':
// "Nothing is being let out" was one word too wide, and the word was load-
// bearing. The holding plane's own contract (netplane/nft.go RenderHoldNft)
// says "No client TRAFFIC reaches the WAN" and then names the exception in
// the same breath: clients can still reach the router's resolver, and
// dnsmasq forwards those queries to the ISP in the clear. That is deliberate
// and cannot be closed — blocking it would also cut the daemon's own name
// resolution, and with it any chance of fetching what it choked on and
// recovering. A panel that rounds "no traffic" up to "nothing" is claiming a
// guarantee the plane below it never made.
return {
variant: 'amber',
headline: 'Traffic blocked — the tunnel is down',
detail:
'Nothing is being let out rather than let out unprotected. Devices can still reach each other and the router, so you can fix it from here.',
'No traffic from your devices is reaching the internet — it’s blocked rather than let out unprotected. Devices can still reach each other and the router, so you can fix it from here. DNS is the exception: lookups sent to the router still go out to your provider in the clear.',
alarm: true,
}
case 'none':
+72
View File
@@ -0,0 +1,72 @@
// carryRulesetFormat — the field the rule-set edit form cannot show and therefore
// must not drop.
//
// Run with `npm test` (node's built-in test runner + native type stripping).
// ruleset.ts has no runtime imports, so this runs against the real module.
//
// THE CASE THIS FILE WAS WRITTEN FOR is the boring one: open a rule-set that was
// given an explicit `Format` over SSH, change nothing but its NAME, press Save.
// The form rebuilt the object from its own controls, `Format` was on none of them,
// and `updateRuleset` swapped the element whole — so the value was gone, with no
// control anywhere in the panel able to put it back. What follows is silent: per
// generate/ruleset.go the list then matches nothing, the rule using it stops
// firing, and its traffic falls through to the next rule with nothing logged.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { carryRulesetFormat } from './ruleset.ts'
import type { Ruleset } from './api.ts'
/** A saved rule-set carrying a format the panel has no control for. */
function saved(over: Partial<Ruleset> = {}): Ruleset {
return { Name: 'blocked-ru', Type: 'domain', Source: 'file', Path: '/etc/shater/ru.lst', Format: 'binary', ...over }
}
/** What the edit form rebuilds from its own controls — no Format among them. */
function rebuilt(over: Partial<Ruleset> = {}): Ruleset {
return { Name: 'blocked-ru', Type: 'domain', Source: 'file', Path: '/etc/shater/ru.lst', ...over }
}
test('renaming a file rule-set keeps its explicit format', () => {
const out = carryRulesetFormat(rebuilt({ Name: 'ru-blocked' }), saved())
assert.equal(out.Format, 'binary')
assert.equal(out.Name, 'ru-blocked')
// Everything else the form rebuilt is untouched.
assert.equal(out.Path, '/etc/shater/ru.lst')
assert.equal(out.Source, 'file')
})
test('a url rule-set keeps it too — that is what stops a .srs being read as text', () => {
const out = carryRulesetFormat(
rebuilt({ Source: 'url', Path: undefined, URL: 'https://example.invalid/list' }),
saved({ Source: 'url', Format: 'binary' }),
)
assert.equal(out.Format, 'binary')
})
test('the sources the generator ignores it on do not get it, so it cannot raise a warning', () => {
for (const Source of ['inline', 'geosite', 'geoip']) {
const out = carryRulesetFormat(rebuilt({ Source, Path: undefined }), saved({ Source }))
assert.equal(out.Format, undefined, `Source=${Source}`)
}
})
test('nothing to carry is not something to invent', () => {
const next = rebuilt()
// No prior rule-set at all (the ADD path).
assert.equal(carryRulesetFormat(next, null).Format, undefined)
assert.equal(carryRulesetFormat(next, undefined).Format, undefined)
// A prior rule-set with an empty or blank format.
assert.equal(carryRulesetFormat(next, saved({ Format: '' })).Format, undefined)
assert.equal(carryRulesetFormat(next, saved({ Format: ' ' })).Format, undefined)
})
test('the inputs are not mutated — the caller keeps a usable `initial`', () => {
const next = rebuilt()
const prev = saved()
const out = carryRulesetFormat(next, prev)
assert.notEqual(out, next)
assert.equal(next.Format, undefined, 'the rebuilt literal must not be written through')
assert.equal(prev.Format, 'binary')
})
+52
View File
@@ -0,0 +1,52 @@
// Rule-set model rules the panel must not get wrong — the ones about fields the
// panel does not show.
//
// This module exists so they can be tested. The forms that apply them live in
// page components (Routing.tsx), which import CSS and React and therefore cannot
// be loaded by `npm test`; the imports here are type-only, exactly like
// planeState.ts, so the test runs against the real code with nothing stubbed.
import type { Ruleset } from './api'
/**
* The sources for which the GENERATOR reads `Ruleset.Format`.
*
* Positive and closed, not "everything except inline". generate/ruleset.go
* consults the field in exactly two branches and WARNS about it in the others
* ("format set on a source that has no use for it"), so a source that grows into
* the model later must be added here deliberately rather than inheriting a
* meaning nobody chose for it.
*/
const FORMAT_BEARING_SOURCES: ReadonlySet<string> = new Set(['url', 'file'])
/**
* Carry `Format` from the ruleset being edited onto the one being saved.
*
* WHY THIS IS NOT JUST A SPREAD. The edit form rebuilds a `Ruleset` literal per
* source branch, and rebuilding is right: switching a list from `url` to `file`
* must not drag the old `URL` and `UpdateInterval` along. But `Format` is not like
* those — the panel has NO CONTROL for it. It cannot be set here and cannot be
* restored here, so a save that dropped it destroyed a value only an SSH session
* could put back, and the trigger was as domestic as RENAMING the list (the rename
* path runs through the same rebuild, and the update replaces the element whole).
*
* WHAT IT COSTS WHEN IT GOES, per generate/ruleset.go:
*
* - `file` — an explicit format OVERRIDES the filename extension, and is the
* only way to load a rule-set whose name says nothing about its contents.
* - `url` — it is what stops a compiled binary `.srs`, served without a
* recognisable extension, from being parsed as a plain-text domain list.
*
* Either way the rule-set silently matches nothing afterwards. Nothing errors: the
* rule that uses it simply stops firing and its traffic falls through to whatever
* rule comes next, which is the failure mode that costs a day to find.
*
* Returns a NEW ruleset; the inputs are not modified. An empty/absent Format, or a
* source the generator ignores it on, yields `next` unchanged.
*/
export function carryRulesetFormat(next: Ruleset, initial: Ruleset | null | undefined): Ruleset {
const format = (initial?.Format ?? '').trim()
if (!format) return next
if (!FORMAT_BEARING_SOURCES.has((next.Source ?? '').trim().toLowerCase())) return next
return { ...next, Format: format }
}
+48 -1
View File
@@ -1,15 +1,62 @@
import { defineConfig } from 'vite'
import type { Plugin } from 'vite'
import react from '@vitejs/plugin-react'
// Minimal ambient for the dev-proxy target override — avoids pulling in @types/node
// just for one env read. Vite runs this file under Node where `process` exists.
declare const process: { env: Record<string, string | undefined> }
/** `src/mock.ts`, as the module graph spells it (POSIX-normalised for Windows). */
const MOCK_MODULE = 'src/mock.ts'
/**
* Refuse to emit a production bundle that contains the offline fixture backend.
*
* `src/mock.ts` describes an invented, healthy router: a full config, 122 nodes
* with 119 of them alive, "Protected". It exists so `npm run dev` renders without
* a daemon. It shipped inside the binary that goes on real hardware, switched on
* by nothing more than a `?dev` on the end of the URL — so a link someone was
* sent, or a bookmark they saved, showed an appliance in perfect health while
* making no request to the appliance at all.
*
* api.ts now loads it behind `import.meta.env.DEV`, which Vite folds to a literal
* `false` for a build, so Rollup drops the dynamic import and the module never
* enters the graph. That is a property of a build tool's optimiser, and an
* optimiser is not a promise: one refactor that makes the condition non-static
* silently puts the fixtures back. So the property is CHECKED rather than
* trusted — if `src/mock.ts` reaches any emitted chunk, the build fails here
* instead of shipping.
*/
function assertNoMockFixtures(): Plugin {
return {
name: 'shater:assert-no-mock-fixtures',
apply: 'build',
generateBundle(_options, bundle) {
const guilty: string[] = []
for (const [file, output] of Object.entries(bundle)) {
if (output.type !== 'chunk') continue
for (const id of output.moduleIds) {
if (id.replace(/\\/g, '/').endsWith(MOCK_MODULE)) guilty.push(`${file} ← ${id}`)
}
}
if (guilty.length > 0) {
this.error(
`the offline fixture backend (${MOCK_MODULE}) reached the production bundle:\n ` +
guilty.join('\n ') +
`\nFixtures describe a router that does not exist. Keep every path to them behind ` +
`\`import.meta.env.DEV\` so Rollup can drop them, and never gate them on a runtime ` +
`flag such as a query parameter.`,
)
}
},
}
}
// The SPA is embedded in the forked sing-box binary and served by the daemon on
// its own port. Relative base so it works under any mount path; single small
// bundle (no code-splitting) keeps the embed simple and the flash budget low.
export default defineConfig({
plugins: [react()],
plugins: [react(), assertNoMockFixtures()],
base: './',
build: {
outDir: 'dist',
+268
View File
@@ -0,0 +1,268 @@
// lx:begin l3-honest-drop
package route
import (
"context"
"net/netip"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
R "github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/stretchr/testify/require"
)
// The contract under test: PreMatch never answers "continue" (nor "bypass") for
// an ICMP flow. adapter.JudgeFlow maps both to tun.ActionAccept, and the TUN
// stack answers Accept by FORGING the echo reply itself
// (sing-tun stack_gvisor_icmp.go ICMPForwarder.HandlePacket, the fallthrough
// under the Flow/Reject/Drop switch). A verdict of "continue" therefore reads to
// the operator as a working ping off a tunnel that never carried the packet.
//
// Every test below has a TCP/UDP twin: the honest drop must not leak into the
// protocols where "continue" really does mean "take the ordinary connection
// route".
// icmpL4Outbound is a minimal L4-only outbound (the vless/vmess/... shape): it
// does NOT implement adapter.FlowOutbound, and Network() lists only TCP/UDP.
// Unused Outbound methods come from the embedded nil interface and are never
// called on the pre-match paths under test.
type icmpL4Outbound struct {
adapter.Outbound
tag string
}
func (o *icmpL4Outbound) Tag() string { return o.tag }
func (o *icmpL4Outbound) Type() string { return "vless" }
func (o *icmpL4Outbound) Network() []string { return []string{N.NetworkTCP, N.NetworkUDP} }
// icmpOutboundManager resolves tags from a fixed map and hands the same L4-only
// outbound out as the default; the rest of the OutboundManager surface is never
// touched by the pre-match walk.
type icmpOutboundManager struct {
adapter.OutboundManager
defaultOutbound adapter.Outbound
outbounds map[string]adapter.Outbound
}
func (m *icmpOutboundManager) Default() adapter.Outbound { return m.defaultOutbound }
func (m *icmpOutboundManager) Outbound(tag string) (adapter.Outbound, bool) {
outbound, loaded := m.outbounds[tag]
return outbound, loaded
}
// icmpDNSRouter / icmpDNSTransportManager implement only what
// prepareMatchMetadata reaches. FakeIP returns nil unless a transport is
// installed, which is how the "fakeip lookup failed" exit is driven below.
type icmpDNSRouter struct {
adapter.DNSRouter
}
func (s *icmpDNSRouter) LookupReverseMapping(netip.Addr) (string, bool) { return "", false }
type icmpDNSTransportManager struct {
adapter.DNSTransportManager
fakeIP adapter.FakeIPTransport
}
func (s *icmpDNSTransportManager) FakeIP() adapter.FakeIPTransport {
if s.fakeIP == nil {
return nil
}
return s.fakeIP
}
// icmpMissingFakeIPTransport claims every address and then fails to look any of
// them up — exactly the "missing fakeip record, try enable
// `experimental.cache_file`" error prepareMatchMetadata returns.
type icmpMissingFakeIPTransport struct {
adapter.FakeIPTransport
}
func (t *icmpMissingFakeIPTransport) Store() adapter.FakeIPStore {
return &icmpMissingFakeIPStore{}
}
type icmpMissingFakeIPStore struct {
adapter.FakeIPStore
}
func (s *icmpMissingFakeIPStore) Contains(netip.Addr) bool { return true }
func (s *icmpMissingFakeIPStore) Lookup(netip.Addr) (string, bool) { return "", false }
type icmpRouterOptions struct {
fakeIP adapter.FakeIPTransport
rules []option.Rule
}
func icmpTestRouter(t *testing.T, options icmpRouterOptions) *Router {
t.Helper()
logger := log.NewNOPFactory().NewLogger("test")
defaultOutbound := &icmpL4Outbound{tag: "proxy-out"}
router := &Router{
ctx: context.Background(),
logger: logger,
dns: &icmpDNSRouter{},
dnsTransport: &icmpDNSTransportManager{fakeIP: options.fakeIP},
outbound: &icmpOutboundManager{
defaultOutbound: defaultOutbound,
outbounds: map[string]adapter.Outbound{defaultOutbound.Tag(): defaultOutbound},
},
}
for i, ruleOptions := range options.rules {
rule, err := R.NewRule(router.ctx, logger, ruleOptions, false)
require.NoError(t, err, "build rule[%d]", i)
router.rules = append(router.rules, rule)
}
return router
}
func icmpTestMetadata(network string) adapter.InboundContext {
return adapter.InboundContext{
Inbound: "l3-in",
InboundType: C.TypeTun,
Network: network,
Source: M.SocksaddrFrom(netip.MustParseAddr("192.168.1.2"), 0),
Destination: M.SocksaddrFrom(netip.MustParseAddr("1.1.1.1"), 0),
}
}
// lanRuleWithAction matches every packet from the test source, so the action is
// what the test is actually about.
func lanRuleWithAction(action option.RuleAction) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceIPCIDR: badoption.Listable[string]{"192.168.1.0/24"},
},
RuleAction: action,
},
}
}
// --- exit 1: an outbound that cannot carry layer 3 --------------------------
func TestPreMatchICMPToL4OutboundDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"ICMP to an L4-only outbound fell through to the ordinary pre-match path: the TUN stack will forge the echo reply and ping will lie about a tunnel that never saw the packet")
}
func TestPreMatchTCPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"TCP to an L4-only outbound must keep taking the ordinary connection route; the ICMP honest-drop must not leak into TCP/UDP pre-match")
}
func TestPreMatchUDPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkUDP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"UDP to an L4-only outbound must keep taking the ordinary connection route")
}
// --- exit 2: prepareMatchMetadata failed before any rule was walked ---------
// This exit arrived with the shared prepareMatchMetadata refactor (upstream
// b911fb078): it returns before the rule walk, so it never reaches preMatchFlow
// where the ICMP override used to live.
func TestPreMatchICMPMetadataErrorDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"a fakeip record that cannot be resolved must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPMetadataErrorContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"for TCP the metadata-error exit must keep meaning `take the ordinary connection route`")
}
// --- exit 3: a rule action the pre-match walk does not handle ---------------
// hijack-dns is one of the actions PreMatch's switch has no arm for, so it lands
// in the default arm. Any future unhandled action lands there too — that is why
// the guard is a funnel on the return value and not a per-arm override.
func TestPreMatchICMPUnhandledRuleActionDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"an unhandled rule action must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPUnhandledRuleActionContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"the unhandled-action exit must stay a continue for TCP")
}
// --- exit 4: an explicit bypass ---------------------------------------------
// sing-tun implements ActionBypass on the nfqueue plane only; on the TUN path it
// falls into the same default arm as Accept (flow_dispatch.go judgeAndInstall,
// and the ICMP forwarder's switch has no Bypass case either), i.e. into the same
// forgery. There is no honest bypass for a packet already inside the engine's
// TUN.
func TestPreMatchICMPBypassDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"bypass degrades to tun.ActionAccept on the TUN path, which is the forged echo reply again")
}
func TestPreMatchTCPBypassIsStillBypass(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchBypass, result.Action,
"the ICMP honest-drop must not turn a TCP bypass rule into a drop")
}
// --- the verdicts that must pass through untouched ---------------------------
// A reject rule already carries its own honest verdict; the funnel must not
// rewrite it (a Reject sends an ICMP unreachable, which is information, not a
// forged liveness signal).
func TestPreMatchICMPRejectIsNotRewritten(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"the ICMP funnel must only rewrite continue/bypass, never an explicit reject")
}
// lx:end l3-honest-drop
+194
View File
@@ -0,0 +1,194 @@
package route
import (
"context"
"net"
"net/netip"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
R "github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/stretchr/testify/require"
)
// The pre-match path (adapter.JudgeFlow -> Router.PreMatch, used by the TUN/
// WireGuard-endpoint flow dispatcher) used to prepare only fakeip and the IP
// version. Everything a rule matches on that the inbound cannot know — the
// connection owner and the neighbor behind the source address — was resolved in
// matchRule only, so a `source_mac_address` / `source_hostname` rule silently
// failed to match in pre-match and the flow fell through to the default
// outbound. Upstream b911fb078 shares one prepareMatchMetadata between both
// paths; these tests pin that.
// stubNeighborResolver answers for exactly one address.
type stubNeighborResolver struct {
address netip.Addr
mac net.HardwareAddr
hostname string
}
func (r *stubNeighborResolver) LookupMAC(address netip.Addr) (net.HardwareAddr, bool) {
if address != r.address || r.mac == nil {
return nil, false
}
return r.mac, true
}
func (r *stubNeighborResolver) LookupHostname(address netip.Addr) (string, bool) {
if address != r.address || r.hostname == "" {
return "", false
}
return r.hostname, true
}
func (r *stubNeighborResolver) LookupAddresses(hostname string) []netip.Addr {
if hostname != r.hostname {
return nil
}
return []netip.Addr{r.address}
}
func (r *stubNeighborResolver) Start() error { return nil }
func (r *stubNeighborResolver) Close() error { return nil }
// stubDNSRouter / stubDNSTransportManager implement only what
// prepareMatchMetadata reaches; every other method is left to the embedded nil
// interface and would panic if it were ever called.
type stubDNSRouter struct {
adapter.DNSRouter
}
func (s *stubDNSRouter) LookupReverseMapping(netip.Addr) (string, bool) { return "", false }
type stubDNSTransportManager struct {
adapter.DNSTransportManager
}
func (s *stubDNSTransportManager) FakeIP() adapter.FakeIPTransport { return nil }
// stubOutboundManager's default outbound supports no network at all, so a flow
// that reaches preMatchFlow bails out with PreMatchContinue instead of nil-
// dereferencing. That is exactly the pre-fix verdict we assert against.
type stubOutboundManager struct {
adapter.OutboundManager
defaultOutbound adapter.Outbound
}
func (s *stubOutboundManager) Default() adapter.Outbound { return s.defaultOutbound }
type stubNoNetworkOutbound struct {
adapter.Outbound
}
func (o *stubNoNetworkOutbound) Tag() string { return "stub" }
func (o *stubNoNetworkOutbound) Type() string { return "direct" }
func (o *stubNoNetworkOutbound) Network() []string { return nil }
func newPreMatchTestRouter(t *testing.T, resolver adapter.NeighborResolver, rules ...option.Rule) *Router {
t.Helper()
logger := log.NewNOPFactory().NewLogger("test")
router := &Router{
ctx: context.Background(),
logger: logger,
dns: &stubDNSRouter{},
dnsTransport: &stubDNSTransportManager{},
outbound: &stubOutboundManager{defaultOutbound: &stubNoNetworkOutbound{}},
neighborResolver: resolver,
needFindNeighbor: true,
}
for i, ruleOptions := range rules {
rule, err := R.NewRule(router.ctx, logger, ruleOptions, false)
require.NoError(t, err, "build rule[%d]", i)
router.rules = append(router.rules, rule)
}
return router
}
func rejectOnSourceMAC(macAddress string) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceMACAddress: badoption.Listable[string]{macAddress},
},
RuleAction: option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
},
},
}
}
func rejectOnSourceHostname(hostname string) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceHostname: badoption.Listable[string]{hostname},
},
RuleAction: option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
},
},
}
}
func preMatchMetadata() adapter.InboundContext {
return adapter.InboundContext{
Inbound: "tun-in",
InboundType: C.TypeTun,
Network: N.NetworkUDP,
Source: M.ParseSocksaddr("192.168.1.5:41234"),
Destination: M.ParseSocksaddr("1.1.1.1:443"),
}
}
func TestPreMatchResolvesNeighborMAC(t *testing.T) {
t.Parallel()
mac, err := net.ParseMAC("de:ad:be:ef:00:01")
require.NoError(t, err)
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.5"),
mac: mac,
hostname: "kitchen-tv",
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceMAC("de:ad:be:ef:00:01"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"source_mac_address rule must match in pre-match; the MAC has to be resolved there too")
}
func TestPreMatchResolvesNeighborHostname(t *testing.T) {
t.Parallel()
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.5"),
hostname: "kitchen-tv",
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceHostname("kitchen-tv"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"source_hostname rule must match in pre-match; the hostname has to be resolved there too")
}
// A source the neighbor resolver does not know must still fall through, not
// match on a half-filled metadata.
func TestPreMatchNeighborMissDoesNotMatch(t *testing.T) {
t.Parallel()
mac, err := net.ParseMAC("de:ad:be:ef:00:01")
require.NoError(t, err)
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.9"),
mac: mac,
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceMAC("de:ad:be:ef:00:01"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action)
}
+74 -25
View File
@@ -314,27 +314,60 @@ func (r *Router) routePacketConnection(ctx context.Context, conn N.PacketConn, m
return nil
}
// lx:begin l3-honest-drop
// PreMatch funnels every verdict of the pre-match walk through one ICMP check.
//
// An ICMP flow has no fallback path, so PreMatchContinue is not "try the
// ordinary connection route" the way it is for TCP and UDP: the TUN stack takes
// the packet back and answers the echo ITSELF (sing-tun stack_gvisor_icmp.go —
// adapter.JudgeFlow maps Continue to tun.ActionAccept, and the ICMP forwarder
// answers Accept by rewriting Echo into EchoReply and swapping the addresses).
// A ping routed to an outbound that cannot carry layer 3 — every proxy
// protocol; only adapter.FlowOutbound can — would therefore return a FORGED
// reply, and the operator would read a working ping off a tunnel that never saw
// the packet. Dropping instead reports the truth.
//
// PreMatchBypass is folded into the same drop because sing-tun implements
// bypass for the nfqueue plane only (`ActionBypass` appears nowhere in
// flow_dispatch.go / stack_gvisor_icmp.go): on the TUN path it degrades to the
// same Accept, i.e. to the same forgery. There is no honest bypass for an ICMP
// packet that is already inside the engine's TUN.
//
// This is a funnel and not an override inside the walk on purpose: the walk has
// several independent exits that say "continue" (the prepareMatchMetadata error
// return, the sniff bail-outs, the un-routable `bypass`, and the default arm of
// the rule-action switch), and an earlier version of this delta guarded only
// the ones that pass through preMatchFlow — leaving the others as narrow paths
// to the forged reply. Guarding the single return value cannot be outgrown by a
// new exit.
func (r *Router) PreMatch(metadata adapter.InboundContext, firstPacket []byte) adapter.PreMatchResult {
result := r.preMatch(metadata, firstPacket)
if metadata.Network == N.NetworkICMP {
switch result.Action {
case adapter.PreMatchContinue, adapter.PreMatchBypass:
return adapter.PreMatchResult{Action: adapter.PreMatchDrop}
}
}
return result
}
// preMatch is upstream's PreMatch body, unchanged; only the name moved, so that
// the funnel above owns the exported entry point. An upstream change to the
// pre-match walk applies to THIS function.
func (r *Router) preMatch(metadata adapter.InboundContext, firstPacket []byte) adapter.PreMatchResult {
// lx:end l3-honest-drop
ctx := log.ContextWithNewID(r.ctx)
metadata.PreMatch = true
continueResult := adapter.PreMatchResult{Action: adapter.PreMatchContinue}
packetDestination := metadata.Destination
if metadata.Destination.Addr.IsValid() && r.dnsTransport.FakeIP() != nil && r.dnsTransport.FakeIP().Store().Contains(metadata.Destination.Addr) {
domain, loaded := r.dnsTransport.FakeIP().Store().Lookup(metadata.Destination.Addr)
if !loaded || domain == "" {
return continueResult
}
metadata.OriginDestination = metadata.Destination
metadata.Destination = M.Socksaddr{
Fqdn: domain,
Port: metadata.Destination.Port,
}
metadata.FakeIP = true
}
if metadata.Destination.IsIPv4() {
metadata.IPVersion = 4
} else if metadata.Destination.IsIPv6() {
metadata.IPVersion = 6
// lx: pre-match used to prepare only fakeip + IP version, so process/neighbor
// rule items (process_name, source_mac_address, source_hostname, …) never had
// their metadata filled here and silently failed to match — they were resolved
// in matchRule only. Both paths now share prepareMatchMetadata (upstream
// b911fb078).
err := r.prepareMatchMetadata(ctx, &metadata)
if err != nil {
return continueResult
}
for currentRuleIndex, currentRule := range r.rules {
metadata.ResetRuleCache()
@@ -448,6 +481,11 @@ func applyRouteOptionsOverride(metadata *adapter.InboundContext, routeOptions *R
func (r *Router) preMatchFlow(ctx context.Context, metadata *adapter.InboundContext, packetDestination M.Socksaddr, matchedRule adapter.Rule, outboundTag string) adapter.PreMatchResult {
continueResult := adapter.PreMatchResult{Action: adapter.PreMatchContinue}
// lx: ICMP does NOT get a local override here any more — the honest drop is
// applied once, to the single return value of PreMatch (see the funnel
// there, marker l3-honest-drop). Overriding continueResult in this function
// covered only the exits that reach it and left the walk's own exits
// forging.
var outbound adapter.Outbound
if outboundTag == "" {
outbound = r.outbound.Default()
@@ -540,13 +578,11 @@ func (r *Router) preMatchFlow(ctx context.Context, metadata *adapter.InboundCont
return result
}
func (r *Router) matchRule(
ctx context.Context, metadata *adapter.InboundContext,
inputConn net.Conn, inputPacketConn N.PacketConn,
) (
selectedRule adapter.Rule, selectedRuleIndex int,
buffers []*buf.Buffer, packetBuffers []*N.PacketBuffer, fatalErr error,
) {
// prepareMatchMetadata fills in everything a rule may match on but the inbound
// cannot know: the connection owner, the neighbor (MAC/hostname) behind the
// source address, the fakeip / reverse-mapped domain and the IP version. Shared
// by matchRule and PreMatch — see the note at the PreMatch call site.
func (r *Router) prepareMatchMetadata(ctx context.Context, metadata *adapter.InboundContext) error {
r.searchProcessInfo(ctx, metadata)
if r.neighborResolver != nil && metadata.SourceMACAddress == nil && metadata.Source.Addr.IsValid() {
mac, macFound := r.neighborResolver.LookupMAC(metadata.Source.Addr)
@@ -568,8 +604,7 @@ func (r *Router) matchRule(
if metadata.Destination.Addr.IsValid() && r.dnsTransport.FakeIP() != nil && r.dnsTransport.FakeIP().Store().Contains(metadata.Destination.Addr) {
domain, loaded := r.dnsTransport.FakeIP().Store().Lookup(metadata.Destination.Addr)
if !loaded {
fatalErr = E.New("missing fakeip record, try enable `experimental.cache_file`")
return
return E.New("missing fakeip record, try enable `experimental.cache_file`")
}
if domain != "" {
metadata.OriginDestination = metadata.Destination
@@ -592,6 +627,20 @@ func (r *Router) matchRule(
} else if metadata.Destination.IsIPv6() {
metadata.IPVersion = 6
}
return nil
}
func (r *Router) matchRule(
ctx context.Context, metadata *adapter.InboundContext,
inputConn net.Conn, inputPacketConn N.PacketConn,
) (
selectedRule adapter.Rule, selectedRuleIndex int,
buffers []*buf.Buffer, packetBuffers []*N.PacketBuffer, fatalErr error,
) {
fatalErr = r.prepareMatchMetadata(ctx, metadata)
if fatalErr != nil {
return
}
match:
for currentRuleIndex, currentRule := range r.rules {
+549 -30
View File
@@ -22,11 +22,34 @@
# 2. It runs on linux. shater/generate has 44 test files on linux against 32 on
# windows/darwin; the linux-only half is where the routing, ruleset, DNS and
# health tests live.
# 3. Nothing is skipped SILENTLY. Two machine checks:
# 3. Nothing is skipped SILENTLY. Six machine checks:
# - the tag set may only ADD test files, never hide them (a test behind
# `//go:build !with_awg` would vanish from the gate — this fails first);
# - every package that has tests must report `ok` by name; a suite that
# compiles down to "no test files" fails the gate instead of passing it.
# compiles down to "no test files" fails the gate instead of passing it;
# - every ORDINARY test that calls t.Skip is named in the output and must
# be DECLARED in SKIP_DECLARED below with the reason it cannot run here;
# an undeclared skip fails the gate. This is why the suites run with -v:
# without it a skipped test prints nothing whatsoever and the package
# still reports `ok`. It was not a hypothetical — shater/apply's
# TestApplyInstallsHoldWhenEngineFailsToStart, the W5 regression for
# "the engine died, the LAN must not be left open", guarded itself with
# a t.Skip whose condition had become permanently true, so it asserted
# nothing at all while the gate reported `ok shater/apply`;
# - every ^TestIntegration under the fork's trees must produce a verdict
# BY NAME ([5/7]). `ok <pkg>` is printed whether the privileged tests in
# that package ran or called t.Skip, so the second check cannot see them
# — and the gate would keep saying "passes every test we own" while the
# tests that need a real kernel never executed;
# - every NON-GO test file in the tree must be claimed by a named runner
# ([6/7]). The four checks above are all built on `go list`/`go test`, so
# a test in another language is invisible to them BY CONSTRUCTION — and
# that is not hypothetical either: openwrt/luci-app-shater/tests/
# status-readout.test.js, 24 assertions over the one screen an operator
# reaches while the LAN is cut off, was run by nothing at all;
# - the non-Go suites this gate owns produce a verdict BY NAME ([7/7]),
# including "did not run: no node here", which then replaces the closing
# banner.
# A guard that silently runs nothing is worse than no guard (same rule as
# scripts/check-router-tags.sh).
#
@@ -38,6 +61,10 @@
# SHATER_GO_IMAGE docker image used to reach linux from a non-linux host
# (default golang:1.26 — keep it >= go.mod's toolchain).
# SHATER_NO_DOCKER=1 fail instead of falling back to docker.
# SHATER_REQUIRE_PRIVILEGED=1
# turn [5/7]'s "did not run here" report into a hard
# failure. Use it on the OpenWrt VM or in any pre-release
# run that must actually have exercised the kernel paths.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
@@ -48,7 +75,7 @@ RACE=1
for a in "$@"; do
case "$a" in
--no-race) RACE=0 ;;
-h|--help) sed -n '2,41p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
-h|--help) sed -n '2,67p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) echo "run-tests: unknown flag: $a" >&2; exit 2 ;;
esac
done
@@ -72,6 +99,12 @@ ROOTS_COMMON=(./common/...)
# rest of common/ can be a real gate instead of a permanently red one. (On
# linux every tlsspoof test is a TestIntegration*, so that package is
# effectively uncovered here; it is covered by the VM runs.)
# The ^TestIntegration prefix is the fork-wide marker for "needs capabilities
# the ordinary gate lacks", and [5/7] below leans on the same convention to
# catch privileged tests inside ROOTS, which are NOT name-filtered and would
# otherwise skip behind a green `ok <pkg>`. ROOTS_COMMON stays out of [5/7]:
# these fail rather than skip without the capability, and that is a decision
# about upstream code, not about the fork's own coverage.
SKIP_COMMON='^TestIntegration'
# SKIP, WITH REASON: the first -race run over this tree (2026-07-26 — nobody
@@ -80,15 +113,133 @@ SKIP_COMMON='^TestIntegration'
# RoutineReceiveIncoming() with no lock, caught by
# transport/wireguard.TestAwgDetourClientBindDelivers; that one has since been
# fixed in client_bind.go and is NOT skipped — it is exactly what this pass is
# for. What is left:
# - shater/alert.TestExpiryDedupWithinDay — the test's own closure
# (expiry_test.go:78) reads a variable the test body writes at :85 while
# Notifier.dispatch's goroutine is still delivering. A test-side bug, ~one
# mutex to fix, but it lives in shater/ and is nobody's blocker to ship.
# Naming it here keeps the gate a gate from day one. It is skipped ONLY in the
# -race pass — it still runs, and still has to pass, in the main pass below.
# DELETE THE ENTRY THE MOMENT THE RACE IS FIXED.
RACE_SKIP='^TestExpiryDedupWithinDay$'
# for. What was left:
# - shater/alert.TestExpiryDedupWithinDay — the test's own closure read a
# variable the test body wrote while Notifier.dispatch's goroutine was
# still delivering. FIXED 2026-07-26 (the simulated clock now has a mutex),
# so the entry is gone and the -race pass covers the whole tree again.
# Nothing is skipped under -race any more. Keep it that way: an entry here is a
# hole in the gate, so add one only with a named reason and delete it the moment
# the race is fixed.
RACE_SKIP='^$'
# --- the ONLY skips this gate accepts ----------------------------------------
# A t.Skip is invisible to every other check here: the test binary exits 0, the
# package prints `ok <pkg>`, and the name of the test that did not run appears
# NOWHERE unless -v is on. That is how shater/apply's W5 regression —
# TestApplyInstallsHoldWhenEngineFailsToStart, the only END-TO-END test between
# "the engine died" and "the LAN forwards to the WAN in the clear" — came to
# assert nothing at all: it broke the engine by pointing a rule-set at
# /nonexistent/nope.srs and stood itself down with t.Skip when that failed to
# break anything, and it stopped breaking anything once LocalRuleSet.reloadFile
# began treating an unreadable file as an empty one. Measured 2026-07-26 in
# golang:1.26: the skip fired unconditionally, and the package still printed
# `ok shater/apply`.
#
# So the suites below run with -v and every `--- SKIP` is matched against this
# list. A skip that is not here fails the gate BY NAME. Skips that genuinely
# cannot run in some environment are not forbidden — they are DECLARED, with the
# reason, and printed on every run so nobody mistakes the gate's silence for
# coverage.
#
# Format: '<extended regexp matched against the full test name>|<reason>'.
# The reason is shown verbatim next to the test on every run; write it for
# someone deciding whether the gate proved what they think it proved.
SKIP_DECLARED=(
'^TestIntegration|privileged: needs root + CAP_NET_ADMIN + /dev/net/tun. NOT waved through — [5/7] below gives every one of these a verdict by name, and an unrunnable one REPLACES the closing banner so this run cannot claim it covered them.'
'^TestCompiledTagsMatchTheShippedSet$|compares the tags COMPILED INTO a binary with router-tags.sh, and needs the harness that builds that binary and sets SHATER_ROUTER_TAG_CHECK=1. The harness is scripts/check-router-tags.sh, which the release tract runs separately; here there is no such binary to read.'
)
# --- every NON-GO test file must be claimed by a runner ----------------------
# The Go half of this gate cannot see a test written in another language, and the
# review of 2026-07-26 found what that costs: openwrt/luci-app-shater/tests/
# status-readout.test.js — 236 lines, six recorded fixtures, 24 assertions, the
# only thing checking what the LuCI dashboard tells an operator while the engine
# is down — was executed by NOTHING. Not by panel/package.json's `test` script
# (`node --test src/*.test.ts`, panel/src only), not by scripts/run-panel-tests.sh
# (same glob), not by any step here. And [1/7] could not report it, because
# `go list` is the instrument and a .js file is invisible to it BY CONSTRUCTION.
#
# So [6/7] enumerates the tree's non-Go test files and requires each to be claimed
# by a runner named HERE. A new .test.js/.test.ts/.spec.*/test_*.py that no runner
# picks up fails the gate by name on the day it is committed, instead of sitting
# there looking like coverage.
#
# Format: '<extended regexp matched against the repo-relative path>|<runner>'.
# Positive and CLOSED on purpose: a file that matches nothing is a failure, not a
# default. The Go files are deliberately NOT in scope — `_test.go` under the
# declared ROOTS is what [1/7]+[2/7] already prove ran, and pulling the rest of
# the upstream tree in here would be a different decision.
NONGO_TEST_RUNNERS=(
'^panel/src/[^/]+\.test\.ts$|scripts/run-panel-tests.sh (node --test via panel/package.json); CI runs it as its own step on node 24, before this script'
'^openwrt/luci-app-shater/tests/[^/]+\.test\.js$|[7/7] of this script'
)
# The non-Go suites [7/7] RUNS, as `node <file>`. panel/src is not here: it has its
# own script with its own npm install, and duplicating it would mean two places to
# keep right. Each entry is a glob; a glob that matches nothing is a failure (a
# renamed file must be reported as that, not as a fast green step).
JS_SUITES=('openwrt/luci-app-shater/tests/*.test.js')
# JS_UNVERIFIED collects the suites that did NOT run, by path, for the closing
# banner — the same treatment PRIV_UNVERIFIED gets, and for the same reason.
JS_UNVERIFIED=""
# js_step runs the declared non-Go suites and prints a verdict for each BY NAME.
# Defined up here because it is called from TWO places: the non-linux re-exec
# below runs it on the HOST, where node usually exists, rather than let the
# golang image (which has none) report "did not run" on every single local run —
# a banner that always fires is a banner nobody reads.
#
# `node <file>`: these are standalone harnesses that exit non-zero on a failed
# assertion, not `node --test` modules.
#
# KNOWN LIMIT, stated rather than papered over: the verdict is the process exit
# code. A harness gutted of its assertions that still exits 0 reads as a pass —
# the same limit scripts/run-panel-tests.sh already names for `node --test`.
# Deletion, rename, a throw and a failed assertion are all caught.
#
# Returns 1 if a suite failed or the globs matched nothing. Sets JS_UNVERIFIED
# when there is no node to run them with.
js_step() { # $1 = where we are, in words, for the "no node" line
local where="$1" f g rc bad=0 tmp
local files=()
shopt -s nullglob
for g in "${JS_SUITES[@]}"; do
files+=($g)
done
shopt -u nullglob
if [ "${#files[@]}" -eq 0 ]; then
echo " FAILED [js]: JS_SUITES matched no file at all. Either the glob is wrong or" >&2
echo " the suite was renamed/deleted — both must be said, not passed over." >&2
return 1
fi
if ! command -v node >/dev/null 2>&1; then
echo " node: NOT AVAILABLE $where — these suites did NOT run:"
for f in "${files[@]}"; do
echo " DID NOT RUN $f"
JS_UNVERIFIED="$JS_UNVERIFIED $f"
done
return 0
fi
echo " node: $(node --version) $where, ${#files[@]} suite(s)"
tmp="$(mktemp)"
for f in "${files[@]}"; do
set +e
node "$f" >"$tmp" 2>&1
rc=$?
set -e
if [ "$rc" -eq 0 ]; then
echo " RAN $f"
else
echo " FAILED $f (exit $rc)" >&2
sed 's/^/ | /' "$tmp" >&2
bad=1
fi
done
rm -f "$tmp"
return "$bad"
}
echo "== shater test gate =="
echo " tags : $SHATER_ROUTER_TAGS"
@@ -111,14 +262,67 @@ if [ "$(go env GOOS)" != "linux" ] && [ "${SHATER_TESTS_IN_DOCKER:-0}" != "1" ];
echo "== re-exec on linux via docker ($image) =="
host_repo="$REPO"
command -v cygpath >/dev/null 2>&1 && host_repo="$(cygpath -w "$REPO")"
# Hand the container CAP_NET_ADMIN and /dev/net/tun when this host's docker
# can. shater/generate's ^TestIntegration tests open a real TUN and stand a
# real engine on it; without the device they skip, and a dev running the gate
# by hand would get a green result that never touched the kernel path the
# branch is about. The dev host CAN give them (Docker Desktop's VM has the tun
# module) — the CI runner cannot, which is what [5/7] exists to say out loud.
# PROBED, never assumed: a docker whose kernel lacks tun refuses --device and
# would take the whole gate down with it.
priv_flags=()
if MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
--cap-add NET_ADMIN --device /dev/net/tun "$image" true >/dev/null 2>&1; then
priv_flags=(--cap-add NET_ADMIN --device /dev/net/tun)
echo " CAP_NET_ADMIN + /dev/net/tun: available — the privileged tests will really run"
else
echo " CAP_NET_ADMIN + /dev/net/tun: NOT available from this docker — [5/7] will report the gap"
fi
# The non-Go suites do not need linux, and this host very likely has node while
# the golang image certainly does not. Run them HERE, so the local loop really
# executes them instead of being told every single time that it did not: a
# banner that always fires is a banner nobody reads, and that is how a report
# stops being a report. Their verdict is folded into this script's exit status
# below, and the container is told not to repeat them.
host_js_rc=0
js_flags=()
if command -v node >/dev/null 2>&1; then
echo "== [7/7] the non-Go suites (node), run on this host before the re-exec =="
js_step "on this host ($image has none)" || host_js_rc=1
js_flags=(-e SHATER_JS_ALREADY_RAN=1)
echo
fi
MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
"${priv_flags[@]+"${priv_flags[@]}"}" \
"${js_flags[@]+"${js_flags[@]}"}" \
-v "$host_repo":/src \
-v shater-tagcheck-gomod:/go/pkg/mod \
-v shater-tagcheck-gocache:/root/.cache/go-build \
-w /src \
-e SHATER_TESTS_IN_DOCKER=1 \
"$image" bash scripts/run-tests.sh "$@"
exit $?
-e SHATER_REQUIRE_PRIVILEGED="${SHATER_REQUIRE_PRIVILEGED:-0}" \
"$image" bash -c '
# netplane.L3SlotFor asks the kernel through `ip link show` and reclaims a
# stale slot through `ip link del`. Without iproute2 EVERY slot reads as
# free, so TestIntegrationL3StaleSlotIsReclaimed refuses to run rather than
# pass while proving the opposite of what it claims — and [5/7] then fails
# the whole gate, correctly. golang:1.26 ships no iproute2, so install it
# here rather than let the image quietly narrow what this gate can verify.
# On a Linux host the script never re-execs, and the router has ip-full as
# a hard dependency, so this is the docker path only.
if ! command -v ip >/dev/null 2>&1; then
echo " iproute2: absent from '"$image"' — installing (the slot-reclaim test needs it)"
apt-get update -qq >/dev/null 2>&1 && apt-get install -y -qq iproute2 >/dev/null 2>&1 \
|| echo " iproute2: INSTALL FAILED — [5/7] will report the gap by name"
fi
exec bash scripts/run-tests.sh "$@"
' _ "$@"
docker_rc=$?
if [ "$host_js_rc" -ne 0 ]; then
echo "== TEST GATE FAILED — a non-Go suite failed on the host (see [7/7] above). ==" >&2
exit 1
fi
exit "$docker_rc"
fi
ALL_ROOTS=("${ROOTS[@]}" "${ROOTS_COMMON[@]}")
@@ -128,7 +332,7 @@ ALL_ROOTS=("${ROOTS[@]}" "${ROOTS_COMMON[@]}")
# shipped tags REMOVES a test file from any package, that test exists but the
# gate would never see it — which is the failure mode this whole script is about,
# just pointed the other way.
echo "== [1/4] no test file is hidden by the shipped tag set =="
echo "== [1/7] no test file is hidden by the shipped tag set =="
LISTFMT='{{.ImportPath}} {{len .TestGoFiles}} {{len .XTestGoFiles}}'
plain="$(go list -f "$LISTFMT" "${ALL_ROOTS[@]}")"
tagged="$(go list -tags "$SHATER_ROUTER_TAGS" -f "$LISTFMT" "${ALL_ROOTS[@]}")"
@@ -162,14 +366,67 @@ fi
echo
# --- the runner --------------------------------------------------------------
# Runs one suite and then PROVES it ran: every package `go list` says has tests
# must appear as `ok <pkg>` in the output. `go test` over a package whose tests
# all vanished behind a build constraint prints "[no test files]" and exits 0 —
# a green run that verified nothing.
# Runs one suite and then PROVES it ran, at BOTH granularities:
# - package: every package `go list` says has tests must appear as `ok <pkg>`.
# `go test` over a package whose tests all vanished behind a build constraint
# prints "[no test files]" and exits 0 — a green run that verified nothing.
# - test: every `--- SKIP` must be declared in SKIP_DECLARED (check_skips).
# `ok <pkg>` is printed whether the tests inside ran or stood themselves down.
LOG="$(mktemp)"
trap 'rm -f "$LOG"' EXIT
FAILED=0
# check_skips reads the per-test verdicts of the suite in $LOG and refuses to let
# a t.Skip through unnamed. Returns non-zero on an undeclared skip.
check_skips() { # $1=label ; reads $LOG
local label="$1" name reason matched entry re bad=0
# THE CONTROL, and it comes first on purpose. Everything below reads `--- SKIP`
# lines, which exist only under `go test -v`. Drop the -v and this function
# reports a clean bill of health over a suite that skipped every test it had —
# a check against silent skipping that is itself silently skipping, which is
# exactly how the first cut of [5/7] shipped (`go test -list` failing to link,
# swallowed by `|| true`, reporting "none declared"). `=== RUN` is printed for
# every test the binary starts, so its absence means the verdicts are not being
# produced at all and this instrument is reading a blank page.
if ! grep -q '^=== RUN ' "$LOG"; then
echo " FAILED [$label]: not one '=== RUN' line in the output — per-test verdicts are" >&2
echo " not being produced (is -v still on?), so the skip check was reading a" >&2
echo " blank page and its silence means nothing." >&2
return 1
fi
while read -r name; do
[ -n "$name" ] || continue
matched=""
for entry in "${SKIP_DECLARED[@]}"; do
re="${entry%%|*}"
reason="${entry#*|}"
if grep -qE "$re" <<<"$name"; then
matched="$reason"
break
fi
done
if [ -n "$matched" ]; then
echo " DECLARED SKIP $name"
echo " -> $matched"
else
echo " UNDECLARED SKIP $name" >&2
bad=1
fi
done < <(sed -n 's/^[[:space:]]*--- SKIP: \([^[:space:]]*\).*/\1/p' "$LOG" | sort -u)
if [ "$bad" -ne 0 ]; then
echo " FAILED [$label]: the test(s) above called t.Skip and are not declared in" >&2
echo " SKIP_DECLARED at the top of this script. A skipped test is a test that" >&2
echo " DID NOT RUN, and the package's 'ok' line says nothing about it. Either" >&2
echo " make it run here, or declare it by name with the reason it cannot —" >&2
echo " the reason is printed on every run, so it has to hold up." >&2
return 1
fi
return 0
}
run_suite() { # $1=label $2=extra go-test flags (may be empty) $3..=packages
local label="$1" extra="$2"
shift 2
@@ -187,17 +444,41 @@ run_suite() { # $1=label $2=extra go-test flags (may be empty) $3..=packages
echo " packages with tests: $(wc -l <<<"$expect" | tr -d ' ')"
set +e
# -v is NOT optional: it is the only way a t.Skip becomes visible at all (see
# check_skips). It costs no test TIME — measured 2026-07-26 over the fork's
# trees, warm cache, three alternating runs each: 38/25/24 s plain against
# 38/24/24 s with -v. What it costs is OUTPUT: 5 KB -> 257 KB, which is why the
# printing below is filtered rather than the flag dropped.
# shellcheck disable=SC2086 # $extra is a deliberate word-split flag list
go test -count=1 $extra \
go test -count=1 -v $extra \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
"${pkgs[@]}" >"$LOG" 2>&1
rc=$?
set -e
sed 's/^/ /' "$LOG"
if [ "$rc" -ne 0 ]; then
# A failure needs the whole story, t.Logf output and all.
sed 's/^/ /' "$LOG"
else
# A green run gets what the non-verbose gate always printed — one line per
# package — plus every skip verdict. The rest of -v's output is a
# `=== RUN`/`--- PASS` pair per test (250 KB a suite); printing it would bury
# the handful of lines anyone reads.
grep -E '^(ok|FAIL|\?)[[:space:]]|^[[:space:]]*--- SKIP: ' "$LOG" | sed 's/^/ /' || true
fi
if [ "$rc" -ne 0 ]; then
echo " FAILED [$label]: go test exited $rc" >&2
FAILED=1
# Name the skips anyway. A suite that failed somewhere else must not become
# a hiding place for a test that did not run — that is the same sin one
# level down, and while a red tree is being fixed is exactly when a skip
# gets added "temporarily". The verdict is already FAILED, so this only
# reports. Guarded on tests having actually run: a BUILD failure produces no
# verdicts to read, and check_skips' own control would then fire and bury
# the compiler error under a complaint about -v.
if grep -q '^=== RUN ' "$LOG"; then
check_skips "$label" || true
fi
return
fi
@@ -214,31 +495,233 @@ run_suite() { # $1=label $2=extra go-test flags (may be empty) $3..=packages
FAILED=1
return
fi
if ! check_skips "$label"; then
FAILED=1
return
fi
echo " OK [$label]"
}
# --- [2/4] the fork's trees, shipped tags, linux -----------------------------
echo "== [2/4] go test — the fork's trees (shipped tags, linux) =="
# --- [2/7] the fork's trees, shipped tags, linux -----------------------------
echo "== [2/7] go test — the fork's trees (shipped tags, linux) =="
run_suite main "" "${ROOTS[@]}"
echo
# --- [3/4] common/, minus the tests that need CAP_NET_ADMIN ------------------
echo "== [3/4] go test — common/ (minus the CAP_NET_ADMIN integration tests) =="
# --- [3/7] common/, minus the tests that need CAP_NET_ADMIN ------------------
echo "== [3/7] go test — common/ (minus the CAP_NET_ADMIN integration tests) =="
run_suite common "-skip $SKIP_COMMON" "${ROOTS_COMMON[@]}"
echo
# --- [4/4] -race over the same trees -----------------------------------------
# --- [4/7] -race over the same trees -----------------------------------------
# Everything, not a subset: shater/netplane alone is ~110 s under -race and it is
# the single most concurrency-critical package we own (the nft data plane), so
# once it is in, adding the rest costs ~40 s more. common/ is left out — it is
# upstream code exercised by upstream CI.
if [ "$RACE" -eq 1 ]; then
echo "== [4/4] go test -race — the fork's trees =="
echo " known-red under -race, skipped BY NAME (fix it and delete from RACE_SKIP):"
echo " TestExpiryDedupWithinDay shater/alert (test-side race, expiry_test.go:78/85)"
echo "== [4/7] go test -race — the fork's trees =="
echo " nothing is skipped under -race"
run_suite race "-race -skip $RACE_SKIP" "${ROOTS[@]}"
else
echo "== [4/4] -race pass skipped (--no-race) =="
echo "== [4/7] -race pass skipped (--no-race) =="
fi
echo
# --- [5/7] the privileged tests may not skip in silence ----------------------
# THE HOLE THIS CLOSES. Some tests can only prove what they claim against a real
# kernel: shater/generate's TestIntegrationL3TunInboundStarts opens /dev/net/tun
# and stands a real engine on it, TestIntegrationL3EgressICMPIsAFlow binds a real
# socket to a real device. Both guard themselves with t.Skip when root or the
# device is missing — the honest thing for a test to do, and completely INVISIBLE
# above: `go test` prints `ok <pkg>` whether they ran or skipped, so [2/7]'s
# per-package `ok` check is satisfied either way and the gate closes by claiming
# it "passes every test we own". That is precisely the failure this whole script
# was written for (115 of 116 test files never running while CI stayed green),
# one level down and harder to see.
#
# The list is DISCOVERED, not hand-kept — `go test -list` over the same ROOTS —
# so a privileged test written next month joins this check on the day it is
# named, with no edit here. It keys on the ^TestIntegration prefix, already this
# fork's marker for "needs capabilities the ordinary gate lacks" (SKIP_COMMON
# above excludes common/tlsspoof's TestIntegration* for exactly that reason).
# Name a privileged test anything else and it is invisible again — so don't.
#
# Verdicts, per test, by name:
# RAN — it executed here; printed so that is visible rather than assumed.
# FAILED — fatal, like any other failure.
# MISSING — `go test -list` named it and the run produced no verdict for it:
# fatal. A test that vanished between listing and running is the
# same class of hole as one hidden by a build tag.
# SKIPPED while this environment HAS root and /dev/net/tun — fatal. The
# capability guard cannot be what skipped it, so something else did
# and only the test knows what.
# SKIPPED because the environment genuinely cannot run it — reported loudly,
# by name, and it REPLACES the closing banner, so the last line of
# the gate can never claim coverage it does not have. Deliberately
# not fatal by default: the act_runner is an LXC guest whose kernel
# has no tun module at all (checked 2026-07-26 on 10.10.10.211 —
# `modprobe tun` answers "Module tun not found", /dev/net does not
# exist, and act_runner runs job containers with privileged:false
# and no container.options), so the device cannot be handed down
# without reconfiguring the Proxmox host. Making it fatal would
# paint CI permanently red and teach everyone to ignore the gate.
# SHATER_REQUIRE_PRIVILEGED=1 makes it fatal for the runs that can.
echo "== [5/7] the privileged tests (^TestIntegration) produced a verdict by name =="
PRIV_RE='^TestIntegration'
PRIV_UNVERIFIED=""
# -ldflags is NOT optional on the discovery call either: `go test -list` LINKS
# each test binary before it can enumerate its tests, and without
# -checklinkname=0 every package that pulls common/badtls fails to link. The
# first cut of this step omitted it, swallowed the error with `2>/dev/null ||
# true`, and reported "none declared" — a check against silent skipping that was
# itself silently skipping. Hence also: the exit status is inspected, and an
# empty list is only ever reported after a SUCCESSFUL enumeration.
set +e
priv_expect_raw="$(go test -list "$PRIV_RE" \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" "${ROOTS[@]}" 2>&1)"
priv_list_rc=$?
set -e
priv_expect="$(grep -E "$PRIV_RE" <<<"$priv_expect_raw" | sort -u || true)"
if [ "$priv_list_rc" -ne 0 ]; then
echo " FAILED [privileged]: could not enumerate the privileged tests (go test -list exited $priv_list_rc)." >&2
echo " An unreadable list is NOT an empty list — this check refuses to" >&2
echo " report 'nothing to verify' on the strength of a failed command." >&2
sed 's/^/ /' <<<"$priv_expect_raw" | grep -vE '^\s+(ok|\?)\s' >&2 || true
FAILED=1
elif [ -z "$priv_expect" ]; then
echo " none declared under the fork's trees — nothing to verify"
else
priv_capable=0
if [ "$(id -u)" = "0" ] && [ -e /dev/net/tun ]; then
priv_capable=1
fi
echo " declared: $(wc -l <<<"$priv_expect" | tr -d ' ')"
echo " this environment: uid=$(id -u), /dev/net/tun $([ -e /dev/net/tun ] && echo present || echo MISSING) => can run them: $([ "$priv_capable" -eq 1 ] && echo yes || echo NO)"
set +e
go test -count=1 -v -run "$PRIV_RE" \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
"${ROOTS[@]}" >"$LOG" 2>&1
priv_rc=$?
set -e
# The verdict lines plus whatever reason the test printed just before them,
# so a skip is readable here and not just counted.
grep -E '^(--- (PASS|SKIP|FAIL): |[[:space:]]+[^[:space:]]+\.go:[0-9]+: )' "$LOG" \
| sed 's/^/ | /' || true
priv_bad=0
while read -r name; do
[ -n "$name" ] || continue
if grep -qE "^--- PASS: ${name}([[:space:]]|\$)" "$LOG"; then
echo " RAN $name"
elif grep -qE "^--- FAIL: ${name}([[:space:]]|\$)" "$LOG"; then
echo " FAILED $name" >&2
priv_bad=1
elif grep -qE "^--- SKIP: ${name}([[:space:]]|\$)" "$LOG"; then
if [ "$priv_capable" -eq 1 ]; then
echo " SKIPPED $name — but this environment HAS root and /dev/net/tun, so the capability guard is NOT what skipped it" >&2
priv_bad=1
else
echo " DID NOT RUN $name — skipped: no root and/or no /dev/net/tun here"
PRIV_UNVERIFIED="$PRIV_UNVERIFIED $name"
fi
else
echo " MISSING $name — go test -list named it, the run produced no verdict for it" >&2
priv_bad=1
fi
done <<<"$priv_expect"
if [ "$priv_rc" -ne 0 ] && [ "$priv_bad" -eq 0 ]; then
echo " FAILED [privileged]: go test exited $priv_rc with every named test accounted for —" >&2
echo " a build or package-level failure, see the log above." >&2
priv_bad=1
fi
if [ "$priv_bad" -ne 0 ]; then
FAILED=1
fi
fi
echo
# --- [6/7] every non-Go test file is claimed by a runner ---------------------
# See NONGO_TEST_RUNNERS above for why this exists. The instrument is `git
# ls-files`, not a filesystem walk: a test file that is not committed is not
# anybody's coverage, and a walk would also drag node_modules in.
#
# The candidate pattern is a CLOSED positive list of the shapes a test file takes
# in this repo and the ones it plausibly will (js/ts/jsx/tsx, python, bats, shell,
# ucode). It is deliberately wider than what exists today: the whole point is to
# catch the file somebody adds next month in a language no step here knows about.
echo "== [6/7] every non-Go test file is claimed by a runner =="
NONGO_TEST_RE='(\.(test|spec)\.(js|mjs|cjs|jsx|ts|tsx)|(^|/)test_[^/]*\.py|_test\.py|\.bats|_test\.sh|_test\.uc)$'
set +e
tracked="$(git ls-files 2>&1)"
tracked_rc=$?
set -e
if [ "$tracked_rc" -ne 0 ]; then
echo " FAILED [claimed]: could not enumerate the tree (git ls-files exited $tracked_rc)." >&2
echo " An unreadable list is NOT an empty list — same rule as [5/7]." >&2
sed 's/^/ /' <<<"$tracked" >&2
FAILED=1
else
nongo="$(grep -E "$NONGO_TEST_RE" <<<"$tracked" | sort || true)"
if [ -z "$nongo" ]; then
echo " FAILED [claimed]: not one non-Go test file found in a tree that has several." >&2
echo " The pattern stopped matching; this check would pass having looked" >&2
echo " at nothing." >&2
FAILED=1
else
echo " non-Go test files: $(wc -l <<<"$nongo" | tr -d ' ')"
unclaimed=0
while read -r f; do
[ -n "$f" ] || continue
owner=""
for entry in "${NONGO_TEST_RUNNERS[@]}"; do
if grep -qE "${entry%%|*}" <<<"$f"; then
owner="${entry#*|}"
break
fi
done
if [ -n "$owner" ]; then
echo " claimed $f"
echo " -> $owner"
else
echo " UNCLAIMED $f — no runner in this gate executes it" >&2
unclaimed=1
fi
done <<<"$nongo"
if [ "$unclaimed" -ne 0 ]; then
echo " FAILED [claimed]: the file(s) above are test files that NOTHING runs." >&2
echo " That is a test which cannot fail — the most expensive kind, because" >&2
echo " it reads as coverage. Either wire a runner (JS_SUITES below, or" >&2
echo " scripts/run-panel-tests.sh) and declare it in NONGO_TEST_RUNNERS, or" >&2
echo " delete the file. Declaring it without wiring one is not an option:" >&2
echo " the runner named there is the one [7/7] reports a verdict for." >&2
FAILED=1
fi
fi
fi
echo
# --- [7/7] the non-Go suites this gate owns, with a verdict by name ----------
# The work is in js_step() at the top of this file; see there for what `node
# <file>` proves and what it cannot.
#
# NODE MAY BE ABSENT, and that is handled the way [5/7] handles a missing
# /dev/net/tun: the files are named, the run says out loud that they DID NOT RUN,
# and that notice REPLACES the closing banner so this script can never end by
# claiming coverage it does not have. Not fatal by default, because the
# golang:1.26 image the non-linux re-exec uses has no node. CI does: both
# .gitea/workflows/test.yml and release.yml run actions/setup-node@v4 (node 24)
# and scripts/run-panel-tests.sh BEFORE this script, in the same job, so on the
# release path node is on PATH here and these really execute.
#
# SHATER_JS_ALREADY_RAN=1 means the re-exec that started this container ran them
# on the host first and will fold their verdict into its own exit status — so
# re-running them here would only be slower and, without node, would print a
# "did not run" that is not true of this gate as a whole.
echo "== [7/7] the non-Go suites (node) produced a verdict by name =="
if [ "${SHATER_JS_ALREADY_RAN:-0}" = "1" ]; then
echo " already run on the host before the re-exec into this container (see above);"
echo " that run's verdict is folded into the exit status of the script that started it."
else
js_step "here" || FAILED=1
fi
echo
@@ -246,4 +729,40 @@ if [ "$FAILED" -ne 0 ]; then
echo "== TEST GATE FAILED — nothing may be published from this run. ==" >&2
exit 1
fi
if [ -n "$PRIV_UNVERIFIED" ] || [ -n "$JS_UNVERIFIED" ]; then
echo "== !! PASSED, BUT NOT FULLY VERIFIED !! =================================="
echo " Every test that COULD run here passed. These did not run at all:"
for t in $PRIV_UNVERIFIED $JS_UNVERIFIED; do
echo " - $t"
done
echo
fi
if [ -n "$JS_UNVERIFIED" ]; then
echo " The file(s) above with a path are non-Go suites and this environment has"
echo " no \`node\`. The golang image the non-linux re-exec uses does not ship one;"
echo " CI does (actions/setup-node@v4, node 24, in the same job before this"
echo " script), so on the release path they DO run. To run them here:"
echo " node openwrt/luci-app-shater/tests/status-readout.test.js"
echo
fi
if [ -n "$PRIV_UNVERIFIED" ]; then
echo " The named tests above need root + CAP_NET_ADMIN + /dev/net/tun, which"
echo " this environment does not have. Nothing about the kernel paths they"
echo " cover was verified by this run. To actually run them, from a host whose"
echo " docker can:"
echo " scripts/run-tests.sh # the re-exec hands the container both"
echo " or directly:"
echo " docker run --rm --cap-add NET_ADMIN --device /dev/net/tun \\"
echo " -v \"\$PWD\":/src -w /src golang:1.26 bash scripts/run-tests.sh"
echo " or on the OpenWrt VM. SHATER_REQUIRE_PRIVILEGED=1 makes this a hard"
echo " failure instead of this notice."
fi
if [ -n "$PRIV_UNVERIFIED" ] || [ -n "$JS_UNVERIFIED" ]; then
echo "=========================================================================="
if [ -n "$PRIV_UNVERIFIED" ] && [ "${SHATER_REQUIRE_PRIVILEGED:-0}" = "1" ]; then
echo "== TEST GATE FAILED: SHATER_REQUIRE_PRIVILEGED=1 and the tests above did not run. ==" >&2
exit 1
fi
exit 0
fi
echo "== OK: the shipped tag set, on linux, passes every test we own. =="
+123
View File
@@ -0,0 +1,123 @@
#!/bin/sh
# scripts/testbed-lao.sh — add a SECOND LAN network ("lao", 10.67.1.0/24) in its
# own fw4 zone, on a testbed router.
#
# WHY IT EXISTS
#
# Almost every zone-related defect in this project is invisible on a router with
# one LAN zone, because "the zone" and "the LAN" are the same thing there. The
# divert set the daemon builds spans every LAN inbound and every `iface:`/`zone:`
# rule source, so on a multi-zone router traffic from the other zones is marked,
# routed, accepted by `inet shater` — and then dropped by fw4's zone policy,
# silently. Reproducing that needs a second zone and nothing else: no second
# physical port, no client, no traffic. This script makes one.
#
# WHY IT IS NOT SHIPPED
#
# It lives in scripts/ and is NOT installed by openwrt/shater-core/Makefile,
# which lists every file it installs by name. That is the whole opt-in mechanism,
# and it was chosen over the alternatives on purpose:
#
# - an extra /etc/uci-defaults/ file would run on EVERY install, handing a
# second network and a second firewall zone to every ordinary user — the one
# thing this must not do;
# - an environment variable read inside 30_shater-core is unreachable in
# practice: that script is deleted after its first successful run, so there
# is no later moment at which an operator could set the variable and re-run it;
# - a separate package would need a feed entry, a build, a release and a
# version, for a file that exists to be scp'd onto one VM.
#
# USAGE
#
# scp scripts/testbed-lao.sh root@testbed:/tmp/ && ssh root@testbed sh /tmp/testbed-lao.sh
# ssh root@testbed sh /tmp/testbed-lao.sh --remove
#
# It is idempotent (every section is NAMED and guarded), purely additive, and
# touches no existing section. Running it three times in a row leaves exactly one
# of everything.
set -e
REMOVE=0
[ "$1" = "--remove" ] && REMOVE=1
if [ "$REMOVE" = 1 ]; then
uci -q delete firewall.lao_fwd
uci -q delete firewall.lao
uci -q delete dhcp.lao
uci -q delete network.lao
uci -q delete network.br_lao
uci -q commit firewall
uci -q commit dhcp
uci -q commit network
/etc/init.d/network reload
/etc/init.d/firewall reload
echo "lao removed"
exit 0
fi
# --- L2: an empty bridge -----------------------------------------------------
# No ports on purpose: the point is a second ROUTED network with its own firewall
# zone, and giving it a switch port would mean re-cabling a testbed for nothing.
#
# bridge_empty is what makes a portless bridge usable. Without it netifd leaves a
# member-less bridge down (no carrier), the `lao` interface never comes up, fw4
# resolves `list network 'lao'` to an EMPTY device set, and the zone silently
# matches nothing — which would make this script a worse instrument than no
# instrument, since it would look set up and prove nothing.
if ! uci -q get network.br_lao >/dev/null; then
uci set network.br_lao=device
uci set network.br_lao.name='br-lao'
uci set network.br_lao.type='bridge'
uci set network.br_lao.bridge_empty='1'
fi
# --- L3: the interface -------------------------------------------------------
# 10.67.1.0/24 is deliberately far from anything a home LAN or a proxy node uses.
# ipaddr+netmask rather than CIDR: CIDR in `ipaddr` is a 25.12 convenience and
# this script should also run on an older testbed image.
if ! uci -q get network.lao >/dev/null; then
uci set network.lao=interface
uci set network.lao.proto='static'
uci set network.lao.device='br-lao'
uci set network.lao.ipaddr='10.67.1.1'
uci set network.lao.netmask='255.255.255.0'
fi
# --- DHCP: same shape as lan -------------------------------------------------
if ! uci -q get dhcp.lao >/dev/null; then
uci set dhcp.lao=dhcp
uci set dhcp.lao.interface='lao'
uci set dhcp.lao.start='100'
uci set dhcp.lao.limit='150'
uci set dhcp.lao.leasetime='12h'
fi
# --- Firewall: its OWN zone, which is the entire point -----------------------
# Same policies as the stock lan zone and its own forwarding to wan, so the
# network behaves like a second LAN. What it does NOT get here is a forwarding
# into shater_l3: seeding that for every zone is the job under test, done by
# /etc/uci-defaults/30_shater-core. If this script seeded it, the test would be
# testing itself.
if ! uci -q get firewall.lao >/dev/null; then
uci set firewall.lao=zone
uci set firewall.lao.name='lao'
uci set firewall.lao.input='ACCEPT'
uci set firewall.lao.output='ACCEPT'
uci set firewall.lao.forward='ACCEPT'
uci add_list firewall.lao.network='lao'
fi
if ! uci -q get firewall.lao_fwd >/dev/null; then
uci set firewall.lao_fwd=forwarding
uci set firewall.lao_fwd.src='lao'
uci set firewall.lao_fwd.dest='wan'
fi
uci commit network
uci commit dhcp
uci commit firewall
/etc/init.d/network reload
/etc/init.d/firewall reload
echo "lao seeded: br-lao 10.67.1.1/24, fw4 zone lao -> wan"
+141
View File
@@ -0,0 +1,141 @@
package alert
import (
"fmt"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
)
// warnLogger records Warn lines so a test can assert that an eviction was
// actually announced. Everything else falls through to the standard logger.
type warnLogger struct {
log.ContextLogger
mu sync.Mutex
warns []string
}
func newWarnLogger() *warnLogger { return &warnLogger{ContextLogger: log.StdLogger()} }
func (l *warnLogger) Warn(args ...any) {
l.mu.Lock()
l.warns = append(l.warns, fmt.Sprint(args...))
l.mu.Unlock()
}
func (l *warnLogger) lines() []string {
l.mu.Lock()
defer l.mu.Unlock()
return append([]string(nil), l.warns...)
}
// notifierAt builds a Notifier with no channels (so nothing is ever delivered —
// only the dedup bookkeeping runs) and a clock the test drives.
func notifierAt(t *testing.T, clock *time.Time) (*Notifier, *warnLogger) {
t.Helper()
lg := newWarnLogger()
n := New([]model.Alert{}, lg)
n.now = func() time.Time { return *clock }
return n, lg
}
func (n *Notifier) dedupSize() int {
n.mu.Lock()
defer n.mu.Unlock()
return len(n.dedup)
}
// TestDedupTableIsBoundedOverTime is the leak itself: a guest network whose
// clients randomise their MAC produces an endless stream of distinct
// "new_device:<MAC>" keys, and nothing ever deleted one. Spread over time — which
// is how it actually happens — the table must stay small, and nothing may be
// reported as evicted, because an entry past the dedup window could no longer
// suppress anything anyway.
func TestDedupTableIsBoundedOverTime(t *testing.T) {
clock := time.Now()
n, lg := notifierAt(t, &clock)
// Ten times the cap, at one incident per second: every key is long past the
// 60s window by the time the next batch arrives.
const fires = maxDedupKeys * 10
for i := 0; i < fires; i++ {
clock = clock.Add(time.Second)
n.FireIncident(Incident{
Events: []string{"new_device"},
Title: "New device on the LAN",
Key: fmt.Sprintf("new_device:02:00:00:%02x:%02x:%02x", i>>16&0xff, i>>8&0xff, i&0xff),
})
}
if got := n.dedupSize(); got > maxDedupKeys {
t.Fatalf("dedup table holds %d entries after %d distinct incidents; cap is %d",
got, fires, maxDedupKeys)
}
if got := n.DedupEvicted(); got != 0 {
t.Fatalf("reported %d LIVE evictions; entries aged out of the window and losing them costs nothing", got)
}
if lines := lg.lines(); len(lines) != 0 {
t.Fatalf("expiry sweep must be silent (it loses nothing), got: %v", lines)
}
}
// TestDedupTableEvictionIsAnnounced is the other half: when the cap genuinely
// bites — more distinct incidents inside ONE dedup window than the table holds —
// live suppression state is lost and repeats may notify twice. That must be said
// out loud, not absorbed.
func TestDedupTableEvictionIsAnnounced(t *testing.T) {
clock := time.Now()
n, lg := notifierAt(t, &clock)
// The clock does not move: every key stays inside its window.
for i := 0; i < maxDedupKeys+10; i++ {
clock = clock.Add(time.Millisecond) // still far inside dedupWindow
n.FireIncident(Incident{
Events: []string{"new_device"},
Title: "New device on the LAN",
Key: fmt.Sprintf("new_device:flood-%d", i),
})
}
if got := n.dedupSize(); got > maxDedupKeys {
t.Fatalf("dedup table grew to %d, above the %d cap", got, maxDedupKeys)
}
if n.DedupEvicted() == 0 {
t.Fatalf("a flood of %d in-window incidents evicted nothing — the cap is not enforced", maxDedupKeys+10)
}
lines := lg.lines()
if len(lines) == 0 {
t.Fatalf("live suppression entries were dropped with no notice")
}
if !strings.Contains(lines[0], "dedup table full") || !strings.Contains(lines[0], "notify twice") {
t.Fatalf("eviction notice does not explain the consequence: %q", lines[0])
}
}
// TestDedupStillSuppressesWithinTheWindow guards the behaviour the bound must not
// break: a repeat inside the window is still collapsed, and one outside it is not.
func TestDedupStillSuppressesWithinTheWindow(t *testing.T) {
clock := time.Now()
n, _ := notifierAt(t, &clock)
fire := func() bool {
n.mu.Lock()
defer n.mu.Unlock()
return n.suppressedLocked("killswitch\x00Kill-switch engaged")
}
if fire() {
t.Fatalf("first fire was suppressed")
}
clock = clock.Add(dedupWindow / 2)
if !fire() {
t.Fatalf("a repeat inside the window was NOT suppressed")
}
clock = clock.Add(dedupWindow)
if fire() {
t.Fatalf("a repeat past the window was suppressed")
}
}
+87
View File
@@ -0,0 +1,87 @@
package alert
import (
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// closingRT is an http.RoundTripper that also implements the CloseIdleConnections
// hook http.Client forwards to, so a test can observe whether the client was ever
// released. http.Transport implements the same hook — this stands in for it.
type closingRT struct {
rt http.RoundTripper
closed atomic.Int32
}
func (c *closingRT) RoundTrip(r *http.Request) (*http.Response, error) { return c.rt.RoundTrip(r) }
func (c *closingRT) CloseIdleConnections() { c.closed.Add(1) }
// TestDetourClientIsClosedAfterDelivery: the detour factory (engine.HTTPClient)
// mints a NEW http.Transport for every call, and that transport's idle connections
// are live proxying sessions through an engine outbound whose object the dial
// closure pins. The notifier used one per delivery and dropped it, so every alert
// left a keep-alive session — and a reference to a possibly-retired engine
// generation — behind for the whole idle timeout.
func TestDetourClientIsClosedAfterDelivery(t *testing.T) {
c, srv := newSink(t)
n := New([]model.Alert{{
Name: "hook", Enabled: true, Type: "webhook", URL: srv.URL,
Events: []string{"killswitch"}, Via: "node:tunnel",
}}, nil)
var made []*closingRT
n.SetClientFactory(func(via string) (*http.Client, error) {
rt := &closingRT{rt: http.DefaultTransport}
made = append(made, rt)
return &http.Client{Transport: rt}, nil
})
n.Fire("killswitch", "Kill-switch engaged", "the tunnel is down")
n.Wait()
if c.n() != 1 {
t.Fatalf("delivery count = %d, want 1", c.n())
}
if len(made) != 1 {
t.Fatalf("factory called %d times, want 1", len(made))
}
if got := made[0].closed.Load(); got == 0 {
t.Fatalf("the per-delivery detour client was never closed — its idle connections " +
"(and the engine outbound its dialer pins) outlive the alert")
}
}
// TestDetourClientIsClosedWhenTheSendFails covers the fallback path: a detour that
// errors and falls back to direct still built a transport, and that one leaked too.
func TestDetourClientIsClosedWhenTheSendFails(t *testing.T) {
// A sink that rejects, so the detour send fails and Fallback kicks in.
reject := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
}))
defer reject.Close()
n := New([]model.Alert{{
Name: "hook", Enabled: true, Type: "webhook", URL: reject.URL,
Events: []string{"killswitch"}, Via: "node:tunnel", Fallback: true,
}}, nil)
var rt *closingRT
n.SetClientFactory(func(via string) (*http.Client, error) {
rt = &closingRT{rt: http.DefaultTransport}
return &http.Client{Transport: rt}, nil
})
n.Fire("killswitch", "Kill-switch engaged", "the tunnel is down")
n.Wait()
if rt == nil {
t.Fatalf("factory was never called")
}
if got := rt.closed.Load(); got == 0 {
t.Fatalf("a failed detour delivery still leaked its transport")
}
}
+21 -5
View File
@@ -5,6 +5,7 @@ import (
"os"
"path/filepath"
"strings"
"sync"
"testing"
"time"
@@ -74,22 +75,37 @@ func TestExpiryDedupWithinDay(t *testing.T) {
// test compresses 24 simulated hours into a few milliseconds of wall clock — so
// without this the second delivery is (correctly) suppressed by that window and
// the test would be measuring the fixture, not the behaviour.
// The clock is shared with the notifier's DELIVERY goroutines (buildPayload
// stamps the payload with n.now()), which are still in flight while this body
// advances the simulated time — so it needs a lock, not a bare variable.
var clockMu sync.Mutex
simNow := now
e.n.now = func() time.Time { return simNow }
setNow := func(t time.Time) {
clockMu.Lock()
simNow = t
clockMu.Unlock()
}
e.n.now = func() time.Time {
clockMu.Lock()
defer clockMu.Unlock()
return simNow
}
if got := e.Run(m, now); got != 1 {
t.Fatalf("first Run fired %d, want 1", got)
}
// Simulate a minute-by-minute reconcile for the next 23 hours.
for i := 1; i <= 23; i++ {
simNow = now.Add(time.Duration(i) * time.Hour)
if got := e.Run(m, simNow); got != 0 {
at := now.Add(time.Duration(i) * time.Hour)
setNow(at)
if got := e.Run(m, at); got != 0 {
t.Fatalf("Run at +%dh fired %d alerts, want 0 (dedup window is 24h)", i, got)
}
}
// Just past the window it may speak again.
simNow = now.Add(24*time.Hour + time.Minute)
if got := e.Run(m, simNow); got != 1 {
past := now.Add(24*time.Hour + time.Minute)
setNow(past)
if got := e.Run(m, past); got != 1 {
t.Errorf("Run just past 24h fired %d, want 1", got)
}
e.n.Wait()
+99
View File
@@ -20,6 +20,7 @@ import (
"fmt"
"io"
"net/http"
"sort"
"strings"
"sync"
"time"
@@ -31,6 +32,29 @@ import (
const (
httpTimeout = 8 * time.Second
dedupWindow = 60 * time.Second
// maxDedupKeys / keepDedupKeys bound the dedup table.
//
// Nothing ever deleted from it. Every key that had EVER fired stayed forever,
// and its highest-cardinality producer is "new_device:<MAC>" — on a guest
// network where clients randomise their MAC per association, that is a fresh
// key per device per join, for the life of a daemon that runs for months.
//
// The real bound is not this cap, it is the window: an entry older than
// dedupWindow (60s) can never suppress anything again, so it is pure garbage
// and sweeping it costs nothing and changes no behaviour. compactDedupLocked
// does that first. The cap only bites when 1024 DISTINCT incidents fired inside
// one 60-second window — and there the eviction is a real loss of suppression
// state, so it is reported rather than done quietly.
//
// Why 1024: the new_device watcher polls every ~45s, so a poll would have to
// discover a thousand previously-unseen MACs at once to reach it — roughly 20x
// the worst guest-network churn this box has seen. At ~100 B per entry (a
// 28-byte "new_device:<MAC>" key plus a time.Time plus map overhead) the full
// table is ~100 KB of a 512 MB router: cheap enough that a generous headroom
// costs nothing.
maxDedupKeys = 1024
keepDedupKeys = 512
)
// telegramAPIBase is the Telegram Bot API root. A package var so tests can point
@@ -43,6 +67,10 @@ type Notifier struct {
mu sync.Mutex
alerts []model.Alert
dedup map[string]time.Time // (event\x00title) -> last fire time
// dedupEvicted counts LIVE dedup entries dropped by the capacity bound (see
// compactDedupLocked). Expired entries swept out are NOT counted: they could no
// longer suppress anything, so dropping them loses nothing.
dedupEvicted uint64
client *http.Client
log log.ContextLogger
@@ -227,10 +255,69 @@ func (n *Notifier) suppressedLocked(key string) bool {
if last, ok := n.dedup[key]; ok && now.Sub(last) < dedupWindow {
return true
}
if len(n.dedup) >= maxDedupKeys {
n.compactDedupLocked(now)
}
n.dedup[key] = now
return false
}
// compactDedupLocked bounds the dedup table. Caller holds n.mu.
//
// Two stages, deliberately distinct because only one of them loses anything:
//
// 1. Drop every entry older than dedupWindow. Such an entry cannot suppress a
// future fire (suppressedLocked already ignores it), so this is garbage
// collection, not eviction: no notification changes, and nothing is reported.
// On any realistic traffic this stage alone keeps the table at "keys seen in
// the last minute".
// 2. If the table is STILL full, a thousand distinct incidents fired inside one
// window. Now eviction is real — the oldest live entries go, and a repeat of
// one of them within its window will notify a second time instead of being
// collapsed. That is a visible change in behaviour, so it is logged, with the
// running total, rather than silently absorbed.
func (n *Notifier) compactDedupLocked(now time.Time) {
for k, last := range n.dedup {
if now.Sub(last) >= dedupWindow {
delete(n.dedup, k)
}
}
if len(n.dedup) < maxDedupKeys {
return
}
type kv struct {
key string
at time.Time
}
all := make([]kv, 0, len(n.dedup))
for k, at := range n.dedup {
all = append(all, kv{k, at})
}
sort.Slice(all, func(i, j int) bool { return all[i].at.Before(all[j].at) })
drop := len(all) - keepDedupKeys
for i := 0; i < drop; i++ {
delete(n.dedup, all[i].key)
}
n.dedupEvicted += uint64(drop)
n.log.Warn("alert: dedup table full (", maxDedupKeys,
" incidents inside one ", dedupWindow, " window) — dropped ", drop,
" live suppression entries (", n.dedupEvicted,
" total); repeats of those incidents may notify twice")
}
// DedupEvicted reports how many LIVE suppression entries the cap has dropped since
// start (stage 2 of compactDedupLocked only — the expiry sweep is not counted,
// because it loses nothing). Nonzero means alerts may have been delivered twice.
func (n *Notifier) DedupEvicted() uint64 {
if n == nil {
return 0
}
n.mu.Lock()
defer n.mu.Unlock()
return n.dedupEvicted
}
// dispatch delivers to a single alert in its own goroutine. Panics are recovered
// and logged; a delivery error is logged. It never crashes the daemon.
func (n *Notifier) dispatch(a model.Alert, event, title, body string) {
@@ -276,6 +363,18 @@ func (n *Notifier) deliver(a model.Alert, event, title, body string) error {
if detour && factory != nil {
client, cerr := factory(via)
if cerr == nil && client != nil {
// The factory (engine.HTTPClient) builds a BRAND NEW http.Transport per
// call, with keep-alive and a 90s idle timeout, and we use it for exactly
// one POST. Dropping it without this leaves the idle connection — a real
// proxying session through an engine outbound, plus its read and write
// loops — alive for the whole idle timeout. Worse, the transport's
// DialContext closure captures that outbound object, so the idle
// connection PINS a retired engine generation whose close budget is 5
// seconds. One alert delivery per minute keeps a permanent rolling set of
// them.
defer client.CloseIdleConnections()
}
if cerr != nil {
if a.Fallback {
n.log.Warn("alert: ", a.Name, " detour ", via, " unavailable (", cerr, ") — falling back to direct")
+788 -55
View File
File diff suppressed because it is too large Load Diff
+156 -19
View File
@@ -1,15 +1,48 @@
package apply
import (
"errors"
"os"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// TestMain takes THIS TEST BINARY off the real network namespace, for the one
// operation of netplane's that destroys something another test binary can be
// using: the L3-ingress TUN devices.
//
// Applier.Teardown calls netplane.TeardownRouting for real here, and its
// device sweep deletes BOTH slots unconditionally. `go test` runs package
// binaries concurrently (-p defaults to GOMAXPROCS) and they all share one
// network namespace, so on a runner with iproute2 and /dev/net/tun this binary
// was issuing 9 `ip link del shater-l3a` and 9 `ip link del shater-l3b` per
// gate run — measured with an `ip` shim on PATH — into the namespace where
// shater/generate's privileged tests hold a live TUN. What that looks like from
// the other side is a red TestIntegrationL3* in a package that did nothing
// wrong:
//
// no [shater-l3a shater-l3b] device exists after a successful Start
//
// netplane.L3StubKernelForTest carries the full measurement and the reasoning.
//
// Scope, stated rather than implied: this diverts ONLY the L3 device deletes.
// The `ip rule del` / `ip route flush` on the reserved tables that teardown also
// performs still run for real from here. They are left alone because nothing
// else in the gate reads those tables, so unlike the devices they have no
// observed victim — not because they are harmless on a machine that matters.
func TestMain(m *testing.M) {
restore := netplane.L3StubKernelForTest()
code := m.Run()
restore()
os.Exit(code)
}
// TestCanRollback pins the signal the panel gates its rollback control on:
// canRollback is false on a fresh applier (no armed commit-confirm snapshot AND
// the engine holds no last-good predecessor), and flips to true once a snapshot
@@ -372,36 +405,140 @@ func TestHoldLockedOpenInstallsNothing(t *testing.T) {
}
}
// TestApplyInstallsHoldWhenEngineFailsToStart is the end-to-end W5 wiring: a
// model the engine cannot come up on must leave the data plane PROTECTING, not
// absent. On the VM the trigger was an unreachable remote rule-set at boot; any
// construction failure takes the same path.
// errEngineStartFailedInTest stands in for what really happens on the router:
// engine.Apply -> newBox -> box.New refuses the config (an unreachable remote
// rule-set at boot, a node the registry cannot build, a busy port), so the engine
// never starts. TestEngineApplyReallyFailsWithoutStarting below is the CONTROL
// that this is a faithful model and not a convenient fiction.
var errEngineStartFailedInTest = errors.New("create instance: initialize outbound[0]: outbound type not found")
// stubEngineApply replaces the engineApply seam — applyLocked's engine-swap step,
// documented in apply.go as existing precisely so a failing stage can be tested
// without a router — and COUNTS the calls.
//
// The counter is not decoration. This is the test that has to survive the way its
// predecessor did not: it used to break the engine by pointing a rule-set at
// /nonexistent/nope.srs and then guard itself with `if err == nil ||
// a.eng.Running() { t.Skip(...) }`. Once LocalRuleSet.reloadFile started treating
// an unreadable file as an EMPTY one, the engine came up fine, the guard fired on
// every platform, and the package still reported `ok` — the single most important
// test of this package asserted nothing at all for months. A seam the test drives
// itself cannot rot that way; the counter closes the one remaining hole, which is
// applyLocked ceasing to go through the seam at all.
func stubEngineApply(t *testing.T, err error) *int {
t.Helper()
var calls int
orig := engineApply
engineApply = func(*Applier, option.Options) (bool, error) {
calls++
return false, err
}
t.Cleanup(func() { engineApply = orig })
return &calls
}
// stubTableExists pins the "is our nft table in the kernel" fact, which Status
// reads to name the plane. A unit test has no kernel table, so without this the
// only reachable verdict is "none" and the "hold" branch — the one the panel shows
// the operator — is never exercised.
func stubTableExists(t *testing.T, loaded bool) {
t.Helper()
orig := tableExists
tableExists = func() bool { return loaded }
t.Cleanup(func() { tableExists = orig })
}
// TestApplyInstallsHoldWhenEngineFailsToStart is the end-to-end W5 wiring: a model
// the engine cannot come up on must leave the data plane PROTECTING, not absent.
// On the VM the trigger was an unreachable remote rule-set at boot; any
// construction failure takes the same path, which is why the failure is injected
// at the seam rather than reproduced through one particular cause — the branch
// under test is applyLocked's, and a cause that stops causing (see stubEngineApply)
// silently retires the test.
//
// Everything except the engine swap is REAL here: the model, generate, the
// kill-switch decision, netplane.RenderHoldNft, the latch and Status. Only the
// load into the kernel is intercepted (withHoldProbe), so what the assertions read
// is the ruleset that would have gone to `nft -f`.
func TestApplyInstallsHoldWhenEngineFailsToStart(t *testing.T) {
loaded := withHoldProbe(t)
calls := stubEngineApply(t, errEngineStartFailedInTest)
stubTableExists(t, true) // the holding plane we install below IS a loaded table
a := New(engine.New(), nil)
m := holdModel("closed")
// A rule-set file that cannot be read: generate warns, and the engine build
// fails for one reason or another on every platform we run on.
m.Rulesets = []model.Ruleset{{
Name: "badfile", Type: "domain", Source: "file",
Path: "/nonexistent/nope.srs", Format: "binary",
}}
_, err := a.Apply(m)
if err == nil || a.eng.Running() {
t.Skip("this platform started the engine anyway; the hold path is covered by TestHoldLockedInstallsBlockingPlane")
// The instrument first: an assertion suite that never reached the branch is
// worth nothing, and saying so by name is the whole lesson of this test.
if *calls != 1 {
t.Fatalf("the engine-swap seam ran %d times, want exactly 1 — applyLocked no longer goes "+
"through engineApply, so this test is NOT exercising the engine-failure branch", *calls)
}
if len(*loaded) == 0 {
t.Fatalf("engine failed to start (%v) but NO holding plane was installed — "+
"the router would forward LAN traffic to the WAN in the clear with kill_switch=closed", err)
if !errors.Is(err, errEngineStartFailedInTest) {
t.Fatalf("Apply returned %v, want the engine failure unwrapped — the caller (cmd/shaterd, "+
"the panel) decides what to tell the operator from this error", err)
}
if !strings.Contains((*loaded)[0], "drop") {
t.Errorf("the installed plane does not drop:\n%s", (*loaded)[0])
if a.eng.Running() {
t.Fatalf("precondition: the engine must not be running after a failed swap")
}
if len(*loaded) != 1 {
t.Fatalf("engine failed to start (%v) but %d holding planes were installed, want 1 — "+
"with none, the router forwards LAN traffic to the WAN in the clear while "+
"kill_switch=closed", err, len(*loaded))
}
rs := (*loaded)[0]
if !strings.Contains(rs, "meta nfproto ipv4 drop") || !strings.Contains(rs, "meta nfproto ipv6 drop") {
t.Errorf("the installed plane does not block forwarded traffic on both families:\n%s", rs)
}
if strings.Contains(rs, "hook input") || strings.Contains(rs, "hook output") {
t.Errorf("the installed plane must only hook forward, or the operator loses management access:\n%s", rs)
}
if strings.Contains(rs, "tproxy ") {
t.Errorf("the installed plane must not divert to a dead engine socket:\n%s", rs)
}
if !a.Holding() {
t.Errorf("Holding() must report true so the panel can show 'protected, not proxying'")
}
if got := a.Status().Plane; got != "hold" && got != "none" {
t.Errorf("Status().Plane = %q, want hold (or none where nft is unavailable)", got)
s := a.Status()
if s.Plane != "hold" {
t.Errorf("Status().Plane = %q, want \"hold\" — with a table loaded and the engine down, "+
"\"full\" would tell the operator traffic is being proxied when nothing is", s.Plane)
}
if s.EngineRunning {
t.Errorf("Status().EngineRunning must be false after a failed engine swap")
}
if s.Traffic.Verdict != "" {
t.Errorf("Traffic.Verdict = %q, want \"\" (unknown): no config of ours is running, and a "+
"leftover verdict is a reassuring lie", s.Traffic.Verdict)
}
}
// TestEngineApplyReallyFailsWithoutStarting is the CONTROL for the stub above: it
// proves that the seam's PRODUCTION twin, engine.Apply, really can return an error
// with the engine left stopped — i.e. that the state
// TestApplyInstallsHoldWhenEngineFailsToStart simulates is a state this fork can
// actually be in. Without it, that test would be an instrument with no proof it
// measures anything real.
//
// The trigger is chosen to be independent of every moving part around it: an
// outbound type no registry can ever hold fails in adapter/outbound.Registry.
// CreateOutbound ("outbound type not found"), before any listener is opened, so it
// needs no root, no network, no TUN and no nft — and it cannot quietly start
// succeeding the way an unreadable rule-set file did.
func TestEngineApplyReallyFailsWithoutStarting(t *testing.T) {
e := engine.New()
_, err := e.Apply(option.Options{
Outbounds: []option.Outbound{{Type: "shater-no-such-outbound-type", Tag: "probe"}},
})
if err == nil {
t.Fatalf("engine.Apply accepted an outbound type that cannot exist; the failure this " +
"package's hold path is built for would never occur and the W5 test above simulates nothing")
}
if e.Running() {
t.Errorf("engine.Apply failed with %v but left the engine RUNNING — the hold path is "+
"gated on !Running(), so it would never engage", err)
}
}
+366
View File
@@ -0,0 +1,366 @@
package apply
import (
"errors"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/generate"
"github.com/sagernet/sing-box/shater/model"
)
// A subscription's fetch_detour is resolved by Applier.HTTPClient, and `chain:<X>`
// is the one form the engine cannot resolve on its own: a chain has no outbound
// named after itself. These tests pin the resolution against the tag grammar the
// GENERATOR actually emits, and pin what happens on every miss — the answer must
// never be "fetch it direct", because that puts the feed and the router's real
// address on the plain WAN, which is what fetch_via=proxy exists to prevent.
// tagSet builds a running-box tag set for the pure resolver.
func tagSet(tags ...string) map[string]bool {
out := make(map[string]bool, len(tags))
for _, t := range tags {
out[t] = true
}
return out
}
// ssURI is a parseable share link so generate can build a real outbound from it.
func ssURI(host, name string) string {
return "ss://YWVzLTI1Ni1nY206cGFzcw@" + host + ":8388#" + name
}
// TestChainDetourTagMatchesGenerator is the tripwire: it runs the REAL generator
// over a model whose rule targets a chain, and requires that the tag resolveViaTag
// hands the engine is one the generator actually emitted — and specifically the
// chain's entry (highest hop index), not the chain's own name.
//
// This is the test that fails on the unfixed code: engine.ViaToTag("chain:work")
// yields "work", which appears nowhere in the emitted config.
func TestChainDetourTagMatchesGenerator(t *testing.T) {
m := &model.Model{
Globals: model.Globals{Enabled: true},
Nodes: []model.Node{
{Name: "n1", Enabled: true, URI: ssURI("1.2.3.4", "n1")},
{Name: "n2", Enabled: true, URI: ssURI("5.6.7.8", "n2")},
},
Chains: []model.Chain{{Name: "work", Hops: []string{"node:n1", "node:n2"}}},
Rules: []model.Rule{
{Name: "r", Enabled: true, Target: "chain:work", DstPort: "443"},
},
}
opts, _, err := generate.GenerateWithWarnings(m)
if err != nil {
t.Fatalf("generate: %v", err)
}
emitted := map[string]bool{}
for _, ob := range opts.Outbounds {
emitted[ob.Tag] = true
}
for _, ep := range opts.Endpoints {
emitted[ep.Tag] = true
}
if !emitted["chain-work-h2"] {
t.Fatalf("fixture broken: generator emitted no chain-work-h2; tags=%v", emitted)
}
got, err := resolveViaTag("chain:work", emitted, m.Chains)
if err != nil {
t.Fatalf("resolveViaTag(chain:work): %v", err)
}
if !emitted[got] {
t.Fatalf("resolved %q, which the generator never emitted (tags=%v)", got, emitted)
}
if got != "chain-work-h2" {
t.Fatalf("resolved %q, want the chain ENTRY chain-work-h2", got)
}
}
// TestChainDetourEntryIsTheLastHop pins entry = highest hop index, and that a group
// hop's member copies (chain-<n>-h<i>-<member>) are not mistaken for hops. Dialling
// a member copy instead of its wrapper is not a harmless near-miss: the wrapper IS
// the group's selector, so the copy pins one member and throws the balancing away.
func TestChainDetourEntryIsTheLastHop(t *testing.T) {
cases := []struct {
name string
tags map[string]bool
want string
}{
{
name: "highest index wins, 10 beats 2",
tags: tagSet(
"direct", "block", "n1", "n2", "n3",
"chain-work-h1", "chain-work-h1-n1",
"chain-work-h2", "chain-work-h2-n2", "chain-work-h2-n3",
"chain-work-h10", "chain-work-h10-n3",
"chain-workshop-h1", // another chain sharing the prefix
),
want: "chain-work-h10",
},
{
// The digits check is what makes this case come out right: every tag here
// starts with "chain-work-h", and only one of them is a hop.
name: "a lone group hop resolves to the WRAPPER, never to a member copy",
tags: tagSet(
"direct", "n1", "n2",
"chain-work-h1",
"chain-work-h1-n1", "chain-work-h1-n2", "chain-work-h1-zzz",
),
want: "chain-work-h1",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, err := resolveViaTag("chain:work", tc.tags, nil)
if err != nil {
t.Fatalf("resolveViaTag: %v", err)
}
if got != tc.want {
t.Fatalf("entry = %q, want %q", got, tc.want)
}
})
}
}
// TestChainDetourNotBuiltRefusesInsteadOfDirect is the leak guard. A chain nothing
// else references is never materialised, so no wrapper exists. The resolver must
// say so by name — and must NOT resolve to "direct" or to a node/group that merely
// shares the chain's name.
func TestChainDetourNotBuiltRefusesInsteadOfDirect(t *testing.T) {
// "work" is ALSO a group tag here: the trap the unfixed code fell into, where
// engine.ViaToTag("chain:work") -> "work" hits an unrelated outbound.
tags := tagSet("direct", "block", "n1", "n2", "work")
chains := []model.Chain{{Name: "work", Hops: []string{"node:n1", "node:n2"}}}
got, err := resolveViaTag("chain:work", tags, chains)
if err == nil {
t.Fatalf("unbuilt chain resolved to %q, want a refusal", got)
}
if got != "" {
t.Fatalf("refusal also returned a tag %q; a caller could dial it", got)
}
if !errors.Is(err, engine.ErrOutboundUnknown) {
t.Fatalf("err = %v, want it to wrap engine.ErrOutboundUnknown (panel maps that to 400)", err)
}
for _, want := range []string{"chain \"work\"", "only built when an enabled rule"} {
if !strings.Contains(err.Error(), want) {
t.Fatalf("err %q does not explain the cause (missing %q)", err, want)
}
}
}
// TestChainDetourUndefinedIsNamed: a fetch_detour pointing at a chain that does not
// exist at all must say exactly that.
func TestChainDetourUndefinedIsNamed(t *testing.T) {
got, err := resolveViaTag("chain:ghost", tagSet("direct", "ghost"), nil)
if err == nil {
t.Fatalf("undefined chain resolved to %q, want a refusal", got)
}
if !errors.Is(err, engine.ErrOutboundUnknown) {
t.Fatalf("err = %v, want engine.ErrOutboundUnknown", err)
}
if !strings.Contains(err.Error(), "no chain named \"ghost\"") {
t.Fatalf("err %q does not name the missing chain", err)
}
}
// TestChainDetourSingleHop covers the shape the generator does NOT wrap: a chain
// that flattens to one hop IS that hop.
func TestChainDetourSingleHop(t *testing.T) {
tags := tagSet("direct", "block", "n1", "grp")
cases := []struct {
name string
chains []model.Chain
via string
want string
}{
{
name: "node hop",
chains: []model.Chain{{Name: "solo", Hops: []string{"node:n1"}}},
via: "chain:solo",
want: "n1",
},
{
name: "group hop",
chains: []model.Chain{{Name: "solo", Hops: []string{"group:grp"}}},
via: "chain:solo",
want: "grp",
},
{
name: "bare hop",
chains: []model.Chain{{Name: "solo", Hops: []string{"n1"}}},
via: "chain:solo",
want: "n1",
},
{
name: "sub-chain hop is spliced",
chains: []model.Chain{
{Name: "outer", Hops: []string{"chain:inner"}},
{Name: "inner", Hops: []string{"node:n1"}},
},
via: "chain:outer",
want: "n1",
},
{
name: "blank hops are ignored",
chains: []model.Chain{{Name: "solo", Hops: []string{"", "node:n1", " "}}},
via: "chain:solo",
want: "n1",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, err := resolveViaTag(tc.via, tags, tc.chains)
if err != nil {
t.Fatalf("resolveViaTag(%q): %v", tc.via, err)
}
if got != tc.want {
t.Fatalf("resolveViaTag(%q) = %q, want %q", tc.via, got, tc.want)
}
})
}
}
// TestChainDetourRefusals pins every remaining miss as a NAMED refusal with no tag,
// so none of them can degrade into a fetch through the wrong outbound.
func TestChainDetourRefusals(t *testing.T) {
tags := tagSet("direct", "block", "n1", "egress-wg0")
cases := []struct {
name string
chains []model.Chain
via string
wantMsg string
}{
{
name: "hop not in the running box",
chains: []model.Chain{{Name: "solo", Hops: []string{"node:gone"}}},
via: "chain:solo",
wantMsg: "holds no outbound for",
},
{
name: "egress entry hop alone has no exit",
chains: []model.Chain{{Name: "solo", Hops: []string{"egress:wg0"}}},
via: "chain:solo",
wantMsg: "no exit to fetch through",
},
{
name: "direct is not a tunnel hop",
chains: []model.Chain{{Name: "solo", Hops: []string{"direct"}}},
via: "chain:solo",
wantMsg: "terminal route target",
},
{
name: "block is not a tunnel hop",
chains: []model.Chain{{Name: "solo", Hops: []string{"block"}}},
via: "chain:solo",
wantMsg: "terminal route target",
},
{
name: "chain with no hops",
chains: []model.Chain{{Name: "solo", Hops: []string{" "}}},
via: "chain:solo",
wantMsg: "has no hops",
},
{
name: "chain: with no name",
chains: nil,
via: "chain:",
wantMsg: "names no chain",
},
{
name: "self-referencing chain terminates",
chains: []model.Chain{{Name: "loop", Hops: []string{"chain:loop"}}},
via: "chain:loop",
wantMsg: "reference cycle",
},
{
name: "mutual cycle terminates",
chains: []model.Chain{
{Name: "a", Hops: []string{"chain:b"}},
{Name: "b", Hops: []string{"chain:a"}},
},
via: "chain:a",
wantMsg: "reference cycle",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got, err := resolveViaTag(tc.via, tags, tc.chains)
if err == nil {
t.Fatalf("resolveViaTag(%q) = %q, want a refusal", tc.via, got)
}
if got != "" {
t.Fatalf("refusal also returned tag %q", got)
}
if got == "direct" {
t.Fatalf("refusal degraded to direct — that is a WAN leak")
}
if !errors.Is(err, engine.ErrOutboundUnknown) {
t.Fatalf("err = %v, want engine.ErrOutboundUnknown", err)
}
if !strings.Contains(err.Error(), tc.wantMsg) {
t.Fatalf("err %q does not contain %q", err, tc.wantMsg)
}
})
}
}
// TestChainDetourCaseInsensitive: engine.ViaToTag accepts "CHAIN:x"; so must this,
// or a spelling the engine would have honoured falls through as a bare tag.
func TestChainDetourCaseInsensitive(t *testing.T) {
tags := tagSet("chain-work-h1", "chain-work-h2")
for _, via := range []string{"chain:work", "Chain:work", "CHAIN:work", " chain:work "} {
got, err := resolveViaTag(via, tags, nil)
if err != nil {
t.Fatalf("resolveViaTag(%q): %v", via, err)
}
if got != "chain-work-h2" {
t.Fatalf("resolveViaTag(%q) = %q, want chain-work-h2", via, got)
}
}
}
// TestNonChainDetoursUnchanged is the regression guard: every other fetch_detour
// form must reach engine.ViaToTag byte-for-byte as configured, so group:/node:/
// egress:/direct/bare keep behaving exactly as before this resolver existed.
func TestNonChainDetoursUnchanged(t *testing.T) {
// A tag set that WOULD satisfy a chain lookup, to catch a resolver that starts
// treating everything as a chain name.
tags := tagSet("direct", "auto", "n1", "egress-wg0", "chain-auto-h1", "chain-n1-h1")
chains := []model.Chain{{Name: "auto", Hops: []string{"node:n1"}}}
for _, via := range []string{"", "direct", "DIRECT", "group:auto", "node:n1", "egress:wg0", "n1", " group:auto "} {
got, err := resolveViaTag(via, tags, chains)
if err != nil {
t.Fatalf("resolveViaTag(%q): %v", via, err)
}
if got != via {
t.Fatalf("resolveViaTag(%q) = %q, want it passed through unchanged for engine.ViaToTag", via, got)
}
}
}
// TestHTTPClientChainOnStoppedEngine: with no box running there are no tags to
// resolve against, and the caller must hear "engine stopped" — not a confusing
// claim about the chain, and not a client dialling anything.
func TestHTTPClientChainOnStoppedEngine(t *testing.T) {
a := New(engine.New(), nil)
for _, via := range []string{"chain:work", "group:auto", "direct", ""} {
c, err := a.HTTPClient(via)
if !errors.Is(err, engine.ErrEngineStopped) {
t.Fatalf("HTTPClient(%q): err = %v, want engine.ErrEngineStopped", via, err)
}
if c != nil {
t.Fatalf("HTTPClient(%q) returned a client %v on a stopped engine", via, c)
}
}
}
// TestHTTPClientNilEngine keeps the pre-existing nil-engine contract.
func TestHTTPClientNilEngine(t *testing.T) {
a := &Applier{}
if _, err := a.HTTPClient("chain:work"); !errors.Is(err, engine.ErrEngineStopped) {
t.Fatalf("err = %v, want engine.ErrEngineStopped", err)
}
}
+270
View File
@@ -0,0 +1,270 @@
package apply
// The hold state must describe THE ROUTER, not this process's memory of what it
// did.
//
// The boot armor (netplane/armor.go) put a fail-closed plane in the kernel that
// the Applier never installs: /etc/init.d/shater-armor loads it at START=21,
// before the daemon exists, and cmd/shaterd reinstates it when the config cannot
// be read. The latch behind Holding() knew nothing about either, so the daemon
// reported `holding=false` and `plane="full"` over a LAN that was blocked — and
// with an unreadable config that state is PERMANENT, because the only thing that
// ever wrote the latch was an apply and no apply can run.
//
// What comes out the other end is an alert (cmd/shaterd fireApplyFail) that
// chooses its wording from exactly this bool and tells the operator "Traffic is
// NOT being blocked" while it is. That is the inverted failure: not a fault
// hidden, but a fault invented — and the obvious response to it is to go and
// dismantle the protection that is doing its job.
import (
"errors"
"os"
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
)
// stubPlaneFacts pins the two outside-world seams — is our table in the kernel,
// is the persisted fail-closed plane on flash — for the duration of a test.
func stubPlaneFacts(t *testing.T, table, armor bool) func() {
t.Helper()
origTable, origArmor := tableExists, bootArmorPresent
tableExists = func() bool { return table }
bootArmorPresent = func() bool { return armor }
return func() { tableExists, bootArmorPresent = origTable, origArmor }
}
// TestHoldingSeesAPlaneThisProcessDidNotInstall is the core regression.
//
// All three cases share the same latch value (false — this Applier has installed
// nothing) and must still produce three different, correct answers.
func TestHoldingSeesAPlaneThisProcessDidNotInstall(t *testing.T) {
a := New(engine.New(), nil)
// (1) No table at all. Nothing is protecting the LAN and nothing may claim to.
restore := stubPlaneFacts(t, false, true)
if a.Holding() {
t.Errorf("Holding() = true with no table loaded")
}
if got := a.Status().Plane; got != "none" {
t.Errorf("Plane = %q with no table loaded, want \"none\"", got)
}
restore()
// (2) The boot armor's plane IS loaded and this engine never started. This is
// what /etc/init.d/shater-armor leaves behind at every boot, and what the
// daemon reinstates when /etc/config/shater cannot be read — the case where
// nothing else will ever correct the answer.
restore = stubPlaneFacts(t, true, true)
if !a.Holding() {
t.Errorf("Holding() = false while the boot armor's fail-closed plane is loaded and the " +
"engine is down — the LAN is blocked and the daemon says it is not; the apply-failure " +
"alert would send the operator to fix a protection that is working")
}
s := a.Status()
if s.Plane != "hold" {
t.Errorf("Plane = %q over a blocked LAN with a dead engine, want \"hold\" "+
"(\"full\" means traffic is diverted into a RUNNING engine)", s.Plane)
}
if !s.Table {
t.Errorf("Table = false although a table is loaded")
}
if s.Running || s.EngineRunning {
t.Errorf("running/engine_running must stay false while holding: %+v", s)
}
restore()
// (3) A table is loaded but there is NO armor on flash. refreshBootArmor
// removes that file exactly when the operator disables the stack or opens the
// kill switch, so what is loaded here is a leftover that drops nothing.
// Claiming a hold would be the new lie: it would tell someone who deliberately
// chose fail-open that their LAN is cut off.
restore = stubPlaneFacts(t, true, false)
if a.Holding() {
t.Errorf("Holding() = true with no fail-closed armor on flash — a leftover plane from a " +
"kill_switch=open config blocks nothing, and saying otherwise is the same lie inverted")
}
restore()
}
// TestSuccessfulApplyStillClearsTheHold: the self-healing path must survive the
// derived half. A read-time inference that ignored the engine would re-assert the
// hold one line after setHolding(false) and pin the router in "protected, not
// proxying" forever — with the table loaded and the armor on flash, which is the
// steady state of every healthy router.
func TestSuccessfulApplyStillClearsTheHold(t *testing.T) {
a := New(engine.New(), nil)
t.Cleanup(func() { _ = a.eng.Close() })
t.Cleanup(func() { _ = os.Remove(ActiveFlag) })
// A REAL started instance: the derived half asks the engine, so a stub would
// not exercise the thing under test.
if _, err := a.eng.Apply(mixedOn(18841)); err != nil {
t.Fatalf("bring a real engine up: %v", err)
}
if !a.eng.Running() {
t.Fatalf("precondition: the engine must be running")
}
// The steady state of a healthy router: our table loaded, armor on flash.
defer stubPlaneFacts(t, true, true)()
a.setHolding(true) // whatever put us on hold before this apply
if !a.Holding() {
t.Fatalf("precondition: the latch must read through")
}
m := holdModel("closed")
m.Globals.GroupHealth = false // no background probing from a unit test
defer stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return true, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
return planeOutcome{changed: true}, nil
})()
a.mu.Lock()
_, err := a.applyLocked(m)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyLocked: %v", err)
}
if a.Holding() {
t.Fatalf("a successful apply did not clear the hold — the router is proxying and the " +
"panel would still show \"protected, not proxying\"")
}
if got := a.Status().Plane; got != "full" {
t.Errorf("Plane = %q after a successful apply with the engine up, want \"full\"", got)
}
}
// TestApplyFailureOverAForeignPlaneReportsBlocked pins the input cmd/shaterd's
// fireApplyFail words its incident from.
//
// The shape is the real one: the engine will not start, so applyLocked calls
// holdLocked — and holdLocked cannot install anything either (nft refuses, the
// overlay is full). The latch therefore stays false. But the boot armor's plane
// is still standing in the kernel, so forwarded traffic IS being dropped, and the
// incident must say so. With only the latch, this is precisely where the daemon
// said "Traffic is NOT being blocked (kill switch is open)" on a router whose
// kill switch was closed and whose LAN was cut off.
func TestApplyFailureOverAForeignPlaneReportsBlocked(t *testing.T) {
a := New(engine.New(), nil)
defer stubPlaneFacts(t, true, true)()
origHold := applyHoldNft
applyHoldNft = func(string) error { return errors.New("nft -f (load) failed: no space left on device") }
defer func() { applyHoldNft = origHold }()
defer stubApplyStages(t,
func(*Applier, option.Options) (bool, error) {
return false, errors.New("start rule-set[geosite]: connection refused")
},
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
t.Errorf("the data-plane stage must not run after the engine stage failed")
return planeOutcome{}, nil
})()
a.mu.Lock()
_, err := a.applyLocked(holdModel("closed"))
a.mu.Unlock()
if err == nil {
t.Fatalf("applyLocked must surface the engine failure")
}
a.stateMu.RLock()
latched := a.holding
a.stateMu.RUnlock()
if latched {
t.Fatalf("precondition: holdLocked could not install a plane, so the latch must be false")
}
// fireApplyFail(notifier, err, applier.Holding()) — this bool picks between
// "forwarded LAN traffic is being dropped" and "Traffic is NOT being blocked".
if !a.Holding() {
t.Fatalf("Holding() = false while a fail-closed plane blocks the LAN: the incident would " +
"read \"Traffic is NOT being blocked\" at the exact moment it is being blocked")
}
}
// TestArmHoldReplacesAPlaneItDidNotInstall: ArmHold used to return early on
// TableExists() and publish nothing, which is how the boot armor's plane came to
// be loaded with holding=false. It now installs its own — rendered from the model
// it was handed, so the state it publishes is a fact about something this process
// did rather than a guess about something it found (and the fresh render also
// replaces a snapshot that may predate an interface rename).
func TestArmHoldReplacesAPlaneItDidNotInstall(t *testing.T) {
loaded := withHoldProbe(t)
defer stubPlaneFacts(t, true, true)() // a table is ALREADY loaded
a := New(engine.New(), nil)
a.ArmHold(holdModel("closed"))
if len(*loaded) != 1 {
t.Fatalf("ArmHold loaded %d rulesets over an existing table, want 1 — it deferred to a "+
"table it cannot inspect instead of installing one it can vouch for", len(*loaded))
}
if !a.Holding() {
t.Errorf("ArmHold installed the holding plane but did not publish it")
}
if got := a.Status().Plane; got != "hold" {
t.Errorf("Plane = %q after ArmHold, want \"hold\"", got)
}
}
// TestArmHoldStillRespectsFailOpen: replacing a foreign table must not become a
// licence to install a plane the operator did not ask for. kill_switch=open and
// globals.enabled=0 are explicit choices and ArmHold must keep obeying both.
func TestArmHoldStillRespectsFailOpen(t *testing.T) {
loaded := withHoldProbe(t)
defer stubPlaneFacts(t, true, false)()
a := New(engine.New(), nil)
a.ArmHold(holdModel("open"))
if len(*loaded) != 0 {
t.Errorf("kill_switch=open loaded %d rulesets, want 0", len(*loaded))
}
disabled := holdModel("closed")
disabled.Globals.Enabled = false
a.ArmHold(disabled)
if len(*loaded) != 0 {
t.Errorf("globals.enabled=0 loaded %d rulesets, want 0", len(*loaded))
}
a.ArmHold(nil)
if len(*loaded) != 0 {
t.Errorf("a nil model loaded %d rulesets, want 0", len(*loaded))
}
if a.Holding() {
t.Errorf("nothing was installed, so nothing may be reported as holding")
}
}
// TestForeignHoldMatrix pins the predicate itself, including the two facts it
// refuses to guess about.
func TestForeignHoldMatrix(t *testing.T) {
for _, tc := range []struct {
name string
table, engine, armor bool
want bool
}{
{"no table", false, false, true, false},
{"engine up: the table is the working full plane", true, true, true, false},
{"armor on flash, engine down: blocked", true, false, true, true},
{"no armor: the last config was disabled or fail-open", true, false, false, false},
{"nothing at all", false, false, false, false},
} {
t.Run(tc.name, func(t *testing.T) {
defer stubPlaneFacts(t, tc.table, tc.armor)()
if got := foreignHold(tc.table, tc.engine); got != tc.want {
t.Errorf("foreignHold(table=%v, engineUp=%v) = %v, want %v",
tc.table, tc.engine, got, tc.want)
}
})
}
}
+576
View File
@@ -0,0 +1,576 @@
package apply
// Regression tests for two more ways this package reported calm over a router that
// was not doing what its config said. Both are the INVERTED failure — not an error
// raised when things are fine, but silence (or an amber "you switched it off")
// while something is really wrong — which is the only kind that gets someone hurt.
//
// 1. a CRITICAL policy-routing finding published by one apply and erased by the
// next no-op reconcile, sixty seconds later, while its cause stood;
// 2. a configuration that cannot be READ published as enabled=false, i.e. as the
// owner's own choice, while the fail-closed plane had the LAN cut off.
import (
"errors"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// The package's tests are written against a ROUTER-LIKE BASELINE, in which the
// configuration can be read. That baseline used to be supplied by accident: on a
// build host there is no `uci`, model.ReadUCI failed on every Status() call, and
// the failure was silent — which is defect 2 itself. Now that the failure is
// published, leaving it in place would make every test in this package assert
// against a router whose configuration is unreadable, which is nobody's intended
// fixture and would have masked the real subject of several of them.
//
// So the baseline is made explicit here rather than left to the absence of a
// binary. A test that wants the UNREADABLE case sets readConfig itself and restores
// it (see TestStatusSaysWhenTheConfigCannotBeRead) — which is also the only way to
// exercise that branch deterministically on a router, where `uci` does exist.
//
// An init() rather than a TestMain on purpose: init functions compose, so another
// test file in this package can add its own without a conflict.
func init() {
readConfig = func() (*model.Model, error) { return &model.Model{Globals: model.DefaultGlobals()}, nil }
}
// noGatewayFinding is netplane/apply.go's real text for the fault this is all
// about, shortened to its load-bearing clause: an egress that was BUILT, reports as
// applied everywhere in the UI, and cannot carry a packet off its own subnet.
// Quoted rather than invented so the test breaks if that warning is ever reworded
// out of the critical class.
const noGatewayFinding = `egress "wan2": no IPv4 gateway could be found for interface "wan2" (device eth1) ` +
`by any means, so its routing table sends traffic straight onto the local segment. The device is not ` +
`point-to-point, so this egress CANNOT REACH ANYTHING outside its own subnet — every node, group and ` +
`rule bound to it will fail to connect.`
// findFinding reports whether the published warning set still carries the finding,
// and at what severity.
func findFinding(ws []Warning, substr string) (Warning, bool) {
for _, w := range ws {
if strings.Contains(w.Message, substr) {
return w, true
}
}
return Warning{}, false
}
// TestStandingRouteWarningSurvivesTheFastPath is the defect-1 regression, stated in
// the terms the owner experiences it.
//
// A second WAN comes up without a gateway. The apply that notices publishes a
// critical finding. A minute later cron reconciles; nothing has changed, so the
// data plane takes its fast path and measures no routing — and applyLocked then
// published `collectWarnings(..., plane.routeWarnings, ...)` UNCONDITIONALLY, with
// routeWarnings nil, replacing the set with one that no longer contained the
// finding. Neither of the two findings that survive to the fast path (no gateway,
// no fail-closed floor) makes RoutingPresent false, so nothing ever brought it
// back: the panel showed zero findings, plane full, green, over an egress carrying
// nothing — or over a routing table whose next ifdown is a leak onto the plain WAN.
//
// THE CONTROL IS THE POINT. Pass 1 proves this instrument can see a live finding at
// all, and pass 3 proves it can see one DISAPPEAR. Without both, "the warning is
// still there after the fast path" would be satisfied by an instrument that always
// says yes, and would equally be satisfied by turning the fix into a latch — which
// is the same defect pointed the other way.
func TestStandingRouteWarningSurvivesTheFastPath(t *testing.T) {
a := New(engine.New(), nil)
m := holdModel("closed")
// The plane stage is driven directly: what is under test is what applyLocked
// PUBLISHES for a given plane outcome, and the outcome itself is produced by
// TestApplyDataPlaneFastPathTakesNoRoutingMeasurement below.
var outcome planeOutcome
restore := stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return false, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
return outcome, nil
})
defer restore()
apply := func(t *testing.T, o planeOutcome) []Warning {
t.Helper()
outcome = o
a.mu.Lock()
_, err := a.applyLocked(m)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyLocked: %v", err)
}
return a.Warnings()
}
// (1) CONTROL — the instrument can see a live finding. The routing was measured
// and it found the egress unable to route.
var measured planeOutcome
measured.measuredRouting([]string{noGatewayFinding})
ws := apply(t, measured)
w, ok := findFinding(ws, "CANNOT REACH ANYTHING outside its own subnet")
if !ok {
t.Fatalf("control failed: a MEASURED route finding is not published at all, so this test "+
"could not detect its loss either; warnings = %+v", ws)
}
if w.Severity != SeverityCritical {
t.Errorf("an egress that cannot reach off its own subnet is critical, got %q", w.Severity)
}
// (2) THE DEFECT — the cron reconcile a minute later. Nothing changed, the fast
// path took no measurement, and the fault is still standing.
ws = apply(t, planeOutcome{}) // routeMeasured false: nothing was checked
if _, ok := findFinding(ws, "CANNOT REACH ANYTHING outside its own subnet"); !ok {
t.Errorf("the fast path erased a CRITICAL finding whose cause is still standing: "+
"the panel now shows a clean, green router over an egress that carries nothing. "+
"warnings = %+v", ws)
}
// (3) CONTROL — the instrument can see the finding go away. The cause was fixed
// (a gateway appeared), something forced a real rebuild, and the routing was
// measured again with nothing to report. A fix that merely latched the warning
// would fail here.
var clean planeOutcome
clean.measuredRouting(nil)
ws = apply(t, clean)
if w, ok := findFinding(ws, "CANNOT REACH ANYTHING outside its own subnet"); ok {
t.Errorf("a re-MEASURED clean routing state must retire the finding, still published: %+v", w)
}
// And the erasure is not permanent either: a later measurement that finds it
// again publishes it again.
ws = apply(t, measured)
if _, ok := findFinding(ws, "CANNOT REACH ANYTHING outside its own subnet"); !ok {
t.Errorf("a finding that recurs must be published again; warnings = %+v", ws)
}
}
// TestStandingRouteWarningSurvivesAPostSwapAbort pins the same property on the
// other publishing path. abortAfterSwap republishes the plane outcome with its own
// critical entry in front; if it were handed the fast path's empty routeWarnings it
// would drop the standing finding exactly like the success path did.
func TestStandingRouteWarningSurvivesAPostSwapAbort(t *testing.T) {
a := New(engine.New(), nil)
m := holdModel("closed")
var outcome planeOutcome
var stageErr error
restore := stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return true, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
return outcome, stageErr
})
defer restore()
// A successful apply records the standing finding.
outcome.measuredRouting([]string{noGatewayFinding})
a.mu.Lock()
_, err := a.applyLocked(m)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyLocked (seed): %v", err)
}
// Now the sysctl stage fails on a pass that fast-pathed the routing.
outcome = planeOutcome{stage: "setting the kernel sysctls"}
stageErr = errors.New("sysctl: read-only file system")
a.mu.Lock()
_, err = a.applyLocked(m)
a.mu.Unlock()
if err == nil {
t.Fatalf("applyLocked must surface the netplane failure")
}
ws := a.Warnings()
if _, ok := findFinding(ws, "could NOT be completed"); !ok {
t.Errorf("the abort must still say the data plane is incomplete; warnings = %+v", ws)
}
if _, ok := findFinding(ws, "CANNOT REACH ANYTHING outside its own subnet"); !ok {
t.Errorf("a half-installed plane must not also erase the standing routing finding; "+
"warnings = %+v", ws)
}
}
// TestStandingRouteWarningsClearedOnTeardown: teardown removes the ip rules and
// tables those findings are ABOUT, so the last measurement stops describing
// anything and must not be carried into the next apply.
func TestStandingRouteWarningsClearedOnTeardown(t *testing.T) {
a := New(engine.New(), nil)
m := holdModel("closed")
var outcome planeOutcome
outcome.measuredRouting([]string{noGatewayFinding})
restore := stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return false, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
return outcome, nil
})
defer restore()
a.mu.Lock()
if _, err := a.applyLocked(m); err != nil {
a.mu.Unlock()
t.Fatalf("applyLocked: %v", err)
}
a.mu.Unlock()
if len(a.lastRouteWarnings) == 0 {
t.Fatalf("control failed: the measurement was not remembered, so this test cannot show it cleared")
}
// Teardown with everything stubbed out: nothing here may touch a real kernel.
origTeardownNft, origTable := teardownNft, tableExists
teardownNft = func() error { return nil }
tableExists = func() bool { return false }
defer func() { teardownNft, tableExists = origTeardownNft, origTable }()
if err := a.Teardown(); err != nil {
t.Fatalf("Teardown: %v", err)
}
if len(a.lastRouteWarnings) != 0 {
t.Errorf("teardown removed the rules and tables those findings describe, but kept them: %+v",
a.lastRouteWarnings)
}
// The next apply's fast path must therefore publish nothing, not the findings of
// a plane that no longer exists.
outcome = planeOutcome{}
a.mu.Lock()
if _, err := a.applyLocked(m); err != nil {
a.mu.Unlock()
t.Fatalf("applyLocked after teardown: %v", err)
}
a.mu.Unlock()
if w, ok := findFinding(a.Warnings(), "CANNOT REACH ANYTHING outside its own subnet"); ok {
t.Errorf("a finding about a torn-down plane is still published: %+v", w)
}
}
// TestApplyDataPlaneFastPathTakesNoRoutingMeasurement is the other half of the
// instrument, and without it the tests above prove nothing about production: they
// drive applyDataPlane through its seam, so they would pass just as happily if the
// REAL data-plane stage marked its fast path as a measurement.
//
// It exercises applyDataPlaneLocked itself, with the ruleset pre-loaded so the nft
// fast path is taken and only the routing decision varies.
func TestApplyDataPlaneFastPathTakesNoRoutingMeasurement(t *testing.T) {
m := holdModel("closed")
now := time.Now()
opts := option.Options{}
a := New(engine.New(), nil)
// Reproduce exactly what the stage will render, so `ruleset == a.lastNft` holds
// and ApplyNft is never reached (there is no nft binary here, and a test must not
// load a ruleset into the host it runs on).
ruleset, _, err := netplane.RenderNftPlanAt(m, a.untunnelablePlanFor(m, opts), now)
if err != nil {
t.Fatalf("RenderNftPlanAt: %v", err)
}
origTable := tableExists
origPresent, origApplyRouting := routingPresent, applyRoutingWithWarnings
origSysctl, origIface := applySysctl, applyIfaceSysctlsAt
tableExists = func() bool { return true }
applySysctl = func() error { return nil }
applyIfaceSysctlsAt = func(*model.Model, time.Time) error { return nil }
defer func() {
tableExists = origTable
routingPresent, applyRoutingWithWarnings = origPresent, origApplyRouting
applySysctl, applyIfaceSysctlsAt = origSysctl, origIface
}()
tests := []struct {
name string
present bool
wantCalled bool
wantMeasured bool
}{
{
name: "routing intact: the fast path measures nothing",
present: true,
wantCalled: false,
wantMeasured: false,
},
{
// The CONTROL for the case above: the same stage, one input flipped, does
// call the routing and does report a measurement. A stage that never
// measured anything would satisfy the first row on its own.
name: "routing missing: the full path measures",
present: false,
wantCalled: true,
wantMeasured: true,
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
called := false
routingPresent = func(model.Globals) bool { return tc.present }
applyRoutingWithWarnings = func(*model.Model) ([]string, error) {
called = true
return []string{noGatewayFinding}, nil
}
a.lastNft = ruleset
a.mu.Lock()
out, err := a.applyDataPlaneLocked(m, opts, now)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyDataPlaneLocked: %v (stage %q)", err, out.stage)
}
// Guards on the setup itself: if the nft fast path was NOT taken, the run
// below says nothing about the routing decision.
if out.stage != "" || out.changed {
t.Fatalf("setup broken: the nft fast path was not taken (stage %q, changed %v)",
out.stage, out.changed)
}
if called != tc.wantCalled {
t.Errorf("ApplyRoutingWithWarnings called = %v, want %v", called, tc.wantCalled)
}
if out.routeMeasured != tc.wantMeasured {
t.Errorf("routeMeasured = %v, want %v — a pass that took no measurement must not "+
"report one, or applyLocked will publish 'nothing found' as if something had looked",
out.routeMeasured, tc.wantMeasured)
}
if tc.wantMeasured && len(out.routeWarnings) != 1 {
t.Errorf("a measured pass must carry what it found, got %+v", out.routeWarnings)
}
})
}
}
// --- defect 3: the engine is down and nothing says why -----------------------
// TestEngineStartFailureIsPublished is the defect-3 regression.
//
// holdLocked wrote the cause of a failed engine start to the LOG and nowhere else:
// it sets traffic to unknown, installs the fail-closed plane, and never calls
// setWarnings. Status.Warnings carries the findings of the last SUCCESSFUL apply,
// and an engine that never started produced none — so `plane: "hold"` with an empty
// findings list was a normal, reachable state of the product. The panel correctly
// said "traffic is blocked, the tunnel is down, fix it from here", and nothing
// anywhere said WHAT to fix. The two workarounds were pressing Apply again to read
// the POST's error body, and reading a log that lives in tmpfs by default — gone
// after the reboot the owner sat down to investigate.
func TestEngineStartFailureIsPublished(t *testing.T) {
a := New(engine.New(), nil)
m := holdModel("closed")
// CONTROL — nothing has failed yet, so nothing is claimed. Without this, an
// implementation that warns unconditionally would satisfy every check below.
if w, ok := findFinding(a.Warnings(), "could NOT be started"); ok {
t.Fatalf("control failed: an engine that has not failed must not be reported as failed: %+v", w)
}
// The holding plane is installed without touching a real kernel.
origHold := applyHoldNft
applyHoldNft = func(string) error { return nil }
defer func() { applyHoldNft = origHold }()
cause := errors.New("start outbound/vless[node-tokyo]: parse server address: invalid IP")
a.mu.Lock()
a.holdLocked(m, cause)
a.mu.Unlock()
if !a.holding {
t.Fatalf("precondition: holdLocked must have installed the holding plane")
}
w, ok := findFinding(a.Warnings(), "could NOT be started")
if !ok {
t.Fatalf("the engine is down and the LAN is held, and NOTHING says why; warnings = %+v",
a.Warnings())
}
if w.Severity != SeverityCritical {
t.Errorf("a router that cannot proxy anything is critical, got %q", w.Severity)
}
if w.Section != "engine" {
t.Errorf("the finding must be attributable to the engine, got section %q", w.Section)
}
if !strings.Contains(w.Message, "parse server address: invalid IP") {
t.Errorf("the finding must carry the CAUSE verbatim — that is the whole point: %q", w.Message)
}
// It must reach Status, which is what the panel actually reads.
origTable := tableExists
tableExists = func() bool { return true }
defer func() { tableExists = origTable }()
s := a.Status()
if s.Plane != "hold" {
t.Fatalf("precondition: plane = %q, want \"hold\" — the state this defect is about", s.Plane)
}
if _, ok := findFinding(s.Warnings, "could NOT be started"); !ok {
t.Errorf("plane=hold with an EMPTY findings list is the defect; warnings = %+v", s.Warnings)
}
// CONTROL — it is a CONDITION, not a log feed. With the same cause still on
// record, an engine that is up must publish nothing: the entry is gated on the
// engine being down right now, which is what makes it self-clearing rather than a
// latch somebody has to remember to reset. (Driven through warningsWith because
// engine liveness is unexported state; Status feeds it the same fact it puts in
// `running`, so the two can never contradict each other.)
if w, ok := findFinding(a.warningsWith(true), "could NOT be started"); ok {
t.Errorf("with the engine UP, a past start failure is not a current condition: %+v", w)
}
if _, ok := findFinding(a.warningsWith(false), "could NOT be started"); !ok {
t.Errorf("control failed: warningsWith(false) must still show it, or the check above proves nothing")
}
// And the record itself is dropped by a successful apply, so a LATER engine death
// cannot resurrect this reason and present it as the current one.
restore := stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return true, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
var out planeOutcome
out.measuredRouting(nil)
return out, nil
})
defer restore()
a.mu.Lock()
_, err := a.applyLocked(m)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyLocked (recovery): %v", err)
}
if w, ok := findFinding(a.Warnings(), "could NOT be started"); ok {
t.Errorf("a start failure must retire itself once the engine runs, still published: %+v", w)
}
}
// TestEngineNotStartedYetIsNotAnAlarm: the boot-time arm (ArmHold) goes through the
// same holdLocked, and the few seconds between the daemon starting and the engine
// coming up are a healthy start-up, not a fault. The panel's alarm banner lights on
// critical and on nothing else, so spending it here would teach the owner that the
// banner means nothing — and the next one, about a node that really cannot be
// parsed, is the one they would not read.
func TestEngineNotStartedYetIsNotAnAlarm(t *testing.T) {
a := New(engine.New(), nil)
m := holdModel("closed")
origHold := applyHoldNft
applyHoldNft = func(string) error { return nil }
defer func() { applyHoldNft = origHold }()
a.ArmHold(m)
w, ok := findFinding(a.Warnings(), "has not started yet")
if !ok {
t.Fatalf("the boot arm blocks the LAN and must still say why; warnings = %+v", a.Warnings())
}
if w.Severity != SeverityWarning {
t.Errorf("a normal start-up must not spend the critical banner, got %q", w.Severity)
}
}
// TestEngineDownCauseClearedOnTeardown: a `stop` is the operator's own decision, and
// a start failure left standing would present it as a malfunction for as long as the
// process lives to answer status calls.
func TestEngineDownCauseClearedOnTeardown(t *testing.T) {
a := New(engine.New(), nil)
origHold, origTeardown, origTable := applyHoldNft, teardownNft, tableExists
applyHoldNft = func(string) error { return nil }
teardownNft = func() error { return nil }
tableExists = func() bool { return false }
defer func() { applyHoldNft, teardownNft, tableExists = origHold, origTeardown, origTable }()
a.mu.Lock()
a.holdLocked(holdModel("closed"), errors.New("port 12345 already in use"))
a.mu.Unlock()
if _, ok := findFinding(a.Warnings(), "could NOT be started"); !ok {
t.Fatalf("control failed: the cause was not published, so this test cannot show it cleared")
}
if err := a.Teardown(); err != nil {
t.Fatalf("Teardown: %v", err)
}
if w, ok := findFinding(a.Warnings(), "could NOT be started"); ok {
t.Errorf("after a deliberate teardown the engine is down BY REQUEST, not by failure: %+v", w)
}
}
// --- defect 2: an unreadable configuration -----------------------------------
// TestStatusSaysWhenTheConfigCannotBeRead is the defect-2 regression.
//
// `if m, err := model.ReadUCI(); err == nil { ... }` left Enabled/KillSwitch/
// PanelPort at their zero values and recorded NOWHERE that the read had failed. The
// panel checks `!status.enabled` before it looks at `plane` and renders "Turned
// off" — amber, alarm:false, "turn the service on in Settings" — so the exact
// situation the boot armor exists for (a full /overlay, a `uci commit` caught
// half-written) came out as the owner's own choice, with the whole LAN cut off and
// the suggested remedy pointing at a settings page backed by the same unreadable
// file.
func TestStatusSaysWhenTheConfigCannotBeRead(t *testing.T) {
a := New(engine.New(), nil)
origRead, origTable := readConfig, tableExists
tableExists = func() bool { return true } // a plane IS loaded: the LAN is being held
defer func() { readConfig, tableExists = origRead, origTable }()
// CONTROL — the readable case. Without it, every assertion below is satisfied by
// a status that reports "unreadable" unconditionally.
m := holdModel("closed")
m.Globals.PanelPort = 8443
readConfig = func() (*model.Model, error) { return m, nil }
s := a.Status()
if !s.ConfigReadable {
t.Fatalf("control failed: a successful read must report config_readable=true")
}
if !s.Enabled || s.KillSwitch != "closed" || s.PanelPort != 8443 {
t.Fatalf("control failed: a successful read must publish the config: %+v", s)
}
if s.ConfigError != "" {
t.Errorf("config_error must be empty on a successful read, got %q", s.ConfigError)
}
if w, ok := findFinding(s.Warnings, "configuration could NOT be read"); ok {
t.Errorf("a readable config must not warn about being unreadable: %+v", w)
}
// THE DEFECT — the read fails.
readConfig = func() (*model.Model, error) { return nil, errors.New("uci export shater: exit status 1") }
s = a.Status()
if s.ConfigReadable {
t.Errorf("config_readable = true after a failed read")
}
if !strings.Contains(s.ConfigError, "exit status 1") {
t.Errorf("config_error must carry the reason, got %q", s.ConfigError)
}
// enabled is still false — it has no honest value to take — so the ONLY thing
// standing between the owner and "Turned off" is that this is distinguishable.
if s.Enabled {
t.Errorf("enabled must not be invented on a failed read")
}
if s.ConfigReadable == !s.Enabled {
// i.e. false == true; guards against a future refactor that makes
// config_readable track enabled and stops distinguishing anything.
t.Errorf("config_readable must be an independent fact from enabled")
}
w, ok := findFinding(s.Warnings, "configuration could NOT be read")
if !ok {
t.Fatalf("an unreadable configuration must be in the warning list — that is the only "+
"channel every consumer already renders; warnings = %+v", s.Warnings)
}
if w.Severity != SeverityCritical {
t.Errorf("an unreadable configuration is critical, got %q", w.Severity)
}
if w.Section != "config" || w.Name != "unreadable" {
t.Errorf("the warning must be attributable (section/name), got %q/%q", w.Section, w.Name)
}
// The one sentence that keeps the owner from "fixing" the fail-closed plane by
// switching the protection off.
if !strings.Contains(w.Message, "NOT the service being switched off") {
t.Errorf("the warning must say the block is not the service being off: %q", w.Message)
}
// And it self-clears: computed at read time, so the next successful read drops it
// without anything having to remember to.
readConfig = func() (*model.Model, error) { return m, nil }
if s = a.Status(); !s.ConfigReadable {
t.Errorf("config_readable must go true again the moment the config is readable")
}
if w, ok := findFinding(s.Warnings, "configuration could NOT be read"); ok {
t.Errorf("the unreadable warning must clear itself, still published: %+v", w)
}
}
+149
View File
@@ -0,0 +1,149 @@
package apply
// The exit path must never leave the LAN uncovered, and "never" is an ORDER, not
// an intention.
//
// WHAT THIS PINS. The daemon's SIGTERM path used to be:
//
// applier.Teardown() // netplane.TeardownNft() -> `nft delete table inet shater`
// armOnExit(handoff) // then, separately, install the fail-closed holding plane
//
// Two nft transactions. Between them the `inet shater` table does not exist, so
// fw4's `lan -> wan ACCEPT` is the only policy on the box and every forwarded LAN
// packet leaves in the clear. MEASURED on the stand at 80-90 ms, reproduced twice
// with a 35 000-sample run at ~1.3 ms resolution — and this is not a boot-time
// window that heals itself: it is every `restart`, every `reload_service` (i.e.
// every LuCI Save & Apply), and every package upgrade.
//
// The old call site even carried a comment explaining the mechanism — "AFTER the
// teardown, never before: Teardown deletes the table, so a plane installed first
// would simply be removed again" — and drew the wrong conclusion from a correct
// observation. The fix is not to arm later, it is to stop deleting: arm first (a
// single `nft -f` that opens with `delete table` and closes with the new one, so
// the kernel replaces rather than removes), then skip the delete.
//
// So the property under test is a SEQUENCE, and the test records the order the
// two seams are called in. A test that only asserted "the table still exists at
// the end" would pass against the broken code.
import (
"errors"
"testing"
"github.com/sagernet/sing-box/shater/engine"
)
// recordTeardownSeams captures the order in which the exit path touches the
// kernel: "arm" when a holding plane is installed, "delete" when the table is
// removed.
func recordTeardownSeams(t *testing.T) (*[]string, func()) {
t.Helper()
var calls []string
orig := teardownNft
teardownNft = func() error {
calls = append(calls, "delete")
return nil
}
return &calls, func() { teardownNft = orig }
}
// TestTeardownExitingReplacesThePlaneInsteadOfRemovingIt is the regression: with a
// successor coming, the table must be swapped and NEVER deleted.
func TestTeardownExitingReplacesThePlaneInsteadOfRemovingIt(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.TeardownExiting(func() bool {
*calls = append(*calls, "arm")
return true
}); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "arm" {
t.Fatalf("exit path did %v, want exactly [arm]: the holding plane must be installed "+
"and the table must NOT be deleted — a `delete` here is the 80-90 ms window in which "+
"fw4's lan->wan ACCEPT is the only policy on the box", *calls)
}
if !a.Holding() && tableExists == nil {
t.Errorf("unreachable; keeps the linter honest about the seam")
}
}
// TestTeardownExitingArmsBeforeItTearsDown pins the ORDER even in the case where
// the table does still get removed. Arming has to be the first thing that touches
// the kernel; if it ran after the delete we would be back to the two-transaction
// gap with extra steps.
func TestTeardownExitingArmsBeforeItTearsDown(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
// arm reports FALSE: nothing was installed (a render failure, or kill_switch=open
// where fail-open is the operator's documented choice). The table must then come
// down exactly as it always did.
if err := a.TeardownExiting(func() bool {
*calls = append(*calls, "arm")
return false
}); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if len(*calls) != 2 || (*calls)[0] != "arm" || (*calls)[1] != "delete" {
t.Fatalf("exit path did %v, want [arm delete]: arming must precede the delete, and a "+
"plane that was NOT installed must not keep the table alive", *calls)
}
}
// TestTeardownStillRemovesEverything is the escape hatch. A deliberate `stop`, and
// a Reconcile that finds the stack disabled, both come through the plain Teardown
// and must dismantle the plane completely — a kill switch that cannot be switched
// off is a brick.
func TestTeardownStillRemovesEverything(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.Teardown(); err != nil {
t.Fatalf("Teardown: %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "delete" {
t.Fatalf("Teardown did %v, want [delete]: an operator's stop must take the table with it", *calls)
}
}
// TestTeardownExitingReportsHolding pins the honesty half: with a plane left
// standing the Applier must not go on saying it installed nothing. Status can
// still be read over the control socket between the swap and the exit, and
// "holding=false over a blocked LAN" is the inverted lie holdstate_test.go is
// about, just reached down a different path.
func TestTeardownExitingReportsHolding(t *testing.T) {
_, restore := recordTeardownSeams(t)
defer restore()
restoreFacts := stubPlaneFacts(t, true, true)
defer restoreFacts()
a := New(engine.New(), nil)
if err := a.TeardownExiting(func() bool { return true }); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if !a.Holding() {
t.Errorf("Holding() = false right after the exit path left a fail-closed plane standing")
}
}
// TestTeardownExitingSurvivesANilArm keeps the plain-Teardown contract explicit:
// a nil arm is "nothing to install", not a panic.
func TestTeardownExitingSurvivesANilArm(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.TeardownExiting(nil); err != nil && !errors.Is(err, nil) {
t.Fatalf("TeardownExiting(nil): %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "delete" {
t.Fatalf("TeardownExiting(nil) did %v, want [delete]", *calls)
}
}
+197
View File
@@ -0,0 +1,197 @@
package apply
// The panel may not tell an operator that traceroute works.
//
// It did, in four different ways, for as long as the untunnelable notes have
// existed: "on Linux and macOS traceroute sends UDP probes instead, which still
// follow your rules", "which ARE tunnelled — the hops it prints are the tunnel's
// path", and twice "Ping and traceroute work everywhere". Measured on the
// production router, a plain `traceroute` from a LAN device prints `* * *` and
// nothing else — under every rung of the ladder, `direct` included, with the L3
// ingress on or off, with the kill switch open or closed. There is no mechanism
// that could print a hop: the UDP probe is delivered LOCALLY to the engine by
// tproxy, local delivery is not forwarding, so the TTL is never decremented and
// no router on the path is provoked into a `time-exceeded`.
//
// That makes the old texts the most expensive kind of wrong: an operator who
// reads "traceroute works" over a screen of stars goes looking for a fault in
// their own network, and there is none to find.
//
// This file is a RATCHET, not a prose test. It walks every branch of
// untunnelablePolicyWarnings and asserts two things about each note:
//
// 1. the retired sentences never come back, in any branch;
// 2. a note that mentions traceroute/tracert at all carries the shared
// udpTracerouteFacts verbatim — so a future edit cannot keep the claim and
// drop the correction, and cannot fork the wording into a second version.
//
// The matrix is exhaustive over the four inputs that select a branch, so a new
// branch added without the facts fails here rather than shipping.
import (
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// retiredTracerouteClaims are the exact phrases that were false. Each is quoted
// from the text it replaced, with the branch it lived in, so a reviewer can see
// this list is a record of what was actually said rather than a guess at what
// someone might say.
var retiredTracerouteClaims = []string{
// the untunnelable_egress branch, L3 off
"which still follow your rules",
// the `block` branch
"which ARE tunnelled — the hops it prints are the tunnel's path",
// `direct`, `icmp`, and the kill-switch-open note
"Ping and traceroute work",
"ping, traceroute and raw VPN passthrough",
// the shape of the claim, not one phrasing of it: any promise that the
// UDP-probe default follows the routing rules or shows a path.
"traceroute sends UDP probes instead",
}
// tracerouteMentions is what makes rule 2 above bite. A note that talks about
// tracing at all has taken on the duty to say what a plain `traceroute` does.
func tracerouteMentions(msg string) bool {
return strings.Contains(msg, "traceroute") || strings.Contains(msg, "tracert")
}
// untunnelableCases is the exhaustive cross-product of the inputs that pick a
// branch in untunnelablePolicyWarnings: the three policy rungs (plus one
// unrecognised value, which EffectiveUntunnelable folds into `block`), the kill
// switch, the L3 ingress, and untunnelable_egress.
func untunnelableCases() []model.Globals {
var out []model.Globals
for _, policy := range []string{
netplane.UntunnelableBlock,
netplane.UntunnelableICMP,
netplane.UntunnelableDirect,
"", // a config written before the option existed
} {
for _, kill := range []string{"closed", "open"} {
for _, l3 := range []bool{true, false} {
for _, egress := range []string{"", "wan2"} {
g := model.DefaultGlobals()
g.Untunnelable = policy
g.KillSwitch = kill
g.L3Tunnel = l3
g.UntunnelableEgress = egress
out = append(out, g)
}
}
}
}
return out
}
func caseLabel(g model.Globals) string {
return "policy=" + netplane.EffectiveUntunnelable(g) +
" kill=" + g.KillSwitch +
" l3=" + map[bool]string{true: "on", false: "off"}[g.L3Tunnel] +
" egress=" + map[bool]string{true: g.UntunnelableEgress, false: "-"}[g.UntunnelableEgress != ""]
}
// TestNoNoteClaimsPlainTracerouteWorks is the ratchet described at the top.
func TestNoNoteClaimsPlainTracerouteWorks(t *testing.T) {
seenMentions := 0
for _, g := range untunnelableCases() {
label := caseLabel(g)
for _, w := range untunnelablePolicyWarnings(g, nil) {
for _, claim := range retiredTracerouteClaims {
if strings.Contains(w.Message, claim) {
t.Errorf("%s: the untunnelable note says %q again.\n"+
"That claim was measured false on the production router: a plain "+
"`traceroute` prints `* * *` and no hops under every setting here. "+
"Say what udpTracerouteFacts says, or say nothing about traceroute.\n"+
"note: %s", label, claim, w.Message)
}
}
if !tracerouteMentions(w.Message) {
continue
}
seenMentions++
if !strings.Contains(w.Message, udpTracerouteFacts) {
t.Errorf("%s: the untunnelable note talks about tracing but does not carry "+
"udpTracerouteFacts.\n"+
"A note that mentions traceroute/tracert has taken on the duty to say that "+
"the Linux/macOS default prints no hops at all, why (local delivery is not "+
"forwarding), and what to use instead (`traceroute -I`). Append the shared "+
"constant rather than re-writing it — netplane/untunnelable.go states the "+
"same fact and the two must not fork.\nnote: %s", label, w.Message)
}
}
}
// Control. Without this, deleting every mention of tracing from every branch
// would leave a silent green test that proves nothing — the same instrument
// that returns "no lies found" when it cannot see a lie.
if seenMentions == 0 {
t.Fatal("no untunnelable note mentioned traceroute or tracert in the whole matrix — " +
"this test then asserts nothing at all. Either the notes stopped talking about " +
"ping diagnostics entirely, or untunnelablePolicyWarnings is no longer being " +
"reached from here.")
}
t.Logf("checked %d notes that mention tracing", seenMentions)
}
// TestUDPTracerouteFactsSayTheThreeThings pins the CONTENT of the shared text,
// not just its presence. Without this the constant could be emptied to "" and
// every assertion above would still pass — strings.Contains(x, "") is true.
func TestUDPTracerouteFactsSayTheThreeThings(t *testing.T) {
for _, want := range []string{
// it prints nothing — the symptom the operator is staring at
"no hops at all",
"* * *",
// under every setting on this page, so nobody goes hunting the knob
"`direct` included",
// the cause, short enough to be read
"local delivery is not forwarding",
// the way out
"traceroute -I",
} {
if !strings.Contains(udpTracerouteFacts, want) {
t.Errorf("udpTracerouteFacts no longer contains %q.\n"+
"The text has to carry all of: the symptom (no hops, only stars), that it is "+
"the same under every rung, the one-clause cause, and the working alternative. "+
"Drop any of them and the note stops being an answer.\ntext: %s", want, udpTracerouteFacts)
}
}
}
// TestL3NoteNamesTheDirectCarriers guards the second correction in this pass.
//
// The L3 notes used to say ping travels the tunnel toward "an outbound that can
// carry plain IP (WireGuard/AmneziaWG)" and that everything else "cannot be
// pinged at all". generate/route.go's l3Target is the authority, and its list is
// exhaustive by adapter registration: a wireguard/AWG node AND the direct
// outbound behind `direct` or an interface/direct egress. In the commonest
// configuration here — tunnel the blocked list, send the rest direct — the
// second half is most of the address space, and those pings answer out of the
// ordinary uplink with its real address. Reading the old text, an operator
// concluded either "tunnelled" or "dropped"; it was neither.
func TestL3NoteNamesTheDirectCarriers(t *testing.T) {
for _, egress := range []string{"", "wan2"} {
g := model.DefaultGlobals()
g.L3Tunnel = true
g.UntunnelableEgress = egress
ws := untunnelablePolicyWarnings(g, nil)
if len(ws) != 1 {
t.Fatalf("egress=%q: expected exactly one untunnelable note, got %d", egress, len(ws))
}
msg := ws[0].Message
if !strings.Contains(msg, "`direct`") || !strings.Contains(msg, "interface egress") {
t.Errorf("egress=%q: the L3 note does not name the other outbounds that carry an "+
"echo.\nPing also answers for addresses routed `direct` or out an interface "+
"egress (generate/route.go l3Target) — and it answers with the real address of "+
"that uplink, which is a disclosure the operator is entitled to read here.\n"+
"note: %s", egress, msg)
}
if !strings.Contains(msg, "real address of that uplink") {
t.Errorf("egress=%q: the L3 note names the direct carriers but not what they cost: "+
"such a ping leaves with the uplink's real address, not the tunnel's.\nnote: %s",
egress, msg)
}
}
}
+184 -13
View File
@@ -328,6 +328,37 @@ func finalizeWarnings(out []Warning) []Warning {
return out
}
// udpTracerouteFacts is the one thing this file is allowed to say about plain
// `traceroute`, shared by every branch below so the panel cannot carry two
// versions of it.
//
// It replaces four claims that were false, and false in the most expensive
// direction: they told an operator whose trace printed nothing that traceroute
// "works", or "still follows your rules", or that the hops it prints "are the
// tunnel's path". Measured on the production router, an ordinary `traceroute`
// from a LAN device prints `* * *` and nothing else — under every rung of the
// untunnelable ladder, `direct` included, with the L3 ingress on or off.
//
// There is no mechanism by which it could print a hop, which is why the text
// says so unconditionally rather than hedging: the UDP probe is diverted by
// tproxy and delivered LOCALLY to the engine's socket, and local delivery is not
// forwarding — the TTL is never decremented, so no router on the path is ever
// provoked into a `time-exceeded`. The engine then opens its OWN connection with
// a fresh TTL, and an ICMP error raised against that has no way back to the
// client's original datagram; the final `port-unreachable` is absorbed in the
// same place. `traceroute -I` and Windows `tracert` are unaffected because they
// are ICMP echo, which the L3 ingress (or the `icmp` rung) carries for real.
//
// netplane/untunnelable.go states the same fact in the same terms where it
// prices `block` (the ECHO bullet). One product, one version of this: change
// one, change both.
const udpTracerouteFacts = "Plain `traceroute` on Linux and macOS is a separate matter, and it reads " +
"the same under every setting here, `direct` included: its UDP probes are tunnelled and DO reach " +
"the target, but it prints no hops at all, only `* * *`. tproxy delivers each probe locally to the " +
"engine, and local delivery is not forwarding, so nothing on the path is ever asked for a " +
"`time-exceeded` — there is no fault at your end to go looking for. Use `traceroute -I` (ICMP " +
"probes, which is what Windows `tracert` already sends) for a trace that prints hops."
// untunnelablePolicyWarnings explains, in terms of what the user will actually
// experience, what the untunnelable-protocol policy costs them.
//
@@ -360,6 +391,142 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
return append(out, Warning{Severity: SeverityInfo, Section: section, Name: name, Message: msg})
}
// The egress carrier owns the whole story the moment untunnelable_egress
// names an egress: the carried protocols get that egress's own mark in
// prerouting and the ROUTING decision sends them out its device with the
// kernel's NAT — the forward chain, where the policy's verdicts live, no
// longer decides their fate. So this branch sits above every other and
// returns its own text; with the option empty, the notes below are
// byte-for-byte what they were. Two phrasing rules here are load-bearing.
// First, the egress is NEVER called a tunnel unconditionally: the option
// accepts any interface/tunnel egress, and on the routers this ships to
// that is at least as often a second WAN — another uplink, whose real
// address the far end sees — as a WireGuard device. Second, the note must
// say out loud that UDP-based VPNs are none of this option's business: an
// operator who reads "VPN passthrough" and enables it for a WireGuard
// client that already worked through the ordinary tunnel has been misled,
// not helped. What the policy still owns is exactly the failure path — a
// name that resolves to no interface/tunnel egress, a rule or route that
// did not come up — and each tail below says what that failure looks like,
// because under `direct` (or an open kill switch) it is a silent leak
// through the normal uplink with the real address, and nothing anywhere
// else would say so.
if egressName := strings.TrimSpace(g.UntunnelableEgress); egressName != "" {
var msg, failure string
if netplane.L3Enabled(g) {
msg = "Ping and Windows tracert keep travelling THROUGH the tunnel, toward every " +
"address your rules send to a WireGuard/AmneziaWG node — the L3 ingress " +
"claims ICMP before this option is consulted. Addresses your rules send " +
"`direct`, or out an interface egress, answer as well, but those pings " +
"leave the way that traffic does, carrying the real address of that uplink " +
"rather than the tunnel's; addresses your rules send anywhere else " +
"(vless/vmess/trojan/shadowsocks and the like) still cannot be pinged at " +
"all, deliberately. " + udpTracerouteFacts + " Everything else the proxy cannot carry — IPsec " +
"(ESP/AH), PPTP/GRE, SCTP and every other protocol that is neither TCP " +
"nor UDP — now leaves through egress \"" + egressName + "\": the kernel " +
"routes it out that interface with that interface's own NAT, and none of " +
"it goes through the proxy or follows your routing rules. "
failure = "when one of the routes is not in place — the L3 route for ping, the " +
"egress route for the rest (a name that matches no interface/tunnel " +
"egress, or a rule or route that failed to come up): "
} else {
msg = "Ping, Windows tracert, IPsec (ESP/AH), PPTP/GRE, SCTP and every other " +
"protocol that is neither TCP nor UDP now leave through egress \"" +
egressName + "\": the kernel routes them out that interface with that " +
"interface's own NAT, and none of it goes through the proxy or follows " +
"your routing rules — the hops Windows tracert and `traceroute -I` print " +
"are that interface's path. " + udpTracerouteFacts + " "
failure = "when the egress route is not in place (a name that matches no " +
"interface/tunnel egress, or a rule or route that failed to come up): "
}
msg += "What that buys depends entirely on what the interface IS: a WireGuard " +
"interface really is a tunnel, but a second WAN is not — it is just another " +
"uplink, and the host on the far end sees that uplink's real address. Two " +
"things this option does NOT do: multicast IPTV does not pass this router " +
"under any setting, and carrying IGMP out an egress cannot change that; and " +
"VPNs that run over UDP (WireGuard, OpenVPN-UDP, IPsec through NAT — IKE on " +
"UDP 500, NAT-T on UDP 4500) never needed it: they are ordinary tunnelled " +
"traffic, keep following your routing rules exactly as before, and gain " +
"nothing from this option. The `untunnelable` policy no longer decides this " +
"traffic's fate — routing settles it before the forward chain gets a say — " +
"and answers only for failure, " + failure
switch {
case !killSwitchClosed(g):
msg += "with the kill switch open nothing is dropped, so whatever loses its " +
"route quietly leaves through your normal uplink with your real IP address."
case policy == netplane.UntunnelableDirect:
msg += "\"direct\" quietly lets it leave through your normal uplink with " +
"your real IP address."
case policy == netplane.UntunnelableICMP:
msg += "\"icmp\" drops it, excepting only ping — which then quietly leaves " +
"with your real IP address instead of failing."
default:
msg += "\"block\" drops it — an honest loss rather than a silent leak."
}
return note(egressName, msg)
}
// The L3 ingress rewrites the ICMP half of every note below, so it gets one
// text of its own rather than four patched variants: echo is marked in
// prerouting and the ROUTING decision carries it into the engine's TUN before
// the forward chain — where the policy accepts and the kill-switch drops
// live — is ever consulted. That holds under all three policy values and
// with the kill switch open alike, which is why this branch sits above the
// kill-switch note: "reaches the internet with your real IP address" stops
// being true for ping the moment the divert exists. What the policy still
// owns is exactly two things, and both are said: the protocols the engine
// cannot ingest at all (raw IPsec, PPTP/GRE), and the fallback path a marked
// packet takes when the L3 route failed to install — under `direct` (or an
// open kill switch) that failure is a SILENT leak with the real address,
// under `block` an honest packet loss. The unpingable-through-proxy sentence
// is deliberate too: those pings used to be answered by the router itself,
// and a fake "alive" is worse than a truthful timeout.
//
// The set of outbounds that carry an echo is NOT "WireGuard/AmneziaWG", and
// saying so was an understatement that hid a disclosure. generate/route.go's
// l3Target is the authority and its list is exhaustive by adapter
// registration: a wireguard/AWG node, AND the direct outbound behind `direct`
// or an interface/direct egress. In the commonest configuration on this router
// — tunnel the blocked list, send the rest direct — that second half is most
// of the address space, and those pings do answer, out of the ordinary uplink
// with its real address. An operator who read the old text concluded either
// "tunnelled" or "dropped", and neither was what their ping was doing.
if netplane.L3Enabled(g) {
msg := "Ping and Windows tracert work and travel THROUGH the tunnel, toward every " +
"address your rules send to a WireGuard/AmneziaWG node. Addresses your rules " +
"send `direct`, or out an interface egress, answer as well — but those pings " +
"leave the way that traffic does, carrying the real address of that uplink " +
"rather than the tunnel's. Addresses your rules send anywhere else " +
"(vless/vmess/trojan/shadowsocks and the like) cannot be pinged at all — " +
"deliberately: those pings used to be answered by the router itself, reporting " +
"hosts alive it had never reached. The hops Windows tracert and `traceroute -I` " +
"print are the tunnel's path, not your own; over IPv6 that same trace shows only " +
"the destination and none of the hops on the way. " + udpTracerouteFacts +
" Raw VPN passthrough (IPsec ESP/AH, PPTP/GRE) cannot " +
"enter the tunnel at all and stays with the untunnelable policy: "
switch {
case !killSwitchClosed(g):
msg += "with the kill switch open none of it is dropped, so it leaves with your " +
"real IP address — and if the L3 route ever fails to come up, ping quietly " +
"does the same instead of failing."
case policy == netplane.UntunnelableDirect:
msg += "\"direct\" lets it out with your real IP address — and if the L3 route " +
"ever fails to come up, ping quietly does the same instead of failing."
case policy == netplane.UntunnelableICMP:
msg += "\"icmp\" drops it, excepting only echo — which now rides the tunnel " +
"anyway, so the exception matters just once: if the L3 route ever fails to " +
"come up, it lets ping quietly leave with your real IP address instead of " +
"failing."
default:
msg += "\"block\" drops it — and if the L3 route ever fails to come up, ping " +
"fails outright rather than leaking."
}
msg += " VPNs that run over UDP (WireGuard, OpenVPN-UDP, IPsec through NAT) are " +
"ordinary tunnelled traffic and are unaffected either way. Multicast IPTV does " +
"not pass this router on any setting; the L3 ingress does not change that."
return note(policy, msg)
}
// With the kill switch open the forward chain has no drops at all, so nothing
// is restricted whatever the policy says. Saying that is more useful than
// repeating a promise which is not being kept.
@@ -377,26 +544,30 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
}
return note(policy,
"This setting has no effect while the kill switch is open: with the kill switch open the "+
"forward chain has no drops at all, so ping, traceroute and raw VPN passthrough "+
"(IPsec ESP/AH, PPTP/GRE) all work — and every one of them reaches the internet with "+
"your real IP address. IPTV is not part of that: multicast does not pass this router "+
"forward chain has no drops at all, so ping, an ICMP trace (`traceroute -I`, Windows "+
"`tracert`) and raw VPN passthrough (IPsec ESP/AH, PPTP/GRE) all work — and every one "+
"of them reaches the internet with your real IP address. "+udpTracerouteFacts+
" IPTV is not part of that: multicast does not pass this router "+
"on any setting, which is a separate matter from this one.")
}
switch policy {
case netplane.UntunnelableDirect:
return note(policy,
"Ping and traceroute work everywhere, and so does raw VPN passthrough (IPsec ESP/AH, "+
"Ping works everywhere, and so do an ICMP trace (`traceroute -I`, Windows `tracert`) and "+
"raw VPN passthrough (IPsec ESP/AH, "+
"PPTP/GRE) — but all of it goes straight out with your real IP address instead of "+
"through the tunnel, because a tunnel cannot carry this kind of traffic. VPNs that "+
"through the tunnel, because a tunnel cannot carry this kind of traffic. "+
udpTracerouteFacts+" VPNs that "+
"run over UDP (WireGuard, OpenVPN-UDP, IPsec through NAT) are ordinary tunnelled "+
"traffic and are unaffected either way. IPTV is not covered by this setting at all: "+
"multicast does not pass this router on any of the three, so switching to `direct` "+
"will not bring it back.")
case netplane.UntunnelableICMP:
return note(policy,
"Ping and traceroute work everywhere, including addresses you send through the tunnel; "+
"the host you ping sees your real IP address. IPTV and VPN passthrough (IPsec/PPTP) "+
"Ping works everywhere, including addresses you send through the tunnel, and so does an "+
"ICMP trace (`traceroute -I`, Windows `tracert`); the host you ping sees your real IP "+
"address. "+udpTracerouteFacts+" IPTV and VPN passthrough (IPsec/PPTP) "+
"work only toward addresses your rules route directly.")
default:
// This text used to say these things "work only toward addresses your rules
@@ -409,13 +580,13 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
// through NAT is the difference between "my VPN broke" and "my VPN is fine".
return note(netplane.UntunnelableBlock,
"Ping, traceroute, IPsec/PPTP VPN passthrough and IPTV do not work from your devices at "+
"all — not even toward addresses your rules route directly. None of this traffic can "+
"travel through a tunnel, so rather than let it out with your real IP address it is "+
"dropped. Concretely: ping and Windows tracert fail (on Linux and macOS traceroute "+
"sends UDP probes instead, which ARE tunnelled — the hops it prints are the tunnel's "+
"path, not your own), and so do raw IPsec (ESP/AH) and PPTP/GRE — a PPTP session will "+
"all — not even toward addresses your rules route directly. None of the traffic this "+
"policy decides can travel through a tunnel, so rather than let it out with your real IP address it is "+
"dropped. Concretely: ping and every ICMP trace (`traceroute -I`, Windows `tracert`) "+
"fail, and so do raw IPsec (ESP/AH) and PPTP/GRE — a PPTP session will "+
"even look connected, because its control channel is TCP and only the payload is "+
"dropped. VPNs that run over UDP are NOT affected: WireGuard, OpenVPN-UDP and IPsec "+
"dropped. "+udpTracerouteFacts+
" VPNs that run over UDP are NOT affected: WireGuard, OpenVPN-UDP and IPsec "+
"through NAT (IKE on UDP 500, NAT-T on UDP 4500) keep working normally. Multicast "+
"IPTV does not cross this router under any setting; that one is not this policy.")
}
+38 -1
View File
@@ -139,9 +139,35 @@ func TestCollectWarningsCapKeepsCriticals(t *testing.T) {
}
}
// withReadableConfig pins the one outside fact this file's warning-COUNT
// assertions depend on: whether Status could read the router's configuration.
//
// It is here because of a defect this test used to have, of the same family as the
// silent skip. The assertion "a fresh Status carries no warnings" was true only by
// accident of environment: a build host has no `uci`, model.ReadUCI failed on every
// Status() call, and that failure was SILENT, so it contributed nothing to count.
// On a router, where `uci` exists and the config reads fine, it was also zero — for
// the opposite reason. The test therefore never stated which world it was in, and
// the moment the unreadable-config failure started publishing a critical warning
// (apply.go, "a finding that is still true may not erase itself") the same source
// line meant two different things in the two environments.
//
// A package-level init() in standing_state_test.go now supplies a router-like
// baseline for the whole package, which is the right home for a package-wide
// fixture. This helper is NOT a duplicate of it: a test that counts warnings must
// not depend on any ambient baseline at all, whoever set it up and whether or not
// it is still there tomorrow. It says what it needs, in its own body.
func withReadableConfig(t *testing.T) {
t.Helper()
orig := readConfig
readConfig = func() (*model.Model, error) { return &model.Model{Globals: model.DefaultGlobals()}, nil }
t.Cleanup(func() { readConfig = orig })
}
// TestStatusWarningsAlwaysNonNil: the panel maps over this array unconditionally,
// so it must serialise as [] and never null.
func TestStatusWarningsAlwaysNonNil(t *testing.T) {
withReadableConfig(t) // counting warnings requires knowing which world we are in
a := New(engine.New(), nil)
s := a.Status()
if s.Warnings == nil {
@@ -216,8 +242,19 @@ func TestWarningsAgainstRealGenerateOutput(t *testing.T) {
if err != nil {
t.Fatalf("GenerateWithWarnings: %v", err)
}
// NOT a skip. This model is built to be unloadable — a binary rule-set at a
// path that does not exist, and a url rule-set with no url — and both are
// protection the operator configured that is not in force. If generate stops
// saying so, the panel goes back to being indistinguishable from "not
// configured" and the user believes ad blocking is on when it is not: that is
// the W7 regression itself, not a reason to stand this test down. (The
// predecessor of this line was a t.Skip, in the same family as the one that
// retired TestApplyInstallsHoldWhenEngineFailsToStart.)
if len(genWarnings) == 0 {
t.Skip("generate produced no warnings for this model; nothing to classify")
t.Fatal("generate produced NO warning for a model whose blocklists cannot load " +
"(/nonexistent/ads.srs does not exist, and the second list has an empty url) — an " +
"unloadable blocklist must always be reported, or nothing in the UI distinguishes " +
"it from one that is working")
}
t.Logf("real generate warnings: %q", genWarnings)
+294
View File
@@ -0,0 +1,294 @@
// The verdict half of `shaterd apply`.
//
// `apply` is not "push the config" — the daemon reconciles anyway, from SIGHUP,
// from the cron/hotplug `shaterd reconcile`, from the panel, and from its own
// startup. What `apply` ADDS, and the only reason to type it, is the safety net:
// snapshot the current last-good, apply, and arm an automatic rollback so a change
// that costs you access to the router undoes itself.
//
// THE DEFECT THIS FILE EXISTS FOR (hit on the live router, 2026-07-26). The verb
// answered `{"changed":false}` and nothing else. That reads as "all good, nothing
// to do". It was not: the operator had edited UCI and run `uci commit`, the
// `config.change` reload trigger had already restarted the daemon, and the fresh
// daemon had applied the new config on startup. By the time `apply` ran there was
// nothing left to apply — and, worse, the last-good it snapshotted as the ROLLBACK
// TARGET was the newly applied config itself. So the auto-rollback was armed onto
// the very configuration it was supposed to protect against: firing it would have
// restored exactly what was already loaded. The safety net was absent, the output
// said nothing about it, and the house went down.
//
// So this file answers ONE question in words, on every apply: IS THERE A SAFETY
// NET, AND IF NOT, WHY NOT. It does not build a new one — that is separate work.
//
// The classification is a pure function of facts the daemon already has, so the
// whole vocabulary is testable on a dev host with no router, no root and no nft.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"os"
"time"
"github.com/sagernet/sing-box/shater/model"
)
// uciConfigPath is the file `uci commit shater` rewrites. A var so tests can point
// the mtime probe at a fixture.
var uciConfigPath = "/etc/config/shater"
// daemonStarted is when THIS process began.
//
// Round(0) strips the monotonic reading on purpose: the only comparison made with
// it is against a FILE mtime, which is wall-clock-only, so keeping a monotonic
// reading here would just make the comparison silently fall back to the wall clock
// anyway. Being explicit is worth a word.
//
// The router has no RTC, so this instant is captured before NTP steps the clock
// (typically forward, by years). That bias is in the recoverable direction for the
// one use below: a start instant that reads OLDER than it was can only make an
// edit look like it happened after the daemon came up — which, for any edit made
// after boot, it did.
var daemonStarted = time.Now().Round(0)
// applyVerdict is the `apply` verb's answer on the control socket.
//
// `changed` and `error` are unchanged from the old ctlResult, so every existing
// consumer keeps working; the rest is the part that was missing. RollbackArmed is
// the load-bearing field: it is the answer to "if this just broke my network, will
// anything undo it?".
type applyVerdict struct {
Changed bool `json:"changed"`
Error string `json:"error,omitempty"`
// RollbackArmed is true ONLY when an automatic rollback was armed AND it would
// take the router somewhere other than where it already is. An armed watcher
// whose target is the running configuration is not a safety net, and is not
// reported as one.
RollbackArmed bool `json:"rollback_armed"`
// Reason is a closed vocabulary (the reason* constants). Closed and positive on
// purpose: an unenumerated outcome must not fall into a bucket that reads
// reassuring.
Reason string `json:"reason"`
// Message says the same thing in the operator's words. Never empty.
Message string `json:"message"`
// ConfirmTimeout is the configured commit-confirm window in seconds, echoed so
// the answer carries the setting it depends on (0 = the feature is off).
ConfirmTimeout int `json:"confirm_timeout"`
}
// The closed reason vocabulary.
const (
// reasonApplyFailed — the apply returned an error; nothing was armed.
reasonApplyFailed = "apply-failed"
// reasonConfigUnreadable — the apply worked but /etc/config/shater could not be
// re-read to learn the confirm timeout, so ArmRollback was never called.
reasonConfigUnreadable = "config-unreadable"
// reasonCommitConfirmOff — globals.confirm_timeout is 0. This is the SHIPPED
// DEFAULT, so on a stock box it is the usual answer.
reasonCommitConfirmOff = "commit-confirm-off"
// reasonDisabled — globals.enabled=0, so the reconcile tore the plane down.
reasonDisabled = "disabled"
// reasonApplied — the running configuration moved and a real rollback target
// was recorded. The only outcome where the net exists.
reasonApplied = "applied"
// reasonAlreadyApplied — nothing moved, and the config file was edited AFTER
// this daemon started: something else applied it before this command ran.
reasonAlreadyApplied = "already-applied"
// reasonNothingToApply — nothing moved, and there is no evidence either way
// about whether the config was edited (a daemon restart erases it).
reasonNothingToApply = "nothing-to-apply"
)
// applyFacts is everything one `apply` learned, as plain values.
type applyFacts struct {
// engineChanged is the Applier's own `changed` — TRUE only when the ENGINE
// config moved. It is deliberately not used to decide whether a rollback target
// is meaningful: a netplane-only change (the kill-switch flipping, a rule
// gaining a schedule window) hashes identical for the engine and still moves
// the router. before/after do that job.
engineChanged bool
err error
// uciRead is whether /etc/config/shater could be re-read after the apply — the
// read that supplies confirmTimeout and gates ArmRollback.
uciRead bool
enabled bool
confirmTimeout int
// before is the last-good model at the instant of Snapshot(), i.e. the rollback
// target this apply recorded. after is the last-good once the apply finished,
// i.e. what is running now. nil means "no successful apply is on record".
before *model.Model
after *model.Model
// configMTime is the mtime of /etc/config/shater, daemonStart is when this
// process began. Zero means unknown, and unknown must never be read as "no".
configMTime time.Time
daemonStart time.Time
}
// classifyApply turns the facts into the verdict. Every branch sets Reason AND
// Message, and every branch that leaves RollbackArmed false says so in words: a
// silent `{"changed":false}` is the exact failure this replaces.
func classifyApply(f applyFacts) applyVerdict {
v := applyVerdict{Changed: f.engineChanged, ConfirmTimeout: f.confirmTimeout}
switch {
case f.err != nil:
v.Error = f.err.Error()
v.Reason = reasonApplyFailed
v.Message = fmt.Sprintf("apply FAILED (%v) and armed NO automatic rollback. "+
"Run `shaterd status`: a failed apply can leave the engine down with the "+
"fail-closed plane holding the LAN.", f.err)
case !f.uciRead:
v.Reason = reasonConfigUnreadable
v.Message = "the apply itself succeeded, but /etc/config/shater could not be re-read " +
"afterwards, so the commit-confirm window was never armed. There is NO automatic " +
"rollback for what was just applied."
case f.confirmTimeout <= 0:
v.Reason = reasonCommitConfirmOff
v.Message = "globals.confirm_timeout is 0, so commit-confirm is switched OFF: this apply " +
"armed NO automatic rollback and nothing will undo it if it cost you access to the " +
"router. Arm it with `uci set shater.@globals[0].confirm_timeout=<seconds>`."
if sameConfig(f.before, f.after) {
v.Message += " Nothing was applied either — the running configuration already " +
"equals /etc/config/shater."
}
case !f.enabled:
// The reconcile tore the plane down. The rollback target is the configuration
// that was running BEFORE, so firing it would switch shater back ON — worth
// saying out loud, because "apply" after a disable looks like a no-op.
if f.before == nil {
v.Reason = reasonDisabled
v.Message = "shater is disabled (globals.enabled=0) and the data plane was torn down. " +
"No earlier configuration is on record, so NO automatic rollback is armed."
break
}
v.RollbackArmed = true
v.Reason = reasonDisabled
v.Message = fmt.Sprintf("shater is disabled (globals.enabled=0) and the data plane was torn "+
"down. An automatic rollback is armed: in %d s the PREVIOUS configuration is re-applied "+
"— i.e. shater switches back on — unless you run `shaterd confirm`.", f.confirmTimeout)
case !sameConfig(f.before, f.after):
v.RollbackArmed = true
v.Reason = reasonApplied
if f.before == nil {
v.Message = fmt.Sprintf("applied. No earlier configuration is on record (this is the "+
"first apply this daemon completed), so the automatic rollback in %d s REMOVES the "+
"shater data plane entirely (safe teardown) unless you run `shaterd confirm`.",
f.confirmTimeout)
break
}
v.Message = fmt.Sprintf("applied. An automatic rollback to the previous configuration is "+
"armed for %d s — run `shaterd confirm` to keep this one.", f.confirmTimeout)
case configEditedSinceStart(f):
v.Reason = reasonAlreadyApplied
v.Message = fmt.Sprintf("nothing was applied: /etc/config/shater was last modified %s, AFTER "+
"this daemon started %s, and the running configuration ALREADY matches it — so something "+
"other than this command applied it (the `config.change` reload trigger, `shaterd "+
"reconcile` from cron/hotplug, the panel, or SIGHUP). NO automatic rollback is armed: the "+
"rollback target recorded here is the configuration that is already running, so nothing "+
"can undo that change.",
stampUTC(f.configMTime), stampUTC(f.daemonStart))
default:
v.Reason = reasonNothingToApply
v.Message = "nothing was applied: the running configuration already equals " +
"/etc/config/shater. NO automatic rollback is armed — the rollback target recorded here " +
"IS the running configuration, so firing it would restore exactly what is loaded now. " +
"That is harmless if you changed nothing. If you DID edit the config, it was applied " +
"before this command ran (a `uci commit` fires the reload trigger, which restarts the " +
"daemon, and the fresh daemon applies on startup) — and then that change is running with " +
"no safety net. shaterd cannot tell those two cases apart across a daemon restart."
}
return v
}
// configEditedSinceStart reports whether /etc/config/shater was written after this
// daemon process began. It is the ONE reliable discriminator between "you changed
// nothing" and "something else applied your change": an edit landing during the
// daemon's lifetime, with the running config already matching it, can only mean a
// reconcile beat this command to it.
//
// It is deliberately one-directional. False does NOT mean "nothing was edited" —
// the `uci commit` reload trigger RESTARTS the daemon, which moves daemonStart
// past the edit — which is why the false branch says so instead of claiming
// everything is fine.
func configEditedSinceStart(f applyFacts) bool {
if f.configMTime.IsZero() || f.daemonStart.IsZero() {
return false
}
return f.configMTime.After(f.daemonStart)
}
// sameConfig reports whether two models are the same desired state — i.e. whether
// a rollback to `a` would leave the router where `b` already has it.
//
// Compared over the WHOLE model, not the engine's option hash: a change the engine
// hashes identical (kill-switch, divert set, DNS intercept) still moves the router
// and still deserves a real rollback target.
func sameConfig(a, b *model.Model) bool {
if a == nil || b == nil {
return a == nil && b == nil
}
return configHash(a) == configHash(b)
}
// configHash is a content hash of a model. encoding/json sorts map keys, so it is
// stable across runs. A marshal failure yields a UNIQUE value rather than a shared
// sentinel: two configs that could not be hashed must not be reported as equal,
// because "equal" is the answer that says the safety net is missing.
func configHash(m *model.Model) string {
b, err := json.Marshal(m)
if err != nil {
return fmt.Sprintf("unhashable-%p", m)
}
sum := sha256.Sum256(b)
return hex.EncodeToString(sum[:])
}
// stampUTC formats an instant for the operator. UTC, like every other timestamp
// this daemon prints (the box has no tzdata).
func stampUTC(t time.Time) string { return t.UTC().Format("2006-01-02 15:04:05 UTC") }
// applyNoticeLine renders the operator-facing stderr line for a control-socket
// reply, or "" when there is nothing to add. Only the `apply` verb has anything to
// say here; every other verb keeps its old, quiet output.
//
// It is the CLI half of the same honesty rule: when no rollback was armed the line
// LEADS with that fact, because the reader of a terminal scans the first words.
func applyNoticeLine(cmd, resp string) string {
if cmd != "apply" {
return ""
}
var v applyVerdict
if json.Unmarshal([]byte(resp), &v) != nil || v.Message == "" {
return ""
}
if v.RollbackArmed {
return "shaterd apply: " + v.Message
}
return "shaterd apply: NO AUTOMATIC ROLLBACK — " + v.Message
}
// uciConfigMTime returns the mtime of /etc/config/shater, or the zero time when it
// cannot be stat'ed (which classifyApply treats as "unknown", never as "not
// edited").
func uciConfigMTime() time.Time {
st, err := os.Stat(uciConfigPath)
if err != nil {
return time.Time{}
}
return st.ModTime()
}
+466
View File
@@ -0,0 +1,466 @@
package main
import (
"encoding/json"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/shater/model"
)
// --- fixtures ---------------------------------------------------------------
var (
fixtureStart = time.Date(2026, 7, 26, 12, 0, 0, 0, time.UTC)
beforeStart = fixtureStart.Add(-1 * time.Hour)
afterStart = fixtureStart.Add(1 * time.Minute)
)
// cfg builds a distinguishable model. killSwitch is varied rather than an engine
// field on purpose: it is a change the ENGINE hashes identical, so a test that
// separates cfg("open") from cfg("closed") also proves the comparison does not
// lean on the engine's own `changed` flag.
func cfg(killSwitch string) *model.Model {
m := &model.Model{Globals: model.DefaultGlobals()}
m.Globals.KillSwitch = killSwitch
m.Nodes = []model.Node{{Name: "n1", Enabled: true, URI: "vless://example"}}
return m
}
// armedFacts is the healthy baseline: enabled, a 30 s confirm window, config read.
func armedFacts() applyFacts {
return applyFacts{
uciRead: true,
enabled: true,
confirmTimeout: 30,
before: cfg("closed"),
after: cfg("closed"),
configMTime: beforeStart,
daemonStart: fixtureStart,
}
}
// --- the two cases the verb must tell apart ---------------------------------
// TestApplyReportsSomethingElseAlreadyApplied is the live-router regression.
//
// Sequence on mini_router (2026-07-26): `uci set` + `uci commit`, which fires the
// `config.change` reload trigger, which restarts the daemon / a cron reconcile
// picks it up — either way the new config is already running by the time `shaterd
// apply` is typed. The verb answered `{"changed":false}`, which reads as "all
// good", while the rollback target it had just snapshotted WAS the newly applied
// config. There was no safety net and nothing said so.
//
// Here the discriminator is available: the file was written after the daemon
// started, so the reconcile that applied it can only have been someone else's.
func TestApplyReportsSomethingElseAlreadyApplied(t *testing.T) {
f := armedFacts()
f.engineChanged = false
f.configMTime = afterStart // edited while this daemon was already up
v := classifyApply(f)
if v.RollbackArmed {
t.Fatalf("rollback_armed = true, but the rollback target is the config that is "+
"already running — nothing could be undone. verdict: %+v", v)
}
if v.Reason != reasonAlreadyApplied {
t.Errorf("reason = %q, want %q", v.Reason, reasonAlreadyApplied)
}
mustSay(t, v.Message,
"NO automatic rollback", // the missing safety net, in words
"ALREADY matches", // WHY nothing was applied
"other than this command",
)
// The evidence the operator needs to believe it must be in the sentence.
mustSay(t, v.Message, stampUTC(afterStart), stampUTC(fixtureStart))
}
// TestApplyUnchangedIsHonestAboutTheAmbiguity covers the other case: nothing moved
// and there is NO evidence either way, because the `uci commit` reload trigger is
// stop+start — it moves the daemon's start past the edit, erasing the mtime
// discriminator.
//
// The verb must not read as "all good". It must say the net is absent, that this
// is harmless if nothing was changed, AND that a change applied before the command
// ran is running unprotected. A formally-true sentence that reads as success is
// the same lie in a nicer suit.
func TestApplyUnchangedIsHonestAboutTheAmbiguity(t *testing.T) {
f := armedFacts()
f.engineChanged = false
f.configMTime = beforeStart // the file predates this daemon: no evidence
v := classifyApply(f)
if v.RollbackArmed {
t.Fatalf("rollback_armed = true with target == running config: %+v", v)
}
if v.Reason != reasonNothingToApply {
t.Errorf("reason = %q, want %q", v.Reason, reasonNothingToApply)
}
mustSay(t, v.Message,
"NO automatic rollback",
"harmless if you changed nothing", // the benign half, named as conditional
"no safety net", // the dangerous half, named
"cannot tell", // the ambiguity, admitted
)
}
// TestApplyArmsWhenTheConfigActuallyMoved is the CONTROL for both tests above: the
// same instrument must be able to report a real safety net, or "no net" proves
// nothing. (Engineering standard: a negative result needs a positive control.)
func TestApplyArmsWhenTheConfigActuallyMoved(t *testing.T) {
f := armedFacts()
f.engineChanged = true
f.before = cfg("open")
f.after = cfg("closed")
v := classifyApply(f)
if !v.RollbackArmed {
t.Fatalf("rollback_armed = false, but the rollback target differs from the "+
"running config — the net is real here: %+v", v)
}
if v.Reason != reasonApplied {
t.Errorf("reason = %q, want %q", v.Reason, reasonApplied)
}
mustSay(t, v.Message, "30 s", "shaterd confirm")
}
// TestApplyArmsOnNetplaneOnlyChange pins that the verdict does not trust the
// engine's `changed`. A kill-switch flip hashes identical for the engine
// (option.Options are unchanged) yet reloads the nft ruleset and really can take
// the LAN off the air — exactly the apply that most needs a rollback.
func TestApplyArmsOnNetplaneOnlyChange(t *testing.T) {
f := armedFacts()
f.engineChanged = false // the engine hash did not move
f.before = cfg("open")
f.after = cfg("closed")
v := classifyApply(f)
if !v.RollbackArmed || v.Reason != reasonApplied {
t.Fatalf("a netplane-only change must still arm a rollback; got %+v", v)
}
if v.Changed {
t.Errorf("changed must keep reporting the ENGINE flag verbatim (false here), got true")
}
}
// --- the other ways the net is silently absent ------------------------------
// TestApplyConfirmTimeoutZeroIsNotSilent covers the SHIPPED DEFAULT:
// openwrt/shater-core/files/etc/config/shater sets `confirm_timeout '0'`, and
// apply.ArmRollback returns immediately for a non-positive timeout. So on a stock
// box every `apply` arms nothing at all, and used to say `{"changed":true}`.
func TestApplyConfirmTimeoutZeroIsNotSilent(t *testing.T) {
f := armedFacts()
f.engineChanged = true
f.confirmTimeout = 0
f.before = cfg("open")
f.after = cfg("closed")
v := classifyApply(f)
if v.RollbackArmed {
t.Fatalf("confirm_timeout=0 arms nothing (ArmRollback returns early); got %+v", v)
}
if v.Reason != reasonCommitConfirmOff {
t.Errorf("reason = %q, want %q", v.Reason, reasonCommitConfirmOff)
}
mustSay(t, v.Message, "confirm_timeout", "NO automatic rollback")
if v.ConfirmTimeout != 0 {
t.Errorf("confirm_timeout echoed as %d, want 0", v.ConfirmTimeout)
}
}
// TestApplyUnreadableConfigIsNotSilent: the daemon arms the window from a SECOND
// model.ReadUCI after the reconcile. When that read fails ArmRollback is never
// called — a hole that produced a plain success answer.
func TestApplyUnreadableConfigIsNotSilent(t *testing.T) {
f := armedFacts()
f.engineChanged = true
f.uciRead = false
f.before = cfg("open")
f.after = cfg("closed")
v := classifyApply(f)
if v.RollbackArmed || v.Reason != reasonConfigUnreadable {
t.Fatalf("an unreadable config after a successful apply arms nothing; got %+v", v)
}
mustSay(t, v.Message, "NO automatic", "could not be re-read")
}
// TestApplyFailureSaysNothingWasArmed: on an error the daemon skips ArmRollback
// entirely, and a failed apply is precisely when an operator assumes a net exists.
func TestApplyFailureSaysNothingWasArmed(t *testing.T) {
f := armedFacts()
f.err = errors.New("engine: address already in use")
v := classifyApply(f)
if v.RollbackArmed || v.Reason != reasonApplyFailed {
t.Fatalf("a failed apply arms nothing; got %+v", v)
}
if v.Error != "engine: address already in use" {
t.Errorf("error = %q, want the applier's error verbatim", v.Error)
}
mustSay(t, v.Message, "NO automatic rollback")
}
// TestApplyFirstEverArmsATeardown: with no last-good on record the rollback target
// is nil, which apply.rollbackTo turns into a SAFE TEARDOWN. That is a real net —
// but a surprising one, so the words must name it.
func TestApplyFirstEverArmsATeardown(t *testing.T) {
f := armedFacts()
f.engineChanged = true
f.before = nil
f.after = cfg("closed")
v := classifyApply(f)
if !v.RollbackArmed || v.Reason != reasonApplied {
t.Fatalf("a first-ever apply still arms a rollback (safe teardown); got %+v", v)
}
mustSay(t, v.Message, "REMOVES the shater data plane")
}
// TestApplyDisabledSaysTheRollbackTurnsItBackOn: `apply` after globals.enabled=0
// tears the plane down and returns changed=false, yet the armed rollback re-applies
// the previous (enabled) config. Reported as a no-op, that is a trap.
func TestApplyDisabledSaysTheRollbackTurnsItBackOn(t *testing.T) {
f := armedFacts()
f.enabled = false
f.engineChanged = false
v := classifyApply(f)
if !v.RollbackArmed || v.Reason != reasonDisabled {
t.Fatalf("a disable leaves a real rollback target armed; got %+v", v)
}
mustSay(t, v.Message, "switches back on", "shaterd confirm")
}
// TestApplyDisabledWithNoHistoryArmsNothing: same path, but nothing was ever
// applied, so the "rollback" is a teardown of a plane that is already down.
func TestApplyDisabledWithNoHistoryArmsNothing(t *testing.T) {
f := armedFacts()
f.enabled = false
f.before = nil
f.after = nil
v := classifyApply(f)
if v.RollbackArmed || v.Reason != reasonDisabled {
t.Fatalf("no history means nothing to roll back to; got %+v", v)
}
mustSay(t, v.Message, "NO automatic rollback")
}
// --- invariants across the whole vocabulary ---------------------------------
// TestEveryVerdictSpeaks is the anti-silence gate: whatever the facts, the answer
// carries a reason from the closed vocabulary and a non-empty message, and any
// answer WITHOUT a rollback says so unmistakably. `{"changed":false}` with nothing
// else must be unreachable.
func TestEveryVerdictSpeaks(t *testing.T) {
known := map[string]bool{
reasonApplyFailed: true, reasonConfigUnreadable: true, reasonCommitConfirmOff: true,
reasonDisabled: true, reasonApplied: true, reasonAlreadyApplied: true,
reasonNothingToApply: true,
}
cases := map[string]applyFacts{}
for _, uciRead := range []bool{true, false} {
for _, enabled := range []bool{true, false} {
for _, timeout := range []int{0, 30} {
for _, mtime := range []time.Time{{}, beforeStart, afterStart} {
for i, pair := range [][2]*model.Model{
{nil, nil}, {nil, cfg("closed")},
{cfg("closed"), cfg("closed")}, {cfg("open"), cfg("closed")},
} {
for _, e := range []error{nil, errors.New("boom")} {
f := applyFacts{
engineChanged: i == 3, err: e, uciRead: uciRead, enabled: enabled,
confirmTimeout: timeout, before: pair[0], after: pair[1],
configMTime: mtime, daemonStart: fixtureStart,
}
name := strings.Join([]string{
boolName(uciRead), boolName(enabled), stampUTC(mtime),
}, "/") + "/" + string(rune('a'+i))
if e != nil {
name += "/err"
}
if timeout == 0 {
name += "/t0"
}
cases[name] = f
}
}
}
}
}
}
for name, f := range cases {
v := classifyApply(f)
if !known[v.Reason] {
t.Errorf("%s: reason %q is outside the closed vocabulary", name, v.Reason)
}
if strings.TrimSpace(v.Message) == "" {
t.Errorf("%s: empty message — this is the silent {\"changed\":false} defect", name)
}
if !v.RollbackArmed && !strings.Contains(strings.ToLower(v.Message), "no automatic rollback") {
t.Errorf("%s: rollback_armed=false but the message never says so: %q", name, v.Message)
}
if v.RollbackArmed && strings.Contains(strings.ToLower(v.Message), "no automatic rollback") {
t.Errorf("%s: rollback_armed=true but the message denies it: %q", name, v.Message)
}
}
}
// TestVerdictNeverClaimsANetOverAnIdenticalTarget is the single load-bearing
// invariant, asserted independently of the branch order: a rollback target that
// equals the running config can NEVER be reported as armed, because firing it
// would restore what is already loaded.
func TestVerdictNeverClaimsANetOverAnIdenticalTarget(t *testing.T) {
for _, timeout := range []int{0, 1, 30, 3600} {
for _, mtime := range []time.Time{{}, beforeStart, afterStart} {
f := armedFacts()
f.confirmTimeout = timeout
f.configMTime = mtime
f.before, f.after = cfg("closed"), cfg("closed")
if v := classifyApply(f); v.RollbackArmed {
t.Errorf("timeout=%d mtime=%v: armed over an identical target: %+v", timeout, mtime, v)
}
}
}
}
// --- the comparison instrument itself ---------------------------------------
// TestSameConfigDetectsBothWays: a comparison that always says "equal" would make
// every apply look unprotected, and one that always says "different" would make
// every apply look safe. Both directions are pinned.
func TestSameConfigDetectsBothWays(t *testing.T) {
if !sameConfig(cfg("closed"), cfg("closed")) {
t.Errorf("identical models must compare equal")
}
if sameConfig(cfg("open"), cfg("closed")) {
t.Errorf("a kill-switch difference must compare different (the engine hashes it identical)")
}
a := cfg("closed")
b := cfg("closed")
b.Nodes[0].URI = "vless://other"
if sameConfig(a, b) {
t.Errorf("a node URI difference must compare different")
}
if !sameConfig(nil, nil) {
t.Errorf("nil/nil is the same state (nothing applied)")
}
if sameConfig(nil, cfg("closed")) || sameConfig(cfg("closed"), nil) {
t.Errorf("nil vs a model must compare different")
}
}
// --- wire contract ----------------------------------------------------------
// TestApplyVerdictWire pins the JSON the control socket emits: `changed` survives
// for existing consumers, and the new fields are always present (rollback_armed is
// NOT omitempty — a missing field would read as "unknown", and this answer is the
// one place that must not be ambiguous).
func TestApplyVerdictWire(t *testing.T) {
f := armedFacts()
f.configMTime = afterStart
b, err := json.Marshal(classifyApply(f))
if err != nil {
t.Fatalf("marshal: %v", err)
}
var got map[string]any
if err := json.Unmarshal(b, &got); err != nil {
t.Fatalf("unmarshal: %v", err)
}
for _, k := range []string{"changed", "rollback_armed", "reason", "message", "confirm_timeout"} {
if _, ok := got[k]; !ok {
t.Errorf("key %q missing from the apply reply: %s", k, b)
}
}
if got["rollback_armed"] != false {
t.Errorf("rollback_armed = %v, want false", got["rollback_armed"])
}
}
// TestApplyNoticeLine pins the terminal line. It must LEAD with the missing net —
// the operator scans the first words — and must stay silent for other verbs.
func TestApplyNoticeLine(t *testing.T) {
unarmed, _ := json.Marshal(applyVerdict{Reason: reasonAlreadyApplied, Message: "already applied."})
line := applyNoticeLine("apply", string(unarmed))
if !strings.HasPrefix(line, "shaterd apply: NO AUTOMATIC ROLLBACK") {
t.Errorf("unarmed notice = %q, want it to lead with the missing rollback", line)
}
armed, _ := json.Marshal(applyVerdict{RollbackArmed: true, Reason: reasonApplied, Message: "applied."})
line = applyNoticeLine("apply", string(armed))
if strings.Contains(line, "NO AUTOMATIC") || !strings.Contains(line, "applied.") {
t.Errorf("armed notice = %q, want the plain message", line)
}
if got := applyNoticeLine("confirm", string(unarmed)); got != "" {
t.Errorf("applyNoticeLine(confirm) = %q, want silence for other verbs", got)
}
if got := applyNoticeLine("apply", "not json"); got != "" {
t.Errorf("applyNoticeLine on garbage = %q, want silence", got)
}
}
// --- the mtime probe --------------------------------------------------------
// TestUCIConfigMTime: a real stat, and an absent file reported as UNKNOWN (zero),
// which configEditedSinceStart must never read as "not edited".
func TestUCIConfigMTime(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "shater")
if err := os.WriteFile(path, []byte("config globals\n"), 0o644); err != nil {
t.Fatal(err)
}
orig := uciConfigPath
uciConfigPath = path
defer func() { uciConfigPath = orig }()
if got := uciConfigMTime(); got.IsZero() {
t.Errorf("mtime of an existing file must not be zero")
}
uciConfigPath = filepath.Join(dir, "absent")
if got := uciConfigMTime(); !got.IsZero() {
t.Errorf("mtime of an absent file = %v, want the zero time (unknown)", got)
}
// Unknown must fall to the CAUTIOUS branch, not to "already applied".
f := armedFacts()
f.configMTime = time.Time{}
if configEditedSinceStart(f) {
t.Errorf("an unknown mtime must not be reported as an edit")
}
if v := classifyApply(f); v.Reason != reasonNothingToApply {
t.Errorf("unknown mtime: reason = %q, want the honest-ambiguity branch %q",
v.Reason, reasonNothingToApply)
}
}
// --- helpers ----------------------------------------------------------------
func mustSay(t *testing.T, msg string, phrases ...string) {
t.Helper()
for _, p := range phrases {
if !strings.Contains(msg, p) {
t.Errorf("message does not contain %q:\n %s", p, msg)
}
}
}
func boolName(b bool) string {
if b {
return "y"
}
return "n"
}
+281
View File
@@ -0,0 +1,281 @@
// Keeping the fail-closed plane alive across the moments the daemon is not.
//
// The daemon owns the `inet shater` table, which means the table exists exactly
// while the daemon does. Three of those moments are not covered by anything else,
// and all three are the same defect wearing different clothes: the protection is
// an in-process thing, and the process is not always there.
//
// BOOT /etc/init.d/shater is START=99. fw4 loaded `lan -> wan ACCEPT` at 19
// and netifd brought the LAN up at 20; the clients that reconnect in
// between are unprotected until the daemon has been decompressed off
// flash, waited out any predecessor, migrated UCI and applied.
// RESTART SIGTERM ran an unconditional Teardown — kill_switch was not so much
// as consulted — and the successor cannot apply until the init's
// shater_wait_stopped loop, `shaterd migrate` and engine start have all
// finished. `reload_service` is stop+start, and so is every package
// upgrade, so this ran on a routine `Save & Apply`.
// NO CONFIG model.ReadUCI failing left the arming call unreached: it sat in the
// else-branch of the successful read. Nothing recovered from it either
// — Reconcile returns before any plane work, and the cron watchdog sees
// a live pidof and its own `uci -q get` fails the same way.
//
// The answer to all three is one artifact: netplane's BOOT ARMOR, a persisted copy
// of the fail-closed holding plane (netplane/armor.go). This file is the daemon's
// half — it keeps that copy honest, and it reinstates it in the two cases the
// daemon is the only one who can.
//
// Everything here is deliberately conservative in ONE direction: it never installs
// a plane the operator did not ask for. `globals.enabled=0` or `kill_switch=open`
// removes the armor and installs nothing, and a deliberate `/etc/init.d/shater
// stop` is a handoff-free exit that leaves nothing behind. Fail-closed is a
// policy, and a kill switch that outlives its own off switch is not one.
package main
import (
"os"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// restartHandoffPath is raised by /etc/init.d/shater around a restart/reload and
// cleared by its start (and by a real stop). Its presence at SIGTERM means "this
// daemon is being REPLACED", as opposed to "this daemon is being switched off".
//
// tmpfs on purpose: a marker that survived a power cut would make the first boot
// after it look like a restart.
//
// A var, not a const, only so tests can point it at a temp dir.
var restartHandoffPath = "/var/run/shater.restarting"
// restartHandoffPending reports whether the init script announced a restart.
//
// Absent is read as "a real stop", which is the SAFE direction to be wrong in: it
// degrades to exactly the behaviour that shipped before this file existed (full
// teardown), whereas the other default would leave a stopped router blocked.
func restartHandoffPending() bool {
_, err := os.Stat(restartHandoffPath)
return err == nil
}
// armorPlan is what to do about the fail-closed plane at a decision point.
type armorPlan int
const (
// armorNothing: install nothing. Either the operator does not want a plane
// (disabled / kill_switch=open), or there is nothing to install from.
armorNothing armorPlan = iota
// armorRender: build the holding plane from the model we just read. Preferred
// whenever a model is readable — it reflects the CURRENT interface set, where a
// snapshot may predate an interface rename.
armorRender
// armorSnapshot: reinstate the persisted boot armor. The only option when the
// config cannot be read, which is precisely when it is needed.
armorSnapshot
)
// armorWanted reports whether m asks for a fail-closed plane at all: the stack is
// enabled AND the kill switch is closed. Both halves are the operator's explicit
// choice and neither may be second-guessed — `kill_switch=open` is a documented
// decision to let traffic through when the engine is down, not an oversight.
func armorWanted(m *model.Model) bool {
return m != nil && m.Globals.Enabled && netplane.KillSwitchClosed(m.Globals)
}
// planStartupArmor decides what to install when the daemon starts and could NOT
// read its config.
//
// It is only ever consulted on the read-failure path: with a readable model the
// applier's own ArmHold does this job (and does it better — it holds the apply
// lock while it works). With no model there is nothing to render from, so the
// persisted snapshot is the entire answer; with no snapshot either, nothing is
// installed, because "this router has never applied an enabled, fail-closed
// config" is then the most likely truth and blacking out a LAN on a guess is not
// a recovery.
func planStartupArmor(readErr error, snapshot bool) armorPlan {
if readErr == nil {
return armorNothing
}
if snapshot {
return armorSnapshot
}
return armorNothing
}
// planExitArmor decides what the daemon leaves behind when it is asked to exit.
//
// The handoff flag is the whole distinction the old code was missing. A restart,
// a reload and a package upgrade all reach this point, and in all three the
// operator has not asked for protection to end — only for this process to be
// replaced. A `stop` has asked for exactly that, and must be obeyed: it is the
// operator's escape hatch, and a kill switch that cannot be switched off is a
// brick.
func planExitArmor(handoff bool, m *model.Model, readErr error, snapshot bool) armorPlan {
if !handoff {
return armorNothing
}
if readErr != nil {
// Being replaced with an unreadable config: the snapshot is the last thing
// this router is known to have wanted, and it is still the honest answer.
if snapshot {
return armorSnapshot
}
return armorNothing
}
if !armorWanted(m) {
return armorNothing
}
return armorRender
}
// refreshBootArmor keeps the persisted holding plane in step with the desired
// state. Called after every successful UCI read, so the snapshot on flash always
// describes the config the router is actually running.
//
// Writing is content-gated inside netplane.SaveBootArmor (this runs once a minute
// under cron; rewriting an identical file that often is how flash dies), and the
// REMOVE half matters just as much as the write: turning the stack off, or opening
// the kill switch, has to disarm the next boot too, or the operator's change would
// silently come back after a power cut.
func refreshBootArmor(m *model.Model, logger log.ContextLogger) {
if !armorWanted(m) {
if netplane.BootArmorPresent() {
if err := netplane.RemoveBootArmor(); err != nil {
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", err)
} else {
logger.Info("boot armor removed (the stack is disabled or the kill switch is open): ",
"the LAN is no longer blocked at boot before the daemon starts")
}
}
return
}
ruleset, err := netplane.RenderHoldNft(m)
if err != nil {
// A transient render failure must not disarm: a stale fail-closed plane is
// recoverable (the daemon replaces it seconds into the next boot), an absent
// one is a leak.
logger.Warn("boot armor: could not render the fail-closed plane: ", err)
return
}
if ruleset == "" {
// No divert devices at all — this config intercepts nothing, so there is
// nothing for a boot-time plane to protect. Blocking the LAN at boot on
// behalf of a config that does not touch it would be a pure outage.
if netplane.BootArmorPresent() {
if rerr := netplane.RemoveBootArmor(); rerr != nil {
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", rerr)
}
}
return
}
changed, serr := netplane.SaveBootArmor(ruleset)
switch {
case serr != nil:
logger.Warn("boot armor: could not write ", netplane.BootArmorPath, ": ", serr,
" — the LAN will be unprotected between boot and this daemon's first apply")
case changed:
logger.Info("boot armor updated (", netplane.BootArmorPath,
"): the LAN is fail-closed from early boot until the engine is up")
}
}
// armFromSnapshot reinstates the persisted holding plane and reports whether a
// plane is now standing. why is a short phrase for the log ("config is
// unreadable", "restart handoff").
func armFromSnapshot(why string, logger log.ContextLogger) bool {
loaded, err := netplane.LoadBootArmor()
switch {
case err != nil:
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): the saved plane ",
netplane.BootArmorPath, " could not be loaded: ", err,
" — LAN traffic may be reaching the WAN unprotected")
return false
case loaded:
logger.Error("fail-closed plane reinstated from ", netplane.BootArmorPath,
" (", why, "): LAN->WAN forwarding is BLOCKED. ",
"SSH, LuCI and the admin panel remain reachable.")
return true
default:
logger.Warn("no saved fail-closed plane at ", netplane.BootArmorPath, " (", why,
"): nothing was installed")
return false
}
}
// armFromModel renders the holding plane for m, installs it, and reports whether
// a plane is now standing.
//
// The return value is load-bearing on the exit path: Applier.TeardownExiting keeps
// the nft table only when a plane really was installed, so a render failure or a
// fail-open config falls back to the old remove-everything behaviour instead of
// leaving whatever the engine happened to have in the kernel.
func armFromModel(m *model.Model, why string, logger log.ContextLogger) bool {
ruleset, err := netplane.RenderHoldNft(m)
if err != nil {
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): render failed: ", err)
return false
}
if ruleset == "" {
// No divert devices: there is nothing this plane would protect.
return false
}
// One `nft -f` that opens with `delete table` and closes with the new table:
// the swap is a single netlink transaction, so this REPLACES whatever plane is
// loaded without the table ever being absent. That property is why the exit
// path can arm before it tears down.
if err := netplane.ApplyNft(ruleset); err != nil {
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): ", err,
" — LAN traffic may be reaching the WAN unprotected")
return false
}
logger.Info("fail-closed plane left in place (", why,
"): LAN->WAN forwarding stays BLOCKED until the next daemon applies. ",
"SSH, LuCI and the admin panel remain reachable.")
return true
}
// armOnUnreadableConfig is the window-3 answer: the daemon is up but cannot read
// its own desired state, so it falls back to the last state it persisted.
//
// A table that is ALREADY loaded is left alone. This runs on every reconcile —
// cron fires one a minute — and a `nft -f` is a delete-and-recreate of the whole
// table plus a DNS conntrack flush, so re-installing an identical plane sixty
// times an hour would be pure churn, and each replacement is itself a brief hole.
// The question this path answers is "is there anything at all standing", and once
// the answer is yes it stays yes until an apply succeeds and replaces it properly.
func armOnUnreadableConfig(logger log.ContextLogger) {
if netplane.TableExists() {
return
}
if planStartupArmor(errUnreadableConfig, netplane.BootArmorPresent()) != armorSnapshot {
logger.Warn("the config could not be read and there is no saved fail-closed plane at ",
netplane.BootArmorPath, " — nothing is protecting the LAN; fix /etc/config/shater ",
"(a full /overlay is the usual cause) and reconcile")
return
}
armFromSnapshot("the config could not be read", logger)
}
// errUnreadableConfig is a stand-in for "the read failed" in the call above,
// where the concrete error has already been logged by the caller.
var errUnreadableConfig = os.ErrInvalid
// armOnExit is the window-2 answer: what this daemon leaves in the kernel when it
// is asked to go away, and whether anything is now standing there. See
// planExitArmor for the policy.
//
// It is called BY Applier.TeardownExiting, before the teardown and under the apply
// lock, so that the plane is swapped rather than removed-then-rebuilt. It must
// therefore never call back into the Applier — everything here goes straight to
// model.ReadUCI and netplane.
func armOnExit(handoff bool, logger log.ContextLogger) bool {
m, err := model.ReadUCI()
switch planExitArmor(handoff, m, err, netplane.BootArmorPresent()) {
case armorRender:
return armFromModel(m, "restart handoff", logger)
case armorSnapshot:
return armFromSnapshot("restart handoff, config unreadable", logger)
}
return false
}
+182
View File
@@ -0,0 +1,182 @@
package main
// The daemon's half of the fail-closed armor: the two decisions that decide
// whether the LAN is protected in the moments this process is not running.
//
// Both are pure functions on purpose. The behaviour they encode can otherwise
// only be observed on a router, with root, by killing a daemon at the right
// moment and reading `nft list ruleset` — i.e. never, in a gate.
import (
"errors"
"os"
"path/filepath"
"strings"
"testing"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
func enabledClosed() *model.Model {
return &model.Model{Globals: model.Globals{
Enabled: true, KillSwitch: "closed", FwmarkBase: 0x2000, TableBase: 0x2000,
}, Inbounds: []model.Inbound{{
Name: "lan", Enabled: true, Type: "tproxy", Network: "lan",
TproxyPort: 12345, TCP: true, UDP: true,
}}}
}
// TestPlanExitArmor is W2. SIGTERM used to run an unconditional Teardown — the
// only place in this codebase that removes the fail-closed plane without so much
// as reading kill_switch — and every `restart`, every `reload_service` (which is
// what a LuCI Save & Apply runs) and every package upgrade went through it. The
// gap that follows is guaranteed non-empty by the init script itself.
//
// So the exit has to know WHY it is exiting. What it must never do is confuse the
// two directions: a restart that leaves nothing behind is a plaintext window, and
// a stop that leaves a block behind is a router the operator cannot un-brick.
//
// RED BEFORE: there was no such decision — the daemon always tore everything down.
func TestPlanExitArmor(t *testing.T) {
open := enabledClosed()
open.Globals.KillSwitch = "open"
disabled := enabledClosed()
disabled.Globals.Enabled = false
cases := []struct {
name string
handoff bool
m *model.Model
readErr error
snapshot bool
want armorPlan
}{
{"restart, enabled + fail-closed => leave the plane standing",
true, enabledClosed(), nil, true, armorRender},
{"restart, no snapshot on disk => still render from the live model",
true, enabledClosed(), nil, false, armorRender},
{"restart, kill_switch=open => the operator chose fail-open; install nothing",
true, open, nil, true, armorNothing},
{"restart, stack disabled => nothing to protect",
true, disabled, nil, true, armorNothing},
{"restart, config unreadable => the persisted plane is the last known truth",
true, nil, errors.New("uci: no such file"), true, armorSnapshot},
{"restart, config unreadable and nothing persisted => nothing to install",
true, nil, errors.New("uci: no such file"), false, armorNothing},
// The escape hatch. A deliberate `/etc/init.d/shater stop` must mean what it
// says in every one of these, or the kill switch has no off switch.
{"stop, enabled + fail-closed => the plane goes away",
false, enabledClosed(), nil, true, armorNothing},
{"stop, config unreadable => still goes away",
false, nil, errors.New("uci: no such file"), true, armorNothing},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := planExitArmor(c.handoff, c.m, c.readErr, c.snapshot); got != c.want {
t.Errorf("planExitArmor(handoff=%v, readErr=%v, snapshot=%v) = %v, want %v",
c.handoff, c.readErr, c.snapshot, got, c.want)
}
})
}
}
// TestPlanStartupArmor is W3. An unreadable /etc/config/shater — a full /overlay
// caught mid `uci commit` is the cause the init script itself documents — left the
// arming call unreached, because it sat in the else-branch of the successful read.
// Nothing recovered from that: Reconcile returns before any plane work and the
// cron watchdog's own `uci -q get` fails identically, so the box sat there with a
// live daemon, an answering panel and no table at all.
//
// RED BEFORE: no decision existed; the read-failure branch only logged.
func TestPlanStartupArmor(t *testing.T) {
readErr := errors.New("uci: cannot read /etc/config/shater")
if got := planStartupArmor(readErr, true); got != armorSnapshot {
t.Errorf("unreadable config with a persisted plane = %v, want armorSnapshot", got)
}
// Nothing persisted means this router has never applied an enabled,
// fail-closed config. Blacking out a LAN on that guess is not a recovery.
if got := planStartupArmor(readErr, false); got != armorNothing {
t.Errorf("unreadable config with no persisted plane = %v, want armorNothing", got)
}
// A readable config is the applier's business (ArmHold), not this path's.
if got := planStartupArmor(nil, true); got != armorNothing {
t.Errorf("readable config = %v, want armorNothing", got)
}
}
// TestRefreshBootArmorTracksDesiredState is W1's durable half: the file
// /etc/init.d/shater-armor loads at START=21 only exists while the operator wants
// it to. Writing it is half the contract; REMOVING it when the stack is switched
// off or the kill switch is opened is the other half, and the more dangerous one
// to get wrong — a stale armor would reinstate, at the next power cut, a block the
// operator had already turned off.
//
// RED BEFORE: neither the file nor this function existed.
func TestRefreshBootArmorTracksDesiredState(t *testing.T) {
orig := netplane.BootArmorPath
netplane.BootArmorPath = filepath.Join(t.TempDir(), "shater", "boot.nft")
defer func() { netplane.BootArmorPath = orig }()
logger := log.StdLogger()
refreshBootArmor(enabledClosed(), logger)
if !netplane.BootArmorPresent() {
t.Fatalf("an enabled, fail-closed config must persist a boot armor")
}
b, err := os.ReadFile(netplane.BootArmorPath)
if err != nil {
t.Fatalf("read: %v", err)
}
// It must be the HOLDING plane — a forward chain that drops — and not the full
// tproxy ruleset, which would reference an engine that is not running at boot.
for _, must := range []string{"table inet shater", "hook forward", "drop"} {
if !strings.Contains(string(b), must) {
t.Errorf("the persisted armor must contain %q; got:\n%s", must, string(b))
}
}
if strings.Contains(string(b), "tproxy") {
t.Errorf("the persisted armor must NOT divert to an engine that is not running:\n%s", string(b))
}
openKS := enabledClosed()
openKS.Globals.KillSwitch = "open"
refreshBootArmor(openKS, logger)
if netplane.BootArmorPresent() {
t.Errorf("kill_switch=open is a documented choice to let traffic through; " +
"the boot armor must be removed, not left to block the next boot")
}
refreshBootArmor(enabledClosed(), logger)
if !netplane.BootArmorPresent() {
t.Fatalf("re-arming after a disarm must work")
}
off := enabledClosed()
off.Globals.Enabled = false
refreshBootArmor(off, logger)
if netplane.BootArmorPresent() {
t.Errorf("globals.enabled=0 must remove the boot armor")
}
}
// TestRestartHandoffPending pins the marker's read side, including the default
// that matters: an ABSENT marker means "a real stop". Defaulting the other way
// would leave a stopped router blocked whenever the init script failed to write
// the flag.
func TestRestartHandoffPending(t *testing.T) {
orig := restartHandoffPath
defer func() { restartHandoffPath = orig }()
dir := t.TempDir()
restartHandoffPath = filepath.Join(dir, "shater.restarting")
if restartHandoffPending() {
t.Errorf("an absent marker must read as a real stop")
}
if err := os.WriteFile(restartHandoffPath, nil, 0o644); err != nil {
t.Fatalf("write marker: %v", err)
}
if !restartHandoffPending() {
t.Errorf("a present marker must read as a restart handoff")
}
}
+169
View File
@@ -0,0 +1,169 @@
// Executing the SHIPPED init script's action classification.
//
// WHY THIS TEST IS SHAPED LIKE THIS
//
// The boot-armor defect that shipped in v0.2.17 was not in any Go file. The Go
// half was correct and fully covered: refreshBootArmor tracked desired state,
// planExitArmor made the right call, LoadBootArmor validated before loading.
// Every one of those tests was green while the feature did not work at all on
// hardware, because the thing that broke it was one shell `case` in
// /etc/init.d/shater whose default arm swept up procd's `shutdown` action — so
// the arm token was deleted on the way down, every reboot, and the boot it
// existed to protect always found no file.
//
// A unit test that cannot see the shell file cannot catch that, and a comment in
// the shell file claiming `shutdown` is handled is precisely what shipped. So
// this runs the real thing: it sources the actual packaged
// openwrt/shater-core/files/etc/init.d/shater in /bin/sh and calls its two
// classification predicates with every action procd actually uses.
//
// Sourcing the whole file is safe and deliberate — at top level it contains only
// variable assignments and function definitions, nothing that touches the system —
// and sourcing the WHOLE file is the point: a test that copy-pasted the `case`
// would pass while the shipped script said something else.
//
// The action names are not invented. They were measured on the target
// (ImmortalWrt 25.12.1 r37978) with a throwaway probe init script:
//
// /etc/init.d/X restart -> stop_service action=[restart]
// /etc/init.d/X stop -> stop_service action=[stop]
// /etc/init.d/X reload -> reload_service action=[reload]
// `reboot` -> stop_service action=[shutdown]
// the boot after it -> start_service action=[boot]
package main
import (
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
)
// initScriptPath is the packaged init script, relative to this package dir.
const initScriptPath = "../../../openwrt/shater-core/files/etc/init.d/shater"
// askInitScript sources the init script in /bin/sh and reports whether fn
// returns true for the given arguments.
func askInitScript(t *testing.T, fn string, args ...string) bool {
t.Helper()
abs, err := filepath.Abs(initScriptPath)
if err != nil {
t.Fatalf("resolve %s: %v", initScriptPath, err)
}
// `. script` then call the predicate. `set -e` is deliberately NOT used: the
// predicates report by exit status, and a false answer is not an error.
script := `. "$1" || exit 3; shift; if ` + fn + ` "$@"; then echo yes; else echo no; fi`
argv := append([]string{"-c", script, "sh", abs}, args...)
out, err := exec.Command("/bin/sh", argv...).CombinedOutput()
if err != nil {
t.Fatalf("%s(%q): %v\n%s", fn, args, err, out)
}
switch strings.TrimSpace(string(out)) {
case "yes":
return true
case "no":
return false
default:
t.Fatalf("%s(%q): unreadable answer %q", fn, args, out)
return false
}
}
// TestInitScriptActionClassification pins the two closed lists. The `shutdown`
// rows are the regression: both must be false, because a reboot is neither a
// handoff (nothing is coming) nor an operator switching the product off.
func TestInitScriptActionClassification(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, tc := range []struct {
action string
disarms bool
handoff bool
why string
}{
{"stop", true, false, "the operator switched the product off"},
{"shutdown", false, false, "REBOOT/POWEROFF — must not disarm; this is the boot the armor exists for"},
{"restart", false, true, "a successor is coming"},
{"reload", false, true, "Save & Apply is stop+start"},
{"boot", false, false, "start side, never reaches stop_service"},
{"start", false, false, "start side"},
{"", false, false, "unknown/empty degrades to changing nothing"},
{"enable", false, false, "not a lifecycle transition"},
{"disable", false, false, "durable off, but handled by shater-armor's rc.d refusal, not here"},
} {
if got := askInitScript(t, "shater_action_disarms", tc.action); got != tc.disarms {
t.Errorf("shater_action_disarms(%q) = %v, want %v (%s)", tc.action, got, tc.disarms, tc.why)
}
if got := askInitScript(t, "shater_action_handoff", tc.action); got != tc.handoff {
t.Errorf("shater_action_handoff(%q) = %v, want %v (%s)", tc.action, got, tc.handoff, tc.why)
}
}
}
// TestInitScriptStopDisarmsOnlyForAPerson pins the rest of the decision: `stop`
// disarms when a PERSON is behind it, or when the product is being removed — and
// not when something is merely replacing it.
//
// base-files' default_prerm reaches stop_service as a plain `stop`:
//
// if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
// "$i" stop
//
// so two different intentions arrive as one action. The rc.d state separates them:
// a removal has already run `disable`, a replacement has not.
//
// On this target an apk UPGRADE turns out never to run default_prerm at all
// (no pre-upgrade script — verified with a real `apk fix --reinstall` while
// sampling the armor file), so the upgrade rows below are defence-in-depth rather
// than a reproduction. The removal rows are live behaviour.
func TestInitScriptStopDisarmsOnlyForAPerson(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, tc := range []struct {
action string
inPkg string
rcEnable string
want bool
why string
}{
{"stop", "0", "1", true, "an operator typed it — the escape hatch must keep working"},
{"stop", "0", "0", true, "an operator typed it on an already-disabled service"},
{"stop", "1", "1", false, "BEING REPLACED — prerm left the service enabled, so something is coming back"},
{"stop", "1", "0", true, "REMOVAL — prerm already ran `disable`; the product is going away"},
{"shutdown", "0", "1", false, "reboot never disarms"},
{"shutdown", "1", "1", false, "reboot never disarms, package manager or not"},
{"shutdown", "1", "0", false, "still a reboot; the action decides first"},
{"restart", "0", "1", false, "a successor is coming"},
{"restart", "1", "0", false, "a successor is coming; action decides before any state"},
{"reload", "1", "1", false, "Save & Apply"},
{"", "1", "0", false, "unknown action changes nothing"},
} {
got := askInitScript(t, "shater_stop_disarms", tc.action, tc.inPkg, tc.rcEnable)
if got != tc.want {
t.Errorf("shater_stop_disarms(%q, in_pkg=%s, rc_enabled=%s) = %v, want %v (%s)",
tc.action, tc.inPkg, tc.rcEnable, got, tc.want, tc.why)
}
}
}
// TestInitScriptsParse is the cheapest possible guard against the class of bug
// that no Go test can otherwise see: a shell file that ships syntactically
// broken. An init script that fails to parse takes the whole service down and
// `go build` is perfectly happy about it.
func TestInitScriptsParse(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, name := range []string{"shater", "shater-armor", "shater-cron"} {
p, err := filepath.Abs(filepath.Join(filepath.Dir(initScriptPath), name))
if err != nil {
t.Fatalf("resolve %s: %v", name, err)
}
if out, err := exec.Command("/bin/sh", "-n", p).CombinedOutput(); err != nil {
t.Errorf("/etc/init.d/%s does not parse: %v\n%s", name, err, out)
}
}
}
+234 -20
View File
@@ -16,7 +16,8 @@
// status | nodes | stats read-side JSON for LuCI/panel
// blocklist update force a DNS-filter refresh (reconcile; no-op if down)
// schedule due re-evaluate time-scheduled rules now (reconcile; no-op if down)
// sub update | ruleset update Phase-2b no-op stubs (cron-safe)
// sub update [<name>] fetch + re-cache subscription nodes (real; exits 1 on failure)
// ruleset update NOT IMPLEMENTED — exits non-zero, does nothing
//
// See docs-shater/PORTING.md "Wave 3 daemon contract" and DECISIONS D11/D12.
package main
@@ -51,7 +52,11 @@ import (
"github.com/sagernet/sing-box/shater/subscribe"
)
const pidfilePath = "/var/run/shaterd.pid"
// pidfilePath is where the daemon publishes its pid and where every other verb
// looks for it. A var, not a const, for the same reason ctlPath is one: the
// daemon-reachable / daemon-absent split is a contract worth testing, and a test
// must be able to stage both without writing to the real /var/run.
var pidfilePath = "/var/run/shaterd.pid"
func main() { os.Exit(dispatch(os.Args[1:])) }
@@ -130,13 +135,30 @@ usage: shaterd <verb>
schedule due re-evaluate time-scheduled rules now (reconcile; no-op if down)
alert test send a test notification to every enabled alert (no daemon needed)
sub update [<name>] fetch + re-cache subscription nodes (all enabled, or one named)
ruleset update (Phase 2b) not yet implemented
ruleset update NOT IMPLEMENTED — exits non-zero without updating anything
`)
}
// notImplExitCode is what an unimplemented verb exits with.
//
// It is NOT 0, and that is the entire point. `ruleset update` used to print a
// note on stderr and exit 0, which made every caller that checks the exit status
// believe the work was done: /etc/init.d/shater-cron runs it as
// `"$SHATERD" ruleset update "$name" >/dev/null 2>&1` and, on success, STAMPS the
// item as freshly updated and sets changed=1 — so every `config ruleset` with
// source=url was permanently reported up to date by a verb that never fetched a
// byte, and each stamp also triggered a reconcile with an unchanged config.
//
// It is also NOT 2: `dispatch` reserves 2 for "I do not know this verb" (usage),
// and a caller must be able to tell a typo from a verb that exists but does not
// work yet.
const notImplExitCode = 1
func notImpl(verb string) int {
fmt.Fprintf(os.Stderr, "shaterd: %s not yet implemented (Phase 2b)\n", verb)
return 0
fmt.Fprintf(os.Stderr, "shaterd: %s is NOT implemented — nothing was updated. "+
"Remote rule-sets are refreshed by the engine itself "+
"(ruleset.update_interval on the running box), not by this verb.\n", verb)
return notImplExitCode
}
// --- run: THE daemon --------------------------------------------------------
@@ -289,10 +311,27 @@ func cmdRun() int {
// reconcile. No-op when no iface-driven profiles are configured.
go watchActiveProfile(applier, logger)
// Keep the persisted fail-closed plane in step with the config BEFORE anything
// is attempted: it is what protects the LAN at the NEXT boot (and across the
// next restart), and an engine start that hangs for a minute must not be what
// stands between a config change and its armor being written.
if readErr == nil {
refreshBootArmor(m, logger)
}
// Initial apply. A failed initial apply must NOT crash-loop the box into a
// blackout: log it and stay up so a later SIGHUP/apply can fix the config.
if readErr != nil {
logger.Error("initial ReadUCI failed (staying up): ", readErr)
// ...but STAYING UP IS NOT THE SAME AS BEING SAFE. This branch used to end
// here, which meant an unreadable /etc/config/shater — a full /overlay caught
// mid `uci commit` is the documented cause — left the router with no table at
// all, permanently: nothing else installs one (Reconcile returns before any
// plane work, and the cron watchdog's own `uci -q get` fails identically), and
// the panel reported a live daemon the whole time. The persisted holding plane
// is the last thing this router is KNOWN to have wanted, and it needs nothing
// readable to be true.
armOnUnreadableConfig(logger)
} else if m.Globals.Enabled {
// ARM FIRST, THEN TRY. Install the fail-closed holding plane BEFORE the
// engine is attempted, so the gap between daemon start and a working engine
@@ -369,9 +408,20 @@ func cmdRun() int {
// learn whether the plane is meant to be up (globals.enabled) so a
// reconcile error can raise the kill-switch alert.
enabled := false
if mm, e := model.ReadUCI(); e == nil {
if mm, e := model.ReadUCI(); e != nil {
// The live config just became unreadable. Same answer as at startup:
// reinstate what this router last persisted, rather than run on with
// whatever the kernel happens to hold.
logger.Error("reconcile: could not read the config: ", e)
armOnUnreadableConfig(logger)
} else {
notifier.Update(mm.Alerts)
enabled = mm.Globals.Enabled
// Keep the persisted fail-closed plane in step with the config the
// operator just changed — including the disarm half, so turning the
// stack off (or opening the kill switch) also stops the next boot from
// blocking the LAN.
refreshBootArmor(mm, logger)
// Pick up a changed stats backend / sizing without a daemon restart.
// A no-op when nothing changed, so a routine reconcile never churns
// the store (and never restarts its log cursors).
@@ -393,8 +443,31 @@ func cmdRun() int {
// DNS-query manager). Re-point the aggregator at the current manager.
statsAgg.Resubscribe()
case syscall.SIGTERM, syscall.SIGINT:
logger.Info("signal ", sig, ": honest teardown + exit")
if err := applier.Teardown(); err != nil {
// Being REPLACED is not the same as being switched off, and until now
// this path could not tell the difference: Teardown does not consult
// kill_switch at all (compare holdLocked, which does), so `restart`,
// `reload_service` — which is stop+start, i.e. every LuCI Save & Apply —
// and every package upgrade dismantled the fail-closed plane and left the
// LAN forwarding in the clear for as long as the successor needed to come
// up. That interval is guaranteed non-empty by the init itself: it waits
// for this process to exit, then runs `shaterd migrate`, then starts the
// daemon, which then has to build an engine.
//
// The init script announces a restart with a tmpfs marker; absent it, this
// is a deliberate stop and the plane goes away for good, which is the
// operator's escape hatch and must keep working.
handoff := restartHandoffPending()
logger.Info("signal ", sig, ": honest teardown + exit (restart handoff: ", handoff, ")")
// BEFORE the teardown, not after. This used to read `Teardown(); armOnExit()`
// on the reasoning that Teardown deletes the table so arming first would be
// undone — true, and the wrong conclusion: it left a measured 80-90 ms window per
// restart (80-90 ms measured) in which no `inet shater` table existed at all and fw4's
// `lan -> wan ACCEPT` was the only policy on the box. TeardownExiting arms
// first (one nft transaction that REPLACES the table) and then skips the
// delete iff a plane really went in.
if err := applier.TeardownExiting(func() bool {
return armOnExit(handoff, logger)
}); err != nil {
logger.Error("teardown: ", err)
}
return 0
@@ -817,11 +890,84 @@ func cmdScheduleDue() int {
return 0
}
// statusDaemonAnsweredKey is the ONE field of `shaterd status` that says, BY
// CONTRACT, where the object under it came from:
//
// true — a running daemon answered over the control socket; every other field
// is that daemon's own Applier.Status().
// false — no daemon answered. What follows is the OFFLINE STUB: the apply.Status
// zero value plus what could be read from UCI and from the kernel. It is
// NOT a status report, and the fields that only a live daemon can know
// (running/engine_running/plane/traffic/hash/warnings/uptime) are
// placeholders, not measurements.
//
// It exists because the stub used to be indistinguishable from an answer: it is
// the same struct, printed by the same marshaller, and `shaterd status` exited 0
// either way. The one thing that happened to differ was `plane` being "" — a
// value a live Applier.Status() cannot produce because it always assigns one of
// three words — and luci-app-shater's dashboard.js was forced to key its
// "daemon: down" verdict off exactly that. Detection by a side effect is not a
// contract: filling `plane` in the stub, for any reason, would silently turn
// "the daemon is dead" into "the daemon is fine" on the LuCI page.
//
// The field is ADDITIVE. Every pre-existing key keeps its name, its value and
// its position, so no current consumer breaks; `plane: ""` in particular is
// still emitted, deliberately, so dashboard.js keeps working until it is moved
// onto this field.
const statusDaemonAnsweredKey = "daemon_answered"
// markStatusOrigin prefixes a status JSON object with the daemon_answered
// verdict, leaving every other key untouched and in place.
//
// It splices rather than re-marshals on purpose: the live branch relays the
// DAEMON's own words, and a decode/encode round-trip would silently reorder
// them and drop any field this build does not know about (a newer daemon's, an
// older CLI's). The input is validated as a JSON object first, so the splice is
// well defined; a reply that is not an object is an error, never something we
// print unmarked.
func markStatusOrigin(obj string, answered bool) (string, error) {
t := strings.TrimSpace(obj)
var probe map[string]json.RawMessage
if err := json.Unmarshal([]byte(t), &probe); err != nil {
return "", fmt.Errorf("status reply is not a JSON object: %w", err)
}
if _, dup := probe[statusDaemonAnsweredKey]; dup {
return "", fmt.Errorf("status reply already carries %q", statusDaemonAnsweredKey)
}
field := `"` + statusDaemonAnsweredKey + `": ` + strconv.FormatBool(answered)
rest := t[1:] // Unmarshal succeeded and t is trimmed, so t[0] is '{'
switch {
case strings.HasPrefix(rest, "\n"):
return "{\n " + field + "," + rest, nil // apply.Status.JSON() is indented
case strings.TrimSpace(rest) == "}":
return "{" + field + "}", nil // the empty object
default:
return "{" + field + "," + rest, nil
}
}
// cmdStatus prints the daemon's status as JSON.
//
// Three outcomes, and a caller can tell them apart without guessing:
//
// exit 0 — a daemon answered; stdout carries its status with
// "daemon_answered": true.
// exit 1 + stdout — no daemon answered; stdout carries the OFFLINE STUB with
// "daemon_answered": false (see statusDaemonAnsweredKey).
// exit 1 + NO stdout — the process is alive but wedged. Printing the stub here
// would assert running=false about a daemon that is running, and would
// hide a data plane that is very probably still installed and still
// enforcing.
func cmdStatus() int {
if _, ok := daemonAlive(); ok {
resp, err := ctlRequest("status")
if err == nil {
fmt.Println(strings.TrimSpace(resp))
out, merr := markStatusOrigin(resp, true)
if merr != nil {
fmt.Fprintf(os.Stderr, "shaterd status: %v\n", merr)
return 1
}
fmt.Println(out)
return 0
}
// The daemon PROCESS exists but did not answer in time. Falling through to
@@ -835,7 +981,13 @@ func cmdStatus() int {
return 1
}
}
// Offline stub: the daemon isn't reachable, so running=false / hash="".
// OFFLINE STUB. No daemon answered, so the only honest fields are the ones
// read from the kernel and from UCI right here; everything else is the
// apply.Status zero value and is marked as such by daemon_answered=false.
//
// `plane` stays "" on purpose: it is the signal luci-app-shater currently
// keys "daemon: down" off, and dropping it would break that page silently.
// It is now a COMPATIBILITY carry-over, not the contract.
s := apply.Status{
Running: false,
Active: apply.ActiveFlagPresent(),
@@ -851,8 +1003,18 @@ func cmdStatus() int {
}
}
b, _ := s.JSON()
fmt.Println(string(b))
return 0
out, merr := markStatusOrigin(string(b), false)
if merr != nil {
fmt.Fprintf(os.Stderr, "shaterd status: %v\n", merr)
return 1
}
// stdout keeps carrying a parseable object (the rpcd plugin reads stdout and
// ignores the exit status), but the exit code no longer reports success for a
// status nobody produced.
fmt.Println(out)
fmt.Fprintln(os.Stderr, "shaterd status: the daemon did not answer — the object on "+
"stdout is the offline stub (daemon_answered=false), not a status report.")
return 1
}
// cmdNodes prints the node inventory as a JSON array (see nodes.go for the
@@ -991,16 +1153,18 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
writeLine(conn, string(b))
case "reconcile":
ch, err := a.Reconcile()
// The panel and the CLI reconcile over this socket, not by SIGHUP, so the
// persisted fail-closed plane has to be refreshed here too — otherwise
// enabling the stack from the panel left the next boot unarmed (and
// DISABLING it left the next boot armed) until some unrelated SIGHUP
// happened along.
if mm, e := model.ReadUCI(); e == nil {
refreshBootArmor(mm, l)
}
writeResult(conn, ch, err)
case "apply":
a.Snapshot()
changed, err := a.Reconcile()
if err == nil {
if m, e := model.ReadUCI(); e == nil {
a.ArmRollback(m.Globals.ConfirmTimeout)
}
}
writeResult(conn, changed, err)
b, _ := json.Marshal(runApplyVerb(a, l))
writeLine(conn, string(b))
case "confirm":
writeResult(conn, false, a.Confirm())
case "rollback":
@@ -1045,6 +1209,49 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
}
}
// runApplyVerb is the daemon side of `shaterd apply`: snapshot the rollback
// target, reconcile, arm the commit-confirm window — and then REPORT whether a
// safety net actually exists. The arming behaviour is byte-for-byte what it was
// (Snapshot, Reconcile, ArmRollback only on success and only with a readable UCI,
// refreshBootArmor either way); what is new is that the answer says what happened.
// See applyverb.go for why silence here cost a house's connectivity.
func runApplyVerb(a *apply.Applier, l log.ContextLogger) applyVerdict {
// The rollback target this apply is about to record. Snapshot() reads the same
// LastGood() under the apply mutex; reading it here first is the only way to see
// it, and the two reads cannot disagree in practice — an Apply holds that mutex
// for its whole (multi-second) duration, so both calls either precede it or
// follow it.
before := a.LastGood()
a.Snapshot()
changed, err := a.Reconcile()
f := applyFacts{
engineChanged: changed,
err: err,
before: before,
after: a.LastGood(),
configMTime: uciConfigMTime(),
daemonStart: daemonStarted,
}
if m, e := model.ReadUCI(); e == nil {
f.uciRead = true
f.enabled = m.Globals.Enabled
f.confirmTimeout = m.Globals.ConfirmTimeout
if err == nil {
a.ArmRollback(m.Globals.ConfirmTimeout)
}
refreshBootArmor(m, l)
}
v := classifyApply(f)
// The CLI answer is read once; logread is what an incident is reconstructed
// from. A missing safety net is a WARNING there, not an info line.
if v.RollbackArmed {
l.Info("apply (", v.Reason, "): ", v.Message)
} else {
l.Warn("apply (", v.Reason, "): NO automatic rollback armed: ", v.Message)
}
return v
}
func writeResult(conn net.Conn, changed bool, err error) {
r := ctlResult{Changed: changed}
if err != nil {
@@ -1081,6 +1288,13 @@ func ctlMutate(cmd string, requireDaemon bool) int {
return 1
}
fmt.Println(strings.TrimSpace(resp))
// `apply` answers with more than {changed,error}: it says whether a
// commit-confirm window was actually armed. That is the entire reason the verb
// exists, so it is also said in plain words on stderr — a one-line JSON object
// on stdout is exactly what let a missing safety net pass for a healthy apply.
if line := applyNoticeLine(cmd, resp); line != "" {
fmt.Fprintln(os.Stderr, line)
}
return 0
}
+284
View File
@@ -0,0 +1,284 @@
//go:build linux
package main
import (
"bufio"
"encoding/json"
"io"
"net"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
)
// What this file pins
//
// 1. `shaterd status` must let its caller tell "a daemon answered" from "no
// daemon answered" BY CONTRACT. Until now the offline stub was the same
// struct printed by the same marshaller with exit 0, and the only thing that
// happened to differ was `plane` being "" — a side effect luci-app-shater was
// forced to key its "daemon down" verdict off.
// 2. `shaterd ruleset update` must not report success for work it does not do.
// shater-cron stamps the item as freshly updated on a zero exit.
// --- helpers ---------------------------------------------------------------
// captureStdout runs fn with os.Stdout redirected and returns what it printed.
func captureStdout(t *testing.T, fn func() int) (string, int) {
t.Helper()
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
orig := os.Stdout
os.Stdout = w
done := make(chan string, 1)
go func() {
b, _ := io.ReadAll(r)
done <- string(b)
}()
code := fn()
_ = w.Close()
os.Stdout = orig
out := <-done
_ = r.Close()
return out, code
}
// noDaemon points the pid lookup at a path that cannot name a live process, so
// cmdStatus takes the offline branch. Restored by t.Cleanup.
func noDaemon(t *testing.T) {
t.Helper()
orig := pidfilePath
pidfilePath = filepath.Join(t.TempDir(), "absent.pid")
t.Cleanup(func() { pidfilePath = orig })
}
// liveDaemon stages a reachable daemon: a pidfile naming THIS process (so
// daemonAlive's kill(pid,0) probe succeeds) and a control socket that answers a
// single `status` request with reply. Restored by t.Cleanup.
func liveDaemon(t *testing.T, reply string) {
t.Helper()
dir := t.TempDir()
origPid := pidfilePath
pidfilePath = filepath.Join(dir, "shaterd.pid")
if err := os.WriteFile(pidfilePath, []byte(strconv.Itoa(os.Getpid())+"\n"), 0o644); err != nil {
t.Fatalf("write pidfile: %v", err)
}
t.Cleanup(func() { pidfilePath = origPid })
origCtl := ctlPath
ctlPath = filepath.Join(dir, "shaterd.ctl")
ln, err := net.Listen("unix", ctlPath)
if err != nil {
t.Fatalf("listen %s: %v", ctlPath, err)
}
t.Cleanup(func() { _ = ln.Close(); ctlPath = origCtl })
go func() {
for {
conn, err := ln.Accept()
if err != nil {
return
}
go func(c net.Conn) {
defer c.Close()
line, _ := bufio.NewReader(c).ReadString('\n')
if strings.TrimSpace(line) != "status" {
_, _ = c.Write([]byte(`{"error":"unknown command"}` + "\n"))
return
}
_, _ = c.Write([]byte(reply + "\n"))
}(conn)
}
}()
}
// liveStatusReply is a realistic Applier.Status() answer: indented (Status.JSON
// uses MarshalIndent) and carrying a plane word only a live daemon can emit.
const liveStatusReply = `{
"running": true,
"enabled": true,
"active": true,
"table": true,
"hash": "abc123",
"kill_switch": "closed",
"panel_port": 8088,
"can_rollback": false,
"engine_running": true,
"plane": "full",
"traffic": {"verdict": "tunnelled"},
"warnings": [],
"started_unix": 1753500000,
"uptime_seconds": 42
}`
// --- defect 1: the offline stub must not pass for a status ------------------
// TestStatusOfflineStubIsMarkedAndFails is the whole contract of the daemon-down
// case: stdout still carries a parseable object (the rpcd plugin reads stdout and
// ignores the exit status, so removing it would blank the LuCI page), but that
// object SAYS it is not a daemon's answer, and the exit code says so too.
func TestStatusOfflineStubIsMarkedAndFails(t *testing.T) {
noDaemon(t)
out, code := captureStdout(t, cmdStatus)
if code == 0 {
t.Fatal("`shaterd status` exited 0 with no daemon to ask — a caller that " +
"checks the exit status is told the status is real")
}
var got map[string]any
if err := json.Unmarshal([]byte(out), &got); err != nil {
t.Fatalf("stdout is not a JSON object (%v):\n%s", err, out)
}
v, present := got[statusDaemonAnsweredKey]
if !present {
t.Fatalf("the offline stub carries no %q field, so it is indistinguishable "+
"from a daemon's answer except by guesswork:\n%s", statusDaemonAnsweredKey, out)
}
if v != false {
t.Fatalf("%s = %v, want false for the offline stub", statusDaemonAnsweredKey, v)
}
// The compatibility carry-over luci-app-shater currently keys off. It is
// deliberately still emitted; if it ever goes away, dashboard.js has to be
// moved onto daemon_answered FIRST.
if plane, ok := got["plane"]; !ok || plane != "" {
t.Fatalf(`the stub no longer emits plane:"" (got %v, present=%v) — `+
"luci-app-shater's dashboard.js derives \"daemon: down\" from exactly "+
"that and would go silent; move it onto %q before removing it",
got["plane"], ok, statusDaemonAnsweredKey)
}
}
// TestStatusLiveAnswerIsMarkedAndRelayedVerbatim is the CONTROL for the test
// above: the same command, the same field, the opposite verdict. Without it,
// "daemon_answered is false" proves nothing — a build that hardcoded false would
// pass. It also pins that the daemon's own words are relayed unchanged: the CLI
// adds one key and reorders/drops nothing, including fields this build has never
// heard of.
func TestStatusLiveAnswerIsMarkedAndRelayedVerbatim(t *testing.T) {
liveDaemon(t, liveStatusReply)
out, code := captureStdout(t, cmdStatus)
if code != 0 {
t.Fatalf("`shaterd status` exited %d while a daemon answered", code)
}
var got map[string]any
if err := json.Unmarshal([]byte(out), &got); err != nil {
t.Fatalf("stdout is not a JSON object (%v):\n%s", err, out)
}
if v, ok := got[statusDaemonAnsweredKey]; !ok || v != true {
t.Fatalf("%s = %v (present=%v), want true for a daemon's answer:\n%s",
statusDaemonAnsweredKey, got[statusDaemonAnsweredKey], ok, out)
}
var want map[string]any
if err := json.Unmarshal([]byte(liveStatusReply), &want); err != nil {
t.Fatalf("fixture: %v", err)
}
for k, wv := range want {
gv, ok := got[k]
if !ok {
t.Errorf("the CLI dropped the daemon's %q field", k)
continue
}
wb, _ := json.Marshal(wv)
gb, _ := json.Marshal(gv)
if string(wb) != string(gb) {
t.Errorf("the CLI changed the daemon's %q: got %s, daemon said %s", k, gb, wb)
}
}
if len(got) != len(want)+1 {
t.Errorf("the CLI added more than the one verdict field: %d keys vs the daemon's %d",
len(got), len(want))
}
// Field order is not JSON semantics, but it IS what an operator reads first.
if !strings.HasPrefix(strings.TrimSpace(out), `{`+"\n \""+statusDaemonAnsweredKey+`"`) {
t.Errorf("the verdict is not the first field of the object:\n%s", out)
}
}
// TestMarkStatusOriginRefusesWhatItCannotMark: a reply that is not a JSON object
// must NOT be printed unmarked. Printing it would put an unlabelled blob on
// stdout, which is the exact ambiguity this field exists to remove.
func TestMarkStatusOriginRefusesWhatItCannotMark(t *testing.T) {
for _, bad := range []string{"", "not json", "[1,2,3]", `"a string"`, "{", `{"a":}`} {
if out, err := markStatusOrigin(bad, true); err == nil {
t.Errorf("markStatusOrigin(%q) accepted it and produced %q", bad, out)
}
}
// A reply that already carries the key is refused too: silently keeping the
// daemon's value would let a future daemon assert its own reachability.
if _, err := markStatusOrigin(`{"`+statusDaemonAnsweredKey+`":false}`, true); err == nil {
t.Error("a reply that already carries the verdict field was accepted")
}
}
// TestMarkStatusOriginShapes covers the object shapes the two producers emit.
func TestMarkStatusOriginShapes(t *testing.T) {
for _, tc := range []struct {
name, in string
answered bool
}{
{"empty object", `{}`, false},
{"compact", `{"a":1,"b":"x"}`, true},
{"indented", "{\n \"a\": 1\n}", false},
} {
t.Run(tc.name, func(t *testing.T) {
out, err := markStatusOrigin(tc.in, tc.answered)
if err != nil {
t.Fatalf("markStatusOrigin: %v", err)
}
var got map[string]any
if err := json.Unmarshal([]byte(out), &got); err != nil {
t.Fatalf("result is not valid JSON (%v): %s", err, out)
}
if got[statusDaemonAnsweredKey] != tc.answered {
t.Fatalf("%s = %v, want %v", statusDaemonAnsweredKey,
got[statusDaemonAnsweredKey], tc.answered)
}
})
}
}
// --- defect 3: an unimplemented verb must not report success ----------------
// TestRulesetUpdateExitsNonZero. /etc/init.d/shater-cron runs
// `"$SHATERD" ruleset update "$name" >/dev/null 2>&1` and, on a ZERO exit,
// stamps the ruleset as freshly updated and sets changed=1. With the old exit 0
// every url ruleset was permanently "just updated" by a verb that fetched
// nothing, and every stamp triggered a reconcile with an identical config.
func TestRulesetUpdateExitsNonZero(t *testing.T) {
if code := dispatch([]string{"ruleset", "update"}); code == 0 {
t.Fatal("`shaterd ruleset update` exited 0 without updating anything — " +
"shater-cron reads that as a successful fetch and stamps the item")
}
if code := dispatch([]string{"ruleset", "update", "geosite"}); code == 0 {
t.Fatal("`shaterd ruleset update <name>` exited 0 without updating anything")
}
}
// TestNotImplNeverReportsSuccess guards the helper itself, so a verb added to it
// later cannot inherit a zero exit. It also pins that "not implemented" is
// distinguishable from "no such verb" (usage, exit 2) — a caller must be able to
// tell a typo from a stub.
func TestNotImplNeverReportsSuccess(t *testing.T) {
if code := notImpl("some future verb"); code == 0 {
t.Fatal("notImpl reports success")
}
if notImplExitCode == 2 {
t.Fatal("the not-implemented exit code collides with the usage exit code (2), " +
"so a caller cannot tell an unknown verb from an unimplemented one")
}
if code := dispatch([]string{"ruleset", "frobnicate"}); code != 2 {
t.Fatalf("an unknown ruleset sub-verb exited %d, want the usage code 2", code)
}
if code := dispatch([]string{"nosuchverb"}); code != 2 {
t.Fatalf("an unknown verb exited %d, want the usage code 2", code)
}
}
+67 -10
View File
@@ -40,6 +40,7 @@ import (
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/route"
"github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing-box/shater/netplane"
"github.com/sagernet/sing-box/shater/registry"
E "github.com/sagernet/sing/common/exceptions"
"github.com/sagernet/sing/common/json"
@@ -73,6 +74,13 @@ type Engine struct {
// PendingCloses when a shutdown overruns its budget.
instanceGen uint64
// l3Device is the L3-ingress TUN slot the RUNNING instance opened, or "" when
// nothing is running (or the running config has no L3 ingress). It is the
// input to the slot choice for the next generation — the one name that
// generation may not take — and it is what the netplane is told to point the
// L3 routing table at. See l3slot.go and netplane/l3.go. Guarded by mu.
l3Device string
// pending holds the retirements in flight (see teardown.go). Its OWN leaf
// lock, deliberately not mu: PendingCloses is read by the status path, and
// the moment that read matters most is while an apply is holding mu waiting
@@ -246,7 +254,7 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// (2) build + validate on its OWN cancellable context. box.New constructs and
// validates every adapter.
nb, nbCancel, err := e.newBox(opts)
nb, nbCancel, nbDev, err := e.newBox(opts)
if err != nil {
// Validation failed: keep the running instance, do not swap.
return false, E.Cause(err, "create instance")
@@ -307,7 +315,7 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// an ERROR line naming it — visible, but not mistaken for a failed apply.
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
}
e.adoptLocked(nb, nbCancel, opts, newHash)
e.adoptLocked(nb, nbCancel, opts, newHash, nbDev)
return true, nil
}
@@ -315,7 +323,15 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// returns the cancel alongside it. Every goroutine the box starts inherits that
// context, so cancelling it is what unwinds the ones Close does not reach; see
// teardown.go for why the shared, never-cancelled context was the defect.
func (e *Engine) newBox(opts option.Options) (*box.Box, context.CancelFunc, error) {
//
// It also picks the L3-ingress TUN slot this box will open and returns it, so
// the caller can publish it on adoption. The generated config carries only a
// placeholder name; substituting it HERE — after the caller hashed the canonical
// options, and once per box actually built — is what stops two generations from
// contending for one TUN device. See l3slot.go. The returned device is "" for
// every config without an L3 ingress, which is all of them by default.
func (e *Engine) newBox(opts option.Options) (*box.Box, context.CancelFunc, string, error) {
opts, dev := l3RetargetForNext(opts, e.l3Device)
ctx, cancel := context.WithCancel(e.ctx)
b, err := box.New(box.Options{
Context: ctx,
@@ -324,20 +340,39 @@ func (e *Engine) newBox(opts option.Options) (*box.Box, context.CancelFunc, erro
})
if err != nil {
cancel()
return nil, nil, err
return nil, nil, "", err
}
return b, cancel, nil
return b, cancel, dev, nil
}
// adoptLocked installs a started box as THE running instance and gives it the
// next generation number. Caller holds e.mu and has already retired whatever was
// running before.
func (e *Engine) adoptLocked(b *box.Box, cancel context.CancelFunc, opts option.Options, hash string) {
//
// opts is the CANONICAL (placeholder-named) config — the one the hash describes
// and the one a later reconcile is compared against. dev is the L3 slot the box
// really opened, kept apart from opts for exactly that reason, and published to
// the netplane so ApplyRouting points the L3 table at the device that exists
// rather than at a name it guessed.
func (e *Engine) adoptLocked(b *box.Box, cancel context.CancelFunc, opts option.Options, hash string, dev string) {
e.instanceGen++
e.instance = b
e.instanceCancel = cancel
e.current = opts
e.hash = hash
e.setL3DeviceLocked(dev)
}
// setL3DeviceLocked records the L3 slot the running instance holds and publishes
// it to the netplane. Caller holds e.mu.
//
// The two live together on purpose: they answered different questions once (the
// engine's "which name may the next generation not take" and the netplane's
// "which device does the route point at") and any state where they disagree is a
// state where one of them is lying about the same device.
func (e *Engine) setL3DeviceLocked(dev string) {
e.l3Device = dev
netplane.RememberL3Device(dev)
}
// closeOldThenStart is the close-old-then-start-new swap. It is taken both
@@ -377,9 +412,17 @@ func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.Cance
prevHasLastGood := e.hasLastGood
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, prevHash)
e.instance, e.instanceCancel = nil, nil
// Nothing is running any more, so no L3 slot is spoken for. Clearing this
// BEFORE the rebuilds below is what gives them a real choice: with the old
// generation's slot still recorded, the rebuild and the restore would each
// have exactly one candidate left, and the restore's would be the slot the
// failed rebuild just released. Cleared, the slot picker takes whichever the
// kernel says is actually free — which after a failed Start is the one that
// has been free the longest. See netplane.L3SlotFor.
e.setL3DeviceLocked("")
// (c) Build a FRESH box for opts (the discarded one cannot be reused).
nb2, cancel2, err := e.newBox(opts)
nb2, cancel2, dev2, err := e.newBox(opts)
if err == nil {
err = nb2.Start()
if err != nil {
@@ -391,14 +434,23 @@ func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.Cance
// (d) Success: the old config we just closed becomes last-good.
e.lastGood = prevOpts
e.hasLastGood = true
e.adoptLocked(nb2, cancel2, opts, newHash)
e.adoptLocked(nb2, cancel2, opts, newHash, dev2)
return true, nil
}
// (e) The fresh box could not come up and the old one is already closed —
// interception is currently down. Try to RESTORE the previous config so we
// do not leave the tunnel dead.
rb, rcancel, rerr := e.newBox(prevOpts)
//
// This rebuild goes through newBox like every other, which means it gets its
// own L3 slot rather than the name prevOpts was originally started under.
// That is the whole point: this path used to fail for the SAME reason it was
// entered — a fixed TUN name that the generation we just closed had not
// finished handing back — so the rescue was wired to the resource whose
// contention it was rescuing from, and the measured outcome on the router was
// "start instance failed and could not restore previous config; engine
// stopped" with the kill-switch closed and the LAN dark.
rb, rcancel, rdev, rerr := e.newBox(prevOpts)
if rerr == nil {
rerr = rb.Start()
if rerr != nil {
@@ -409,7 +461,7 @@ func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.Cance
if rerr == nil {
// Old config restored: keep current/hash/last-good exactly as they were
// (do NOT advance them). Report that opts was not applied.
e.adoptLocked(rb, rcancel, prevOpts, prevHash)
e.adoptLocked(rb, rcancel, prevOpts, prevHash, rdev)
e.lastGood = prevLastGood
e.hasLastGood = prevHasLastGood
return false, E.Cause(err, "start instance (config not applied; previous config restored)")
@@ -420,6 +472,7 @@ func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.Cance
// safe (no unproxied leak) even though interception is down.
e.instance, e.instanceCancel = nil, nil
e.hash = ""
e.setL3DeviceLocked("")
return false, E.Cause(E.Errors(err, rerr), "start instance failed and could not restore previous config; engine stopped")
}
@@ -535,6 +588,10 @@ func (e *Engine) Close() error {
err := e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
e.instance, e.instanceCancel = nil, nil
e.hash = ""
// No instance, no slot. Leaving the old name published would make the next
// generation avoid a device nobody holds, and would make the netplane point
// the L3 route at a device that is on its way out.
e.setL3DeviceLocked("")
return err
}
+22 -3
View File
@@ -104,6 +104,18 @@ func (e *Engine) HTTPClient(via string) (*http.Client, error) {
// resolved one (the group test, grouptest.go) guarantee the request cannot end up
// anywhere else — in particular not on the direct outbound, which would report the
// ISP's address as the tunnel's exit address.
//
// # Every caller MUST call CloseIdleConnections on the returned client
//
// A fresh http.Transport is built per call and belongs to that call alone. Its idle
// connections are not ordinary sockets: each is a live proxying session through an
// engine outbound, with a read loop and a write loop of its own, and the DialContext
// closure above captures the outbound OBJECT — so an idle connection keeps a whole
// retired engine generation reachable long after Apply swapped it out and its 5s
// close budget expired. Dropping the client without closing it therefore leaks far
// more than a socket.
//
// grouptest.go and shater/generate/ruleset.go get this right; copy them.
func httpClientVia(ob adapter.Outbound, timeout time.Duration) *http.Client {
transport := &http.Transport{
// Dial the underlying TCP connection through the selected outbound. The
@@ -112,9 +124,16 @@ func httpClientVia(ob adapter.Outbound, timeout time.Duration) *http.Client {
DialContext: func(ctx context.Context, network, addr string) (net.Conn, error) {
return ob.DialContext(ctx, network, M.ParseSocksaddr(addr))
},
ForceAttemptHTTP2: true,
MaxIdleConns: 8,
IdleConnTimeout: 90 * time.Second,
ForceAttemptHTTP2: true,
MaxIdleConns: 8,
// The backstop for a caller that forgets to close, not the intended
// mechanism. 90s (the net/http default this used to carry) is eighteen times
// the engine's 5s close budget, so a single forgotten client could pin a dead
// generation through more than a minute and a half of it. 15s is still ample
// for the reuse this actually buys — a redirect chain or the second request of
// a subscription fetch, both within seconds — while bounding the damage of a
// leak to about one apply cycle.
IdleConnTimeout: 15 * time.Second,
TLSHandshakeTimeout: 15 * time.Second,
ExpectContinueTimeout: 1 * time.Second,
}
+86
View File
@@ -0,0 +1,86 @@
package engine
// Which L3-ingress TUN device a NEW generation opens.
//
// generate emits the L3 inbound with netplane.L3DeviceBase as a placeholder;
// this file substitutes a real slot just before box.New. netplane/l3.go carries
// the full argument for the two-slot design and for why the choice belongs
// here rather than in generate. The short version, because it is the part that
// is easy to "simplify" back into a bug:
//
// - generate runs on every reconcile, including the once-a-minute no-ops, and
// Apply's fast path is a hash of what generate produced. A device name that
// alternated in generate would change that hash every minute and rebuild the
// whole engine — so the name in the CONFIG has to be stable.
// - the name the KERNEL gets must not be stable, because the fresh box and the
// box it replaces are alive at the same time (or the old one is still being
// unregistered), and one name for both is TUNSETIFF EBUSY.
//
// Those two requirements are only compatible if the substitution happens after
// the hash and before box.New. That is exactly here.
//
// The retarget deliberately does NOT mutate the caller's options: the canonical
// (placeholder) form is what the engine stores as e.current and what every later
// hash is compared against, so a mutation would make the next identical
// reconcile look like a change and swap the engine for nothing.
import (
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/netplane"
)
// l3InboundIndex returns the index of the L3-ingress TUN inbound in opts, or -1
// when this config has none (l3_tunnel off, or the inbound was skipped).
//
// It matches on the TUN type AND on the device name being one this project
// owns (netplane.L3IsManagedName) rather than on the tag: the tag is generate's
// private constant, and importing it here would be an import cycle — the
// generate package's own tests import this package.
func l3InboundIndex(opts option.Options) int {
for i, in := range opts.Inbounds {
if in.Type != C.TypeTun {
continue
}
to, ok := in.Options.(*option.TunInboundOptions)
if !ok || !netplane.L3IsManagedName(to.InterfaceName) {
continue
}
return i
}
return -1
}
// l3Retarget returns a copy of opts whose L3-ingress TUN inbound opens dev, and
// reports the device that copy will actually create ("" when opts has no L3
// inbound, in which case opts is returned untouched).
//
// The copy is shallow except for the two things that must not be shared: the
// Inbounds slice (a slice header copy still aliases the backing array) and the
// TunInboundOptions struct the entry points at (it is a pointer behind an
// `any`, so writing through it would reach every holder of the original —
// including e.current, e.lastGood and the applier's own copy).
func l3Retarget(opts option.Options, dev string) (option.Options, string) {
idx := l3InboundIndex(opts)
if idx < 0 {
return opts, ""
}
inbounds := make([]option.Inbound, len(opts.Inbounds))
copy(inbounds, opts.Inbounds)
tun := *(inbounds[idx].Options.(*option.TunInboundOptions))
tun.InterfaceName = dev
inbounds[idx].Options = &tun
opts.Inbounds = inbounds
return opts, dev
}
// l3RetargetForNext rewrites opts to open the slot the NEXT generation may use,
// given the device the currently running generation holds (current, "" when
// nothing is running). It returns the rewritten options and the device chosen.
func l3RetargetForNext(opts option.Options, current string) (option.Options, string) {
if l3InboundIndex(opts) < 0 {
return opts, ""
}
return l3Retarget(opts, netplane.L3SlotFor(current))
}
+223
View File
@@ -0,0 +1,223 @@
package engine
// The slot choice is what stops two generations of the engine from wanting one
// TUN device. Everything here is about the two ways that guarantee can be lost:
// picking the name the running generation holds, and letting the retarget leak
// back into the options the hash gate compares.
import (
"os"
"testing"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/netplane"
)
// TestMain takes THIS TEST BINARY off the real network namespace.
//
// The tests below call l3RetargetForNext for its return value, and that reaches
// netplane.L3SlotFor — which does not merely ask the kernel about a device, it
// DELETES one it finds occupying a candidate slot. `go test` runs package
// binaries concurrently (-p defaults to GOMAXPROCS) and every one of them shares
// the host's network namespace, so on a runner that has iproute2 and a real
// /dev/net/tun these unit tests were deleting the TUN device shater/generate's
// privileged tests had just opened:
//
// post-start inbound/tun[l3-in]: starting TUN interface: find tun interface: Link not found
//
// which is a red gate in a package that did nothing wrong. netplane.L3StubKernelForTest
// carries the measurement and the reasoning; netplane's own
// TestL3StubKernelTakesTheSlotChoiceOffTheKernel is the control that the hook
// still diverts.
//
// The fake kernel starts EMPTY, so every slot reads as free and no reclaim is
// ever attempted from here. That costs this file nothing: the reclaim is
// netplane's subject (TestL3SlotForReclaimsARetiredSlot, against netplane's own
// exec fake), the live-kernel proof is shater/generate's
// TestIntegrationL3StaleSlotIsReclaimed, and what the tests below are about —
// that the slot handed to the next generation is never the running one, and
// that the substitution does not leak into the canonical options — is answered
// by the real L3SlotFor either way.
//
// It is a TestMain rather than a per-test helper deliberately: a helper is
// something the next test added here can forget, and the failure that causes
// lands in a DIFFERENT package, on some runs only.
func TestMain(m *testing.M) {
restore := netplane.L3StubKernelForTest()
code := m.Run()
restore()
os.Exit(code)
}
// l3Opts is a config carrying one L3-ingress TUN inbound exactly as generate
// emits it: the canonical placeholder name, which is the only name the retarget
// is allowed to recognise.
func l3Opts() option.Options {
return option.Options{
Inbounds: []option.Inbound{
{
Type: C.TypeTun,
Tag: "l3-in",
Options: &option.TunInboundOptions{InterfaceName: netplane.L3DeviceBase, MTU: 65535},
},
},
}
}
func tunName(t *testing.T, opts option.Options, idx int) string {
t.Helper()
to, ok := opts.Inbounds[idx].Options.(*option.TunInboundOptions)
if !ok {
t.Fatalf("inbound %d options are %T, want *option.TunInboundOptions", idx, opts.Inbounds[idx].Options)
}
return to.InterfaceName
}
// TestL3SlotNeverCollidesWithTheRunningGeneration is the A1 invariant in one
// assertion: whatever the running generation holds, the next one is given
// something else.
//
// This is the whole fix. With a single fixed name, an apply built a box that had
// to open the device the running box still held (or that the kernel had not
// finished unregistering), TUNSETIFF answered EBUSY, and — because the recovery
// path rebuilds the PREVIOUS config, which named the same device — the rescue
// failed for the very reason it was needed. Measured on the production router:
// "start instance failed and could not restore previous config; engine stopped",
// then `plane: hold`, i.e. the whole LAN offline until a manual restart.
func TestL3SlotNeverCollidesWithTheRunningGeneration(t *testing.T) {
for _, current := range append([]string{"", netplane.L3DeviceBase}, netplane.L3Slots[:]...) {
got, dev := l3RetargetForNext(l3Opts(), current)
if dev == "" {
t.Fatalf("current=%q: no device was chosen for a config that HAS an L3 inbound — the engine would then start it under the placeholder name, which is the single-name collision this design removes", current)
}
if dev == current {
t.Errorf("current=%q: the next generation was handed the SAME device %q the running one holds. Start would answer `TUNSETIFF: device or resource busy`, and on the router that is a LAN-wide outage, not a failed apply", current, dev)
}
if name := tunName(t, got, 0); name != dev {
t.Errorf("current=%q: chose %q but the returned options still say %q", current, dev, name)
}
if !netplane.L3IsManagedName(dev) || dev == netplane.L3DeviceBase {
t.Errorf("current=%q: chose %q, which is not one of the slots %v — the fw4 zone and our nft accepts match the slot prefix, and a name outside it is a device the firewall drops", current, dev, netplane.L3Slots)
}
}
}
// TestL3RetargetLeavesTheCanonicalOptionsAlone guards the OTHER half of the
// design, and it is the half that is easy to lose in a "simplification": the
// substitution must not reach the options the caller keeps.
//
// applyLocked hashes opts and stores it as e.current; every later reconcile
// compares a freshly generated (canonical) config against that stored one. If
// the retarget wrote through the shared *TunInboundOptions pointer, the stored
// config would carry a per-generation device name, the next identical reconcile
// would hash differently, and the engine would rebuild itself — dropping every
// connection through the tunnel — once a minute, forever.
func TestL3RetargetLeavesTheCanonicalOptionsAlone(t *testing.T) {
canonical := l3Opts()
origOptions := canonical.Inbounds[0].Options
got, dev := l3RetargetForNext(canonical, "")
if dev == netplane.L3DeviceBase || dev == "" {
t.Fatalf("retarget chose %q — nothing below would prove anything", dev)
}
if name := tunName(t, canonical, 0); name != netplane.L3DeviceBase {
t.Errorf("the caller's options were MUTATED to %q. That name then becomes e.current, every later reconcile hashes differently against it, and the engine rebuilds on every one-minute no-op reconcile.", name)
}
if canonical.Inbounds[0].Options == got.Inbounds[0].Options {
t.Error("the returned options share the TunInboundOptions pointer with the caller's — they are one struct behind two `any`s, so the next retarget writes through both")
}
if origOptions != canonical.Inbounds[0].Options {
t.Error("the caller's inbound now points at a different options struct")
}
// The backing array too: copying only the slice header still aliases it.
if len(canonical.Inbounds) > 0 && len(got.Inbounds) > 0 &&
&canonical.Inbounds[0] == &got.Inbounds[0] {
t.Error("the Inbounds slices share a backing array — writing the retargeted inbound reached the caller's slice")
}
}
// TestL3RetargetIgnoresInboundsThatAreNotOurs: the retarget matches on the TUN
// type AND on a name this project owns. A tun inbound that is not the L3
// ingress must be left exactly as configured — renaming a device out from under
// its owner is a bigger failure than not renaming ours.
func TestL3RetargetIgnoresInboundsThatAreNotOurs(t *testing.T) {
cases := []struct {
name string
in option.Inbound
}{
{"foreign tun device", option.Inbound{Type: C.TypeTun, Tag: "vpn", Options: &option.TunInboundOptions{InterfaceName: "tun0"}}},
{"not a tun at all", option.Inbound{Type: C.TypeTProxy, Tag: "lan", Options: &option.TProxyInboundOptions{}}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
opts := option.Options{Inbounds: []option.Inbound{tc.in}}
got, dev := l3RetargetForNext(opts, "")
if dev != "" {
t.Errorf("chose device %q for a config with no L3 ingress", dev)
}
if to, ok := got.Inbounds[0].Options.(*option.TunInboundOptions); ok && to.InterfaceName != "tun0" {
t.Errorf("renamed a foreign tun device to %q", to.InterfaceName)
}
})
}
}
// TestL3RetargetNoInboundIsNotADevice: with l3_tunnel off (the default) there is
// no TUN inbound, so there is no slot to claim and nothing to publish to the
// netplane. Reporting a device here would make addL3Routing point the L3 table
// at an interface nothing ever creates.
func TestL3RetargetNoInboundIsNotADevice(t *testing.T) {
opts := option.Options{Inbounds: []option.Inbound{
{Type: C.TypeTProxy, Tag: "lan", Options: &option.TProxyInboundOptions{}},
}}
if _, dev := l3RetargetForNext(opts, netplane.L3Slots[0]); dev != "" {
t.Errorf("device = %q for a config without an L3 ingress, want \"\"", dev)
}
}
// TestEngineSlotHandoffSurvivesAFailedSwap walks the exact sequence
// closeOldThenStart performs, through the ENGINE's own state, and asserts the
// property the production outage violated: no step is ever handed the device
// the step before it was using.
//
// The sequence, with the router log it reproduces:
//
// gen1 running on some slot (19:10 healthy)
// an apply builds gen2 while gen1 still holds its ← must differ, or TUNSETIFF
// gen1 is closed, gen2's Start fails (19:14:02)
// the RESTORE rebuilds the previous config ← must not be handed gen2's
//
// The last step is the one that turned a failed apply into an outage: it used to
// rebuild a config naming the one fixed device, so it collided with the teardown
// still in flight and the engine stopped with the kill-switch closed.
func TestEngineSlotHandoffSurvivesAFailedSwap(t *testing.T) {
e := &Engine{}
_, gen1 := l3RetargetForNext(l3Opts(), e.l3Device)
e.l3Device = gen1 // gen1 adopted and running
_, gen2 := l3RetargetForNext(l3Opts(), e.l3Device)
if gen2 == gen1 {
t.Fatalf("gen2 was handed gen1's device %q while gen1 is still running — this is the TUNSETIFF EBUSY at 19:10:33", gen1)
}
// closeOldThenStart: gen1 retired, nothing running. The engine clears the
// slot here precisely so the two rebuilds below get a real choice instead of
// inheriting a single forced candidate.
e.l3Device = ""
_, restore := l3RetargetForNext(l3Opts(), e.l3Device)
if restore == "" {
t.Fatal("the restore path was given no device at all")
}
// The restore must not be forced onto the device the failed generation was
// using. Without the clear above, gen1's name would still be recorded, the
// only remaining candidate would be gen2's — the one that just failed — and
// the rescue would again depend on the resource it is rescuing from.
if e.l3Device != "" {
t.Fatalf("the engine still claims device %q with nothing running", e.l3Device)
}
}
+1 -1
View File
@@ -27,7 +27,7 @@ func outboundByTag(opts option.Options, tag string) *option.Outbound {
func byedpiModel(port int) *model.Model {
return &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: port}},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "egress:bd"},
@@ -117,7 +117,7 @@ func TestChainDeadExitMarkedByObservatory(t *testing.T) {
eng.StopObservatory()
_ = eng.Close()
})
if _, err := eng.Apply(opts); err != nil {
if _, err := eng.Apply(withoutL3Ingress(t, opts)); err != nil {
t.Fatalf("engine.Apply (box.New + start): %v", err)
}
@@ -174,7 +174,7 @@ func TestChainDeadExitDialFailsClosed(t *testing.T) {
}
eng := engine.New()
t.Cleanup(func() { _ = eng.Close() })
if _, err := eng.Apply(opts); err != nil {
if _, err := eng.Apply(withoutL3Ingress(t, opts)); err != nil {
t.Fatalf("engine.Apply: %v", err)
}
inst := eng.Instance()
+3 -3
View File
@@ -47,7 +47,7 @@ func generalRouteOutbound(rt *option.RouteOptions) (string, bool) {
func threeNodeChainModel() *model.Model {
return &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Nodes: []model.Node{
{Name: "a", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#a"},
{Name: "b", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.2:8388#b"},
@@ -122,7 +122,7 @@ func TestChainMultiHopDetourWiring(t *testing.T) {
// (no wrapper, no detour outbound).
func TestChainSingleHopResolvesToHop(t *testing.T) {
m := &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Nodes: []model.Node{
{Name: "a", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#a"},
},
@@ -210,7 +210,7 @@ func TestChainEmptyHopsWarnsBlocks(t *testing.T) {
// through the previous hop; the next hop detours into the group wrapper.
func TestChainGroupHop(t *testing.T) {
m := &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Nodes: []model.Node{
{Name: "a", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#a"},
{Name: "b", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.2:8388#b"},
+12 -2
View File
@@ -275,7 +275,14 @@ func TestDNSFilterLiveAnswers(t *testing.T) {
var mgr deprecated.Manager = &recordingDeprecated{}
ctx = service.ContextWith(ctx, mgr)
ctx = registry.Context(ctx)
b, err := box.New(box.Options{Context: ctx, Options: opts})
// withoutL3Ingress for the reasons given at its definition, plus one that is
// sharper here than anywhere else in the suite: this test builds the box
// DIRECTLY, so it never reaches engine.newBox and never gets a slot — it
// would open the device under the PLACEHOLDER name, which is the one name
// every generation wants and therefore the one name that must never exist.
// That is not a hypothesis: it is where the `shater-l3` device in
// TestIntegrationL3TunInboundStarts' failure came from.
b, err := box.New(box.Options{Context: ctx, Options: withoutL3Ingress(t, opts)})
if err != nil {
t.Fatalf("box.New: %v", err)
}
@@ -393,7 +400,10 @@ func applyWithDeprecations(t *testing.T, m *model.Model) (option.Options, []depr
var mgr deprecated.Manager = rec
ctx = service.ContextWith(ctx, mgr)
ctx = registry.Context(ctx)
b, err := box.New(box.Options{Context: ctx, Options: opts})
// Direct box.New, so the L3 ingress must come out first — see the note at
// the identical call in TestDNSFilterLiveAnswers, and withoutL3Ingress for
// the whole argument.
b, err := box.New(box.Options{Context: ctx, Options: withoutL3Ingress(t, opts)})
if err != nil {
t.Fatalf("box.New failed: %v\nwarnings: %v", err, warns)
}
+1 -1
View File
@@ -35,7 +35,7 @@ func routeActionFor(opts option.Options, tag string) *option.RouteActionOptions
func dpiModel(dpi string) *model.Model {
return &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Egresses: []model.Egress{{Name: "frag", Type: "direct", DPI: dpi}},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "egress:frag"},
+148
View File
@@ -0,0 +1,148 @@
// Defect 1: an `interface` egress with no interface used to be bound to the LAN
// bridge.
//
// The chain, because none of it is visible in the generated JSON: this generator
// resolved the bind device with netplane.IfaceDevice(eg.Interface), and
// IfaceDevice("") falls back to "br-lan" (the right default for an INBOUND with
// no network). netplane.EgressDevice — the resolution the DATA plane uses —
// returns "" for the same egress on purpose, and calls br-lan "catastrophic
// here", so addEgressRouting installed no `ip rule` and no routing table for that
// egress's mark, and the prerouting marking and the forward-chain accept skipped
// it too.
//
// The result was an outbound with SO_BINDTODEVICE=br-lan and a routing mark
// nothing routed: every node, group and rule bound to that egress dialled public
// addresses out of the LAN bridge. Not a leak — the bind pins the socket to the
// LAN — but a total, silent black hole, with the panel showing a configured,
// applied egress and no findings at all.
package generate
import (
"strings"
"testing"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// deviceLessEgressModel is the config at issue: one interface egress with no
// interface, and a rule that sends a LAN subnet through it.
func deviceLessEgressModel(iface string) *model.Model {
return &model.Model{
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12345, TCP: true, UDP: true},
},
Egresses: []model.Egress{{Name: "hole", Type: "interface", Interface: iface}},
Rules: []model.Rule{
{Name: "via-hole", Enabled: true, Order: 10, Src: []string{"192.168.9.0/24"}, Target: "egress:hole"},
},
}
}
// bindInterfaces returns every BindInterface the generated outbounds carry,
// keyed by outbound tag. Only direct outbounds can carry one.
func bindInterfaces(opts option.Options) map[string]string {
out := map[string]string{}
for i := range opts.Outbounds {
if do, ok := opts.Outbounds[i].Options.(*option.DirectOutboundOptions); ok {
out[opts.Outbounds[i].Tag] = do.BindInterface
}
}
return out
}
// TestInterfaceEgressWithNoInterfaceEmitsNoOutbound is the defect proper.
func TestInterfaceEgressWithNoInterfaceEmitsNoOutbound(t *testing.T) {
opts, warns, err := GenerateWithWarnings(deviceLessEgressModel(""))
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
tag := netplane.EgressOutboundTag("hole")
binds := bindInterfaces(opts)
if dev, ok := binds[tag]; ok {
t.Errorf("outbound %q was emitted, bound to device %q. netplane.EgressDevice resolves this egress to "+
"NO device, so addEgressRouting installs neither its `ip rule` nor its routing table and the "+
"prerouting mark and forward accept skip it: this outbound is bound to a device the router does not "+
"route for, under a mark that leads nowhere. Every node, group and rule bound to this egress "+
"disappears into it. Emit nothing instead, so the binding fails closed and is visible.", tag, dev)
}
for otag, dev := range binds {
if dev == "br-lan" {
t.Errorf("outbound %q is bound to %q — the LAN BRIDGE. That is IfaceDevice's empty-name fallback "+
"leaking into an egress bind: the socket is pinned to the LAN and dials public addresses out of "+
"it. Nothing leaves the house, and nothing works, and nothing says why.", otag, dev)
}
}
// The skip must be LOUD. A silent skip only moves the black hole from the data
// plane into the panel.
if !hasEgressWarning(warns, "hole") {
t.Errorf("no warning names egress %q. The egress is configured, the panel lists it, rules are bound to "+
"it, and it carries nothing — the operator has to discover that by noticing their traffic stopped. "+
"Warnings were:\n %s", "hole", strings.Join(warns, "\n "))
}
}
// TestInterfaceEgressWithBlankInterfaceEmitsNoOutbound is the same defect through
// the whitespace door — the exact shape that had already produced one
// validator/data-plane divergence in this codebase (see EgressDevice's comment).
// IfaceDevice(" ") hands back " ", which is neither empty nor a device, so the
// old code bound the socket to a device name made of spaces.
func TestInterfaceEgressWithBlankInterfaceEmitsNoOutbound(t *testing.T) {
opts, warns, err := GenerateWithWarnings(deviceLessEgressModel(" "))
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
tag := netplane.EgressOutboundTag("hole")
if dev, ok := bindInterfaces(opts)[tag]; ok {
t.Errorf("outbound %q was emitted, bound to %q. netplane.EgressDevice trims and resolves this to no "+
"device, so the router routes nothing for this egress's mark.", tag, dev)
}
if !hasEgressWarning(warns, "hole") {
t.Errorf("blank interface skipped silently; warnings were:\n %s", strings.Join(warns, "\n "))
}
}
// TestInterfaceEgressWithADeviceStillEmitsItsOutbound is the CONTROL. Without it
// the two tests above are equally satisfied by a generator that emits no egress
// outbound ever — an instrument that cannot produce a positive proves nothing by
// producing a negative.
func TestInterfaceEgressWithADeviceStillEmitsItsOutbound(t *testing.T) {
m := deviceLessEgressModel("wan2")
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
tag := netplane.EgressOutboundTag("hole")
dev, ok := bindInterfaces(opts)[tag]
if !ok {
t.Fatalf("no outbound %q for an egress that DOES resolve to a device — the fail-closed skip has swallowed "+
"a healthy egress, and every rule bound to it is now blocked. Warnings: %v", tag, warns)
}
if want := netplane.EgressDevice(m.Egresses[0]); dev != want {
t.Errorf("BindInterface = %q, want %q — the bind device must be the one netplane.EgressDevice resolves, "+
"because that is the device addEgressRouting builds the mark's routing table around. Any other "+
"string binds the socket to one device while the kernel routes its mark out another.", dev, want)
}
if hasEgressWarning(warns, "hole") {
t.Errorf("a healthy interface egress must not be reported as device-less; warnings were:\n %s",
strings.Join(warns, "\n "))
}
}
// hasEgressWarning reports whether any warning is about this egress AND about it
// having no device — matching on the entity prefix apply's normaliser parses
// (`egress "name": ...`) plus the substance, so an unrelated egress warning (a
// stray port, an unknown type) cannot pass for this one.
func hasEgressWarning(warns []string, name string) bool {
for _, w := range warns {
if strings.HasPrefix(w, `egress "`+name+`": `) && strings.Contains(w, "no device") {
return true
}
}
return false
}
+2 -2
View File
@@ -13,7 +13,7 @@ import (
// the egress tag (multi-WAN). No copies: the binding lands on the node itself.
func TestNodeEgressBindsOutbound(t *testing.T) {
m := &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Egresses: []model.Egress{{Name: "wan2", Type: "direct"}},
Nodes: []model.Node{
{Name: "a", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#a", Egress: "wan2"},
@@ -80,7 +80,7 @@ func TestNodeEgressMissingIsFailClosed(t *testing.T) {
// outbound (instead of dialing directly from the router).
func TestChainEgressEntryHop(t *testing.T) {
m := &model.Model{
Globals: nonDNSGlobals(),
Globals: plainGlobals(),
Egresses: []model.Egress{{Name: "wan2", Type: "direct"}},
Nodes: []model.Node{
{Name: "x", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#x"},

Some files were not shown because too many files have changed in this diff Show More