Compare commits

..
Author SHA1 Message Date
omarandClaude Opus 5 1267d20fb8 docs: drop the L3 handoff note — it is merged, and it said to
test / go + panel tests (push) Successful in 8m33s
release / test gate (push) Successful in 8m8s
release / apk aarch64_cortex-a53 (push) Successful in 6m33s
release / apk x86_64 (push) Successful in 3m45s
release / release apk (push) Successful in 8s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 20:00:46 +03:00
omarandClaude Opus 5 35f697ed08 docs(openwrt): say why mtu_fix is inert instead of claiming an MTU we no longer set
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:52:39 +03:00
omarandClaude Opus 5 d0fb6befb1 fix(l3): the l3-in MTU is not a tunnel budget — 1420 was a forgery generator
shater-l3 was created at 1420, the WireGuard payload budget, copied one
layer too far out. It bought nothing: what actually goes into the tunnel
is sized by sing-tun's forwardToPort against Port.PortMTU(), which
already fragments to the outbound MTU without DF and answers a
well-formed `fragmentation needed` quoting it with DF. All 1420 did was
make the KERNEL split every packet above 1392 bytes of payload on its
way into the device -- and a fragment is the one thing sing-tun will not
judge. Dispatch returns on parsed.fragment before calling JudgeFlow, the
fragments reach the gVisor stack, it reassembles them, and the ICMP
forwarder's installFlow demands an unspecified port address that a
WireGuard endpoint never has. So it declined and answered the echo
itself. `ping -s 1392` honest, `ping -s 1393` a lie, and only for the
outbounds the feature exists for.

65535 rather than merely "large": no IP datagram can exceed it, so the
kernel cannot fragment at this device for any packet ever. Anything
smaller leaves a band open and re-opens the class. It is also sing-box's
own default TUN MTU on Linux.

Memory was measured, not argued. Three paired runs of the integration
test under -test.memprofilerate=1 allocate 5.41/5.48/5.47 MB at 65535
against 5.76/5.46/5.70 MB at 1420, and a -diff_base profile puts every
difference in netlink interface enumeration. Nothing in the read path
scales with the MTU: gVisor reads through fdbased.BufConfig, which
sing-tun pins to one 65535-byte view regardless. I predicted a ~1.8 MB
saving from GSO switching off above 49152 and was wrong -- protocol/tun
turns GSO back on at StartStateStart whenever a FlowOutbound exists, so
the GRO scaffolding is there at both values. The corrected reasoning is
in the constant's comment so the next reader does not redo the mistake.

The integration test now reads the MTU back off the real kernel device,
which is the assertion the value exists for: a kernel that clamped it
would restore the forgery without changing a generated byte.

D25's KNOWN HOLE block is replaced with what is genuinely left. Chiefly:
a big non-DF ping does not start WORKING, it starts failing HONESTLY --
classifyReturn declines fragments on the way back too, so the packet
really leaves, the far host really answers, and the reply is not NAT'd
home. And a client that fragments on the wire itself is still uncovered;
that is the nft carve-out's job, with a warning that conntrack defrag
may reassemble in prerouting and leave such a rule unable to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:50:08 +03:00
omarandClaude Opus 5 81c96019b5 fix(panel): let a routing rule say ICMP, instead of calling one broken
The Proto picker was a closed list of the two transports and the ten
sniffed L7 labels, and anything else drew "<value> — never matches".
The engine now routes ICMP by rule (Rule.Proto accepts icmp, icmpv4,
icmpv6), so a working ping rule was rendered as a dead one and could not
be created here at all — the operator had to hand-edit /etc/config/shater
and then watch the panel call the result broken.

Adds a third group, "Layer 3". All three spellings are offered: they are
not synonyms — icmpv4/icmpv6 pin the rule's ip_version — so hiding the
narrowing would both strand a capability outside the UI and silently
widen such a rule the first time someone edited it here.

The doc comment no longer claims the list IS generate/route.go's
sniffedProtocols; only the middle group is. ICMP goes to the emitted
rule's `network`, never to `protocol`, which is the whole reason it never
matched as a sniffed label.

An unknown value is still kept and offered as written, but the
never-matches flag is now judged on the lower-cased value, the way the
engine judges it — a hand-written `ICMP` is a live rule, not an inert one.

Verified: npm run build clean (tsc --noEmit + vite build); an icmp rule
added through the panel renders as a plain "PROTO icmp" chip; no
horizontal overflow at 360px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:26:07 +03:00
omarandClaude Opus 5 baed8ff8f2 fix(model): fwmark_base 0x7f routes the engine's own traffic into its own TUN
The panel offers fwmark_base and table_base as free hex fields under
"Advanced" and nothing has ever checked them. What makes that more than a
footgun is that the derived values are invisible from the number typed: the
L3 mark is base+0x80, so 0x7f lands it exactly on 0xff — the loop-guard mark
the engine stamps on its OWN traffic — and `ip rule fwmark 0xff lookup 8200`
then captures everything the engine sends and routes it into the engine's
TUN. The router loses the internet the moment l3_tunnel is switched on, for
a reason nothing on screen connects to a collapsed section. fwmark_base 0xff
had produced the same failure since long before the L3 offset existed.

table_base is worse and got the same treatment: its derived values can land
on the kernel's own table ids, and teardown does `ip route flush table <n>`.
It is count-sensitive (egress #i uses base+0x10+i), so the check takes the
egresses rather than living in ValidateGlobals.

Written as "derive every value this layout produces, then look for
duplicates and reserved ids" rather than as a blacklist, so a future offset
is covered by construction. The layout constants are duplicated from
netplane (the import only runs one way) and pinned by netplane's
TestMarkLayoutConstantsLockstep.

Warn-only, like every check in this file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 f80fb4dd1b fix(netplane): give every mark-driven table a floor, and check the L3 pair
Two halves of the same omission.

1. A fwmark lookup that finds an empty table does not fail — it falls
   through to main. Every mark-driven table now gets an `unreachable
   default` at the maximum metric: it loses to any real default route while
   one exists, it has no device so the kernel never garbage-collects it, and
   it turns "lookup failed, try main" into "lookup succeeded: unreachable".
   The fallthrough stops depending on somebody reading a warning at the
   moment an interface goes down. Deliberately not gated on the kill-switch:
   that switch decides whether traffic may escape the tunnel, while an egress
   binding is a statement about WHICH UPLINK, and silently substituting a
   different one is not what "fail open" was meant to permit.

   RoutingPresent's "does this table have a default route" test is tightened
   in the same breath, or the floor would answer it and turn the safety net
   into a blindfold.

2. RoutingPresent had never heard of addL3Routing. This is the same defect
   its own comment describes as already caught twice ("a presence check must
   cover everything its Apply counterpart installs"), committed a third time
   — and its trigger needs no interface to go down: editing a node URI
   restarts the engine, the kernel destroys shater-l3 and takes `default dev
   shater-l3 table 8200` with it, the rendered nft text is unchanged, so the
   fast-path skipped ApplyRouting forever and LAN ping stayed dead until
   someone restarted the daemon.

TestRoutingPresentSeesL3Table, TestEgressTableGetsFailClosedFloor and
TestEveryStampedMarkIsRoutedAndVerified all fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 b71b793681 fix(netplane): a mark says where a packet was sent, not where it went
The forward chain let untunnelable-egress traffic past the kill-switch on
the strength of its fwmark alone. `ip rule fwmark X lookup N` does not
deliver the packet to table N, it delivers the LOOKUP there — and a lookup
that finds nothing falls through to main. So when the egress interface goes
down and the kernel garbage-collects its default route, every non-TCP/UDP
packet from the LAN is still stamped, still accepted here (above the
fail-closed drop), and leaves out the plain WAN with the router's real
address. Nothing we render changes, so no apply runs and nothing notices.

Ordinary egress traffic never had this hole: the engine binds those sockets
to the device, and a dead device fails the socket. The untunnelable-egress
path is made of nothing but a mark, so the accept now carries the second
opinion instead — `meta mark X oifname "dev"`, strictly narrower than either
half, true only when the routing did what the mark asked. The comment being
replaced argued correctly that oifname ALONE would be too loose, then drew
from that the conclusion that oifname should be dropped rather than added.

Same conjunction in the holding plane, where it is theory (that plane stamps
nothing) but where a bare mark accept has no business sitting.

Also folds the egress device resolution into one EgressDevice(), because the
binding and model.ValidateUntunnelableEgress had already drifted: the
validator trimmed the interface name and the binding did not, so `option
interface '   '` gave a panel saying "the option is ignored" over a data
plane that was marking packets for a table nobody built.

TestUntunnelableEgressAcceptIsBoundToItsDevice and
TestUntunnelableEgressResolutionLockstep fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:32 +03:00
omarandClaude Opus 5 61c87ad1d9 fix(l3): guard the ICMP honest-drop at PreMatch, not inside the walk
The drop that keeps a ping from reading as tunnelled lived in
preMatchFlow, overriding the pre-declared continueResult. That covered
every exit of THAT function and none of the walk above it: the
prepareMatchMetadata error return (which arrived later, with the shared
metadata refactor), the sniff bail-outs, and the default: arm of the
rule-action switch all returned PreMatchContinue on their own.
adapter.JudgeFlow maps Continue to tun.ActionAccept, and sing-tun answers
Accept by rewriting Echo into EchoReply itself -- the exact forgery this
delta exists to remove. Narrow paths, but paths.

PreMatch is now a funnel over the renamed preMatch walk, so the guard
sits on the single return value and cannot be outgrown by a new exit.
PreMatchBypass joins the drop: sing-tun implements ActionBypass on the
nfqueue plane only, so on the TUN path it lands in the same default: arm
as Accept and forges too.

Every ICMP case has an explicit TCP/UDP twin; the JudgeFlow mapping
table is pinned outright, including the one fix that must NOT be made
there -- refusing ActionFlow for a port whose address is not unspecified
would drop every ping through WireGuard/AWG, because the forward
dispatcher and the ICMP forwarder share that function with identical
arguments and only the latter needs an unspecified address.

That leaves a real hole open, now named in D25 rather than papered over:
a FRAGMENTED echo to a WireGuard/AWG outbound is still answered by the
router. The dispatcher returns before asking for a verdict at all when
the packet is a fragment, and the reassembled packet reaches the ICMP
forwarder, whose installFlow demands the unspecified address a WireGuard
endpoint never has. The two fixes that would close it both live outside
pre-match and are written down; the Consequence paragraph is scoped
until one lands.

The stack comment in generate/inbound.go repeated the "only gvisor
really forwards ICMP" argument that D25 itself retracts -- both stacks
run the same ForwardDispatcher first. Brought in line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:21:29 +03:00
omar 4dee508e12 fix(route): let a rule say "icmp", and say when saying it is a lie
`icmp` fell through ruleMatchers' proto switch into RawDefaultRule.Protocol —
the SNIFFED-L7 field, compared against what the sniffers labelled a connection.
Nothing ever labels a flow "icmp" (PreMatch skips the sniff action for an ICMP
flow outright), so the rule was structurally valid and permanently dead. That
made the whole L3 ingress unusable on a real config: with no way to write "ICMP
goes here", every ping fell to the catch-all, which resolves to the chain's last
hop — a group of VLESS nodes that cannot carry layer 3 at all.

icmp is a NETWORK. NetworkItem.Match is a map lookup over metadata.Network, and
adapter.JudgeFlow sets that to N.NetworkICMP for BOTH ICMPv4 and ICMPv6 (one
case covers both protocol numbers), so there is exactly one network value and it
covers both families. `icmpv4`/`icmpv6` narrow that same network with an
ip_version item instead of inventing a second one: metadata.IPVersion comes from
the destination address, and an ICMPv6 packet always has an IPv6 destination —
no false positives, no false negatives.

An ICMP rule that cannot fire is not a dead setting: ICMP has no fall-through,
so route.preMatchFlow DROPS it. Four ways to get that silently are now reported:
l3_tunnel off (nothing enters the engine at all), icmpv6 with ipv6 off (neither
the nft mark nor the TUN address exists), a port matcher next to it (JudgeFlow
zeroes both ports), and a target that cannot carry layer 3 — decidable from the
model, because the capability is fixed by the outbound TYPE: only wireguard/AWG
endpoints and the direct outbound behind direct/interface egresses declare
N.NetworkICMP. A mixed group gets its own text (the answer follows group.Now()),
`block` gets none (dropping the ping IS the policy), and an unresolved target
gets none either (ruleKillFallback already said the louder thing).

Wording stays clear of shater/apply's criticalMarkers on purpose: a failed ping
is fail-CLOSED, and a cosmetic alarm is how the real one stops being read.
2026-07-26 19:18:06 +03:00
omarandClaude Opus 5 76da5134ef test(gate): the two tests that need a kernel may not skip in silence
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.

Two halves.

Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.

Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.

The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 18:50:46 +03:00
omar 4c630c9a13 docs: handoff note for the L3 branch
Transient, to be deleted when omp/work merges. Everything meant to outlive the
merge is already in D25/D26 and the lx changelog; this file is the part that is
only useful while the branch is still a branch — the verification commands, the
testbed recipe, what was proven on hardware and what was not, and the six files
that will conflict on rebase.
2026-07-26 18:31:59 +03:00
omar d8dbefcd07 docs: record the AWG site-to-site path as declined, not impossible
D26's "no port-like selector" line disposes of NAT-based forwarding and nothing
else, and read alone it says "impossible" — which is false and would be
re-derived at the cost of another research pass. The endpoint is protocol-blind
in both directions, so ESP could ride it untouched with the client's own source
address and no NAT whatsoever. That was declined for two reasons worth naming:
lx-owned code in the forward hot path, and a server-side AllowedIPs prerequisite
that turns a router option into a deployment contract.
2026-07-26 18:31:59 +03:00
omar 974208fc05 docs: record the kernel egress, and retract the reason D25 gave for the ceiling
D26 writes down where the engine's boundary actually is, because the intuitive
answer is wrong and someone will look for it again: the WG/AWG forward path
never consults gVisor in either direction, so the limit is sing-tun's
ForwardDispatcher — its parser and its port-shaped NAT — and the kernel egress
was chosen because it clears that limit without a line of new hot-path code, not
because userspace "cannot". Tailscale documents the same boundary for their
userspace mode and is quoted as corroboration, with the caveat that ours sits at
the dispatcher rather than the stack.

D25 said two things that do not survive checking, and both are corrected in
place rather than left for the next reader to trip over. It blamed the netstack
for the ICMP-echo ceiling; that was the dispatcher. And it called `stack: gvisor`
mandatory because the system stack fakes ping — the system stack runs the very
same dispatcher first and only forges an echo for packets the dispatcher
declined, so gvisor is a deliberate choice (already linked via with_wireguard,
and the combination the integration test exercises), not a necessity.

The operator note says what the option buys and refuses to call an egress a
tunnel on its own say-so: with a WireGuard device it is one, with a second WAN
the destination sees that uplink's address. It also says what the option does
not fix — multicast IPTV stays broken — and that IPsec through NAT-T is ordinary
UDP that never needed any of this.
2026-07-26 18:31:59 +03:00
omar 2eb71e8244 feat(netplane,model): hand the protocols the engine will not dispatch to the kernel
ESP, AH, GRE, IGMP and SCTP cannot enter the engine, and the reason is not the
one that looks obvious. A WireGuard or AmneziaWG endpoint forwards straight past
its gVisor stack — WritePackets reads the IP version and the destination address
and hands the raw bytes to the device, and the return path offers every
decrypted packet back before the stack sees it. WireGuard would carry ESP today
if anything handed it one. What refuses is sing-tun's ForwardDispatcher: its
parser recognises TCP, UDP and ICMP echo, and its NAT wants a port-shaped
selector that ESP, AH and GRE do not have. The retracted rationale is corrected
where it was written down, not quietly dropped.

So these protocols go to the kernel instead. untunnelable_egress names an
interface or tunnel egress; prerouting stamps that egress's OWN mark on
everything that is not TCP or UDP, and addEgressRouting has already bound that
mark to a table whose default route leaves via the device. Every protocol works
because nothing in the path has to understand any of them. No new mark, no new
table, no new code in the hot path.

Whether that is a tunnel depends on the device, and nothing here claims
otherwise: a WireGuard interface is one, a second WAN is a different uplink
whose real address the far end sees.

The wide `!= { tcp, udp }` filter is safe here and stays banned for the L3
ingress, for the same reason stated in both places: there the receiver is a
dispatcher that knows four protocols, here it is the kernel. ICMP is claimed by
the L3 ingress first when both are on. The local plane keeps its exclusions —
router-addressed traffic, private destinations, ICMPv6 ND/RA — and with IPv6 off
the marking is scoped to v4, because addEgressRouting installs no v6 rule then
and a marked v6 packet would fall into the main table.

An interface egress with an empty `interface` no longer resolves: IfaceDevice
defaults to br-lan, so it passed the binding while addEgressRouting skipped it —
mark set, no rule, straight past a closed kill switch and out the default WAN.
2026-07-26 18:31:59 +03:00
omar 668cccbf24 test(generate): the L3 device name is a singleton, so wait for the kernel to take it back
Both gated tests stand an engine up on shater-l3. Run together, the second met
`TUNSETIFF: device or resource busy` and failed for a reason that had nothing to
do with what it asserts — the first had closed its box and yielded while
unregister_netdevice was still catching up. Each passed alone, which is the
shape of a fixture bug that gets rediscovered rather than fixed.

The poll that already guarded the first test is now a shared helper both call.
It stays a poll rather than a sleep for the reason it always was: the removal is
usually immediate and a fixed wait would be either flaky or slow.
2026-07-26 18:31:59 +03:00
omar 4ea4585402 test(generate): pin that ping through an interface egress is real, and byedpi's is not
An interface egress is a direct outbound carrying BindInterface and a routing
mark, and direct builds its ICMP port from the very same dialer control — so
ping routed at that egress leaves through that device, marked, like every other
packet bound to it. Nothing said so. Both halves of that sentence are one
`common.Cast[*dialer.DefaultDialer]` away from being false: if the dialer ever
stops being a DefaultDialer, icmpPort is nil, PreMatchFlow declines, and ping
through the egress degrades to a drop without a single generated byte changing.
The gated test asserts the live outbound, not the config, because that is where
the cast happens.

The failure the codegen half guards is worse than a broken ping: losing
BindInterface or the mark does not stop the echo, it sends it out the main table
over the plain WAN with the real address, which is the one thing an egress
exists to prevent.

byedpi is a SOCKS outbound and cannot be a tun.Port, so ICMP aimed at it is
dropped. That is the honest end of l3-honest-drop and it is pinned too, because
the alternative the TUN stack offers is a forged reply.
2026-07-26 18:31:59 +03:00
omar dc6d102473 docs: put a number on the second netstack, and say what it does not bound
Measured on a throwaway harness in a container: peak RSS of a process that
brought the engine up went from ~26 MB to ~28 MB with l3_tunnel on, three
paired runs. It is x86_64, idle, with an empty ICMP NAT table, so it stays
listed as unverified for the router — an indicative figure is more useful than
silence only if it says loudly what it is not.
2026-07-26 18:31:58 +03:00
omar 683afc0a47 docs: record how ping got through the tunnel, and where it stops
D25 writes down the reasoning that is expensive to reconstruct: why a TUN rather
than TPROXY, why the interface is its own with auto_route off, why gvisor is
mandatory rather than preferred, and why the ceiling is ICMP echo — a boundary
in sing-tun's flow parser and gVisor's protocol set, not an unfinished edge of
ours. It also records what carries layer 3 and what does not, that masque could
and does not, and the two things still unproven: the live-router path end to
end, and what a second gVisor NIC costs in memory on the hardware.

D17 gains one line: its claim that TPROXY cannot carry ICMP is still true, and
is no longer the end of the story.
2026-07-26 18:31:58 +03:00
omar 2c3e20512e feat(openwrt): let fw4 know the L3 tunnel device before it exists
Both nft tables run and a drop in either one wins, so our forward accept for
shater-l3 decides nothing on its own: fw4 sees a device in no zone and drops the
forward, and the feature fails with exactly the symptom it was built to fix —
ping does not work, and nothing says why.

The zone names the device directly rather than a network. fw4 resolves a zone's
networks through netifd, and a proto-none interface for a device the daemon
creates is never up and contributes nothing, so list network would compile to an
empty device set. list device compiles to a plain iifname/oifname match that is
valid before the TUN exists and starts matching the moment shaterd creates it,
with no firewall reload at enable time.

It is seeded unconditionally, not gated on l3_tunnel: uci-defaults run once, and
a zone naming an absent device is inert. Gating it would mean the option could
be switched on and never take effect. The sections are named so a re-run is a
no-op instead of a second zone, and kmod-tun joins DEPENDS because /dev/net/tun
is not on a stock image.
2026-07-26 18:31:58 +03:00
omar 51b2f04672 feat(netplane,generate): carry LAN ping through the tunnel, on a TUN of its own
Kernel TPROXY needs a socket to hand a packet to, so it moves TCP and UDP and
nothing else. Everything else reached the forward chain and met the untunnelable
policy, whose best answer was "let it out with your real address" and whose
default was "drop it" — so on a stock install ping simply did not work, and the
setting that fixed it did so by leaking.

The engine has been able to do better for a while: sing-tun's ForwardDispatcher
does real ICMP forwarding with NAT on the echo id, and a WireGuard or AmneziaWG
endpoint is a tun.Port that carries the packet for real. What was missing was a
way in, because nothing on the router could hand it an IP packet.

l3_tunnel (opt-in, off by default) adds one: the generator emits an "l3-in" TUN
inbound and prerouting fwmarks LAN ICMP into it. The interface is its own and
auto_route is off, so the main routing table is never touched and the fwmark
plus addL3Routing's ip rule are the only entrance — the TPROXY plane is byte for
byte what it was. gvisor is not a preference: the system stack forges echo
replies locally, which is the very thing this is meant to end.

Only icmp and ipv6-icmp are ever marked, and only after the local plane is out
of the way — the router itself, private destinations, and ICMPv6 ND/RA, which
mean nothing off-link and take v6 down if one neighbour probe is tunnelled.
ESP, AH, GRE, IGMP and SCTP are deliberately left alone: sing-tun's parser and
gVisor's stack know no such protocol, so marking them would black-hole the
traffic while looking like a feature. They stay with the untunnelable policy,
which also keeps its say over what happens if the ip rule fails to install.

Ping and Windows tracert now cross the tunnel; IPv6 traceroute shows only the
destination, because the return path recognises TimeExceeded for v4 alone.
2026-07-26 18:31:58 +03:00
omar f190c8251e feat(lx): stop answering ping on behalf of a tunnel that never saw it
PreMatchContinue is not "fall back to the ordinary route" the way it is for TCP
and UDP. An ICMP flow has no ordinary route: the TUN stack takes the packet back
and answers the echo itself, swapping the addresses and writing a reply
(sing-tun stack_gvisor_icmp.go). So a ping routed to any outbound that cannot
carry layer 3 — every proxy protocol; only adapter.FlowOutbound can — came back
successful, and the operator read a working tunnel off a packet that was never
sent.

That is worse than the packet loss it replaced. Loss is a fault the operator can
see and chase; a forged reply is a fault that reports itself as health, and it
reports it on the one tool anyone reaches for first.

preMatchFlow now overrides continueResult once, at the top, for
N.NetworkICMP. One hunk covers every exit that used to fall through — no such
outbound, a group whose selection is gone, an outbound whose Network() omits
icmp, an outbound that is not a FlowOutbound — and keeps the diff to three lines
against a function upstream will keep editing. JudgeFlow carries the same
verdict in its !isPort branch, because FlowOutbound and tun.Port are separate
interfaces and drift between them must not reopen the forgery.

TCP and UDP are untouched, and the test pins that as hard as it pins the drop.
2026-07-26 18:31:58 +03:00
omarandClaude Opus 5 1945404eaa fix(armor): a reboot is not someone switching the product off
test / go + panel tests (push) Successful in 5m24s
release / test gate (push) Successful in 5m24s
release / apk aarch64_cortex-a53 (push) Successful in 3m9s
release / apk x86_64 (push) Successful in 3m9s
release / release apk (push) Successful in 8s
The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:

    case $action in restart|reload) keep;; *) DISARM;; esac

`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:

    13:28  /etc/shater/boot.nft present
    ----   reboot
    18s    at_S22: NO_TABLE  armor_file=NO_FILE

It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.

Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.

Also closed, found while proving the above:

  * Every restart left the LAN in the clear for 80-90ms. The exit path was
    `Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
    transactions with no `inet shater` between them, leaving fw4's
    `lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
    & Apply. TeardownExiting arms first under the apply lock and skips the
    delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
    REPLACES the table, so the kernel never observes its absence.
    35k-sample instrument: 7 and 6 no-table hits before, 0 across three
    runs after.

  * SaveBootArmor fsynced the payload but not the directory, so a power cut
    could lose the rename that publishes it -- a boot with no armor and no
    error anywhere.

`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.

Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 17:33:53 +03:00
omarandClaude Opus 5 6476722372 fix(panel): stop shipping a fabricated router in the binary
test / go + panel tests (push) Successful in 5m26s
release / test gate (push) Successful in 5m28s
release / apk aarch64_cortex-a53 (push) Successful in 6m7s
release / apk x86_64 (push) Successful in 3m5s
release / release apk (push) Successful in 7s
mock.ts was a static import and the mock switch was read from the query string at
runtime, so the bundle that ships inside the daemon carried a complete fictional
router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive,
without a single request to the daemon. The only tell was a line in the footer.
That is worse than any wrong number — there is no data at all and nothing says
so. It is out of the production bundle now, which is 21 kB smaller for it.

Unknown state stopped reading as good news in two more places. The kill-switch
tile treated an absent plane as armed, because the check was "not none" and
undefined satisfies it — the contract in the API types says the opposite. And the
apply page announced "daemon auto-rolled back" from its own timer, while the
daemon, seeing the state generation move, disarms and says it is NOT rolling back
in the log only.

Alerts moved to Settings. They are about the kill switch, apply failures, new
devices and subscription expiry, and they lived at the bottom of the DNS page,
while Settings mentioned them in prose with nothing to click.

Findings truncation is visible now: the notice that says how many were suppressed
arrives as info, and the attention list keeps only critical and warning, so past
fifty findings the operator saw forty-nine and no hint of the rest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:39 +03:00
omarandClaude Opus 5 4078334d85 fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.

They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.

Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.

The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.

And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.

The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:21 +03:00
omarandClaude Opus 5 0a6689b29e fix(quic,v2ray): close the sockets quic-go was never going to close
DialEarly with a packet conn the caller made sets a flag that means quic-go does
not own it: closing the transport only stops reading from the socket. Neither DNS
transport closed it. On the QUIC one it was closed on a failed handshake and
never on success, so every redial — idle timeout, retry error, engine reload —
left a UDP socket for the life of the process. On the HTTP/3 one the library
drives its own reconnects, so the leak compounds without anything in our code
looking wrong.

That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on
every reconnect without closing the previous one. Both are now owned by a watcher
tied to the connection's own context, so the socket lives exactly as long as the
connection does.

This matters more than it did last week: the shipped resolvers are DoH, and DNS
is intercepted by default now, so the whole network's query stream rides this
path on a router with 512 MB.

The same upstream commit fixes both halves. We had taken the v2ray half and not
the DNS one — the third time this session a paired fix arrived half-applied, and
the first of those cost a day of debugging. These two files are now byte-identical
to upstream so a rebase cannot reopen it.

Also from that family: websocket and httpupgrade leaked their conn on failed
handshakes, and a QUIC stream's Close did not release a blocked write.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:59 +03:00
omarandClaude Opus 5 ef22167b1a fix(apply): report the hold when the plane was armed by someone else
Booting with the armor loaded, or restarting through the handoff, left the status
saying the LAN was not being held while it was being dropped. Transient after a
successful apply, but permanent on the unreadable-config path — and there the
apply-failure alert words itself "traffic is NOT being blocked" at the exact
moment it is. That sends the operator to fix something that is not broken, past
the protection that is holding.

The table cannot be identified from here — netplane exposes no read-back and nft
does not keep comments — but identifying it is the wrong question. Holding does
not claim the holding plane is the object in the kernel; it claims the engine is
down and forwarded traffic is being dropped. A leftover full ruleset does that
too: with no engine socket the tproxy statement breaks its own rule before the
accept, so the packet reaches the forward chain unmarked and meets the primary
drop. What decides it is whether the last applied config was enabled and
fail-closed, which is exactly what the boot armor's presence already means.

So it is derived at read time rather than latched. A latch set from an inference
would have to be remembered in order to be cleared, which is the trap the active
flag already taught us. ArmHold also stops deferring to a table it cannot
inspect and installs its own render instead — the honest answer to "do not claim
a foreign table blindly" is to make it ours, and a fresh render beats a snapshot
that predates an interface rename.

Also closes the last of the detour transports: the subscription fetch took a
client and dropped it, and the exits that leak are the error ones, retried by
cron forever against a broken feed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:39 +03:00
omarandClaude Opus 5 cbda0fee0a fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.

The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.

The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.

An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.

A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:16 +03:00
95 changed files with 12752 additions and 897 deletions
+145
View File
@@ -0,0 +1,145 @@
// lx:begin l3-honest-drop
package adapter
import (
"net/netip"
"testing"
"github.com/sagernet/sing-tun"
"github.com/sagernet/sing-tun/gtcpip/header"
"github.com/stretchr/testify/require"
)
// judgeFlowRouter answers PreMatch with a canned verdict; JudgeFlow reads
// nothing else off the Router.
type judgeFlowRouter struct {
Router
result PreMatchResult
}
func (r *judgeFlowRouter) PreMatch(InboundContext, []byte) PreMatchResult { return r.result }
// judgeFlowPort is the tun.Port half of a FlowOutbound. inet4 is what
// PortAddresses reports for IPv4 — the one field the two ICMP consumers in
// sing-tun disagree about (see the comment on
// TestJudgeFlowICMPToBoundPortStaysAFlow).
type judgeFlowPort struct {
Outbound
inet4 netip.Addr
}
func (o *judgeFlowPort) Tag() string { return "wg-out" }
func (o *judgeFlowPort) Type() string { return "wireguard" }
func (o *judgeFlowPort) PortAddresses() (netip.Addr, netip.Addr) {
return o.inet4, netip.Addr{}
}
func (o *judgeFlowPort) PortMTU() uint32 { return 1420 }
func (o *judgeFlowPort) AttachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) DetachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) WritePackets(packets [][]byte) error { return nil }
// judgeFlowNonPort is a FlowOutbound-shaped result that is NOT a tun.Port — the
// interface drift the second line of defense in JudgeFlow exists for.
type judgeFlowNonPort struct {
Outbound
}
func (o *judgeFlowNonPort) Tag() string { return "drifted" }
func (o *judgeFlowNonPort) Type() string { return "drifted" }
func judgeFlow(t *testing.T, protocol uint8, result PreMatchResult) tun.FlowVerdict {
t.Helper()
return JudgeFlow(
&judgeFlowRouter{result: result},
"l3-in", "tun", protocol,
netip.MustParseAddrPort("192.168.1.2:1234"),
netip.MustParseAddrPort("1.1.1.1:1234"),
nil,
)
}
const (
judgeFlowICMP = uint8(header.ICMPv4ProtocolNumber)
judgeFlowTCP = uint8(header.TCPProtocolNumber)
)
// TestJudgeFlowICMPToBoundPortStaysAFlow is the guard on the ONE fix that must
// not be made here.
//
// sing-tun has two ICMP consumers with different requirements on the port:
//
// - ForwardDispatcher.createFlow (flow_dispatch.go) needs only a VALID port
// address — it NATs the echo identifier and rewrites the source to that
// address. This is the path every unfragmented LAN ping takes, and it is
// what makes ping-through-WireGuard/AWG work at all.
// - ICMPForwarder.installFlow (stack_gvisor_icmp.go) additionally requires the
// address to be UNSPECIFIED, because it writes the packet to the port
// unmodified. A WireGuard endpoint reports its concrete interface address
// (transport/wireguard/port.go), so installFlow declines and HandlePacket
// falls through to forging the echo reply.
//
// The tempting fix — "for ICMP, refuse ActionFlow when PortAddresses() is not
// unspecified, so the verdict becomes a drop and the forgery is unreachable" —
// is applied HERE, in the one function both consumers share, with byte-identical
// arguments from either. It would therefore kill the working path too: every
// ping through WireGuard/AWG, fragmented or not, would drop, and l3_tunnel would
// carry nothing but `direct`. Keep this test failing loudly if anyone tries.
func TestJudgeFlowICMPToBoundPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.MustParseAddr("10.2.0.2")}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action,
"ICMP to a WireGuard/AWG endpoint must stay a flow: the forward dispatcher NATs it by echo identifier and this is the whole point of l3_tunnel")
require.Same(t, tun.Port(port), verdict.Port)
}
// The `direct` shape: an unspecified port address. Both consumers accept it.
func TestJudgeFlowICMPToUnspecifiedPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.IPv4Unspecified()}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action)
require.Same(t, tun.Port(port), verdict.Port)
}
// PreMatchDrop is the honest verdict and must arrive as ActionDrop: it is the
// only value (besides Reject) that stops ICMPForwarder.HandlePacket before the
// Echo -> EchoReply rewrite.
func TestJudgeFlowICMPDropReachesTheStackAsDrop(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchDrop})
require.Equal(t, tun.ActionDrop, verdict.Action)
}
// The second line of defense: a PreMatchFlow whose outbound is not a tun.Port
// must not degrade ICMP to ActionAccept, because Accept is the forged reply.
func TestJudgeFlowICMPNonPortOutboundDrops(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionDrop, verdict.Action,
"FlowOutbound and tun.Port are distinct interfaces; a drift between them must not silently re-enable the echo forger")
}
func TestJudgeFlowTCPNonPortOutboundAccepts(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionAccept, verdict.Action,
"for TCP, falling back to Accept is upstream behaviour and must stay untouched")
}
// TCP keeps every mapping it had, including the Continue -> Accept default that
// is a forgery only for ICMP.
func TestJudgeFlowTCPContinueStaysAccept(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchContinue})
require.Equal(t, tun.ActionAccept, verdict.Action)
}
func TestJudgeFlowTCPBypassStaysBypass(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchBypass})
require.Equal(t, tun.ActionBypass, verdict.Action)
}
// lx:end l3-honest-drop
+11
View File
@@ -75,7 +75,18 @@ func JudgeFlow(router Router, inbound string, inboundType string, network uint8,
case PreMatchFlow:
port, isPort := result.Outbound.(tun.Port)
if !isPort {
// lx:begin l3-honest-drop
// Second line of defense behind route.(*Router).preMatchFlow: a
// PreMatchFlow result already implies the outbound is an
// adapter.FlowOutbound, but FlowOutbound and tun.Port are distinct
// interfaces, and a drift between them must not degrade ICMP to
// ActionAccept — the TUN stack would then forge the echo reply
// itself instead of admitting the tunnel cannot carry the packet.
if networkName == N.NetworkICMP {
return tun.FlowVerdict{Action: tun.ActionDrop}
}
return tun.FlowVerdict{Action: tun.ActionAccept}
// lx:end l3-honest-drop
}
verdict := tun.FlowVerdict{Action: tun.ActionFlow, Port: port, UDPTimeout: result.UDPTimeout, NewTracker: result.NewTracker}
if result.Destination.IsValid() {
+141
View File
@@ -0,0 +1,141 @@
// lx:begin health-board
package urltest
import (
"strconv"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/adapter"
)
// captureEvictions swaps the eviction notice sink for the duration of a test and
// returns a func that reads back everything reported.
func captureEvictions(t *testing.T) func() []string {
t.Helper()
var (
mu sync.Mutex
msgs []string
)
orig := boardEvictionLog
boardEvictionLog = func(m string) {
mu.Lock()
msgs = append(msgs, m)
mu.Unlock()
}
t.Cleanup(func() { boardEvictionLog = orig })
return func() []string {
mu.Lock()
defer mu.Unlock()
return append([]string(nil), msgs...)
}
}
// TestBoardHoldsAGenerationWithoutEvicting is the "what it holds" half of the
// bound. A live generation on this box is ~1200 tags (≈380 nodes plus their
// per-group egress copies and chain hops); the board must carry that — and a
// second generation's worth of overlap during a subscription rename — with no
// eviction at all, or the ceiling would be silently degrading real health data.
func TestBoardHoldsAGenerationWithoutEvicting(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
const generation = 1200
for gen := 0; gen < 2; gen++ {
for i := 0; i < generation; i++ {
s.StoreURLTestHistory("gen"+strconv.Itoa(gen)+"-node-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 20})
}
}
if got := s.Evicted(); got != 0 {
t.Fatalf("two full generations (%d tags) evicted %d entries; the board must hold them",
2*generation, got)
}
if msgs := read(); len(msgs) != 0 {
t.Fatalf("unexpected eviction notices: %v", msgs)
}
// Everything is still readable.
if s.LoadURLTestHistory("gen0-node-0") == nil {
t.Fatalf("the first tag of the first generation was lost without an eviction")
}
}
// TestBoardEvictsOldestAndSaysSo is the "what happens when it overflows" half.
// Overflow must (a) actually bound the map, (b) drop the LEAST RECENTLY MEASURED
// tags — on this box, exactly the ones no config names any more — and (c) be
// audible: a silent eviction is a health board quietly forgetting nodes it is
// still being asked about.
func TestBoardEvictsOldestAndSaysSo(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
base := time.Now().Add(-24 * time.Hour)
// Stale generation first: measured a day ago, nothing since.
const stale = 1500
for i := 0; i < stale; i++ {
s.StoreURLTestHistory("stale-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: base.Add(time.Duration(i) * time.Millisecond), Delay: 30})
}
if s.Evicted() != 0 {
t.Fatalf("evicted before the ceiling was reached")
}
// Now push past the ceiling with fresh measurements.
for i := 0; i <= maxBoardEntries; i++ {
s.StoreURLTestHistory("fresh-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 15})
}
if got := s.Evicted(); got == 0 {
t.Fatalf("board grew past %d entries without evicting anything — it is still unbounded", maxBoardEntries)
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("board holds %d entries, above the %d ceiling", size, maxBoardEntries)
}
// The day-old generation is what went, not the fresh one.
if s.LoadURLTestHistory("stale-0") != nil {
t.Fatalf("the oldest observation survived while newer ones were dropped")
}
if s.LoadURLTestHistory("fresh-"+strconv.Itoa(maxBoardEntries)) == nil {
t.Fatalf("the newest measurement was evicted")
}
msgs := read()
if len(msgs) == 0 {
t.Fatalf("entries were evicted with no notice — eviction must never be silent")
}
m := msgs[0]
for _, want := range []string{"health board full", "evicted", "re-probed"} {
if !strings.Contains(m, want) {
t.Fatalf("eviction notice %q does not say %q", m, want)
}
}
}
// TestBoardEvictionThroughMarkFailed pins the OTHER write path. MarkFailed is how
// a dead node is recorded, and a flood of dead renamed nodes is exactly the shape
// of the leak — so it has to prune too, not just the success path.
func TestBoardEvictionThroughMarkFailed(t *testing.T) {
captureEvictions(t)
s := NewHistoryStorage()
for i := 0; i <= maxBoardEntries; i++ {
s.MarkFailed("dead-" + strconv.Itoa(i))
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("MarkFailed grew the board to %d, above the %d ceiling", size, maxBoardEntries)
}
if s.Evicted() == 0 {
t.Fatalf("MarkFailed never prunes — the failure path is still unbounded")
}
}
// lx:end health-board
+118
View File
@@ -10,11 +10,128 @@
package urltest
import (
"sort"
"strconv"
"sync"
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/log"
)
// --- board capacity ---------------------------------------------------------
//
// The board is the one structure in the daemon whose key space is chosen by
// somebody else. Its keys are outbound TAGS, and on this box a tag is a node
// NAME straight out of the subscription — plus the derived per-group egress
// copies ("group-<g>-m<i>-<node>") and per-chain hop copies the probe planner
// creates for the same nodes. Providers rename their nodes freely, so a daily
// subscription refresh introduces a whole new generation of keys, while the
// store itself is pinned to the ENGINE's context (shater/engine.New) and so
// outlives every generation and every Apply — by design, so health survives a
// config change.
//
// Nothing ever removed a key. DeleteURLTestHistory exists but no shater path
// calls it (only daemon/ and clashapi/, which this fork does not run), so the
// map was strictly append-only for the life of the process — and the process is
// expected to live for months.
//
// The arithmetic: ~380 nodes, and a config with a couple of egress-bound groups
// plus a handful of chains puts a LIVE generation at roughly 380 base tags +
// 2x380 group copies + ~100 chain copies ≈ 1200 keys. One new generation per day
// is ~440k keys a year, at ~200 B per entry (map bucket + a tag string that is
// routinely 30-50 B with flag emoji, + a 56 B URLTestHistory) ≈ 88 MB of a
// 512 MB box — spent entirely on nodes that no longer exist.
const (
// maxBoardEntries is the hard ceiling. 4096 is ~3.4 live generations, so the
// board comfortably holds the current config plus the overlap while a
// subscription refresh swaps names, and still costs under a megabyte. A tighter
// bound would start evicting tags the running config actually uses; a looser one
// would stop being a bound in any useful sense.
maxBoardEntries = 4096
// keepBoardEntries is the prune target: drop a quarter at a time so the
// O(n log n) selection is amortised over ~1024 inserts instead of running on
// every probe once the board is full.
keepBoardEntries = 3072
)
// boardEvictionLog reports an eviction. A package var so tests can capture it;
// production leaves it writing to the process log, which under procd is the same
// syslog/logsink stream every other daemon line lands in.
//
// Eviction is NEVER silent. It is not free either: an evicted tag reverts to
// "untested" and its next probe re-measures it, so a board that evicts entries
// belonging to the LIVE config is a board whose ceiling is too low — and the only
// way anyone finds that out is this line.
var boardEvictionLog = func(msg string) { boardLogger().Warn(msg) }
// pruneLocked drops the least-recently-OBSERVED entries when the board exceeds
// maxBoardEntries. "Least recently observed" is max(LastOK, LastFail): the entry
// nothing has measured for the longest is, on this box, precisely a tag that no
// longer exists in any config — a renamed node, a removed group copy, a retired
// chain hop. Caller holds access.
func (s *HistoryStorage) pruneLocked() {
if len(s.delayHistory) <= maxBoardEntries {
return
}
type kv struct {
tag string
seen time.Time
}
all := make([]kv, 0, len(s.delayHistory))
for tag, h := range s.delayHistory {
seen := h.LastOK
if h.LastFail.After(seen) {
seen = h.LastFail
}
all = append(all, kv{tag, seen})
}
sort.Slice(all, func(i, j int) bool { return all[i].seen.Before(all[j].seen) })
drop := len(all) - keepBoardEntries
var oldest time.Time
for i := 0; i < drop; i++ {
if i == 0 {
oldest = all[i].seen
}
delete(s.delayHistory, all[i].tag)
}
s.evicted += uint64(drop)
msg := "urltest: health board full (" + strconv.Itoa(maxBoardEntries) + " tags) — evicted " +
strconv.Itoa(drop) + " least-recently-measured entries (" + strconv.FormatUint(s.evicted, 10) +
" total since start); they revert to untested and will be re-probed"
if !oldest.IsZero() {
msg += "; oldest observation was " + time.Since(oldest).Truncate(time.Second).String() + " ago"
}
boardEvictionLog(msg)
}
// Evicted reports how many entries the capacity bound has dropped since the store
// was created. Nonzero means the board reached maxBoardEntries at least once.
func (s *HistoryStorage) Evicted() uint64 {
if s == nil {
return 0
}
s.access.RLock()
defer s.access.RUnlock()
return s.evicted
}
// boardLogger is the process-wide fallback logger for eviction notices. The store
// is built from a plain constructor with no logger in sight (box.New, the daemon,
// shater/engine all call NewHistoryStorage()), so rather than change that
// signature everywhere the notice goes to the standard logger — which on the
// router is the daemon's own stderr, i.e. the same sink logsink owns.
var (
boardLogOnce sync.Once
boardLog log.ContextLogger
)
func boardLogger() log.ContextLogger {
boardLogOnce.Do(func() { boardLog = log.StdLogger() })
return boardLog
}
// HealthVerdict classifies a stored history entry at read time.
type HealthVerdict int
@@ -54,6 +171,7 @@ func (s *HistoryStorage) MarkFailed(tag string) {
updated.Delay = previous.Delay
}
s.delayHistory[tag] = updated
s.pruneLocked()
s.notifyUpdated()
s.access.Unlock()
}
+9
View File
@@ -21,6 +21,10 @@ type HistoryStorage struct {
access sync.RWMutex
delayHistory map[string]*adapter.URLTestHistory
updateHooks []*observable.Subscriber[struct{}]
// evicted counts entries dropped by the capacity bound (board_lx.go). The map
// is keyed by outbound tags chosen by a subscription provider, so it needs a
// ceiling; see the comment on maxBoardEntries.
evicted uint64
}
func NewHistoryStorage() *HistoryStorage {
@@ -71,6 +75,11 @@ func (s *HistoryStorage) StoreURLTestHistory(tag string, history *adapter.URLTes
}
// lx:end health-board
s.delayHistory[tag] = history
// lx:begin health-board — the map is keyed by provider-chosen tags and the
// store outlives every engine generation, so it must bound itself here: no
// shater path ever calls DeleteURLTestHistory. See maxBoardEntries.
s.pruneLocked()
// lx:end health-board
s.notifyUpdated()
s.access.Unlock()
}
+6
View File
@@ -126,6 +126,12 @@ func (t *HTTP3Transport) newTransport() *http3.Transport {
conn.Close()
return nil, dialErr
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-quicConn.Context().Done()
conn.Close()
}()
return quicConn, nil
},
TLSClientConfig: t.tlsConfig,
+351
View File
@@ -0,0 +1,351 @@
package quic
import (
"context"
"crypto/tls"
"net"
"net/http"
"net/url"
"sync"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
mDNS "github.com/miekg/dns"
)
var _ N.Dialer = (*trackingDialer)(nil)
// These tests pin down who owns the UDP socket handed to quic-go.
//
// quic-go's Dial/DialEarly take a net.PacketConn but do NOT take ownership of
// it: quic.setupTransport() builds a Transport with createdConn=false, and
// Transport.Close() then only calls conn.SetReadDeadline(time.Now()) instead of
// conn.Close(). So every QUIC connection torn down here — idle timeout, a
// retryable error, an engine reload calling Reset() — used to strand the UDP
// socket that carried it for the rest of the process's life. On a router that
// resolves through DoQ/DoH3 for months that is an unbounded fd leak.
//
// Both tests reconnect once and assert the socket from the FIRST connection is
// actually closed. Without the `<-conn.Context().Done() -> rawConn.Close()`
// watchdogs in quic.go / http3.go they fail on that assertion.
type trackedConn struct {
net.Conn
closeOnce sync.Once
closed chan struct{}
}
func (c *trackedConn) Close() error {
c.closeOnce.Do(func() { close(c.closed) })
return c.Conn.Close()
}
// trackingDialer hands out real UDP sockets and remembers every one of them.
type trackingDialer struct {
access sync.Mutex
conns []*trackedConn
}
func (d *trackingDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn, closed: make(chan struct{})}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *trackingDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func (d *trackingDialer) count() int {
d.access.Lock()
defer d.access.Unlock()
return len(d.conns)
}
func (d *trackingDialer) at(index int) *trackedConn {
d.access.Lock()
defer d.access.Unlock()
return d.conns[index]
}
func (d *trackingDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
func requireClosed(t *testing.T, conn *trackedConn, what string) {
t.Helper()
select {
case <-conn.closed:
case <-time.After(5 * time.Second):
t.Fatalf("%s: the UDP socket of the retired QUIC connection was never closed — quic-go does not own it, we must", what)
}
}
func requireDialed(t *testing.T, dialer *trackingDialer, want int) {
t.Helper()
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if dialer.count() >= want {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Fatalf("expected at least %d dial(s), got %d", want, dialer.count())
}
func testServerTLSConfig(t *testing.T, nextProtos []string) *tls.Config {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
return &tls.Config{
Certificates: []tls.Certificate{*certificate},
NextProtos: nextProtos,
MinVersion: tls.VersionTLS13,
}
}
func testClientTLSConfig(t *testing.T, nextProtos []string) sbTLS.Config {
t.Helper()
config, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
config.SetNextProtos(nextProtos)
return config
}
// startDoQServer serves a minimal DoQ responder and returns its address.
func startDoQServer(t *testing.T) M.Socksaddr {
t.Helper()
listener, err := quic.ListenAddr("127.0.0.1:0", testServerTLSConfig(t, []string{"doq"}), nil)
if err != nil {
t.Fatal(err)
}
ctx, cancel := context.WithCancel(context.Background())
t.Cleanup(func() {
cancel()
listener.Close()
})
go func() {
for {
conn, acceptErr := listener.Accept(ctx)
if acceptErr != nil {
return
}
go func(conn *quic.Conn) {
for {
stream, streamErr := conn.AcceptStream(ctx)
if streamErr != nil {
return
}
go func(stream *quic.Stream) {
defer stream.Close()
request, readErr := transport.ReadMessage(stream)
if readErr != nil {
return
}
response := new(mDNS.Msg)
response.SetReply(request)
transport.WriteMessage(stream, 0, response)
}(stream)
}
}(conn)
}
}()
return M.ParseSocksaddr(listener.Addr().String())
}
func testQuery() *mDNS.Msg {
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
return message
}
func TestQUICTransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
serverAddr := startDoQServer(t)
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeQUIC, "test-doq", nil),
dialer: dialer,
serverAddr: serverAddr,
tlsConfig: testClientTLSConfig(t, []string{"doq"}),
connection: transport.NewConnPool(transport.ConnPoolOptions[*quic.Conn]{
Mode: transport.ConnPoolSingle,
IsAlive: func(conn *quic.Conn) bool {
return conn != nil && !common.Done(conn.Context())
},
Close: func(conn *quic.Conn, _ error) {
conn.CloseWithError(0, "")
},
}),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
// Retire the connection the way a retryable error or an engine reload does.
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
// The reconnect must still work, on a fresh socket.
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err := dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func TestHTTP3TransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
message, err := readRequestMessage(request)
if err != nil {
writer.WriteHeader(http.StatusBadRequest)
return
}
response := new(mDNS.Msg)
response.SetReply(message)
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
writer.Write(rawResponse)
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
serverAddr := M.ParseSocksaddr(listener.Addr().String())
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
stdConfig := &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
}
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: serverAddr,
tlsConfig: stdConfig,
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err = dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func readRequestMessage(request *http.Request) (*mDNS.Msg, error) {
defer request.Body.Close()
rawMessage := make([]byte, 4096)
n, err := readFull(request.Body, rawMessage)
if err != nil {
return nil, err
}
var message mDNS.Msg
err = message.Unpack(rawMessage[:n])
if err != nil {
return nil, err
}
return &message, nil
}
func readFull(reader interface{ Read([]byte) (int, error) }, buffer []byte) (int, error) {
var total int
for total < len(buffer) {
n, err := reader.Read(buffer[total:])
total += n
if err != nil {
if total > 0 {
return total, nil
}
return total, err
}
}
return total, nil
}
+12
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"os"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/sing-box/adapter"
@@ -117,6 +118,12 @@ func (t *Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg,
rawConn.Close()
return nil, E.Cause(err, "establish QUIC connection")
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-earlyConnection.Context().Done()
rawConn.Close()
}()
return earlyConnection, nil
})
if err != nil {
@@ -144,6 +151,11 @@ func (t *Transport) exchange(ctx context.Context, message *mDNS.Msg, conn *quic.
return nil, E.Cause(err, "open stream")
}
defer stream.CancelRead(0)
stopWatch := context.AfterFunc(ctx, func() {
stream.CancelRead(0)
_ = stream.SetWriteDeadline(time.Now())
})
defer stopWatch()
err = transport.WriteMessage(stream, 0, message)
if err != nil {
stream.Close()
+43
View File
@@ -12,6 +12,49 @@ as GitHub **pre-releases** and never become "Latest".
#### Unreleased (shater)
**`l3-honest-drop` — ICMP routed to an L4-only outbound is dropped, not
forged** — ships with `shaterd` (part of the shater L3 ingress,
`docs-shater/DECISIONS.md` D25), not as an lx release tag; recorded here because
it edits two upstream files. Without it the TUN stack answers an unroutable echo
ITSELF — sing-tun's `ICMPForwarder.HandlePacket` rewrites Echo→EchoReply
whenever the flow judgment comes back Accept (`stack_gvisor_icmp.go`) — so a
ping routed to vless/vmess/… would read as a working tunnel while the packet
never left the router.
* **`route/route.go` (`PreMatch`)** — the pre-match walk was renamed to
`preMatch` and the exported `PreMatch` became a thin FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for
`N.NetworkICMP`. An earlier version overrode `continueResult` inside
`preMatchFlow` instead; that covered only the exits reaching that function and
left three of the walk's own exits forging — the `prepareMatchMetadata` error
return, the sniff bail-outs, and the `default:` arm of the rule-action switch
(every action pre-match has no arm for: `hijack-dns`, `direct`, …). A guard on
the single return value cannot be outgrown by a new exit. `PreMatchBypass` is
folded in because sing-tun implements `ActionBypass` on the nfqueue plane only
— on the TUN path it lands in the same `default:` arm as Accept, i.e. forges.
* **`adapter/router.go` (`JudgeFlow`, the `!isPort` branch)** — ICMP returns
`ActionDrop` where it fell through to `ActionAccept`. Second line of defense:
`adapter.FlowOutbound` and `tun.Port` are distinct interfaces, and a drift
between them must not quietly re-enable the forged reply.
* **TCP/UDP behaviour is unchanged** — `PreMatchContinue` still means "take the
ordinary connection route" for both, `PreMatchBypass` still means bypass, and
the `!isPort` fallthrough still returns `ActionAccept` for them; pinned by
`route/prematch_icmp_lx_test.go` and `adapter/judgeflow_icmp_lx_test.go`
(both inside the marker), each ICMP case having an explicit TCP/UDP twin.
* **NOT covered: a FRAGMENTED echo to a WireGuard/AWG outbound is still
forged** — sing-tun's `ForwardDispatcher.Dispatch` returns before asking for a
verdict at all when `parsed.fragment`, and the reassembled packet reaches
`ICMPForwarder.HandlePacket`, whose `installFlow` demands an UNSPECIFIED port
address that a WireGuard endpoint never has. Fixing it inside `JudgeFlow`
is NOT possible — both consumers call it with identical arguments and the
working path needs the concrete address. Full chain, the two viable fixes and
the trap are in `docs-shater/DECISIONS.md` D25, "KNOWN HOLE".
* **Rebase cost: two small marked blocks** (`lx:begin/end l3-honest-drop`, a
wrapper function in `route/route.go` and one branch body in
`adapter/router.go`) plus the two self-contained test files — carried across
an upstream rebase by eye. Note that `PreMatch`'s own body now lives in
`preMatch`, so an upstream change to the walk applies to that function.
**Fork-layer + control-plane rework of proxy health** — ships with `shaterd`
(the shater router daemon), not as an lx release tag; recorded here because the
load-bearing half lives in fork zones (`common/urltest`, `protocol/group`).
+430
View File
@@ -319,6 +319,10 @@ Three values, not two, because the leaks differ in *kind*: an ICMP echo is ephem
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
*(Refined 2026-07-26 by D25: still true of TPROXY — but ICMP echo now has an
opt-in data plane of its own, the dedicated L3 TUN, so the policy no longer
speaks alone for ping; it keeps sole charge of ESP/GRE/IGMP and of the degraded
paths.)*
**Fail-open degradations must be visible in the panel, not only in `logread`.** The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
@@ -772,3 +776,429 @@ server, which restores exactly the pre-D24 behaviour and clears the notice. Unti
that lands, an operator can get the same result by setting `endpoint_resolver` to a
direct resolver. Note the hazard is **not** created by D24 — any config with two
resolvers has it today; the default merely makes it universal.
## D25 — L3 ingress: LAN ICMP rides a dedicated TUN through the tunnel, not a policy verdict
Decided 2026-07-26. D17 made everything TPROXY cannot divert an explicit policy
(`Globals.Untunnelable` = block | icmp | direct) — and its premise still holds:
kernel TPROXY delivers a packet by handing it to a listening SOCKET, and sockets
exist for TCP and UDP only, so an ICMP echo has nothing to be handed to. But a
policy can only choose between losing the packet and leaking it with the
client's real source address; neither ever puts a ping THROUGH the tunnel. This
decision adds the data plane D17 could not have: **`globals.l3_tunnel` (opt-in,
default off; `model.Globals.L3Tunnel`) opens a second, dedicated ingress — a TUN
device — and LAN ICMP enters the engine as raw IP packets**, where the ordinary
route rules pick an outbound exactly as for any flow. The policy is refined, not
repealed: it keeps sole charge of the protocols the engine cannot ingest at all,
and of the degraded paths (both below).
**The whole mechanism is one mark, one rule, one device — the TPROXY plane is
untouched.** The nft prerouting chain stamps `L3Mark` (= fwmark_base + 0x80,
`netplane/nft.go` `l3MarkOffset`) on LAN `ip protocol icmp` / `meta l4proto
ipv6-icmp` ONLY, and only after every local plane was already accepted
(fib-local, RFC1918/link-local/multicast daddr sets) and — for v6 — after a
unicast ND/NA carve-out, because one tunnelled neighbour probe is enough to take
the LAN's v6 plane down (`renderNft`, the L3 block). `addL3Routing`
(`netplane/apply.go`) binds that mark to a table (= table_base + 0x08) whose
only content is `default dev shater-l3`; del-then-add idempotent, and a failed
rule or route is a NAMED operator warning, never an apply abort. `generate`
emits the synthetic `l3-in` TUN inbound bound to exactly `netplane.L3Device`,
MTU 65535 (the largest IP datagram there can be, so the KERNEL can never
fragment on the way in — see "the device MTU is not a tunnel budget" below),
point-to-point /30 + /126 addresses from private space,
the v6 one only when `globals.ipv6` is on — and only next to a tproxy inbound:
the ingress rides the same LAN divert plane, and without one the TUN would sit
dark while the config claims ICMP is tunnelled, so it is skipped with a warning
(`generate/inbound.go`, `appendL3TunInbound`). `shater/registry` registers the
`tun` inbound type; that costs no new build tag and no meaningful size because
`with_wireguard` already requires `with_gvisor` (D23, `scripts/router-tags.sh`).
**`auto_route: false` is load-bearing, not a default we happened to keep.**
sing-box's auto_route rewrites the router's MAIN routing table — it would drag
everything the router itself sends (WAN traffic, DNS, the tunnel's own underlay)
into this TUN. The fwmark rule + dedicated table above is deliberately the ONLY
entrance, and disabling the feature can never strand a stale default route in
main (`generate/inbound.go`; pinned by `TestL3TunnelEmitsTunInbound`).
**`stack: "gvisor"` is a deliberate choice, and the tempting reason for it is
wrong.** It is TRUE that sing-tun's system stack answers an ICMP echo LOCALLY —
`processIPv4ICMP` rewrites Echo→EchoReply in place and swaps the addresses
(sing-tun `stack_system.go:648`; the v6 twin sits right under it). It is FALSE
that this makes the system stack unusable here: `dispatchIPv4`
(`stack_system.go:355-372`) hands the packet to the SAME `ForwardDispatcher`
first and only falls through to that forger for packets addressed to the TUN
itself, exactly as the gVisor filter does (`stack_gvisor_filter.go:52-113`).
Both stacks would forward. gvisor is chosen because it is already linked —
`with_wireguard` requires `with_gvisor` (D23), so it costs no tag and no new
code path — and because it is the combination the integration test actually
exercises. Do not re-derive this as "the system stack fakes ping": it fakes ping
only where the dispatcher declined the packet.
**The ceiling is ICMP echo, and it is upstream's dispatcher — NOT the netstack.**
This distinction matters because the netstack answer is the intuitive one and it
is wrong. On the forward path a WireGuard/AWG endpoint never consults gVisor at
all: `Endpoint.WritePackets` (`transport/wireguard/port.go:21-58`) reads the IP
version and the destination address and hands the raw bytes to
`wgDevice.InputPackets` — the protocol byte is never examined — and
`returnDeviceWrapper.Write` (`:127-157`) offers every decrypted packet to
`returnPath.ReturnPackets` before the stack sees it. WireGuard would carry ESP
today if anything handed it one. What refuses is `ForwardDispatcher`: its parser
sets `hasFlow` for TCP, UDP and ICMP echo alone (`flow_parse.go`,
`parseTransport`, the echo identifier serving as the pseudo-port), and
`createFlow` NATs through a port-shaped selector (`flow_dispatch.go:325`,
`allocateSelector`) that ESP, AH and GRE do not have. So ESP/AH/GRE/IGMP/SCTP
cannot enter the engine in ANY configuration and REMAIN on the D17 policy —
or on the kernel egress of D26, which sidesteps the dispatcher entirely. The nft
plane encodes the same boundary on purpose: it marks `icmp`/`ipv6-icmp` only,
never `l4proto != { tcp, udp }`, because a marked ESP packet would enter the
device and vanish — a black hole wearing a tunnel's name — instead of receiving
the policy's honest verdict (`netplane/nft.go`, the prerouting L3 comment).
**What works and what does not, read off the upstream source.** ping v4/v6 —
yes. Windows `tracert` — yes: the gVisor return path recognises
`ICMPv4TimeExceeded` and `ICMPv4DstUnreachable` alongside EchoReply and NATs
them back to the LAN client (`stack_gvisor_icmp.go:341+`, `returnPacket`). IPv6
traceroute — intermediate hops stay invisible: the v6 branch of the same
function accepts EchoReply only, so just the final destination answers. Several
LAN clients behind the one tunnel address are already solved upstream:
`ForwardDispatcher` NATs by echo identifier and rewrites the source to the
outbound's port address (`flow_dispatch.go:325+`, `createFlow`; `icmpFlowKey`) —
we wrote no NAT of our own.
**Which outbounds can carry it.** The contract is `adapter.FlowOutbound`
(= `Outbound` + `tun.Port` + `PreMatchFlow`, `adapter/outbound.go`). In-tree
implementors: the WireGuard/AWG endpoint (`protocol/wireguard`), `direct`
(`protocol/direct`), `bridge` (`protocol/bridge`), `tailscale`
(`protocol/tailscale`). Of those, the shaterd registry can construct only
WireGuard/AWG and direct (`shater/registry/registry.go` — bridge and tailscale
are not registered). Every proxy protocol — vless/vmess/trojan/shadowsocks/
hysteria2/tuic/socks/http/shadowtls — is L4-only and cannot. Recorded as a known
gap: `masque` is L3 by nature (CONNECT-IP; it builds a userspace gVisor stack
per tunnel, `protocol/masque/outbound.go`) but implements no `tun.Port` and is
not in the shater registry, so today it cannot carry the ingress. Wiring it up
is possible future work, not a promise.
**ICMP to an L4-only outbound is DROPPED, and that took patching upstream files
(the `lx:l3-honest-drop` delta — see `docs-lx/lx-changelog.md`).** In the gVisor
stack the fallthrough verdict is a forgery: `ICMPForwarder.HandlePacket` answers
the echo ITSELF (Echo→EchoReply + address swap) whenever the flow judgment comes
back Accept (`stack_gvisor_icmp.go:120`), and upstream maps "no flow route" to
exactly that Accept — so a ping routed to vless would read as tunnelled while
the packet died on the router. Two small marked hunks make the truth observable:
`route/route.go` wraps the whole pre-match walk — the walk itself became
`preMatch`, and the exported `PreMatch` is now a FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for `N.NetworkICMP` —
and `adapter/router.go` (`JudgeFlow`, the `!isPort` branch) returns `ActionDrop`
for ICMP where it fell through to `ActionAccept` — the second line of defense,
because `FlowOutbound` and `tun.Port` are distinct interfaces and a drift
between them must not quietly re-enable the forger. TCP/UDP verdicts are
byte-identical; `route/prematch_icmp_lx_test.go` and
`adapter/judgeflow_icmp_lx_test.go` pin both directions. The operator-facing
text says the same out loud (`shater/apply/warnings.go`): proxy-routed addresses
"cannot be pinged at all — deliberately".
> **Why a funnel and not an override inside the walk.** The first version of
> this delta overrode the pre-declared `continueResult` inside `preMatchFlow`
> and claimed to cover "every exit point of the function at once". It covered
> every exit of THAT function; the walk above it has exits of its own that never
> reach it — the `prepareMatchMetadata` error return (which arrived later, with
> the shared-metadata refactor, upstream `b911fb078`), the sniff bail-outs, and
> the `default:` arm of the rule-action switch, which catches every action
> pre-match has no arm for (`hijack-dns`, `direct`, and whatever upstream adds
> next). Each of those returned `PreMatchContinue`, i.e. `tun.ActionAccept`,
> i.e. the forged reply. A guard on the single return value cannot be outgrown
> by a new exit. `PreMatchBypass` joined the drop for the same reason: sing-tun
> implements `ActionBypass` on the nfqueue plane only — the name appears nowhere
> in `flow_dispatch.go` or `stack_gvisor_icmp.go` — so on the TUN path it lands
> in the same `default:` arm as Accept and forges too. There is no honest bypass
> for a packet that is already inside the engine's TUN.
**The device MTU is NOT a tunnel budget, and pretending it was manufactured
forged replies.** `l3-in` is created with MTU **65535**, not the tunnel's 1420,
and the maximum is the whole argument. This MTU governs exactly one thing:
whether the KERNEL splits a packet on its way INTO the device. What the engine
then puts into the tunnel is sized separately and correctly, against the
OUTBOUND's MTU — `ForwardDispatcher.forwardToPort` (`flow_dispatch.go:445-481`)
measures every forwarded packet against `Port.PortMTU()` and either fragments to
it (no DF, `fragmentIPv4Packet`) or answers a well-formed `fragmentation needed`
quoting it (DF, `buildFragmentationNeeded`, source = the far host, so PMTU
discovery works end to end). That machinery was always there; it was simply
never handed a whole packet.
At 1420 it wasn't. Anything above 1392 bytes of payload was fragmented by the
kernel at this device, and a fragment is the one thing sing-tun will not judge:
`Dispatch` (`flow_dispatch.go:176-177`) returns on `parsed.fragment` BEFORE
calling `JudgeFlow` at all. The fragments fell through to the gVisor stack —
promiscuous and spoofing (`stack_gvisor.go:219-223`) — which reassembled them
and handed the echo to `ICMPForwarder.HandlePacket` (`stack_gvisor_icmp.go:105+`),
whose `installFlow` (`:233-244`) writes to the port UNMODIFIED and therefore
demands a port address that is valid **and UNSPECIFIED**. `direct` qualifies
(`IPv4Unspecified()`); a WireGuard/AWG endpoint reports its concrete interface
address (`transport/wireguard/port.go:13`) and does not. So it declined, and
`HandlePacket` fell past the switch and FORGED the reply: `SetType(EchoReply)` +
address swap. Net effect on the operator's bench: `ping -s 1392` honest,
`ping -s 1393` a lie told by the router — and the lie was, of course, only for
the outbounds this feature exists for. (Upstream applies the very same
unspecified test and answers it honestly in the cloudflared ICMP handler,
`protocol/cloudflare/inbound.go:163-167`: it drops. Only the TUN path forges.)
65535 rather than "big enough": no IP datagram can exceed it, so the kernel
CANNOT fragment at this device, for any packet, ever. Any smaller value leaves
a band of sizes open and re-opens the class. It is also sing-box's own default
TUN MTU on Linux. Pinned by `TestL3TunnelMTULeavesNothingForTheKernelToFragment`
and `TestL3TunnelMTUIsNotATunnelBudget` (`generate/l3mtu_test.go`), and — the
assertion that matters — by the integration test reading the MTU back off the
real kernel device, since a kernel that clamped it would restore the forgery
without changing a generated byte.
Memory was MEASURED, not reasoned about: three paired runs of
`TestIntegrationL3TunInboundStarts` under `-test.memprofilerate=1` (exact
accounting, not sampled) allocate 5.41 / 5.48 / 5.47 MB at 65535 against
5.76 / 5.46 / 5.70 MB at 1420, and a `-diff_base` profile attributes every
difference to netlink interface enumeration. Nothing in the read path scales
with the MTU: gVisor reads through `fdbased.BufConfig`, which sing-tun's `init`
pins to a single 65535-byte view regardless of MTU, and `fdbased` keeps `mtu`
only to return it from `MTU()`. Two adjacent facts, recorded because both are
easy to derive wrongly: (a) `protocol/tun` computes
`enableGSO = stack == gvisor && mtu < 49152`, so this MTU turns GSO off there —
and then `StartStateStart` turns it back ON unconditionally because an
`adapter.FlowOutbound` exists in the config, so the ~1.98 MB of TCP/UDP GRO
scaffolding is present at BOTH MTUs and is priced by the flow-capable outbound,
not by this number; (b) the `mtu_fix` on the `shater_l3` fw4 zone is now inert —
only ICMP is ever marked into the device — and its uci-defaults comment still
says "the tunnel MTU is 1420".
**What is still NOT covered, said plainly.**
1. **A big ping does not start WORKING — it starts FAILING HONESTLY.** Upstream's
ICMP NAT is unfragmented-only in BOTH directions: `classifyReturn`
(`flow_dispatch.go:703-710`) returns `returnPass` on `parsed.fragment` exactly
as the forward path does. So a non-DF `ping -s 2000` now genuinely leaves the
router (fragmented to the tunnel MTU by `forwardToPort`), the far host really
answers, and the reply — fragmented by the peer to fit the tunnel — is not
NAT'd back to the LAN client. The operator sees a timeout. That is the
feature's promise ("travels or fails honestly"), not a capability claim.
Carrying oversized ICMP end to end would need reassembly upstream does not
have; it is not planned.
2. **A client that puts fragments on the wire ITSELF.** The device MTU cannot
un-fragment what already arrived fragmented, so such packets still reach the
gVisor stack, still get reassembled there, and still receive a forged reply
when the outbound is WireGuard/AWG. This is the residue the planned
`ip frag-off & 0x3fff != 0` prerouting carve-out (`netplane/nft.go`) is for.
**Whoever writes that rule must first check whether it can ever match:** fw4's
ruleset uses conntrack, conntrack pulls in `nf_defrag_ipv4`/`nf_defrag_ipv6`,
and defrag REASSEMBLES in PREROUTING before our marking rules run. Where
defrag is active the case does not arise (the MTU covers it) and the rule is
dead; where it is not, the rule is the only cover. Verify on the bench with
`nft list ruleset | grep -c ct` and a fragment counter, do not assume.
3. **The DF path changed hands and is untested on hardware.** It used to be the
kernel that answered `fragmentation needed` (from the router's LAN address,
MTU 1420); it is now the engine (from the far host's address, quoting
`Port.PortMTU()`). Both are correct PMTUD; only the first has ever run on a
real router.
**fw4 has to be told about the device, and `list device` is the only spelling
that works.** nftables runs EVERY table on every packet and a drop in any one of
them wins — an accept in `inet shater` cannot override fw4, and fw4 WILL reject
this forward: netifd never learns about a device the daemon creates at runtime,
so `shater-l3` belongs to no zone and falls into fw4's zone-less defaults. Hence
a real fw4 zone `shater_l3` + a lan→shater_l3 forwarding, seeded idempotently
(NAMED sections) and unconditionally in uci-defaults
(`openwrt/shater-core/files/etc/uci-defaults/30_shater-core`, `seed_l3_zone`),
with `mtu_fix` set. That `mtu_fix` is now inert and should be read as such: it
clamps forwarded TCP MSS to the route MTU, the device MTU is 65535, and nothing
but ICMP is ever marked into this device — the uci-defaults comment still says
"the tunnel MTU is 1420" and is stale. The device is attached
via `list device`, deliberately NOT `list network`: fw4 resolves a zone's
networks through netifd, which yields an EMPTY device set for a runtime-created
TUN (a proto-none stub would have to be brought UP to contribute an l3_device,
and nothing ever brings it up), while `list device` compiles to a plain
iifname/oifname string match — valid before the TUN exists, matching from the
moment shaterd creates it. `kmod-tun` joined DEPENDS so a slimmed image cannot
lose `/dev/net/tun` (`openwrt/shater-core/Makefile`). Our own forward chain
accepts both TUN legs ahead of the fail-closed drops — accepts that speak for
OUR table only (`netplane/nft.go`, forward chain step 4).
**What the policy still owns, and the one combination that now warns.** With the
ingress on, the mark is stamped in prerouting and the ROUTING decision carries
echo into the TUN before the forward chain — where the policy's verdicts live —
is ever consulted; that holds under every `untunnelable` value. The policy
therefore governs exactly two things: the never-markable protocols above, and
the fallback when the L3 rule/route did not come up (engine down, partial apply)
— `block` turns that failure into an honest loss, `direct` into a silent leak
with the real address. That is why `l3_tunnel` + `untunnelable=direct` draws a
validation warning naming the safe choice (`model/validate.go`), why every
rule/route failure surfaces as a named panel warning rather than an abort
(`addL3Routing`), and why the D17 HOLDING plane never marks: the TUN is created
BY the engine, and the holding plane exists precisely because the engine is not
running — marking would dead-end ping in a device that does not exist
(`netplane/nft.go`, hold comment).
- **Rejected: `auto_route` / letting the engine own the routing.** It rewrites
the main table and intercepts the router's own WAN/DNS/underlay traffic; the
blast radius of a toggle meant for LAN ping would be the whole router.
- **Rejected: marking all `l4proto != { tcp, udp }` into the TUN.** ESP/AH/GRE/
IGMP/SCTP cannot be parsed into flows upstream; they would vanish inside the
device. A drop with a name (the policy's) beats a silent black hole.
- **Rejected: keeping upstream's accept-and-forge for unroutable ICMP.** A ping
that "works" without leaving the router is the inverted lie this project keeps
deleting (D17's fiction purge, D23's dead WireGuard, D24's obedient-client
leak).
**Proven, and not proven, said plainly.** The cold start is PROVEN, not assumed:
`TestIntegrationL3TunInboundStarts`
(`shater/generate/l3_integration_linux_test.go`, run as root with NET_ADMIN and
`/dev/net/tun`, PASS) drives an `l3_tunnel=1` config through the SLIM registry
(`registry.Context`, not upstream's `include.Context`) under the shipped router
tag set: `box.New` + `Start` accept it, the kernel really ends up with the
`shater-l3` device at the contract MTU 65535 — the assertion the value exists
for, since a kernel that clamped it would silently restore the forged-reply
band — and Close removes it; precisely
the "built with X, verified with Y" gap class D23 exists for (a lost
`tun.RegisterInbound` or a trimmed `with_gvisor` changes no generated byte and
would otherwise surface only on the operator's router). Each layer contract is
pinned besides (`generate` `TestL3Tunnel*`, `netplane` `TestL3Ingress*`,
`route/prematch_icmp_lx_test.go`). Exactly two things remain UNVERIFIED:
(a) the end-to-end path on live hardware — LAN client → prerouting mark →
ip rule → TUN → WireGuard peer → reply back to the client — has not been
exercised on a real router; (b) the steady-state memory cost of the second
gVisor netstack (the `l3-in` TUN beside the WireGuard endpoint's own) is
unmeasured on the target hardware. An indicative figure exists and is only
that: on x86_64 in a container, idle and carrying no flows, peak RSS of a
process that brought the same engine up went from ~26.0-26.8 MB without
`l3_tunnel` to ~28.3-28.7 MB with it over three paired runs — about +2.2 MB.
That was measured on a throwaway harness, not on aarch64, not under load, and
with an empty ICMP NAT table, so it bounds nothing on the router. Neither
item is folded into any claim above.
Consequence: a ping from the LAN either genuinely travels through the tunnel
(WireGuard/AWG, direct) or fails honestly, at every size the router itself can
put into the device — and a router that never opts in renders the pre-feature
plane byte-for-byte (`TestL3IngressOptIn` pins the off-state render). Read
"fails honestly" strictly: above the tunnel MTU a non-DF ping now leaves the
router for real and then times out, because upstream's ICMP NAT does not carry
fragments back either. The one qualifier left is item 2 above — a client that
puts fragments on the wire ITSELF, on a router where conntrack defrag is not
reassembling them first. This paragraph has been overclaimed twice already;
extend it only against a bench result, never against a reading.
## D26 — What the engine cannot carry, the kernel carries: `untunnelable_egress`
Decided 2026-07-26. D25 ended with ESP/AH/GRE/IGMP/SCTP still owned by the D17
policy — that is, with a choice between dropping them and leaking them out the
WAN, never a data plane. This decision gives them one, and deliberately NOT
ours: **`globals.untunnelable_egress` (default empty;
`model.Globals.UntunnelableEgress`) names an existing egress of type
interface/tunnel, and LAN traffic that is neither TCP nor UDP is stamped in
prerouting with that egress's own mark, so the KERNEL routes it out that
egress's device with the kernel's own NAT.** No proxy, no engine, no userspace
stack ever touches the packet — which is exactly why every protocol works.
**Where the engine's boundary actually is — recorded so nobody digs for it
twice.** It is NOT the gVisor stack, and it is not WireGuard: on the forward
path the WG/AWG endpoint never consults gVisor at all. `Endpoint.WritePackets`
(`transport/wireguard/port.go:21-58`) takes the raw IP packet bytes, reads
exactly the IP version and the destination address, and hands
`device.InputPacketRef`s to `wgDevice.InputPackets` — the protocol byte is
never read; on the way back (`port.go:127-157`) `returnDeviceWrapper.Write`
offers every decrypted packet to `returnPath.ReturnPackets` first and only the
unconsumed remainder falls through to the gVisor device. gVisor serves
`DialContext`/`ListenPacket` — traffic the ENGINE originates — while forwarded
traffic bypasses the stack in both directions, indifferent to protocol. The
real ceiling sits one step earlier, in sing-tun's `ForwardDispatcher`:
`parseTransport` (`flow_parse.go:106-153`) sets `hasFlow` for exactly TCP, UDP,
ICMPv4 Echo/EchoReply and ICMPv6 EchoRequest/EchoReply — a packet of any other
protocol is never dispatched as a flow — and `createFlow`
(`flow_dispatch.go:325`) builds its NAT through
`allocateSelector(packet.protocol, …, packet.source.Port())` (line 355), which
needs a port-like selector that ESP/AH/GRE simply do not have (SCTP has ports,
but the parser above never grants it a flow either). Tailscale documents the
same frontier for its own userspace mode — "Any IP protocol other than TCP or
UDP (such as SCTP) is not supported in userspace mode… All IP protocols are
supported" in kernel mode
(https://tailscale.com/docs/reference/kernel-vs-userspace-routers) — useful as
external corroboration of where userspace data planes generally end, though OUR
boundary is the dispatcher, not the stack. The kernel egress was therefore
chosen not because userspace "cannot" in principle, but because the kernel
delivers all protocols with zero new code on the hot path.
**The mechanism already existed; the feature is one binding and one marking
step.** `addEgressRouting` (`netplane/apply.go`) has always installed, for
every interface/tunnel egress, an `ip rule fwmark <EgressMark> lookup
<EgressTable>` plus a `default dev <device>` route in that table — per-rule
egress selection rides on it. The only missing piece was that nothing ever
marked non-TCP/UDP traffic: `untunnelable=direct` merely ACCEPTED it in the
forward chain, so it left over the main table, i.e. the WAN.
`UntunnelableEgressBinding` (`netplane/nft.go`) resolves the option to the
egress's index, its OWN mark and its OWN device — deliberately no third
mark/table pair to keep coherent — and the prerouting chain stamps that mark on
the untunnelable protocols. A name that does not resolve to an interface/tunnel
egress with a device renders nothing and is reported: the D17 policy stays in
sole charge, which is the fail-closed reading of a typo.
**Why `l4proto != { tcp, udp }` is safe here when D25 banned it.** D25 rejected
the broad filter because the receiving side was the `ForwardDispatcher`, which
classifies nothing beyond TCP/UDP/ICMP echo — a marked ESP packet would enter
the TUN and vanish, a black hole wearing a tunnel's name. Here the receiving
side is the kernel, which forwards ANY IP protocol and NATs what it has
machinery for: SCTP carries ports and NATs like TCP/UDP; GRE is NATed only
through the PPTP helper keyed on the call-id — the kernel's own comment calls
GRE "generally not very suited for NAT, as it has no protocol-specific part as
port numbers" (`net/netfilter/nf_conntrack_proto_gre.c`); ESP/AH pass as plain
routed IP. Nothing on this path can silently swallow a protocol it does not
understand, which was the entire objection.
**Order against D25: the L3 ingress claims ICMP first.** With `l3_tunnel` on,
LAN ICMP is marked into the engine's TUN before the egress carrier is consulted
— the engine path routes ping by the operator's rules, which a kernel egress
cannot do — and only the remaining protocols go to the egress. With `l3_tunnel`
off, ICMP goes to the egress with everything else. In both shapes marked
traffic is settled by ROUTING before the forward chain speaks, so the D17
policy now governs exactly the failure case — the rule or route that did not
come up — the same division D25 already established for the L3 mark.
**What the feature refuses to promise — and the operator text refuses with it
(`shater/apply/warnings.go`, the egress-carrier note).** (a) It is not a tunnel
per se: the option accepts any interface/tunnel egress, and on the target
routers a WireGuard device is the exception (`kmod-wireguard` is usually
absent) while a second WAN is routine. Through a WireGuard egress this
genuinely is a tunnel; through a second WAN it is simply another uplink, and
the destination sees that uplink's real address. No text, comment or doc line
may call it a tunnel unconditionally. (b) It does not revive IPTV: IGMP is
LAN-side multicast group management, WireGuard is L3 point-to-point and carries
no multicast, and multicast never crossed this router under any setting —
routing IGMP out an egress restores nothing, and no wording may hint otherwise.
(c) IPsec through NAT-T never needed it: RFC 3948 encapsulates ESP in UDP/4500,
so a modern IPsec client behind NAT is ordinary UDP that already follows the
routing rules; the raw-ESP case this feature carries is the no-NAT-T remainder.
- **Rejected: teaching the engine these protocols.** Extending `parseTransport`
and the selector NAT upstream would be new hot-path code in an
actively-maintained adversarial area, for protocols the kernel already
forwards for free — and for ESP/AH/GRE there is no port-like selector to NAT
by in the first place.
- **Rejected 2026-07-26: carrying them through the userspace AWG endpoint
site-to-site, with no NAT at all.** This is the alternative the "no port-like
selector" line above does NOT dispose of, and it is written down because the
obvious reading of that line — "impossible" — is wrong and would be
re-derived. The endpoint is already protocol-blind in both directions
(`transport/wireguard/port.go:21-58`, `:127-157`), so an ESP packet could be
forwarded UNTOUCHED, keeping the LAN client's own source address, and the
reply would come back addressed to that client and need only be written to the
TUN. No selector, no NAT, every protocol. It needs two things we declined to
take on: lx-owned code in the forward hot path, bypassing `ForwardDispatcher`
on both legs — precisely the surface CONSTITUTION §2 exists to keep small on
an actively-maintained upstream — and a SERVER-side prerequisite (our LAN
prefix in the peer's `AllowedIPs`, plus a route back), which turns a router
option into a deployment contract. The kernel egress above buys the same
protocols with zero hot-path code, so this stays a design on file, not a gap.
- **Rejected: a dedicated mark/table pair for the carrier.** `addEgressRouting`
already binds `EgressMark`/`EgressTable` to the device; a third pair would be
a second copy of the same route that could drift from the first.
**Not verified, said plainly.** The end-to-end path — LAN client → prerouting
mark → ip rule → egress device → far end and back — has not been exercised with
real ESP or GRE on live hardware. Nothing above claims it has.
Consequence: raw IPsec, PPTP/GRE, SCTP — and ICMP when the L3 ingress is off —
leave through an egress the operator explicitly named, under kernel routing and
kernel NAT, instead of being dropped or silently leaking out the WAN; and with
the option empty (the default) the plane renders byte-for-byte as before, with
the D17 policy in sole charge.
+22
View File
@@ -12,6 +12,28 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Transparent proxying & routing
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** **L3 ingress for ICMP** (`globals.l3_tunnel`, opt-in, default off):
LAN ping travels THROUGH the tunnel instead of being dropped or answered by a
forged local reply. The engine opens a dedicated TUN (`shater-l3`, gVisor
stack, `auto_route` off); nft marks LAN icmp/icmpv6 only and a scoped
`ip rule` routes it in — the TPROXY plane and the main routing table stay
untouched (D25). Carried only by L3-capable egresses (WireGuard/AmneziaWG,
direct); ICMP routed to vless/vmess/… is honestly dropped, never faked.
Ceiling is upstream sing-tun's: ICMP echo only — Windows tracert works, IPv6
traceroute shows just the destination; ESP/AH/GRE/IGMP stay with the
`untunnelable` policy (D17) unless `untunnelable_egress` carries them (D26).
- **[MVP]** **Kernel egress for untunnelable protocols**
(`globals.untunnelable_egress`, opt-in, default empty): names an existing
interface/tunnel egress, and IPsec (ESP/AH), PPTP/GRE, SCTP — everything that
is neither TCP nor UDP, plus ICMP when the L3 ingress is off — is routed out
that egress's device by the KERNEL with kernel NAT, reusing the egress's own
fwmark/table from `addEgressRouting`; the proxy never sees a byte, which is
why every protocol works (D26). What that buys depends on the device: a
WireGuard interface really is a tunnel, a second WAN is just another uplink
whose real address the destination sees. It does not revive multicast IPTV,
and UDP-based VPNs (WireGuard, OpenVPN-UDP, IPsec NAT-T) never needed it —
they follow the routing rules as before. The `untunnelable` policy (D17)
keeps only the failure case: a route that did not come up.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
+9 -1
View File
@@ -40,6 +40,10 @@ define Package/shater-core
# shaterd : the daemon our init supervises (`shaterd run`)
# kmod-nft-tproxy : kernel TPROXY (shaterd emits the `inet shater` rules)
# kmod-nft-socket : socket match used by the tproxy divert chain
# kmod-tun : /dev/net/tun — the daemon opens the `shater-l3` TUN
# for L3 ingress (globals.l3_tunnel); usually built-in
# on stock images, but a slimmed image without it would
# make the option fail with a cryptic open() error.
# ip-full : `ip rule`/`ip route`/rt_tables for policy routing
# nftables-json : shaterd shells out to `nft`, and netplane/stats.go
# parses `nft -j list ...` — the JSON output only exists
@@ -49,7 +53,7 @@ define Package/shater-core
# ca-bundle : the daemon is CGO_ENABLED=0, so crypto/x509 has no
# host cert fallback — without /etc/ssl/certs every
# HTTPS subscription / .srs ruleset fetch fails.
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full +nftables-json +ca-bundle
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle
PKGARCH:=all
endef
@@ -80,6 +84,10 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/init.d
$(INSTALL_BIN) ./files/etc/init.d/shater $(1)/etc/init.d/shater
$(INSTALL_BIN) ./files/etc/init.d/shater-cron $(1)/etc/init.d/shater-cron
# START=21 one-shot that loads the persisted fail-closed plane before fw4's
# `lan -> wan ACCEPT` can be the only thing on the box (the main init is
# START=99, i.e. seconds of plaintext forwarding on every boot).
$(INSTALL_BIN) ./files/etc/init.d/shater-armor $(1)/etc/init.d/shater-armor
$(INSTALL_DIR) $(1)/etc/hotplug.d/iface
$(INSTALL_BIN) ./files/etc/hotplug.d/iface/99-shater $(1)/etc/hotplug.d/iface/99-shater
+259 -1
View File
@@ -33,6 +33,30 @@
# be running. `start` raises ACTIVE_FLAG, `stop` clears it; hotplug/cron
# reconcile ONLY while the flag is up, so an admin `stop` STICKS — no
# background actor may resurrect interception behind a stopped daemon.
# * BEING REPLACED IS NOT BEING SWITCHED OFF. `restart` and `reload` (which is
# stop+start, i.e. every LuCI Save & Apply) both run through `stop`, and the
# daemon's SIGTERM teardown removes the fail-closed table unconditionally — it
# does not consult kill_switch at all. Between that teardown and the
# successor's first apply the init GUARANTEES a gap: it waits for the old
# process to exit (shater_wait_stopped), then runs `shaterd migrate`, then
# starts a daemon that still has to build an engine. So a restart is announced
# with RESTART_FLAG, which tells the outgoing daemon to leave the fail-closed
# holding plane STANDING — apply.TeardownExiting swaps it in with one nft
# transaction and then skips the delete, so the table is never absent, not even
# for the 80-90 ms the old arm-after-teardown order measured. A real `stop`
# raises no flag and therefore still means what it says.
# (A package UPGRADE does not come through here at all on apk v3: shater-core's
# script table is post-install / pre-deinstall / post-upgrade, with no
# pre-upgrade, so default_prerm — and its `stop` — runs only on REMOVAL.)
# * The FAIL-CLOSED PLANE MUST ALSO EXIST BEFORE THIS SCRIPT DOES. START=99 is
# after fw4 (19) and netifd (20), so at every boot the LAN forwards to the WAN
# in the clear for as long as it takes procd to decompress the daemon off
# flash and get an engine up. /etc/init.d/shater-armor (START=21) loads
# BOOT_ARMOR — a copy of the holding plane the daemon persists on every apply
# — to close that window. This script owns the DISARM half, and it owns it
# with a CLOSED LIST: an operator's `stop`, or a removal, and nothing else.
# Powering the box down must not — `shutdown` reaches stop_service too, and it
# is not a person switching the product off (see shater_stop_disarms).
# * The engine must never be permanently abandoned while interception stands:
# respawn retries are infinite (procd never gives up); a sustained-dead
# daemon is additionally escalated by the shater-cron watchdog.
@@ -50,11 +74,166 @@ ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Raised around a restart/reload, read by the OUTGOING `shaterd run` at SIGTERM:
# present => "you are being replaced, leave the fail-closed plane standing";
# absent => "you are being switched off, take everything down". tmpfs, so a
# power cut can never make the next boot look like a restart.
RESTART_FLAG=/var/run/shater.restarting
# The persisted fail-closed holding plane. Written by the daemon on every apply,
# loaded by /etc/init.d/shater-armor at boot. Its PRESENCE is the arm token, so
# removing it here is how a deliberate stop stops the next boot from blocking.
BOOT_ARMOR=/etc/shater/boot.nft
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# WHICH ACTION rc.common was invoked with, frozen at source time.
#
# rc.common does, in this order:
# initscript=$1; action=${2:-help}; shift 2; ...; . "$initscript"; $action "$@"
# so `action` is ALREADY assigned when this file is sourced, and every action then
# runs as a function in THAT SAME shell. MEASURED on the target (ImmortalWrt
# 25.12.1 r37978) with a throwaway probe init script, not read off documentation:
#
# /etc/init.d/X restart -> stop_service action=[restart], start_service [restart]
# /etc/init.d/X stop -> stop_service action=[stop]
# /etc/init.d/X reload -> reload_service action=[reload]
# `reboot` -> stop_service action=[SHUTDOWN] <-- see below
# the boot after it -> start_service action=[boot]
#
# A previous probe reported this variable EMPTY and the emptiness was written up as
# the defect. It was the probe: `sh -x /etc/init.d/shater restart` bypasses the
# `#!/bin/sh /etc/rc.common` shebang, so rc.common never runs, never assigns
# `action`, and the variable reads empty no matter what this file does.
#
# Frozen into our own variable because `action` is a short, generic name that other
# framework helpers also use as a local; a snapshot taken before any function runs
# cannot be shadowed later.
SHATER_RC_ACTION="$action"
# --- what an action MEANS --------------------------------------------------
#
# THE BUG THESE TWO PREDICATES REPLACE (v0.2.17, measured on the live router).
# The old stop_service was `case $action in restart|reload) keep;; *) DISARM;; esac`
# — an open default that swept up every action nobody had enumerated. `reboot` is
# one of them: procd runs the K-links with the action `shutdown`, so the shutdown
# path deleted the arm token on the way down and the next boot had nothing to load.
# The mechanism destroyed itself at exactly the moment it exists for. Instrument
# reading from the router, one minute apart across a reboot:
#
# 13:28 /etc/shater/boot.nft present
# ---- reboot (stop_service action=[shutdown] -> old `*` branch -> rm)
# 18s at_S22: NO_TABLE armor_file=NO_FILE
#
# So both lists below are POSITIVE and CLOSED. An action nobody thought about —
# `shutdown` above all, but also whatever a future procd invents — falls through
# both and changes nothing. The default now fails in the recoverable direction: at
# worst a boot arms when it need not have, which costs the second before the daemon
# applies and is still gated by shater-armor's own four state refusals. The old
# default failed in the direction of the plaintext window the feature was built to
# close.
#
# They are predicates rather than an inline `case` so the test gate can execute the
# real thing: it sources THIS FILE in /bin/sh and calls them with every action procd
# actually uses (shater/cmd/shaterd/initscript_test.go). A comment claiming
# `shutdown` is handled is what shipped last time.
# True only for the ONE action that means "the operator switched the product off".
# Deliberately not `shutdown`: powering a router down is not turning a feature off.
#
# NOT sufficient on its own — see shater_stop_disarms. `stop` is also how the
# package manager's plumbing reaches us, and a package manager is not a person.
shater_action_disarms() {
case "$1" in
stop) return 0 ;;
*) return 1 ;;
esac
}
# Is a package manager in the middle of a transaction RIGHT NOW?
#
# This is a state, read at the moment the decision is made, exactly like
# shater-armor's four refusals — not a record of an event. The same question is
# already asked (for the same reason: prerm/postinst plumbing is not a user
# action) by the detached bring-up in /etc/uci-defaults/30_shater-core.
shater_pkg_transaction() {
pidof apk >/dev/null 2>&1 && return 0
pidof opkg >/dev/null 2>&1 && return 0
return 1
}
# Is the main service still enabled at boot? Same glob, and for the same reason,
# as shater-armor's own check: `/etc/init.d/shater enabled` would source procd.sh
# and take a blocking flock, which is not something to do from inside a package
# manager's transaction.
shater_rc_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
# THE ACTUAL DISARM DECISION.
# $1 = action
# $2 = 1 when a package transaction is in flight
# $3 = 1 when the service is still enabled in rc.d
# All three are passed in rather than read inside, so the gate can drive every
# combination without a package manager or an /etc/rc.d.
#
# WHY IT IS NOT JUST THE ACTION. base-files' default_prerm runs, in this order:
#
# if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
# "$i" stop
#
# so a package manager reaches stop_service wearing the operator's clothes. Two
# different intentions arrive as the same action, and the difference between them
# is readable at the moment of the decision:
#
# REMOVAL — prerm has ALREADY run `disable`, so S99shater is gone. The product
# is going away; the armor goes with it. (It is belt-and-braces even
# so: shater-armor refuses to arm without that symlink, and the whole
# init script is about to be deleted anyway.)
# REPLACED — the service is still enabled, so something intends to bring it
# back. That is not an operator switching anything off, and deleting
# the armor here would leave the next boot unprotected. "The next
# apply will rewrite it" is not an answer: the armor exists precisely
# to cover a reboot, and a reboot between an update and the first
# apply is how this product is deployed.
#
# MEASURED, because the paragraph above is about a path I got wrong once already.
# On THIS target (apk-tools 3.0.5, ImmortalWrt 25.12.1) shater-core's script table
# is post-install / pre-deinstall / post-upgrade, with NO pre-upgrade — so an apk
# UPGRADE never executes default_prerm and never calls `stop` at all. Verified with
# a real `apk fix --reinstall shater-core` while sampling the armor file: 245 625
# samples, zero disappearances, even with this guard mutated off. The upgrade half
# of this predicate is therefore defence-in-depth for a shape that is one
# `pre-upgrade` script (or a returning opkg lane) away, NOT a fix for an observed
# failure. The removal half is live today.
shater_stop_disarms() {
shater_action_disarms "$1" || return 1
# No package manager involved => a person typed it. The escape hatch must work.
[ "$2" = "1" ] || return 0
# A package transaction that has NOT disabled the service is replacing it.
[ "$3" = "1" ] && return 1
return 0
}
# True when a successor is coming, so the outgoing daemon should leave the
# fail-closed holding plane standing instead of removing it.
#
# `shutdown` is deliberately NOT a handoff either: nothing is coming, and the
# kernel that would hold the plane is going away with it. Leaving the flag down
# there also keeps the marker's meaning exact — it says "you are being replaced",
# and at shutdown nothing is.
shater_action_handoff() {
case "$1" in
restart|reload) return 0 ;;
*) return 1 ;;
esac
}
# --- helpers ---------------------------------------------------------------
# True only when the stack is explicitly enabled in UCI.
@@ -73,6 +252,25 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Announce/withdraw "this daemon is being replaced, not switched off". Read by
# `shaterd run` when it receives SIGTERM.
shater_mark_restart() {
mkdir -p "$(dirname "$RESTART_FLAG")" 2>/dev/null
: > "$RESTART_FLAG"
}
shater_clear_restart() { rm -f "$RESTART_FLAG"; }
# Remove the persisted boot armor, so the LAN is NOT blocked at the next boot
# before the daemon starts. Called from exactly two places, both of which are a
# statement about the PRODUCT rather than about this process: an operator typing
# `stop`, and a daemon binary that is no longer on the box. In neither case is
# anything going to come along and replace the armor with a real data plane, and a
# kill switch with nothing behind it is just a brick.
#
# NOT called on the shutdown path. That is the whole fix — see
# shater_action_disarms.
shater_disarm_boot() { rm -f "$BOOT_ARMOR"; }
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
@@ -136,9 +334,20 @@ start_service() {
# Guard: never claim to run without the daemon binary. A half-removed/failed
# shaterd upgrade must degrade to "plugin off", not to a box that thinks
# interception is live with nothing behind it.
#
# "Plugin off" now has to include DISARMING. With the boot armor in play, a
# missing binary is the one case where the fail-closed plane could stand
# forever with nothing able to replace it: the armor loads at START=21, the
# daemon never starts, and every later boot repeats it. The product being gone
# is not a security event — it is an uninstall — so the plane comes down and
# the LAN returns to plain routing, loudly.
if [ ! -x "$PROG" ]; then
shater_clear_restart
shater_disarm_boot
rm -f "$ACTIVE_FLAG"
nft delete table inet shater 2>/dev/null
_slog -p daemon.err \
"shaterd binary missing/not executable at $PROG — refusing to start (LAN stays on plain routing)"
"shaterd binary missing/not executable at $PROG — refusing to start; the fail-closed plane and its boot armor have been REMOVED (LAN back to plain routing, unprotected). Reinstall shaterd."
return 0
fi
@@ -150,6 +359,11 @@ start_service() {
# running, which is the boot case.
shater_wait_stopped
# The predecessor is gone and has already consumed the flag (it reads it in its
# SIGTERM handler). Withdraw it now, so a LATER `stop` is unambiguous even if
# this start fails further down.
shater_clear_restart
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
#
@@ -214,6 +428,45 @@ start_service() {
}
stop_service() {
# Say WHY we are stopping before procd sends the signal, because the daemon
# cannot tell from the signal alone and the answer changes what it leaves in
# the kernel. Two INDEPENDENT questions, and the old code conflated them into
# one two-armed `case` whose else-branch answered both wrongly for `shutdown`:
#
# 1. IS A SUCCESSOR COMING (this process only)? restart / reload.
# Raise RESTART_FLAG so the outgoing daemon replaces its data plane with
# the fail-closed HOLDING plane instead of removing it. The gap until the
# successor applies is not a moment: this script waits out the old
# process, runs `shaterd migrate`, then starts a daemon that must build an
# engine — all of it, before this flag existed, with `lan -> wan ACCEPT`
# and nothing else.
#
# 2. IS THE PRODUCT BEING SWITCHED OFF (across boots)? `stop` — and only
# `stop`, and only when a PERSON is behind it (shater_stop_disarms; the
# package manager reaches us through `stop` too). Then the boot armor goes
# with it, so the next boot does not quietly reinstate what the operator
# just switched off — the same rule ACTIVE_FLAG has always enforced for
# hotplug/cron.
#
# `shutdown` answers NO to both, which is the defect this replaced: a reboot is
# not a successor and it is certainly not an operator switching the product off.
# It is the boot the armor exists for. An upgrade answers NO to the second for
# the same kind of reason.
if shater_action_handoff "$SHATER_RC_ACTION"; then
shater_mark_restart
else
shater_clear_restart
fi
local in_pkg=0 rc_en=0
shater_pkg_transaction && in_pkg=1
shater_rc_enabled && rc_en=1
if shater_stop_disarms "$SHATER_RC_ACTION" "$in_pkg" "$rc_en"; then
shater_disarm_boot
elif [ "$in_pkg" = "1" ] && shater_action_disarms "$SHATER_RC_ACTION"; then
_slog -p daemon.info \
"stop came from a package transaction that left the service enabled — keeping the boot armor, so being replaced cannot leave the next boot unprotected"
fi
# Drop the live-flag FIRST so a concurrent hotplug/cron tick cannot rebuild
# what we are about to tear down. procd then sends SIGTERM to `shaterd run`,
# which runs its OWN honest teardown (engine.Close + netplane restore) — we
@@ -233,6 +486,11 @@ reload_service() {
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
#
# Marked EXPLICITLY as well as via SHATER_RC_ACTION: this is the path a routine
# Save & Apply takes, so it is the one that must not depend on reading an
# rc.common variable correctly. Belt and braces, one line.
shater_mark_restart
stop
start
}
@@ -0,0 +1,162 @@
#!/bin/sh /etc/rc.common
# /etc/init.d/shater-armor — the fail-closed plane, before the daemon exists.
#
# WHAT THIS CLOSES
#
# /etc/init.d/shater is START=99. By then fw4 (START=19) has long since loaded
# `lan -> wan ACCEPT` and netifd (START=20) has brought the LAN bridge up, so the
# router forwards LAN traffic to the WAN in the clear from the moment the link
# comes up until `shaterd run` has been decompressed off flash, has waited out any
# predecessor, has migrated UCI, has read the config and has installed its first
# table. On router-class hardware with a UPX-packed binary that is seconds — and
# they are exactly the seconds in which Wi-Fi finishes associating and every
# client on the network reconnects and starts talking. `kill_switch=closed` was
# configured the whole time and covered none of it.
#
# There was nothing in the package that could cover it either: no /etc/nftables.d
# include, no `nft -f` in uci-defaults. Protection existed only inside a Go
# process that had not started yet.
#
# HOW
#
# The daemon persists a copy of its fail-closed HOLDING plane (the same ruleset it
# installs when the engine is down: one forward chain, LAN-to-LAN and router
# traffic accepted, everything else from the diverted devices dropped) to
# $ARMOR on every apply. This script loads it early. When the daemon comes up it
# replaces the table atomically — the ruleset begins with `delete table` and adds
# its own in one netlink transaction — so there is never a moment with no table.
#
# `iifname` matches by NAME at packet time, not by ifindex at load time, so
# loading this before netifd has created br-lan is fine: the rules simply start
# matching when the device appears. That is why START can sit here rather than
# racing netifd.
#
# START=21: after fw4 (19) and netifd (20), because fw4's own start tears its
# table down and rebuilds it and we do not want to be in the middle of that, and
# because there is nothing to protect before the LAN device is being created. The
# residual exposure is the fraction of a second between netifd's `ifup` and this
# script, against seconds-to-a-minute before.
#
# THE ESCAPE HATCHES (a kill switch that cannot be switched off is a brick)
#
# These are STATE checks, evaluated here, at the moment of arming — not a record
# of something that happened on the way down. That distinction is the whole
# lesson of v0.2.17: the arm token was deleted by an EVENT on the shutdown path
# ("this looks like a stop"), and since `reboot` also runs the K-links, the
# mechanism reliably erased itself on the one transition it was built for. An
# event on the way down cannot be trusted to describe the world on the way up; a
# question asked on the way up can be.
#
# * $ARMOR only exists while the daemon's last applied config was BOTH enabled
# and fail-closed. `globals.enabled=0` and `kill_switch=open` each remove it
# at the next apply, and an operator typing `/etc/init.d/shater stop` removes
# it there and then. Powering the box off does NOT.
# * We refuse to arm when the main service is disabled in rc.d, or when the
# daemon binary is gone — in either case nothing would ever come along to
# replace the armor with a real data plane. These two are what makes a
# genuinely uninstalled/disabled product safe REGARDLESS of what the file
# says, which is why they are checked here rather than trusted to have been
# acted on earlier.
# * We refuse to arm when UCI can be read AND says the stack is disabled. A
# config that cannot be read is NOT a refusal: that case is precisely why the
# armor is a file rather than a query.
# * The chain hooks `forward` only, so SSH, LuCI and the admin panel (all input
# hook, to the router's own addresses) stay reachable. The operator can always
# get in and undo this.
#
# Note what a bare `/etc/init.d/shater stop` does NOT mean: it does not survive a
# reboot, because S99shater is still linked and procd starts the daemon again. So
# "stopped" is not a durable off-state and this script must not be designed as if
# it were — the durable ones are `disable` (no S??shater) and `globals.enabled=0`,
# and those are the two refusals above.
#
# busybox ash only — no bashisms.
START=21 # after firewall (19) and network (20), long before shater (99)
STOP=89
ARMOR=/etc/shater/boot.nft
PROG=/usr/bin/shaterd
# Syslog line that honors globals.log_syslog, like the other two inits. An
# unreadable UCI leaves the option empty => ON, which is what we want here: the
# one boot where the config cannot be read is the boot worth logging.
_slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater-armor "$@"
}
# Is the MAIN service enabled at boot? Answered by looking for its rc.d symlink
# rather than by running `/etc/init.d/shater enabled`: that is a USE_PROCD script,
# so every action of it sources procd.sh, which takes a blocking flock — and this
# runs at START=21, in the middle of boot, for a question a glob answers exactly
# as well. The START number is not hardcoded; any S<NN>shater counts.
shater_service_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
start() {
# No saved plane => the stack has never applied an enabled, fail-closed config
# (or it was explicitly switched off). Nothing to do, and nothing to say.
[ -f "$ARMOR" ] || return 0
[ -s "$ARMOR" ] || {
_slog -p daemon.err "$ARMOR is empty — NOT arming; the LAN is unprotected until shaterd starts"
return 0
}
# Never arm something nothing can disarm.
[ -x "$PROG" ] || {
_slog -p daemon.err \
"$PROG is missing — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
shater_service_enabled || {
_slog -p daemon.warn \
"the shater service is disabled in rc.d — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
# A READABLE config that says "off" wins over the saved plane (it means the
# daemon was stopped before it could disarm). An UNREADABLE config does not:
# that is the case this whole mechanism exists for.
en=$(uci -q get shater.globals.enabled 2>/dev/null)
if [ -n "$en" ] && [ "$en" != "1" ]; then
rm -f "$ARMOR"
_slog -p daemon.info "globals.enabled=$en — boot armor removed, not arming"
return 0
fi
command -v nft >/dev/null 2>&1 || {
_slog -p daemon.err "nft is not installed — cannot arm; the LAN is unprotected until shaterd starts"
return 0
}
# Validate before loading: a truncated/incompatible snapshot must not leave a
# half-built table behind on the one boot it is needed.
if ! nft -c -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.err \
"$ARMOR did not validate (nft -c) — NOT arming; the LAN is unprotected until shaterd starts"
return 0
fi
if nft -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.warn \
"fail-closed plane armed from $ARMOR: LAN->WAN forwarding is BLOCKED until shaterd applies. SSH, LuCI and the admin panel stay reachable."
else
_slog -p daemon.err \
"could not load $ARMOR — the LAN is unprotected until shaterd starts"
fi
return 0
}
stop() {
# Deliberately a NO-OP. By the time anything stops this service the daemon owns
# `inet shater`, and deleting the table here would dismantle a LIVE data plane
# on the strength of a service that only ever ran for one second at boot. The
# disarm paths that matter live where the decision is actually made:
# /etc/init.d/shater stop (operator switched it off) and the daemon itself
# (globals.enabled=0 / kill_switch=open).
return 0
}
@@ -59,6 +59,68 @@ if uci -q get shater.globals >/dev/null 2>&1 || [ -f /etc/config/shater ]; then
uci -q commit shater
fi
# Introduce the daemon-created `shater-l3` TUN to fw4 (L3 ingress, D-L3). The
# daemon policy-routes LAN ICMP into that device from OUR nft table
# `inet shater`, but nftables runs EVERY table on every packet and a drop in
# any one of them wins — an accept in `inet shater` cannot override fw4. And
# fw4 WILL drop this forward: netifd knows nothing about a device the daemon
# creates at runtime, so it belongs to no zone and falls into fw4's zone-less
# defaults (REJECT). The device has to be declared to fw4 itself; it cannot be
# fixed from our own table.
#
# Seeded UNCONDITIONALLY (not gated on globals.l3_tunnel): uci-defaults run
# once, so gating on the option would require re-running this script when the
# option is flipped later — which never happens. An idle zone is harmless: its
# device match is a plain iifname/oifname STRING compare that simply never hits
# while the TUN does not exist.
#
# Idempotency: `config zone`/`config forwarding` are normally ANONYMOUS
# sections, and a naive `uci add firewall zone` would append a duplicate on
# every re-run (uci-defaults re-run on package upgrade/reinstall). All sections
# here are NAMED instead, guarded by an existence check — a re-run re-finds the
# section and touches nothing.
seed_l3_zone() {
# No fw4 on this image (bare nftables build) => nothing drops the forward
# on fw4's behalf and there is nothing to punch through.
[ -f /etc/config/firewall ] || return 0
if ! uci -q get firewall.shater_l3 >/dev/null; then
uci set firewall.shater_l3=zone
uci set firewall.shater_l3.name='shater_l3'
uci set firewall.shater_l3.input='REJECT'
uci set firewall.shater_l3.output='ACCEPT'
uci set firewall.shater_l3.forward='REJECT'
uci set firewall.shater_l3.masq='0'
# INERT TODAY, kept for the day it is not. mtu_fix clamps forwarded TCP
# MSS to the route MTU — but shater-l3 is 65535 (deliberately: at any
# smaller value the kernel fragments into the device, and the flow
# dispatcher refuses to judge a fragment and lets the stack forge the
# echo reply — see l3MTU in shater/generate/inbound.go), so the clamp has
# nothing to clamp to. And only ICMP is ever marked into this device, so
# no TCP rides here to be clamped in the first place. It earns its keep
# the moment either of those changes; removing it would make that day
# silent.
uci set firewall.shater_l3.mtu_fix='1'
# `list device`, deliberately NOT the usual `list network`: fw4
# resolves a zone's networks through netifd, and netifd never learns
# about a device the daemon creates at runtime — a stub interface
# (proto none) would need to be brought UP to contribute an l3_device,
# and nothing ever brings it up, so `list network` resolves to an
# EMPTY device set and fw4 keeps dropping the forward. `list device`
# instead compiles to an iifname/oifname STRING match, valid before
# the TUN exists and matching from the moment shaterd creates it —
# no netifd involvement and no firewall reload at enable time. Do not
# "normalize" this to `list network` in a refactor; it breaks silently.
uci add_list firewall.shater_l3.device='shater-l3'
fi
if ! uci -q get firewall.shater_l3_fwd >/dev/null; then
uci set firewall.shater_l3_fwd=forwarding
uci set firewall.shater_l3_fwd.src='lan'
uci set firewall.shater_l3_fwd.dest='shater_l3'
fi
uci -q commit firewall
}
seed_l3_zone
# Bring the UCI schema forward on upgrade (idempotent; refuses a newer schema).
[ -x /usr/bin/shaterd ] && /usr/bin/shaterd migrate >/dev/null 2>&1
@@ -110,8 +172,27 @@ SHATER_BRINGUP='
done
[ -x /etc/init.d/shater ] && /etc/init.d/shater enable
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron enable
# The boot-time fail-closed armor. `enable` only — it is a one-shot that loads
# the persisted holding plane at START=21, and running it NOW would install a
# block on a live box moments before the daemon replaces it anyway. It has to
# be enabled here regardless of whether the stack is on: the file it loads only
# exists while the daemon wants it to, so an enabled-but-unarmed service is a
# no-op, and enabling it later would mean the first boot after an upgrade is
# the one boot still exposed.
[ -x /etc/init.d/shater-armor ] && /etc/init.d/shater-armor enable
[ -x /etc/init.d/shater ] && /etc/init.d/shater restart
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron restart
# Fold the seeded shater_l3 zone into the LIVE ruleset — matters on a live
# opkg/apk install only, where firewall started long before our commit and
# nothing else would re-read it until the next reboot. Gated on the fw4
# table actually being loaded: at FIRST boot this job can run before the
# S19 firewall start, and an early reload would install a ruleset built
# from a half-initialized netifd AND make the later start a no-op (fw4
# start skips when its table already exists). No table => the pending S19
# start reads the committed config by itself, no reload needed.
if nft list tables 2>/dev/null | grep -q "inet fw4"; then
[ -x /etc/init.d/firewall ] && /etc/init.d/firewall reload
fi
exit 0
'
SHATER_TMO=""
+66
View File
@@ -262,6 +262,65 @@
}
}
/* ---- fixture band (dev builds only; see App.tsx MockBanner) ----
Deliberately outside the crit/amber vocabulary: nothing is wrong with the
router, there is no router. The hazard hatch is the service-sticker language a
piece of network hardware already uses for "this unit is not in service". */
.mock-band {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
margin-top: calc(var(--u, 8px) * 2);
padding: 10px 14px;
border: 1px dashed var(--faint);
border-radius: 9px;
background: repeating-linear-gradient(
-45deg,
var(--sink),
var(--sink) 9px,
var(--panel) 9px,
var(--panel) 18px
);
}
.mock-band-tag {
flex-shrink: 0;
align-self: flex-start;
padding: 3px 7px;
border: 1px solid var(--faint);
border-radius: 4px;
background: var(--raised);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 700;
letter-spacing: 0.14em;
color: var(--dim);
}
.mock-band-copy {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 2px;
}
.mock-band-headline {
font-family: var(--font-mono);
font-size: 12.5px;
font-weight: 700;
letter-spacing: 0.02em;
color: var(--ink);
}
.mock-band-detail {
font-size: 12.5px;
line-height: 1.5;
color: var(--dim);
max-width: 76ch;
}
.mock-band-detail code {
font-family: var(--font-mono);
font-size: 11.5px;
color: var(--ink);
}
/* ---- commit-confirm band (every page except Apply, which has the full panel) ----
Same plate as the protection banner so the two read as one family; the seconds
are the loud element because they are the only thing that is running out. */
@@ -397,6 +456,13 @@
.finding--warning {
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
}
/* The daemon's "the list is capped" disclosure. Dashed, because the row is about
what ISN'T here — it must not read as one more finding to work through. */
.finding--truncated {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
background: var(--panel);
}
.finding-copy {
flex: 1;
min-width: 0;
+31 -1
View File
@@ -8,6 +8,7 @@ import { usePendingConfirm } from './pendingConfirm'
import { bootstrapSession } from './session'
import { ROUTES, navigate, useRoute } from './router'
import { engineState, protectionState } from './planeState'
import { truncationNote } from './findings'
import type { Route } from './router'
import { Overview, Placeholder, Nodes, Routing, Apply, DNS, Devices, Targets, Settings, Profiles, Insights, Networks } from './pages'
@@ -100,6 +101,7 @@ export function App() {
footer={<StatusBar status={status} />}
>
<Nav route={route} />
<MockBanner />
<PlaneBanner status={status} route={route} />
<ConfirmBand route={route} onChanged={() => void refreshStatus()} />
<Page route={route} status={status} onStatusChange={() => void refreshStatus()} />
@@ -170,6 +172,31 @@ function ConfirmBand({ route, onChanged }: { route: Route; onChanged: () => void
)
}
/**
* Says, on every page, that nothing on screen came from a router.
*
* Only a DEV build can ever render this — the fixtures are not in a production
* bundle (api.ts initMockBackend), so an operator cannot reach this state at all.
* It is here for the person who CAN: a footer line reading "DEMO DATA" is easy to
* work past for an afternoon and then screenshot into a bug report, and every
* number above it is invented.
*/
function MockBanner() {
if (!MOCK) return null
return (
<div className="mock-band" role="status">
<span className="mock-band-tag">FIXTURES</span>
<div className="mock-band-copy">
<span className="mock-band-headline">No router is being read</span>
<span className="mock-band-detail">
Every reading on this page is invented by <code>src/mock.ts</code> for offline
development. Drop <code>?mock</code> from the address to talk to a daemon.
</span>
</div>
</div>
)
}
/**
* The protection state, pinned under the nav on every page EXCEPT Overview
* (which shows the same state as its own headline readout — see planeState.ts).
@@ -195,12 +222,15 @@ function PlaneBanner({ status, route }: { status: Status | null; route: Route })
const criticals = (status.warnings ?? []).filter((w) => w.severity === 'critical').length
const state = protectionState(status)
// The published list is capped at 50, so with a note attached the count is a
// floor. Say "at least" rather than quoting a total the daemon didn't send.
const atLeast = truncationNote(status.warnings) ? 'At least ' : ''
// Wording comes from the shared source of truth so the banner and Overview can
// never describe the same router differently.
const headline = state.alarm
? state.headline
: `${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
: `${atLeast}${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
const detail = state.alarm
? state.detail
: 'Something you configured isn’t in effect. Review the findings before relying on it.'
+78 -27
View File
@@ -13,8 +13,12 @@
// serves in-memory fixtures instead of hitting the network, so `npm run dev`
// and screenshot runs render without a live backend. A real backend in dev is
// reachable instead via the Vite proxy in vite.config.ts (no flag ⇒ real fetch).
//
// THE FIXTURES ARE A DEV-BUILD-ONLY ARTEFACT — see initMockBackend below. They
// used to be a plain static import, decided at RUNTIME off `location.search`, so
// the invented router shipped inside the binary that goes on real hardware and a
// link ending in `?dev` painted a healthy appliance without making one request.
import * as mock from './mock'
import { armPendingConfirm, clearPendingConfirm, noteConfirmTimeout } from './pendingConfirm'
// --- error type -------------------------------------------------------------
@@ -1120,12 +1124,59 @@ export interface Model {
// --- transport --------------------------------------------------------------
/** True when the URL asks for the offline fixture backend (?mock or ?dev). */
export const MOCK: boolean = (() => {
// --- the offline fixture backend (dev builds only) ---------------------------
/**
* True when the in-memory fixtures are serving this session instead of the
* daemon. ALWAYS false in a production build — see {@link initMockBackend}.
*
* A live binding, not a constant: it is decided once during boot, before the
* first render, and every importer sees the same value for the whole session.
*/
export let MOCK = false
/** The loaded fixture module. `null` unless a dev build was asked for `?mock`. */
let fixtures: typeof import('./mock') | null = null
/**
* Load the fixture backend, if this build has one and the URL asks for it.
* Call ONCE from the entry point and await it before the first render — the
* pages read {@link MOCK} while they render, so flipping it afterwards would
* leave a half-mocked screen.
*
* Two gates, and the order matters. `import.meta.env.DEV` is folded to a literal
* `false` by Vite at build time, so in a production build the whole body is
* unreachable, `import('./mock')` is tree-shaken out of the module graph, and the
* fixtures are not in the emitted bundle AT ALL — not lazily, not behind a flag.
* `vite.config.ts` fails the build if that ever stops being true.
*
* This is deliberately stronger than "hide the mock behind a query flag". The
* flag was the bug: `?dev` on a production URL rendered an invented healthy
* router — 119 of 122 nodes alive, "Protected" — with no request made and one
* line of small print in the footer to say so. A person cannot audit a bundle;
* the only honest guarantee is that the invented data is not in it.
*/
export async function initMockBackend(): Promise<boolean> {
if (import.meta.env.DEV && mockRequested()) {
fixtures = await import('./mock')
MOCK = true
}
return MOCK
}
/** Does the URL ask for the offline fixture backend (`?mock` or `?dev`)? */
function mockRequested(): boolean {
if (typeof location === 'undefined') return false
const q = new URLSearchParams(location.search)
return q.has('mock') || q.has('dev')
})()
}
/** The fixture backend, for the `MOCK ? … : …` branches below. Throws rather
* than inventing data if it is ever reached without having been loaded. */
function mock(): NonNullable<typeof fixtures> {
if (!fixtures) throw new Error('mock backend not loaded — call initMockBackend() first')
return fixtures
}
/** A decoded response plus the raw Headers, for endpoints whose contract puts
* pagination metadata outside the JSON body (see the stats log endpoints). */
@@ -1174,11 +1225,11 @@ async function req<T>(path: string, init?: RequestInit): Promise<T> {
// --- endpoints --------------------------------------------------------------
export function getStatus(): Promise<Status> {
return MOCK ? mock.getStatus() : req<Status>('api/status')
return MOCK ? mock().getStatus() : req<Status>('api/status')
}
export async function getConfig(): Promise<Model> {
const m = await (MOCK ? mock.getConfig() : req<Model>('api/config'))
const m = await (MOCK ? mock().getConfig() : req<Model>('api/config'))
// Every page reads the config, and the commit-confirm window's length is the
// only thing needed to arm a countdown — so it is captured here once instead of
// being threaded through eight pages. See pendingConfirm.ts.
@@ -1188,7 +1239,7 @@ export async function getConfig(): Promise<Model> {
export function putConfig(m: Model): Promise<{ ok: boolean; applied: boolean }> {
return MOCK
? mock.putConfig(m)
? mock().putConfig(m)
: req('api/config', { method: 'PUT', body: JSON.stringify(m) })
}
@@ -1203,25 +1254,25 @@ export function putConfig(m: Model): Promise<{ ok: boolean; applied: boolean }>
* state. See pendingConfirm.ts.
*/
export async function apply(): Promise<ApplyResult> {
const r = await (MOCK ? mock.apply() : req<ApplyResult>('api/apply', { method: 'POST' }))
const r = await (MOCK ? mock().apply() : req<ApplyResult>('api/apply', { method: 'POST' }))
if (!r.error && r.changed) armPendingConfirm()
return r
}
export async function confirm(): Promise<ApplyResult> {
const r = await (MOCK ? mock.confirm() : req<ApplyResult>('api/confirm', { method: 'POST' }))
const r = await (MOCK ? mock().confirm() : req<ApplyResult>('api/confirm', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export async function rollback(): Promise<ApplyResult> {
const r = await (MOCK ? mock.rollback() : req<ApplyResult>('api/rollback', { method: 'POST' }))
const r = await (MOCK ? mock().rollback() : req<ApplyResult>('api/rollback', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export function getStats(): Promise<Stats> {
return MOCK ? mock.getStats() : req<Stats>('api/stats')
return MOCK ? mock().getStats() : req<Stats>('api/stats')
}
// --- daemon log download ------------------------------------------------------
@@ -1273,7 +1324,7 @@ function saveBlob(blob: Blob, filename: string): void {
*/
export async function downloadLog(range: LogRange): Promise<void> {
if (MOCK) {
saveBlob(new Blob([mock.getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
saveBlob(new Blob([mock().getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
return
}
let res: Response
@@ -1352,14 +1403,14 @@ function logPage<T>(env: { body: T[] | null; headers: Headers }): StatsLogPage<T
/** GET /api/stats/log — one page of the DNS query log with its cursor metadata. */
export function getStatsLogPage(q: StatsLogQuery = {}): Promise<StatsLogPage<QueryLogEntry>> {
return MOCK
? mock.getStatsLogPage(q)
? mock().getStatsLogPage(q)
: reqFull<QueryLogEntry[] | null>(`api/stats/log${statsLogQS(q)}`).then(logPage)
}
/** GET /api/stats/conns — one page of the connection log with its cursor metadata. */
export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<ConnLogEntry>> {
return MOCK
? mock.getStatsConnsPage(q)
? mock().getStatsConnsPage(q)
: reqFull<ConnLogEntry[] | null>(`api/stats/conns${statsLogQS(q)}`).then(logPage)
}
@@ -1367,13 +1418,13 @@ export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<C
* Rows only; callers that tail the stream want {@link getStatsLogPage} instead. */
export function getStatsLog(q: number | StatsLogQuery = {}): Promise<QueryLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
return MOCK ? mock().getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
}
/** GET /api/stats/conns — the live connection-event log (device→dest), newest first. */
export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
return MOCK ? mock().getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
@@ -1428,12 +1479,12 @@ export interface RulesReachability {
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
return MOCK ? mock().getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
return MOCK ? mock().getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
}
/**
@@ -1444,7 +1495,7 @@ export function getRulesetStatus(): Promise<RulesetStatus[]> {
*/
export function updateRuleset(tag: string): Promise<RulesetStatus | { ok: boolean }> {
return MOCK
? mock.updateRuleset(tag)
? mock().updateRuleset(tag)
: req('api/ruleset/update', { method: 'POST', body: JSON.stringify({ tag }) })
}
@@ -1472,7 +1523,7 @@ export interface RulesetCheck {
*/
export function checkRulesetCategory(source: string, category: string): Promise<RulesetCheck> {
return MOCK
? mock.checkRulesetCategory(source, category)
? mock().checkRulesetCategory(source, category)
: req<RulesetCheck>('api/ruleset/check', {
method: 'POST',
body: JSON.stringify({ source, category }),
@@ -1501,18 +1552,18 @@ export interface RulesetCategories {
*/
export function getRulesetCategories(source: string): Promise<RulesetCategories> {
return MOCK
? mock.getRulesetCategories(source)
? mock().getRulesetCategories(source)
: req<RulesetCategories>(`api/ruleset/categories?source=${encodeURIComponent(source)}`)
}
/** GET /api/devices — discovered LAN clients merged with per-device config. */
export function getDevices(): Promise<DiscoveredDevice[]> {
return MOCK ? mock.getDevices() : req<DiscoveredDevice[]>('api/devices')
return MOCK ? mock().getDevices() : req<DiscoveredDevice[]>('api/devices')
}
/** GET /api/interfaces — the router's UCI network interfaces for the egress picker. */
export function getInterfaces(): Promise<Interface[]> {
return MOCK ? mock.getInterfaces() : req<Interface[]>('api/interfaces')
return MOCK ? mock().getInterfaces() : req<Interface[]>('api/interfaces')
}
/** POST /api/session — exchange a single-use handoff token for a session cookie. */
@@ -1548,7 +1599,7 @@ export function importWg(conf: string): Promise<{ uri: string; name: string }> {
*/
export function updateSubscription(name: string): Promise<{ added: number }> {
return MOCK
? mock.updateSubscription(name)
? mock().updateSubscription(name)
: req('api/subscription/update', { method: 'POST', body: JSON.stringify({ name }) })
}
@@ -1568,7 +1619,7 @@ export function updateSubscription(name: string): Promise<{ added: number }> {
export function getGroupsHealth(
opts: { group?: string; members?: boolean } = {},
): Promise<GroupsHealth> {
if (MOCK) return mock.getGroupsHealth(opts)
if (MOCK) return mock().getGroupsHealth(opts)
const p = new URLSearchParams()
if (opts.group) p.set('group', opts.group)
if (opts.members) p.set('members', '1')
@@ -1683,11 +1734,11 @@ export interface GroupTestStart {
*/
export function postGroupsTest(name = ''): Promise<GroupTestStart> {
return MOCK
? mock.postGroupsTest(name)
? mock().postGroupsTest(name)
: req<GroupTestStart>('api/groups/test', { method: 'POST', body: JSON.stringify({ name }) })
}
/** GET /api/groups/test — progress + results of the current/last group test. */
export function getGroupsTest(): Promise<GroupTestStatus> {
return MOCK ? mock.getGroupsTest() : req<GroupTestStatus>('api/groups/test')
return MOCK ? mock().getGroupsTest() : req<GroupTestStatus>('api/groups/test')
}
+127
View File
@@ -0,0 +1,127 @@
// findings.ts — which apply-time finding is shown where.
//
// Run with `npm test` (node's built-in test runner + native TypeScript
// stripping; no test dependency is added to the SPA, which ships inside the
// daemon binary).
//
// Two defects are pinned here.
//
// 1. THE TRUNCATION NOTE WAS UNREACHABLE. The daemon caps Status.warnings at 50
// and overwrites the last slot with an `info` note counting what it dropped.
// Overview filtered `info` away wholesale, and the settings-page route keys on
// a section (`generate`) that no page owns — so the single line telling the
// operator "you are not seeing all of it" reached no screen at all.
//
// 2. FINDINGS ABOUT AN ENTITY NEVER REACHED THAT ENTITY'S PAGE. The generator
// drops a node it cannot build and names it; the Nodes page rendered that node
// as an ordinary row with a green toggle, because it never read the findings.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
attentionFindings,
entityFindings,
findingsByName,
sectionNotes,
truncationNote,
worstSeverity,
} from './findings.ts'
import type { StatusWarning } from './api.ts'
const crit = (section: string, name: string, message = 'broken'): StatusWarning => ({
severity: 'critical',
section,
name,
message,
})
const warn = (section: string, name: string, message = 'degraded'): StatusWarning => ({
severity: 'warning',
section,
name,
message,
})
const info = (section: string, name: string, message: string): StatusWarning => ({
severity: 'info',
section,
name,
message,
})
/** Verbatim from apply/warnings.go finalizeWarnings. */
const SUPPRESSED = info(
'generate',
'',
'7 further warning(s) suppressed; run `logread -e shater` for the full list',
)
// --- the truncation note ----------------------------------------------------
test('the truncation note is found, whatever else is in the list', () => {
const note = truncationNote([crit('rule', 'a'), warn('node', 'b'), SUPPRESSED])
assert.notEqual(note, null)
assert.match(note!.message, /7 further warning/)
})
test('a whole list has no truncation note', () => {
assert.equal(truncationNote([crit('rule', 'a'), warn('node', 'b')]), null)
assert.equal(truncationNote([]), null)
assert.equal(truncationNote(undefined), null)
})
test('an ordinary info note is not mistaken for the truncation note', () => {
const notes = [info('untunnelable', 'block', 'Ping and traceroute do not work…')]
assert.equal(truncationNote(notes), null)
})
test('the truncation note is kept out of the settings-page notes it would pollute', () => {
const all = [info('generate', '', 'cache: moved to /overlay'), SUPPRESSED]
const notes = sectionNotes(all, 'generate')
assert.equal(notes.length, 1)
assert.match(notes[0].message, /cache:/)
})
test('the attention list still carries only critical and warning', () => {
const all = [crit('rule', 'a'), warn('node', 'b'), info('untunnelable', 'block', 'x'), SUPPRESSED]
const attention = attentionFindings(all)
assert.equal(attention.length, 2)
assert.ok(attention.every((w) => w.severity !== 'info'))
})
// --- per-entity findings ----------------------------------------------------
test('a page takes only the sections it owns', () => {
const all = [
crit('node', 'tokyo-01', 'parse share-link: bad scheme (skipped)'),
warn('subscription', 'qomar', 'fetch failed'),
crit('rule', 'default', 'never applies'),
info('generate', '', 'cache: x'),
]
const mine = entityFindings(all, ['node', 'subscription'])
assert.deepEqual(
mine.map((w) => w.name),
['tokyo-01', 'qomar'],
)
})
test('entity findings never include info notes', () => {
const all = [info('node', 'tokyo-01', 'just a note'), SUPPRESSED]
assert.equal(entityFindings(all, ['node', 'generate']).length, 0)
})
test('findings index by name, and global (unnamed) ones are left out', () => {
const all = [
crit('node', 'tokyo-01', 'first'),
warn('node', 'tokyo-01', 'second'),
crit('node', '', 'global to the section'),
]
const byName = findingsByName(entityFindings(all, ['node']))
assert.equal(byName.size, 1)
assert.equal(byName.get('tokyo-01')!.length, 2)
})
test('one lamp per row takes the loudest severity', () => {
assert.equal(worstSeverity([warn('node', 'a'), crit('node', 'a')]), 'critical')
assert.equal(worstSeverity([warn('node', 'a')]), 'warning')
assert.equal(worstSeverity([]), null)
})
+91 -2
View File
@@ -6,7 +6,8 @@
//
// critical / warning — something needs attention: a protection promise is
// broken, or something you configured isn't in effect. These belong on
// Overview, where the operator looks first.
// Overview, where the operator looks first — and, when they name an entity,
// ALSO on the page that owns that entity (see `entityFindings`).
//
// info — a statement ABOUT the configuration, not a problem. It never clears,
// because nothing is wrong: it is simply describing a choice that was made.
@@ -16,9 +17,49 @@
// page that never goes away and never asks for anything trains people to skim
// the list — which is exactly how a real critical finding gets missed. Anything
// standing in the findings list should be something you could act on.
//
// The one exception is carved out below: the daemon's own note that it dropped
// findings to fit the cap. It is `info` by severity and unactionable by nature,
// and it is the single most important line in the list, because it is the list
// telling you it is not the whole list.
import type { StatusWarning } from './api'
/**
* The daemon's truncation disclosure, verbatim from apply/warnings.go
* finalizeWarnings:
*
* "%d further warning(s) suppressed; run `logread -e shater` for the full list"
*
* Matched on the stable clause rather than the whole sentence so a reworded tail
* still registers. If this ever stops matching, the failure mode is a list that
* silently claims to be complete — which is why `truncationNote` is tested.
*/
const SUPPRESSED_RE = /further warning\(s\) suppressed/
/**
* The daemon's "this list is incomplete" note, or null when the list is whole.
*
* Status.warnings is capped at 50, sorted critical-first, and the last slot is
* REPLACED by an `info` note counting what was dropped. That note therefore
* arrives on the one channel the panel filtered away wholesale: `info` never
* reached Overview, and the settings-page route (`sectionNotes`) keys on
* section `generate`, which no page owns. So the single line saying "there are
* findings you are not being shown" was the only one guaranteed to be invisible.
*
* Callers must render this WITH the attention list, not instead of it.
*/
export function truncationNote(warnings: StatusWarning[] | undefined): StatusWarning | null {
return (
(warnings ?? []).find((w) => w.severity === 'info' && SUPPRESSED_RE.test(w.message)) ?? null
)
}
/** Is this the truncation disclosure rather than an ordinary note? */
function isTruncationNote(w: StatusWarning): boolean {
return w.severity === 'info' && SUPPRESSED_RE.test(w.message)
}
/** Findings that need attention — the Overview list. Info notes are excluded. */
export function attentionFindings(warnings: StatusWarning[] | undefined): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'critical' || w.severity === 'warning')
@@ -29,10 +70,58 @@ export function attentionFindings(warnings: StatusWarning[] | undefined): Status
* (e.g. `untunnelable` → the Networks page's "Other traffic" section). Only info:
* a critical/warning is an attention item and stays on Overview, so it can't be
* quietly buried on a settings page instead.
*
* The truncation note is excluded: it is about the LIST, not about any section,
* and it has its own home beside the list ({@link truncationNote}).
*/
export function sectionNotes(
warnings: StatusWarning[] | undefined,
section: string,
): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'info' && w.section === section)
return (warnings ?? []).filter(
(w) => w.severity === 'info' && w.section === section && !isTruncationNote(w),
)
}
/**
* The attention findings about entities ONE page owns — for that page to show
* beside the entities themselves.
*
* Overview is where you look when you already suspect something; a page like
* Nodes is where you look when you don't. The generator drops a node it cannot
* build — an unparseable share link, a WireGuard key materialised twice — and
* says so by name ("node \"x\": parse share-link: … (skipped)"), yet that node
* kept rendering as an ordinary row with a green toggle, because the page never
* read the findings at all. The switch says on; the engine has no such outbound.
*
* This does NOT move anything off Overview: the same finding appears in both
* places, which is correct — one list is "what is wrong with this router", the
* other is "what is wrong with this node".
*/
export function entityFindings(
warnings: StatusWarning[] | undefined,
sections: readonly string[],
): StatusWarning[] {
const want = new Set(sections)
return attentionFindings(warnings).filter((w) => want.has(w.section))
}
/** Index attention findings by entity name, for badging a row directly. Entries
* with an empty `name` are global to their section and are left out. */
export function findingsByName(findings: StatusWarning[]): Map<string, StatusWarning[]> {
const out = new Map<string, StatusWarning[]>()
for (const f of findings) {
if (!f.name) continue
const list = out.get(f.name)
if (list) list.push(f)
else out.set(f.name, [f])
}
return out
}
/** The loudest severity in a set — for a row badge that has room for one lamp. */
export function worstSeverity(findings: StatusWarning[]): 'critical' | 'warning' | null {
if (findings.some((f) => f.severity === 'critical')) return 'critical'
if (findings.length > 0) return 'warning'
return null
}
+25 -9
View File
@@ -2,17 +2,33 @@ import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './tokens.css'
import { App } from './App'
import { initMockBackend } from './api'
import { ConfirmProvider } from './components'
const rootEl = document.getElementById('root')
if (!rootEl) throw new Error('#root not found')
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
// Settle the fixture question BEFORE the first render: pages read `MOCK` while
// they render, so a backend that arrives afterwards would paint half a screen
// from the daemon and half from fixtures. In a production build this resolves
// immediately and to `false` — the fixtures are not in the bundle to load (see
// api.ts initMockBackend and the assertNoMockFixtures plugin in vite.config.ts).
function mount() {
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl!).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
}
// A fixture module that fails to load is a broken dev checkout, not a reason to
// hand the operator a blank plate — mount anyway and let the shell report that it
// cannot reach a daemon, which by then is the truth.
void initMockBackend().then(mount, (e) => {
console.error('mock backend failed to load; continuing against the real API', e)
mount()
})
+52 -4
View File
@@ -500,7 +500,15 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
// the field case the readout used to call "Protected" (one
// rule, `default → direct`); `unknown` is a daemon too old to
// report. Default: tunnel.
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
// ?mock&plane=unreported → a daemon that sends NO `plane` field. The panel then
// knows nothing about what is installed, which is the state
// the Kill-switch module used to render as a green "ARMED"
// (`undefined !== 'none'` is true).
function mockPlane(): {
plane: 'full' | 'hold' | 'none' | undefined
engine: boolean
killSwitch: string
} {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
const killSwitch = params.get('ks') === 'open' ? 'open' : 'closed'
@@ -508,6 +516,7 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
if (p === 'hold') return { plane: 'hold', engine: false, killSwitch: 'closed' }
if (p === 'none') return { plane: 'none', engine: false, killSwitch }
if (p === 'open') return { plane: 'none', engine: false, killSwitch: 'open' }
if (p === 'unreported') return { plane: undefined, engine: true, killSwitch }
return { plane: 'full', engine: true, killSwitch }
}
@@ -515,7 +524,7 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
// meaningful with the plane installed: with the engine down there is no running
// config to judge, and the daemon reports the unknown/zero value — so do the same
// here rather than leaving a stale "tunnel" behind a dead engine.
function mockTraffic(plane: 'full' | 'hold' | 'none'): Traffic | undefined {
function mockTraffic(plane: 'full' | 'hold' | 'none' | undefined): Traffic | undefined {
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
switch (params.get('traffic')) {
@@ -560,6 +569,29 @@ const MOCK_WARNINGS: StatusWarning[] = [
name: 'fakeip-pool',
message: 'fake-IP resolver cannot be used as a fallback; the failover chain was not built',
},
// Two findings the generator attributes to a NODE by name — the class that the
// Nodes page never showed, leaving a node the engine threw away rendered as an
// ordinary row with a green toggle. Both name real fixture nodes so the row
// badge, the collapsed-bucket "N flagged" count and the per-row strip all fire.
{
severity: 'warning',
section: 'node',
name: 'fi-trojan',
message: 'parse share-link: unsupported scheme "trojan+ws" (skipped)',
},
{
severity: 'warning',
section: 'node',
name: 'home-wg',
message:
'this WireGuard node is materialised twice in the engine config — as "home-wg" and as "group-stealth-m1-home-wg" — and traffic can reach both. A WireGuard peer keeps ONE session per public key, so two devices built from one private key evict each other continuously and NEITHER tunnel passes traffic. Only "home-wg" is kept; everything that routed through "group-stealth-m1-home-wg" is fail-closed (blocked) instead of leaving over the plain WAN',
},
{
severity: 'warning',
section: 'subscription',
name: 'backup',
message: 'fetch failed: dial tcp 203.0.113.9:443: i/o timeout — serving the nodes cached earlier',
},
{
severity: 'info',
section: 'generate',
@@ -568,6 +600,19 @@ const MOCK_WARNINGS: StatusWarning[] = [
},
]
/**
* The daemon's truncation disclosure, exactly as apply/warnings.go writes it when
* the published set overflows the 50-entry cap. Served under `?mock&trunc` so the
* "this list is incomplete" rendering is exercisable — it used to be dropped
* wholesale by the panel's `info` filter and reached no screen at all.
*/
const MOCK_TRUNCATION: StatusWarning = {
severity: 'info',
section: 'generate',
name: '',
message: '7 further warning(s) suppressed; run `logread -e shater` for the full list',
}
/**
* The standing `untunnelable` note the daemon reports. It is INFO, never a
* problem: it states a correct, chosen configuration. Two shapes, mirroring the
@@ -614,9 +659,12 @@ function mockWarnings(killSwitch: string): StatusWarning[] {
const params = new URLSearchParams(q)
const mode = (CONFIG.Globals as { Untunnelable?: string }).Untunnelable ?? 'block'
const notes = untunnelableNote(mode, killSwitch)
// `?trunc` adds the daemon's "the published list is capped" disclosure, which
// it appends IN PLACE OF the last entry it had room for.
const trunc = params.has('trunc') ? [{ ...MOCK_TRUNCATION }] : []
// A degraded plane always comes with the findings that explain it.
if (params.has('warn') || params.get('plane')) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes]
if (params.has('warn') || params.get('plane') || trunc.length > 0) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes, ...trunc]
}
return notes
}
+443
View File
@@ -0,0 +1,443 @@
/* Alerts section (rendered on Settings) — inherits the Faceplate tokens and the
* shared page chrome from App.css (.toast, .mono). Every rule below is a
* one-to-one copy of the DNS.css rule the markup used before the section moved
* here, renamed `dns-*` → `alr-*` so nothing collides. Orange stays an accent. */
/* ---- section shell (matches the Settings group plates one-to-one) ---- */
.alr-section {
margin-top: calc(var(--u, 8px) * 3.5);
}
.alr-sec-hd {
display: flex;
align-items: baseline;
gap: 12px;
padding-bottom: 10px;
border-bottom: 1px solid var(--groove);
}
.alr-sec-title {
margin: 0;
font-family: var(--font-mono);
font-size: 13px;
font-weight: 700;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--dim);
}
.alr-sec-count {
font-size: 11px;
letter-spacing: 0.06em;
color: var(--faint);
}
.alr-sec-note {
margin: 10px 2px 0;
font-family: var(--font-sans);
font-size: 12.5px;
line-height: 1.55;
color: var(--dim);
max-width: 56ch;
}
/* ---- add form ---- */
.alr-add {
display: flex;
flex-direction: column;
gap: 10px;
margin-top: calc(var(--u, 8px) * 2);
}
.alr-add-top {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px;
}
.alr-input {
min-width: 0;
padding: 9px 12px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 12.5px;
letter-spacing: 0.02em;
box-shadow: 0 1px 2px var(--shadow) inset;
transition: border-color 0.15s, box-shadow 0.15s;
}
.alr-input::placeholder {
color: var(--faint);
}
.alr-input:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-input:disabled {
opacity: 0.55;
}
.alr-input--name {
flex: 0 1 14rem;
}
/* segmented type picker */
.alr-seg {
display: inline-flex;
border: 1px solid var(--groove);
border-radius: 7px;
overflow: hidden;
background: var(--sink);
}
.alr-seg-btn {
padding: 8px 14px;
border: 0;
background: transparent;
color: var(--dim);
font-family: var(--font-mono);
font-size: 11px;
letter-spacing: 0.08em;
text-transform: uppercase;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.alr-seg-btn + .alr-seg-btn {
border-left: 1px solid var(--groove);
}
.alr-seg-btn.on {
background: var(--accent);
color: #fff;
}
.alr-seg-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
.alr-resp {
display: inline-flex;
align-items: center;
gap: 8px;
}
.alr-resp-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-select {
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 11.5px;
letter-spacing: 0.04em;
cursor: pointer;
}
.alr-select:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-add-actions {
display: flex;
align-items: center;
justify-content: flex-end;
gap: 14px;
flex-wrap: wrap;
}
.alr-field-err {
flex: 1;
min-width: 0;
margin: 0;
font-family: var(--font-mono);
font-size: 11.5px;
line-height: 1.5;
color: var(--crit);
}
/* ---- rows ---- */
.alr-rows {
list-style: none;
margin: calc(var(--u, 8px) * 2) 0 0;
padding: 0;
display: flex;
flex-direction: column;
gap: 8px;
}
.alr-row {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
padding: 12px 14px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(
180deg,
var(--raised),
color-mix(in srgb, var(--raised) 82%, var(--panel))
);
box-shadow: 0 1px 0 var(--edge) inset;
}
.alr-row-main {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 4px;
}
.alr-row-l1 {
display: flex;
align-items: center;
gap: 8px;
flex-wrap: wrap;
}
.alr-row-name {
font-family: var(--font-mono);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
color: var(--ink);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 24ch;
}
.alr-row-l2 {
display: flex;
align-items: center;
gap: 10px;
flex-wrap: wrap;
font-size: 11.5px;
letter-spacing: 0.02em;
}
.alr-row-detail {
color: var(--dim);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 40ch;
}
/* badge — groove-bordered, not orange (accent stays reserved) */
.alr-badge {
display: inline-block;
padding: 2px 7px;
border: 1px solid var(--groove);
border-radius: 5px;
background: color-mix(in srgb, var(--sink) 60%, transparent);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 600;
letter-spacing: 0.1em;
text-transform: uppercase;
color: var(--dim);
white-space: nowrap;
}
.alr-badge--accent {
border-color: color-mix(in srgb, var(--accent) 55%, var(--groove));
color: var(--accent);
}
.alr-masked {
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: 0.08em;
color: var(--faint);
text-transform: uppercase;
cursor: help;
}
.alr-del {
flex: none;
padding: 6px 12px;
font-size: 10.5px;
}
/* ---- empty plate ---- */
.alr-empty {
margin-top: calc(var(--u, 8px) * 2);
padding: calc(var(--u, 8px) * 3);
border: 1px dashed var(--groove);
border-radius: 9px;
background: color-mix(in srgb, var(--raised) 55%, transparent);
text-align: center;
}
.alr-empty-title {
display: block;
font-size: 13px;
font-weight: 700;
letter-spacing: 0.06em;
color: var(--dim);
}
.alr-empty-body {
margin: 8px auto 0;
max-width: 48ch;
font-family: var(--font-sans);
font-size: 13px;
line-height: 1.55;
color: var(--dim);
}
/* ---- loading skeleton ---- */
.alr-skel {
height: 62px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(90deg, var(--raised), var(--sink), var(--raised));
background-size: 200% 100%;
animation: alr-skel-shift 1.4s ease-in-out infinite;
}
@keyframes alr-skel-shift {
from {
background-position: 200% 0;
}
to {
background-position: -200% 0;
}
}
/* the per-row delivery picker sits inline in the row */
.alr-detour {
flex: none;
display: flex;
flex-direction: column;
gap: 5px;
min-width: 0;
}
.alr-detour-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-detour-select {
max-width: 22rem;
}
/* current delivery-path readout on the row */
.alr-path {
color: var(--faint);
white-space: nowrap;
}
.alr-path[data-active='on'] {
color: var(--dim);
}
.alr-path-name {
color: var(--led-on);
font-weight: 600;
}
.alr-path[data-missing='y'] .alr-path-name {
color: var(--amber);
}
.alr-path-flag {
color: var(--amber);
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.alr-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.alr-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.alr-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-row delivery controls sit inline in the row (like .alr-detour) */
.alr-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.alr-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.alr-note--row {
margin-top: 2px;
}
/* alert event checkboxes */
.alr-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.alr-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.alr-event input {
accent-color: var(--fp-accent, currentColor);
}
/* ---- responsive ---- */
@media (max-width: 640px) {
.alr-row {
flex-wrap: wrap;
}
.alr-row-main {
flex-basis: calc(100% - 90px);
}
.alr-del {
margin-left: auto;
}
.alr-input--name {
flex-basis: 100%;
}
.alr-detour {
flex-basis: 100%;
order: 3;
flex-wrap: wrap;
}
/* A <select> won't shrink below its widest option unless it's allowed to:
without min-width:0 the long detour labels push the page into a horizontal
scroll at 390px. Let them fill the row and clip instead. */
.alr-detour-select,
.alr-resp .alr-select {
max-width: 100%;
width: 100%;
min-width: 0;
}
.alr-resp {
display: flex;
flex-wrap: wrap;
max-width: 100%;
}
.alr-ctl {
flex-basis: 100%;
order: 3;
}
}
@media (prefers-reduced-motion: reduce) {
.alr-skel {
animation: none;
}
.alr-input,
.alr-seg-btn {
transition: none;
}
}
+690
View File
@@ -0,0 +1,690 @@
import './Alerts.css'
import { useCallback, useMemo, useState } from 'react'
import { Button, Toggle, useConfirm } from '../components'
import type { Alert, Model } from '../api'
// The Alerts section — out-of-band notifications (Telegram bot / webhook) for
// kill-switch trips, apply failures, new devices and subscription expiry. It
// lived at the bottom of the DNS page, which is the last place an operator
// looking for "tell me when the tunnel dies" would think to look; it now renders
// as a group on Settings. The component owns no I/O: every mutation goes through
// the `onSave` prop so Settings keeps a single dirty banner and a single toast.
//
// NOTE on duplication: the detour helpers below (DetourCatalog, canonDetour,
// detourValues, describeDetour, DetourSelect) plus asArray / uniqueName /
// maskUrl / EmptyPlate are deliberate copies of the ones in DNS.tsx. DNS keeps
// its own for resolvers and DNS rules; extracting a shared module would couple
// two pages that otherwise share nothing, and that refactor is out of scope
// here. If a third consumer ever appears, promote them then.
// ---- local Model extension --------------------------------------------------
/** The Model with the Alerts slice surfaced (index-signature passthrough). */
type AlertsModel = Model & { Alerts?: Alert[] | null }
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers (copies of DNS.tsx — see the header note) ----------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
const HTTP_RE = /^https?:\/\//i
/** A remote URL often carries a token in its query/path — show host only. */
function maskUrl(url: string): { host: string; masked: boolean } {
try {
const u = new URL(url)
return { host: u.host, masked: u.search !== '' || u.pathname.replace(/\/+$/, '') !== '' }
} catch {
return { host: url || '—', masked: false }
}
}
function uniqueName(base: string, taken: Set<string>): string {
const seed = base.trim() || 'alert'
if (!taken.has(seed)) return seed
let i = 2
while (taken.has(`${seed}-${i}`)) i++
return `${seed}-${i}`
}
/** The live targets an alert's delivery can be pinned to (the picker). */
interface DetourCatalog {
groups: string[]
chains: string[]
egresses: { name: string; type: string }[]
nodes: string[]
}
/**
* Normalise a stored `Via` to a picker option value. Empty/`direct` ⇒
* `direct`; already-prefixed values (`group:`/`chain:`/`egress:`/`node:`) pass
* through; a bare legacy name is resolved against the catalog so a still-valid
* setup isn't mislabelled; anything unresolved is kept verbatim (shown stale).
*/
function canonDetour(raw: string | undefined, cat: DetourCatalog): string {
const d = (raw ?? '').trim()
if (!d || d.toLowerCase() === 'direct') return 'direct'
if (/^(node|group|chain|egress):/i.test(d)) return d
if (cat.egresses.some((e) => e.name === d)) return `egress:${d}`
if (cat.groups.includes(d)) return `group:${d}`
if (cat.chains.includes(d)) return `chain:${d}`
if (cat.nodes.includes(d)) return `node:${d}`
return d
}
/** Every valid option value for a catalog, including `direct`. */
function detourValues(cat: DetourCatalog): Set<string> {
const s = new Set<string>(['direct'])
for (const g of cat.groups) s.add(`group:${g}`)
for (const c of cat.chains) s.add(`chain:${c}`)
for (const e of cat.egresses) s.add(`egress:${e.name}`)
for (const n of cat.nodes) s.add(`node:${n}`)
return s
}
/** Describe a canonical detour value for the row readout. */
function describeDetour(
canon: string,
cat: DetourCatalog,
valid: Set<string>,
): { direct: boolean; prefix: string; name: string; missing: boolean } {
if (canon === 'direct') return { direct: true, prefix: '', name: '', missing: false }
const i = canon.indexOf(':')
const kind = i === -1 ? '' : canon.slice(0, i)
const name = i === -1 ? canon : canon.slice(i + 1)
const missing = !valid.has(canon)
let prefix = 'via'
if (kind === 'group') prefix = 'via group'
else if (kind === 'chain') prefix = 'via chain'
else if (kind === 'node') prefix = 'via node'
else if (kind === 'egress') {
const eg = cat.egresses.find((e) => e.name === name)
prefix = eg?.type === 'interface' ? 'via interface' : 'via egress'
}
return { direct: false, prefix, name, missing }
}
// ---- section ----------------------------------------------------------------
export function AlertsSection({
config,
busy,
loading,
onSave,
}: {
/** Full desired-state model; null until it has loaded. */
config: Model | null
/** A save/apply is in flight — controls lock. */
busy: boolean
/** The config is still loading — show a skeleton row. */
loading: boolean
/** Persist the whole next model; resolves true on success (Settings' `save`). */
onSave: (next: Model, okMsg: string) => Promise<boolean>
}): JSX.Element {
const confirm = useConfirm()
const model = config as AlertsModel | null
const alerts = useMemo<Alert[]>(() => asArray(model?.Alerts), [model])
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Built straight from the Model with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({
groups: asArray(config?.Groups).map((g) => g.Name),
chains: [],
egresses: asArray(config?.Egresses).map((e) => ({ name: e.Name, type: e.Type })),
nodes: asArray(config?.Nodes).map((n) => n.Name),
}),
[config],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
// ---- mutations — all writes go through onSave -----------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!model) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return onSave({ ...model, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[model, alerts, onSave],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void onSave({ ...model, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[model, alerts, onSave],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!model) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void onSave({ ...model, Alerts: next }, `Deleted ${target.Name}`)
},
[model, alerts, onSave, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!model) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void onSave(
{ ...model, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[model, alerts, onSave],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void onSave(
{ ...model, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[model, alerts, onSave],
)
return (
<div className="alr-section" aria-label="Alerts">
<header className="alr-sec-hd">
<h2 className="alr-sec-title">Alerts</h2>
<span className="alr-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="alr-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="alr-rows" aria-hidden="true">
<li className="alr-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="alr-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
)
}
// ---- alert add form + row ----------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="alr-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="alr-add-top">
<input
className="alr-input alr-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="alr-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="alr-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="alr-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="alr-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="alr-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="alr-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="alr-delivery">
<label className="alr-resp">
<span className="alr-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="alr-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="alr-add-actions">
{err && (
<p className="alr-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="alr-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="alr-row-main">
<div className="alr-row-l1">
<span className="alr-row-name">{alert.Name}</span>
<span className="alr-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="alr-badge alr-badge--accent">
{e}
</span>
))}
</div>
<div className="alr-row-l2 mono">
<span className="alr-row-detail">{detail.text}</span>
{detail.masked && (
<span className="alr-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="alr-path">direct</span>
) : (
<span className="alr-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="alr-path-name">{route.name}</strong>
{route.missing && <span className="alr-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="alr-note alr-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="alr-ctl">
<label className="alr-detour">
<span className="alr-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="alr-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
/** The live delivery picker: option list built from the Model's targets. */
function DetourSelect({
value,
catalog,
valid,
busy,
disabled,
ariaLabel,
onChange,
directLabel = 'Direct (no proxy)',
}: {
value: string // canonical value
catalog: DetourCatalog
valid: Set<string>
busy: boolean
disabled: boolean
ariaLabel: string
onChange: (v: string) => void
directLabel?: string
}) {
const missing = value !== 'direct' && !valid.has(value)
return (
<select
className="alr-select alr-detour-select"
value={value}
onChange={(e) => onChange(e.target.value)}
disabled={busy || disabled}
aria-label={ariaLabel}
>
<option value="direct">{directLabel}</option>
{catalog.groups.length > 0 && (
<optgroup label="Groups">
{catalog.groups.map((g) => (
<option key={g} value={`group:${g}`}>
Group {g} (balancer)
</option>
))}
</optgroup>
)}
{catalog.chains.length > 0 && (
<optgroup label="Chains">
{catalog.chains.map((c) => (
<option key={c} value={`chain:${c}`}>
Chain {c}
</option>
))}
</optgroup>
)}
{catalog.egresses.length > 0 && (
<optgroup label="Interfaces / egresses">
{catalog.egresses.map((e) => (
<option key={e.name} value={`egress:${e.name}`}>
Interface/egress {e.name}
{e.type ? ` (${e.type})` : ''}
</option>
))}
</optgroup>
)}
{catalog.nodes.length > 0 && (
<optgroup label="Nodes">
{catalog.nodes.map((n) => (
<option key={n} value={`node:${n}`}>
Node {n}
</option>
))}
</optgroup>
)}
{missing && <option value={value}>{value} (missing)</option>}
</select>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="alr-empty">
<span className="alr-empty-title mono">{title}</span>
<p className="alr-empty-body">{body}</p>
</div>
)
}
+76 -32
View File
@@ -11,7 +11,7 @@ import {
ApiError,
} from '../api'
import type { Globals, Status } from '../api'
import { engineReadout } from '../planeState'
import { engineReadout, killSwitchReadout } from '../planeState'
import { onPendingConfirmExpire, usePendingConfirm } from '../pendingConfirm'
// Short, readable config hash — drops the "sha256:" prefix like the footer does.
@@ -103,30 +103,64 @@ export default function Apply() {
void loadConfig()
}, [loadConfig])
// The window running out is the daemon reverting on its own — observe it and
// say so. The countdown itself ticks inside usePendingConfirm; this only reacts
// to the end of it, and the store makes sure that fires exactly once even with
// the app-wide band mounted alongside.
// The window running out does NOT mean the daemon rolled back.
//
// apply.ArmRollback captures the data-plane generation when it arms, and on
// expiry it compares. If anything re-applied the plane in between — another
// panel apply, SIGHUP, a hotplug or the once-a-minute cron reconcile, the WAN
// profile auto-switch — it disarms and KEEPS the running config, logging "NOT
// rolling back" and nothing else. That is the common case on a production
// router, and this page used to print "daemon auto-rolled back to last-good
// config" for it: a confident report of an event that did not happen, with a
// hash pair underneath that quietly said "unchanged".
//
// The panel cannot see which branch ran — the daemon says so only in its log.
// So it reports the one thing it CAN observe, the live config hash, and waits
// for the revert to land before reading it (a rollback is a full re-apply and
// does not complete the instant the timer fires).
const liveHashRef = useRef('')
liveHashRef.current = status?.hash ?? ''
useEffect(
() =>
onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Auto-rolled back')
void (async () => {
const after = (await refreshStatus())?.hash ?? ''
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed — daemon auto-rolled back to last-good config.',
before,
after,
})
})()
}),
[flash, refreshStatus],
)
useEffect(() => {
let cancelled = false
const off = onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Confirm window elapsed')
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed. Reading what the daemon did…',
before,
after: before,
})
void (async () => {
let after = before
for (let i = 0; i < 4 && !cancelled; i++) {
await new Promise((r) => window.setTimeout(r, 1500))
if (cancelled) return
after = (await refreshStatus())?.hash ?? after
if (after !== before) break
}
if (cancelled) return
setResult({
kind: 'expire',
tone: 'warn',
text:
after !== before
? 'Confirm window elapsed and the live config changed — the daemon reverted to its last-good config.'
: 'Confirm window elapsed and the live config has not changed, so this config is still running. ' +
'The daemon only reverts if nothing else re-applied the data plane while the window was open; ' +
'otherwise it stands down and keeps what is live. Which one happened is in the daemon log — ' +
'download it from Settings, or run `logread -e shater`.',
before,
after,
})
})()
})
return () => {
cancelled = true
off()
}
}, [flash, refreshStatus])
const confirmWindow = globals?.ConfirmTimeout ?? 0
@@ -241,6 +275,17 @@ export default function Apply() {
// is a status readout, and the config on disk can already differ from what is
// installed. Falls back to the config only while /api/status is unread.
const killArmed = (status?.kill_switch ?? globals?.KillSwitch ?? 'closed') === 'closed'
// Whether that setting is actually installed — same three-plus-unknown reading
// as Overview, so the two pages cannot disagree about the same router.
const kill = killSwitchReadout(status, globals?.KillSwitch)
const killWord =
kill.state === 'open'
? 'open'
: kill.state === 'armed'
? 'fail-closed'
: kill.state === 'inert'
? 'closed · not in effect'
: 'closed · not reported'
// Every engine mark on this page comes from ONE reading, and that reading is
// able to say "stopped" — see planeState.engineState for why `status.running`
// could not. This page is where someone lands when the network is down; three
@@ -282,11 +327,7 @@ export default function Apply() {
variant={dataVariant}
value={status?.table ? 'nft installed' : 'no table'}
/>
<StatusPip
label="Kill-switch"
variant={killArmed ? 'on' : 'amber'}
value={killArmed ? 'fail-closed' : 'open'}
/>
<StatusPip label="Kill-switch" variant={kill.variant} value={killWord} />
</div>
{statusError && (
@@ -320,7 +361,7 @@ export default function Apply() {
v: status?.table ? 'nft installed' : 'no table',
hot: dataVariant === 'crit',
},
{ k: 'kill-switch', v: killArmed ? 'fail-closed' : 'open', hot: !killArmed },
{ k: 'kill-switch', v: killWord, hot: kill.variant === 'crit' || !killArmed },
]}
/>
<Module
@@ -374,8 +415,9 @@ export default function Apply() {
is what "live but not kept" means — so the readout survives a
reload instead of depending on what this tab remembers. */}
Applied config <span className="mono">{liveHash}</span> is live but not yet kept.
Confirm to keep it — otherwise the daemon rolls back to the last-good config when
the timer hits zero.
Confirm to keep it. At zero the daemon rolls back to the last-good config — unless
something else re-applies the data plane first, in which case it stands down and
keeps whatever is live.
</p>
<div className="cc-bar" aria-hidden="true">
<span className="cc-bar-fill" style={{ width: `${pct}%` }} />
@@ -492,7 +534,9 @@ function labelFor(kind: ActionKind): string {
case 'rollback':
return 'Rollback'
case 'expire':
return 'Auto-rollback'
// NOT "Auto-rollback": on expiry the daemon either reverts or stands down,
// and this page cannot tell which. Name the event it did observe.
return 'Window elapsed'
}
}
-67
View File
@@ -648,73 +648,6 @@
}
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.dns-alert-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.dns-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.dns-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-alert-row delivery controls sit inline in the row (like .dns-detour) */
.dns-alert-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.dns-alert-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.dns-alert-note--row {
margin-top: 2px;
}
@media (max-width: 640px) {
.dns-alert-ctl {
flex-basis: 100%;
order: 3;
}
}
/* alert event checkboxes */
.dns-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.dns-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.dns-event input {
accent-color: var(--fp-accent, currentColor);
}
@media (prefers-reduced-motion: reduce) {
.dns-skel {
animation: none;
+1 -480
View File
@@ -9,7 +9,7 @@ import {
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { Alert, DNSRule, Model, Resolver, RulesetStatus } from '../api'
import type { DNSRule, Model, Resolver, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
// The DNS / Blocklists page is a thin editor over the desired-state Model —
@@ -60,28 +60,9 @@ type GlobalsX = Model['Globals'] & { DNSFilter?: boolean }
type DNSModel = Model & {
Blocklists?: Blocklist[] | null
Allowlists?: Allowlist[] | null
Alerts?: Alert[] | null
DNSRules?: DNSRule[] | null
}
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers ---------------------------------------------------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
@@ -310,7 +291,6 @@ export default function DNS() {
const blocklists = useMemo(() => asArray(config?.Blocklists), [config])
const allowlists = useMemo(() => asArray(config?.Allowlists), [config])
const resolvers = useMemo<Resolver[]>(() => asArray(config?.Resolvers), [config])
const alerts = useMemo<Alert[]>(() => asArray(config?.Alerts), [config])
// Ascending Order — the engine evaluates DNS rules first-match, so the list is
// shown and edited in the order it actually runs.
const dnsRules = useMemo<DNSRule[]>(
@@ -719,78 +699,6 @@ export default function DNS() {
[config, dnsRules, save, confirm],
)
// ---- alert mutations ------------------------------------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!config) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return save({ ...config, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[config, alerts, save],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void save({ ...config, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[config, alerts, save],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!config) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void save({ ...config, Alerts: next }, `Deleted ${target.Name}`)
},
[config, alerts, save, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!config) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void save(
{ ...config, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[config, alerts, save],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void save(
{ ...config, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[config, alerts, save],
)
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Reuse the resolver detour catalog with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({ ...detourCatalog, chains: [] }),
[detourCatalog],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
const loading = config === null && loadError === null
return (
@@ -1211,57 +1119,6 @@ export default function DNS() {
)}
</div>
{/* ---- 5. ALERTS ---- */}
<div className="dns-section" aria-label="Alerts">
<header className="dns-sec-hd">
<h2 className="dns-sec-title">Alerts</h2>
<span className="dns-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="dns-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="dns-rows" aria-hidden="true">
<li className="dns-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="dns-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
{toast && (
<div className="toast" role="status">
{toast}
@@ -1271,342 +1128,6 @@ export default function DNS() {
)
}
// ---- alert add form + row --------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="dns-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="dns-add-top">
<input
className="dns-input dns-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="dns-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="dns-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="dns-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="dns-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="dns-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="dns-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="dns-alert-delivery">
<label className="dns-resp">
<span className="dns-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="dns-alert-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="dns-add-actions">
{err && (
<p className="dns-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="dns-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="dns-row-main">
<div className="dns-row-l1">
<span className="dns-row-name">{alert.Name}</span>
<span className="dns-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="dns-badge dns-badge--accent">
{e}
</span>
))}
</div>
<div className="dns-row-l2 mono">
<span className="dns-row-detail">{detail.text}</span>
{detail.masked && (
<span className="dns-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="dns-path">direct</span>
) : (
<span className="dns-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="dns-path-name">{route.name}</strong>
{route.missing && <span className="dns-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="dns-alert-note dns-alert-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="dns-alert-ctl">
<label className="dns-detour">
<span className="dns-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="dns-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
// ---- add form --------------------------------------------------------------
interface AddDraft {
+60
View File
@@ -196,6 +196,52 @@
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
color: var(--amber);
}
/* "not built" — the saved switch says on and the engine has no such outbound. */
.badge--crit {
border-color: color-mix(in srgb, var(--crit) 55%, var(--groove));
color: var(--crit);
}
/* ---- last-apply findings, attached to the row they are about ----
Sits under the row's own two lines, inside the row plate, so a node the
generator threw away cannot read as an ordinary enabled node. Severity carries
the colour; the accent stays reserved for controls. */
.row-findings {
margin: 6px 0 0;
padding: 0;
list-style: none;
display: flex;
flex-direction: column;
gap: 5px;
}
.row-finding {
display: flex;
align-items: flex-start;
gap: 8px;
padding: 7px 9px;
border: 1px solid color-mix(in srgb, var(--amber) 40%, var(--groove));
border-radius: 6px;
background: color-mix(in srgb, var(--sink) 35%, transparent);
}
.row-finding--critical {
border-color: color-mix(in srgb, var(--crit) 45%, var(--groove));
}
.row-finding-msg {
flex: 1;
min-width: 0;
font-size: 12px;
line-height: 1.5;
color: var(--ink);
max-width: 82ch;
overflow-wrap: anywhere;
}
/* Findings that belong to no single row (see Nodes.tsx globalFindings). */
.node-findings {
margin-bottom: calc(var(--u, 8px) * 2);
}
.node-findings .row-findings {
margin-top: 0;
}
/* masked-credential marker */
.masked {
@@ -641,6 +687,20 @@ select.fp-input {
letter-spacing: 0.06em;
color: var(--faint);
}
/* A collapsed bucket has to carry its own bad news: a 300-node subscription is
closed by default, and the per-row findings inside it are otherwise unreachable
without knowing to look. */
.group-flagged {
flex: none;
display: inline-flex;
align-items: center;
gap: 6px;
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--amber);
}
.group-rows {
margin-top: 8px;
}
+131 -2
View File
@@ -6,12 +6,14 @@ import { Button, Led, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
getStatus,
putConfig,
importWg,
updateSubscription,
ApiError,
} from '../api'
import type { Model, Node as NodeCfg, Subscription } from '../api'
import type { Model, Node as NodeCfg, StatusWarning, Subscription } from '../api'
import { entityFindings, findingsByName } from '../findings'
import { fmtBytes, fmtDate, fmtUntil } from '../format'
// The whole page is a thin editor over the desired-state Model: every mutation
@@ -519,6 +521,43 @@ export default function Nodes() {
void loadConfig()
}, [loadConfig])
// ---- what the last apply said about these nodes ---------------------------
//
// The generator drops a node it cannot build and names it: an unparseable
// share link (generate/outbound.go), a name colliding with a reserved tag, a
// WireGuard private key materialised twice (generate/wgdedup.go). Until now
// this page never read /api/status, so a node the engine had thrown away
// rendered as an ordinary row with a green toggle — the switch said on and
// there was no such outbound anywhere in the running config.
//
// Findings are attached to the ROWS, not summarised at the top: a 300-node
// subscription makes a list of names useless, and the row is where the false
// reassurance was.
const [findings, setFindings] = useState<StatusWarning[]>([])
const loadFindings = useCallback(async () => {
try {
const s = await getStatus()
setFindings(entityFindings(s.warnings, ['node', 'subscription']))
} catch {
// Status is a supplement here, not the page. Keep the last set rather than
// clearing it — a dropped poll is not the same as "the problem is fixed".
}
}, [])
useEffect(() => {
void loadFindings()
}, [loadFindings])
const nodeFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'node')),
[findings],
)
const subFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'subscription')),
[findings],
)
// Findings about nodes/subscriptions in general, which belong to no single row.
const globalFindings = useMemo(() => findings.filter((w) => !w.name), [findings])
// ---- toast + persistent apply banner --------------------------------------
const [toast, setToast] = useState<string | null>(null)
const toastTimer = useRef<number | undefined>(undefined)
@@ -573,8 +612,11 @@ export default function Nodes() {
flash(`Apply failed — ${errText(e)}`)
} finally {
setApplying(false)
// An apply is exactly what rewrites the findings — including clearing the
// ones the operator just fixed.
void loadFindings()
}
}, [flash, loadConfig])
}, [flash, loadConfig, loadFindings])
// ---- node mutations -------------------------------------------------------
const nodes = useMemo(() => asArray(config?.Nodes), [config])
@@ -994,6 +1036,14 @@ export default function Nodes() {
</div>
)}
{/* Findings about nodes in general — no single row owns them, so they sit
above the lists rather than being dropped for having no name. */}
{globalFindings.length > 0 && (
<div className="node-findings">
<RowFindings findings={globalFindings} />
</div>
)}
{/* ---- NODES ---- */}
<div className="node-section" aria-label="Nodes">
<header className="sec-hd">
@@ -1149,6 +1199,7 @@ export default function Nodes() {
<NodeGroup
key={g.key || '__manual__'}
group={g}
findings={nodeFindings}
open={isGroupOpen(g)}
busy={busy}
egressNames={egressNames}
@@ -1238,6 +1289,7 @@ export default function Nodes() {
busy={busy}
catalog={detourCatalog}
valid={detourValid}
findings={subFindings.get(s.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => toggleSub(i, on)}
onDelete={() => removeSub(i)}
onEdit={(patch) => editSub(i, patch)}
@@ -1264,6 +1316,7 @@ function NodeGroup({
open,
busy,
egressNames,
findings,
onToggle,
onToggleNode,
onRemoveNode,
@@ -1274,6 +1327,8 @@ function NodeGroup({
open: boolean
busy: boolean
egressNames: string[]
/** Last-apply findings per node name (findings.ts findingsByName). */
findings: Map<string, StatusWarning[]>
onToggle: () => void
onToggleNode: (idx: number, on: boolean) => void
onRemoveNode: (idx: number) => void
@@ -1283,6 +1338,13 @@ function NodeGroup({
const panelId = `node-group-${group.key || 'manual'}`
// The same inventory count as the section header, scoped to this bucket.
const count = useMemo(() => fmtEnabled(group.items.map((i) => i.node)), [group.items])
// How many nodes in this bucket the last apply had something to say about —
// shown on the COLLAPSED header, because a subscription of 300 nodes is
// collapsed by default and the row badge below would never be seen otherwise.
const flagged = useMemo(
() => group.items.filter(({ node }) => findings.has(node.Name)).length,
[group.items, findings],
)
return (
<section className={`node-group${open ? ' node-group--open' : ''}`}>
<h3 className="group-hd-wrap">
@@ -1296,6 +1358,12 @@ function NodeGroup({
<span className="group-caret" aria-hidden="true" />
<span className="group-name">{group.label}</span>
<span className="group-count mono">{count}</span>
{flagged > 0 && (
<span className="group-flagged" title="Findings from the last apply">
<Led variant="amber" />
{flagged} flagged
</span>
)}
</button>
</h3>
{open && group.key !== '' && (
@@ -1312,6 +1380,7 @@ function NodeGroup({
node={node}
busy={busy}
egressNames={egressNames}
findings={findings.get(node.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => onToggleNode(idx, on)}
onDelete={() => onRemoveNode(idx)}
onRename={(name, onError) => onRenameNode(idx, name, onError)}
@@ -1324,10 +1393,29 @@ function NodeGroup({
)
}
/** One shared empty array, so a clean row doesn't get a fresh identity per render. */
const EMPTY_FINDINGS: StatusWarning[] = []
/**
* Did the generator say it left this entity OUT of the engine config?
*
* The producers all end the sentence with the same word — "(skipped)" for an
* unparseable share link or a bad WireGuard endpoint (generate/outbound.go),
* "skipped" for a name colliding with a reserved tag — and wgdedup says only one
* of the duplicates "is kept". Read the daemon's word rather than inventing a
* verdict: a finding that does NOT say this may well be about a node that is
* running perfectly, and badging it "not built" would be a new lie in place of
* the old one.
*/
function skipped(findings: StatusWarning[]): boolean {
return findings.some((f) => /\bskipped\b|\bis kept\b/i.test(f.message))
}
function NodeRow({
node,
busy,
egressNames,
findings,
onToggle,
onDelete,
onRename,
@@ -1336,6 +1424,8 @@ function NodeRow({
node: NodeCfg
busy: boolean
egressNames: string[]
/** What the last apply said about THIS node; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onRename: (name: string, onError: (msg: string) => void) => Promise<boolean>
@@ -1483,6 +1573,16 @@ function NodeRow({
)}
<span className="badge">{proto}</span>
{node.Stale && <span className="badge badge--warn">stale</span>}
{/* The toggle above is the SAVED state. When the last apply couldn't
build this node the engine has no such outbound, and the two
disagree — so the row says which, rather than leaving a green
switch to imply the node is carrying traffic. The word is the
daemon's own where it used one. */}
{findings.length > 0 && (
<span className={`badge badge--${skipped(findings) ? 'crit' : 'warn'}`}>
{skipped(findings) ? 'not built' : 'flagged'}
</span>
)}
</div>
{renameErr && (
<p className="row-err" role="alert">
@@ -1506,6 +1606,7 @@ function NodeRow({
</span>
)}
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
{canPin && (
@@ -1625,6 +1726,7 @@ function SubRow({
busy,
catalog,
valid,
findings,
onToggle,
onDelete,
onEdit,
@@ -1635,6 +1737,8 @@ function SubRow({
busy: boolean
catalog: DetourCatalog
valid: Set<string>
/** What the last apply said about THIS subscription; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onEdit: (patch: Subscription) => Promise<boolean>
@@ -1659,6 +1763,7 @@ function SubRow({
<span className="row-name">{sub.Name}</span>
{sub.Format && sub.Format !== 'auto' && <span className="badge">{sub.Format}</span>}
{sub.FetchVia === 'proxy' && <span className="badge">via proxy</span>}
{findings.length > 0 && <span className="badge badge--warn">flagged</span>}
</div>
<div className="row-line2 mono">
<span className="row-host">{host}</span>
@@ -1671,6 +1776,7 @@ function SubRow({
every {interval} · {count} node{count === 1 ? '' : 's'}
</span>
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
<Button
@@ -2142,6 +2248,29 @@ function HeaderRows({
)
}
/**
* What the last apply said about THIS row, under the row it is about.
*
* Deliberately inside the row rather than in a list at the top of the page: the
* failure being fixed is a node that looks fine, and a name in a summary three
* screens up does not fix that. The wording is the daemon's own — these messages
* already name the entity and say what was done about it ("(skipped)", "only X
* is kept"), so paraphrasing them here would only invent a second vocabulary.
*/
function RowFindings({ findings }: { findings: StatusWarning[] }) {
if (findings.length === 0) return null
return (
<ul className="row-findings" aria-label="Findings from the last apply">
{findings.map((f, i) => (
<li key={i} className={`row-finding row-finding--${f.severity}`}>
<Led variant={f.severity === 'critical' ? 'crit' : 'amber'} />
<span className="row-finding-msg">{f.message}</span>
</li>
))}
</ul>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="empty-plate">
+51 -14
View File
@@ -14,8 +14,8 @@ import type { Model, Stats, Status, StatusWarning } from '../api'
import { confirmTimeout } from '../pendingConfirm'
import { navigate } from '../router'
import type { Route } from '../router'
import { attentionFindings } from '../findings'
import { engineReadout, protectionState } from '../planeState'
import { attentionFindings, truncationNote } from '../findings'
import { engineReadout, killSwitchReadout, protectionState } from '../planeState'
// null-safe length for a Go slice that may arrive as null.
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
@@ -309,15 +309,19 @@ export function Overview({
const engineVariant: LedVariant = engine.variant
const protection = protectionState(status)
// Configured fail-closed AND actually enforcing it. `none` means nothing is
// installed, so the setting is inert no matter what it says.
const killInEffect = killArmed && status?.plane !== 'none'
// Configured fail-closed, actually enforcing it, or not known — three answers,
// and the third is not folded into the first. See planeState.killSwitchReadout.
const kill = killSwitchReadout(status, g?.KillSwitch)
// Findings that need attention. `info` notes are statements about the config,
// not problems, so they live beside the setting they describe (see findings.ts)
// — keeping this list to things someone could actually act on.
const warnings = attentionFindings(status?.warnings)
const criticalCount = warnings.filter((w) => w.severity === 'critical').length
// The daemon caps the published list at 50 and says so in an `info` note — the
// one channel this page filters away. Carried separately so the list can admit
// it is not the whole list. See findings.ts truncationNote.
const truncated = truncationNote(status?.warnings)
return (
<section className="page" aria-label="Overview">
@@ -340,7 +344,7 @@ export function Overview({
</p>
)}
<Findings warnings={warnings} criticalCount={criticalCount} />
<Findings warnings={warnings} criticalCount={criticalCount} truncated={truncated} />
<div className="grid">
{/* Groups, not nodes: a group is where a dial path is defined, so it is the
@@ -447,17 +451,17 @@ export function Overview({
/>
{/* A kill-switch set to fail-closed is only ARMED if something is actually
installed to enforce it. With no plane it is configured but inert, and
saying "ARMED" there would be a false reassurance next to a readout
that says nothing is protected. */}
installed to enforce it, and "we haven't been told" is neither. With no
plane it is configured but inert; with no reading the lamp stays unlit
rather than joining the healthy branch by default. */}
<Module
name="Kill-switch"
value={killInEffect ? 'ARMED' : killArmed ? 'NOT IN EFFECT' : 'OPEN'}
led={{ variant: killInEffect ? 'on' : killArmed ? 'crit' : 'amber' }}
value={kill.value}
led={{ variant: kill.variant }}
rows={[
{ k: 'setting', v: killArmed ? 'fail-closed' : 'fail-open', hot: !killArmed },
...(killArmed && !killInEffect
? [{ k: 'blocking now', v: 'no — nothing installed', hot: true }]
...(kill.blockingNow
? [{ k: 'blocking now', v: kill.blockingNow, hot: kill.hot }]
: [{ k: 'ipv6', v: g?.IPv6 ? 'covered' : 'off' }]),
{ k: 'confirm', v: g?.ConfirmTimeout ? `${g.ConfirmTimeout}s window` : 'no auto-rollback' },
]}
@@ -536,11 +540,23 @@ const SECTION_ROUTE: Record<string, Route> = {
rule: 'routing',
ruleset: 'routing',
blocklist: 'dns',
allowlist: 'dns',
resolver: 'dns',
dns_rule: 'dns',
device: 'devices',
chain: 'targets',
group: 'targets',
// A node the generator dropped (unparseable share link, duplicate WireGuard
// key, name colliding with a reserved tag) is reported under `node` — and had
// nowhere to jump to, so the one page that could show it a green toggle was
// also the one page the finding could not reach.
node: 'nodes',
subscription: 'nodes',
egress: 'targets',
inbound: 'networks',
interface: 'networks',
profile: 'profiles',
alert: 'settings',
// The standing note about non-TCP/UDP traffic — its control lives on Networks.
untunnelable: 'networks',
}
@@ -558,11 +574,14 @@ const SECTION_ROUTE: Record<string, Route> = {
function Findings({
warnings,
criticalCount,
truncated,
}: {
warnings: StatusWarning[]
criticalCount: number
/** The daemon's "N further suppressed" note, when the list was capped. */
truncated: StatusWarning | null
}) {
if (warnings.length === 0) return null
if (warnings.length === 0 && !truncated) return null
const rank = { critical: 0, warning: 1, info: 2 } as const
const sorted = [...warnings].sort((a, b) => rank[a.severity] - rank[b.severity])
@@ -572,6 +591,10 @@ function Findings({
<header className="findings-hd">
<h2 className="findings-title">Last apply</h2>
<span className="findings-count mono">
{/* "at least" whenever the list was capped: the counts below it are a
floor, not a total, and the cap drops the least severe FIRST — so
on a config with fifty criticals the thing it drops is a critical. */}
{truncated ? 'at least ' : ''}
{criticalCount > 0
? `${criticalCount} critical · ${warnings.length} total`
: `${warnings.length} note${warnings.length === 1 ? '' : 's'}`}
@@ -610,6 +633,20 @@ function Findings({
</li>
)
})}
{/* The list saying it is not the whole list. Last, because it is about
everything above it — and never filtered out with the other `info`
notes, which is where it used to disappear. */}
{truncated && (
<li className="finding finding--truncated">
<Led variant="amber" />
<div className="finding-copy">
<span className="finding-where mono">list truncated</span>
<span className="finding-msg">
Some findings are missing from this list. {truncated.message}
</span>
</div>
</li>
)}
</ul>
</section>
)
+51 -11
View File
@@ -60,16 +60,31 @@ type RuleForce = {
}
/**
* Everything `Proto` can match, and nothing else. The engine understands two
* transports and exactly ten application protocols its sniffers can name
* (generate/route.go sniffedProtocols); a value outside this set builds a rule
* that is perfectly valid and can never fire — so its traffic quietly falls
* Everything `Proto` can match, and nothing else. A value outside this set builds
* a rule that is perfectly valid and can never fire — so its traffic quietly falls
* through to whatever rule sits below it. That is why this is a closed list and
* not a text box.
*
* Split into two groups because they answer different questions: the transport is
* known the moment a packet arrives, while an app protocol is only known once the
* first bytes have been read and labelled.
* Three groups, because the engine reads them through three different matchers
* (generate/route.go, ruleMatchers) and they answer different questions:
*
* - Transport — the L4 network. Known the moment a packet arrives.
* - Detected protocol — the L7 label a sniffer puts on a connection once its
* first bytes have been read. This group, and ONLY this group, is the engine's
* `sniffedProtocols` set; anything else routed into that matcher is inert.
* - Layer 3 — ICMP. Not a sniffed label: it lands in the emitted rule's
* `network`, never in `protocol` (the sniffers are skipped outright for an
* ICMP flow, so they never report "icmp"). All three spellings are the SAME
* one network; `icmpv4`/`icmpv6` additionally pin `ip_version`, which the
* engine derives from the destination address.
*
* ICMP carries caveats the picker deliberately does not try to enforce, because
* the daemon reports each one against the whole config on apply: it reaches the
* engine only while globals l3_tunnel is on, it has no ports (a port matcher
* beside it can never be satisfied), `icmpv6` also needs globals ipv6 on, and it
* is DROPPED rather than falling through when routed at a target that cannot
* carry layer 3 — i.e. every proxy protocol. Only wireguard/AmneziaWG nodes and
* direct/interface egresses can carry a ping.
*/
const PROTO_TRANSPORT: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'tcp', label: 'TCP' },
@@ -87,10 +102,27 @@ const PROTO_APP: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'rdp', label: 'RDP' },
{ id: 'ntp', label: 'NTP' },
]
const PROTO_VALUES = new Set([...PROTO_TRANSPORT, ...PROTO_APP].map((p) => p.id))
/**
* The family-qualified spellings are offered next to plain `icmp` rather than
* hidden behind it: the engine treats them as first-class and the difference is
* observable (an `ip_version` item on the same rule), so hiding them would leave a
* capability reachable only by hand-editing /etc/config/shater — and would mean
* that anyone who edited such a rule here lost the narrowing on the next save.
*/
const PROTO_L3: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'icmp', label: 'ICMP (ping)' },
{ id: 'icmpv4', label: 'ICMP — IPv4 only' },
{ id: 'icmpv6', label: 'ICMP — IPv6 only' },
]
const PROTO_VALUES = new Set(
[...PROTO_TRANSPORT, ...PROTO_APP, ...PROTO_L3].map((p) => p.id),
)
/** The Proto picker's option list — shared by the inline add row and the editor. */
function ProtoOptions({ value }: { value: string }) {
// The engine lower-cases `Proto` before matching it, so a hand-written `ICMP`
// is a working rule; judge it the same way and flag only what really is inert.
const matches = PROTO_VALUES.has(value.trim().toLowerCase())
return (
<>
<option value="">any</option>
@@ -108,10 +140,18 @@ function ProtoOptions({ value }: { value: string }) {
</option>
))}
</optgroup>
{/* A stored value the engine can't detect is kept and flagged, never
silently rewritten — the rule it belongs to is live right now. */}
<optgroup label="Layer 3">
{PROTO_L3.map((p) => (
<option key={p.id} value={p.id}>
{p.label}
</option>
))}
</optgroup>
{/* A stored value none of the groups spells verbatim is kept and offered as
written, never silently rewritten — the rule it belongs to is live right
now. It is flagged only when the engine cannot match it either. */}
{value !== '' && !PROTO_VALUES.has(value) && (
<option value={value}>{value} — never matches</option>
<option value={value}>{matches ? value : `${value} — never matches`}</option>
)}
</>
)
+9 -1
View File
@@ -2,6 +2,7 @@ import './Settings.css'
import { useCallback, useEffect, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { AlertsSection } from './Alerts'
import { apply as apiApply, downloadLog, getConfig, putConfig, ApiError } from '../api'
import type { Globals, LogRange, Model } from '../api'
@@ -406,7 +407,7 @@ export default function Settings() {
<Field
label="Log level"
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level."
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level — set up where they go in the Alerts section below."
>
<Select
value={globals?.LogLevel || 'warning'}
@@ -620,6 +621,13 @@ export default function Settings() {
</Field>
</Group>
{/* ---- ALERTS ---- */}
{/* Extracted from the DNS page — out-of-band notifications belong with
the appliance-wide knobs, next to the log level whose note points
here. Renders its own section header (same plate as a Group); all
writes go through `save`, so the dirty banner and toast stay one. */}
<AlertsSection config={config} busy={busy} loading={loading} onSave={save} />
{/* ---- STATISTICS & LOGGING ---- */}
<Group
title="Statistics &amp; logging"
+63 -1
View File
@@ -15,7 +15,7 @@
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { engineReadout, engineState, protectionState } from './planeState.ts'
import { engineReadout, engineState, killSwitchReadout, protectionState } from './planeState.ts'
import type { Status, Traffic } from './api.ts'
/** A healthy, fully-installed router; `traffic` is what each case varies. */
@@ -189,3 +189,65 @@ test('no status at all is unknown, not down', () => {
assert.equal(engineReadout(null).variant, 'off')
assert.equal(engineReadout(null).word, 'checking…')
})
// --- killSwitchReadout: "I don't know" is not "it's armed" -------------------
//
// The Overview module read `killArmed && status?.plane !== 'none'`, and
// `undefined !== 'none'` is true — so a daemon that never reported `plane`, and
// the seconds before the first status arrives, both lit a green lamp over the
// word ARMED. These pin the fourth answer that expression could not express.
test('a daemon that does not report `plane` reads as not reported, never ARMED', () => {
const { plane, ...noPlane } = status()
void plane
const k = killSwitchReadout(noPlane as Status)
assert.equal(k.state, 'unknown')
assert.notEqual(k.value, 'ARMED')
assert.equal(k.variant, 'off')
assert.notEqual(k.variant, 'on')
assert.equal(k.blockingNow, 'not known')
})
test('no status at all is unknown too, and says there is no reading', () => {
const k = killSwitchReadout(null)
assert.equal(k.state, 'unknown')
assert.equal(k.variant, 'off')
assert.equal(k.blockingNow, 'no reading yet')
})
test('fail-closed with a plane installed is armed', () => {
for (const plane of ['full', 'hold'] as const) {
const k = killSwitchReadout(status({ plane }))
assert.equal(k.state, 'armed')
assert.equal(k.value, 'ARMED')
assert.equal(k.variant, 'on')
assert.equal(k.blockingNow, null)
}
})
test('fail-closed with no plane is configured but blocking nothing', () => {
const k = killSwitchReadout(status({ plane: 'none', table: false }))
assert.equal(k.state, 'inert')
assert.equal(k.value, 'NOT IN EFFECT')
assert.equal(k.variant, 'crit')
assert.equal(k.hot, true)
})
test('fail-open is the operator’s choice — amber, and never a plane question', () => {
for (const plane of ['full', 'none', undefined] as const) {
const k = killSwitchReadout(status({ kill_switch: 'open', plane }))
assert.equal(k.state, 'open')
assert.equal(k.value, 'OPEN')
assert.equal(k.variant, 'amber')
}
})
test('the live kill_switch wins over the saved one; the saved one only fills a gap', () => {
const { kill_switch, ...noKill } = status()
void kill_switch
// Live says open, config says closed → live wins.
assert.equal(killSwitchReadout(status({ kill_switch: 'open' }), 'closed').state, 'open')
// Nothing live → fall back to the saved policy.
assert.equal(killSwitchReadout(noKill as Status, 'open').state, 'open')
assert.equal(killSwitchReadout(noKill as Status, 'closed').state, 'armed')
})
+79
View File
@@ -71,6 +71,85 @@ export function engineReadout(status: Status | null): { variant: LedVariant; wor
}
}
// ---------------------------------------------------------------------------
// Is the kill-switch actually blocking anything?
// ---------------------------------------------------------------------------
/**
* Four answers, and "unknown" is one of them.
*
* armed — configured fail-closed AND a data plane is installed to enforce it.
* inert — configured fail-closed, but there is no plane. Nothing is blocking.
* unknown — the daemon has not said how much plane is installed, so whether the
* setting is in force is not known. NEVER paint this green.
* open — configured fail-open. Nothing is meant to be blocked.
*/
export type KillSwitchState = 'armed' | 'inert' | 'unknown' | 'open'
export interface KillSwitchReadout {
state: KillSwitchState
/** The word the module puts in its readout. */
value: string
variant: LedVariant
/** Is it blocking right now — the row under the readout. `null` ⇒ nothing to add. */
blockingNow: string | null
/** True when `blockingNow` is bad news and should be drawn hot. */
hot: boolean
}
/**
* THE UNKNOWN BRANCH IS THE WHOLE POINT. This used to be
*
* killArmed && status?.plane !== 'none'
*
* and `undefined !== 'none'` is true — so a daemon that had not reported `plane`
* at all, and a panel that had not yet received its first status, both landed in
* the "ARMED" branch under a green lamp. Every other unknown in this file is an
* unlit socket for exactly this reason (see engineReadout): the kill-switch is
* the last thing standing between the LAN and the plain WAN, and "I don't know
* whether it is installed" must never be dressed as "it is".
*
* `configured` is the SAVED policy from /api/config, used only while
* /api/status has not reported one. The live value wins wherever it exists, as
* everywhere else in the panel: this is a status readout, and the config on disk
* can already differ from what is installed.
*/
export function killSwitchReadout(
status: Status | null,
configured?: string,
): KillSwitchReadout {
const closed = (status?.kill_switch ?? configured ?? 'closed') === 'closed'
if (!closed) {
return { state: 'open', value: 'OPEN', variant: 'amber', blockingNow: null, hot: false }
}
switch (status?.plane) {
case 'full':
case 'hold':
// Something is installed, so the fail-closed guard is really in the path.
return { state: 'armed', value: 'ARMED', variant: 'on', blockingNow: null, hot: false }
case 'none':
return {
state: 'inert',
value: 'NOT IN EFFECT',
variant: 'crit',
blockingNow: 'no — nothing installed',
hot: true,
}
default:
return {
state: 'unknown',
value: 'NOT REPORTED',
variant: 'off',
// Terse on purpose: this is a two-column readout row, and the long form
// wrapped onto three lines beside a one-word key.
blockingNow: status ? 'not known' : 'no reading yet',
hot: false,
}
}
}
/**
* `plane` + `engine_running` express the state more precisely than the three
* booleans the old status strip exposed (engine active / config enabled / nft
+48 -1
View File
@@ -1,15 +1,62 @@
import { defineConfig } from 'vite'
import type { Plugin } from 'vite'
import react from '@vitejs/plugin-react'
// Minimal ambient for the dev-proxy target override — avoids pulling in @types/node
// just for one env read. Vite runs this file under Node where `process` exists.
declare const process: { env: Record<string, string | undefined> }
/** `src/mock.ts`, as the module graph spells it (POSIX-normalised for Windows). */
const MOCK_MODULE = 'src/mock.ts'
/**
* Refuse to emit a production bundle that contains the offline fixture backend.
*
* `src/mock.ts` describes an invented, healthy router: a full config, 122 nodes
* with 119 of them alive, "Protected". It exists so `npm run dev` renders without
* a daemon. It shipped inside the binary that goes on real hardware, switched on
* by nothing more than a `?dev` on the end of the URL — so a link someone was
* sent, or a bookmark they saved, showed an appliance in perfect health while
* making no request to the appliance at all.
*
* api.ts now loads it behind `import.meta.env.DEV`, which Vite folds to a literal
* `false` for a build, so Rollup drops the dynamic import and the module never
* enters the graph. That is a property of a build tool's optimiser, and an
* optimiser is not a promise: one refactor that makes the condition non-static
* silently puts the fixtures back. So the property is CHECKED rather than
* trusted — if `src/mock.ts` reaches any emitted chunk, the build fails here
* instead of shipping.
*/
function assertNoMockFixtures(): Plugin {
return {
name: 'shater:assert-no-mock-fixtures',
apply: 'build',
generateBundle(_options, bundle) {
const guilty: string[] = []
for (const [file, output] of Object.entries(bundle)) {
if (output.type !== 'chunk') continue
for (const id of output.moduleIds) {
if (id.replace(/\\/g, '/').endsWith(MOCK_MODULE)) guilty.push(`${file} ← ${id}`)
}
}
if (guilty.length > 0) {
this.error(
`the offline fixture backend (${MOCK_MODULE}) reached the production bundle:\n ` +
guilty.join('\n ') +
`\nFixtures describe a router that does not exist. Keep every path to them behind ` +
`\`import.meta.env.DEV\` so Rollup can drop them, and never gate them on a runtime ` +
`flag such as a query parameter.`,
)
}
},
}
}
// The SPA is embedded in the forked sing-box binary and served by the daemon on
// its own port. Relative base so it works under any mount path; single small
// bundle (no code-splitting) keeps the embed simple and the flash budget low.
export default defineConfig({
plugins: [react()],
plugins: [react(), assertNoMockFixtures()],
base: './',
build: {
outDir: 'dist',
+268
View File
@@ -0,0 +1,268 @@
// lx:begin l3-honest-drop
package route
import (
"context"
"net/netip"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
R "github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/stretchr/testify/require"
)
// The contract under test: PreMatch never answers "continue" (nor "bypass") for
// an ICMP flow. adapter.JudgeFlow maps both to tun.ActionAccept, and the TUN
// stack answers Accept by FORGING the echo reply itself
// (sing-tun stack_gvisor_icmp.go ICMPForwarder.HandlePacket, the fallthrough
// under the Flow/Reject/Drop switch). A verdict of "continue" therefore reads to
// the operator as a working ping off a tunnel that never carried the packet.
//
// Every test below has a TCP/UDP twin: the honest drop must not leak into the
// protocols where "continue" really does mean "take the ordinary connection
// route".
// icmpL4Outbound is a minimal L4-only outbound (the vless/vmess/... shape): it
// does NOT implement adapter.FlowOutbound, and Network() lists only TCP/UDP.
// Unused Outbound methods come from the embedded nil interface and are never
// called on the pre-match paths under test.
type icmpL4Outbound struct {
adapter.Outbound
tag string
}
func (o *icmpL4Outbound) Tag() string { return o.tag }
func (o *icmpL4Outbound) Type() string { return "vless" }
func (o *icmpL4Outbound) Network() []string { return []string{N.NetworkTCP, N.NetworkUDP} }
// icmpOutboundManager resolves tags from a fixed map and hands the same L4-only
// outbound out as the default; the rest of the OutboundManager surface is never
// touched by the pre-match walk.
type icmpOutboundManager struct {
adapter.OutboundManager
defaultOutbound adapter.Outbound
outbounds map[string]adapter.Outbound
}
func (m *icmpOutboundManager) Default() adapter.Outbound { return m.defaultOutbound }
func (m *icmpOutboundManager) Outbound(tag string) (adapter.Outbound, bool) {
outbound, loaded := m.outbounds[tag]
return outbound, loaded
}
// icmpDNSRouter / icmpDNSTransportManager implement only what
// prepareMatchMetadata reaches. FakeIP returns nil unless a transport is
// installed, which is how the "fakeip lookup failed" exit is driven below.
type icmpDNSRouter struct {
adapter.DNSRouter
}
func (s *icmpDNSRouter) LookupReverseMapping(netip.Addr) (string, bool) { return "", false }
type icmpDNSTransportManager struct {
adapter.DNSTransportManager
fakeIP adapter.FakeIPTransport
}
func (s *icmpDNSTransportManager) FakeIP() adapter.FakeIPTransport {
if s.fakeIP == nil {
return nil
}
return s.fakeIP
}
// icmpMissingFakeIPTransport claims every address and then fails to look any of
// them up — exactly the "missing fakeip record, try enable
// `experimental.cache_file`" error prepareMatchMetadata returns.
type icmpMissingFakeIPTransport struct {
adapter.FakeIPTransport
}
func (t *icmpMissingFakeIPTransport) Store() adapter.FakeIPStore {
return &icmpMissingFakeIPStore{}
}
type icmpMissingFakeIPStore struct {
adapter.FakeIPStore
}
func (s *icmpMissingFakeIPStore) Contains(netip.Addr) bool { return true }
func (s *icmpMissingFakeIPStore) Lookup(netip.Addr) (string, bool) { return "", false }
type icmpRouterOptions struct {
fakeIP adapter.FakeIPTransport
rules []option.Rule
}
func icmpTestRouter(t *testing.T, options icmpRouterOptions) *Router {
t.Helper()
logger := log.NewNOPFactory().NewLogger("test")
defaultOutbound := &icmpL4Outbound{tag: "proxy-out"}
router := &Router{
ctx: context.Background(),
logger: logger,
dns: &icmpDNSRouter{},
dnsTransport: &icmpDNSTransportManager{fakeIP: options.fakeIP},
outbound: &icmpOutboundManager{
defaultOutbound: defaultOutbound,
outbounds: map[string]adapter.Outbound{defaultOutbound.Tag(): defaultOutbound},
},
}
for i, ruleOptions := range options.rules {
rule, err := R.NewRule(router.ctx, logger, ruleOptions, false)
require.NoError(t, err, "build rule[%d]", i)
router.rules = append(router.rules, rule)
}
return router
}
func icmpTestMetadata(network string) adapter.InboundContext {
return adapter.InboundContext{
Inbound: "l3-in",
InboundType: C.TypeTun,
Network: network,
Source: M.SocksaddrFrom(netip.MustParseAddr("192.168.1.2"), 0),
Destination: M.SocksaddrFrom(netip.MustParseAddr("1.1.1.1"), 0),
}
}
// lanRuleWithAction matches every packet from the test source, so the action is
// what the test is actually about.
func lanRuleWithAction(action option.RuleAction) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceIPCIDR: badoption.Listable[string]{"192.168.1.0/24"},
},
RuleAction: action,
},
}
}
// --- exit 1: an outbound that cannot carry layer 3 --------------------------
func TestPreMatchICMPToL4OutboundDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"ICMP to an L4-only outbound fell through to the ordinary pre-match path: the TUN stack will forge the echo reply and ping will lie about a tunnel that never saw the packet")
}
func TestPreMatchTCPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"TCP to an L4-only outbound must keep taking the ordinary connection route; the ICMP honest-drop must not leak into TCP/UDP pre-match")
}
func TestPreMatchUDPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkUDP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"UDP to an L4-only outbound must keep taking the ordinary connection route")
}
// --- exit 2: prepareMatchMetadata failed before any rule was walked ---------
// This exit arrived with the shared prepareMatchMetadata refactor (upstream
// b911fb078): it returns before the rule walk, so it never reaches preMatchFlow
// where the ICMP override used to live.
func TestPreMatchICMPMetadataErrorDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"a fakeip record that cannot be resolved must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPMetadataErrorContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"for TCP the metadata-error exit must keep meaning `take the ordinary connection route`")
}
// --- exit 3: a rule action the pre-match walk does not handle ---------------
// hijack-dns is one of the actions PreMatch's switch has no arm for, so it lands
// in the default arm. Any future unhandled action lands there too — that is why
// the guard is a funnel on the return value and not a per-arm override.
func TestPreMatchICMPUnhandledRuleActionDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"an unhandled rule action must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPUnhandledRuleActionContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"the unhandled-action exit must stay a continue for TCP")
}
// --- exit 4: an explicit bypass ---------------------------------------------
// sing-tun implements ActionBypass on the nfqueue plane only; on the TUN path it
// falls into the same default arm as Accept (flow_dispatch.go judgeAndInstall,
// and the ICMP forwarder's switch has no Bypass case either), i.e. into the same
// forgery. There is no honest bypass for a packet already inside the engine's
// TUN.
func TestPreMatchICMPBypassDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"bypass degrades to tun.ActionAccept on the TUN path, which is the forged echo reply again")
}
func TestPreMatchTCPBypassIsStillBypass(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchBypass, result.Action,
"the ICMP honest-drop must not turn a TCP bypass rule into a drop")
}
// --- the verdicts that must pass through untouched ---------------------------
// A reject rule already carries its own honest verdict; the funnel must not
// rewrite it (a Reject sends an ICMP unreachable, which is information, not a
// forged liveness signal).
func TestPreMatchICMPRejectIsNotRewritten(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"the ICMP funnel must only rewrite continue/bypass, never an explicit reject")
}
// lx:end l3-honest-drop
+194
View File
@@ -0,0 +1,194 @@
package route
import (
"context"
"net"
"net/netip"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
R "github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/stretchr/testify/require"
)
// The pre-match path (adapter.JudgeFlow -> Router.PreMatch, used by the TUN/
// WireGuard-endpoint flow dispatcher) used to prepare only fakeip and the IP
// version. Everything a rule matches on that the inbound cannot know — the
// connection owner and the neighbor behind the source address — was resolved in
// matchRule only, so a `source_mac_address` / `source_hostname` rule silently
// failed to match in pre-match and the flow fell through to the default
// outbound. Upstream b911fb078 shares one prepareMatchMetadata between both
// paths; these tests pin that.
// stubNeighborResolver answers for exactly one address.
type stubNeighborResolver struct {
address netip.Addr
mac net.HardwareAddr
hostname string
}
func (r *stubNeighborResolver) LookupMAC(address netip.Addr) (net.HardwareAddr, bool) {
if address != r.address || r.mac == nil {
return nil, false
}
return r.mac, true
}
func (r *stubNeighborResolver) LookupHostname(address netip.Addr) (string, bool) {
if address != r.address || r.hostname == "" {
return "", false
}
return r.hostname, true
}
func (r *stubNeighborResolver) LookupAddresses(hostname string) []netip.Addr {
if hostname != r.hostname {
return nil
}
return []netip.Addr{r.address}
}
func (r *stubNeighborResolver) Start() error { return nil }
func (r *stubNeighborResolver) Close() error { return nil }
// stubDNSRouter / stubDNSTransportManager implement only what
// prepareMatchMetadata reaches; every other method is left to the embedded nil
// interface and would panic if it were ever called.
type stubDNSRouter struct {
adapter.DNSRouter
}
func (s *stubDNSRouter) LookupReverseMapping(netip.Addr) (string, bool) { return "", false }
type stubDNSTransportManager struct {
adapter.DNSTransportManager
}
func (s *stubDNSTransportManager) FakeIP() adapter.FakeIPTransport { return nil }
// stubOutboundManager's default outbound supports no network at all, so a flow
// that reaches preMatchFlow bails out with PreMatchContinue instead of nil-
// dereferencing. That is exactly the pre-fix verdict we assert against.
type stubOutboundManager struct {
adapter.OutboundManager
defaultOutbound adapter.Outbound
}
func (s *stubOutboundManager) Default() adapter.Outbound { return s.defaultOutbound }
type stubNoNetworkOutbound struct {
adapter.Outbound
}
func (o *stubNoNetworkOutbound) Tag() string { return "stub" }
func (o *stubNoNetworkOutbound) Type() string { return "direct" }
func (o *stubNoNetworkOutbound) Network() []string { return nil }
func newPreMatchTestRouter(t *testing.T, resolver adapter.NeighborResolver, rules ...option.Rule) *Router {
t.Helper()
logger := log.NewNOPFactory().NewLogger("test")
router := &Router{
ctx: context.Background(),
logger: logger,
dns: &stubDNSRouter{},
dnsTransport: &stubDNSTransportManager{},
outbound: &stubOutboundManager{defaultOutbound: &stubNoNetworkOutbound{}},
neighborResolver: resolver,
needFindNeighbor: true,
}
for i, ruleOptions := range rules {
rule, err := R.NewRule(router.ctx, logger, ruleOptions, false)
require.NoError(t, err, "build rule[%d]", i)
router.rules = append(router.rules, rule)
}
return router
}
func rejectOnSourceMAC(macAddress string) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceMACAddress: badoption.Listable[string]{macAddress},
},
RuleAction: option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
},
},
}
}
func rejectOnSourceHostname(hostname string) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceHostname: badoption.Listable[string]{hostname},
},
RuleAction: option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
},
},
}
}
func preMatchMetadata() adapter.InboundContext {
return adapter.InboundContext{
Inbound: "tun-in",
InboundType: C.TypeTun,
Network: N.NetworkUDP,
Source: M.ParseSocksaddr("192.168.1.5:41234"),
Destination: M.ParseSocksaddr("1.1.1.1:443"),
}
}
func TestPreMatchResolvesNeighborMAC(t *testing.T) {
t.Parallel()
mac, err := net.ParseMAC("de:ad:be:ef:00:01")
require.NoError(t, err)
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.5"),
mac: mac,
hostname: "kitchen-tv",
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceMAC("de:ad:be:ef:00:01"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"source_mac_address rule must match in pre-match; the MAC has to be resolved there too")
}
func TestPreMatchResolvesNeighborHostname(t *testing.T) {
t.Parallel()
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.5"),
hostname: "kitchen-tv",
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceHostname("kitchen-tv"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"source_hostname rule must match in pre-match; the hostname has to be resolved there too")
}
// A source the neighbor resolver does not know must still fall through, not
// match on a half-filled metadata.
func TestPreMatchNeighborMissDoesNotMatch(t *testing.T) {
t.Parallel()
mac, err := net.ParseMAC("de:ad:be:ef:00:01")
require.NoError(t, err)
resolver := &stubNeighborResolver{
address: netip.MustParseAddr("192.168.1.9"),
mac: mac,
}
router := newPreMatchTestRouter(t, resolver, rejectOnSourceMAC("de:ad:be:ef:00:01"))
result := router.PreMatch(preMatchMetadata(), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action)
}
+74 -25
View File
@@ -314,27 +314,60 @@ func (r *Router) routePacketConnection(ctx context.Context, conn N.PacketConn, m
return nil
}
// lx:begin l3-honest-drop
// PreMatch funnels every verdict of the pre-match walk through one ICMP check.
//
// An ICMP flow has no fallback path, so PreMatchContinue is not "try the
// ordinary connection route" the way it is for TCP and UDP: the TUN stack takes
// the packet back and answers the echo ITSELF (sing-tun stack_gvisor_icmp.go —
// adapter.JudgeFlow maps Continue to tun.ActionAccept, and the ICMP forwarder
// answers Accept by rewriting Echo into EchoReply and swapping the addresses).
// A ping routed to an outbound that cannot carry layer 3 — every proxy
// protocol; only adapter.FlowOutbound can — would therefore return a FORGED
// reply, and the operator would read a working ping off a tunnel that never saw
// the packet. Dropping instead reports the truth.
//
// PreMatchBypass is folded into the same drop because sing-tun implements
// bypass for the nfqueue plane only (`ActionBypass` appears nowhere in
// flow_dispatch.go / stack_gvisor_icmp.go): on the TUN path it degrades to the
// same Accept, i.e. to the same forgery. There is no honest bypass for an ICMP
// packet that is already inside the engine's TUN.
//
// This is a funnel and not an override inside the walk on purpose: the walk has
// several independent exits that say "continue" (the prepareMatchMetadata error
// return, the sniff bail-outs, the un-routable `bypass`, and the default arm of
// the rule-action switch), and an earlier version of this delta guarded only
// the ones that pass through preMatchFlow — leaving the others as narrow paths
// to the forged reply. Guarding the single return value cannot be outgrown by a
// new exit.
func (r *Router) PreMatch(metadata adapter.InboundContext, firstPacket []byte) adapter.PreMatchResult {
result := r.preMatch(metadata, firstPacket)
if metadata.Network == N.NetworkICMP {
switch result.Action {
case adapter.PreMatchContinue, adapter.PreMatchBypass:
return adapter.PreMatchResult{Action: adapter.PreMatchDrop}
}
}
return result
}
// preMatch is upstream's PreMatch body, unchanged; only the name moved, so that
// the funnel above owns the exported entry point. An upstream change to the
// pre-match walk applies to THIS function.
func (r *Router) preMatch(metadata adapter.InboundContext, firstPacket []byte) adapter.PreMatchResult {
// lx:end l3-honest-drop
ctx := log.ContextWithNewID(r.ctx)
metadata.PreMatch = true
continueResult := adapter.PreMatchResult{Action: adapter.PreMatchContinue}
packetDestination := metadata.Destination
if metadata.Destination.Addr.IsValid() && r.dnsTransport.FakeIP() != nil && r.dnsTransport.FakeIP().Store().Contains(metadata.Destination.Addr) {
domain, loaded := r.dnsTransport.FakeIP().Store().Lookup(metadata.Destination.Addr)
if !loaded || domain == "" {
return continueResult
}
metadata.OriginDestination = metadata.Destination
metadata.Destination = M.Socksaddr{
Fqdn: domain,
Port: metadata.Destination.Port,
}
metadata.FakeIP = true
}
if metadata.Destination.IsIPv4() {
metadata.IPVersion = 4
} else if metadata.Destination.IsIPv6() {
metadata.IPVersion = 6
// lx: pre-match used to prepare only fakeip + IP version, so process/neighbor
// rule items (process_name, source_mac_address, source_hostname, …) never had
// their metadata filled here and silently failed to match — they were resolved
// in matchRule only. Both paths now share prepareMatchMetadata (upstream
// b911fb078).
err := r.prepareMatchMetadata(ctx, &metadata)
if err != nil {
return continueResult
}
for currentRuleIndex, currentRule := range r.rules {
metadata.ResetRuleCache()
@@ -448,6 +481,11 @@ func applyRouteOptionsOverride(metadata *adapter.InboundContext, routeOptions *R
func (r *Router) preMatchFlow(ctx context.Context, metadata *adapter.InboundContext, packetDestination M.Socksaddr, matchedRule adapter.Rule, outboundTag string) adapter.PreMatchResult {
continueResult := adapter.PreMatchResult{Action: adapter.PreMatchContinue}
// lx: ICMP does NOT get a local override here any more — the honest drop is
// applied once, to the single return value of PreMatch (see the funnel
// there, marker l3-honest-drop). Overriding continueResult in this function
// covered only the exits that reach it and left the walk's own exits
// forging.
var outbound adapter.Outbound
if outboundTag == "" {
outbound = r.outbound.Default()
@@ -540,13 +578,11 @@ func (r *Router) preMatchFlow(ctx context.Context, metadata *adapter.InboundCont
return result
}
func (r *Router) matchRule(
ctx context.Context, metadata *adapter.InboundContext,
inputConn net.Conn, inputPacketConn N.PacketConn,
) (
selectedRule adapter.Rule, selectedRuleIndex int,
buffers []*buf.Buffer, packetBuffers []*N.PacketBuffer, fatalErr error,
) {
// prepareMatchMetadata fills in everything a rule may match on but the inbound
// cannot know: the connection owner, the neighbor (MAC/hostname) behind the
// source address, the fakeip / reverse-mapped domain and the IP version. Shared
// by matchRule and PreMatch — see the note at the PreMatch call site.
func (r *Router) prepareMatchMetadata(ctx context.Context, metadata *adapter.InboundContext) error {
r.searchProcessInfo(ctx, metadata)
if r.neighborResolver != nil && metadata.SourceMACAddress == nil && metadata.Source.Addr.IsValid() {
mac, macFound := r.neighborResolver.LookupMAC(metadata.Source.Addr)
@@ -568,8 +604,7 @@ func (r *Router) matchRule(
if metadata.Destination.Addr.IsValid() && r.dnsTransport.FakeIP() != nil && r.dnsTransport.FakeIP().Store().Contains(metadata.Destination.Addr) {
domain, loaded := r.dnsTransport.FakeIP().Store().Lookup(metadata.Destination.Addr)
if !loaded {
fatalErr = E.New("missing fakeip record, try enable `experimental.cache_file`")
return
return E.New("missing fakeip record, try enable `experimental.cache_file`")
}
if domain != "" {
metadata.OriginDestination = metadata.Destination
@@ -592,6 +627,20 @@ func (r *Router) matchRule(
} else if metadata.Destination.IsIPv6() {
metadata.IPVersion = 6
}
return nil
}
func (r *Router) matchRule(
ctx context.Context, metadata *adapter.InboundContext,
inputConn net.Conn, inputPacketConn N.PacketConn,
) (
selectedRule adapter.Rule, selectedRuleIndex int,
buffers []*buf.Buffer, packetBuffers []*N.PacketBuffer, fatalErr error,
) {
fatalErr = r.prepareMatchMetadata(ctx, metadata)
if fatalErr != nil {
return
}
match:
for currentRuleIndex, currentRule := range r.rules {
+190 -22
View File
@@ -22,11 +22,16 @@
# 2. It runs on linux. shater/generate has 44 test files on linux against 32 on
# windows/darwin; the linux-only half is where the routing, ruleset, DNS and
# health tests live.
# 3. Nothing is skipped SILENTLY. Two machine checks:
# 3. Nothing is skipped SILENTLY. Three machine checks:
# - the tag set may only ADD test files, never hide them (a test behind
# `//go:build !with_awg` would vanish from the gate — this fails first);
# - every package that has tests must report `ok` by name; a suite that
# compiles down to "no test files" fails the gate instead of passing it.
# compiles down to "no test files" fails the gate instead of passing it;
# - every ^TestIntegration under the fork's trees must produce a verdict
# BY NAME ([5/5]). `ok <pkg>` is printed whether the privileged tests in
# that package ran or called t.Skip, so the second check cannot see them
# — and the gate would keep saying "passes every test we own" while the
# tests that need a real kernel never executed.
# A guard that silently runs nothing is worse than no guard (same rule as
# scripts/check-router-tags.sh).
#
@@ -38,6 +43,10 @@
# SHATER_GO_IMAGE docker image used to reach linux from a non-linux host
# (default golang:1.26 — keep it >= go.mod's toolchain).
# SHATER_NO_DOCKER=1 fail instead of falling back to docker.
# SHATER_REQUIRE_PRIVILEGED=1
# turn [5/5]'s "did not run here" report into a hard
# failure. Use it on the OpenWrt VM or in any pre-release
# run that must actually have exercised the kernel paths.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
@@ -48,7 +57,7 @@ RACE=1
for a in "$@"; do
case "$a" in
--no-race) RACE=0 ;;
-h|--help) sed -n '2,41p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
-h|--help) sed -n '2,49p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) echo "run-tests: unknown flag: $a" >&2; exit 2 ;;
esac
done
@@ -72,6 +81,12 @@ ROOTS_COMMON=(./common/...)
# rest of common/ can be a real gate instead of a permanently red one. (On
# linux every tlsspoof test is a TestIntegration*, so that package is
# effectively uncovered here; it is covered by the VM runs.)
# The ^TestIntegration prefix is the fork-wide marker for "needs capabilities
# the ordinary gate lacks", and [5/5] below leans on the same convention to
# catch privileged tests inside ROOTS, which are NOT name-filtered and would
# otherwise skip behind a green `ok <pkg>`. ROOTS_COMMON stays out of [5/5]:
# these fail rather than skip without the capability, and that is a decision
# about upstream code, not about the fork's own coverage.
SKIP_COMMON='^TestIntegration'
# SKIP, WITH REASON: the first -race run over this tree (2026-07-26 — nobody
@@ -80,15 +95,15 @@ SKIP_COMMON='^TestIntegration'
# RoutineReceiveIncoming() with no lock, caught by
# transport/wireguard.TestAwgDetourClientBindDelivers; that one has since been
# fixed in client_bind.go and is NOT skipped — it is exactly what this pass is
# for. What is left:
# - shater/alert.TestExpiryDedupWithinDay — the test's own closure
# (expiry_test.go:78) reads a variable the test body writes at :85 while
# Notifier.dispatch's goroutine is still delivering. A test-side bug, ~one
# mutex to fix, but it lives in shater/ and is nobody's blocker to ship.
# Naming it here keeps the gate a gate from day one. It is skipped ONLY in the
# -race pass — it still runs, and still has to pass, in the main pass below.
# DELETE THE ENTRY THE MOMENT THE RACE IS FIXED.
RACE_SKIP='^TestExpiryDedupWithinDay$'
# for. What was left:
# - shater/alert.TestExpiryDedupWithinDay — the test's own closure read a
# variable the test body wrote while Notifier.dispatch's goroutine was
# still delivering. FIXED 2026-07-26 (the simulated clock now has a mutex),
# so the entry is gone and the -race pass covers the whole tree again.
# Nothing is skipped under -race any more. Keep it that way: an entry here is a
# hole in the gate, so add one only with a named reason and delete it the moment
# the race is fixed.
RACE_SKIP='^$'
echo "== shater test gate =="
echo " tags : $SHATER_ROUTER_TAGS"
@@ -111,12 +126,30 @@ if [ "$(go env GOOS)" != "linux" ] && [ "${SHATER_TESTS_IN_DOCKER:-0}" != "1" ];
echo "== re-exec on linux via docker ($image) =="
host_repo="$REPO"
command -v cygpath >/dev/null 2>&1 && host_repo="$(cygpath -w "$REPO")"
# Hand the container CAP_NET_ADMIN and /dev/net/tun when this host's docker
# can. shater/generate's ^TestIntegration tests open a real TUN and stand a
# real engine on it; without the device they skip, and a dev running the gate
# by hand would get a green result that never touched the kernel path the
# branch is about. The dev host CAN give them (Docker Desktop's VM has the tun
# module) — the CI runner cannot, which is what [5/5] exists to say out loud.
# PROBED, never assumed: a docker whose kernel lacks tun refuses --device and
# would take the whole gate down with it.
priv_flags=()
if MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
--cap-add NET_ADMIN --device /dev/net/tun "$image" true >/dev/null 2>&1; then
priv_flags=(--cap-add NET_ADMIN --device /dev/net/tun)
echo " CAP_NET_ADMIN + /dev/net/tun: available — the privileged tests will really run"
else
echo " CAP_NET_ADMIN + /dev/net/tun: NOT available from this docker — [5/5] will report the gap"
fi
MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
"${priv_flags[@]+"${priv_flags[@]}"}" \
-v "$host_repo":/src \
-v shater-tagcheck-gomod:/go/pkg/mod \
-v shater-tagcheck-gocache:/root/.cache/go-build \
-w /src \
-e SHATER_TESTS_IN_DOCKER=1 \
-e SHATER_REQUIRE_PRIVILEGED="${SHATER_REQUIRE_PRIVILEGED:-0}" \
"$image" bash scripts/run-tests.sh "$@"
exit $?
fi
@@ -128,7 +161,7 @@ ALL_ROOTS=("${ROOTS[@]}" "${ROOTS_COMMON[@]}")
# shipped tags REMOVES a test file from any package, that test exists but the
# gate would never see it — which is the failure mode this whole script is about,
# just pointed the other way.
echo "== [1/4] no test file is hidden by the shipped tag set =="
echo "== [1/5] no test file is hidden by the shipped tag set =="
LISTFMT='{{.ImportPath}} {{len .TestGoFiles}} {{len .XTestGoFiles}}'
plain="$(go list -f "$LISTFMT" "${ALL_ROOTS[@]}")"
tagged="$(go list -tags "$SHATER_ROUTER_TAGS" -f "$LISTFMT" "${ALL_ROOTS[@]}")"
@@ -217,28 +250,140 @@ run_suite() { # $1=label $2=extra go-test flags (may be empty) $3..=packages
echo " OK [$label]"
}
# --- [2/4] the fork's trees, shipped tags, linux -----------------------------
echo "== [2/4] go test — the fork's trees (shipped tags, linux) =="
# --- [2/5] the fork's trees, shipped tags, linux -----------------------------
echo "== [2/5] go test — the fork's trees (shipped tags, linux) =="
run_suite main "" "${ROOTS[@]}"
echo
# --- [3/4] common/, minus the tests that need CAP_NET_ADMIN ------------------
echo "== [3/4] go test — common/ (minus the CAP_NET_ADMIN integration tests) =="
# --- [3/5] common/, minus the tests that need CAP_NET_ADMIN ------------------
echo "== [3/5] go test — common/ (minus the CAP_NET_ADMIN integration tests) =="
run_suite common "-skip $SKIP_COMMON" "${ROOTS_COMMON[@]}"
echo
# --- [4/4] -race over the same trees -----------------------------------------
# --- [4/5] -race over the same trees -----------------------------------------
# Everything, not a subset: shater/netplane alone is ~110 s under -race and it is
# the single most concurrency-critical package we own (the nft data plane), so
# once it is in, adding the rest costs ~40 s more. common/ is left out — it is
# upstream code exercised by upstream CI.
if [ "$RACE" -eq 1 ]; then
echo "== [4/4] go test -race — the fork's trees =="
echo " known-red under -race, skipped BY NAME (fix it and delete from RACE_SKIP):"
echo " TestExpiryDedupWithinDay shater/alert (test-side race, expiry_test.go:78/85)"
echo "== [4/5] go test -race — the fork's trees =="
echo " nothing is skipped under -race"
run_suite race "-race -skip $RACE_SKIP" "${ROOTS[@]}"
else
echo "== [4/4] -race pass skipped (--no-race) =="
echo "== [4/5] -race pass skipped (--no-race) =="
fi
echo
# --- [5/5] the privileged tests may not skip in silence ----------------------
# THE HOLE THIS CLOSES. Some tests can only prove what they claim against a real
# kernel: shater/generate's TestIntegrationL3TunInboundStarts opens /dev/net/tun
# and stands a real engine on it, TestIntegrationL3EgressICMPIsAFlow binds a real
# socket to a real device. Both guard themselves with t.Skip when root or the
# device is missing — the honest thing for a test to do, and completely INVISIBLE
# above: `go test` prints `ok <pkg>` whether they ran or skipped, so [2/5]'s
# per-package `ok` check is satisfied either way and the gate closes by claiming
# it "passes every test we own". That is precisely the failure this whole script
# was written for (115 of 116 test files never running while CI stayed green),
# one level down and harder to see.
#
# The list is DISCOVERED, not hand-kept — `go test -list` over the same ROOTS —
# so a privileged test written next month joins this check on the day it is
# named, with no edit here. It keys on the ^TestIntegration prefix, already this
# fork's marker for "needs capabilities the ordinary gate lacks" (SKIP_COMMON
# above excludes common/tlsspoof's TestIntegration* for exactly that reason).
# Name a privileged test anything else and it is invisible again — so don't.
#
# Verdicts, per test, by name:
# RAN — it executed here; printed so that is visible rather than assumed.
# FAILED — fatal, like any other failure.
# MISSING — `go test -list` named it and the run produced no verdict for it:
# fatal. A test that vanished between listing and running is the
# same class of hole as one hidden by a build tag.
# SKIPPED while this environment HAS root and /dev/net/tun — fatal. The
# capability guard cannot be what skipped it, so something else did
# and only the test knows what.
# SKIPPED because the environment genuinely cannot run it — reported loudly,
# by name, and it REPLACES the closing banner, so the last line of
# the gate can never claim coverage it does not have. Deliberately
# not fatal by default: the act_runner is an LXC guest whose kernel
# has no tun module at all (checked 2026-07-26 on 10.10.10.211 —
# `modprobe tun` answers "Module tun not found", /dev/net does not
# exist, and act_runner runs job containers with privileged:false
# and no container.options), so the device cannot be handed down
# without reconfiguring the Proxmox host. Making it fatal would
# paint CI permanently red and teach everyone to ignore the gate.
# SHATER_REQUIRE_PRIVILEGED=1 makes it fatal for the runs that can.
echo "== [5/5] the privileged tests (^TestIntegration) produced a verdict by name =="
PRIV_RE='^TestIntegration'
PRIV_UNVERIFIED=""
# -ldflags is NOT optional on the discovery call either: `go test -list` LINKS
# each test binary before it can enumerate its tests, and without
# -checklinkname=0 every package that pulls common/badtls fails to link. The
# first cut of this step omitted it, swallowed the error with `2>/dev/null ||
# true`, and reported "none declared" — a check against silent skipping that was
# itself silently skipping. Hence also: the exit status is inspected, and an
# empty list is only ever reported after a SUCCESSFUL enumeration.
set +e
priv_expect_raw="$(go test -list "$PRIV_RE" \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" "${ROOTS[@]}" 2>&1)"
priv_list_rc=$?
set -e
priv_expect="$(grep -E "$PRIV_RE" <<<"$priv_expect_raw" | sort -u || true)"
if [ "$priv_list_rc" -ne 0 ]; then
echo " FAILED [privileged]: could not enumerate the privileged tests (go test -list exited $priv_list_rc)." >&2
echo " An unreadable list is NOT an empty list — this check refuses to" >&2
echo " report 'nothing to verify' on the strength of a failed command." >&2
sed 's/^/ /' <<<"$priv_expect_raw" | grep -vE '^\s+(ok|\?)\s' >&2 || true
FAILED=1
elif [ -z "$priv_expect" ]; then
echo " none declared under the fork's trees — nothing to verify"
else
priv_capable=0
if [ "$(id -u)" = "0" ] && [ -e /dev/net/tun ]; then
priv_capable=1
fi
echo " declared: $(wc -l <<<"$priv_expect" | tr -d ' ')"
echo " this environment: uid=$(id -u), /dev/net/tun $([ -e /dev/net/tun ] && echo present || echo MISSING) => can run them: $([ "$priv_capable" -eq 1 ] && echo yes || echo NO)"
set +e
go test -count=1 -v -run "$PRIV_RE" \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
"${ROOTS[@]}" >"$LOG" 2>&1
priv_rc=$?
set -e
# The verdict lines plus whatever reason the test printed just before them,
# so a skip is readable here and not just counted.
grep -E '^(--- (PASS|SKIP|FAIL): |[[:space:]]+[^[:space:]]+\.go:[0-9]+: )' "$LOG" \
| sed 's/^/ | /' || true
priv_bad=0
while read -r name; do
[ -n "$name" ] || continue
if grep -qE "^--- PASS: ${name}([[:space:]]|\$)" "$LOG"; then
echo " RAN $name"
elif grep -qE "^--- FAIL: ${name}([[:space:]]|\$)" "$LOG"; then
echo " FAILED $name" >&2
priv_bad=1
elif grep -qE "^--- SKIP: ${name}([[:space:]]|\$)" "$LOG"; then
if [ "$priv_capable" -eq 1 ]; then
echo " SKIPPED $name — but this environment HAS root and /dev/net/tun, so the capability guard is NOT what skipped it" >&2
priv_bad=1
else
echo " DID NOT RUN $name — skipped: no root and/or no /dev/net/tun here"
PRIV_UNVERIFIED="$PRIV_UNVERIFIED $name"
fi
else
echo " MISSING $name — go test -list named it, the run produced no verdict for it" >&2
priv_bad=1
fi
done <<<"$priv_expect"
if [ "$priv_rc" -ne 0 ] && [ "$priv_bad" -eq 0 ]; then
echo " FAILED [privileged]: go test exited $priv_rc with every named test accounted for —" >&2
echo " a build or package-level failure, see the log above." >&2
priv_bad=1
fi
if [ "$priv_bad" -ne 0 ]; then
FAILED=1
fi
fi
echo
@@ -246,4 +391,27 @@ if [ "$FAILED" -ne 0 ]; then
echo "== TEST GATE FAILED — nothing may be published from this run. ==" >&2
exit 1
fi
if [ -n "$PRIV_UNVERIFIED" ]; then
echo "== !! PASSED, BUT NOT FULLY VERIFIED !! =================================="
echo " Every test that COULD run here passed. These did not run at all:"
for t in $PRIV_UNVERIFIED; do
echo " - $t"
done
echo
echo " They need root + CAP_NET_ADMIN + /dev/net/tun, which this environment"
echo " does not have. Nothing about the kernel paths they cover was verified"
echo " by this run. To actually run them, from a host whose docker can:"
echo " scripts/run-tests.sh # the re-exec hands the container both"
echo " or directly:"
echo " docker run --rm --cap-add NET_ADMIN --device /dev/net/tun \\"
echo " -v \"\$PWD\":/src -w /src golang:1.26 bash scripts/run-tests.sh"
echo " or on the OpenWrt VM. SHATER_REQUIRE_PRIVILEGED=1 makes this a hard"
echo " failure instead of this notice."
echo "=========================================================================="
if [ "${SHATER_REQUIRE_PRIVILEGED:-0}" = "1" ]; then
echo "== TEST GATE FAILED: SHATER_REQUIRE_PRIVILEGED=1 and the tests above did not run. ==" >&2
exit 1
fi
exit 0
fi
echo "== OK: the shipped tag set, on linux, passes every test we own. =="
+141
View File
@@ -0,0 +1,141 @@
package alert
import (
"fmt"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
)
// warnLogger records Warn lines so a test can assert that an eviction was
// actually announced. Everything else falls through to the standard logger.
type warnLogger struct {
log.ContextLogger
mu sync.Mutex
warns []string
}
func newWarnLogger() *warnLogger { return &warnLogger{ContextLogger: log.StdLogger()} }
func (l *warnLogger) Warn(args ...any) {
l.mu.Lock()
l.warns = append(l.warns, fmt.Sprint(args...))
l.mu.Unlock()
}
func (l *warnLogger) lines() []string {
l.mu.Lock()
defer l.mu.Unlock()
return append([]string(nil), l.warns...)
}
// notifierAt builds a Notifier with no channels (so nothing is ever delivered —
// only the dedup bookkeeping runs) and a clock the test drives.
func notifierAt(t *testing.T, clock *time.Time) (*Notifier, *warnLogger) {
t.Helper()
lg := newWarnLogger()
n := New([]model.Alert{}, lg)
n.now = func() time.Time { return *clock }
return n, lg
}
func (n *Notifier) dedupSize() int {
n.mu.Lock()
defer n.mu.Unlock()
return len(n.dedup)
}
// TestDedupTableIsBoundedOverTime is the leak itself: a guest network whose
// clients randomise their MAC produces an endless stream of distinct
// "new_device:<MAC>" keys, and nothing ever deleted one. Spread over time — which
// is how it actually happens — the table must stay small, and nothing may be
// reported as evicted, because an entry past the dedup window could no longer
// suppress anything anyway.
func TestDedupTableIsBoundedOverTime(t *testing.T) {
clock := time.Now()
n, lg := notifierAt(t, &clock)
// Ten times the cap, at one incident per second: every key is long past the
// 60s window by the time the next batch arrives.
const fires = maxDedupKeys * 10
for i := 0; i < fires; i++ {
clock = clock.Add(time.Second)
n.FireIncident(Incident{
Events: []string{"new_device"},
Title: "New device on the LAN",
Key: fmt.Sprintf("new_device:02:00:00:%02x:%02x:%02x", i>>16&0xff, i>>8&0xff, i&0xff),
})
}
if got := n.dedupSize(); got > maxDedupKeys {
t.Fatalf("dedup table holds %d entries after %d distinct incidents; cap is %d",
got, fires, maxDedupKeys)
}
if got := n.DedupEvicted(); got != 0 {
t.Fatalf("reported %d LIVE evictions; entries aged out of the window and losing them costs nothing", got)
}
if lines := lg.lines(); len(lines) != 0 {
t.Fatalf("expiry sweep must be silent (it loses nothing), got: %v", lines)
}
}
// TestDedupTableEvictionIsAnnounced is the other half: when the cap genuinely
// bites — more distinct incidents inside ONE dedup window than the table holds —
// live suppression state is lost and repeats may notify twice. That must be said
// out loud, not absorbed.
func TestDedupTableEvictionIsAnnounced(t *testing.T) {
clock := time.Now()
n, lg := notifierAt(t, &clock)
// The clock does not move: every key stays inside its window.
for i := 0; i < maxDedupKeys+10; i++ {
clock = clock.Add(time.Millisecond) // still far inside dedupWindow
n.FireIncident(Incident{
Events: []string{"new_device"},
Title: "New device on the LAN",
Key: fmt.Sprintf("new_device:flood-%d", i),
})
}
if got := n.dedupSize(); got > maxDedupKeys {
t.Fatalf("dedup table grew to %d, above the %d cap", got, maxDedupKeys)
}
if n.DedupEvicted() == 0 {
t.Fatalf("a flood of %d in-window incidents evicted nothing — the cap is not enforced", maxDedupKeys+10)
}
lines := lg.lines()
if len(lines) == 0 {
t.Fatalf("live suppression entries were dropped with no notice")
}
if !strings.Contains(lines[0], "dedup table full") || !strings.Contains(lines[0], "notify twice") {
t.Fatalf("eviction notice does not explain the consequence: %q", lines[0])
}
}
// TestDedupStillSuppressesWithinTheWindow guards the behaviour the bound must not
// break: a repeat inside the window is still collapsed, and one outside it is not.
func TestDedupStillSuppressesWithinTheWindow(t *testing.T) {
clock := time.Now()
n, _ := notifierAt(t, &clock)
fire := func() bool {
n.mu.Lock()
defer n.mu.Unlock()
return n.suppressedLocked("killswitch\x00Kill-switch engaged")
}
if fire() {
t.Fatalf("first fire was suppressed")
}
clock = clock.Add(dedupWindow / 2)
if !fire() {
t.Fatalf("a repeat inside the window was NOT suppressed")
}
clock = clock.Add(dedupWindow)
if fire() {
t.Fatalf("a repeat past the window was suppressed")
}
}
+87
View File
@@ -0,0 +1,87 @@
package alert
import (
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// closingRT is an http.RoundTripper that also implements the CloseIdleConnections
// hook http.Client forwards to, so a test can observe whether the client was ever
// released. http.Transport implements the same hook — this stands in for it.
type closingRT struct {
rt http.RoundTripper
closed atomic.Int32
}
func (c *closingRT) RoundTrip(r *http.Request) (*http.Response, error) { return c.rt.RoundTrip(r) }
func (c *closingRT) CloseIdleConnections() { c.closed.Add(1) }
// TestDetourClientIsClosedAfterDelivery: the detour factory (engine.HTTPClient)
// mints a NEW http.Transport for every call, and that transport's idle connections
// are live proxying sessions through an engine outbound whose object the dial
// closure pins. The notifier used one per delivery and dropped it, so every alert
// left a keep-alive session — and a reference to a possibly-retired engine
// generation — behind for the whole idle timeout.
func TestDetourClientIsClosedAfterDelivery(t *testing.T) {
c, srv := newSink(t)
n := New([]model.Alert{{
Name: "hook", Enabled: true, Type: "webhook", URL: srv.URL,
Events: []string{"killswitch"}, Via: "node:tunnel",
}}, nil)
var made []*closingRT
n.SetClientFactory(func(via string) (*http.Client, error) {
rt := &closingRT{rt: http.DefaultTransport}
made = append(made, rt)
return &http.Client{Transport: rt}, nil
})
n.Fire("killswitch", "Kill-switch engaged", "the tunnel is down")
n.Wait()
if c.n() != 1 {
t.Fatalf("delivery count = %d, want 1", c.n())
}
if len(made) != 1 {
t.Fatalf("factory called %d times, want 1", len(made))
}
if got := made[0].closed.Load(); got == 0 {
t.Fatalf("the per-delivery detour client was never closed — its idle connections " +
"(and the engine outbound its dialer pins) outlive the alert")
}
}
// TestDetourClientIsClosedWhenTheSendFails covers the fallback path: a detour that
// errors and falls back to direct still built a transport, and that one leaked too.
func TestDetourClientIsClosedWhenTheSendFails(t *testing.T) {
// A sink that rejects, so the detour send fails and Fallback kicks in.
reject := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
}))
defer reject.Close()
n := New([]model.Alert{{
Name: "hook", Enabled: true, Type: "webhook", URL: reject.URL,
Events: []string{"killswitch"}, Via: "node:tunnel", Fallback: true,
}}, nil)
var rt *closingRT
n.SetClientFactory(func(via string) (*http.Client, error) {
rt = &closingRT{rt: http.DefaultTransport}
return &http.Client{Transport: rt}, nil
})
n.Fire("killswitch", "Kill-switch engaged", "the tunnel is down")
n.Wait()
if rt == nil {
t.Fatalf("factory was never called")
}
if got := rt.closed.Load(); got == 0 {
t.Fatalf("a failed detour delivery still leaked its transport")
}
}
+21 -5
View File
@@ -5,6 +5,7 @@ import (
"os"
"path/filepath"
"strings"
"sync"
"testing"
"time"
@@ -74,22 +75,37 @@ func TestExpiryDedupWithinDay(t *testing.T) {
// test compresses 24 simulated hours into a few milliseconds of wall clock — so
// without this the second delivery is (correctly) suppressed by that window and
// the test would be measuring the fixture, not the behaviour.
// The clock is shared with the notifier's DELIVERY goroutines (buildPayload
// stamps the payload with n.now()), which are still in flight while this body
// advances the simulated time — so it needs a lock, not a bare variable.
var clockMu sync.Mutex
simNow := now
e.n.now = func() time.Time { return simNow }
setNow := func(t time.Time) {
clockMu.Lock()
simNow = t
clockMu.Unlock()
}
e.n.now = func() time.Time {
clockMu.Lock()
defer clockMu.Unlock()
return simNow
}
if got := e.Run(m, now); got != 1 {
t.Fatalf("first Run fired %d, want 1", got)
}
// Simulate a minute-by-minute reconcile for the next 23 hours.
for i := 1; i <= 23; i++ {
simNow = now.Add(time.Duration(i) * time.Hour)
if got := e.Run(m, simNow); got != 0 {
at := now.Add(time.Duration(i) * time.Hour)
setNow(at)
if got := e.Run(m, at); got != 0 {
t.Fatalf("Run at +%dh fired %d alerts, want 0 (dedup window is 24h)", i, got)
}
}
// Just past the window it may speak again.
simNow = now.Add(24*time.Hour + time.Minute)
if got := e.Run(m, simNow); got != 1 {
past := now.Add(24*time.Hour + time.Minute)
setNow(past)
if got := e.Run(m, past); got != 1 {
t.Errorf("Run just past 24h fired %d, want 1", got)
}
e.n.Wait()
+99
View File
@@ -20,6 +20,7 @@ import (
"fmt"
"io"
"net/http"
"sort"
"strings"
"sync"
"time"
@@ -31,6 +32,29 @@ import (
const (
httpTimeout = 8 * time.Second
dedupWindow = 60 * time.Second
// maxDedupKeys / keepDedupKeys bound the dedup table.
//
// Nothing ever deleted from it. Every key that had EVER fired stayed forever,
// and its highest-cardinality producer is "new_device:<MAC>" — on a guest
// network where clients randomise their MAC per association, that is a fresh
// key per device per join, for the life of a daemon that runs for months.
//
// The real bound is not this cap, it is the window: an entry older than
// dedupWindow (60s) can never suppress anything again, so it is pure garbage
// and sweeping it costs nothing and changes no behaviour. compactDedupLocked
// does that first. The cap only bites when 1024 DISTINCT incidents fired inside
// one 60-second window — and there the eviction is a real loss of suppression
// state, so it is reported rather than done quietly.
//
// Why 1024: the new_device watcher polls every ~45s, so a poll would have to
// discover a thousand previously-unseen MACs at once to reach it — roughly 20x
// the worst guest-network churn this box has seen. At ~100 B per entry (a
// 28-byte "new_device:<MAC>" key plus a time.Time plus map overhead) the full
// table is ~100 KB of a 512 MB router: cheap enough that a generous headroom
// costs nothing.
maxDedupKeys = 1024
keepDedupKeys = 512
)
// telegramAPIBase is the Telegram Bot API root. A package var so tests can point
@@ -43,6 +67,10 @@ type Notifier struct {
mu sync.Mutex
alerts []model.Alert
dedup map[string]time.Time // (event\x00title) -> last fire time
// dedupEvicted counts LIVE dedup entries dropped by the capacity bound (see
// compactDedupLocked). Expired entries swept out are NOT counted: they could no
// longer suppress anything, so dropping them loses nothing.
dedupEvicted uint64
client *http.Client
log log.ContextLogger
@@ -227,10 +255,69 @@ func (n *Notifier) suppressedLocked(key string) bool {
if last, ok := n.dedup[key]; ok && now.Sub(last) < dedupWindow {
return true
}
if len(n.dedup) >= maxDedupKeys {
n.compactDedupLocked(now)
}
n.dedup[key] = now
return false
}
// compactDedupLocked bounds the dedup table. Caller holds n.mu.
//
// Two stages, deliberately distinct because only one of them loses anything:
//
// 1. Drop every entry older than dedupWindow. Such an entry cannot suppress a
// future fire (suppressedLocked already ignores it), so this is garbage
// collection, not eviction: no notification changes, and nothing is reported.
// On any realistic traffic this stage alone keeps the table at "keys seen in
// the last minute".
// 2. If the table is STILL full, a thousand distinct incidents fired inside one
// window. Now eviction is real — the oldest live entries go, and a repeat of
// one of them within its window will notify a second time instead of being
// collapsed. That is a visible change in behaviour, so it is logged, with the
// running total, rather than silently absorbed.
func (n *Notifier) compactDedupLocked(now time.Time) {
for k, last := range n.dedup {
if now.Sub(last) >= dedupWindow {
delete(n.dedup, k)
}
}
if len(n.dedup) < maxDedupKeys {
return
}
type kv struct {
key string
at time.Time
}
all := make([]kv, 0, len(n.dedup))
for k, at := range n.dedup {
all = append(all, kv{k, at})
}
sort.Slice(all, func(i, j int) bool { return all[i].at.Before(all[j].at) })
drop := len(all) - keepDedupKeys
for i := 0; i < drop; i++ {
delete(n.dedup, all[i].key)
}
n.dedupEvicted += uint64(drop)
n.log.Warn("alert: dedup table full (", maxDedupKeys,
" incidents inside one ", dedupWindow, " window) — dropped ", drop,
" live suppression entries (", n.dedupEvicted,
" total); repeats of those incidents may notify twice")
}
// DedupEvicted reports how many LIVE suppression entries the cap has dropped since
// start (stage 2 of compactDedupLocked only — the expiry sweep is not counted,
// because it loses nothing). Nonzero means alerts may have been delivered twice.
func (n *Notifier) DedupEvicted() uint64 {
if n == nil {
return 0
}
n.mu.Lock()
defer n.mu.Unlock()
return n.dedupEvicted
}
// dispatch delivers to a single alert in its own goroutine. Panics are recovered
// and logged; a delivery error is logged. It never crashes the daemon.
func (n *Notifier) dispatch(a model.Alert, event, title, body string) {
@@ -276,6 +363,18 @@ func (n *Notifier) deliver(a model.Alert, event, title, body string) error {
if detour && factory != nil {
client, cerr := factory(via)
if cerr == nil && client != nil {
// The factory (engine.HTTPClient) builds a BRAND NEW http.Transport per
// call, with keep-alive and a 90s idle timeout, and we use it for exactly
// one POST. Dropping it without this leaves the idle connection — a real
// proxying session through an engine outbound, plus its read and write
// loops — alive for the whole idle timeout. Worse, the transport's
// DialContext closure captures that outbound object, so the idle
// connection PINS a retired engine generation whose close budget is 5
// seconds. One alert delivery per minute keeps a permanent rolling set of
// them.
defer client.CloseIdleConnections()
}
if cerr != nil {
if a.Fallback {
n.log.Warn("alert: ", a.Name, " detour ", via, " unavailable (", cerr, ") — falling back to direct")
+219 -22
View File
@@ -61,10 +61,16 @@ type Applier struct {
// silently skipped. "" until the first successful nft load.
lastNft string
// holding is true while the fail-closed HOLDING PLANE is installed: the engine
// holding is the LATCH half of the hold state: true while a fail-closed HOLDING
// PLANE that THIS Applier installed (holdLocked) is in the kernel — the engine
// is down and LAN->WAN forwarding is blocked. Surfaced in Status so the panel
// can say "protected but not proxying" instead of showing a healthy-looking UI
// over an unprotected router.
//
// It is deliberately NOT the whole answer: a plane installed by the boot armor
// or left by a predecessor blocks the LAN just as hard and never touches this
// field. Read it through Holding()/holdingWith(), which add the derived half;
// writing an inference into this latch would create a claim nothing clears.
// holding + lastWarnings live behind their OWN lock, NOT the apply mutex.
// Status() reads them, and Status is polled by the LuCI dashboard, the cron
// watchdog and hotplug; making those reads queue behind a running apply
@@ -342,6 +348,19 @@ func (a *Applier) UpdateSubscription(name string) (added int, err error) {
if err != nil {
return 0, fmt.Errorf("subscription %q detour %q: %w", name, sub.FetchDetour, err)
}
// engine.HTTPClient builds a FRESH http.Transport per call, and its dialer
// closes over the resolved outbound. Dropping the client on the floor leaves
// that transport's idle keep-alives open — those are real proxied sessions
// through the engine, a goroutine pair each — and the closure PINS the engine
// generation they were dialled on. A generation's shutdown budget is five
// seconds; an idle connection outlives it, which is how a retired box stays
// alive with its WireGuard devices still held.
//
// Deferred rather than closed at the end on purpose: every exit from here is
// covered, including the four error returns below, and those are the ones that
// leak in the field — a subscription whose feed is broken is retried by cron
// once a refresh interval, forever.
defer client.CloseIdleConnections()
}
// FetchWithInfo additionally parses the provider's `subscription-userinfo`
@@ -542,14 +561,14 @@ func (a *Applier) applyDataPlaneLocked(m *model.Model, opts option.Options, now
out.stage = "rendering the nft ruleset"
return out, err
}
nftCurrent := ruleset == a.lastNft && netplane.TableExists()
nftCurrent := ruleset == a.lastNft && tableExists()
if !nftCurrent {
if err := netplane.ApplyNft(ruleset); err != nil {
// The engine may be up, but with no table loaded nothing is diverted into
// it — LAN traffic goes straight out the WAN. That is the same silent
// fail-open as a dead engine, so it gets the same answer: if there is no
// table at all, hold the line rather than leave the LAN exposed.
if !netplane.TableExists() {
if !tableExists() {
a.holdLocked(m, err)
}
out.stage = "loading the nft ruleset"
@@ -726,10 +745,35 @@ func killSwitchClosed(g model.Globals) bool {
// attempted, so the window between daemon start and a successful engine start is
// protected rather than open. The daemon calls it at startup.
//
// It is a no-op when the plane is disabled, the kill-switch is open, or a table
// is already loaded. A successful apply replaces the holding plane atomically
// (the full ruleset is a single `delete table` + `table` transaction), so the
// cost on a healthy boot is a few seconds of blocked forwarding — which is
// It is a no-op when the plane is disabled or the kill-switch is open: fail-open
// is the operator's documented choice and must not be quietly overridden.
//
// It is NO LONGER a no-op when a table is already loaded, and that is the change.
// The early return on netplane.TableExists() cost two things:
//
// - It deferred to a table this process cannot inspect. Since the boot armor
// landed, the table found here is usually /etc/init.d/shater-armor's snapshot
// — which may have been rendered before an interface rename and therefore
// protects a device that no longer exists. cmd/shaterd's own armorRender
// preference says a fresh render beats a snapshot for exactly this reason.
// - It skipped setHolding entirely, so with the boot armor loaded the daemon
// reported holding=false over a LAN that really was cut off. Status then
// published plane="full" and the apply-failure alert (cmd/shaterd
// fireApplyFail) told the operator traffic was NOT being blocked at the very
// moment it was — sending them to dismantle a protection that was working.
//
// Rather than adopt a table it cannot vouch for, ArmHold installs its own,
// rendered from the model it was handed: what it then publishes is a fact about
// something this process did, not a guess about something it found. That is also
// why the fix is not simply "call setHolding(true) when a table is present" — see
// foreignHold for the case where installing is impossible and a claim has to be
// derived instead.
//
// Replacing a loaded table costs nothing in exposure: `nft -f` commits as ONE
// netlink transaction (netplane.runNftStdin), so there is no window between the
// delete and the new table, and a ruleset that fails to load leaves the previous
// one in place. A successful apply then replaces the holding plane the same way,
// so the cost on a healthy boot is a few seconds of blocked forwarding — which is
// precisely what fail-closed is supposed to mean.
func (a *Applier) ArmHold(m *model.Model) {
if m == nil || !m.Globals.Enabled || !killSwitchClosed(m.Globals) {
@@ -742,9 +786,6 @@ func (a *Applier) ArmHold(m *model.Model) {
defer release()
a.mu.Lock()
defer a.mu.Unlock()
if netplane.TableExists() {
return
}
a.holdLocked(m, errEngineNotStartedYet)
}
@@ -755,13 +796,107 @@ var errEngineNotStartedYet = errors.New("engine has not started yet")
// running nft.
var applyHoldNft = netplane.ApplyNft
// Holding reports whether the fail-closed holding plane is currently installed
// (engine down, LAN->WAN blocked). Takes only the leaf lock, so it never waits
// on a running apply.
// tableExists and bootArmorPresent are the two OUTSIDE-WORLD facts this file's
// honesty now rests on: is our table in the kernel, and does the persisted
// fail-closed plane exist on flash. Seams for the same reason as engineApply and
// applyDataPlane — what the daemon SAYS about a plane it did not install can
// otherwise only be exercised on a router, with root, a real nft and a real
// /etc/shater, i.e. never, in the gate. Production reads netplane.
var (
tableExists = netplane.TableExists
bootArmorPresent = netplane.BootArmorPresent
// teardownNft is a seam for the same reason: whether the table is REMOVED or
// REPLACED on the way out is the whole of the restart-gap fix, and it can
// otherwise only be observed on a router with a real nft.
teardownNft = netplane.TeardownNft
)
// Holding reports whether forwarded LAN traffic is currently being BLOCKED by a
// fail-closed plane while the engine is down. It never waits on a running apply.
//
// It has two sources, and the second one is the whole point.
//
// The LATCH (a.holding) is written by holdLocked, i.e. when THIS process
// installed the plane. It used to be the only source, which quietly made this a
// record of what the process had DONE rather than a description of the router.
// The boot armor (netplane/armor.go) opened the gap: /etc/init.d/shater-armor
// loads the persisted holding plane at START=21, long before the daemon exists,
// and the daemon's own unreadable-config path (cmd/shaterd armOnUnreadableConfig)
// reinstates it without going anywhere near the apply pipeline. In both cases the
// LAN really is cut off and the latch reads false. With an unreadable config
// NOTHING ever corrects it, because the correction only happens on an apply and
// no apply can run: Status publishes plane="full" over a blocked LAN, and
// fireApplyFail words its incident as "traffic is NOT being blocked" at the exact
// moment it is. That is the inverted lie — it does not hide a fault, it invents
// one, and the obvious response to it is to tear down the protection that works.
//
// The DERIVED half (foreignHold) closes that, and it self-clears: it is computed
// at read time from the kernel and the engine, so it goes false the instant
// either fact changes, whereas a latch set from an inference would have to be
// remembered to be cleared. Same reasoning as Warnings' live half.
func (a *Applier) Holding() bool {
return a.holdingWith(tableExists(), a.eng != nil && a.eng.Running())
}
// holdingWith is Holding against facts the caller has already established, so
// Status does not shell out to `nft list table` twice for one poll.
func (a *Applier) holdingWith(tableLoaded, engineUp bool) bool {
a.stateMu.RLock()
defer a.stateMu.RUnlock()
return a.holding
latched := a.holding
a.stateMu.RUnlock()
return latched || foreignHold(tableLoaded, engineUp)
}
// foreignHold answers: is a plane THIS PROCESS DID NOT INSTALL holding the LAN?
//
// It deliberately does not try to identify the loaded table, because it cannot.
// netplane exposes no read-back of the loaded ruleset, `#` comments do not
// survive `nft -f`, and the two candidates — our holding plane, or a full ruleset
// left behind by a generation that died without tearing down — are not
// distinguishable from here. Guessing which one it is would be a second lie
// inside the fix for the first.
//
// It does not need to. What `holding` asserts is not "the holding plane is the
// object in the kernel", it is "the engine is down and forwarded traffic is being
// dropped" — and BOTH candidates do that, provided they were rendered from an
// enabled, fail-closed config:
//
// - the holding plane drops by construction: netplane.RenderHoldNftAt ends in
// `<iif> meta nfproto ipv4 drop` and its v6 twin;
// - a leftover FULL ruleset drops too, by design. With no engine socket the
// `tproxy` statement returns NFT_BREAK, which aborts its own rule before the
// trailing `meta mark set ... accept`, so the packet reaches the forward chain
// unmarked and meets the PRIMARY FAIL-CLOSED DROP there (netplane/nft.go).
// That is the kill-switch leak this project fixed in the forward chain, and
// netplane/armor.go restates the property in prose.
//
// So the question that actually decides the claim is whether the last config this
// router applied was enabled AND fail-closed — and the boot armor's PRESENCE is
// exactly that fact, by its own contract ("the file IS the arm token",
// netplane/armor.go): written after every successful config read while enabled
// with the kill switch closed, removed the moment either stops being true. A
// leftover full plane from a kill_switch=open config — the one that really does
// drop nothing — therefore reads FALSE here, which is the answer the operator
// needs and the case condition (1) of this fix exists to protect.
//
// engineUp must be false. With the engine up, a loaded table is the working full
// plane doing its job and nothing is being held; this is also what keeps a
// successful apply's setHolding(false) from being undone one line later.
//
// Two residual gaps, named rather than papered over:
//
// - a full plane rendered under kill_switch=open, still loaded after the
// operator closed the switch and the armor was rewritten, reads as holding
// while it is not. Any reconcile closes it, and cron runs one a minute.
// - a boot armor whose device set predates an interface rename protects the old
// name. That is a property of the snapshot, not of this predicate, and it is
// why ArmHold now replaces it with a fresh render the moment this daemon has a
// readable model.
func foreignHold(tableLoaded, engineUp bool) bool {
if !tableLoaded || engineUp {
return false
}
return bootArmorPresent()
}
func (a *Applier) setHolding(v bool) {
@@ -836,7 +971,42 @@ func (a *Applier) Reconcile() (changed bool, err error) {
// Teardown is the honest teardown: engine.Close + netplane routing/nft teardown +
// clear ACTIVE_FLAG, under the flock. Safe to call when nothing is up (idempotent).
func (a *Applier) Teardown() error {
//
// This is the "everything goes" form — the operator disabled the stack or stopped
// the service. A process that is being REPLACED wants TeardownExiting instead.
func (a *Applier) Teardown() error { return a.teardown(nil) }
// TeardownExiting is Teardown for a daemon that is going away, with the one
// ordering that never leaves the LAN uncovered.
//
// THE GAP THIS CLOSES (MEASURED on the stand: 80-90 ms, twice). The exit path used
// to be `Teardown(); armOnExit()` — TeardownNft DELETED the table, and only then
// was the fail-closed holding plane installed. Two nft transactions, and between
// them the `inet shater` table does not exist at all, so fw4's `lan -> wan ACCEPT`
// is the only policy on the box and the whole LAN forwards in the clear. That is
// not a boot-time window: it is every `restart`, every `reload_service` (i.e.
// every LuCI Save & Apply) and every package upgrade. The width is two `nft`
// invocations, so it does NOT grow with the engine — eng.Close runs before the
// table is touched — but it is the whole LAN, in the clear, every time.
//
// The comment that used to sit on the call site — "AFTER the teardown, never
// before: Teardown deletes the table, so a plane installed first would simply be
// removed again" — described the mechanism correctly and drew the wrong conclusion
// from it: the answer is not to arm later, it is to stop deleting.
//
// So arm FIRST and then skip the delete. RenderHoldNft's output is a single
// `nft -f` script that opens with `table inet shater` / `delete table inet shater`
// / `table inet shater { ... }` — one netlink transaction, in which the table is
// REPLACED rather than removed and re-added. The kernel never observes its
// absence, so a sampler cannot either.
//
// arm reports whether it actually installed a plane. When it did NOT — a real
// `stop`, or a handoff with kill_switch=open, where fail-open is the operator's
// documented choice — the table is removed exactly as before. "Keep the table"
// therefore follows from "a plane is standing", never from the caller's intent.
func (a *Applier) TeardownExiting(arm func() bool) error { return a.teardown(arm) }
func (a *Applier) teardown(arm func() bool) error {
release, err := lockExclusive()
if err != nil {
return err
@@ -845,6 +1015,15 @@ func (a *Applier) Teardown() error {
a.mu.Lock()
defer a.mu.Unlock()
// Install the successor plane BEFORE anything is dismantled, and do it while
// still holding the apply lock, so a concurrent apply cannot slip between the
// swap and the teardown. arm must not call back into the Applier (cmd/shaterd's
// armOnExit goes straight to model + netplane) or this deadlocks.
kept := false
if arm != nil {
kept = arm()
}
// Stop the observatory BEFORE closing the engine, and wait for an in-flight
// tick: a teardown must not leave probe dials racing a box that is going away.
a.eng.StopObservatory()
@@ -867,8 +1046,14 @@ func (a *Applier) Teardown() error {
if err := netplane.TeardownRouting(m); err != nil && firstErr == nil {
firstErr = err
}
if err := netplane.TeardownNft(); err != nil && firstErr == nil {
firstErr = err
// The policy routing above is safe to remove either way: the holding plane is a
// single `forward` chain of accepts and drops and consults no routing table, so
// it keeps working with the ip rules gone. The TABLE is the one thing that must
// not be removed out from under it.
if !kept {
if err := teardownNft(); err != nil && firstErr == nil {
firstErr = err
}
}
// Put the per-ingress-iface knobs back the way we found them. With the table
// and the policy routing gone, a lingering accept_local=1 / rp_filter=0 on a
@@ -880,7 +1065,11 @@ func (a *Applier) Teardown() error {
clearActiveFlag(a.log)
a.lastGood = nil
a.lastNft = ""
a.setHolding(false)
// Not a blanket false any more: with a holding plane standing, forwarded LAN
// traffic really IS being blocked, and saying otherwise here is the inverted lie
// Holding()'s doc comment is about — the process is exiting, but Status can
// still be read over the control socket before it does.
a.setHolding(kept)
a.setTraffic(generate.Traffic{})
a.setWarnings(nil)
// The plane is gone, so the logged set no longer describes anything. Forget it,
@@ -1344,10 +1533,14 @@ func processUptime(now time.Time) (startedUnix, uptimeSeconds int64) {
// traffic verdict — which is the state this field exists to make expressible.
func (a *Applier) Status() Status {
engineUp := a.eng != nil && a.eng.Running()
// Read the kernel ONCE and hand the same fact to both Table and the hold
// verdict. Asking twice is not just an extra `nft` fork per poll: the two reads
// could disagree across a teardown and publish table=false with plane="hold".
tableLoaded := tableExists()
s := Status{
Running: engineUp,
Active: ActiveFlagPresent(),
Table: netplane.TableExists(),
Table: tableLoaded,
Hash: a.eng.Hash(),
CanRollback: a.canRollback(),
EngineRunning: engineUp,
@@ -1356,9 +1549,13 @@ func (a *Applier) Status() Status {
}
s.StartedUnix, s.UptimeSeconds = processUptime(time.Now())
switch {
case !s.Table:
case !tableLoaded:
s.Plane = "none"
case a.Holding():
case a.holdingWith(tableLoaded, engineUp):
// Includes the plane THIS PROCESS DID NOT INSTALL — the boot armor loaded by
// /etc/init.d/shater-armor, or what a predecessor left behind. Without that
// the branch below claimed "full" (documented as "traffic is diverted into a
// RUNNING engine") over a dead engine and a blocked LAN.
s.Plane = "hold"
default:
s.Plane = "full"
+270
View File
@@ -0,0 +1,270 @@
package apply
// The hold state must describe THE ROUTER, not this process's memory of what it
// did.
//
// The boot armor (netplane/armor.go) put a fail-closed plane in the kernel that
// the Applier never installs: /etc/init.d/shater-armor loads it at START=21,
// before the daemon exists, and cmd/shaterd reinstates it when the config cannot
// be read. The latch behind Holding() knew nothing about either, so the daemon
// reported `holding=false` and `plane="full"` over a LAN that was blocked — and
// with an unreadable config that state is PERMANENT, because the only thing that
// ever wrote the latch was an apply and no apply can run.
//
// What comes out the other end is an alert (cmd/shaterd fireApplyFail) that
// chooses its wording from exactly this bool and tells the operator "Traffic is
// NOT being blocked" while it is. That is the inverted failure: not a fault
// hidden, but a fault invented — and the obvious response to it is to go and
// dismantle the protection that is doing its job.
import (
"errors"
"os"
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
)
// stubPlaneFacts pins the two outside-world seams — is our table in the kernel,
// is the persisted fail-closed plane on flash — for the duration of a test.
func stubPlaneFacts(t *testing.T, table, armor bool) func() {
t.Helper()
origTable, origArmor := tableExists, bootArmorPresent
tableExists = func() bool { return table }
bootArmorPresent = func() bool { return armor }
return func() { tableExists, bootArmorPresent = origTable, origArmor }
}
// TestHoldingSeesAPlaneThisProcessDidNotInstall is the core regression.
//
// All three cases share the same latch value (false — this Applier has installed
// nothing) and must still produce three different, correct answers.
func TestHoldingSeesAPlaneThisProcessDidNotInstall(t *testing.T) {
a := New(engine.New(), nil)
// (1) No table at all. Nothing is protecting the LAN and nothing may claim to.
restore := stubPlaneFacts(t, false, true)
if a.Holding() {
t.Errorf("Holding() = true with no table loaded")
}
if got := a.Status().Plane; got != "none" {
t.Errorf("Plane = %q with no table loaded, want \"none\"", got)
}
restore()
// (2) The boot armor's plane IS loaded and this engine never started. This is
// what /etc/init.d/shater-armor leaves behind at every boot, and what the
// daemon reinstates when /etc/config/shater cannot be read — the case where
// nothing else will ever correct the answer.
restore = stubPlaneFacts(t, true, true)
if !a.Holding() {
t.Errorf("Holding() = false while the boot armor's fail-closed plane is loaded and the " +
"engine is down — the LAN is blocked and the daemon says it is not; the apply-failure " +
"alert would send the operator to fix a protection that is working")
}
s := a.Status()
if s.Plane != "hold" {
t.Errorf("Plane = %q over a blocked LAN with a dead engine, want \"hold\" "+
"(\"full\" means traffic is diverted into a RUNNING engine)", s.Plane)
}
if !s.Table {
t.Errorf("Table = false although a table is loaded")
}
if s.Running || s.EngineRunning {
t.Errorf("running/engine_running must stay false while holding: %+v", s)
}
restore()
// (3) A table is loaded but there is NO armor on flash. refreshBootArmor
// removes that file exactly when the operator disables the stack or opens the
// kill switch, so what is loaded here is a leftover that drops nothing.
// Claiming a hold would be the new lie: it would tell someone who deliberately
// chose fail-open that their LAN is cut off.
restore = stubPlaneFacts(t, true, false)
if a.Holding() {
t.Errorf("Holding() = true with no fail-closed armor on flash — a leftover plane from a " +
"kill_switch=open config blocks nothing, and saying otherwise is the same lie inverted")
}
restore()
}
// TestSuccessfulApplyStillClearsTheHold: the self-healing path must survive the
// derived half. A read-time inference that ignored the engine would re-assert the
// hold one line after setHolding(false) and pin the router in "protected, not
// proxying" forever — with the table loaded and the armor on flash, which is the
// steady state of every healthy router.
func TestSuccessfulApplyStillClearsTheHold(t *testing.T) {
a := New(engine.New(), nil)
t.Cleanup(func() { _ = a.eng.Close() })
t.Cleanup(func() { _ = os.Remove(ActiveFlag) })
// A REAL started instance: the derived half asks the engine, so a stub would
// not exercise the thing under test.
if _, err := a.eng.Apply(mixedOn(18841)); err != nil {
t.Fatalf("bring a real engine up: %v", err)
}
if !a.eng.Running() {
t.Fatalf("precondition: the engine must be running")
}
// The steady state of a healthy router: our table loaded, armor on flash.
defer stubPlaneFacts(t, true, true)()
a.setHolding(true) // whatever put us on hold before this apply
if !a.Holding() {
t.Fatalf("precondition: the latch must read through")
}
m := holdModel("closed")
m.Globals.GroupHealth = false // no background probing from a unit test
defer stubApplyStages(t,
func(*Applier, option.Options) (bool, error) { return true, nil },
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
return planeOutcome{changed: true}, nil
})()
a.mu.Lock()
_, err := a.applyLocked(m)
a.mu.Unlock()
if err != nil {
t.Fatalf("applyLocked: %v", err)
}
if a.Holding() {
t.Fatalf("a successful apply did not clear the hold — the router is proxying and the " +
"panel would still show \"protected, not proxying\"")
}
if got := a.Status().Plane; got != "full" {
t.Errorf("Plane = %q after a successful apply with the engine up, want \"full\"", got)
}
}
// TestApplyFailureOverAForeignPlaneReportsBlocked pins the input cmd/shaterd's
// fireApplyFail words its incident from.
//
// The shape is the real one: the engine will not start, so applyLocked calls
// holdLocked — and holdLocked cannot install anything either (nft refuses, the
// overlay is full). The latch therefore stays false. But the boot armor's plane
// is still standing in the kernel, so forwarded traffic IS being dropped, and the
// incident must say so. With only the latch, this is precisely where the daemon
// said "Traffic is NOT being blocked (kill switch is open)" on a router whose
// kill switch was closed and whose LAN was cut off.
func TestApplyFailureOverAForeignPlaneReportsBlocked(t *testing.T) {
a := New(engine.New(), nil)
defer stubPlaneFacts(t, true, true)()
origHold := applyHoldNft
applyHoldNft = func(string) error { return errors.New("nft -f (load) failed: no space left on device") }
defer func() { applyHoldNft = origHold }()
defer stubApplyStages(t,
func(*Applier, option.Options) (bool, error) {
return false, errors.New("start rule-set[geosite]: connection refused")
},
func(*Applier, *model.Model, option.Options, time.Time) (planeOutcome, error) {
t.Errorf("the data-plane stage must not run after the engine stage failed")
return planeOutcome{}, nil
})()
a.mu.Lock()
_, err := a.applyLocked(holdModel("closed"))
a.mu.Unlock()
if err == nil {
t.Fatalf("applyLocked must surface the engine failure")
}
a.stateMu.RLock()
latched := a.holding
a.stateMu.RUnlock()
if latched {
t.Fatalf("precondition: holdLocked could not install a plane, so the latch must be false")
}
// fireApplyFail(notifier, err, applier.Holding()) — this bool picks between
// "forwarded LAN traffic is being dropped" and "Traffic is NOT being blocked".
if !a.Holding() {
t.Fatalf("Holding() = false while a fail-closed plane blocks the LAN: the incident would " +
"read \"Traffic is NOT being blocked\" at the exact moment it is being blocked")
}
}
// TestArmHoldReplacesAPlaneItDidNotInstall: ArmHold used to return early on
// TableExists() and publish nothing, which is how the boot armor's plane came to
// be loaded with holding=false. It now installs its own — rendered from the model
// it was handed, so the state it publishes is a fact about something this process
// did rather than a guess about something it found (and the fresh render also
// replaces a snapshot that may predate an interface rename).
func TestArmHoldReplacesAPlaneItDidNotInstall(t *testing.T) {
loaded := withHoldProbe(t)
defer stubPlaneFacts(t, true, true)() // a table is ALREADY loaded
a := New(engine.New(), nil)
a.ArmHold(holdModel("closed"))
if len(*loaded) != 1 {
t.Fatalf("ArmHold loaded %d rulesets over an existing table, want 1 — it deferred to a "+
"table it cannot inspect instead of installing one it can vouch for", len(*loaded))
}
if !a.Holding() {
t.Errorf("ArmHold installed the holding plane but did not publish it")
}
if got := a.Status().Plane; got != "hold" {
t.Errorf("Plane = %q after ArmHold, want \"hold\"", got)
}
}
// TestArmHoldStillRespectsFailOpen: replacing a foreign table must not become a
// licence to install a plane the operator did not ask for. kill_switch=open and
// globals.enabled=0 are explicit choices and ArmHold must keep obeying both.
func TestArmHoldStillRespectsFailOpen(t *testing.T) {
loaded := withHoldProbe(t)
defer stubPlaneFacts(t, true, false)()
a := New(engine.New(), nil)
a.ArmHold(holdModel("open"))
if len(*loaded) != 0 {
t.Errorf("kill_switch=open loaded %d rulesets, want 0", len(*loaded))
}
disabled := holdModel("closed")
disabled.Globals.Enabled = false
a.ArmHold(disabled)
if len(*loaded) != 0 {
t.Errorf("globals.enabled=0 loaded %d rulesets, want 0", len(*loaded))
}
a.ArmHold(nil)
if len(*loaded) != 0 {
t.Errorf("a nil model loaded %d rulesets, want 0", len(*loaded))
}
if a.Holding() {
t.Errorf("nothing was installed, so nothing may be reported as holding")
}
}
// TestForeignHoldMatrix pins the predicate itself, including the two facts it
// refuses to guess about.
func TestForeignHoldMatrix(t *testing.T) {
for _, tc := range []struct {
name string
table, engine, armor bool
want bool
}{
{"no table", false, false, true, false},
{"engine up: the table is the working full plane", true, true, true, false},
{"armor on flash, engine down: blocked", true, false, true, true},
{"no armor: the last config was disabled or fail-open", true, false, false, false},
{"nothing at all", false, false, false, false},
} {
t.Run(tc.name, func(t *testing.T) {
defer stubPlaneFacts(t, tc.table, tc.armor)()
if got := foreignHold(tc.table, tc.engine); got != tc.want {
t.Errorf("foreignHold(table=%v, engineUp=%v) = %v, want %v",
tc.table, tc.engine, got, tc.want)
}
})
}
}
+149
View File
@@ -0,0 +1,149 @@
package apply
// The exit path must never leave the LAN uncovered, and "never" is an ORDER, not
// an intention.
//
// WHAT THIS PINS. The daemon's SIGTERM path used to be:
//
// applier.Teardown() // netplane.TeardownNft() -> `nft delete table inet shater`
// armOnExit(handoff) // then, separately, install the fail-closed holding plane
//
// Two nft transactions. Between them the `inet shater` table does not exist, so
// fw4's `lan -> wan ACCEPT` is the only policy on the box and every forwarded LAN
// packet leaves in the clear. MEASURED on the stand at 80-90 ms, reproduced twice
// with a 35 000-sample run at ~1.3 ms resolution — and this is not a boot-time
// window that heals itself: it is every `restart`, every `reload_service` (i.e.
// every LuCI Save & Apply), and every package upgrade.
//
// The old call site even carried a comment explaining the mechanism — "AFTER the
// teardown, never before: Teardown deletes the table, so a plane installed first
// would simply be removed again" — and drew the wrong conclusion from a correct
// observation. The fix is not to arm later, it is to stop deleting: arm first (a
// single `nft -f` that opens with `delete table` and closes with the new one, so
// the kernel replaces rather than removes), then skip the delete.
//
// So the property under test is a SEQUENCE, and the test records the order the
// two seams are called in. A test that only asserted "the table still exists at
// the end" would pass against the broken code.
import (
"errors"
"testing"
"github.com/sagernet/sing-box/shater/engine"
)
// recordTeardownSeams captures the order in which the exit path touches the
// kernel: "arm" when a holding plane is installed, "delete" when the table is
// removed.
func recordTeardownSeams(t *testing.T) (*[]string, func()) {
t.Helper()
var calls []string
orig := teardownNft
teardownNft = func() error {
calls = append(calls, "delete")
return nil
}
return &calls, func() { teardownNft = orig }
}
// TestTeardownExitingReplacesThePlaneInsteadOfRemovingIt is the regression: with a
// successor coming, the table must be swapped and NEVER deleted.
func TestTeardownExitingReplacesThePlaneInsteadOfRemovingIt(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.TeardownExiting(func() bool {
*calls = append(*calls, "arm")
return true
}); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "arm" {
t.Fatalf("exit path did %v, want exactly [arm]: the holding plane must be installed "+
"and the table must NOT be deleted — a `delete` here is the 80-90 ms window in which "+
"fw4's lan->wan ACCEPT is the only policy on the box", *calls)
}
if !a.Holding() && tableExists == nil {
t.Errorf("unreachable; keeps the linter honest about the seam")
}
}
// TestTeardownExitingArmsBeforeItTearsDown pins the ORDER even in the case where
// the table does still get removed. Arming has to be the first thing that touches
// the kernel; if it ran after the delete we would be back to the two-transaction
// gap with extra steps.
func TestTeardownExitingArmsBeforeItTearsDown(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
// arm reports FALSE: nothing was installed (a render failure, or kill_switch=open
// where fail-open is the operator's documented choice). The table must then come
// down exactly as it always did.
if err := a.TeardownExiting(func() bool {
*calls = append(*calls, "arm")
return false
}); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if len(*calls) != 2 || (*calls)[0] != "arm" || (*calls)[1] != "delete" {
t.Fatalf("exit path did %v, want [arm delete]: arming must precede the delete, and a "+
"plane that was NOT installed must not keep the table alive", *calls)
}
}
// TestTeardownStillRemovesEverything is the escape hatch. A deliberate `stop`, and
// a Reconcile that finds the stack disabled, both come through the plain Teardown
// and must dismantle the plane completely — a kill switch that cannot be switched
// off is a brick.
func TestTeardownStillRemovesEverything(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.Teardown(); err != nil {
t.Fatalf("Teardown: %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "delete" {
t.Fatalf("Teardown did %v, want [delete]: an operator's stop must take the table with it", *calls)
}
}
// TestTeardownExitingReportsHolding pins the honesty half: with a plane left
// standing the Applier must not go on saying it installed nothing. Status can
// still be read over the control socket between the swap and the exit, and
// "holding=false over a blocked LAN" is the inverted lie holdstate_test.go is
// about, just reached down a different path.
func TestTeardownExitingReportsHolding(t *testing.T) {
_, restore := recordTeardownSeams(t)
defer restore()
restoreFacts := stubPlaneFacts(t, true, true)
defer restoreFacts()
a := New(engine.New(), nil)
if err := a.TeardownExiting(func() bool { return true }); err != nil {
t.Fatalf("TeardownExiting: %v", err)
}
if !a.Holding() {
t.Errorf("Holding() = false right after the exit path left a fail-closed plane standing")
}
}
// TestTeardownExitingSurvivesANilArm keeps the plain-Teardown contract explicit:
// a nil arm is "nothing to install", not a panic.
func TestTeardownExitingSurvivesANilArm(t *testing.T) {
calls, restore := recordTeardownSeams(t)
defer restore()
a := New(engine.New(), nil)
if err := a.TeardownExiting(nil); err != nil && !errors.Is(err, nil) {
t.Fatalf("TeardownExiting(nil): %v", err)
}
if len(*calls) != 1 || (*calls)[0] != "delete" {
t.Fatalf("TeardownExiting(nil) did %v, want [delete]", *calls)
}
}
+122
View File
@@ -360,6 +360,128 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
return append(out, Warning{Severity: SeverityInfo, Section: section, Name: name, Message: msg})
}
// The egress carrier owns the whole story the moment untunnelable_egress
// names an egress: the carried protocols get that egress's own mark in
// prerouting and the ROUTING decision sends them out its device with the
// kernel's NAT — the forward chain, where the policy's verdicts live, no
// longer decides their fate. So this branch sits above every other and
// returns its own text; with the option empty, the notes below are
// byte-for-byte what they were. Two phrasing rules here are load-bearing.
// First, the egress is NEVER called a tunnel unconditionally: the option
// accepts any interface/tunnel egress, and on the routers this ships to
// that is at least as often a second WAN — another uplink, whose real
// address the far end sees — as a WireGuard device. Second, the note must
// say out loud that UDP-based VPNs are none of this option's business: an
// operator who reads "VPN passthrough" and enables it for a WireGuard
// client that already worked through the ordinary tunnel has been misled,
// not helped. What the policy still owns is exactly the failure path — a
// name that resolves to no interface/tunnel egress, a rule or route that
// did not come up — and each tail below says what that failure looks like,
// because under `direct` (or an open kill switch) it is a silent leak
// through the normal uplink with the real address, and nothing anywhere
// else would say so.
if egressName := strings.TrimSpace(g.UntunnelableEgress); egressName != "" {
var msg, failure string
if netplane.L3Enabled(g) {
msg = "Ping and Windows tracert keep travelling THROUGH the tunnel, toward every " +
"address your rules send to an outbound that can carry plain IP " +
"(WireGuard/AmneziaWG) — the L3 ingress claims ICMP before this option is " +
"consulted, and addresses your rules send anywhere else " +
"(vless/vmess/trojan/shadowsocks and the like) still cannot be pinged at " +
"all, deliberately. Everything else the proxy cannot carry — IPsec " +
"(ESP/AH), PPTP/GRE, SCTP and every other protocol that is neither TCP " +
"nor UDP — now leaves through egress \"" + egressName + "\": the kernel " +
"routes it out that interface with that interface's own NAT, and none of " +
"it goes through the proxy or follows your routing rules. "
failure = "when one of the routes is not in place — the L3 route for ping, the " +
"egress route for the rest (a name that matches no interface/tunnel " +
"egress, or a rule or route that failed to come up): "
} else {
msg = "Ping, Windows tracert, IPsec (ESP/AH), PPTP/GRE, SCTP and every other " +
"protocol that is neither TCP nor UDP now leave through egress \"" +
egressName + "\": the kernel routes them out that interface with that " +
"interface's own NAT, and none of it goes through the proxy or follows " +
"your routing rules — the hops tracert prints are that interface's path " +
"(on Linux and macOS traceroute sends UDP probes instead, which still " +
"follow your rules). "
failure = "when the egress route is not in place (a name that matches no " +
"interface/tunnel egress, or a rule or route that failed to come up): "
}
msg += "What that buys depends entirely on what the interface IS: a WireGuard " +
"interface really is a tunnel, but a second WAN is not — it is just another " +
"uplink, and the host on the far end sees that uplink's real address. Two " +
"things this option does NOT do: multicast IPTV does not pass this router " +
"under any setting, and carrying IGMP out an egress cannot change that; and " +
"VPNs that run over UDP (WireGuard, OpenVPN-UDP, IPsec through NAT — IKE on " +
"UDP 500, NAT-T on UDP 4500) never needed it: they are ordinary tunnelled " +
"traffic, keep following your routing rules exactly as before, and gain " +
"nothing from this option. The `untunnelable` policy no longer decides this " +
"traffic's fate — routing settles it before the forward chain gets a say — " +
"and answers only for failure, " + failure
switch {
case !killSwitchClosed(g):
msg += "with the kill switch open nothing is dropped, so whatever loses its " +
"route quietly leaves through your normal uplink with your real IP address."
case policy == netplane.UntunnelableDirect:
msg += "\"direct\" quietly lets it leave through your normal uplink with " +
"your real IP address."
case policy == netplane.UntunnelableICMP:
msg += "\"icmp\" drops it, excepting only ping — which then quietly leaves " +
"with your real IP address instead of failing."
default:
msg += "\"block\" drops it — an honest loss rather than a silent leak."
}
return note(egressName, msg)
}
// The L3 ingress rewrites the ICMP half of every note below, so it gets one
// text of its own rather than four patched variants: echo is marked in
// prerouting and the ROUTING decision carries it into the engine's TUN before
// the forward chain — where the policy accepts and the kill-switch drops
// live — is ever consulted. That holds under all three policy values and
// with the kill switch open alike, which is why this branch sits above the
// kill-switch note: "reaches the internet with your real IP address" stops
// being true for ping the moment the divert exists. What the policy still
// owns is exactly two things, and both are said: the protocols the engine
// cannot ingest at all (raw IPsec, PPTP/GRE), and the fallback path a marked
// packet takes when the L3 route failed to install — under `direct` (or an
// open kill switch) that failure is a SILENT leak with the real address,
// under `block` an honest packet loss. The unpingable-through-proxy sentence
// is deliberate too: those pings used to be answered by the router itself,
// and a fake "alive" is worse than a truthful timeout.
if netplane.L3Enabled(g) {
msg := "Ping and Windows tracert work and travel THROUGH the tunnel, toward every " +
"address your rules send to an outbound that can carry plain IP " +
"(WireGuard/AmneziaWG). Addresses your rules send anywhere else " +
"(vless/vmess/trojan/shadowsocks and the like) cannot be pinged at all — " +
"deliberately: those pings used to be answered by the router itself, reporting " +
"hosts alive it had never reached. The hops tracert prints are the tunnel's " +
"path, not your own, and IPv6 traceroute shows only the destination, none of " +
"the hops on the way. Raw VPN passthrough (IPsec ESP/AH, PPTP/GRE) cannot " +
"enter the tunnel at all and stays with the untunnelable policy: "
switch {
case !killSwitchClosed(g):
msg += "with the kill switch open none of it is dropped, so it leaves with your " +
"real IP address — and if the L3 route ever fails to come up, ping quietly " +
"does the same instead of failing."
case policy == netplane.UntunnelableDirect:
msg += "\"direct\" lets it out with your real IP address — and if the L3 route " +
"ever fails to come up, ping quietly does the same instead of failing."
case policy == netplane.UntunnelableICMP:
msg += "\"icmp\" drops it, excepting only echo — which now rides the tunnel " +
"anyway, so the exception matters just once: if the L3 route ever fails to " +
"come up, it lets ping quietly leave with your real IP address instead of " +
"failing."
default:
msg += "\"block\" drops it — and if the L3 route ever fails to come up, ping " +
"fails outright rather than leaking."
}
msg += " VPNs that run over UDP (WireGuard, OpenVPN-UDP, IPsec through NAT) are " +
"ordinary tunnelled traffic and are unaffected either way. Multicast IPTV does " +
"not pass this router on any setting; the L3 ingress does not change that."
return note(policy, msg)
}
// With the kill switch open the forward chain has no drops at all, so nothing
// is restricted whatever the policy says. Saying that is more useful than
// repeating a promise which is not being kept.
+281
View File
@@ -0,0 +1,281 @@
// Keeping the fail-closed plane alive across the moments the daemon is not.
//
// The daemon owns the `inet shater` table, which means the table exists exactly
// while the daemon does. Three of those moments are not covered by anything else,
// and all three are the same defect wearing different clothes: the protection is
// an in-process thing, and the process is not always there.
//
// BOOT /etc/init.d/shater is START=99. fw4 loaded `lan -> wan ACCEPT` at 19
// and netifd brought the LAN up at 20; the clients that reconnect in
// between are unprotected until the daemon has been decompressed off
// flash, waited out any predecessor, migrated UCI and applied.
// RESTART SIGTERM ran an unconditional Teardown — kill_switch was not so much
// as consulted — and the successor cannot apply until the init's
// shater_wait_stopped loop, `shaterd migrate` and engine start have all
// finished. `reload_service` is stop+start, and so is every package
// upgrade, so this ran on a routine `Save & Apply`.
// NO CONFIG model.ReadUCI failing left the arming call unreached: it sat in the
// else-branch of the successful read. Nothing recovered from it either
// — Reconcile returns before any plane work, and the cron watchdog sees
// a live pidof and its own `uci -q get` fails the same way.
//
// The answer to all three is one artifact: netplane's BOOT ARMOR, a persisted copy
// of the fail-closed holding plane (netplane/armor.go). This file is the daemon's
// half — it keeps that copy honest, and it reinstates it in the two cases the
// daemon is the only one who can.
//
// Everything here is deliberately conservative in ONE direction: it never installs
// a plane the operator did not ask for. `globals.enabled=0` or `kill_switch=open`
// removes the armor and installs nothing, and a deliberate `/etc/init.d/shater
// stop` is a handoff-free exit that leaves nothing behind. Fail-closed is a
// policy, and a kill switch that outlives its own off switch is not one.
package main
import (
"os"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// restartHandoffPath is raised by /etc/init.d/shater around a restart/reload and
// cleared by its start (and by a real stop). Its presence at SIGTERM means "this
// daemon is being REPLACED", as opposed to "this daemon is being switched off".
//
// tmpfs on purpose: a marker that survived a power cut would make the first boot
// after it look like a restart.
//
// A var, not a const, only so tests can point it at a temp dir.
var restartHandoffPath = "/var/run/shater.restarting"
// restartHandoffPending reports whether the init script announced a restart.
//
// Absent is read as "a real stop", which is the SAFE direction to be wrong in: it
// degrades to exactly the behaviour that shipped before this file existed (full
// teardown), whereas the other default would leave a stopped router blocked.
func restartHandoffPending() bool {
_, err := os.Stat(restartHandoffPath)
return err == nil
}
// armorPlan is what to do about the fail-closed plane at a decision point.
type armorPlan int
const (
// armorNothing: install nothing. Either the operator does not want a plane
// (disabled / kill_switch=open), or there is nothing to install from.
armorNothing armorPlan = iota
// armorRender: build the holding plane from the model we just read. Preferred
// whenever a model is readable — it reflects the CURRENT interface set, where a
// snapshot may predate an interface rename.
armorRender
// armorSnapshot: reinstate the persisted boot armor. The only option when the
// config cannot be read, which is precisely when it is needed.
armorSnapshot
)
// armorWanted reports whether m asks for a fail-closed plane at all: the stack is
// enabled AND the kill switch is closed. Both halves are the operator's explicit
// choice and neither may be second-guessed — `kill_switch=open` is a documented
// decision to let traffic through when the engine is down, not an oversight.
func armorWanted(m *model.Model) bool {
return m != nil && m.Globals.Enabled && netplane.KillSwitchClosed(m.Globals)
}
// planStartupArmor decides what to install when the daemon starts and could NOT
// read its config.
//
// It is only ever consulted on the read-failure path: with a readable model the
// applier's own ArmHold does this job (and does it better — it holds the apply
// lock while it works). With no model there is nothing to render from, so the
// persisted snapshot is the entire answer; with no snapshot either, nothing is
// installed, because "this router has never applied an enabled, fail-closed
// config" is then the most likely truth and blacking out a LAN on a guess is not
// a recovery.
func planStartupArmor(readErr error, snapshot bool) armorPlan {
if readErr == nil {
return armorNothing
}
if snapshot {
return armorSnapshot
}
return armorNothing
}
// planExitArmor decides what the daemon leaves behind when it is asked to exit.
//
// The handoff flag is the whole distinction the old code was missing. A restart,
// a reload and a package upgrade all reach this point, and in all three the
// operator has not asked for protection to end — only for this process to be
// replaced. A `stop` has asked for exactly that, and must be obeyed: it is the
// operator's escape hatch, and a kill switch that cannot be switched off is a
// brick.
func planExitArmor(handoff bool, m *model.Model, readErr error, snapshot bool) armorPlan {
if !handoff {
return armorNothing
}
if readErr != nil {
// Being replaced with an unreadable config: the snapshot is the last thing
// this router is known to have wanted, and it is still the honest answer.
if snapshot {
return armorSnapshot
}
return armorNothing
}
if !armorWanted(m) {
return armorNothing
}
return armorRender
}
// refreshBootArmor keeps the persisted holding plane in step with the desired
// state. Called after every successful UCI read, so the snapshot on flash always
// describes the config the router is actually running.
//
// Writing is content-gated inside netplane.SaveBootArmor (this runs once a minute
// under cron; rewriting an identical file that often is how flash dies), and the
// REMOVE half matters just as much as the write: turning the stack off, or opening
// the kill switch, has to disarm the next boot too, or the operator's change would
// silently come back after a power cut.
func refreshBootArmor(m *model.Model, logger log.ContextLogger) {
if !armorWanted(m) {
if netplane.BootArmorPresent() {
if err := netplane.RemoveBootArmor(); err != nil {
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", err)
} else {
logger.Info("boot armor removed (the stack is disabled or the kill switch is open): ",
"the LAN is no longer blocked at boot before the daemon starts")
}
}
return
}
ruleset, err := netplane.RenderHoldNft(m)
if err != nil {
// A transient render failure must not disarm: a stale fail-closed plane is
// recoverable (the daemon replaces it seconds into the next boot), an absent
// one is a leak.
logger.Warn("boot armor: could not render the fail-closed plane: ", err)
return
}
if ruleset == "" {
// No divert devices at all — this config intercepts nothing, so there is
// nothing for a boot-time plane to protect. Blocking the LAN at boot on
// behalf of a config that does not touch it would be a pure outage.
if netplane.BootArmorPresent() {
if rerr := netplane.RemoveBootArmor(); rerr != nil {
logger.Warn("boot armor: could not remove ", netplane.BootArmorPath, ": ", rerr)
}
}
return
}
changed, serr := netplane.SaveBootArmor(ruleset)
switch {
case serr != nil:
logger.Warn("boot armor: could not write ", netplane.BootArmorPath, ": ", serr,
" — the LAN will be unprotected between boot and this daemon's first apply")
case changed:
logger.Info("boot armor updated (", netplane.BootArmorPath,
"): the LAN is fail-closed from early boot until the engine is up")
}
}
// armFromSnapshot reinstates the persisted holding plane and reports whether a
// plane is now standing. why is a short phrase for the log ("config is
// unreadable", "restart handoff").
func armFromSnapshot(why string, logger log.ContextLogger) bool {
loaded, err := netplane.LoadBootArmor()
switch {
case err != nil:
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): the saved plane ",
netplane.BootArmorPath, " could not be loaded: ", err,
" — LAN traffic may be reaching the WAN unprotected")
return false
case loaded:
logger.Error("fail-closed plane reinstated from ", netplane.BootArmorPath,
" (", why, "): LAN->WAN forwarding is BLOCKED. ",
"SSH, LuCI and the admin panel remain reachable.")
return true
default:
logger.Warn("no saved fail-closed plane at ", netplane.BootArmorPath, " (", why,
"): nothing was installed")
return false
}
}
// armFromModel renders the holding plane for m, installs it, and reports whether
// a plane is now standing.
//
// The return value is load-bearing on the exit path: Applier.TeardownExiting keeps
// the nft table only when a plane really was installed, so a render failure or a
// fail-open config falls back to the old remove-everything behaviour instead of
// leaving whatever the engine happened to have in the kernel.
func armFromModel(m *model.Model, why string, logger log.ContextLogger) bool {
ruleset, err := netplane.RenderHoldNft(m)
if err != nil {
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): render failed: ", err)
return false
}
if ruleset == "" {
// No divert devices: there is nothing this plane would protect.
return false
}
// One `nft -f` that opens with `delete table` and closes with the new table:
// the swap is a single netlink transaction, so this REPLACES whatever plane is
// loaded without the table ever being absent. That property is why the exit
// path can arm before it tears down.
if err := netplane.ApplyNft(ruleset); err != nil {
logger.Error("FAIL-CLOSED PLANE NOT INSTALLED (", why, "): ", err,
" — LAN traffic may be reaching the WAN unprotected")
return false
}
logger.Info("fail-closed plane left in place (", why,
"): LAN->WAN forwarding stays BLOCKED until the next daemon applies. ",
"SSH, LuCI and the admin panel remain reachable.")
return true
}
// armOnUnreadableConfig is the window-3 answer: the daemon is up but cannot read
// its own desired state, so it falls back to the last state it persisted.
//
// A table that is ALREADY loaded is left alone. This runs on every reconcile —
// cron fires one a minute — and a `nft -f` is a delete-and-recreate of the whole
// table plus a DNS conntrack flush, so re-installing an identical plane sixty
// times an hour would be pure churn, and each replacement is itself a brief hole.
// The question this path answers is "is there anything at all standing", and once
// the answer is yes it stays yes until an apply succeeds and replaces it properly.
func armOnUnreadableConfig(logger log.ContextLogger) {
if netplane.TableExists() {
return
}
if planStartupArmor(errUnreadableConfig, netplane.BootArmorPresent()) != armorSnapshot {
logger.Warn("the config could not be read and there is no saved fail-closed plane at ",
netplane.BootArmorPath, " — nothing is protecting the LAN; fix /etc/config/shater ",
"(a full /overlay is the usual cause) and reconcile")
return
}
armFromSnapshot("the config could not be read", logger)
}
// errUnreadableConfig is a stand-in for "the read failed" in the call above,
// where the concrete error has already been logged by the caller.
var errUnreadableConfig = os.ErrInvalid
// armOnExit is the window-2 answer: what this daemon leaves in the kernel when it
// is asked to go away, and whether anything is now standing there. See
// planExitArmor for the policy.
//
// It is called BY Applier.TeardownExiting, before the teardown and under the apply
// lock, so that the plane is swapped rather than removed-then-rebuilt. It must
// therefore never call back into the Applier — everything here goes straight to
// model.ReadUCI and netplane.
func armOnExit(handoff bool, logger log.ContextLogger) bool {
m, err := model.ReadUCI()
switch planExitArmor(handoff, m, err, netplane.BootArmorPresent()) {
case armorRender:
return armFromModel(m, "restart handoff", logger)
case armorSnapshot:
return armFromSnapshot("restart handoff, config unreadable", logger)
}
return false
}
+182
View File
@@ -0,0 +1,182 @@
package main
// The daemon's half of the fail-closed armor: the two decisions that decide
// whether the LAN is protected in the moments this process is not running.
//
// Both are pure functions on purpose. The behaviour they encode can otherwise
// only be observed on a router, with root, by killing a daemon at the right
// moment and reading `nft list ruleset` — i.e. never, in a gate.
import (
"errors"
"os"
"path/filepath"
"strings"
"testing"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
func enabledClosed() *model.Model {
return &model.Model{Globals: model.Globals{
Enabled: true, KillSwitch: "closed", FwmarkBase: 0x2000, TableBase: 0x2000,
}, Inbounds: []model.Inbound{{
Name: "lan", Enabled: true, Type: "tproxy", Network: "lan",
TproxyPort: 12345, TCP: true, UDP: true,
}}}
}
// TestPlanExitArmor is W2. SIGTERM used to run an unconditional Teardown — the
// only place in this codebase that removes the fail-closed plane without so much
// as reading kill_switch — and every `restart`, every `reload_service` (which is
// what a LuCI Save & Apply runs) and every package upgrade went through it. The
// gap that follows is guaranteed non-empty by the init script itself.
//
// So the exit has to know WHY it is exiting. What it must never do is confuse the
// two directions: a restart that leaves nothing behind is a plaintext window, and
// a stop that leaves a block behind is a router the operator cannot un-brick.
//
// RED BEFORE: there was no such decision — the daemon always tore everything down.
func TestPlanExitArmor(t *testing.T) {
open := enabledClosed()
open.Globals.KillSwitch = "open"
disabled := enabledClosed()
disabled.Globals.Enabled = false
cases := []struct {
name string
handoff bool
m *model.Model
readErr error
snapshot bool
want armorPlan
}{
{"restart, enabled + fail-closed => leave the plane standing",
true, enabledClosed(), nil, true, armorRender},
{"restart, no snapshot on disk => still render from the live model",
true, enabledClosed(), nil, false, armorRender},
{"restart, kill_switch=open => the operator chose fail-open; install nothing",
true, open, nil, true, armorNothing},
{"restart, stack disabled => nothing to protect",
true, disabled, nil, true, armorNothing},
{"restart, config unreadable => the persisted plane is the last known truth",
true, nil, errors.New("uci: no such file"), true, armorSnapshot},
{"restart, config unreadable and nothing persisted => nothing to install",
true, nil, errors.New("uci: no such file"), false, armorNothing},
// The escape hatch. A deliberate `/etc/init.d/shater stop` must mean what it
// says in every one of these, or the kill switch has no off switch.
{"stop, enabled + fail-closed => the plane goes away",
false, enabledClosed(), nil, true, armorNothing},
{"stop, config unreadable => still goes away",
false, nil, errors.New("uci: no such file"), true, armorNothing},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := planExitArmor(c.handoff, c.m, c.readErr, c.snapshot); got != c.want {
t.Errorf("planExitArmor(handoff=%v, readErr=%v, snapshot=%v) = %v, want %v",
c.handoff, c.readErr, c.snapshot, got, c.want)
}
})
}
}
// TestPlanStartupArmor is W3. An unreadable /etc/config/shater — a full /overlay
// caught mid `uci commit` is the cause the init script itself documents — left the
// arming call unreached, because it sat in the else-branch of the successful read.
// Nothing recovered from that: Reconcile returns before any plane work and the
// cron watchdog's own `uci -q get` fails identically, so the box sat there with a
// live daemon, an answering panel and no table at all.
//
// RED BEFORE: no decision existed; the read-failure branch only logged.
func TestPlanStartupArmor(t *testing.T) {
readErr := errors.New("uci: cannot read /etc/config/shater")
if got := planStartupArmor(readErr, true); got != armorSnapshot {
t.Errorf("unreadable config with a persisted plane = %v, want armorSnapshot", got)
}
// Nothing persisted means this router has never applied an enabled,
// fail-closed config. Blacking out a LAN on that guess is not a recovery.
if got := planStartupArmor(readErr, false); got != armorNothing {
t.Errorf("unreadable config with no persisted plane = %v, want armorNothing", got)
}
// A readable config is the applier's business (ArmHold), not this path's.
if got := planStartupArmor(nil, true); got != armorNothing {
t.Errorf("readable config = %v, want armorNothing", got)
}
}
// TestRefreshBootArmorTracksDesiredState is W1's durable half: the file
// /etc/init.d/shater-armor loads at START=21 only exists while the operator wants
// it to. Writing it is half the contract; REMOVING it when the stack is switched
// off or the kill switch is opened is the other half, and the more dangerous one
// to get wrong — a stale armor would reinstate, at the next power cut, a block the
// operator had already turned off.
//
// RED BEFORE: neither the file nor this function existed.
func TestRefreshBootArmorTracksDesiredState(t *testing.T) {
orig := netplane.BootArmorPath
netplane.BootArmorPath = filepath.Join(t.TempDir(), "shater", "boot.nft")
defer func() { netplane.BootArmorPath = orig }()
logger := log.StdLogger()
refreshBootArmor(enabledClosed(), logger)
if !netplane.BootArmorPresent() {
t.Fatalf("an enabled, fail-closed config must persist a boot armor")
}
b, err := os.ReadFile(netplane.BootArmorPath)
if err != nil {
t.Fatalf("read: %v", err)
}
// It must be the HOLDING plane — a forward chain that drops — and not the full
// tproxy ruleset, which would reference an engine that is not running at boot.
for _, must := range []string{"table inet shater", "hook forward", "drop"} {
if !strings.Contains(string(b), must) {
t.Errorf("the persisted armor must contain %q; got:\n%s", must, string(b))
}
}
if strings.Contains(string(b), "tproxy") {
t.Errorf("the persisted armor must NOT divert to an engine that is not running:\n%s", string(b))
}
openKS := enabledClosed()
openKS.Globals.KillSwitch = "open"
refreshBootArmor(openKS, logger)
if netplane.BootArmorPresent() {
t.Errorf("kill_switch=open is a documented choice to let traffic through; " +
"the boot armor must be removed, not left to block the next boot")
}
refreshBootArmor(enabledClosed(), logger)
if !netplane.BootArmorPresent() {
t.Fatalf("re-arming after a disarm must work")
}
off := enabledClosed()
off.Globals.Enabled = false
refreshBootArmor(off, logger)
if netplane.BootArmorPresent() {
t.Errorf("globals.enabled=0 must remove the boot armor")
}
}
// TestRestartHandoffPending pins the marker's read side, including the default
// that matters: an ABSENT marker means "a real stop". Defaulting the other way
// would leave a stopped router blocked whenever the init script failed to write
// the flag.
func TestRestartHandoffPending(t *testing.T) {
orig := restartHandoffPath
defer func() { restartHandoffPath = orig }()
dir := t.TempDir()
restartHandoffPath = filepath.Join(dir, "shater.restarting")
if restartHandoffPending() {
t.Errorf("an absent marker must read as a real stop")
}
if err := os.WriteFile(restartHandoffPath, nil, 0o644); err != nil {
t.Fatalf("write marker: %v", err)
}
if !restartHandoffPending() {
t.Errorf("a present marker must read as a restart handoff")
}
}
+169
View File
@@ -0,0 +1,169 @@
// Executing the SHIPPED init script's action classification.
//
// WHY THIS TEST IS SHAPED LIKE THIS
//
// The boot-armor defect that shipped in v0.2.17 was not in any Go file. The Go
// half was correct and fully covered: refreshBootArmor tracked desired state,
// planExitArmor made the right call, LoadBootArmor validated before loading.
// Every one of those tests was green while the feature did not work at all on
// hardware, because the thing that broke it was one shell `case` in
// /etc/init.d/shater whose default arm swept up procd's `shutdown` action — so
// the arm token was deleted on the way down, every reboot, and the boot it
// existed to protect always found no file.
//
// A unit test that cannot see the shell file cannot catch that, and a comment in
// the shell file claiming `shutdown` is handled is precisely what shipped. So
// this runs the real thing: it sources the actual packaged
// openwrt/shater-core/files/etc/init.d/shater in /bin/sh and calls its two
// classification predicates with every action procd actually uses.
//
// Sourcing the whole file is safe and deliberate — at top level it contains only
// variable assignments and function definitions, nothing that touches the system —
// and sourcing the WHOLE file is the point: a test that copy-pasted the `case`
// would pass while the shipped script said something else.
//
// The action names are not invented. They were measured on the target
// (ImmortalWrt 25.12.1 r37978) with a throwaway probe init script:
//
// /etc/init.d/X restart -> stop_service action=[restart]
// /etc/init.d/X stop -> stop_service action=[stop]
// /etc/init.d/X reload -> reload_service action=[reload]
// `reboot` -> stop_service action=[shutdown]
// the boot after it -> start_service action=[boot]
package main
import (
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
)
// initScriptPath is the packaged init script, relative to this package dir.
const initScriptPath = "../../../openwrt/shater-core/files/etc/init.d/shater"
// askInitScript sources the init script in /bin/sh and reports whether fn
// returns true for the given arguments.
func askInitScript(t *testing.T, fn string, args ...string) bool {
t.Helper()
abs, err := filepath.Abs(initScriptPath)
if err != nil {
t.Fatalf("resolve %s: %v", initScriptPath, err)
}
// `. script` then call the predicate. `set -e` is deliberately NOT used: the
// predicates report by exit status, and a false answer is not an error.
script := `. "$1" || exit 3; shift; if ` + fn + ` "$@"; then echo yes; else echo no; fi`
argv := append([]string{"-c", script, "sh", abs}, args...)
out, err := exec.Command("/bin/sh", argv...).CombinedOutput()
if err != nil {
t.Fatalf("%s(%q): %v\n%s", fn, args, err, out)
}
switch strings.TrimSpace(string(out)) {
case "yes":
return true
case "no":
return false
default:
t.Fatalf("%s(%q): unreadable answer %q", fn, args, out)
return false
}
}
// TestInitScriptActionClassification pins the two closed lists. The `shutdown`
// rows are the regression: both must be false, because a reboot is neither a
// handoff (nothing is coming) nor an operator switching the product off.
func TestInitScriptActionClassification(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, tc := range []struct {
action string
disarms bool
handoff bool
why string
}{
{"stop", true, false, "the operator switched the product off"},
{"shutdown", false, false, "REBOOT/POWEROFF — must not disarm; this is the boot the armor exists for"},
{"restart", false, true, "a successor is coming"},
{"reload", false, true, "Save & Apply is stop+start"},
{"boot", false, false, "start side, never reaches stop_service"},
{"start", false, false, "start side"},
{"", false, false, "unknown/empty degrades to changing nothing"},
{"enable", false, false, "not a lifecycle transition"},
{"disable", false, false, "durable off, but handled by shater-armor's rc.d refusal, not here"},
} {
if got := askInitScript(t, "shater_action_disarms", tc.action); got != tc.disarms {
t.Errorf("shater_action_disarms(%q) = %v, want %v (%s)", tc.action, got, tc.disarms, tc.why)
}
if got := askInitScript(t, "shater_action_handoff", tc.action); got != tc.handoff {
t.Errorf("shater_action_handoff(%q) = %v, want %v (%s)", tc.action, got, tc.handoff, tc.why)
}
}
}
// TestInitScriptStopDisarmsOnlyForAPerson pins the rest of the decision: `stop`
// disarms when a PERSON is behind it, or when the product is being removed — and
// not when something is merely replacing it.
//
// base-files' default_prerm reaches stop_service as a plain `stop`:
//
// if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
// "$i" stop
//
// so two different intentions arrive as one action. The rc.d state separates them:
// a removal has already run `disable`, a replacement has not.
//
// On this target an apk UPGRADE turns out never to run default_prerm at all
// (no pre-upgrade script — verified with a real `apk fix --reinstall` while
// sampling the armor file), so the upgrade rows below are defence-in-depth rather
// than a reproduction. The removal rows are live behaviour.
func TestInitScriptStopDisarmsOnlyForAPerson(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, tc := range []struct {
action string
inPkg string
rcEnable string
want bool
why string
}{
{"stop", "0", "1", true, "an operator typed it — the escape hatch must keep working"},
{"stop", "0", "0", true, "an operator typed it on an already-disabled service"},
{"stop", "1", "1", false, "BEING REPLACED — prerm left the service enabled, so something is coming back"},
{"stop", "1", "0", true, "REMOVAL — prerm already ran `disable`; the product is going away"},
{"shutdown", "0", "1", false, "reboot never disarms"},
{"shutdown", "1", "1", false, "reboot never disarms, package manager or not"},
{"shutdown", "1", "0", false, "still a reboot; the action decides first"},
{"restart", "0", "1", false, "a successor is coming"},
{"restart", "1", "0", false, "a successor is coming; action decides before any state"},
{"reload", "1", "1", false, "Save & Apply"},
{"", "1", "0", false, "unknown action changes nothing"},
} {
got := askInitScript(t, "shater_stop_disarms", tc.action, tc.inPkg, tc.rcEnable)
if got != tc.want {
t.Errorf("shater_stop_disarms(%q, in_pkg=%s, rc_enabled=%s) = %v, want %v (%s)",
tc.action, tc.inPkg, tc.rcEnable, got, tc.want, tc.why)
}
}
}
// TestInitScriptsParse is the cheapest possible guard against the class of bug
// that no Go test can otherwise see: a shell file that ships syntactically
// broken. An init script that fails to parse takes the whole service down and
// `go build` is perfectly happy about it.
func TestInitScriptsParse(t *testing.T) {
if runtime.GOOS == "windows" {
t.Skip("needs a POSIX /bin/sh; the gate runs on linux")
}
for _, name := range []string{"shater", "shater-armor", "shater-cron"} {
p, err := filepath.Abs(filepath.Join(filepath.Dir(initScriptPath), name))
if err != nil {
t.Fatalf("resolve %s: %v", name, err)
}
if out, err := exec.Command("/bin/sh", "-n", p).CombinedOutput(); err != nil {
t.Errorf("/etc/init.d/%s does not parse: %v\n%s", name, err, out)
}
}
}
+65 -5
View File
@@ -289,10 +289,27 @@ func cmdRun() int {
// reconcile. No-op when no iface-driven profiles are configured.
go watchActiveProfile(applier, logger)
// Keep the persisted fail-closed plane in step with the config BEFORE anything
// is attempted: it is what protects the LAN at the NEXT boot (and across the
// next restart), and an engine start that hangs for a minute must not be what
// stands between a config change and its armor being written.
if readErr == nil {
refreshBootArmor(m, logger)
}
// Initial apply. A failed initial apply must NOT crash-loop the box into a
// blackout: log it and stay up so a later SIGHUP/apply can fix the config.
if readErr != nil {
logger.Error("initial ReadUCI failed (staying up): ", readErr)
// ...but STAYING UP IS NOT THE SAME AS BEING SAFE. This branch used to end
// here, which meant an unreadable /etc/config/shater — a full /overlay caught
// mid `uci commit` is the documented cause — left the router with no table at
// all, permanently: nothing else installs one (Reconcile returns before any
// plane work, and the cron watchdog's own `uci -q get` fails identically), and
// the panel reported a live daemon the whole time. The persisted holding plane
// is the last thing this router is KNOWN to have wanted, and it needs nothing
// readable to be true.
armOnUnreadableConfig(logger)
} else if m.Globals.Enabled {
// ARM FIRST, THEN TRY. Install the fail-closed holding plane BEFORE the
// engine is attempted, so the gap between daemon start and a working engine
@@ -369,9 +386,20 @@ func cmdRun() int {
// learn whether the plane is meant to be up (globals.enabled) so a
// reconcile error can raise the kill-switch alert.
enabled := false
if mm, e := model.ReadUCI(); e == nil {
if mm, e := model.ReadUCI(); e != nil {
// The live config just became unreadable. Same answer as at startup:
// reinstate what this router last persisted, rather than run on with
// whatever the kernel happens to hold.
logger.Error("reconcile: could not read the config: ", e)
armOnUnreadableConfig(logger)
} else {
notifier.Update(mm.Alerts)
enabled = mm.Globals.Enabled
// Keep the persisted fail-closed plane in step with the config the
// operator just changed — including the disarm half, so turning the
// stack off (or opening the kill switch) also stops the next boot from
// blocking the LAN.
refreshBootArmor(mm, logger)
// Pick up a changed stats backend / sizing without a daemon restart.
// A no-op when nothing changed, so a routine reconcile never churns
// the store (and never restarts its log cursors).
@@ -393,8 +421,31 @@ func cmdRun() int {
// DNS-query manager). Re-point the aggregator at the current manager.
statsAgg.Resubscribe()
case syscall.SIGTERM, syscall.SIGINT:
logger.Info("signal ", sig, ": honest teardown + exit")
if err := applier.Teardown(); err != nil {
// Being REPLACED is not the same as being switched off, and until now
// this path could not tell the difference: Teardown does not consult
// kill_switch at all (compare holdLocked, which does), so `restart`,
// `reload_service` — which is stop+start, i.e. every LuCI Save & Apply —
// and every package upgrade dismantled the fail-closed plane and left the
// LAN forwarding in the clear for as long as the successor needed to come
// up. That interval is guaranteed non-empty by the init itself: it waits
// for this process to exit, then runs `shaterd migrate`, then starts the
// daemon, which then has to build an engine.
//
// The init script announces a restart with a tmpfs marker; absent it, this
// is a deliberate stop and the plane goes away for good, which is the
// operator's escape hatch and must keep working.
handoff := restartHandoffPending()
logger.Info("signal ", sig, ": honest teardown + exit (restart handoff: ", handoff, ")")
// BEFORE the teardown, not after. This used to read `Teardown(); armOnExit()`
// on the reasoning that Teardown deletes the table so arming first would be
// undone — true, and the wrong conclusion: it left a measured 80-90 ms window per
// restart (80-90 ms measured) in which no `inet shater` table existed at all and fw4's
// `lan -> wan ACCEPT` was the only policy on the box. TeardownExiting arms
// first (one nft transaction that REPLACES the table) and then skips the
// delete iff a plane really went in.
if err := applier.TeardownExiting(func() bool {
return armOnExit(handoff, logger)
}); err != nil {
logger.Error("teardown: ", err)
}
return 0
@@ -991,14 +1042,23 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
writeLine(conn, string(b))
case "reconcile":
ch, err := a.Reconcile()
// The panel and the CLI reconcile over this socket, not by SIGHUP, so the
// persisted fail-closed plane has to be refreshed here too — otherwise
// enabling the stack from the panel left the next boot unarmed (and
// DISABLING it left the next boot armed) until some unrelated SIGHUP
// happened along.
if mm, e := model.ReadUCI(); e == nil {
refreshBootArmor(mm, l)
}
writeResult(conn, ch, err)
case "apply":
a.Snapshot()
changed, err := a.Reconcile()
if err == nil {
if m, e := model.ReadUCI(); e == nil {
if m, e := model.ReadUCI(); e == nil {
if err == nil {
a.ArmRollback(m.Globals.ConfirmTimeout)
}
refreshBootArmor(m, l)
}
writeResult(conn, changed, err)
case "confirm":
+22 -3
View File
@@ -104,6 +104,18 @@ func (e *Engine) HTTPClient(via string) (*http.Client, error) {
// resolved one (the group test, grouptest.go) guarantee the request cannot end up
// anywhere else — in particular not on the direct outbound, which would report the
// ISP's address as the tunnel's exit address.
//
// # Every caller MUST call CloseIdleConnections on the returned client
//
// A fresh http.Transport is built per call and belongs to that call alone. Its idle
// connections are not ordinary sockets: each is a live proxying session through an
// engine outbound, with a read loop and a write loop of its own, and the DialContext
// closure above captures the outbound OBJECT — so an idle connection keeps a whole
// retired engine generation reachable long after Apply swapped it out and its 5s
// close budget expired. Dropping the client without closing it therefore leaks far
// more than a socket.
//
// grouptest.go and shater/generate/ruleset.go get this right; copy them.
func httpClientVia(ob adapter.Outbound, timeout time.Duration) *http.Client {
transport := &http.Transport{
// Dial the underlying TCP connection through the selected outbound. The
@@ -112,9 +124,16 @@ func httpClientVia(ob adapter.Outbound, timeout time.Duration) *http.Client {
DialContext: func(ctx context.Context, network, addr string) (net.Conn, error) {
return ob.DialContext(ctx, network, M.ParseSocksaddr(addr))
},
ForceAttemptHTTP2: true,
MaxIdleConns: 8,
IdleConnTimeout: 90 * time.Second,
ForceAttemptHTTP2: true,
MaxIdleConns: 8,
// The backstop for a caller that forgets to close, not the intended
// mechanism. 90s (the net/http default this used to carry) is eighteen times
// the engine's 5s close budget, so a single forgotten client could pin a dead
// generation through more than a minute and a half of it. 15s is still ample
// for the reuse this actually buys — a redirect chain or the second request of
// a subscription fetch, both within seconds — while bounding the damage of a
// leak to about one apply cycle.
IdleConnTimeout: 15 * time.Second,
TLSHandshakeTimeout: 15 * time.Second,
ExpectContinueTimeout: 1 * time.Second,
}
+2
View File
@@ -17,6 +17,8 @@
// # What this package covers (Phase-2 MVP gate)
//
// - tproxy inbound (+ mixed/socks/dokodemo local listeners)
// - the synthetic "l3-in" TUN inbound (globals l3_tunnel): the L3 ingress that
// lets the engine carry ICMP, which kernel TPROXY cannot divert at all
// - outbounds: vless, vmess, trojan, shadowsocks, hysteria2, tuic, shadowtls
// with the shared TLS/Reality/uTLS container and ws/grpc/httpupgrade/http/
// quic/xhttp transports, plus per-node multiplex
+145 -2
View File
@@ -9,11 +9,15 @@ import (
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/auth"
"github.com/sagernet/sing/common/json/badoption"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// buildInbounds maps every enabled model.Inbound to a typed sing-box inbound.
// buildInbounds maps every enabled model.Inbound to a typed sing-box inbound,
// then appends the synthetic L3-ingress TUN inbound when globals l3_tunnel
// calls for one (see appendL3TunInbound).
//
// # What each Type really produces
//
@@ -87,7 +91,146 @@ func (b *builder) buildInbounds() []option.Inbound {
seenTag[built.Tag] = true
inbounds = append(inbounds, built)
}
return inbounds
return b.appendL3TunInbound(inbounds, seenTag)
}
// The L3-ingress TUN parameters below are a fixed contract with the netplane:
// the nft prerouting chain stamps netplane.L3Mark on LAN traffic the tunnel can
// only carry at layer 3, ApplyRouting points the L3 table's default route at
// netplane.L3Device, and this package opens the engine's end of that device.
// None of it is operator-tunable, which is why the inbound is synthesised from
// globals instead of being a model.Inbound.
const (
// l3InboundTag is how route rules and diagnostics address the L3 ingress.
l3InboundTag = "l3-in"
// l3MTU is 65535 — the largest total length an IPv4 datagram can carry —
// and that maximum IS the reason, not a round number.
//
// This MTU is NOT a tunnel budget. It decides exactly one thing: whether
// the KERNEL fragments a packet on its way INTO the device. What the
// engine then puts into the tunnel is sized against the OUTBOUND's own
// MTU: sing-tun's ForwardDispatcher.forwardToPort measures every forwarded
// packet against Port.PortMTU() — the WireGuard/AWG endpoint's MTU — and
// either fragments to it (no DF) or answers a proper `fragmentation
// needed` quoting it (DF). It already does that work correctly; it only
// has to be handed a WHOLE packet to do it.
//
// The previous value, 1420, was the WireGuard payload budget copied one
// layer too far out, and it was not merely useless — it was a forgery
// generator. Anything bigger was fragmented by the kernel at this device,
// and sing-tun's dispatcher returns on `parsed.fragment` BEFORE asking for
// a routing verdict at all (flow_dispatch.go). The fragments then reached
// the gVisor stack, which reassembled them and handed the echo to
// ICMPForwarder.HandlePacket, whose installFlow demands an UNSPECIFIED
// port address that a WireGuard endpoint never has — so it fell through
// and ANSWERED THE ECHO ITSELF. Net effect: `ping -s 1392` was honest and
// `ping -s 1393` was a lie told by the router. See D25.
//
// 65535 is chosen over any other large value because no IP datagram can
// exceed it: the kernel therefore CANNOT fragment at this device, for any
// packet, ever. Any smaller value leaves a band open and re-opens the bug.
// It is also sing-box's own default TUN MTU on Linux
// (protocol/tun/inbound.go), so it is a well-trodden value.
//
// It costs NO memory, and that was measured rather than assumed: three
// paired runs of TestIntegrationL3TunInboundStarts under
// -test.memprofilerate=1 (exact accounting, not sampled) allocate 5.41 /
// 5.48 / 5.47 MB at 65535 against 5.76 / 5.46 / 5.70 MB at 1420, and a
// -diff_base profile attributes every difference to netlink interface
// enumeration, not to the MTU. Nothing in the read path scales with it:
// gVisor reads through fdbased.BufConfig, which sing-tun's init pins to a
// single 65535-byte view regardless of MTU, and fdbased keeps `mtu` only
// to return it from MTU().
//
// Do NOT expect a saving from GSO either, tempting as the arithmetic is.
// protocol/tun computes `enableGSO = stack == gvisor && mtu < 49152`, so
// this MTU turns it off there — and then StartStateStart turns it back ON
// unconditionally because an adapter.FlowOutbound exists in the config
// (protocol/tun/inbound.go, the outbound/endpoint scan). The ~1.98 MB of
// TCP/UDP GRO scaffolding is therefore present at BOTH MTUs; it is priced
// by the presence of a flow-capable outbound, not by this number.
l3MTU = 65535
)
// l3Addr4/l3Addr6 are the device's point-to-point addresses — the kernel only
// routes over an interface that has one. The prefixes are deliberately tiny
// (/30, /126) and from private space no sane LAN uses, so they cannot shadow a
// real subnet.
var (
l3Addr4 = netip.MustParsePrefix("172.19.242.1/30")
l3Addr6 = netip.MustParsePrefix("fdfe:d3ad:b33f::1/126")
)
// appendL3TunInbound appends the synthetic L3-ingress TUN inbound when globals
// l3_tunnel opted in. Kernel TPROXY diverts nothing but TCP/UDP, so the rest of
// the LAN's traffic (in practice: ICMP echo) can reach the engine only as raw
// IP packets through a TUN device; the netplane policy-routes such packets into
// netplane.L3Device and the engine picks them up here.
//
// There is deliberately no loop-guard RoutingMark: a TUN inbound is not a
// socket (option.TunInboundOptions carries no ListenOptions), and the loop
// guard already lives on every outbound dialer (DialerOptions.RoutingMark =
// LoopMark), which is what keeps the engine's own egress out of the divert.
//
// Warnings carry the `icmp "tunnel"` entity prefix — the same entity the
// netplane's L3 warnings use — so the panel groups everything about the L3
// ingress under one section instead of scattering it.
func (b *builder) appendL3TunInbound(inbounds []option.Inbound, seenTag map[string]bool) []option.Inbound {
if !b.m.Globals.L3Tunnel {
return inbounds
}
hasTproxy := false
for _, in := range inbounds {
if in.Type == C.TypeTProxy {
hasTproxy = true
break
}
}
if !hasTproxy {
// The L3 ingress rides the same LAN divert plane as tproxy: with no
// tproxy inbound the netplane raises no divert chain and marks nothing,
// so the TUN would sit dark while the config claims ICMP is tunnelled.
b.warnf("icmp \"tunnel\": l3_tunnel is on but no tproxy inbound is enabled — the L3 ingress only receives LAN traffic the tproxy divert plane marks for it, so it would carry nothing; the %q inbound is not started (enable a tproxy inbound to restore it)", l3InboundTag)
return inbounds
}
if seenTag[l3InboundTag] {
// Unreachable today — inboundTag prefixes every model-derived tag with
// "in-" — but kept on the same fail-degraded principle as the guards in
// buildInbounds: losing the L3 ingress is strictly smaller damage than
// two inbounds fighting over one tag.
b.warnf("icmp \"tunnel\": inbound tag %q is already taken — the L3 ingress inbound is skipped so the existing listener keeps working (rename the clashing inbound to restore it)", l3InboundTag)
return inbounds
}
addrs := badoption.Listable[netip.Prefix]{l3Addr4}
if b.m.Globals.IPv6 {
addrs = append(addrs, l3Addr6)
}
return append(inbounds, option.Inbound{
Type: C.TypeTun,
Tag: l3InboundTag,
Options: &option.TunInboundOptions{
InterfaceName: netplane.L3Device,
MTU: l3MTU,
Address: addrs,
// AutoRoute MUST stay false. auto_route rewrites the router's MAIN
// routing table and would drag everything the router itself sends —
// WAN traffic, DNS, the tunnel's own underlay — into this TUN. The
// netplane installs the scoped rule/route (fwmark L3Mark -> L3Table
// -> L3Device) itself; that is the whole routing story.
AutoRoute: false,
// gvisor is a CHOICE, not a necessity, and the tempting reason for
// it is wrong: BOTH sing-tun stacks forward ICMP through the same
// ForwardDispatcher first, and both forge an echo reply only for
// what that dispatcher declined (system: stack_system.go
// dispatchIPv4 -> processIPv4ICMP; gvisor: stack_gvisor_filter.go
// -> ICMPForwarder.HandlePacket). Do not re-derive this as "the
// system stack fakes ping" — D25 says so explicitly. gvisor is
// picked because it is already linked (with_wireguard requires
// with_gvisor, D23), so it costs no build tag and no new code path,
// and because it is the combination the integration test exercises.
Stack: "gvisor",
},
})
}
// listenKey returns the "addr:port" an inbound binds, for the clash guard. ok is
+126 -1
View File
@@ -17,12 +17,20 @@ import (
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// buildInboundsOf runs buildInbounds over a single inbound and returns the
// emitted inbounds plus the warnings.
func buildInboundsOf(in model.Inbound) ([]option.Inbound, []string) {
b := newBuilder(&model.Model{Globals: model.DefaultGlobals(), Inbounds: []model.Inbound{in}})
return buildInboundsWith(model.DefaultGlobals(), in)
}
// buildInboundsWith is buildInboundsOf with explicit globals and any number of
// inbounds — for the knobs (l3_tunnel, ipv6) that decide WHICH inbounds exist
// rather than how a single one is shaped.
func buildInboundsWith(g model.Globals, ins ...model.Inbound) ([]option.Inbound, []string) {
b := newBuilder(&model.Model{Globals: g, Inbounds: ins})
return b.buildInbounds(), b.warnings
}
@@ -305,3 +313,120 @@ func TestSniffIsNotAnInboundField(t *testing.T) {
t.Fatal("legacy inbound sniff fields must stay empty (rejected by sing-box 1.13+)")
}
}
// tunInbounds filters the emitted inbounds down to the TUN ones. The synthetic
// L3 ingress is their only source — TestInboundTypeMapping pins that a user
// inbound of type "tun" stays rejected.
func tunInbounds(ins []option.Inbound) []option.Inbound {
var tuns []option.Inbound
for _, in := range ins {
if in.Type == C.TypeTun {
tuns = append(tuns, in)
}
}
return tuns
}
// lanTproxy is the minimal enabled tproxy inbound the L3 tests pair with: the
// L3 ingress only exists alongside the LAN divert plane.
func lanTproxy() model.Inbound {
return model.Inbound{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12345, TCP: true, UDP: true}
}
// TestL3TunnelOffEmitsNoTunInbound: l3_tunnel is OPT-IN. The synthetic TUN
// inbound brings a device, a gVisor netstack and new routing behaviour with it;
// a router whose operator never asked must come through an upgrade unchanged.
func TestL3TunnelOffEmitsNoTunInbound(t *testing.T) {
ins, _ := buildInboundsWith(model.DefaultGlobals(), lanTproxy())
if got := tunInbounds(ins); len(got) != 0 {
t.Fatalf("l3_tunnel defaults to off, yet a TUN inbound was emitted — the L3 ingress grew out of an upgrade nobody opted into")
}
}
// TestL3TunnelEmitsTunInbound pins the data-plane contract of the L3 ingress.
// Every field is load-bearing: the netplane points the L3 table's default
// route at exactly netplane.L3Device, auto_route off keeps sing-box away from
// the router's MAIN routing table, and only the gVisor stack actually forwards
// ICMP — the system stack forges echo replies locally.
func TestL3TunnelEmitsTunInbound(t *testing.T) {
g := model.DefaultGlobals()
g.L3Tunnel = true
ins, warns := buildInboundsWith(g, lanTproxy())
tuns := tunInbounds(ins)
if len(tuns) != 1 {
t.Fatalf("want exactly one L3 TUN inbound, got %d (warns=%v)", len(tuns), warns)
}
if tuns[0].Tag != "l3-in" {
t.Fatalf("tag = %q, want %q — route rules address the L3 ingress by exactly this tag, so any other spelling detaches it from its routing", tuns[0].Tag, "l3-in")
}
to, ok := tuns[0].Options.(*option.TunInboundOptions)
if !ok {
t.Fatalf("options are %T, want *option.TunInboundOptions", tuns[0].Options)
}
if to.InterfaceName != netplane.L3Device {
t.Fatalf("interface_name = %q, want netplane.L3Device (%q) — the netplane routes marked LAN traffic into that exact device, so any other name leaves a TUN nothing feeds", to.InterfaceName, netplane.L3Device)
}
if to.AutoRoute {
t.Fatal("auto_route is on — sing-box would rewrite the router's MAIN routing table and drag the router's own WAN/DNS/underlay traffic into the tunnel; the netplane owns the scoped L3 rules")
}
if to.Stack != "gvisor" {
t.Fatalf("stack = %q, want gvisor — the system stack answers ICMP echo locally instead of forwarding it, which is the exact forged reply l3_tunnel exists to remove", to.Stack)
}
if to.MTU != 65535 {
t.Fatalf("mtu = %d, want 65535 — see TestL3TunnelMTULeavesNothingForTheKernelToFragment for why the number is the maximum and not a tunnel budget", to.MTU)
}
if len(to.Address) != 2 || to.Address[0].String() != "172.19.242.1/30" || to.Address[1].String() != "fdfe:d3ad:b33f::1/126" {
t.Fatalf("address = %v, want [172.19.242.1/30 fdfe:d3ad:b33f::1/126] — the netplane's routes are built against exactly these prefixes", to.Address)
}
}
// TestL3TunnelIPv6OffKeepsV4Only: with globals ipv6 off the netplane installs
// no v6 rules, so a v6 prefix on the TUN would advertise an ICMPv6 path that
// dead-ends inside the device.
func TestL3TunnelIPv6OffKeepsV4Only(t *testing.T) {
g := model.DefaultGlobals()
g.L3Tunnel = true
g.IPv6 = false
ins, _ := buildInboundsWith(g, lanTproxy())
tuns := tunInbounds(ins)
if len(tuns) != 1 {
t.Fatalf("want the L3 TUN inbound, got %d", len(tuns))
}
to, ok := tuns[0].Options.(*option.TunInboundOptions)
if !ok {
t.Fatalf("options are %T, want *option.TunInboundOptions", tuns[0].Options)
}
if len(to.Address) != 1 || !to.Address[0].Addr().Is4() {
t.Fatalf("address = %v — on an ipv6=off router only the v4 prefix may remain; a v6 address advertises an ICMPv6 path the netplane never routes", to.Address)
}
}
// TestL3TunnelWithoutTproxySkipped: the L3 ingress rides the LAN divert plane
// that only exists alongside a tproxy inbound. Without one the TUN would sit
// dark while the config claims ICMP is tunnelled — so it is skipped, and the
// skip is said out loud.
func TestL3TunnelWithoutTproxySkipped(t *testing.T) {
g := model.DefaultGlobals()
g.L3Tunnel = true
ins, warns := buildInboundsWith(g, model.Inbound{
Name: "sock", Enabled: true, Type: "socks", Listen: "127.0.0.1", Port: 1080, TCP: true, UDP: true,
})
if got := tunInbounds(ins); len(got) != 0 {
t.Fatalf("no tproxy inbound is enabled, yet the L3 TUN inbound was emitted — it would carry nothing while claiming ICMP coverage")
}
if !warnsHave(warns, "no tproxy inbound") {
t.Fatalf("expected a warning explaining the skipped L3 ingress, got %v", warns)
}
if !warnsHave(warns, `icmp "tunnel": `) {
t.Fatalf("the L3 warning must carry the `icmp \"tunnel\"` entity prefix — without it the panel cannot group it with the data plane's L3 warnings and it lands nameless under a bare \"generate\" section; got %v", warns)
}
// The same holds for a config with no inbounds at all (every one disabled).
ins, warns = buildInboundsWith(g)
if got := tunInbounds(ins); len(got) != 0 {
t.Fatalf("an inbound-less config emitted the L3 TUN inbound — it would carry nothing while claiming ICMP coverage")
}
if !warnsHave(warns, "no tproxy inbound") {
t.Fatalf("expected a warning explaining the skipped L3 ingress, got %v", warns)
}
}
+160
View File
@@ -0,0 +1,160 @@
//go:build linux
// The egress half of the L3 ICMP story, live-engine part: not the SHAPE of the
// config (l3_egress_test.go pins that portably) but what the running engine
// BELIEVES about the two egress kinds. The belief is the whole feature:
//
// - An interface egress must come up as an adapter.FlowOutbound whose
// PreMatchFlow answers Flow for ICMP. That answer exists only if
// protocol/direct's constructor actually built its ping.Port, which it does
// only when dialer.NewWithOptions hands back a *dialer.DefaultDialer for a
// dialer that carries BindInterface + RoutingMark. Nothing in the portable
// suite can see this: if that cast ever stops holding (an upstream bump
// wrapping the bound dialer, say), the generated JSON stays byte-identical,
// every codegen test stays green, and ping through every interface egress
// silently degrades from "leaves via the second WAN" to "dropped".
// - A byedpi egress must NOT look ICMP-capable: route.preMatchFlow
// (l3-honest-drop) drops an ICMP flow whose outbound either lacks icmp in
// Network() or is not a FlowOutbound, and SOCKS satisfies both refusals. If
// it ever stops refusing, the drop stops happening — and the TUN stack's
// alternative is forging the echo reply itself.
//
// Gating mirrors l3_integration_linux_test.go, whose comment carries the full
// argument: the model opts into l3_tunnel, so Start opens /dev/net/tun and
// needs root + CAP_NET_ADMIN, which the ordinary gate's containers do not
// expose; the TestIntegration prefix and the honest skips below keep the gate
// green while telling a human exactly how to run this for real.
package generate
import (
"net/netip"
"os"
"slices"
"testing"
"github.com/sagernet/sing-box/adapter"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// l3EgressICMPModel is the smallest l3_tunnel=1 config that carries both
// egress kinds: the mandatory tproxy divert plane, a resolver, one interface
// egress and one byedpi egress, each with a rule routing into it.
//
// The interface egress binds to `lo` — deliberately: BindInterface must name a
// device that EXISTS on the runner (netplane.IfaceDevice passes an
// unresolvable name through unchanged, and every kernel has lo), and this test
// never dials, so nothing actually leaves through it. IPv6 is off for the same
// reason as the sibling model: a disable_ipv6=1 host must not masquerade as
// the regression this test hunts.
func l3EgressICMPModel() *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
g.ResolverDefault = "cf"
g.L3Tunnel = true
g.IPv6 = false
return &model.Model{
Globals: g,
Inbounds: []model.Inbound{
// 12404: next free port above the package's hand-allocated tproxy
// band (12403 belongs to l3_integration_linux_test.go). Start
// binds for real, so a clash with a sibling would fail this test
// for reasons that have nothing to do with the egresses.
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12404, TCP: true, UDP: true},
},
Resolvers: []model.Resolver{
{Name: "cf", Type: "doh", Address: "https://1.1.1.1/dns-query", Detour: "direct"},
},
Egresses: []model.Egress{
{Name: "lo", Type: "interface", Interface: "lo"},
{Name: "bd", Type: "byedpi", Port: 1080},
},
Rules: []model.Rule{
{Name: "ping-lo", Enabled: true, Order: 10, Src: []string{"192.168.88.0/24"}, Target: "egress:lo"},
{Name: "desync", Enabled: true, Order: 20, Src: []string{"192.168.89.0/24"}, Target: "egress:bd"},
},
}
}
// TestIntegrationL3EgressICMPIsAFlow proves the live engine's verdict on ICMP
// through each egress kind, in failure-mode order:
//
// 1. the interface egress outbound is an adapter.FlowOutbound advertising
// icmp — the two static gates route.preMatchFlow checks before it even
// asks the outbound;
// 2. its PreMatchFlow(icmp) answers Flow — the dynamic gate, true only when
// the ping.Port was really constructed despite BindInterface+RoutingMark
// on the dialer. THE assertion of this file: its failure mode is a ping
// that silently turns into a drop with not one generated byte changed;
// 3. the byedpi egress outbound fails at least one of the same static gates,
// which is precisely what makes l3-honest-drop DROP a ping routed at it
// instead of the TUN stack forging the echo reply.
func TestIntegrationL3EgressICMPIsAFlow(t *testing.T) {
if os.Geteuid() != 0 {
t.Skipf("needs root to open and configure a TUN device (euid=%d) — run as root with CAP_NET_ADMIN and /dev/net/tun, e.g. on the OpenWrt VM or via `docker run --cap-add NET_ADMIN --device /dev/net/tun`", os.Geteuid())
}
if _, err := os.Stat("/dev/net/tun"); err != nil {
t.Skipf("/dev/net/tun is not available (%v) — expose it (modprobe tun; in docker: --device /dev/net/tun --cap-add NET_ADMIN) and run as root", err)
}
opts, warns, err := GenerateWithWarnings(l3EgressICMPModel())
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
if len(warns) != 0 {
t.Fatalf("generate degraded the model (warnings: %v) — the egress and L3 skip paths warn instead of failing, so a warning here usually means an egress outbound or the TUN inbound was silently dropped and the engine verdicts below would prove nothing", warns)
}
e := engine.New()
// Idempotent; covers every Fatalf below. The wait is NOT decoration: this
// test opens netplane.L3Device, and the name is a singleton, so yielding
// before the kernel has taken it back leaves the next test in this package
// to meet `TUNSETIFF: device or resource busy` and fail for a reason that
// has nothing to do with what it asserts.
t.Cleanup(func() {
_ = e.Close()
l3WaitDeviceGone(t)
})
if _, err := e.Apply(opts); err != nil {
t.Fatalf("engine.Apply (box.New validate + Start) rejected the config: %v\n%s", err, l3StartFailureHint(err))
}
om := e.Instance().Outbound()
if om == nil {
t.Fatalf("running box has no OutboundManager — nothing below could prove anything")
}
// 1+2. The interface egress: the engine must consider it ICMP-capable, and
// capable FOR REAL (the ping.Port exists), not just by interface shape.
loTag := netplane.EgressOutboundTag("lo")
loOb, ok := om.Outbound(loTag)
if !ok {
t.Fatalf("running box has no outbound %q — every rule bound to this egress resolved into nothing, so its traffic is fail-closed blocked and the second-WAN path this feature sells does not exist", loTag)
}
if !slices.Contains(loOb.Network(), N.NetworkICMP) {
t.Fatalf("outbound %q Network() = %v, without %q — route.preMatchFlow refuses the flow at its first static gate, so every ping routed through an interface egress is dropped while TCP/UDP keep flowing", loTag, loOb.Network(), N.NetworkICMP)
}
flow, isFlow := loOb.(adapter.FlowOutbound)
if !isFlow {
t.Fatalf("outbound %q (%T) is not an adapter.FlowOutbound — route.preMatchFlow can then never answer Flow for it, so every ping routed through an interface egress is dropped while the config still claims the egress carries L3", loTag, loOb)
}
if got := flow.PreMatchFlow(N.NetworkICMP, netip.MustParseAddr("203.0.113.9")); got != adapter.PreMatchFlow {
t.Fatalf("PreMatchFlow(icmp) = %v, want adapter.PreMatchFlow — the direct outbound started WITHOUT its ping.Port, i.e. dialer.NewWithOptions no longer yields a *dialer.DefaultDialer once BindInterface+RoutingMark are set (protocol/direct only builds icmpPort behind that cast); ping through every interface egress then silently turns into a drop with not one generated byte changed, so only this live check can catch it", got)
}
// 3. The byedpi egress: at least one static gate must refuse it. Both
// refusing is today's reality (SOCKS advertises no icmp and is no
// FlowOutbound); the regression is BOTH passing, because then
// l3-honest-drop stops dropping and the TUN stack answers the echo itself
// — a forged reply from a desync hop that never saw the packet.
bdTag := netplane.EgressOutboundTag("bd")
bdOb, ok := om.Outbound(bdTag)
if !ok {
t.Fatalf("running box has no outbound %q — every rule bound to this egress resolved into nothing, so its domains lost the desync entirely", bdTag)
}
if _, isFlow := bdOb.(adapter.FlowOutbound); isFlow && slices.Contains(bdOb.Network(), N.NetworkICMP) {
t.Fatalf("outbound %q (%T, networks %v) passes both of route.preMatchFlow's static gates for ICMP — l3-honest-drop then no longer drops a ping routed at the byedpi egress, and the TUN stack forges the echo reply locally: the operator reads a working ping off a SOCKS hop that cannot carry the packet", bdTag, bdOb, bdOb.Network())
}
}
+123
View File
@@ -0,0 +1,123 @@
// The egress half of the L3 ICMP story, portable codegen part: WHAT the
// generator must emit so a ping routed to `egress:<name>` behaves honestly.
//
// The chain these tests pin (all upstream sing-box internals, none of them
// visible in the generated JSON): protocol/direct is the ONLY proxy outbound
// in the shipped registry that implements adapter.FlowOutbound — at
// construction it wraps a ping.Port around the dialer's Control chain, and
// common/dialer/default.go appends BindInterface + RoutingMark to exactly
// that chain (dialer.Control -> DefaultDialer.dialer4 ->
// DialerForICMPDestination -> the raw ICMP socket). So the SHAPE asserted
// here — a direct outbound carrying the egress device and the deterministic
// egress mark — is precisely what makes a tunnelled ping leave through the
// right interface under the right policy table. A SOCKS outbound (byedpi)
// sits on the other side of the same line: it cannot implement tun.Port, so
// route.preMatchFlow's l3-honest-drop block DROPS ICMP routed at it instead
// of letting the TUN stack forge an echo reply locally.
//
// Portable (no box.New): these assert on the generated option.Options only.
// The live-engine proof that the direct outbound REALLY constructs its
// icmpPort despite bind+mark lives in l3_egress_linux_test.go.
package generate
import (
"testing"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// l3EgressOutbound finds the outbound emitted under tag. A local twin of
// generate_test.go's findOutbound, which lives in the linux-gated engine
// suite and does not exist on other platforms.
func l3EgressOutbound(opts option.Options, tag string) *option.Outbound {
for i := range opts.Outbounds {
if opts.Outbounds[i].Tag == tag {
return &opts.Outbounds[i]
}
}
return nil
}
// TestEgressInterfaceOutboundCarriesICMP pins the three fields that decide
// whether a ping routed to an interface egress ACTUALLY leaves through that
// interface:
//
// - Type direct — the one registered outbound type whose constructor builds
// a ping.Port (adapter.FlowOutbound); any other type demotes ICMP through
// this egress to the honest drop.
// - BindInterface = the egress device — appended to dialer.Control, which
// DefaultDialer.dialer4 carries and DialerForICMPDestination hands to the
// raw ICMP socket.
// - RoutingMark = the deterministic egress mark — same Control chain; it is
// what the netplane's policy rule matches to steer the packet into the
// egress table and past the tproxy divert.
func TestEgressInterfaceOutboundCarriesICMP(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12345, TCP: true, UDP: true},
},
Egresses: []model.Egress{{Name: "wan2", Type: "interface", Interface: "wan2"}},
Rules: []model.Rule{
{Name: "via-wan2", Enabled: true, Order: 10, Src: []string{"192.168.2.0/24"}, Target: "egress:wan2"},
},
}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
tag := netplane.EgressOutboundTag("wan2")
ob := l3EgressOutbound(opts, tag)
if ob == nil {
t.Fatalf("no outbound %q was emitted (warnings: %v) — every rule bound to this egress then resolves through the fail-closed path (egressDetourOrBlock) and the operator's second WAN silently carries nothing", tag, warns)
}
if ob.Type != C.TypeDirect {
t.Fatalf("egress outbound type = %q, want %q — direct is the only proxy outbound in the shipped registry that implements adapter.FlowOutbound (its constructor builds the ping.Port), so any other type turns every ping routed through this egress into a drop while TCP/UDP keep flowing, and nobody can tell the second WAN's L3 path is dead", ob.Type, C.TypeDirect)
}
do, ok := ob.Options.(*option.DirectOutboundOptions)
if !ok {
t.Fatalf("egress outbound options are %T, want *option.DirectOutboundOptions — without the typed dialer options there is no BindInterface/RoutingMark to reach the ICMP socket's Control chain at all", ob.Options)
}
if want := netplane.IfaceDevice("wan2"); do.BindInterface != want {
t.Fatalf("BindInterface = %q, want %q — this field is appended to dialer.Control (common/dialer/default.go), which lands in DefaultDialer.dialer4, whose Control DialerForICMPDestination hands to the ICMP socket; losing it sends the echo over the MAIN routing table, i.e. out the plain default WAN instead of this egress — a silent leak, not a visible failure", do.BindInterface, want)
}
if want := option.FwMark(netplane.EgressMark(m.Globals, 0)); do.RoutingMark != want {
t.Fatalf("RoutingMark = %#x, want %#x — the mark rides the same Control chain into the ICMP socket, and it is what the netplane's per-egress policy rule matches; without it the egress's own packets are routed by the main table (leaking past the egress) or re-caught by the tproxy divert (a routing loop)", uint32(do.RoutingMark), uint32(want))
}
}
// TestEgressByeDPIOutboundCannotCarryICMP pins the OTHER side of the line: a
// byedpi egress is a SOCKS5 hop into the local ciadpi desync proxy, and SOCKS
// does not (and cannot) implement tun.Port, so route.preMatchFlow's
// l3-honest-drop block DROPS ICMP routed at it. That drop is the feature: the
// only alternative the TUN stack offers is answering the echo ITSELF
// (stack_gvisor_icmp.go), i.e. a forged reply from a path that never saw the
// packet. Pinning the SOCKS type here pins the reason the drop happens.
func TestEgressByeDPIOutboundCannotCarryICMP(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12345, TCP: true, UDP: true},
},
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: 1080}},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, Src: []string{"192.168.3.0/24"}, Target: "egress:bd"},
},
}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
tag := netplane.EgressOutboundTag("bd")
ob := l3EgressOutbound(opts, tag)
if ob == nil {
t.Fatalf("no outbound %q was emitted (warnings: %v) — every rule bound to this egress then resolves through the fail-closed path and the desync stops covering its domains", tag, warns)
}
if ob.Type != C.TypeSOCKS {
t.Fatalf("byedpi egress outbound type = %q, want %q — the type is load-bearing twice over: only a SOCKS hop actually reaches the local ciadpi process (anything else skips the desync entirely), and its inability to implement tun.Port is exactly what makes the l3-honest-drop block DROP a ping routed here instead of the TUN stack forging an echo reply from a path that never carried the packet", ob.Type, C.TypeSOCKS)
}
}
@@ -0,0 +1,217 @@
//go:build linux
// The engine half of the L3-ingress proof — the one that needs a real kernel.
//
// Everything else about l3_tunnel is already pinned by unprivileged tests: the
// portable codegen suite (inbound_test.go TestL3Tunnel*) fixes the SHAPE of the
// synthetic "l3-in" TUN inbound (tag, netplane.L3Device name, auto_route off,
// gvisor stack, MTU, addresses), and the netplane tests fix the nft marks and
// the ip rule/route plan around it. What NONE of them prove is that the engine
// actually accepts that config on the router build: shaterd swaps upstream's
// include.Context for the slim shater/registry (engine.New wires
// registry.Context), where tun.RegisterInbound is a deliberate hand-kept
// entry, and the gVisor netstack the inbound demands (stack: gvisor) is only
// compiled in because with_wireguard drags with_gvisor along
// (scripts/router-tags.sh). Losing either — the registry entry deleted as
// "unused", the tag trimmed for size — changes no generated byte, so every
// codegen assertion stays green, and the failure surfaces as a dead engine on
// the operator's router at the first l3_tunnel=1 apply (the 2026-07-25
// WireGuard outage was exactly this "built with X, verified with Y" class).
// This file closes that gap: box.New + Start must accept the generated config
// under the slim registry, the kernel must end up with the shater-l3 device
// the netplane routes into (at the contract MTU), and Close must remove it.
//
// Separate file, TestIntegration name, honest skips: opening /dev/net/tun and
// configuring the device needs root + CAP_NET_ADMIN, which the ordinary gate
// does not have — scripts/run-tests.sh reaches linux through containers (the
// act_runner job container, or the docker re-exec from a dev host) that expose
// no /dev/net/tun, and its SKIP_COMMON already excludes common/tlsspoof's
// TestIntegration* for the same capability reason; the TestIntegration prefix
// keeps this test inside that naming convention. The shater/... roots are not
// name-filtered, so what keeps the ordinary gate green here are the guards
// below: no root or no /dev/net/tun means a loud skip that says how to run it
// for real (the OpenWrt VM, or
// `docker run --cap-add NET_ADMIN --device /dev/net/tun`).
package generate
import (
"net"
"os"
"strings"
"testing"
"time"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
// l3GoneTimeout bounds how long assertion 4 waits for the kernel to drop the
// device after Close. Removal is normally immediate with the last fd, but
// unregister_netdevice may defer briefly under load; polling keeps the check
// honest without a flaky fixed sleep.
const l3GoneTimeout = 5 * time.Second
// l3IntegrationModel is the smallest realistic l3_tunnel=1 config: one enabled
// tproxy inbound (the L3 ingress refuses to exist without the divert plane it
// rides), one resolver, one shadowsocks node on TEST-NET and a rule into it —
// the same skeleton as the other *_linux_test.go models, plus the L3 opt-in.
//
// IPv6 is off deliberately: the v6 address half of the TUN contract is pinned
// by the portable codegen tests, and carrying it here would couple THIS proof
// (the tun inbound starts under the slim registry) to the runner kernel's
// ipv6 sysctls — a disable_ipv6=1 host would fail address configuration and
// masquerade as the registry/tag regression this test hunts.
func l3IntegrationModel() *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
g.ResolverDefault = "cf"
g.L3Tunnel = true
g.IPv6 = false
return &model.Model{
Globals: g,
Inbounds: []model.Inbound{
// 12403: first port above the package's hand-allocated tproxy band
// (currently topping out at 12402). Start binds for real, so a
// clash with a sibling's port would fail this test for reasons
// that have nothing to do with the TUN.
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12403, TCP: true, UDP: true},
},
Nodes: []model.Node{
{Name: "exit", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#exit"},
},
Resolvers: []model.Resolver{
{Name: "cf", Type: "doh", Address: "https://1.1.1.1/dns-query", Detour: "direct"},
},
Rules: []model.Rule{
{Name: "via-exit", Enabled: true, Order: 10, DstPort: "443", Target: "node:exit"},
},
}
}
// TestIntegrationL3TunInboundStarts proves the generated l3_tunnel config is
// not merely well-formed but ALIVE: the engine (slim registry, compiled tag
// set) accepts it, the kernel ends up with the device the netplane routes
// into, and Close returns the name. Four assertions, in failure-mode order:
//
// 1. engine.Apply (box.New + Start) succeeds — a rejection here is the slim
// registry losing tun.RegisterInbound or the build losing the gVisor
// stack, see l3StartFailureHint;
// 2. net.InterfaceByName(netplane.L3Device) finds the device — the L3 table's
// default route points at exactly that name, so no device means marked
// LAN ICMP blackholes while the config claims it is tunnelled;
// 3. the kernel ACCEPTS the contract MTU 65535 on a TUN — the whole point of
// the value is that no IP datagram can exceed it, so the kernel can never
// fragment on the way in (D25); a kernel that clamped it would silently
// restore the forged-reply band;
// 4. after Close the device is GONE — a leak would jam every later apply
// (each swap reopens the same name) until shaterd itself is restarted.
func TestIntegrationL3TunInboundStarts(t *testing.T) {
if os.Geteuid() != 0 {
t.Skipf("needs root to open and configure a TUN device (euid=%d) — run as root with CAP_NET_ADMIN and /dev/net/tun, e.g. on the OpenWrt VM or via `docker run --cap-add NET_ADMIN --device /dev/net/tun`", os.Geteuid())
}
if _, err := os.Stat("/dev/net/tun"); err != nil {
t.Skipf("/dev/net/tun is not available (%v) — expose it (modprobe tun; in docker: --device /dev/net/tun --cap-add NET_ADMIN) and run as root", err)
}
opts, warns, err := GenerateWithWarnings(l3IntegrationModel())
if err != nil {
t.Fatalf("Generate: unexpected error: %v", err)
}
if len(warns) != 0 {
t.Fatalf("generate degraded the model (warnings: %v) — the L3 skip paths warn instead of failing, so a warning here usually means the TUN inbound was silently dropped and the engine run below would prove nothing", warns)
}
// Precondition, not the point: the SHAPE of the tun inbound is
// inbound_test.go's job. If it is missing here, fail with the right
// address instead of a misleading "no such interface" three steps later.
hasTun := false
for _, in := range opts.Inbounds {
if in.Type == C.TypeTun {
hasTun = true
break
}
}
if !hasTun {
t.Fatalf("generated options carry no tun inbound — a codegen regression (inbound_test.go TestL3TunnelEmitsTunInbound should be red too), not an engine one; nothing below could prove anything")
}
// 1. The engine — slim registry, the tag set this binary was built with —
// takes the config and starts it.
e := engine.New()
t.Cleanup(func() { _ = e.Close() }) // idempotent; covers every Fatalf below
changed, err := e.Apply(opts)
if err != nil {
t.Fatalf("engine.Apply (box.New validate + Start) rejected the l3_tunnel config: %v\n%s", err, l3StartFailureHint(err))
}
if !changed {
t.Fatalf("expected Apply changed==true on a fresh engine")
}
// 2. The device is REAL. Start constructs the TUN synchronously
// (protocol/tun StartStateStart -> tun.New), so no settling loop is
// needed on this side.
iface, err := net.InterfaceByName(netplane.L3Device)
if err != nil {
t.Fatalf("Start reported success but %q does not exist (%v) — the netplane's L3 table points its default route at exactly that name, so marked LAN ICMP would blackhole while the config claims it is tunnelled", netplane.L3Device, err)
}
// 3. The kernel device carries the contract MTU. This is the assertion the
// value was chosen for: at 65535 no IP datagram can exceed the device MTU,
// so the kernel cannot fragment on the way in and sing-tun'''s dispatcher
// always gets a whole packet to judge. A kernel that silently clamped this
// (or a driver with a lower max_mtu) would put the forged-reply band back
// without changing a single generated byte.
if iface.MTU != 65535 {
t.Fatalf("%s MTU = %d, want 65535 — the kernel did not take the contract MTU; anything smaller means the kernel fragments packets above it INTO this device, sing-tun'''s ForwardDispatcher declines fragments before asking for a verdict, and the gVisor ICMP forwarder forges the echo reply for WireGuard/AWG (D25)", netplane.L3Device, iface.MTU)
}
// 4. Close returns the device name to the kernel.
if err := e.Close(); err != nil {
t.Fatalf("engine Close failed: %v — a box that does not stop keeps %s open, and every later apply that reopens the name starts dead", err, netplane.L3Device)
}
l3WaitDeviceGone(t)
}
// l3WaitDeviceGone blocks until the L3 TUN is back in the kernel's hands.
//
// It is shared rather than inlined because the name is a SINGLETON: two tests
// in this package each stand an engine up on netplane.L3Device, and the second
// one meets `TUNSETIFF: device or resource busy` if the first only closed its
// box and moved on. Removal is normally immediate with the last fd, but
// unregister_netdevice may defer briefly under load, so every test that opens
// the device MUST wait here before yielding it — polling rather than a fixed
// sleep keeps that honest instead of flaky.
func l3WaitDeviceGone(t *testing.T) {
t.Helper()
deadline := time.Now().Add(l3GoneTimeout)
for {
if _, err := net.InterfaceByName(netplane.L3Device); err != nil {
return // gone: the kernel dropped the device with the engine's fd
}
if time.Now().After(deadline) {
t.Fatalf("%s still exists %v after Close — the engine leaked the TUN; every subsequent apply swaps in a fresh box that must reopen this exact name, so from the first leak onward the L3 ingress comes up dead until shaterd is restarted (and a re-run of this suite would inherit the stale device)", netplane.L3Device, l3GoneTimeout)
}
time.Sleep(50 * time.Millisecond)
}
}
// l3StartFailureHint names the specific regression (or environment defect) an
// Apply failure most likely is, so the gate output points at the fix instead
// of a bare engine error. Substring matching is the same pragmatism engine.go
// itself applies to swap conflicts (isAddrInUse/isCacheLockConflict): the
// upstream errors carry no exported sentinels.
func l3StartFailureHint(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "type not found: tun"):
return "hint: the slim registry no longer registers the tun inbound (shater/registry InboundRegistry must call tun.RegisterInbound) — on the router, l3_tunnel=1 then leaves the engine DOWN, and with the fail-closed nft plane that blackholes the whole LAN, not just ICMP"
case strings.Contains(msg, "not included in this build"):
return "hint: the gVisor netstack is compiled out — the tag set lost with_gvisor (scripts/router-tags.sh keeps it via with_wireguard); this is the 2026-07-25 outage class: green codegen, dead engine at the first Start on the operator's router"
case strings.Contains(msg, "operation not permitted"):
return "hint: environment, not code — this runner has root and /dev/net/tun but the kernel refused the device (missing CAP_NET_ADMIN? seccomp?); rerun with --cap-add NET_ADMIN or on the OpenWrt VM"
default:
return "hint: whatever the cause, on the router this means applying l3_tunnel=1 leaves the engine down, and with the fail-closed nft plane that is a LAN-wide outage"
}
}
+88
View File
@@ -0,0 +1,88 @@
package generate
import (
"testing"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// maxIPv4Datagram is the largest value the 16-bit IPv4 Total Length field can
// hold, i.e. the largest IP datagram that can exist on the wire at all. It is
// spelled out here rather than written as a literal because it IS the argument:
// a device whose MTU is this large cannot be fragmented into.
const maxIPv4Datagram = 65535
// TestL3TunnelMTULeavesNothingForTheKernelToFragment is the reason the number
// is the number. Read this before changing l3MTU.
//
// The MTU of `shater-l3` is NOT a tunnel budget. It governs exactly one thing:
// whether the KERNEL splits a packet on its way INTO the device. What the
// engine then puts into the tunnel is sized separately and correctly, against
// the OUTBOUND's own MTU — sing-tun's ForwardDispatcher.forwardToPort measures
// each forwarded packet against Port.PortMTU() and either fragments to it (no
// DF) or answers a `fragmentation needed` quoting it (DF).
//
// The value used to be 1420, the WireGuard payload budget, copied one layer too
// far out. That did not make pings fit the tunnel; it made the kernel fragment
// everything above 1392 bytes of payload right here, and a fragment is the one
// thing sing-tun's dispatcher will not judge: Dispatch returns on
// `parsed.fragment` BEFORE calling JudgeFlow, the fragments fall through to the
// gVisor stack, which reassembles them and hands the echo to
// ICMPForwarder.HandlePacket — whose installFlow requires an UNSPECIFIED port
// address that a WireGuard/AWG endpoint never has. It declines, and HandlePacket
// FORGES the echo reply. So `ping -s 1392` was honest and `ping -s 1393` was a
// lie told by the router (D25).
//
// Hence the invariant, not merely the constant: the MTU must be at least the
// largest datagram that can exist, so that NO packet can ever be fragmented
// into this device. Anything smaller re-opens a band of sizes where a ping
// reads as tunnelled without leaving the router.
func TestL3TunnelMTULeavesNothingForTheKernelToFragment(t *testing.T) {
to := l3TunOptions(t)
if to.MTU < maxIPv4Datagram {
t.Fatalf("l3-in MTU = %d, must be >= %d (the largest IP datagram there can be).\n"+
"Below that the kernel fragments packets between the MTU and the datagram size on their way INTO shater-l3, sing-tun's ForwardDispatcher declines fragments without ever asking for a routing verdict, and the gVisor ICMP forwarder answers the echo ITSELF for any outbound whose port address is not unspecified — every WireGuard/AWG endpoint.\n"+
"The result is not a dropped ping, it is a FORGED reply: the operator reads a working tunnel off a packet that died on the router. Do not size this to the tunnel MTU — forwardToPort already sizes against Port.PortMTU().",
to.MTU, maxIPv4Datagram)
}
// The concrete value, so that a change is a decision and not a drift. 65535
// is also sing-box's own default TUN MTU on Linux (protocol/tun/inbound.go).
if to.MTU != maxIPv4Datagram {
t.Fatalf("l3-in MTU = %d, want exactly %d — larger is impossible on the wire and buys nothing; if you have a reason, put it in the l3MTU comment and change this test deliberately", to.MTU, maxIPv4Datagram)
}
}
// TestL3TunnelMTUIsNotATunnelBudget guards the specific regression: someone
// reading "the tunnel is 1420" and "re-aligning" the device to it. The two
// numbers are unrelated, and making them equal is what produced the forged
// replies in the first place.
func TestL3TunnelMTUIsNotATunnelBudget(t *testing.T) {
to := l3TunOptions(t)
// 1500 is the largest MTU a plain Ethernet LAN hands the router in one
// piece; every plausible tunnel budget (1420 for WireGuard, 1280 for a
// conservative v6 path, 1412 for AWG with headers) sits below it. An l3-in
// MTU anywhere in that range means the kernel is fragmenting into the
// device again.
if to.MTU <= 1500 {
t.Fatalf("l3-in MTU = %d — that is tunnel/LAN-sized, so the kernel will fragment into shater-l3 and the gVisor ICMP forwarder will forge echo replies for WireGuard/AWG. The device MTU and the tunnel MTU are NOT the same number; see l3MTU's comment and D25.", to.MTU)
}
}
// l3TunOptions builds the l3_tunnel-on config and returns the TUN inbound's
// options, failing with a useful address if the inbound is missing entirely.
func l3TunOptions(t *testing.T) *option.TunInboundOptions {
t.Helper()
g := model.DefaultGlobals()
g.L3Tunnel = true
ins, warns := buildInboundsWith(g, lanTproxy())
tuns := tunInbounds(ins)
if len(tuns) != 1 {
t.Fatalf("want exactly one L3 TUN inbound, got %d (warns=%v)", len(tuns), warns)
}
to, ok := tuns[0].Options.(*option.TunInboundOptions)
if !ok {
t.Fatalf("options are %T, want *option.TunInboundOptions", tuns[0].Options)
}
return to
}
+301 -2
View File
@@ -7,9 +7,11 @@ import (
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/json/badoption"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
"github.com/sagernet/sing-box/shater/parse"
)
// buildRoute assembles option.RouteOptions:
@@ -117,6 +119,15 @@ func (b *builder) buildRoute() *option.RouteOptions {
continue
}
// An ICMP rule is the one shape whose traffic reaches the engine at layer 3,
// and layer 3 has prerequisites the rule itself cannot state. Diagnose it
// only when its target RESOLVED: an unresolved one already earned the much
// louder ruleKillFallback warning above, and the kill fallback is a policy
// decision about a broken reference, not about ICMP.
if isICMPProto(r.Proto) && ok {
b.warnICMPRule(r, want)
}
route := option.RouteActionOptions{Outbound: target}
b.applyDPI(&route, target, r.Name)
general = append(general, option.Rule{
@@ -378,6 +389,35 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
switch p {
case "tcp", "udp":
raw.Network = badoption.Listable[string]{p}
case protoICMP, protoICMPv4, protoICMPv6:
// ICMP is a NETWORK, not a sniffed L7 label. Routed through the default
// branch below it landed in RawDefaultRule.Protocol, where it is compared
// against what the sniffers reported — and the sniffers never report
// "icmp" (route.go's pre-match skips the sniff action for an ICMP flow
// outright), so the rule was valid, warned about, and dead. The L3 ingress
// therefore had no way to express its target at all: every ping fell to
// whatever the catch-all resolved to, which on a real config is a group of
// proxy nodes that cannot carry layer 3.
//
// NetworkItem.Match is a plain map lookup over metadata.Network, which
// adapter.JudgeFlow sets to N.NetworkICMP for BOTH ICMPv4 and ICMPv6
// (header.ICMPv4ProtocolNumber and header.ICMPv6ProtocolNumber share the
// one case there). So there is exactly ONE network value here and `icmp`
// covers both families — a separate `icmpv6` network would match nothing,
// forever.
//
// The family is still expressible, and precisely: metadata.IPVersion is
// derived from the destination address (route.go prepareMatchMetadata,
// which PreMatch runs before the rules), and an ICMPv6 packet always
// carries an IPv6 destination. `icmpv4`/`icmpv6` therefore narrow the SAME
// network with an ip_version item rather than inventing a second network —
// no false positives and no false negatives, unlike an inert Protocol
// matcher (which would also silently widen to both families if we mapped
// it onto plain `icmp`).
raw.Network = badoption.Listable[string]{N.NetworkICMP}
if v := icmpProtoIPVersion(p); v != 0 {
raw.IPVersion = v
}
default:
// Sniffed L7 protocol. route/rule.NewProtocolItem does NO validation —
// it just compares the string against what the sniffers reported — so an
@@ -386,7 +426,7 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
// the rule to ALL traffic, which for a `direct` target is a leak), but
// the operator is told it is inert.
if !sniffedProtocols[p] {
b.warnf("rule %q: proto %q is not something this engine can detect — the rule is kept but can NEVER match, so its traffic silently follows the rules below it. Use tcp, udp or one of: %s", r.Name, r.Proto, sniffedProtocolList())
b.warnf("rule %q: proto %q is not something this engine can detect — the rule is kept but can NEVER match, so its traffic silently follows the rules below it. Use tcp, udp, icmp (icmpv4/icmpv6 narrow it to one family) or one of: %s", r.Name, r.Proto, sniffedProtocolList())
}
raw.Protocol = badoption.Listable[string]{p}
}
@@ -411,10 +451,269 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
return raw, matched
}
// The ICMP spellings `Rule.Proto` accepts. All three become the SAME engine
// network (N.NetworkICMP); the two family-qualified ones additionally pin
// RawDefaultRule.IPVersion. See the switch in ruleMatchers for why there is only
// one network and why the family is an ip_version item rather than a second one.
const (
protoICMP = "icmp"
protoICMPv4 = "icmpv4"
protoICMPv6 = "icmpv6"
)
// isICMPProto reports whether a Rule.Proto value asks for the L3 ingress.
func isICMPProto(proto string) bool {
switch strings.TrimSpace(strings.ToLower(proto)) {
case protoICMP, protoICMPv4, protoICMPv6:
return true
}
return false
}
// icmpProtoIPVersion is the ip_version an ICMP proto narrows to; 0 = both
// families (plain `icmp`) or not an ICMP proto at all.
func icmpProtoIPVersion(proto string) int {
switch strings.TrimSpace(strings.ToLower(proto)) {
case protoICMPv4:
return 4
case protoICMPv6:
return 6
}
return 0
}
// warnICMPRule reports the ways an emitted `proto icmp` rule can still be a
// no-op, none of which is visible anywhere else.
//
// A rule that matches nothing is normally cheap: its traffic falls through to
// the rules below it. Not here. ICMP has no fall-through — the packet either
// reaches an L3-capable outbound or is DROPPED (route.preMatchFlow's
// l3-honest-drop block, which exists so the TUN stack cannot forge an echo reply
// for a path that never carried the packet). So an ICMP rule that cannot fire is
// not a dead setting, it is ping that stops working, with a rule in the UI that
// says it should.
//
// Deliberately worded clear of shater/apply's criticalMarkers: a dropped ping is
// not a protection gap (nothing leaks — the failure is fail-CLOSED), so these are
// warnings, and phrases like "never applies" / "NOT emitted" would light the
// panel's alarm banner for something that costs the operator ping and nothing else.
func (b *builder) warnICMPRule(r model.Rule, want string) {
if !b.m.Globals.L3Tunnel {
// The prerequisite, and the only one whose absence makes the other checks
// moot: with no L3 ingress the packet never enters the engine, so the target
// is not consulted at all.
b.warnf("rule %q: proto %s matches only traffic that reaches the engine at layer 3, and nothing does while globals l3_tunnel is off — kernel TPROXY diverts TCP and UDP and nothing else, so no ping ever enters the engine and this rule cannot fire. ICMP is governed by globals untunnelable (%s) instead, whatever this rule's target %q says. Turn l3_tunnel on to make the rule live", r.Name, strings.ToLower(strings.TrimSpace(r.Proto)), untunnelablePolicyName(b.m.Globals.Untunnelable), want)
return
}
if icmpProtoIPVersion(r.Proto) == 6 && !b.m.Globals.IPv6 {
// netplane/nft.go marks ipv6-icmp into the L3 TUN only under Globals.IPv6,
// and generate/inbound.go gives that TUN an IPv6 address on the same
// condition. With IPv6 off the v6 half of the ingress simply does not exist.
b.warnf("rule %q: proto icmpv6 needs globals ipv6 on — the L3 ingress marks ipv6-icmp into the engine's TUN only then, and the TUN is given no IPv6 address either, so with IPv6 off this rule cannot fire. Use proto icmp (one value, both families) or turn ipv6 on", r.Name)
return
}
// A port matcher and ICMP are mutually exclusive by construction:
// adapter.JudgeFlow zeroes source and destination ports for an ICMP flow
// before PreMatch runs, so a PortItem next to the network item can never be
// satisfied. The rule is emitted (dropping the port would WIDEN it) but it is
// dead.
if single, ranges, _ := splitPorts(r.DstPort); len(single)+len(ranges) > 0 {
b.warnf("rule %q: proto %s together with a port matcher can never match — an ICMP packet has no port, and the engine zeroes both ports on an ICMP flow before matching it (adapter.JudgeFlow). Drop the port from this rule, or move the port part into a rule of its own", r.Name, strings.ToLower(strings.TrimSpace(r.Proto)))
}
switch b.l3Target(want, 0) {
case l3Drops:
b.warnf("rule %q: proto %s is routed to %q, which cannot carry a layer-3 packet — only wireguard/AmneziaWG nodes and direct/interface egresses reach the engine's flow path; every proxy protocol (vless, vmess, trojan, shadowsocks, hysteria2, tuic, shadowtls) and the byedpi SOCKS egress cannot, because none of them can implement it. The engine DROPS a ping routed at such a target rather than let the TUN stack answer the echo itself from a path that never carried the packet, so this rule makes those pings fail instead of tunnelling them. Point it at a wireguard/AmneziaWG node, at an interface egress, or at a chain whose LAST hop is one of those", r.Name, strings.ToLower(strings.TrimSpace(r.Proto)), want)
case l3Partial:
b.warnf("rule %q: proto %s is routed to %q, whose members disagree about layer 3 — the ping is carried while the group's current member is a wireguard/AmneziaWG node and dropped while it is a proxy-protocol one, and nothing in the UI says which is in force right now. Point the rule at the L3-capable node itself (or at a group holding only those) if ping must behave the same from one minute to the next", r.Name, strings.ToLower(strings.TrimSpace(r.Proto)), want)
}
}
// untunnelablePolicyName spells the Untunnelable policy for a diagnostic, naming
// the empty value as the "block" it means (model.Globals.Untunnelable).
func untunnelablePolicyName(policy string) string {
if p := strings.ToLower(strings.TrimSpace(policy)); p != "" {
return p
}
return "block"
}
// l3Verdict is what a route target does with a layer-3 packet.
type l3Verdict int
const (
// l3Unknown: undecidable from the model alone — say nothing rather than guess.
l3Unknown l3Verdict = iota
// l3Carries: the emitted outbound implements adapter.FlowOutbound.
l3Carries
// l3Drops: it does not, so route.preMatchFlow drops the packet.
l3Drops
// l3Blocks: `block`. It drops the packet too, but that IS the stated policy,
// so it is never reported.
l3Blocks
// l3Partial: a group whose members disagree; the answer changes with the pick.
l3Partial
)
// l3Target answers, from the MODEL, whether a rule target ends up at an outbound
// that can carry a layer-3 packet.
//
// # Why this is decidable here at all
//
// The capability is not a runtime property to be discovered: it is fixed by the
// outbound TYPE this package is about to emit, and adapter registration makes the
// list exhaustive. Only `direct` (protocol/direct/outbound.go:66), `wireguard`
// (protocol/wireguard/endpoint.go:73), `tailscale` and `bridge` declare
// N.NetworkICMP among their networks, and of those exactly two are reachable from
// a shater model: a wireguard/AmneziaWG node, and the direct outbound behind
// `direct` or an interface/direct egress. Everything else this generator emits —
// vless, vmess, trojan, shadowsocks, hysteria2, tuic, shadowtls and the byedpi
// SOCKS egress — cannot.
//
// # The predicate that MUST change in lockstep
//
// "Is this node an endpoint?" is decided in outbound.go by
// `p.WG != nil || p.Protocol == "wireguard"`. l3Node repeats that test. If one
// side learns a new L3-capable node kind and the other does not, this diagnosis
// starts lying in whichever direction the drift went — a false alarm on a working
// ping, or silence on a broken one.
//
// depth bounds the chain->chain recursion; expandHops already refuses cycles, so
// it is a belt on top of a brace.
func (b *builder) l3Target(want string, depth int) l3Verdict {
if depth > 8 {
return l3Unknown
}
// The kind switch mirrors resolveTarget exactly — the two must agree about what
// a target string means, or this diagnoses a different outbound than the one the
// rule is routed to.
kind, name := model.SplitTarget(want)
switch strings.ToLower(kind) {
case tagDirect, "":
if strings.EqualFold(want, tagBlock) {
return l3Blocks
}
return l3Carries
case tagBlock:
return l3Blocks
case "node":
return b.l3Node(name)
case "group":
return b.l3Group(name)
case "egress":
return b.l3Egress(name)
case "chain":
return b.l3Chain(name, depth)
default:
// Bare name: node then group, same order resolveTarget uses.
if b.nodeTags[want] {
return b.l3Node(want)
}
if b.groupTags[want] {
return b.l3Group(want)
}
return l3Unknown
}
}
// l3Node: a node carries layer 3 exactly when it is emitted as a WireGuard/
// AmneziaWG ENDPOINT rather than a proxy outbound.
func (b *builder) l3Node(name string) l3Verdict {
for i := range b.m.Nodes {
n := b.m.Nodes[i]
if n.Name != name || !n.Enabled {
continue
}
p, err := parse.ParseShareLink(n.URI)
if err != nil {
// Unreachable from warnICMPRule (an unparseable node has no tag, so the
// target would not have resolved), and not ours to report twice anyway.
return l3Unknown
}
if p.WG != nil || p.Protocol == "wireguard" {
return l3Carries
}
return l3Drops
}
return l3Unknown
}
// l3Group folds its members' verdicts. route.preMatchFlow unwraps a group through
// group.Now() before testing the outbound, so the group's answer IS its current
// member's — which is why a mixed group is reported as its own case rather than
// rounded to either side.
func (b *builder) l3Group(name string) l3Verdict {
g, found := b.findGroup(name)
if !found {
return l3Unknown
}
var members []string
// groupMembers re-emits the member diagnostics buildGroups already surfaced.
b.withoutNewWarnings(func() { members = b.groupMembers(g) })
carries, drops := 0, 0
for _, member := range members {
switch b.l3Node(member) {
case l3Carries:
carries++
case l3Drops:
drops++
}
}
switch {
case carries == 0 && drops == 0:
return l3Unknown
case carries == 0:
return l3Drops
case drops == 0:
return l3Carries
default:
return l3Partial
}
}
// l3Egress: an interface/direct egress is a DIRECT outbound (the one proxy
// outbound in the shipped registry that builds a ping.Port), byedpi is SOCKS, and
// any other type emits no outbound at all — so it never reaches this diagnosis.
func (b *builder) l3Egress(name string) l3Verdict {
for _, eg := range b.m.Egresses {
if eg.Name != name {
continue
}
switch strings.ToLower(strings.TrimSpace(eg.Type)) {
case "interface", "direct", "":
return l3Carries
default:
return l3Drops
}
}
return l3Unknown
}
// l3Chain: a rule routed at a chain enters at the LAST hop's wrapper (chain.go —
// Detour points backwards, so Ln is the outbound the rule is handed to and the
// exit on the wire). That wrapper is a copy of Ln's own outbound, so Ln's
// capability is the chain's.
func (b *builder) l3Chain(name string, depth int) l3Verdict {
var (
hops []string
baseDetour string
expanded bool
)
// expandHops flattens sub-chains and reports cycles/undefined references;
// resolveChain already surfaced all of that for this same chain.
b.withoutNewWarnings(func() {
expanded = b.expandHops(name, map[string]bool{}, &hops, &baseDetour, true)
})
if !expanded || len(hops) == 0 {
return l3Unknown
}
return b.l3Target(hops[len(hops)-1], depth+1)
}
// sniffedProtocols is exactly the set of L7 labels this engine's sniffers can
// ever put on a connection (route/rule.RuleActionSniff.build + the default
// stream/packet sniffer sets in route/route.go). A `proto` outside this set —
// and outside tcp/udp — matches nothing, forever.
// and outside the tcp/udp/icmp NETWORKS handled ahead of it in ruleMatchers —
// matches nothing, forever.
var sniffedProtocols = map[string]bool{
C.ProtocolTLS: true,
C.ProtocolHTTP: true,
+377
View File
@@ -0,0 +1,377 @@
// The rule side of the L3 ingress: whether `proto icmp` can be SAID at all, and
// whether saying it is honest.
//
// Before this, `icmp` fell through ruleMatchers' proto switch into
// RawDefaultRule.Protocol — the sniffed-L7 field — where it was compared against
// labels the sniffers report. They never report "icmp" (route.go's pre-match
// skips the sniff action for an ICMP flow outright), so the rule was structurally
// valid and permanently dead. The consequence was not a dead setting but a dead
// FEATURE: with no way to write "ICMP goes here", every ping fell to whatever the
// catch-all resolved to, and on the configuration this was built for that is a
// group of VLESS nodes, which cannot carry layer 3 at all.
//
// Portable (no box.New): these assert on the generated option.Options and on the
// warning texts only. The live-engine proofs that a direct/wireguard outbound
// really carries the flow live in l3_egress_linux_test.go and
// l3_integration_linux_test.go.
package generate
import (
"strings"
"testing"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// --- fixture ----------------------------------------------------------------
// icmpModel builds a model carrying one of everything the L3 verdict has to tell
// apart: a WireGuard node (an ENDPOINT — carries layer 3), a Shadowsocks node (a
// proxy outbound — cannot), an interface egress (a direct outbound — carries), a
// byedpi egress (SOCKS — cannot), three groups spanning all-proxy/all-L3/mixed,
// and two chains that differ only in which kind of node they EXIT through.
func icmpModel(l3Tunnel bool, rules ...model.Rule) *model.Model {
g := model.DefaultGlobals()
g.DNSIntercept = false // not the subject; keeps the DNS diagnostics out
g.KillSwitch = "closed"
g.L3Tunnel = l3Tunnel
return &model.Model{
Globals: g,
Inbounds: []model.Inbound{
{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12345, TCP: true, UDP: true},
},
Nodes: []model.Node{
{Name: "wg1", Enabled: true, URI: wgDedupURI(wgDedupKey(1), wgDedupKey(2), "203.0.113.10", 51820)},
{Name: "wg2", Enabled: true, URI: wgDedupURI(wgDedupKey(3), wgDedupKey(4), "203.0.113.11", 51820)},
{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#ss1"},
},
Egresses: []model.Egress{
{Name: "wan2", Type: "interface", Interface: "wan2"},
{Name: "bd", Type: "byedpi", Port: 1080},
},
Groups: []model.Group{
{Name: "proxies", Source: "manual", Nodes: []string{"ss1"}},
{Name: "l3only", Source: "manual", Nodes: []string{"wg1", "wg2"}},
{Name: "mixed", Source: "manual", Nodes: []string{"ss1", "wg1"}},
},
Chains: []model.Chain{
{Name: "exit-proxy", Hops: []string{"node:wg1", "node:ss1"}},
{Name: "exit-wg", Hops: []string{"node:ss1", "node:wg1"}},
},
Rules: rules,
}
}
// genICMP generates the fixture and returns the model-derived route rules plus
// the warnings.
func genICMP(t *testing.T, l3Tunnel bool, rules ...model.Rule) ([]option.Rule, []string) {
t.Helper()
opts, warns, err := GenerateWithWarnings(icmpModel(l3Tunnel, rules...))
if err != nil {
t.Fatalf("Generate: %v", err)
}
return generalRules(opts.Route), warns
}
// oneICMPRule generates a single-rule model and returns that rule's matchers.
func oneICMPRule(t *testing.T, l3Tunnel bool, r model.Rule) (option.RawDefaultRule, []string) {
t.Helper()
rules, warns := genICMP(t, l3Tunnel, r)
if len(rules) != 1 {
t.Fatalf("want exactly 1 emitted route rule, got %d (warns=%v) — a rule the generator drops carries no target at all, and for ICMP that is not a fall-through but a drop", len(rules), warns)
}
return rules[0].DefaultOptions.RawDefaultRule, warns
}
// icmpRule is a `proto <p>` rule scoped to one LAN source, pointed at target.
func icmpRule(name, proto, target string) model.Rule {
return model.Rule{
Name: name, Enabled: true, Order: 10,
Src: []string{"192.168.1.0/24"},
Proto: proto,
Target: target,
}
}
// --- what the matcher must become -------------------------------------------
// TestProtoICMPBecomesTheNetworkNotTheSniffedProtocol is the fix itself.
//
// route/rule.NewNetworkItem matches metadata.Network with a plain map lookup, and
// adapter.JudgeFlow sets that to N.NetworkICMP ("icmp") for an ICMP flow — so the
// value belongs in Network. RawDefaultRule.Protocol is the sniffed-L7 field
// (NewProtocolItem compares against what the sniffers labelled the connection),
// and nothing ever labels a flow "icmp": route.go's PreMatch skips the sniff
// action for ICMP before it starts. A value there is a rule that cannot fire.
func TestProtoICMPBecomesTheNetworkNotTheSniffedProtocol(t *testing.T) {
raw, warns := oneICMPRule(t, true, icmpRule("ping", "icmp", "node:wg1"))
if len(raw.Protocol) != 0 {
t.Fatalf("proto icmp landed in RawDefaultRule.Protocol = %v — that field is matched against the SNIFFED protocol label, and no sniffer ever produces \"icmp\" (PreMatch skips sniffing for an ICMP flow), so the rule is valid, silent and permanently dead: every ping keeps falling to the catch-all instead of the target the operator wrote", raw.Protocol)
}
if len(raw.Network) != 1 || raw.Network[0] != "icmp" {
t.Fatalf("RawDefaultRule.Network = %v, want [icmp] — NetworkItem.Match is a map lookup over metadata.Network, which adapter.JudgeFlow sets to exactly this string for an ICMP flow; anything else and the L3 ingress has no expressible target at all", raw.Network)
}
if raw.IPVersion != 0 {
t.Fatalf("RawDefaultRule.IPVersion = %d, want 0 — plain `icmp` must cover BOTH families (JudgeFlow maps ICMPv4 and ICMPv6 to the one network), so narrowing it here would silently drop half the pings the operator asked for", raw.IPVersion)
}
if routeWarnsHave(warns, "is not something this engine can detect") {
t.Fatalf("icmp was reported as an undetectable sniffed protocol: %v — it is a NETWORK, and the warning tells the operator to stop using the only spelling that works", warns)
}
}
// TestProtoICMPFamilySpellingsNarrowByIPVersion pins the answer to "is ICMPv6 a
// second network?": it is not. adapter.JudgeFlow folds
// header.ICMPv4ProtocolNumber and header.ICMPv6ProtocolNumber into ONE case and
// sets N.NetworkICMP for both, so an `icmpv6` network value would match nothing,
// forever. The family is still expressible, and exactly: metadata.IPVersion comes
// from the destination address (prepareMatchMetadata, which PreMatch runs), and
// an ICMPv6 packet always has an IPv6 destination. So the family spellings narrow
// the same network with ip_version instead of inventing a second one.
func TestProtoICMPFamilySpellingsNarrowByIPVersion(t *testing.T) {
for _, tc := range []struct {
proto string
want int
}{{"icmpv4", 4}, {"icmpv6", 6}} {
raw, warns := oneICMPRule(t, true, icmpRule("ping", tc.proto, "node:wg1"))
if len(raw.Network) != 1 || raw.Network[0] != "icmp" {
t.Fatalf("proto %s: Network = %v, want [icmp] — there is no separate icmpv6 network in this engine; emitting one produces a rule that can never match", tc.proto, raw.Network)
}
if raw.IPVersion != tc.want {
t.Fatalf("proto %s: IPVersion = %d, want %d — without it the rule silently covers the OTHER family too, which is a different rule than the one written", tc.proto, raw.IPVersion, tc.want)
}
if len(raw.Protocol) != 0 {
t.Fatalf("proto %s: landed in Protocol = %v (warns=%v) — the sniffed-L7 field, where it can never match", tc.proto, raw.Protocol, warns)
}
}
}
// TestProtoTCPUDPAndSniffedAreUnchanged: the new branch must not move any value
// that already worked. tcp/udp stay networks, a sniffed L7 label stays in
// Protocol, and neither picks up an ip_version it never had.
func TestProtoTCPUDPAndSniffedAreUnchanged(t *testing.T) {
for _, network := range []string{"tcp", "udp"} {
raw, _ := oneICMPRule(t, true, icmpRule("t", network, "node:ss1"))
if len(raw.Network) != 1 || raw.Network[0] != network {
t.Fatalf("proto %s: Network = %v, want [%s]", network, raw.Network, network)
}
if len(raw.Protocol) != 0 || raw.IPVersion != 0 {
t.Fatalf("proto %s: Protocol = %v / IPVersion = %d, want empty/0 — the ICMP branch leaked into the transport branch", network, raw.Protocol, raw.IPVersion)
}
}
raw, warns := oneICMPRule(t, true, icmpRule("t", "tls", "node:ss1"))
if len(raw.Protocol) != 1 || raw.Protocol[0] != "tls" {
t.Fatalf("proto tls: Protocol = %v, want [tls] — a sniffed L7 label belongs in the sniffed field", raw.Protocol)
}
if len(raw.Network) != 0 || raw.IPVersion != 0 {
t.Fatalf("proto tls: Network = %v / IPVersion = %d, want empty/0", raw.Network, raw.IPVersion)
}
if routeWarnsHave(warns, "is not something this engine can detect") {
t.Fatalf("tls is sniffable and was reported as not: %v", warns)
}
// The vocabulary the undetectable-proto warning offers must now name icmp —
// it is the list an operator reads when their spelling was rejected, and
// leaving icmp out of it points them away from the only working value.
_, warns = oneICMPRule(t, true, icmpRule("t", "ping", "node:ss1"))
if !routeWarnsHave(warns, "is not something this engine can detect") {
t.Fatalf("proto \"ping\" is neither a network nor a sniffed label, yet nothing was reported: %v", warns)
}
if !routeWarnsHave(warns, "Use tcp, udp, icmp") {
t.Fatalf("the undetectable-proto warning still offers only tcp/udp + sniffed labels: %v — an operator who wrote \"ping\" is told every value EXCEPT the one that would work", warns)
}
}
// --- the two prerequisites the rule cannot state itself ----------------------
// TestProtoICMPWithoutL3TunnelIsReported: with globals l3_tunnel off, kernel
// TPROXY diverts TCP and UDP and nothing else, so no ICMP packet ever enters the
// engine and the rule — perfectly well-formed — matches nothing at all. Silence
// here is the worst kind: the panel shows a rule that says ping is tunnelled.
func TestProtoICMPWithoutL3TunnelIsReported(t *testing.T) {
_, warns := oneICMPRule(t, false, icmpRule("ping", "icmp", "node:wg1"))
if !routeWarnsHave(warns, "globals l3_tunnel is off") {
t.Fatalf("a proto icmp rule under l3_tunnel=0 was accepted in silence: %v — the L3 ingress is the ONLY path an ICMP packet has into the engine, so without it the rule is decoration", warns)
}
// And the opposite: turning the ingress on must silence it, or the warning is
// noise that trains the operator to ignore the section.
_, warns = oneICMPRule(t, true, icmpRule("ping", "icmp", "node:wg1"))
if routeWarnsHave(warns, "globals l3_tunnel is off") {
t.Fatalf("l3_tunnel is ON and the rule was still reported as dead: %v", warns)
}
}
// TestProtoICMPv6WithoutIPv6IsReported: the v6 half of the ingress is gated twice
// on globals ipv6 — netplane/nft.go marks ipv6-icmp into the TUN only then, and
// generate/inbound.go gives the TUN an IPv6 address on the same condition. An
// icmpv6 rule with IPv6 off is therefore inert for a reason nothing else states.
func TestProtoICMPv6WithoutIPv6IsReported(t *testing.T) {
m := icmpModel(true, icmpRule("ping6", "icmpv6", "node:wg1"))
m.Globals.IPv6 = false
_, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if !routeWarnsHave(warns, "proto icmpv6 needs globals ipv6 on") {
t.Fatalf("an icmpv6 rule under ipv6=0 was accepted in silence: %v — neither the nft marking nor the TUN address exists in that configuration, so the rule cannot fire", warns)
}
_, warns = oneICMPRule(t, true, icmpRule("ping6", "icmpv6", "node:wg1"))
if routeWarnsHave(warns, "proto icmpv6 needs globals ipv6 on") {
t.Fatalf("ipv6 is on and the icmpv6 rule was still reported: %v", warns)
}
}
// TestProtoICMPWithPortMatcherIsReported: adapter.JudgeFlow zeroes both ports on
// an ICMP flow before PreMatch runs, so a port item sitting next to the network
// item can never be satisfied. The rule is still EMITTED — dropping the port
// matcher would widen it — but it is dead, and only this says so.
func TestProtoICMPWithPortMatcherIsReported(t *testing.T) {
r := icmpRule("ping", "icmp", "node:wg1")
r.DstPort = "443"
_, warns := oneICMPRule(t, true, r)
if !routeWarnsHave(warns, "an ICMP packet has no port") {
t.Fatalf("proto icmp + dst_port was accepted in silence: %v — the two matchers are mutually exclusive by construction, so the rule can never fire", warns)
}
_, warns = oneICMPRule(t, true, icmpRule("ping", "icmp", "node:wg1"))
if routeWarnsHave(warns, "an ICMP packet has no port") {
t.Fatalf("a portless icmp rule was reported as having a port: %v", warns)
}
}
// --- does the TARGET carry layer 3? -----------------------------------------
// TestProtoICMPToAnL4TargetIsReported is the honesty half. An ICMP flow has no
// fall-through: route.preMatchFlow's l3-honest-drop block DROPS a ping routed at
// an outbound that cannot carry layer 3, precisely so the TUN stack cannot forge
// an echo reply for a path that never saw the packet. So "ICMP -> group of VLESS
// nodes" is not a dead setting, it is ping that stops working — and the operator
// wrote the rule believing the opposite.
//
// Every target below is decidable from the model, because the capability is fixed
// by the outbound TYPE this package emits: only direct (protocol/direct) and the
// wireguard/AmneziaWG endpoints declare N.NetworkICMP.
func TestProtoICMPToAnL4TargetIsReported(t *testing.T) {
for _, target := range []string{
"node:ss1", // a proxy outbound
"group:proxies", // a group of nothing but proxy outbounds
"egress:bd", // byedpi: a SOCKS hop, which cannot implement tun.Port
"chain:exit-proxy", // the rule enters at the LAST hop, and that one is ss1
} {
_, warns := oneICMPRule(t, true, icmpRule("ping", "icmp", target))
if !routeWarnsHave(warns, "cannot carry a layer-3 packet") {
t.Fatalf("proto icmp -> %s was accepted in silence: %v — the engine DROPS that ping (route.preMatchFlow, l3-honest-drop) and nothing else in the UI says so; the operator reads a rule claiming ping is tunnelled and a ping that fails", target, warns)
}
}
}
// TestProtoICMPToAnL3TargetIsSilent: the same check must stay quiet for every
// target that really does carry the packet, including `block` — a blocked ping is
// a stated policy, not an accident — and a chain that merely PASSES THROUGH a
// proxy hop on its way to a WireGuard exit.
func TestProtoICMPToAnL3TargetIsSilent(t *testing.T) {
for _, target := range []string{
"node:wg1", // a wireguard endpoint
"group:l3only", // a group of nothing else
"egress:wan2", // an interface egress = a direct outbound
"direct", // the baseline direct outbound
"block", // dropping the ping IS the policy here
"chain:exit-wg", // enters at the LAST hop, which is wg1
} {
_, warns := oneICMPRule(t, true, icmpRule("ping", "icmp", target))
if routeWarnsHave(warns, "cannot carry a layer-3 packet") {
t.Fatalf("proto icmp -> %s was reported as unable to carry layer 3: %v — a false alarm on a working path teaches the operator to ignore the one that is real", target, warns)
}
if routeWarnsHave(warns, "disagree about layer 3") {
t.Fatalf("proto icmp -> %s was reported as a mixed group: %v", target, warns)
}
}
}
// TestProtoICMPToAMixedGroupIsReported: route.preMatchFlow unwraps a group
// through group.Now() before testing the outbound, so a group holding both kinds
// answers differently from one pick to the next — ping works, then does not, with
// nothing in the UI to say which member is in force. That is its own report, not
// a rounding of either side.
func TestProtoICMPToAMixedGroupIsReported(t *testing.T) {
_, warns := oneICMPRule(t, true, icmpRule("ping", "icmp", "group:mixed"))
if !routeWarnsHave(warns, "disagree about layer 3") {
t.Fatalf("proto icmp -> a group of one wireguard and one proxy node was accepted in silence: %v — the ping's fate follows the group's current pick", warns)
}
if routeWarnsHave(warns, "cannot carry a layer-3 packet") {
t.Fatalf("a MIXED group was reported as unable to carry layer 3 at all: %v — it can, half the time, and telling the operator otherwise sends them to change a target that is only unstable", warns)
}
}
// TestProtoICMPDiagnosticsAreQuietForNonICMPRules: none of the above may fire on
// the tcp/udp/sniffed rules that make up every existing config. A byedpi egress
// carrying TCP is exactly right, and saying otherwise would flood the panel.
func TestProtoICMPDiagnosticsAreQuietForNonICMPRules(t *testing.T) {
_, warns := genICMP(t, true,
icmpRule("a", "tcp", "egress:bd"),
icmpRule("b", "udp", "group:proxies"),
icmpRule("c", "tls", "node:ss1"),
)
for _, marker := range []string{
"cannot carry a layer-3 packet",
"disagree about layer 3",
"globals l3_tunnel is off",
"an ICMP packet has no port",
} {
if routeWarnsHave(warns, marker) {
t.Fatalf("a non-ICMP rule drew the ICMP diagnostic %q: %v", marker, warns)
}
}
}
// TestProtoICMPUnresolvedTargetIsNotDoubleReported: a rule whose target does not
// resolve already earns ruleKillFallback's much louder warning and is routed to
// block. A capability verdict on the target it never reaches would compete with
// it — the operator's problem is the broken reference, and it is the one sentence
// that must be read first.
//
// The shape below is the one where the two actually collide: a group named after
// a node is SKIPPED by buildGroups (the outbound manager last-wins on a duplicate
// tag), so `group:ss1` does not resolve — while the group DEFINITION is still in
// the model, holding a proxy member, so a verdict is perfectly computable for it.
func TestProtoICMPUnresolvedTargetIsNotDoubleReported(t *testing.T) {
m := icmpModel(true, icmpRule("ping", "icmp", "group:ss1"))
m.Groups = append(m.Groups, model.Group{Name: "ss1", Source: "manual", Nodes: []string{"ss1"}})
_, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if !routeWarnsHave(warns, "unresolved target") {
t.Fatalf("a rule pointing at a non-existent node lost its unresolved-target warning: %v", warns)
}
if routeWarnsHave(warns, "cannot carry a layer-3 packet") {
t.Fatalf("the unresolved target was ALSO reported as unable to carry layer 3: %v — the operator's problem is the missing node, and the second sentence competes with the first", warns)
}
}
// TestProtoICMPWarningsStayBelowTheAlarmThreshold: shater/apply grades generate's
// free text, and a handful of substrings ("never applies", "NOT emitted", …) light
// the panel's alarm banner. A ping that fails is fail-CLOSED — nothing leaks — so
// these belong under `warning`, and the wording must not drift into the markers.
// Checked here rather than in shater/apply because that package imports this one.
func TestProtoICMPWarningsStayBelowTheAlarmThreshold(t *testing.T) {
var texts []string
_, w := oneICMPRule(t, false, icmpRule("a", "icmp", "node:wg1"))
texts = append(texts, w...)
_, w = oneICMPRule(t, true, icmpRule("b", "icmp", "node:ss1"))
texts = append(texts, w...)
_, w = oneICMPRule(t, true, icmpRule("c", "icmp", "group:mixed"))
texts = append(texts, w...)
// The exact substrings shater/apply/warnings.go greps for.
markers := []string{"NOT emitted", "inert", "has NO effect", "never applies", "in the clear"}
for _, text := range texts {
if !strings.Contains(text, "layer 3") && !strings.Contains(text, "layer-3") && !strings.Contains(text, "l3_tunnel") {
continue // somebody else's warning
}
for _, m := range markers {
if strings.Contains(text, m) {
t.Fatalf("ICMP warning %q contains the critical marker %q — it would light the panel's alarm banner for a fail-CLOSED ping failure, and every cosmetic critical teaches the operator to skip the real one", text, m)
}
}
}
}
+152
View File
@@ -0,0 +1,152 @@
package model
// fwmark_base / table_base: two free-form hex fields in the panel's "Advanced"
// section that nothing had ever validated, and whose DERIVED values are not
// visible from the number typed.
//
// The panel offers them with the note "only change these if another app on the
// router already uses the same range", which reads like a courtesy. It is not:
// every mark and every routing table the data plane owns is computed from these
// two numbers by adding a fixed offset, and some of the results land on things
// the kernel or the engine already owns. Nothing downstream refuses them,
// nothing warns, and the failures they produce (the router losing its uplink,
// the main routing table being flushed) point nowhere near a number in a
// collapsed section.
import (
"strings"
"testing"
)
func eg(n int) []Egress {
out := make([]Egress, n)
for i := range out {
out[i] = Egress{Name: "e", Type: "interface", Interface: "eth1"}
}
return out
}
func joined(ws []Warning) string {
var b strings.Builder
for _, w := range ws {
b.WriteString(w.Error())
b.WriteString("\n")
}
return b.String()
}
// TestValidateMarkBasesFwmark.
//
// 0x7f is the case the L3 ingress added and the reason this check exists: the L3
// mark is fwmark_base + 0x80, so 0x7f puts it exactly on 0xff — the loop-guard
// mark the engine stamps on its OWN outgoing traffic. `ip rule fwmark 0xff
// lookup <l3 table>` then captures everything the engine sends and routes it
// into the engine's own TUN. The router loses the internet the moment l3_tunnel
// is switched on, and nothing connects that to a hex field.
//
// 0xff is the same failure by the older route (the divert mark itself lands on
// the loop guard) and predates the L3 offset; adding an offset simply added a
// second base that does it, which is exactly why the check is written as "derive
// every value and look for collisions" rather than as a blacklist of two numbers.
func TestValidateMarkBasesFwmark(t *testing.T) {
cases := []struct {
name string
base uint32
egress int
wantBad bool
want string // substring the message must carry
}{
{"unset means the 0x2000 default", 0, 2, false, ""},
{"the shipped default", 0x2000, 8, false, ""},
{"a sane custom base", 0x30000, 4, false, ""},
// The L3 offset lands the derived mark on the engine's own loop guard.
{"0x7f: L3 mark becomes 0xff", 0x7f, 0, true, "fwmark_base + 0x80"},
{"0x7f with egresses too", 0x7f, 3, true, "0xff"},
// The divert mark itself is the loop guard.
{"0xff: the divert mark IS the loop guard", 0xff, 0, true, "fwmark_base itself"},
// Past the top of a 32-bit mark the arithmetic wraps onto another mark.
{"wraps past 32 bits", 0xffffff80, 4, true, "wraps around"},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
ws := ValidateMarkBases(Globals{FwmarkBase: c.base}, eg(c.egress))
if got := len(ws) > 0; got != c.wantBad {
t.Fatalf("fwmark_base 0x%x with %d egresses: warned = %v, want %v.\n%s",
c.base, c.egress, got, c.wantBad, joined(ws))
}
if c.want != "" && !strings.Contains(joined(ws), c.want) {
t.Errorf("the warning must name the derived value that collides (%q):\n%s", c.want, joined(ws))
}
for _, w := range ws {
if w.Section != "globals" {
t.Errorf("warnings must be scoped to globals, got %q", w.Section)
}
}
})
}
}
// TestValidateMarkBasesTable: the same arithmetic decides routing table ids, and
// there the reserved values are the kernel's own. Table 254 is `main` — and this
// codebase does `ip route flush table <n>` on teardown, so a base whose derived
// table lands on it does not merely collide, it takes the router off the network.
//
// The egress count matters, which is why this check is not in ValidateGlobals:
// egress #i uses table_base + 0x10 + i, so whether a base is safe depends on how
// many egresses are configured. table_base 239 is fine on a router with no
// egresses and flushes the kernel's `local` table on one with a single egress.
func TestValidateMarkBasesTable(t *testing.T) {
cases := []struct {
name string
base uint32
egress int
wantBad bool
want string
}{
{"unset", 0, 2, false, ""},
{"the shipped default", 0x2000, 8, false, ""},
{"the divert table IS main", 254, 0, true, "MAIN routing table"},
{"the L3 table becomes main", 246, 0, true, "table_base + 0x8"},
{"the L3 table becomes `default`", 245, 0, true, "253"},
// Count-sensitive: harmless with no egresses, fatal with one.
{"no egress, no derived egress table", 239, 0, false, ""},
{"one egress lands on `local`", 239, 1, true, "local"},
{"two egresses reach main", 238, 2, true, "MAIN routing table"},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
ws := ValidateMarkBases(Globals{TableBase: c.base}, eg(c.egress))
if got := len(ws) > 0; got != c.wantBad {
t.Fatalf("table_base %d with %d egresses: warned = %v, want %v.\n%s",
c.base, c.egress, got, c.wantBad, joined(ws))
}
if c.want != "" && !strings.Contains(joined(ws), c.want) {
t.Errorf("the warning must say WHICH kernel table is being taken over (%q):\n%s",
c.want, joined(ws))
}
})
}
}
// The check has to be reachable from the one entry point the daemon and the
// panel actually call, or it is a function nobody runs.
func TestValidateSurfacesMarkBaseCollisions(t *testing.T) {
m := &Model{
Globals: Globals{FwmarkBase: 0x7f, TableBase: 0x2000},
Egresses: []Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}},
}
ws := m.Validate(nil)
if !strings.Contains(joined(ws), "0xff") {
t.Fatalf("Model.Validate does not run ValidateMarkBases, so the panel and the log never see it:\n%s",
joined(ws))
}
}
// A base whose derived values are all distinct and all clear of the reserved
// ones must stay silent even with a lot of egresses — the check must not become
// noise that gets ignored.
func TestValidateMarkBasesQuietOnTheDefault(t *testing.T) {
if ws := ValidateMarkBases(Globals{FwmarkBase: 0x2000, TableBase: 0x2000}, eg(64)); len(ws) != 0 {
t.Fatalf("the shipped defaults with 64 egresses must warn about nothing:\n%s", joined(ws))
}
}
+42
View File
@@ -260,6 +260,48 @@ type Globals struct {
// current behaviour exactly.
Untunnelable string
// L3Tunnel opts the router into the L3 ingress: the engine opens a TUN device
// (netplane.L3Device) and the nft prerouting chain policy-routes LAN traffic
// the tunnel can only carry at layer 3 into it. There the engine's ordinary
// route rules decide the outbound, and an L3-capable one (wireguard/AWG —
// anything implementing adapter.FlowOutbound) carries the packet for real,
// NAT'd onto the tunnel's own address.
//
// TODAY THAT MEANS ICMP ECHO AND NOTHING ELSE, and the reason is worth stating
// precisely because the obvious one is wrong. It is NOT the gVisor stack: a
// WireGuard/AWG endpoint forwards at layer 3 straight past it — WritePackets
// reads only the IP version and the destination address before handing the raw
// bytes to the device (transport/wireguard/port.go:21-58), and the return path
// offers every decrypted packet back before the stack ever sees it (:127-157).
// The limit is sing-tun's dispatcher: flow_parse.go sets hasFlow for TCP, UDP
// and ICMP echo alone, so nothing else is dispatched, and flow_dispatch.go's
// createFlow NATs through a port-shaped selector that ESP, AH and GRE do not
// have. Those stay governed by Untunnelable and UntunnelableEgress below. Ping
// and Windows `tracert` through the tunnel are the whole of what this buys.
//
// Default false: without it nothing changes, and Untunnelable alone decides
// what happens to non-TCP/UDP traffic.
L3Tunnel bool
// UntunnelableEgress names an egress of type interface/tunnel that carries the
// protocols the engine will not dispatch: ESP, AH, GRE, IGMP, SCTP. The
// blocker is sing-tun's flow dispatcher (see L3Tunnel above), not the tunnel
// itself, but the effect is the same — L3Tunnel cannot help them.
//
// It is not a tunnel of ours. The kernel simply routes those packets out that
// egress's device, with the kernel's own NAT, and every protocol works because
// nothing in the path has to understand any of them. What that BUYS depends
// entirely on what the device is: a WireGuard interface really is a tunnel, a
// second WAN is just a different uplink and the destination sees that uplink's
// real address. The panel says which, at apply time, from whether the device is
// point-to-point.
//
// Empty (the default) leaves Untunnelable above in sole charge, exactly as
// before. A name that does not resolve to an interface/tunnel egress is
// reported and ignored — the policy applies unchanged, which is the
// fail-closed reading.
UntunnelableEgress string
// StatsBackend selects the statistics storage backend (pluggable stats store).
// Values: "off" | "memory" | "sqlite" (default "memory"). Contract: stats.NewStore.
// off => no-op store: nothing is subscribed/aggregated (zero per-query overhead),
+2
View File
@@ -89,6 +89,8 @@ func RenderUCIExport(m *Model) string {
w.boolOpt("block_doh", g.BlockDoH)
w.boolOpt("group_health", g.GroupHealth)
w.strOpt("untunnelable", g.Untunnelable)
w.boolOpt("l3_tunnel", g.L3Tunnel)
w.strOpt("untunnelable_egress", g.UntunnelableEgress)
w.strOpt("stats_backend", g.StatsBackend)
// The three stats-size knobs use intOptAlways (NOT the omit-zero intOpt) because 0
// means UNLIMITED, a meaningful value that MUST survive the round-trip. If they were
+70 -19
View File
@@ -17,25 +17,27 @@ import (
func richModel() *Model {
return &Model{
Globals: Globals{
Enabled: false, // exercise false via always-emit bool
LogLevel: "debug",
KillSwitch: "open",
Untunnelable: "direct",
IPv6: false,
FwmarkBase: 0x2000,
TableBase: 0x3000,
ConfirmTimeout: 30,
ResolverDefault: "cf",
ResolverFallback: "fake",
ProbeURL: "http://gstatic.com/generate_204",
ProbeInterval: "60s",
SchemaVersion: 1,
ActiveProfile: "home",
PanelPort: 8090,
DNSFilter: true,
DNSIntercept: true,
BlockDoH: true,
StatsBackend: "off",
Enabled: false, // exercise false via always-emit bool
LogLevel: "debug",
KillSwitch: "open",
Untunnelable: "direct",
L3Tunnel: true, // opt-in; exercise the non-default via round-trip
UntunnelableEgress: "frag", // exercise the non-default via round-trip
IPv6: false,
FwmarkBase: 0x2000,
TableBase: 0x3000,
ConfirmTimeout: 30,
ResolverDefault: "cf",
ResolverFallback: "fake",
ProbeURL: "http://gstatic.com/generate_204",
ProbeInterval: "60s",
SchemaVersion: 1,
ActiveProfile: "home",
PanelPort: 8090,
DNSFilter: true,
DNSIntercept: true,
BlockDoH: true,
StatsBackend: "off",
},
Inbounds: []Inbound{
{
@@ -312,6 +314,55 @@ func TestGroupHealthRoundTrip(t *testing.T) {
}
}
// TestL3TunnelRoundTrip pins the l3_tunnel opt-in: an EXPLICIT true (L3 ingress
// on) survives WriteUCI->ReadUCI, and an ABSENT option stays false (opt-in,
// never accidentally on).
func TestL3TunnelRoundTrip(t *testing.T) {
for _, v := range []bool{true, false} {
g := DefaultGlobals()
g.L3Tunnel = v
got, err := ParseUCIExport(RenderUCIExport(&Model{Globals: g}))
if err != nil {
t.Fatalf("v=%v parse: %v", v, err)
}
if got.Globals.L3Tunnel != v {
t.Fatalf("v=%v round-trip: l3_tunnel=%v, want %v", v, got.Globals.L3Tunnel, v)
}
}
// Absent option => off (opt-in default, the Go zero value).
got, err := ParseUCIExport("package shater\n")
if err != nil {
t.Fatal(err)
}
if got.Globals.L3Tunnel {
t.Fatalf("absent l3_tunnel = %v, want false (default, opt-in)", got.Globals.L3Tunnel)
}
}
// TestUntunnelableEgressRoundTrip pins the untunnelable_egress option: an
// EXPLICIT name survives WriteUCI->ReadUCI (else the kernel carrier for
// ESP/AH/GRE/IGMP/SCTP silently switches off on the next re-render), and an
// ABSENT option stays "" so the Untunnelable policy remains in sole charge.
func TestUntunnelableEgressRoundTrip(t *testing.T) {
g := DefaultGlobals()
g.UntunnelableEgress = "wan2"
got, err := ParseUCIExport(RenderUCIExport(&Model{Globals: g}))
if err != nil {
t.Fatalf("parse: %v", err)
}
if got.Globals.UntunnelableEgress != "wan2" {
t.Fatalf("round-trip: untunnelable_egress=%q, want %q", got.Globals.UntunnelableEgress, "wan2")
}
// Absent option => "" (default: no kernel carrier, policy alone decides).
got, err = ParseUCIExport("package shater\n")
if err != nil {
t.Fatal(err)
}
if got.Globals.UntunnelableEgress != "" {
t.Fatalf("absent untunnelable_egress = %q, want \"\" (default, opt-in)", got.Globals.UntunnelableEgress)
}
}
// TestNodeEgressRoundTrip pins the node-level egress binding (multi-WAN): a
// node's `egress` option survives WriteUCI->ReadUCI so the generator can bind the
// node's own upstream to that egress outbound. An absent option parses to "".
+7
View File
@@ -346,6 +346,13 @@ func applyGlobals(g *Globals, s uciSection) {
// stays ON (opt-out); an EXPLICIT "0" disables our background probing.
g.GroupHealth = s.optBool("group_health", g.GroupHealth)
g.Untunnelable = s.optOr("untunnelable", g.Untunnelable)
// Opt-in, and the default is the Go zero value (DefaultGlobals does not seed
// it): an ABSENT option keeps the L3 ingress OFF; only an EXPLICIT "1" opens it.
g.L3Tunnel = s.optBool("l3_tunnel", g.L3Tunnel)
// Which egress carries the protocols the engine cannot (ESP/AH/GRE/IGMP/SCTP).
// Default "" (DefaultGlobals does not seed it): the Untunnelable policy above
// stays in sole charge, exactly as before the option existed.
g.UntunnelableEgress = s.optOr("untunnelable_egress", g.UntunnelableEgress)
g.StatsBackend = s.optOr("stats_backend", g.StatsBackend)
// Stats-size knobs: 0 = UNLIMITED, N = limit. The default (g.*) is the DefaultGlobals
// seed (200/60/5000), so an ABSENT option stays bounded; an EXPLICIT "0" parses to 0
+235
View File
@@ -116,6 +116,17 @@ func isWANIfaceName(name string) bool { return strings.HasPrefix(name, "wan") }
// inbound connections — port forwards stop working, and nothing in the UI explains
// why.
//
// THE REVERSE CHECK IS NOT HERE, ON PURPOSE. "This router has a LAN the config
// never mentions" is the mirror image of this one and matters more (the factory
// config ships a single `lan` inbound, and a VLAN added in LuCI is then proxied by
// nothing and covered by no kill switch), but it cannot be answered from this
// package: it needs the ENUMERATION of the router's interfaces and the fw4
// forwarding graph, and this package is a stdlib-only leaf that netplane imports
// — the dependency runs the wrong way, and NetClassifier deliberately answers one
// question about one name rather than exposing an inventory. It lives in
// netplane/coverage.go, where both the inventory and the divert set already are,
// and reaches Status on the netplane warning channel.
//
// classify may be nil, in which case only the name heuristic applies.
func ValidateRules(rules []Rule, classify NetClassifier) []Warning {
if classify == nil {
@@ -309,6 +320,28 @@ func ValidateGlobals(g Globals) []Warning {
"(non-TCP/UDP traffic is dropped); valid values are block/icmp/direct.", g.Untunnelable))
}
// l3_tunnel + untunnelable=direct is a conflict of intent, not a typo. The L3
// mark is stamped in prerouting and the routing decision carries echo into the
// engine's TUN before the forward chain — where the `direct` accept lives — is
// ever consulted, so `direct` no longer buys the operator anything for ping.
// What it still does is exactly what l3_tunnel was turned on to stop: it lets
// the never-markable protocols (raw IPsec ESP/AH, PPTP/GRE) out with the real
// address, and it turns any failure to install the L3 route (interface down,
// partial apply) into a SILENT leak — the marked echo falls through to the
// main table, reaches forward, and `direct` waves it out — where `block`
// makes the same failure an honest packet loss. Both values are individually
// valid, which is exactly why nothing else says a word about the combination.
// Warn only — like every check here, the config is never edited behind the
// operator's back. The normalisation mirrors netplane.EffectiveUntunnelable
// (unimportable from here: netplane imports model).
if g.L3Tunnel && strings.ToLower(strings.TrimSpace(g.Untunnelable)) == "direct" {
add("l3_tunnel is on, so ping already travels through the tunnel whatever this policy " +
"says — \"direct\" no longer buys you anything for ping, but it still sends raw VPN " +
"passthrough (IPsec ESP/AH, PPTP/GRE) out with your real IP address, and if the L3 " +
"route ever fails to come up your pings silently fall back to leaking directly " +
"instead of failing. With \"block\" the same failure is honest packet loss.")
}
switch strings.ToLower(strings.TrimSpace(g.StatsBackend)) {
case "", "off", "memory", "sqlite":
default:
@@ -393,6 +426,206 @@ func ValidateGroups(groups []Group) []Warning {
return out
}
// ValidateUntunnelableEgress reports an untunnelable_egress option that cannot
// do what it says. It is a cross-section check: the option lives in globals but
// only means something when it resolves to an interface/tunnel egress with a
// device, so neither ValidateGlobals (which sees no egresses) nor a per-egress
// check (which sees no globals) can catch it.
//
// The stakes are higher than a cosmetic typo. The data plane resolves the name
// with netplane.UntunnelableEgressBinding, which fails CLOSED: an unresolvable
// name renders nothing and the plain Untunnelable policy stays in sole charge.
// The operator who set the option believes ESP/AH/GRE/IGMP/SCTP now leave
// through the named egress; in reality they are still governed by the policy —
// dropped, or let out directly — and nothing else says so. The resolution here
// mirrors that binding (unimportable from here: netplane imports model), except
// the device lookup: only a statically empty Interface is checkable in this
// leaf package, a name that fails to resolve at apply time is netplane's to
// report.
func ValidateUntunnelableEgress(g Globals, egresses []Egress) []Warning {
name := strings.TrimSpace(g.UntunnelableEgress)
if name == "" {
return nil
}
add := func(msg string) []Warning {
return []Warning{{Section: "globals", Message: msg}}
}
// DUPLICATE OF netplane.UntunnelableEgressBinding by necessity, and it MUST
// change in lockstep with it: if the two resolutions diverge, this warning
// lies exactly in the configs where it matters most.
for _, eg := range egresses {
if !strings.EqualFold(strings.TrimSpace(eg.Name), name) {
continue
}
if t := strings.ToLower(strings.TrimSpace(eg.Type)); t != "interface" && t != "tunnel" {
return add(fmt.Sprintf("untunnelable_egress %q is an egress of type %q, not "+
"interface/tunnel: it has no device and no routing table of its own, so the "+
"kernel has nowhere to route ESP/AH/GRE/IGMP/SCTP. The option is ignored and "+
"those protocols stay on the untunnelable policy.", g.UntunnelableEgress, eg.Type))
}
if strings.TrimSpace(eg.Interface) == "" {
return add(fmt.Sprintf("untunnelable_egress %q names an egress with no interface, "+
"so there is no device to route out of. The option is ignored and "+
"ESP/AH/GRE/IGMP/SCTP stay on the untunnelable policy.", g.UntunnelableEgress))
}
return nil
}
return add(fmt.Sprintf("untunnelable_egress %q does not match any configured egress, so "+
"the option is ignored: ESP/AH/GRE/IGMP/SCTP stay on the untunnelable policy instead "+
"of leaving through that egress.", g.UntunnelableEgress))
}
// The mark/table layout the netplane derives from fwmark_base and table_base.
//
// DUPLICATED FROM shater/netplane (unimportable from here: netplane imports
// model) and MUST change in lockstep with it. Unlike a duplicated string
// comparison, a drift here is silent in both directions — the validator would
// bless a base the data plane cannot use, or condemn one it can. netplane's
// TestMarkLayoutConstantsLockstep asserts every constant below against the
// originals, so the drift is a build-time failure and not a field report.
const (
// MarkLoopGuard is the fixed mark on the engine's OWN traffic. It is not
// derived from fwmark_base and cannot move, so every derived mark has to
// stay off it.
MarkLoopGuard = 0xff
// MarkL3Offset / MarkEgressOffset: L3 mark = base+0x80, egress #i = base+0x100+i.
MarkL3Offset = 0x80
MarkEgressOffset = 0x100
// TableL3Offset / TableEgressOffset: L3 table = base+0x08, egress #i = base+0x10+i.
TableL3Offset = 0x08
TableEgressOffset = 0x10
// maxFwmark is the width of an nft/SO_MARK mark and of a routing table id.
maxFwmark = 0xffffffff
)
// reservedTables are the routing table ids the kernel owns. Writing our default
// route into one of them is not a collision, it is an amputation: table 254 is
// `main`, and this codebase does `ip route flush table <n>` on teardown.
var reservedTables = map[uint64]string{
0: "unspec (not a usable table)",
253: "the kernel's `default` table",
254: "the kernel's MAIN routing table — flushing it takes the whole router off the network",
255: "the kernel's `local` table — flushing it breaks the router's own addresses",
}
// derivedValue is one number the mark/table layout computes from a base, paired
// with an operator-facing description of where it came from.
type derivedValue struct {
what string
v uint64
}
// ValidateMarkBases reports a fwmark_base or table_base whose DERIVED values
// collide with something that already exists.
//
// The panel offers both under "Advanced", as free hex fields with no bounds at
// all, and neither the model, the netplane nor the applier has ever checked
// them. What makes that more than a footgun is that the derived values are not
// obvious from the number typed: fwmark_base 0x7f looks harmless and is not —
// the L3 mark is base+0x80, so it lands exactly on 0xff, the loop-guard mark the
// engine stamps on its OWN traffic, and `ip rule fwmark 0xff lookup <l3 table>`
// then routes every packet the engine sends into the engine's own TUN. The
// router loses the internet the moment l3_tunnel is switched on, for a reason
// nothing on screen connects to a number in a collapsed "Advanced" section.
// fwmark_base 0xff had produced the same class of failure since long before the
// L3 offset existed; adding an offset simply added a second value that does it.
//
// It takes the egresses because the derived range depends on how many there are
// (egress #i uses base+0x100+i), which Globals alone cannot say.
//
// The collision test is deliberately written as "build every value this layout
// derives, then look for duplicates and reserved ids" rather than as a list of
// known-bad bases. A future offset added to the layout is then covered by
// construction — it only has to be listed in the two loops below, and
// netplane's constant-lockstep test is what makes sure it is.
func ValidateMarkBases(g Globals, egresses []Egress) []Warning {
var out []Warning
add := func(msg string) { out = append(out, Warning{Section: "globals", Message: msg}) }
// 0 means "unset" for both fields (netplane's effFwmark/effTable substitute
// the 0x2000 contract default), so it is not a value to judge.
if b := uint64(g.FwmarkBase); b != 0 {
marks := []derivedValue{
{"the divert mark (fwmark_base itself)", b},
{"the L3-tunnel mark (fwmark_base + 0x80)", b + MarkL3Offset},
}
for i := range egresses {
marks = append(marks, derivedValue{
fmt.Sprintf("egress %q's mark (fwmark_base + 0x100 + %d)", egresses[i].Name, i),
b + MarkEgressOffset + uint64(i),
})
}
for _, d := range marks {
if d.v > maxFwmark {
add(fmt.Sprintf("fwmark_base 0x%x is too large: %s would be 0x%x, past the 32-bit "+
"limit of a firewall mark, so it wraps around onto another mark. Use a base below "+
"0x%x.", g.FwmarkBase, d.what, d.v, uint64(maxFwmark)-MarkEgressOffset-uint64(len(egresses))))
break
}
if d.v == MarkLoopGuard {
add(fmt.Sprintf("fwmark_base 0x%x is unusable: %s comes out as 0x%x, which is the mark "+
"the engine stamps on its OWN outgoing traffic. The policy rule for that mark would "+
"then capture everything the engine sends and route it back into shater's own tables — "+
"the router loses its internet connection and the panel cannot explain why. Pick "+
"another base (the default 0x2000 is clear of everything).",
g.FwmarkBase, d.what, d.v))
}
}
if a, b2, dup := firstDuplicate(marks); dup {
add(fmt.Sprintf("fwmark_base 0x%x makes two different things share one mark: %s and %s are "+
"both 0x%x, so the policy rule for one of them routes the other's traffic as well.",
g.FwmarkBase, a.what, b2.what, a.v))
}
}
if b := uint64(g.TableBase); b != 0 {
tables := []derivedValue{
{"the divert table (table_base itself)", b},
{"the L3-tunnel table (table_base + 0x8)", b + TableL3Offset},
}
for i := range egresses {
tables = append(tables, derivedValue{
fmt.Sprintf("egress %q's table (table_base + 0x10 + %d)", egresses[i].Name, i),
b + TableEgressOffset + uint64(i),
})
}
for _, d := range tables {
if d.v > maxFwmark {
add(fmt.Sprintf("table_base 0x%x is too large: %s would be %d, past the 32-bit limit of a "+
"routing table id.", g.TableBase, d.what, d.v))
break
}
if why, bad := reservedTables[d.v]; bad {
add(fmt.Sprintf("table_base 0x%x is unusable: %s comes out as table %d, which is %s. "+
"shater owns the tables it derives — it writes a default route into them and flushes "+
"them on teardown — so this one must be a table nothing else uses. Pick another base "+
"(the default 0x2000 is clear of everything).", g.TableBase, d.what, d.v, why))
}
}
if a, b2, dup := firstDuplicate(tables); dup {
add(fmt.Sprintf("table_base 0x%x makes two different things share one routing table: %s and %s "+
"are both table %d, so whichever is applied last overwrites the other's default route.",
g.TableBase, a.what, b2.what, a.v))
}
}
return out
}
// firstDuplicate reports the first pair of derived values that collide. The
// current layout cannot produce one without overflowing first; it is checked
// anyway because "the offsets do not overlap" is a property of the arithmetic
// that a future offset can quietly break, and this is where that would surface.
func firstDuplicate(vals []derivedValue) (derivedValue, derivedValue, bool) {
seen := map[uint64]derivedValue{}
for _, d := range vals {
if prev, ok := seen[d.v]; ok {
return prev, d, true
}
seen[d.v] = d
}
return derivedValue{}, derivedValue{}, false
}
// Validate runs every model-level check and returns the combined warnings. It
// never fails: the result is advisory, for the log and the panel.
func (m *Model) Validate(classify NetClassifier) []Warning {
@@ -401,6 +634,8 @@ func (m *Model) Validate(classify NetClassifier) []Warning {
}
out := ValidateRules(m.Rules, classify)
out = append(out, ValidateGlobals(m.Globals)...)
out = append(out, ValidateUntunnelableEgress(m.Globals, m.Egresses)...)
out = append(out, ValidateMarkBases(m.Globals, m.Egresses)...)
out = append(out, ValidateSubscriptions(m.Subscriptions)...)
out = append(out, ValidateGroups(m.Groups)...)
out = append(out, ValidateAlerts(m.Alerts)...)
+87
View File
@@ -222,6 +222,93 @@ func TestValidateGlobalsUnknownEnums(t *testing.T) {
}
}
// TestValidateGlobalsL3TunnelDirectConflict: two individually-valid options whose
// combination quietly downgrades the L3 ingress — `direct` keeps letting raw
// IPsec/GRE out with the real address and turns an L3-route failure into a silent
// ping leak instead of an honest loss. The validator must say so, and must say
// nothing for either option alone or for l3_tunnel with the safe policies.
func TestValidateGlobalsL3TunnelDirectConflict(t *testing.T) {
ws := ValidateGlobals(Globals{L3Tunnel: true, Untunnelable: "direct"})
if !hasWarning(ws, "globals", "l3_tunnel") {
t.Fatalf("l3_tunnel + untunnelable=direct must warn, got %v", ws)
}
// The message has to name the safe alternative, or it sends the user hunting.
if !hasWarning(ws, "globals", "\"block\"") {
t.Errorf("the warning must point at \"block\" as the safe policy, got %v", ws)
}
// Normalisation mirrors netplane.EffectiveUntunnelable: case and whitespace
// must not hide the conflict.
if ws := ValidateGlobals(Globals{L3Tunnel: true, Untunnelable: " Direct "}); !hasWarning(ws, "globals", "l3_tunnel") {
t.Errorf("normalised \" Direct \" must still warn, got %v", ws)
}
// Only the combination warns: each option alone, and l3_tunnel with the
// policies that keep the failure mode honest, stay quiet.
for _, g := range []Globals{
{Untunnelable: "direct"},
{L3Tunnel: true},
{L3Tunnel: true, Untunnelable: "block"},
{L3Tunnel: true, Untunnelable: "icmp"},
} {
if ws := ValidateGlobals(g); len(ws) != 0 {
t.Errorf("globals %+v must not warn, got %v", g, ws)
}
}
}
// TestValidateUntunnelableEgress: the option fails CLOSED in the data plane —
// netplane.UntunnelableEgressBinding renders nothing for a name that does not
// resolve to an interface/tunnel egress with a device, and the plain
// untunnelable policy quietly stays in charge. The operator believes
// ESP/AH/GRE/IGMP/SCTP now leave through the named egress; without this warning
// nothing tells them the option did nothing. Each of the three ways to
// mis-point the option must be named, and a correctly-pointed option must stay
// silent — a false warning on a working config teaches operators to ignore the
// real ones.
func TestValidateUntunnelableEgress(t *testing.T) {
egresses := []Egress{
{Name: "wan2", Type: "interface", Interface: "wan2"},
{Name: "dpi", Type: "byedpi", Port: 1080},
{Name: "nodev", Type: "tunnel"},
}
// The name does not match any egress: the option is dropped on the floor.
ws := ValidateUntunnelableEgress(Globals{UntunnelableEgress: "wan3"}, egresses)
if !hasWarning(ws, "globals", "does not match any configured egress") {
t.Fatalf("an unresolvable name must warn (the option silently does nothing), got %v", ws)
}
// The name matches, but the egress type has no device or routing table.
ws = ValidateUntunnelableEgress(Globals{UntunnelableEgress: "dpi"}, egresses)
if !hasWarning(ws, "globals", "not interface/tunnel") {
t.Fatalf("a byedpi/direct egress must warn (nowhere to route to), got %v", ws)
}
// The message has to name the offending type, or the operator cannot see
// which of their egresses they mis-picked.
if !hasWarning(ws, "globals", "\"byedpi\"") {
t.Errorf("the warning must name the egress's actual type, got %v", ws)
}
// Right type, but no interface: still no device to route out of.
ws = ValidateUntunnelableEgress(Globals{UntunnelableEgress: "nodev"}, egresses)
if !hasWarning(ws, "globals", "no interface") {
t.Fatalf("an interface-less egress must warn (no device to route out of), got %v", ws)
}
// Resolution mirrors netplane.UntunnelableEgressBinding: case and whitespace
// must not hide a warning the apply stage would act on.
if ws := ValidateUntunnelableEgress(Globals{UntunnelableEgress: " DPI "}, egresses); !hasWarning(ws, "globals", "not interface/tunnel") {
t.Errorf("normalised \" DPI \" must still warn, got %v", ws)
}
// The quiet cases: an unset option, and one pointing at a well-formed
// interface egress (any casing), produce no warnings at all.
for _, name := range []string{"", "wan2", " WAN2 "} {
if ws := ValidateUntunnelableEgress(Globals{UntunnelableEgress: name}, egresses); len(ws) != 0 {
t.Errorf("untunnelable_egress=%q must not warn, got %v", name, ws)
}
}
}
// TestDeadKnobsAreGoneNotDocumented is the deletion regression. dns_mode and the
// two xudp_* node options were confirmed to have no reader anywhere, so they were
// REMOVED rather than left with a "not used" comment: a value that survives a
+253 -15
View File
@@ -10,6 +10,7 @@ import (
"encoding/json"
"fmt"
"net/netip"
"os"
"os/exec"
"strings"
"sync"
@@ -108,23 +109,33 @@ func ApplyRouting(m *model.Model) error {
// ApplyRoutingWithWarnings reconciles ip rule/route: fwmark(FwmarkBase) ->
// table(TableBase) with `local default dev lo`, so tproxy-marked packets are delivered
// locally, plus the per-egress mark->table->device bindings. Idempotent (del-then-add).
// locally, plus the per-egress mark->table->device bindings and the L3-ingress
// binding (L3Mark -> L3Table -> L3Device). Idempotent (del-then-add).
//
// It mirrors RenderNftWithWarnings: the warnings are operator-facing statements about an
// egress that WAS built but cannot carry traffic, which apply folds into the status
// It mirrors RenderNftWithWarnings: the warnings are operator-facing statements about a
// binding that WAS built but cannot carry traffic, which apply folds into the status
// warning list so the panel shows them.
func ApplyRoutingWithWarnings(m *model.Model) ([]string, error) {
if err := addRouting(effFwmark(m.Globals), effTable(m.Globals), m.Globals.IPv6); err != nil {
return nil, err
}
return addEgressRouting(m)
warnings, err := addEgressRouting(m)
if err != nil {
return warnings, err
}
l3Warnings, err := addL3Routing(m)
return append(warnings, l3Warnings...), err
}
// TeardownRouting removes the main tproxy rule/table and every reserved per-egress
// rule/table. Safe when nothing is set up.
// TeardownRouting removes the main tproxy rule/table, every reserved per-egress
// rule/table, and the L3-ingress rule/table. Safe when nothing is set up. The L3
// pair is removed UNCONDITIONALLY (not gated on L3Enabled) for the same reason
// the egress block is swept wholesale: teardown must clean up what a PREVIOUS
// config installed, and the current model cannot testify about the past.
func TeardownRouting(m *model.Model) error {
removeRouting(effFwmark(m.Globals), effTable(m.Globals))
removeEgressRouting(m)
removeRouting(L3Mark(m.Globals), L3Table(m.Globals))
return nil
}
@@ -239,6 +250,51 @@ func runIP(args ...string) error {
return nil
}
// unreachableFloorMetric is the metric of the `unreachable default` route every
// mark-driven table gets as its last resort. Routes to the same prefix are
// ordered by metric, so any real default route (the kernel gives them metric 0
// unless asked otherwise) always wins while it exists; this one is only ever
// consulted once the real one is gone. It is deliberately absurd rather than
// merely large, so nothing a person might plausibly configure can outrank it.
const unreachableFloorMetric = "4294967295"
// addUnreachableFloor installs the fail-closed floor of a mark-driven routing
// table: a metric-maximal `unreachable default`.
//
// # Why a fallthrough is not a safety net
//
// `ip rule fwmark X lookup N` does not send the packet to table N. It sends the
// LOOKUP to table N, and a lookup that finds nothing there simply continues down
// the rule list — to `main`. So an empty table N is not "this traffic is stuck",
// it is "this traffic is routed exactly as if it had never been marked": out the
// default WAN, with the router's real address, wearing a mark the firewall was
// told to trust. Every leak in this file's history is a variation on that one
// sentence.
//
// The table empties for entirely ordinary reasons and, crucially, WITHOUT
// anything we control changing: an egress rides an interface, `ifdown wg0` or a
// `network restart` happens, and the kernel garbage-collects every route through
// that device. The L3 table empties the same way when the engine is restarted
// and its TUN is destroyed and recreated. The nft ruleset is byte-identical
// across all of that, so the applier's idempotence check sees nothing to do.
//
// An `unreachable` route has no device, so the kernel has no reason to remove it
// when a device dies — which is exactly the property needed. Once it is in the
// table the lookup can no longer FAIL, only succeed with "unreachable", and the
// fallthrough to main becomes structurally impossible instead of merely
// unlikely. That is worth more than the warning it replaces: warnings depend on
// somebody reading them, at the moment the interface goes down, which is not the
// moment anybody is reading anything.
//
// This is deliberately unconditional — not gated on the kill-switch. The
// kill-switch decides whether traffic may escape the TUNNEL; an egress binding
// is a statement about WHICH UPLINK, and silently substituting a different one
// is never what "fail open" was meant to permit.
func addUnreachableFloor(fam, tableArg string) error {
return runIP(fam, "route", "add", "unreachable", "default",
"metric", unreachableFloorMetric, "table", tableArg)
}
// addEgressRouting realises the policy-routing half of an interface/tunnel egress
// (the generator emits the SO_BINDTODEVICE+SO_MARK outbound; this binds the mark
// to a table whose default route leaves via the egress device). Idempotent.
@@ -269,11 +325,15 @@ func addEgressRouting(m *model.Model) ([]string, error) {
var warnings []string
var bindings []egressBinding
for i, eg := range m.Egresses {
t := strings.ToLower(eg.Type)
if (t != "interface" && t != "tunnel") || eg.Interface == "" {
// ONE resolution, shared with the prerouting marking and the forward
// accept (see EgressDevice): an egress this skips is an egress whose mark
// is neither stamped nor accepted anywhere, and vice versa. When these
// drifted apart, the difference was a marked packet with no table.
dev := EgressDevice(eg)
if dev == "" {
continue
}
dev := IfaceDevice(eg.Interface)
iface := strings.TrimSpace(eg.Interface)
mark := EgressMark(m.Globals, i)
table := EgressTable(m.Globals, i)
fams := []string{"-4"}
@@ -298,10 +358,18 @@ func addEgressRouting(m *model.Model) ([]string, error) {
continue
}
run(fam, "route", "flush", "table", tableArg)
if err := addUnreachableFloor(fam, tableArg); err != nil {
warnings = append(warnings, fmt.Sprintf(
"egress %q: its %s routing table %d could not be given a fail-closed floor (%v), so if "+
"interface %q (device %s) ever goes down the kernel deletes the route in that table, the "+
"lookup falls through to the main table, and traffic bound to this egress leaves over the "+
"plain WAN with your real IP address instead of failing.",
eg.Name, famLabel(fam), table, err, iface, dev))
}
// A gateway'd interface (WAN) needs `via <gw>`; a point-to-point tunnel
// (wg/awg) has no gateway and routes straight out the device.
gw := egressGateway(fam, eg.Interface, dev)
gw := egressGateway(fam, iface, dev)
var rerr error
if gw != "" {
rerr = runIP(fam, "route", "add", "default", "via", gw, "dev", dev, "table", tableArg)
@@ -314,7 +382,7 @@ func addEgressRouting(m *model.Model) ([]string, error) {
"and this egress is NOT applied: every node, group and rule bound to it is marked for an "+
"empty table, falls through to the main table and leaves over the plain WAN with your real "+
"IP address. Check that interface %q (device %s) exists and is up. Reason: %v",
eg.Name, famLabel(fam), table, eg.Interface, dev, rerr))
eg.Name, famLabel(fam), table, iface, dev, rerr))
continue
}
if gw != "" {
@@ -332,7 +400,7 @@ func addEgressRouting(m *model.Model) ([]string, error) {
"CANNOT REACH ANYTHING outside its own subnet — every node, group and "+
"rule bound to it will fail to connect. Check that the interface is up "+
"and has a lease, or set a static nexthop (network.%s.%s).",
eg.Name, famLabel(fam), eg.Interface, dev, eg.Interface, uciGatewayOption(fam)))
eg.Name, famLabel(fam), iface, dev, iface, uciGatewayOption(fam)))
}
}
}
@@ -340,6 +408,66 @@ func addEgressRouting(m *model.Model) ([]string, error) {
return warnings, nil
}
// addL3Routing realises the policy-routing half of the L3 ingress (Globals.
// L3Tunnel): the nft prerouting chain stamps L3Mark on LAN ICMP/ICMPv6, and this
// binds that mark to a table whose only content is `default dev shater-l3`. That
// rule+table pair is deliberately the ONLY path into the TUN — the engine opens
// it with auto_route off, so the main routing table is never touched and
// disabling the feature can never strand a stale default route there. Idempotent
// (del-then-add), same shape as addEgressRouting above.
//
// Failures are WARNINGS, not errors, for addEgressRouting's documented reason:
// returning would abort the whole apply — sysctls, conntrack flush and all —
// over one broken binding. The `icmp "tunnel":` prefix is the `kind "name":`
// shape apply's warning normaliser parses, so this lands in the panel as a named
// warning (the kind says what the operator loses: ICMP), and the wording leads
// with the observable consequence rather than the mechanics.
func addL3Routing(m *model.Model) ([]string, error) {
if !L3Enabled(m.Globals) {
return nil, nil
}
run := func(args ...string) { _ = execCommand("ip", args...).Run() }
var warnings []string
markArg := fmt.Sprintf("0x%x", L3Mark(m.Globals))
tableArg := fmt.Sprintf("%d", L3Table(m.Globals))
fams := []string{"-4"}
if m.Globals.IPv6 {
fams = append(fams, "-6")
}
for _, fam := range fams {
run(fam, "rule", "del", "fwmark", markArg, "lookup", tableArg)
if err := runIP(fam, "rule", "add", "fwmark", markArg, "lookup", tableArg); err != nil {
warnings = append(warnings, fmt.Sprintf(
"icmp \"tunnel\": %s ping from LAN clients will NOT work: their ICMP is marked for the L3 "+
"tunnel, but the policy rule (fwmark %s -> table %s) could not be installed, so the marked "+
"packets fall through to the main routing table and the kill-switch treats them like any "+
"escaped traffic — dropped when closed, out the plain WAN with your real IP address when "+
"open. Reason: %v",
famLabel(fam), markArg, tableArg, err))
continue
}
run(fam, "route", "flush", "table", tableArg)
if err := addUnreachableFloor(fam, tableArg); err != nil {
warnings = append(warnings, fmt.Sprintf(
"icmp \"tunnel\": %s table %s could not be given a fail-closed floor (%v), so whenever the "+
"engine restarts and its %s device is destroyed, the kernel empties that table and the "+
"marked ICMP falls through to the main routing table — dropped when the kill-switch is "+
"closed, out the plain WAN with your real IP address when it is open.",
famLabel(fam), tableArg, err, L3Device))
}
if err := runIP(fam, "route", "add", "default", "dev", L3Device, "table", tableArg); err != nil {
warnings = append(warnings, fmt.Sprintf(
"icmp \"tunnel\": %s ping from LAN clients will NOT work: their ICMP is marked for the L3 "+
"tunnel, but table %s got no route into %s (is the engine running and its TUN up?), so the "+
"marked packets fall through to the main routing table and the kill-switch treats them like "+
"any escaped traffic — dropped when closed, out the plain WAN with your real IP address when "+
"open. Reason: %v",
famLabel(fam), tableArg, L3Device, err))
}
}
return warnings, nil
}
// famLabel renders an `ip` family flag for an operator-facing message.
func famLabel(fam string) string {
if fam == "-6" {
@@ -563,6 +691,23 @@ func isPointToPoint(dev string) bool {
// Which bindings must exist comes from the record addEgressRouting keeps (see
// egressBindings): Globals alone cannot say how many egresses there are, and a
// deleted rule leaves nothing on the system to count.
//
// # And the L3 ingress, for the third time
//
// addL3Routing installs a third `fwmark -> table` + default-route pair
// (Globals.L3Tunnel), and it was added without being added here — the exact
// defect the two paragraphs above describe as already caught twice, committed a
// third time. Its trigger is not even an interface going down: the engine
// recreates the shater-l3 TUN on any config change that restarts it (a node URI
// edited, a subscription refreshed), the kernel deletes `default dev shater-l3`
// with the old device and does not restore it with the new one, and the nft text
// is unchanged so the fast-path skips ApplyRouting forever. LAN ping then stays
// dead until somebody restarts the daemon, and nothing on screen says why.
//
// Unlike the egress bindings this one needs no record: L3Mark, L3Table and
// L3Enabled are all derivable from Globals, which is what this function is
// handed. TestEveryStampedMarkIsRoutedAndVerified is what stops the FOURTH mark
// from being added without a line here.
func RoutingPresent(g model.Globals) bool {
mark, table := effFwmark(g), effTable(g)
fams := []string{"-4"}
@@ -574,9 +719,63 @@ func RoutingPresent(g model.Globals) bool {
return false
}
}
if !l3RoutingPresent(g, fams) {
return false
}
return egressRoutingPresent()
}
// l3RoutingPresent reports whether the L3-ingress binding is installed for every
// family the model asks for: the `fwmark -> table` rule AND a real default route
// inside that table. With Globals.L3Tunnel off nothing is installed by design,
// so the answer is trivially yes.
//
// A false answer while the engine's TUN is genuinely absent (the engine is
// down, or coming up) is intended, not a wart: it makes every reconcile re-run
// ApplyRouting, so the route reappears by itself the moment the device does.
// That is the same self-repair contract the per-egress bindings have, and the
// main tproxy pair is guarded against the churn by addRouting's own presence
// check.
func l3RoutingPresent(g model.Globals, fams []string) bool {
if !L3Enabled(g) {
return true
}
mark, table := L3Mark(g), L3Table(g)
for _, fam := range fams {
out, err := execCommand("ip", fam, "rule", "show").Output()
if err != nil || !fwmarkRulePresent(string(out), mark) {
return false
}
rout, rerr := execCommand("ip", fam, "route", "show", "table", fmt.Sprintf("%d", table)).Output()
if rerr != nil || !realDefaultRoutePresent(string(rout)) {
return false
}
}
return true
}
// realDefaultRoutePresent reports whether a table holds a default route that
// actually FORWARDS something: `default via <gw> dev X`, `default dev X`.
//
// It exists because addUnreachableFloor puts an `unreachable default` in every
// mark-driven table, and a substring test for "default" would be satisfied by
// that floor alone — turning the safety net into a blindfold, and reporting the
// precise state it was installed to survive (device gone, real route deleted) as
// a healthy plane that needs no repair. The floor keeps the packet from leaking;
// this keeps the reconcile from giving up on fixing it.
//
// `ip route show` prints the route type as the FIRST token for every non-unicast
// type (`unreachable default ...`, `blackhole default ...`) and omits it for
// unicast, so a line that starts with "default" is a real one.
func realDefaultRoutePresent(out string) bool {
for _, line := range strings.Split(out, "\n") {
if strings.HasPrefix(strings.TrimSpace(line), "default") {
return true
}
}
return false
}
// mainRoutingPresentFam reports whether BOTH halves of the main tproxy plane are
// installed for one family: the fwmark rule and the `local default dev lo` route
// inside the table it points at.
@@ -618,11 +817,14 @@ func egressRoutingPresent() bool {
if !fwmarkRulePresent(out, b.mark) {
return false
}
// Any default route will do: `default via <gw> dev X` for a gateway'd
// Any REAL default route will do: `default via <gw> dev X` for a gateway'd
// uplink, `default dev X` for a point-to-point tunnel. What must never pass
// is an EMPTY table, which is what an ifdown leaves behind.
// is an EMPTY table — nor a table holding only addUnreachableFloor's
// `unreachable default`, which is the shape an ifdown leaves behind now
// that the floor is installed, and which means the very thing this check
// exists to catch.
rout, err := execCommand("ip", b.fam, "route", "show", "table", fmt.Sprintf("%d", b.table)).Output()
if err != nil || !strings.Contains(string(rout), "default") {
if err != nil || !realDefaultRoutePresent(string(rout)) {
return false
}
}
@@ -657,11 +859,47 @@ func fwmarkRulePresent(out string, mark uint32) bool {
return false
}
// sysClassNet is the kernel's list of network devices. A var only so a test can
// point DeviceExists at a fixture directory.
var sysClassNet = "/sys/class/net"
// DeviceExists reports whether the kernel actually has a network device by this
// name.
//
// It exists because IfaceDevice's last resort is to hand back the name it was
// given (see below), and nothing downstream could tell that apart from a real
// device: nftValidDev checks the CHARACTERS of a name, not its existence, so
// `lan` passes as readily as `br-lan` and produces rules that load cleanly and
// match nothing. This is the one cheap question that separates the two, and it
// asks the kernel rather than netifd so a device netifd does not manage still
// counts. Used by coverage.go; deliberately not consulted by IfaceDevice itself,
// which must keep returning a usable string for every caller.
func DeviceExists(dev string) bool {
dev = strings.TrimSpace(dev)
if dev == "" || strings.ContainsAny(dev, "/\\") {
return false
}
_, err := os.Stat(sysClassNet + "/" + dev)
return err == nil
}
// IfaceDevice resolves a UCI interface name (e.g. "lan") to its actual L3 device
// (e.g. "br-lan") for nft iifname matching / SO_BINDTODEVICE — bridges/VLANs have
// device != name, so matching the raw UCI name would silently intercept nothing.
// Empty name defaults to br-lan; an unresolvable name is returned unchanged (it
// may already be a device).
//
// THAT LAST FALLBACK IS A KNOWN TRAP, and it is deliberate. Returning the raw
// name is right for the case it was written for — an operator who wrote a device
// name where an interface name was expected — and it is a silent leak for the
// case where the name resolves to nothing at all: the caller renders
// `iifname "lan" ... drop`, the kernel accepts it, and that network is neither
// diverted nor failed closed on. The fallback is kept (refusing here would strand
// every legitimate device-name user) and the ambiguity is resolved one level up
// instead: divertRef carries what the operator wrote alongside what it resolved
// to, and coverage.go asks DeviceExists whether the result is real, reporting the
// gap to the operator. Callers that need the distinction must go through there,
// not re-derive it from the string.
func IfaceDevice(name string) string {
if name == "" {
return "br-lan"
+182
View File
@@ -0,0 +1,182 @@
package netplane
// The BOOT ARMOR: the fail-closed holding plane, persisted to flash so that
// something OTHER than the running daemon can put it back.
//
// # Why a file exists at all
//
// Every protection this package installs lives inside one long-lived Go process.
// That is fine while the process is running, and it is exactly nothing in the two
// windows where it is not:
//
// - BOOT. /etc/init.d/shater is START=99, i.e. after fw4 (START=19) has already
// loaded `lan -> wan ACCEPT` and after netifd (START=20) has brought the LAN
// bridge up. Wi-Fi associates and every client reconnects in that window, and
// the daemon only reaches applier.ArmHold after procd has decompressed a
// UPX-packed binary off flash, waited out any predecessor, run `shaterd
// migrate` and read UCI. On a small router that is seconds — every one of
// them with `kill_switch=closed` configured and the LAN forwarding to the WAN
// in the clear.
// - AN UNREADABLE CONFIG. model.ReadUCI failing (a full /overlay caught mid
// `uci commit` is the documented cause — see /etc/init.d/shater) used to mean
// NO plane was ever installed: the arming call sat in the else-branch of the
// successful read, Reconcile returns before any of it, and the cron watchdog
// sees a live pidof and its own `uci -q get` fails too. Nothing in the system
// could recover, and nothing said so.
//
// Both need a description of the fail-closed plane that survives the process. So
// the daemon writes RenderHoldNft's output here on every apply, and two readers
// use it: /etc/init.d/shater-armor at boot (before the daemon exists), and the
// daemon itself when it cannot read its own config.
//
// # Why the HOLDING plane and not the full ruleset
//
// The full ruleset diverts to the engine's tproxy port. With no engine listening,
// every `tproxy` statement returns NFT_BREAK and the packets fall through to the
// forward chain, which does drop them — so it would "work". But it also needs
// kmod-nft-tproxy loaded, declares dynamic sets and counters, and encodes a
// tproxy port that may no longer be the configured one. The holding plane is a
// single forward chain of accepts and two drops: it needs nothing but nf_tables,
// it is what the daemon itself installs when the engine is down, and it says
// exactly one thing — this device's forwarded traffic does not leave.
//
// # The file IS the arm token
//
// Its presence means "the last configuration this router applied was enabled and
// fail-closed". Nothing else has to be readable for that to be true, which is the
// whole point in the unreadable-config case. Correspondingly it is REMOVED the
// moment that stops being true: globals.enabled=0, kill_switch=open, or a
// deliberate `/etc/init.d/shater stop`. An operator who turned the stack off must
// not find it reinstated by the next power cut.
import (
"fmt"
"os"
"path/filepath"
"strings"
"github.com/sagernet/sing-box/shater/model"
)
// BootArmorPath is where the persisted holding plane lives. It is a var, not a
// const, ONLY so tests can point it at a temp dir; production never assigns it.
//
// /etc/shater, not /var/run: it has to survive the reboot it exists for. The
// daemon already creates that directory for the engine's cache DB (D16).
var BootArmorPath = "/etc/shater/boot.nft"
// KillSwitchClosed reports the EFFECTIVE kill-switch policy: closed (fail-closed)
// unless the operator explicitly wrote "open". Exported so the daemon decides with
// the same rule the ruleset renderer does — a second copy of this predicate is how
// the plane and the policy end up disagreeing.
func KillSwitchClosed(g model.Globals) bool { return genGlobalClosed(g) }
// BootArmorPresent reports whether a persisted holding plane exists.
func BootArmorPresent() bool {
st, err := os.Stat(BootArmorPath)
return err == nil && st.Mode().IsRegular() && st.Size() > 0
}
// SaveBootArmor persists ruleset as the boot armor and reports whether the file's
// contents actually changed.
//
// The no-change fast path is not an optimisation, it is a durability requirement:
// this is called on EVERY apply, cron reconciles once a minute, and the target is
// raw flash on a router that is expected to run for years. Rewriting an identical
// 700-byte file 525 600 times a year is how an eMMC/NAND block wears out for
// nothing.
//
// The write is atomic (temp file in the same directory, then rename), so a power
// cut mid-write leaves either the previous armor or none — never a half-written
// ruleset that `nft -f` would reject at the next boot, which is precisely the
// boot where it matters.
func SaveBootArmor(ruleset string) (changed bool, err error) {
if strings.TrimSpace(ruleset) == "" {
return false, fmt.Errorf("refusing to save an empty boot armor")
}
if old, rerr := os.ReadFile(BootArmorPath); rerr == nil && string(old) == ruleset {
return false, nil
}
dir := filepath.Dir(BootArmorPath)
if err := os.MkdirAll(dir, 0o755); err != nil {
return false, err
}
tmp, err := os.CreateTemp(dir, ".boot.nft.*")
if err != nil {
return false, err
}
name := tmp.Name()
defer func() {
if err != nil {
_ = os.Remove(name)
}
}()
if _, err = tmp.WriteString(ruleset); err != nil {
_ = tmp.Close()
return false, err
}
// Sync before the rename: the rename is what makes the file visible, and on
// the flash filesystems this ships on an unsynced payload can survive the
// rename as zeroes.
if err = tmp.Sync(); err != nil {
_ = tmp.Close()
return false, err
}
if err = tmp.Close(); err != nil {
return false, err
}
if err = os.Chmod(name, 0o644); err != nil {
return false, err
}
if err = os.Rename(name, BootArmorPath); err != nil {
return false, err
}
// And fsync the DIRECTORY. Syncing the file only guarantees its contents; the
// rename that publishes them is a directory operation, and on the flash
// filesystems this ships on (jffs2/f2fs/ubifs, and ext4 on the x86 images) an
// unsynced rename can be lost across a power cut while the file's data is not.
// The result would be a boot with no armor and no error anywhere — i.e. exactly
// the failure this file exists to prevent, on exactly the boot it exists for.
// A best-effort sync: a filesystem that will not open its own directory is not
// a reason to report a write that did happen as failed.
if d, derr := os.Open(dir); derr == nil {
_ = d.Sync()
_ = d.Close()
}
return true, nil
}
// RemoveBootArmor deletes the persisted holding plane. Absent is success: this is
// called on every apply that finds the stack disabled or fail-open, and "there was
// nothing to remove" is the normal case, not a fault.
func RemoveBootArmor() error {
if err := os.Remove(BootArmorPath); err != nil && !os.IsNotExist(err) {
return err
}
return nil
}
// LoadBootArmor installs the persisted holding plane into the kernel and reports
// whether there was one to install.
//
// It goes through ApplyNft rather than `nft -f <path>` so the snapshot is
// VALIDATED (`nft -c`) before it is loaded and so a wedged nft is killed rather
// than left blocking — the same treatment every other ruleset in this package
// gets. A snapshot that no longer parses (an nft downgrade, a truncated file) is
// therefore an error the caller can report, not a half-loaded table.
func LoadBootArmor() (loaded bool, err error) {
b, rerr := os.ReadFile(BootArmorPath)
if rerr != nil {
if os.IsNotExist(rerr) {
return false, nil
}
return false, rerr
}
if strings.TrimSpace(string(b)) == "" {
return false, fmt.Errorf("boot armor %s is empty", BootArmorPath)
}
if err := ApplyNft(string(b)); err != nil {
return false, err
}
return true, nil
}
+178
View File
@@ -0,0 +1,178 @@
package netplane
// The BOOT ARMOR's storage contract. It is the artifact three separate windows
// depend on (early boot, restart handoff, unreadable config), so the properties
// that matter are the durability ones: it must survive, it must not wear flash
// out, and a power cut must never leave a half-written ruleset behind for the one
// boot that reads it.
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
func armorTempPath(t *testing.T) string {
t.Helper()
p := filepath.Join(t.TempDir(), "shater", "boot.nft")
orig := BootArmorPath
BootArmorPath = p
t.Cleanup(func() { BootArmorPath = orig })
return p
}
// TestBootArmorSaveIsContentGated: the daemon calls this on every apply and cron
// reconciles once a minute, so an unconditional write is ~525 000 rewrites a year
// of an identical file onto raw flash. An unchanged ruleset must not touch the
// file at all.
func TestBootArmorSaveIsContentGated(t *testing.T) {
path := armorTempPath(t)
if BootArmorPresent() {
t.Fatalf("BootArmorPresent must be false before anything is written")
}
changed, err := SaveBootArmor("table inet shater { }\n")
if err != nil || !changed {
t.Fatalf("first SaveBootArmor = (%v, %v), want (true, nil)", changed, err)
}
if !BootArmorPresent() {
t.Fatalf("BootArmorPresent must be true after a save")
}
st1, err := os.Stat(path)
if err != nil {
t.Fatalf("stat: %v", err)
}
changed, err = SaveBootArmor("table inet shater { }\n")
if err != nil || changed {
t.Fatalf("identical SaveBootArmor = (%v, %v), want (false, nil) — an unchanged "+
"armor must not be rewritten", changed, err)
}
st2, err := os.Stat(path)
if err != nil {
t.Fatalf("stat: %v", err)
}
if !st1.ModTime().Equal(st2.ModTime()) {
t.Errorf("the file was rewritten for identical content (mtime %v -> %v)",
st1.ModTime(), st2.ModTime())
}
if changed, err = SaveBootArmor("table inet shater { chain forward { } }\n"); err != nil || !changed {
t.Fatalf("changed SaveBootArmor = (%v, %v), want (true, nil)", changed, err)
}
b, err := os.ReadFile(path)
if err != nil || !strings.Contains(string(b), "chain forward") {
t.Fatalf("file does not hold the new ruleset: %q (%v)", string(b), err)
}
// An empty armor is worse than none: shater-armor would load a file that
// blocks nothing while the log claims the LAN is protected.
if _, err := SaveBootArmor(" \n"); err == nil {
t.Errorf("SaveBootArmor must refuse an empty ruleset")
}
if err := RemoveBootArmor(); err != nil {
t.Fatalf("RemoveBootArmor: %v", err)
}
if BootArmorPresent() {
t.Fatalf("BootArmorPresent must be false after removal")
}
// Removing what is not there is the normal case (every apply of a disabled
// stack), not a fault.
if err := RemoveBootArmor(); err != nil {
t.Errorf("RemoveBootArmor on an absent file = %v, want nil", err)
}
}
// TestBootArmorSaveIsAtomic: nothing but the finished file may ever be visible at
// BootArmorPath. The one boot that reads it is the boot after a power cut.
func TestBootArmorSaveIsAtomic(t *testing.T) {
path := armorTempPath(t)
if _, err := SaveBootArmor("table inet shater { }\n"); err != nil {
t.Fatalf("SaveBootArmor: %v", err)
}
entries, err := os.ReadDir(filepath.Dir(path))
if err != nil {
t.Fatalf("readdir: %v", err)
}
if len(entries) != 1 || entries[0].Name() != filepath.Base(path) {
var names []string
for _, e := range entries {
names = append(names, e.Name())
}
t.Errorf("the save left temp files behind: %v", names)
}
}
// TestLoadBootArmorFeedsNft proves the reinstate path actually reaches the kernel
// through the SAME validated loader every other ruleset uses (nft -c, then nft
// -f), rather than a bare `nft -f <file>` that could half-load a truncated
// snapshot.
func TestLoadBootArmorFeedsNft(t *testing.T) {
armorTempPath(t)
// Nothing saved: not an error, just nothing to do.
loaded, err := LoadBootArmor()
if err != nil || loaded {
t.Fatalf("LoadBootArmor with no armor = (%v, %v), want (false, nil)", loaded, err)
}
const ruleset = "table inet shater { chain forward { } }\n"
if _, err := SaveBootArmor(ruleset); err != nil {
t.Fatalf("SaveBootArmor: %v", err)
}
var rec []string
orig := execCommand
defer func() { execCommand = orig }()
execCommand = func(name string, arg ...string) *exec.Cmd {
rec = append(rec, strings.Join(append([]string{name}, arg...), " "))
cs := append([]string{"-test.run=TestNetplaneHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1")
return cmd
}
loaded, err = LoadBootArmor()
if err != nil || !loaded {
t.Fatalf("LoadBootArmor = (%v, %v), want (true, nil)", loaded, err)
}
var sawCheck, sawLoad bool
for _, c := range rec {
switch c {
case "nft -c -f -":
sawCheck = true
case "nft -f -":
sawLoad = true
}
}
if !sawCheck || !sawLoad {
t.Errorf("LoadBootArmor must VALIDATE then load (nft -c -f -, nft -f -); commands: %v", rec)
}
}
// TestKillSwitchClosed pins that the daemon's arming decision and the renderer's
// drop decision come from one predicate. Two copies of "is the kill switch
// closed" is how an armed boot ends up protecting a router the operator asked to
// fail open (or the reverse).
func TestKillSwitchClosed(t *testing.T) {
cases := []struct {
val string
want bool
}{
{"", true}, {"closed", true}, {"CLOSED", true}, {"nonsense", true},
{"open", false}, {"OPEN", false}, {" open ", false},
}
for _, c := range cases {
if got := KillSwitchClosed(model.Globals{KillSwitch: c.val}); got != c.want {
t.Errorf("KillSwitchClosed(%q) = %v, want %v", c.val, got, c.want)
}
if got := genGlobalClosed(model.Globals{KillSwitch: c.val}); got != c.want {
t.Errorf("genGlobalClosed(%q) = %v, want %v — the two must not diverge", c.val, got, c.want)
}
}
}
+234
View File
@@ -0,0 +1,234 @@
package netplane
// COVERAGE: the gap between the networks this router carries and the networks
// this data plane actually touches.
//
// Everything else in this package answers "what does the plan say?". This file
// answers the question nothing asked before: "what does the plan leave out?".
// Two omissions, both silent, both leaks:
//
// 1. A DEVICE NAME THAT NAMES NOTHING. IfaceDevice returns the raw UCI name when
// neither ubus nor uci can resolve it, and nftValidDev happily accepts "lan"
// as a syntactically fine device name. The resulting `iifname "lan" ... drop`
// loads into the kernel without a murmur and matches no packet ever, so that
// network is neither diverted nor failed closed on. This is the failure mode
// nftValidDev's own doc comment warns about at length — but the warning it
// emits only fires for names it REJECTED, never for names it accepted that
// resolve to nothing.
//
// 2. A CLIENT NETWORK NOBODY NAMED. The divert set is built exclusively from
// what the operator listed: `config inbound` plus rules with `iface:`/`zone:`
// sources. The factory config ships ONE inbound, `lan`. Add a VLAN or a guest
// SSID in LuCI and its clients route to the internet through this router
// without ever entering the engine — no proxy, and no kill switch either.
// netplane.Interfaces() has been able to enumerate exactly this since the
// egress picker was built; the result went to the panel and nowhere else.
//
// # Why these WARN and do not refuse
//
// The rejected-name path refuses the whole apply when the kill switch is closed,
// and that is right there: a name we cannot express is a stable configuration
// error, refusing leaves the PREVIOUS ruleset loaded and protecting, and the
// operator has something concrete to fix.
//
// Neither of these has that shape.
//
// An unresolved name has no fail-closed action available at all — we do not know
// the device, so there is nothing to write a drop against. Refusing would abort
// the apply and leave the LAN on plain routing, i.e. it would produce the very
// leak it is objecting to, plus an outage. And it is not necessarily stable: at
// boot, netifd may not have answered yet.
//
// An uncovered network must not be closed silently for a much simpler reason:
// this router does not know whether the operator MEANT to proxy it. A guest SSID
// or an IoT VLAN deliberately kept off the tunnel is a legitimate configuration,
// and dropping its forwarded traffic because our inbound list does not mention it
// would cut a working network on a guess. So it is named, loudly, and the
// decision stays with the person who built the network.
//
// Both travel on the netplane warning channel, which apply grades CRITICAL
// wholesale — which is exactly the severity apply/warnings.go defines for "an
// interface the kill-switch does not cover", in those words.
//
// # Why the whole file is silent off a router
//
// Every check here is gated on Interfaces() returning something. With no ubus
// there is no inventory, and a coverage claim built on no inventory is a guess.
// That also keeps the checks out of the unit suite's hair: on a build machine the
// probe fails, the gate closes, and rendering behaves exactly as it did before.
import (
"fmt"
"strings"
"github.com/sagernet/sing-box/shater/model"
)
// coverageWarnings reports what this plan does not cover. refs is the divert set
// with provenance (nftDivertRefs); m supplies the egress list, which is the one
// thing that tells a client network apart from an uplink we deliberately send
// traffic OUT of.
//
// Returns nil when the router cannot be enumerated — see the package comment.
func coverageWarnings(m *model.Model, refs []divertRef) []string {
infos := Interfaces()
if len(infos) == 0 {
return nil
}
var out []string
out = append(out, unresolvedDeviceWarnings(refs, infos)...)
out = append(out, uncoveredNetworkWarnings(m, refs, infos)...)
return out
}
// unresolvedDeviceWarnings reports each diverted device that the kernel does not
// have.
//
// "Does not have" is answered from two independent sources, and a hit in either
// clears the device: the netifd inventory (which is how a normal interface is
// resolved) and /sys/class/net (which is how a raw device pulled straight out of
// a firewall zone's `list device`, or any device netifd does not manage, still
// counts). Only a device unknown to BOTH is reported — an over-eager warning here
// would be indistinguishable from the real one and would train the operator to
// ignore both.
func unresolvedDeviceWarnings(refs []divertRef, infos []IfaceInfo) []string {
known := make(map[string]bool, len(infos))
for _, i := range infos {
if i.Device != "" {
known[i.Device] = true
}
}
var out []string
for _, r := range refs {
if known[r.Dev] || DeviceExists(r.Dev) {
continue
}
out = append(out, fmt.Sprintf(
"interface %q: this router has no network device called %q, and that is the name the "+
"ruleset was written against — every rule scoped to it (the TPROXY divert, the "+
"fail-closed drop, the accept_local ingress knob) loads into the kernel cleanly and "+
"matches NOTHING. The traffic of this network is therefore neither proxied nor "+
"blocked when the engine is down: it goes straight out the WAN in the clear. It "+
"happens when a `config inbound`'s network (or a rule's iface:/zone: source) names "+
"something netifd does not know — a typo, a deleted interface, or a device name used "+
"where a UCI interface name belongs. Point it at an existing interface, or remove it.",
r.Src, r.Dev))
}
return out
}
// uncoveredNetworkWarnings reports every client network of this router whose
// traffic reaches the internet without passing through the plan.
//
// # What counts as a client network
//
// Not "has a private address" — that catches tunnels and management links. The
// test is the firewall's own: the interface sits in a zone that is FORWARDED TO A
// WAN ZONE. That is precisely the statement "the devices on this network use this
// router to reach the internet", written by the operator, in the file that
// decides it. A zone with no such forwarding (a management VLAN, an egress tunnel
// zone) reaches nothing through us and is not this check's business.
//
// On top of that, four exclusions, each removing a class of false positive:
//
// - the WAN zones themselves (traffic arriving there is not a client's);
// - interfaces already in the divert set (the whole point);
// - interfaces this model uses as an EGRESS — we send traffic out of those on
// purpose, and their zone may well forward to wan;
// - interfaces netifd reports as down, because they carry no clients right now
// and the ifup hotplug reconciles the moment they do.
func uncoveredNetworkWarnings(m *model.Model, refs []divertRef, infos []IfaceInfo) []string {
forwarded := internetForwardedZones()
if len(forwarded) == 0 {
// Either there is no firewall config we can read, or nothing forwards to a
// WAN zone at all. Both mean we cannot say a network reaches the internet
// through us, and this check refuses to guess.
return nil
}
covered := make(map[string]bool, len(refs))
for _, r := range refs {
covered[r.Dev] = true
}
egress := map[string]bool{}
for _, eg := range m.Egresses {
if n := strings.TrimSpace(eg.Interface); n != "" {
egress[n] = true
}
}
var out []string
for _, i := range infos {
zone := strings.ToLower(strings.TrimSpace(i.Zone))
switch {
case covered[i.Device], !i.Up, zone == "", model.WANZoneNames[zone],
!forwarded[zone], egress[i.Name]:
continue
}
where := i.Device
if i.Subnet != "" {
where = fmt.Sprintf("%s, %s", i.Device, i.Subnet)
}
out = append(out, fmt.Sprintf(
"interface %q: this is a client network of this router (%s; firewall zone %q, which is "+
"forwarded to the internet) and NOTHING in this configuration touches it. Its devices "+
"are not proxied, their DNS is not filtered, and the kill switch does not cover them "+
"— when the engine is down their traffic keeps flowing to the WAN in the clear, "+
"because the fail-closed drop is scoped to the interfaces named in the config and "+
"this is not one of them. This is not closed automatically: doing so would cut a "+
"network you may have meant to leave off the tunnel. Decide explicitly — add a "+
"`config inbound` for it (or a rule with source iface:%s, which may still target "+
"`direct`) to bring it under the plane, or move it into a firewall zone that is not "+
"forwarded to the WAN.",
i.Name, where, i.Zone, i.Name))
}
return out
}
// internetForwardedZones returns the set of firewall zone names (lower-cased)
// that are forwarded to a WAN zone, i.e. the zones whose members use this router
// to reach the internet.
//
// It reads the SAME `uci -q export firewall` the zone map does — one more read of
// a plane we already talk to, rather than a second notion of which zone is which.
// Fail-open: an unreadable or zone-less firewall yields an empty set, and the
// caller then reports nothing.
func internetForwardedZones() map[string]bool {
out, err := execCommand("uci", "-q", "export", "firewall").Output()
if err != nil {
return nil
}
return parseInternetForwardedZones(string(out))
}
// parseInternetForwardedZones is internetForwardedZones' pure half.
//
// A zone is a WAN zone when its NAME is one of the conventional uplink names
// (model.WANZoneNames — the same list rule-source validation uses). A `config
// forwarding` whose dest is such a zone marks its src zone as internet-forwarded.
func parseInternetForwardedZones(exportText string) map[string]bool {
secs := scanUCISections(exportText)
wan := map[string]bool{}
for _, s := range secs {
if s.typ != "zone" {
continue
}
if n := strings.ToLower(strings.TrimSpace(s.options["name"])); n != "" && model.WANZoneNames[n] {
wan[n] = true
}
}
if len(wan) == 0 {
return nil
}
fwd := map[string]bool{}
for _, s := range secs {
if s.typ != "forwarding" {
continue
}
src := strings.ToLower(strings.TrimSpace(s.options["src"]))
dest := strings.ToLower(strings.TrimSpace(s.options["dest"]))
if src != "" && wan[dest] {
fwd[src] = true
}
}
return fwd
}
+283
View File
@@ -0,0 +1,283 @@
package netplane
// W4 + the adjacent defect: what the plan does NOT cover has to reach the
// operator, not just the panel's interface picker.
//
// Both tests drive the real RenderNftWithWarningsAt, because the warning channel
// it returns is the one apply folds into Status — a check that only exercised the
// helper would prove nothing about whether the news ever leaves netplane.
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/shater/model"
)
// A three-network router of exactly the shape this defect was found on: two LAN
// bridges in the `lan` firewall zone, one uplink, and a forwarding that says the
// lan zone reaches the internet through this box.
const coverageFirewallExport = `package firewall
config defaults
option input 'ACCEPT'
config zone
option name 'lan'
list network 'lan'
list network 'sex'
option input 'ACCEPT'
config zone
option name 'wan'
list network 'wan0'
option masq '1'
config forwarding
option src 'lan'
option dest 'wan'
`
const coverageIfaceDump = `{"interface":[
{"interface":"loopback","up":true,"l3_device":"lo"},
{"interface":"lan","up":true,"l3_device":"br-lan","ipv4-address":[{"address":"192.168.1.1","mask":24}]},
{"interface":"sex","up":true,"l3_device":"br-sex","ipv4-address":[{"address":"10.67.0.1","mask":24}]},
{"interface":"wan0","up":true,"l3_device":"eth1","ipv4-address":[{"address":"10.0.2.15","mask":24}]}
]}`
// coverageExec answers the three probes the coverage checks make (the interface
// dump, the firewall export, and IfaceDevice's per-interface status) and returns
// empty output for everything else, which is what every other netplane test sees.
func coverageExec(status map[string]string) func(string, ...string) *exec.Cmd {
return func(name string, arg ...string) *exec.Cmd {
joined := strings.Join(append([]string{name}, arg...), " ")
out := ""
switch {
case joined == "ubus call network.interface dump":
out = coverageIfaceDump
case joined == "uci -q export firewall":
out = coverageFirewallExport
case strings.HasPrefix(joined, "ubus call network.interface.") &&
strings.HasSuffix(joined, " status"):
iface := strings.TrimSuffix(
strings.TrimPrefix(joined, "ubus call network.interface."), " status")
out = status[iface]
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
return cmd
}
}
// fakeSysClassNet points DeviceExists at a directory holding exactly the devices
// the kernel is pretending to have.
func fakeSysClassNet(t *testing.T, devs ...string) {
t.Helper()
dir := t.TempDir()
for _, d := range devs {
if err := os.MkdirAll(filepath.Join(dir, d), 0o755); err != nil {
t.Fatalf("fixture %s: %v", d, err)
}
}
orig := sysClassNet
sysClassNet = dir
t.Cleanup(func() { sysClassNet = orig })
}
func coverageModel() *model.Model {
return &model.Model{
Globals: model.Globals{
Enabled: true, KillSwitch: "closed", FwmarkBase: 0x2000, TableBase: 0x2000,
},
Inbounds: []model.Inbound{{
Name: "lan", Enabled: true, Type: "tproxy", Network: "lan",
TproxyPort: 12345, TCP: true, UDP: true,
}},
}
}
func warningsAbout(ws []string, iface string) []string {
var out []string
for _, w := range ws {
if strings.HasPrefix(w, "interface \""+iface+"\":") {
out = append(out, w)
}
}
return out
}
// TestRenderNftWarnsUncoveredLANNetwork is W4. `sex` is a client network of this
// router — it sits in the `lan` firewall zone, which is forwarded to `wan`, so its
// devices reach the internet through this box — and NOTHING in the config mentions
// it. Its traffic is neither proxied nor covered by the fail-closed drop, and
// before this the only place that was visible was the panel's interface picker.
//
// RED BEFORE: renderNft produced warnings only for device names it had REJECTED,
// so this returned an empty warning list.
func TestRenderNftWarnsUncoveredLANNetwork(t *testing.T) {
orig := execCommand
defer func() { execCommand = orig }()
execCommand = coverageExec(map[string]string{
"lan": `{"l3_device":"br-lan"}`,
"sex": `{"l3_device":"br-sex"}`,
"wan0": `{"l3_device":"eth1"}`,
})
fakeSysClassNet(t, "br-lan", "br-sex", "eth1", "lo")
rs, warnings, err := RenderNftWithWarningsAt(coverageModel(), time.Now())
if err != nil {
t.Fatalf("RenderNftWithWarningsAt: %v", err)
}
if !strings.Contains(rs, `iifname "br-lan"`) {
t.Fatalf("the plan should still divert br-lan; ruleset:\n%s", rs)
}
got := warningsAbout(warnings, "sex")
if len(got) != 1 {
t.Fatalf("want exactly one warning about the uncovered network \"sex\", got %d: %v",
len(got), warnings)
}
for _, must := range []string{"client network", "kill switch does not cover", "iface:sex"} {
if !strings.Contains(got[0], must) {
t.Errorf("the uncovered-network warning must say %q; got:\n%s", must, got[0])
}
}
// The covered network and the uplink must NOT be reported: a check that also
// fires on the interfaces it is fine with teaches the operator to ignore it.
if w := warningsAbout(warnings, "lan"); len(w) != 0 {
t.Errorf("the DIVERTED network must not be reported as uncovered: %v", w)
}
if w := warningsAbout(warnings, "wan0"); len(w) != 0 {
t.Errorf("the WAN uplink is not a client network and must not be reported: %v", w)
}
}
// TestRenderNftWarnsUnresolvedDivertDevice is the adjacent defect. The inbound
// names an interface that does not exist, IfaceDevice hands the raw UCI name back
// (its documented last resort), nftValidDev accepts it as a syntactically fine
// device name — and the ruleset loads with `iifname "lann"` matching nothing at
// all. The divert does nothing, the fail-closed drop does nothing, and the apply
// reports success.
//
// RED BEFORE: nftValidDev checks characters, not existence, so nothing anywhere
// noticed. The warning must name what the operator wrote AND what it resolved to.
func TestRenderNftWarnsUnresolvedDivertDevice(t *testing.T) {
orig := execCommand
defer func() { execCommand = orig }()
// No status entry for "lann": ubus knows nothing about it, and neither does
// `uci get network.lann.device` — the fallback-to-raw-name path.
execCommand = coverageExec(map[string]string{
"lan": `{"l3_device":"br-lan"}`,
"sex": `{"l3_device":"br-sex"}`,
"wan0": `{"l3_device":"eth1"}`,
})
fakeSysClassNet(t, "br-lan", "br-sex", "eth1", "lo")
m := coverageModel()
m.Inbounds[0].Network = "lann"
rs, warnings, err := RenderNftWithWarningsAt(m, time.Now())
if err != nil {
t.Fatalf("RenderNftWithWarningsAt: %v", err)
}
// The point of the defect: this loads cleanly. The test pins that the plan is
// still produced, so the only thing standing between the operator and a silent
// leak is the warning.
if !strings.Contains(rs, `iifname "lann"`) {
t.Fatalf("expected the (useless) rule to still be rendered; ruleset:\n%s", rs)
}
got := warningsAbout(warnings, "lann")
if len(got) != 1 {
t.Fatalf("want exactly one warning about the unresolved device, got %d: %v",
len(got), warnings)
}
for _, must := range []string{"no network device called", `"lann"`, "matches NOTHING"} {
if !strings.Contains(got[0], must) {
t.Errorf("the unresolved-device warning must say %q; got:\n%s", must, got[0])
}
}
}
// TestCoverageSilentWithoutInventory is the fail-open half of the contract: with
// no ubus to enumerate the router (a build host, a stripped image, a boot before
// netifd answers) a coverage claim would be a guess, so nothing is said at all.
func TestCoverageSilentWithoutInventory(t *testing.T) {
orig := execCommand
defer func() { execCommand = orig }()
var rec []string
execCommand = fakeExec(&rec) // every probe succeeds with EMPTY output
fakeSysClassNet(t) // and no device exists either
_, warnings, err := RenderNftWithWarningsAt(coverageModel(), time.Now())
if err != nil {
t.Fatalf("RenderNftWithWarningsAt: %v", err)
}
if len(warnings) != 0 {
t.Fatalf("with no interface inventory the coverage checks must stay silent, got: %v", warnings)
}
}
// TestParseInternetForwardedZones pins what "a client network" means: a zone the
// operator forwards to the uplink, read out of the firewall config rather than
// guessed from addresses. A vpn zone that is not forwarded to wan is NOT one.
func TestParseInternetForwardedZones(t *testing.T) {
const export = `package firewall
config zone
option name 'lan'
list network 'lan'
config zone
option name 'guest'
list network 'guest'
config zone
option name 'mgmt'
list network 'mgmt'
config zone
option name 'vpn'
list network 'awg0'
config zone
option name 'wan'
list network 'wan0'
config forwarding
option src 'lan'
option dest 'wan'
config forwarding
option src 'guest'
option dest 'wan'
config forwarding
option src 'lan'
option dest 'vpn'
`
got := parseInternetForwardedZones(export)
for _, want := range []string{"lan", "guest"} {
if !got[want] {
t.Errorf("zone %q forwards to wan and must count as a client zone; got %v", want, got)
}
}
for _, notWant := range []string{"mgmt", "vpn", "wan"} {
if got[notWant] {
t.Errorf("zone %q does not forward to the uplink and must NOT count; got %v", notWant, got)
}
}
// No WAN zone at all => nothing can be said about reaching the internet.
if m := parseInternetForwardedZones("package firewall\n\nconfig zone\n\toption name 'lan'\n"); len(m) != 0 {
t.Errorf("with no wan zone the set must be empty, got %v", m)
}
if m := parseInternetForwardedZones(""); len(m) != 0 {
t.Errorf("empty input must yield an empty set, got %v", m)
}
}
+5 -2
View File
@@ -202,8 +202,11 @@ func TestFwmarkRulePresentIsExact(t *testing.T) {
// over the ordinary WAN with the router's real address.
func TestEgressRouteAddFailureIsReported(t *testing.T) {
f := wgEgressNet()
// Only the REAL default route fails; `route add unreachable default` (the
// table's fail-closed floor) is left working, so the single warning asserted
// below is the one this test is about.
f.failIf = func(name string, arg []string) bool {
return name == "ip" && len(arg) >= 3 && arg[1] == "route" && arg[2] == "add"
return name == "ip" && len(arg) >= 4 && arg[1] == "route" && arg[2] == "add" && arg[3] == "default"
}
f.install(t)
@@ -269,7 +272,7 @@ func TestEgressRuleAddFailureIsReported(t *testing.T) {
func TestFailedEgressKeepsRoutingAbsent(t *testing.T) {
f := wgEgressNet()
f.failIf = func(name string, arg []string) bool {
return name == "ip" && len(arg) >= 3 && arg[1] == "route" && arg[2] == "add"
return name == "ip" && len(arg) >= 4 && arg[1] == "route" && arg[2] == "add" && arg[3] == "default"
}
f.install(t)
if _, err := addEgressRouting(wgEgressModel()); err != nil {
+602
View File
@@ -0,0 +1,602 @@
package netplane
// The mark plane: every fwmark this ruleset STAMPS must be routed somewhere, and
// every place that trusts a mark must be unable to trust it once the routing
// stopped happening.
//
// All four defects below share one shape. `ip rule fwmark X lookup N` does not
// deliver a packet to table N — it delivers the LOOKUP to table N, and a lookup
// that finds nothing there continues down the rule list to `main`. So an empty
// table N does not mean "this traffic is stuck", it means "this traffic is
// routed as if it had never been marked": out the default WAN, with the router's
// real address, still carrying a mark the firewall was told to trust. Tables go
// empty for entirely ordinary reasons and without anything we render changing —
// an ifdown, an engine restart destroying its TUN — so neither the nft
// idempotence check nor the operator has any reason to notice.
import (
"fmt"
"regexp"
"sort"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// --- defect 1: a dead egress leaked non-TCP/UDP past a CLOSED kill-switch ----
// TestUntunnelableEgressAcceptIsBoundToItsDevice is the security regression.
//
// SCENARIO (reproduced by the auditor against a live kernel, then reduced to
// this): untunnelable_egress='wwan', untunnelable='block', kill_switch='closed'.
// The interface goes down. The kernel deletes every route through its device,
// including `default dev wwan0 table 8208`. Nothing we render changes, so no
// apply runs. nft keeps stamping 0x2100 on every non-TCP/UDP packet from the
// LAN; `ip rule fwmark 0x2100 lookup 8208` now finds an empty table and falls
// through to main; the packet leaves out the default WAN — and the forward chain
// waved it there itself, because its accept looked only at the mark, and that
// accept sits ABOVE the fail-closed drop.
//
// The fix is the conjunction: the mark says where the packet was SENT, oifname
// says where it is actually GOING, and only both together mean the routing did
// what the mark asked. `meta mark X oifname "dev"` is strictly narrower than the
// mark alone, so it can only ever remove leaks, never add one.
//
// (The comment this replaced argued, correctly, that an oifname accept ALONE
// would be too loose — it would bless whatever the routing table pushed at that
// device, from wherever. That is why the mark stays. It is not an argument for
// dropping oifname, which is the conclusion that was drawn from it.)
func TestUntunnelableEgressAcceptIsBoundToItsDevice(t *testing.T) {
m := realisticModel() // kill_switch=closed, egress "wwan" on wwan0 at idx 0
m.Globals.UntunnelableEgress = "wwan"
m.Globals.Untunnelable = "block"
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
// Precondition: prerouting really is stamping the mark, or the rest is moot.
if pre := preroutingChain(t, rs); !strings.Contains(pre, "meta mark set 0x2100 accept") {
t.Fatalf("prerouting no longer stamps the egress mark; this test is testing nothing:\n%s", pre)
}
fwd := forwardChain(t, rs)
for _, line := range strings.Split(fwd, "\n") {
line = strings.TrimSpace(line)
if !strings.Contains(line, "meta mark 0x2100") || !strings.HasSuffix(line, "accept") {
continue
}
if !strings.Contains(line, `oifname "wwan0"`) {
t.Errorf("the forward chain accepts the untunnelable-egress mark on the mark ALONE:\n"+
" %s\n"+
"An ifdown empties table 8208, the fwmark lookup falls through to main, and this line "+
"then accepts the packet on its way out the DEFAULT WAN — above the fail-closed drop, "+
"with the kill-switch closed. Bind the accept to the egress device as well: "+
"`meta mark 0x2100 oifname \"wwan0\" accept`.", line)
}
}
// And the accept must still exist and still precede the drop, or the feature
// is dead instead of leaky.
acc := strings.Index(fwd, `meta mark 0x2100 oifname "wwan0" accept`)
drop := strings.Index(fwd, "meta nfproto ipv4 drop")
if acc < 0 {
t.Fatalf("the device-bound egress accept is missing entirely — marked ESP/GRE is now dropped "+
"by the kill-switch on its way out the egress:\n%s", fwd)
}
if drop >= 0 && acc > drop {
t.Errorf("the egress accept must precede the fail-closed drop or it never matches:\n%s", fwd)
}
}
// TestEgressTableGetsFailClosedFloor is the other half of the same defect, and
// the half that survives an OPEN kill-switch (where there is no forward drop to
// be saved by).
//
// The nft conjunction above stops the leaked packet from being ACCEPTED; this
// stops it from being ROUTED at all. An `unreachable default` at the maximum
// metric loses to any real default route while one exists, has no device so the
// kernel never garbage-collects it, and turns "lookup failed, try main" into
// "lookup succeeded: unreachable". The fallthrough stops being a matter of
// noticing a warning in time and becomes structurally impossible.
func TestEgressTableGetsFailClosedFloor(t *testing.T) {
f := wgEgressNet()
f.install(t)
if _, err := addEgressRouting(wgEgressModel()); err != nil {
t.Fatalf("addEgressRouting: %v", err)
}
var floor, real string
for _, c := range f.calls {
if strings.Contains(c, "route add unreachable default") && strings.HasSuffix(c, "table 8208") {
floor = c
}
if strings.Contains(c, "route add default") && strings.HasSuffix(c, "table 8208") {
real = c
}
}
if floor == "" {
t.Fatalf("no `ip route add unreachable default ... table 8208` was issued.\n"+
"Without it, an ifdown empties the egress table, the fwmark lookup falls through to the "+
"MAIN table, and everything bound to this egress leaves over the plain WAN with the "+
"router's real address. Calls were:\n %s", strings.Join(f.calls, "\n "))
}
if !strings.Contains(floor, "metric "+unreachableFloorMetric) {
t.Errorf("the floor must carry the maximal metric or it can outrank the real default route: %q", floor)
}
if real == "" {
t.Fatalf("the real default route disappeared along with the fix:\n %s", strings.Join(f.calls, "\n "))
}
}
// TestFloorDoesNotMaskAnEmptyTable: the floor keeps the packet from leaking, and
// must not therefore keep the RECONCILE from repairing the egress. A table
// holding only `unreachable default` is the exact state an ifdown now leaves
// behind, i.e. the state the presence check exists to catch — a substring test
// for "default" would read it as healthy and turn the safety net into a
// blindfold.
func TestFloorDoesNotMaskAnEmptyTable(t *testing.T) {
const floorOnly = "unreachable default metric 4294967295"
f := wgEgressNet()
f.install(t)
if _, err := addEgressRouting(wgEgressModel()); err != nil {
t.Fatalf("addEgressRouting: %v", err)
}
f.rules = map[string]string{"-4": "32765:\tfrom all fwmark 0x2000 lookup 8192\n" +
"32764:\tfrom all fwmark 0x2100 lookup 8208"}
f.routes = map[string]string{
"-4 8192": "local default dev lo scope host",
"-4 8208": floorOnly,
}
if RoutingPresent(egressPlaneGlobals()) {
t.Fatal("a table holding ONLY the unreachable floor is an egress whose real route is gone; " +
"reporting it present makes the reconcile skip ApplyRouting and the egress never comes back")
}
// Sanity: the same table WITH a real route is present again.
f.routes["-4 8208"] = "default dev wg0 scope link\n" + floorOnly
if !RoutingPresent(egressPlaneGlobals()) {
t.Fatal("a healthy egress table must still report present with the floor installed")
}
}
// --- defect 2: RoutingPresent had never heard of the L3 ingress -------------
// l3Globals pins the marks/tables the L3 cases hard-code: mark 0x2080, table 8200.
func l3Globals() model.Globals {
return model.Globals{FwmarkBase: 0x2000, TableBase: 0x2000, L3Tunnel: true}
}
// TestRoutingPresentSeesL3Table.
//
// addL3Routing installs a THIRD `fwmark -> table` + default-route pair and
// RoutingPresent was never told about it — it checked the main tproxy pair and
// the per-egress bindings and returned. This is the same defect that had already
// been found twice (main table, then egress tables) and whose own comment in
// apply.go says "a presence check must cover everything its Apply counterpart
// installs, or the idempotent fast-path becomes a trap".
//
// Its trigger does not even need an interface to go down. Edit a node's URI: the
// engine is recreated, the kernel destroys the shater-l3 device and with it
// `default dev shater-l3 table 8200`, and the new device arrives without it. The
// rendered nft text is unchanged, so nftCurrent stays true; RoutingPresent said
// true on the strength of main+egress; applyLocked therefore skipped
// ApplyRoutingWithWarnings and NOTHING ever reinstalled the route. LAN ping was
// dead until somebody restarted the daemon, permanently, at plane=full, with no
// warning — and with untunnelable=direct or an open kill-switch it was not dead
// but leaking.
func TestRoutingPresentSeesL3Table(t *testing.T) {
const (
mainRule = "32765:\tfrom all fwmark 0x2000 lookup 8192"
l3Rule = "32763:\tfrom all fwmark 0x2080 lookup 8200"
mainRoute = "local default dev lo scope host"
l3Route = "default dev shater-l3 scope link"
floor = "unreachable default metric 4294967295"
)
cases := []struct {
name string
rules string
l3Table string
want bool
}{
{"complete", mainRule + "\n" + l3Rule, l3Route + "\n" + floor, true},
// THE REGRESSION: the engine restarted, the TUN was destroyed and
// recreated, the kernel took the route with it. The rule survived.
{"l3 table emptied by an engine restart", mainRule + "\n" + l3Rule, floor, false},
// The other half: the rule itself went away (a `network reload` flushes
// `ip rule` wholesale).
{"l3 rule gone", mainRule, l3Route, false},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
f := &egressPlaneNet{
rules: map[string]string{"-4": c.rules},
routes: map[string]string{"-4 8192": mainRoute, "-4 8200": c.l3Table},
}
f.install(t)
if got := RoutingPresent(l3Globals()); got != c.want {
t.Fatalf("RoutingPresent = %v, want %v.\n"+
"A presence check blind to the L3 pair lets the fast-path skip ApplyRouting forever "+
"while LAN ICMP is marked for a table that no longer routes it.", got, c.want)
}
})
}
}
// With l3_tunnel OFF nothing is installed by design, so its absence must not be
// read as a broken plane — the L3 half may only ever ADD a reason to re-apply.
func TestRoutingPresentUnaffectedWithoutL3(t *testing.T) {
f := &egressPlaneNet{
rules: map[string]string{"-4": "32765:\tfrom all fwmark 0x2000 lookup 8192"},
routes: map[string]string{"-4 8192": "local default dev lo scope host"},
}
f.install(t)
g := l3Globals()
g.L3Tunnel = false
if !RoutingPresent(g) {
t.Fatal("a healthy main pair with l3_tunnel off must report present")
}
}
// Both families, for the reason the main pair checks both: `ip -6 rule` is
// flushed independently of `ip -4 rule`, so a v6 half can go missing on its own
// and a v4-only check would call the plane intact.
func TestRoutingPresentChecksL3PerFamily(t *testing.T) {
f := &egressPlaneNet{
rules: map[string]string{
"-4": "32765:\tfrom all fwmark 0x2000 lookup 8192\n32763:\tfrom all fwmark 0x2080 lookup 8200",
"-6": "32765:\tfrom all fwmark 0x2000 lookup 8192", // the v6 L3 rule is gone
},
routes: map[string]string{
"-4 8192": "local default dev lo scope host",
"-6 8192": "local default dev lo metric 1024 pref medium",
"-4 8200": "default dev shater-l3 scope link",
"-6 8200": "default dev shater-l3 metric 1024 pref medium",
},
}
f.install(t)
g := l3Globals()
g.IPv6 = true
if RoutingPresent(g) {
t.Fatal("the v6 L3 rule is missing; reporting present leaves v6 ICMP marked with nowhere to go")
}
}
// The L3 table gets the same fail-closed floor as the egress tables — the engine
// restarting is a far more frequent event than an interface going down.
func TestL3TableGetsFailClosedFloor(t *testing.T) {
f := &egressPlaneNet{}
f.install(t)
warns, err := addL3Routing(&model.Model{Globals: l3Globals()})
if err != nil {
t.Fatalf("addL3Routing: %v", err)
}
if len(warns) != 0 {
t.Fatalf("a healthy install must warn about nothing: %v", warns)
}
var floor string
for _, c := range f.calls {
if strings.Contains(c, "route add unreachable default") && strings.HasSuffix(c, "table 8200") {
floor = c
}
}
if floor == "" {
t.Fatalf("no fail-closed floor in the L3 table: when the engine recreates its TUN the table is "+
"empty, the fwmark lookup falls through to main, and marked ICMP is dropped by a closed "+
"kill-switch or leaks past an open one. Calls were:\n %s", strings.Join(f.calls, "\n "))
}
}
// --- defect 3: the validator and the data plane disagreed about a name ------
// TestUntunnelableEgressResolutionLockstep.
//
// model.ValidateUntunnelableEgress duplicates netplane.UntunnelableEgressBinding
// because the import can only run one way (netplane imports model). The
// duplication is unavoidable; the DIVERGENCE was not, and a comment saying "MUST
// change in lockstep" had already failed to prevent one: the validator rejected
// an egress on `strings.TrimSpace(Interface) == ""` and the binding on
// `Interface == ""`, so `option interface ' '` produced a panel saying "the
// option is ignored" over a data plane that was stamping the mark — straight
// into defect 1 above, with the operator explicitly told it could not happen.
//
// One table, both functions, one verdict. A future divergence is a test failure
// rather than a warning that lies exactly where it matters most.
func TestUntunnelableEgressResolutionLockstep(t *testing.T) {
cases := []struct {
name string
option string
egresses []model.Egress
resolves bool
}{
{"empty option", "", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, false},
{"blank option", " ", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, false},
{"plain match", "wan2", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, true},
{"tunnel type", "vpn", []model.Egress{{Name: "vpn", Type: "tunnel", Interface: "wg0"}}, true},
{"case-insensitive name", "WAN2", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, true},
{"padded option", " wan2 ", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, true},
{"padded egress name", "wan2", []model.Egress{{Name: " wan2 ", Type: "interface", Interface: "eth1"}}, true},
{"padded type", "wan2", []model.Egress{{Name: "wan2", Type: " Interface ", Interface: "eth1"}}, true},
{"no such egress", "wan9", []model.Egress{{Name: "wan2", Type: "interface", Interface: "eth1"}}, false},
{"wrong type (direct)", "d", []model.Egress{{Name: "d", Type: "direct"}}, false},
{"wrong type (byedpi)", "d", []model.Egress{{Name: "d", Type: "byedpi"}}, false},
{"no interface", "hole", []model.Egress{{Name: "hole", Type: "interface"}}, false},
// THE DIVERGENCE. IfaceDevice(" ") hands back " ", which is neither
// empty nor a device: the binding used to say ok, the validator used to
// say "ignored", and the packets were marked for a table nobody built.
{"whitespace interface", "hole", []model.Egress{{Name: "hole", Type: "interface", Interface: " "}}, false},
{"whitespace interface, tunnel", "hole", []model.Egress{{Name: "hole", Type: "tunnel", Interface: "\t"}}, false},
// A padded but real interface name is a name, not an absence.
{"padded interface", "wan2", []model.Egress{{Name: "wan2", Type: "interface", Interface: " eth1 "}}, true},
// First name match decides, in both.
{"second egress matches", "vpn", []model.Egress{
{Name: "wan2", Type: "interface", Interface: "eth1"},
{Name: "vpn", Type: "tunnel", Interface: "wg0"},
}, true},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
g := model.Globals{FwmarkBase: 0x2000, TableBase: 0x2000, UntunnelableEgress: c.option}
_, _, dev, ok := UntunnelableEgressBinding(&model.Model{Globals: g, Egresses: c.egresses})
warns := model.ValidateUntunnelableEgress(g, c.egresses)
// The validator is silent when the option is empty (nothing was
// asked for) and when it resolves; it warns when it was asked for
// and cannot be honoured.
validatorAccepts := len(warns) == 0
wantValidatorAccepts := c.resolves || strings.TrimSpace(c.option) == ""
if ok != c.resolves {
t.Errorf("netplane.UntunnelableEgressBinding ok = %v (dev %q), want %v", ok, dev, c.resolves)
}
if validatorAccepts != wantValidatorAccepts {
t.Errorf("model.ValidateUntunnelableEgress silent = %v, want %v: %v",
validatorAccepts, wantValidatorAccepts, warns)
}
// The lockstep assertion proper: for a non-empty option the two must
// agree. This is the one that fires when they drift.
if strings.TrimSpace(c.option) != "" && ok != validatorAccepts {
t.Errorf("VERDICTS DIVERGE for option %q: the data plane %s the egress while the "+
"validator %s it. Whichever way round, the panel is telling the operator the "+
"opposite of what the packets do — and the marked-but-unrouted direction is a "+
"kill-switch bypass (see TestUntunnelableEgressAcceptIsBoundToItsDevice).",
c.option, resolvedWord(ok), resolvedWord(validatorAccepts))
}
if ok && strings.TrimSpace(dev) == "" {
t.Errorf("a resolved binding must yield a usable device, got %q", dev)
}
})
}
}
func resolvedWord(ok bool) string {
if ok {
return "ACCEPTS"
}
return "REJECTS"
}
// --- the general rule the three defects above are instances of --------------
var stampedMarkRe = regexp.MustCompile(`meta mark set 0x([0-9a-fA-F]+)`)
var ruleAddRe = regexp.MustCompile(`^ip (-[46]) rule add fwmark 0x([0-9a-fA-F]+) lookup (\d+)$`)
// TestEveryStampedMarkIsRoutedAndVerified is the checklist, executable.
//
// Both blocking defects were the same omission twice: a new mark was added to
// prerouting and one of the three things a mark needs was not added with it.
// Rather than trusting the next author to remember, this test derives the marks
// from the RENDERED RULESET and demands all three of each:
//
// (a) it is ROUTED — ApplyRouting issues an `ip rule fwmark <M> lookup <T>`;
// (b) its table has a FAIL-CLOSED FLOOR, so an emptied table cannot fall
// through to main;
// (c) the forward chain, if it trusts the mark at all, trusts it only in
// conjunction with an oifname;
// (d) RoutingPresent NOTICES when that table loses its real route, so the
// reconcile repairs it instead of skipping forever.
//
// A mark added to prerouting without any of these fails here without anyone
// having to think of writing a test for it.
func TestEveryStampedMarkIsRoutedAndVerified(t *testing.T) {
// Everything on at once, so every mark this plane can stamp is stamped.
m := realisticModel()
m.Globals.IPv6 = true
m.Globals.L3Tunnel = true
m.Globals.DNSIntercept = true
m.Globals.UntunnelableEgress = "wwan"
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre, fwd := preroutingChain(t, rs), forwardChain(t, rs)
stamped := map[uint64]bool{}
for _, mt := range stampedMarkRe.FindAllStringSubmatch(pre, -1) {
var v uint64
if _, err := fmt.Sscanf(mt[1], "%x", &v); err == nil {
stamped[v] = true
}
}
// The divert mark is the one documented exception: its "routing" is the main
// tproxy pair (mark -> table -> `local default dev lo`), which RoutingPresent
// checks separately and which cannot go stale the way a device-backed table
// can, because `lo` never goes down. It is also never accepted in forward.
// The exception is verified, not asserted — see the end of this test.
divert := uint64(effFwmark(m.Globals))
delete(stamped, divert)
if len(stamped) < 2 {
t.Fatalf("expected at least the L3 and untunnelable-egress marks to be stamped, found %v.\n"+
"If the marks moved, move this test with them; if they stopped being stamped, this "+
"whole file is testing nothing.", stamped)
}
f := &egressPlaneNet{}
f.install(t)
if _, err := ApplyRoutingWithWarnings(m); err != nil {
t.Fatalf("ApplyRoutingWithWarnings: %v", err)
}
// (a) every stamped mark is bound to a table, per family.
routed := map[uint64][]markBindingView{}
for _, c := range f.calls {
mt := ruleAddRe.FindStringSubmatch(c)
if mt == nil {
continue
}
var v uint64
if _, err := fmt.Sscanf(mt[2], "%x", &v); err != nil {
continue
}
routed[v] = append(routed[v], markBindingView{mt[1], mt[3]})
}
for _, mark := range sortedMarks(stamped) {
bs := routed[mark]
if len(bs) == 0 {
t.Errorf("(a) mark 0x%x is STAMPED in prerouting but no `ip rule fwmark 0x%x lookup ...` is "+
"ever installed. A mark with no rule is not a route, it is a main-table fallthrough: "+
"the packet leaves out the default WAN carrying a mark the firewall trusts.", mark, mark)
continue
}
for _, b := range bs {
// (b) the table it points at has a fail-closed floor.
want := fmt.Sprintf("ip %s route add unreachable default metric %s table %s",
b.fam, unreachableFloorMetric, b.table)
if !containsCall(f.calls, want) {
t.Errorf("(b) table %s (mark 0x%x, %s) gets no `unreachable default` floor. When the "+
"device behind it disappears the kernel empties the table, the fwmark lookup "+
"fails over to main, and the traffic leaves over the plain WAN. Expected call: %q",
b.table, mark, b.fam, want)
}
}
// (c) forward may trust the mark only together with an oifname.
for _, line := range strings.Split(fwd, "\n") {
line = strings.TrimSpace(line)
if !strings.Contains(line, fmt.Sprintf("meta mark 0x%x", mark)) || !strings.HasSuffix(line, "accept") {
continue
}
if !strings.Contains(line, "oifname") {
t.Errorf("(c) the forward chain accepts mark 0x%x on the mark alone:\n %s\n"+
"The mark says where the packet was sent; only oifname says where it went. "+
"Once the mark's table is empty those differ, and this line waves the packet "+
"out the default WAN above the fail-closed drop.", mark, line)
}
}
}
// (d) RoutingPresent must notice each table losing its real route.
healthy := healthyPlaneView(routed, uint64(effTable(m.Globals)))
f.rules, f.routes = healthy.rules, healthy.routes
if !RoutingPresent(m.Globals) {
t.Fatalf("the synthetic healthy plane must report present, or every case below is vacuous.\n"+
"rules: %v\nroutes: %v", f.rules, f.routes)
}
for _, mark := range sortedMarks(stamped) {
for _, b := range routed[mark] {
key := b.fam + " " + b.table
saved := f.routes[key]
f.routes[key] = "unreachable default metric " + unreachableFloorMetric // real route gone
if RoutingPresent(m.Globals) {
t.Errorf("(d) table %s (mark 0x%x, %s) lost its real route and RoutingPresent still says "+
"present. The fast-path then skips ApplyRouting forever and nothing ever "+
"reinstalls it — the failure is permanent and a reconcile cannot repair it.",
b.table, mark, b.fam)
}
f.routes[key] = saved
}
}
// The exception is checked too: the main tproxy table is covered by
// mainRoutingPresentFam, so "skip the divert mark" is a delegation and not a
// hole in the sweep above.
mainKey := fmt.Sprintf("-4 %d", effTable(m.Globals))
saved := f.routes[mainKey]
f.routes[mainKey] = ""
if RoutingPresent(m.Globals) {
t.Errorf("the main tproxy table (mark 0x%x) is the one mark this test skips; it must be "+
"covered by mainRoutingPresentFam instead, and it is not", divert)
}
f.routes[mainKey] = saved
}
// markBindingView is one `ip <fam> rule add fwmark <M> lookup <T>` the apply issued.
type markBindingView struct{ fam, table string }
type planeView struct {
rules map[string]string
routes map[string]string
}
// healthyPlaneView synthesises the `ip rule show` / `ip route show table N`
// output of a fully installed plane from the rule-adds the apply issued.
func healthyPlaneView(routed map[uint64][]markBindingView, mainTable uint64) planeView {
v := planeView{rules: map[string]string{}, routes: map[string]string{}}
pref := 32765
for _, mark := range sortedMarks(markSet(routed)) {
for _, b := range routed[mark] {
v.rules[b.fam] += fmt.Sprintf("%d:\tfrom all fwmark 0x%x lookup %s\n", pref, mark, b.table)
pref--
if b.table == fmt.Sprintf("%d", mainTable) {
v.routes[b.fam+" "+b.table] = "local default dev lo scope host"
continue
}
v.routes[b.fam+" "+b.table] = "default dev somedev scope link\n" +
"unreachable default metric " + unreachableFloorMetric
}
}
return v
}
func markSet(routed map[uint64][]markBindingView) map[uint64]bool {
out := map[uint64]bool{}
for k := range routed {
out[k] = true
}
return out
}
func sortedMarks(set map[uint64]bool) []uint64 {
out := make([]uint64, 0, len(set))
for k := range set {
out = append(out, k)
}
sort.Slice(out, func(i, j int) bool { return out[i] < out[j] })
return out
}
func containsCall(calls []string, want string) bool {
for _, c := range calls {
if c == want {
return true
}
}
return false
}
// --- defect 4's other half: the duplicated layout constants -----------------
// TestMarkLayoutConstantsLockstep pins model's copy of the mark/table layout to
// the netplane functions that actually compute it. model.ValidateMarkBases has
// to know the offsets to tell an operator that fwmark_base 0x7f puts the L3 mark
// on 0xff; it cannot import netplane (the dependency runs the other way), so it
// carries its own copy. A drift would make the validator bless a base the data
// plane cannot use, or condemn one it can — silently, in both directions.
func TestMarkLayoutConstantsLockstep(t *testing.T) {
g := model.Globals{FwmarkBase: 0x4000, TableBase: 0x5000}
checks := []struct {
what string
derived, model uint32
}{
{"loop-guard mark", loopMark, model.MarkLoopGuard},
{"L3 mark offset", L3Mark(g) - g.FwmarkBase, model.MarkL3Offset},
{"egress mark offset", EgressMark(g, 0) - g.FwmarkBase, model.MarkEgressOffset},
{"egress mark stride", EgressMark(g, 7) - EgressMark(g, 0), 7},
{"L3 table offset", L3Table(g) - g.TableBase, model.TableL3Offset},
{"egress table offset", EgressTable(g, 0) - g.TableBase, model.TableEgressOffset},
{"egress table stride", EgressTable(g, 7) - EgressTable(g, 0), 7},
}
for _, c := range checks {
if c.derived != c.model {
t.Errorf("%s: netplane computes 0x%x, model.ValidateMarkBases assumes 0x%x — the validator "+
"is now judging a layout that does not exist", c.what, c.derived, c.model)
}
}
}
+355 -23
View File
@@ -45,6 +45,12 @@ const (
egEgressMarkOffset = 0x100
egEgressTableOffset = 0x10
// l3MarkOffset / l3TableOffset carve the L3-ingress mark and routing table
// out of the same reserved block, BELOW the per-egress range so no number of
// egresses can ever collide with them.
l3MarkOffset = 0x80
l3TableOffset = 0x08
// defaultFwmark / defaultTable are the model contract defaults when a raw
// Globals arrives with the fields unset (0).
defaultFwmark = 0x2000
@@ -81,6 +87,90 @@ func EgressTable(g model.Globals, idx int) uint32 {
return effTable(g) + egEgressTableOffset + uint32(idx)
}
// L3Device is the TUN interface the engine opens for the L3 ingress. Non-TCP/UDP
// LAN traffic (today: ICMP echo) is policy-routed into it so the engine can carry
// it through an L3-capable outbound. The generate stage emits an inbound bound to
// this exact name; ApplyRouting points the L3 table's default route at it.
const L3Device = "shater-l3"
// L3Mark is the fwmark the nft prerouting chain stamps on LAN traffic the tunnel
// can only carry at layer 3, and which ApplyRouting binds to L3Table.
func L3Mark(g model.Globals) uint32 {
return effFwmark(g) + l3MarkOffset
}
// L3Table is the routing table whose default route leaves via L3Device.
func L3Table(g model.Globals) uint32 {
return effTable(g) + l3TableOffset
}
// L3Enabled reports whether the L3 ingress is configured. It is opt-in: without
// it the untunnelable policy behaves exactly as before.
func L3Enabled(g model.Globals) bool { return g.L3Tunnel }
// EgressDevice resolves an egress to the L3 device its mark is routed out of, or
// "" when this egress has no device at all (wrong type, or no interface).
//
// It is the SINGLE resolution that everything touching an egress mark must
// agree on, because those things are not independent: the prerouting marking
// (UntunnelableEgressBinding), the forward-chain accept that lets that mark past
// the kill-switch, and addEgressRouting, which installs the mark's `ip rule` and
// routing table. A mark that one of them stamps and another does not route is
// not a cosmetic inconsistency — it is a packet that falls through to the main
// table and leaves over the plain WAN with the router's real address.
//
// They HAD already diverged: the binding rejected an egress on `Interface == ""`
// while model.ValidateUntunnelableEgress rejected it on
// `strings.TrimSpace(Interface) == ""`, so `option interface ' '` produced a
// panel that said "the option is ignored" over a data plane that was marking
// packets for a table nobody built. Trimming here is what the operator meant and
// what the validator already assumed; TestUntunnelableEgressResolutionLockstep
// pins the two verdicts together so the next drift is a test failure and not a
// leak.
func EgressDevice(eg model.Egress) string {
if t := strings.ToLower(strings.TrimSpace(eg.Type)); t != "interface" && t != "tunnel" {
return ""
}
// IfaceDevice("") falls back to br-lan — the right default for an INBOUND
// with no network, catastrophic here: the LAN bridge would impersonate the
// egress device while addEgressRouting never installs the mark's routing, so
// marked packets would fall through to the main table and leave over the
// plain WAN. An egress without an interface has no device.
iface := strings.TrimSpace(eg.Interface)
if iface == "" {
return ""
}
return IfaceDevice(iface)
}
// UntunnelableEgressBinding resolves Globals.UntunnelableEgress to the egress's
// index, its nft mark and its device. ok is false when the option is empty or
// names something that is not an interface/tunnel egress with a device — the
// caller then renders nothing and the Untunnelable policy stays in sole charge,
// which is the fail-closed reading of a typo.
//
// The mark and table are the egress's OWN (EgressMark/EgressTable), not a third
// pair: addEgressRouting already binds them to that device, so routing the
// untunnelable protocols there is a matter of stamping the existing mark in
// prerouting and nothing else.
func UntunnelableEgressBinding(m *model.Model) (idx int, mark uint32, device string, ok bool) {
name := strings.TrimSpace(m.Globals.UntunnelableEgress)
if name == "" {
return 0, 0, "", false
}
for i, eg := range m.Egresses {
if !strings.EqualFold(strings.TrimSpace(eg.Name), name) {
continue
}
dev := EgressDevice(eg)
if dev == "" {
return 0, 0, "", false
}
return i, EgressMark(m.Globals, i), dev, true
}
return 0, 0, "", false
}
// EgressOutboundTag is the outbound tag the generator emits for an
// interface/tunnel egress (coordination contract with the generate stage).
func EgressOutboundTag(name string) string { return "egress-" + name }
@@ -358,11 +448,33 @@ func nftPlanRules(m *model.Model, _ time.Time) []model.Rule {
return sortedRules(rules)
}
// nftDivertDevs returns EVERY LAN-ingress device whose traffic this ruleset
// diverts into the engine: the enabled tproxy inbounds' devices PLUS the devices
// named by any enabled rule's `iface:`/`zone:` source. rules is the EFFECTIVE
// rule list (profile overrides applied — nftPlanRules), never raw m.Rules: the
// divert set must match what RenderNft actually emits per-rule.
// divertRef is one device this ruleset diverts, together with WHAT THE OPERATOR
// WROTE to bring it in — the inbound's `network`, an `iface:` source, a `zone:`
// source.
//
// The provenance is not decoration. IfaceDevice returns the UCI name unchanged
// when it cannot resolve it (see its doc comment), so a divert set is a list of
// strings in which "br-lan" and "lan" are indistinguishable — one is a device,
// the other is a name that will match nothing. Reporting that to an operator as
// `device "lan" does not exist` is useless; reporting it as `interface "lan"
// resolved to no device` is actionable, and only the reference knows which is
// which. coverage.go is the only consumer.
type divertRef struct {
// Src is the operator-facing name: the inbound's UCI network, the bare name
// from an `iface:` source, or "zone:<name>" for a device pulled out of a
// firewall zone.
Src string
// Dev is what IfaceDevice/nftZoneDevices resolved Src to — the string that
// actually appears in `iifname "..."`.
Dev string
}
// nftDivertRefs returns EVERY LAN-ingress device whose traffic this ruleset
// diverts into the engine, with its provenance: the enabled tproxy inbounds'
// devices PLUS the devices named by any enabled rule's `iface:`/`zone:` source.
// rules is the EFFECTIVE rule list (profile overrides applied — nftPlanRules),
// never raw m.Rules: the divert set must match what RenderNft actually emits
// per-rule.
//
// The distinction matters for fail-closed correctness. RenderNft emits per-rule
// tproxy lines for `iface:`/`zone:` sources, and those devices need NOT be tproxy
@@ -374,15 +486,22 @@ func nftPlanRules(m *model.Model, _ time.Time) []model.Rule {
// the tproxied packet to the engine socket at all.
//
// So: everything we divert, we must also be able to fail closed on and set the
// ingress sysctls for. Result is deduped and order-stable (inbounds first).
func nftDivertDevs(m *model.Model, rules []model.Rule) []string {
devs := nftEnabledInboundDevs(m)
if _, ok := nftPrimaryInbound(m); !ok {
// No tproxy inbound => RenderNft emits no per-rule divert lines either.
return devs
// ingress sysctls for. Result is deduped BY DEVICE and order-stable (inbounds
// first), so the first reference that produced a device is the one reported.
func nftDivertRefs(m *model.Model, rules []model.Rule) []divertRef {
var refs []divertRef
for _, in := range m.Inbounds {
if model.IsTproxyInbound(in) {
refs = append(refs, divertRef{
Src: orDefault(strings.TrimSpace(in.Network), "lan"),
Dev: IfaceDevice(in.Network),
})
}
}
if len(devs) == 0 {
return devs
// No tproxy inbound => RenderNft emits no per-rule divert lines either; and
// with no inbound device there is nothing for a rule source to add to.
if _, ok := nftPrimaryInbound(m); !ok || len(refs) == 0 {
return dedupRefs(refs)
}
for _, r := range rules {
if !r.Enabled {
@@ -392,12 +511,44 @@ func nftDivertDevs(m *model.Model, rules []model.Rule) []string {
s := strings.TrimSpace(raw)
switch {
case strings.HasPrefix(s, "iface:"):
devs = append(devs, IfaceDevice(strings.TrimPrefix(s, "iface:")))
name := strings.TrimPrefix(s, "iface:")
refs = append(refs, divertRef{Src: name, Dev: IfaceDevice(name)})
case strings.HasPrefix(s, "zone:"):
devs = append(devs, nftZoneDevices(strings.TrimPrefix(s, "zone:"))...)
zone := strings.TrimPrefix(s, "zone:")
for _, dev := range nftZoneDevices(zone) {
refs = append(refs, divertRef{Src: "zone:" + zone, Dev: dev})
}
}
}
}
return dedupRefs(refs)
}
// dedupRefs drops references whose device was already seen (and blank devices),
// preserving order.
func dedupRefs(in []divertRef) []divertRef {
seen := map[string]bool{}
var out []divertRef
for _, r := range in {
dev := strings.TrimSpace(r.Dev)
if dev == "" || seen[dev] {
continue
}
seen[dev] = true
r.Dev = dev
out = append(out, r)
}
return out
}
// nftDivertDevs is nftDivertRefs projected onto the device names, which is what
// every renderer and the sysctl/fail-closed scoping consume.
func nftDivertDevs(m *model.Model, rules []model.Rule) []string {
refs := nftDivertRefs(m, rules)
devs := make([]string, 0, len(refs))
for _, r := range refs {
devs = append(devs, r.Dev)
}
return nftDedupStr(devs)
}
@@ -502,12 +653,20 @@ func RenderHoldNftAt(m *model.Model, now time.Time) (string, error) {
sb.WriteString("\t\ttype filter hook forward priority filter; policy accept;\n")
sb.WriteString("\t\t# HOLDING PLANE: the engine is down; LAN->WAN is blocked (fail-closed).\n")
// Router-origin / egress-marked traffic survives, so the daemon can still
// reach the network and recover on its own.
// reach the network and recover on its own. The egress accepts are the same
// mark+oifname conjunction the full plane uses (see the essay there): the
// mark says where a packet was sent, oifname says where it actually went,
// and only both together mean the egress's routing table did its job. Here
// the difference is theory rather than practice — this plane stamps no marks
// at all — but a bare mark accept in the plane whose entire purpose is
// failing closed would be exactly the wrong thing to leave lying around.
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", loopMark))
for i, eg := range m.Egresses {
if t := strings.ToLower(eg.Type); t == "interface" || t == "tunnel" {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
dev := EgressDevice(eg)
if dev == "" {
continue
}
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x oifname %q accept\n", EgressMark(m.Globals, i), dev))
}
// Keep the LAN itself usable.
sb.WriteString("\t\t" + iif + " ip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16 } accept\n")
@@ -521,6 +680,28 @@ func RenderHoldNftAt(m *model.Model, now time.Time) (string, error) {
// every destination is treated as tunnelled. `direct` still lets everything out
// (it needs no knowledge); `block` and `icmp` degrade to their conservative form.
sb.WriteString(untunnelableRules(m.Globals, iif, nil))
// Deliberately NO L3-ingress marking here (Globals.L3Tunnel): the TUN it
// routes into is created BY the engine, and this plane exists precisely
// because the engine is not running. Marking ICMP at a device that does not
// exist would not keep ping alive — it would dead-end the packets in an
// empty routing table while LOOKING like a feature. While holding, the
// honest verdicts are the policy above and the drops below.
//
// The untunnelable-egress marking (Globals.UntunnelableEgress) is absent
// for a harder reason than the TUN's: the egress DEVICE may well exist
// while the engine is down, but this plane is installed precisely when an
// apply aborted before the netplane stage — possibly before
// addEgressRouting ever installed the `ip rule fwmark -> table` pair
// (first boot after a power cut, engine choking on an undownloaded
// rule-set, is the canonical case). A packet marked without that pair
// falls through to the MAIN table, and the egress-mark accept above would
// wave it out the DEFAULT WAN — a kill-switch bypass wearing the egress's
// name. It cannot be done from here anyway: this plane is one forward
// chain by design, and a forward-hook mark cannot re-route a packet whose
// routing decision already happened. While holding, non-TCP/UDP traffic
// gets the Untunnelable policy verdict above — the same honest
// degradation the L3 ingress chose.
//
// Everything else from the protected devices is dropped, both families —
// IPv6 unconditionally, because with no engine there is no v6 divert either.
sb.WriteString("\t\t" + iif + " meta nfproto ipv4 drop\n")
@@ -598,7 +779,11 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
// Validate every device the plan wants to divert BEFORE rendering a single
// line, so a name we cannot express never reaches the kernel as a rule that
// matches nothing. See the refusal contract in the doc comment.
divertDevs := nftDivertDevs(m, planRules)
divertRefs := nftDivertRefs(m, planRules)
divertDevs := make([]string, 0, len(divertRefs))
for _, r := range divertRefs {
divertDevs = append(divertDevs, r.Dev)
}
validDevs, rejectedDevs := nftSplitDevs(divertDevs)
var warnings []string
for _, bad := range rejectedDevs {
@@ -606,6 +791,13 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
"interface name %q is not a usable device name and was REJECTED: it is not covered by the "+
"fail-closed guard and its traffic is not diverted", bad))
}
// What this plan does NOT cover: a device name that resolves to nothing the
// kernel has, and a client network of this router that the plan never mentions.
// Both are silent leaks of exactly the kind the rejected-name check above
// exists for, and both were previously visible only in the panel's interface
// picker (or nowhere at all). Collected BEFORE the refusal below, so a refused
// plan still tells the operator everything that is wrong with it.
warnings = append(warnings, coverageWarnings(m, divertRefs)...)
if len(rejectedDevs) > 0 && genGlobalClosed(m.Globals) {
return "", warnings, fmt.Errorf(
"refusing to apply: %d interface name(s) cannot be expressed in the ruleset (%s), so the "+
@@ -776,6 +968,11 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
}
}
// An L3-marked packet must never be re-diverted or re-marked either
// (Globals.L3Tunnel; same contract as the loop-guard/egress marks above).
if L3Enabled(m.Globals) {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", L3Mark(m.Globals)))
}
// --- DNS force-intercept (Globals.DNSIntercept) ---
// Divert ALL LAN plaintext :53 into the engine, INCLUDING queries addressed to
// the router itself, which the fib-local bypass below would otherwise hand to
@@ -794,6 +991,83 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
sb.WriteString("\t\tfib daddr type local accept\n")
sb.WriteString("\t\tip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 127.0.0.0/8, 169.254.0.0/16, 224.0.0.0/4, 255.255.255.255 } accept\n")
sb.WriteString("\t\tip6 daddr { ::1, fc00::/7, fe80::/10, ff00::/8 } accept\n")
// --- L3 ingress (Globals.L3Tunnel): route what TPROXY cannot carry ---
// Kernel TPROXY needs a socket to hand the packet to, so everything below
// moves TCP and UDP only. With the L3 ingress enabled, LAN ICMP/ICMPv6 bound
// for a PUBLIC destination (every local plane was accepted above: router
// itself via fib-local, RFC1918/link-local/multicast via the daddr sets) is
// fwmark-routed into the engine's TUN instead of falling to the forward-chain
// verdicts. The mark plus addL3Routing's `ip rule` are the ONLY way into that
// device: the TUN runs with auto_route off and the main routing table is
// never touched, by design — see L3Device.
//
// Deliberately ONLY icmp/ipv6-icmp, never `l4proto != { tcp, udp }`. The
// boundary is sing-tun's ForwardDispatcher, not the outbound: flow_parse.go
// classifies TCP, UDP and ICMP echo and nothing else, and flow_dispatch.go
// builds its NAT flows from port-like selectors, so a marked
// ESP/AH/GRE/IGMP/SCTP packet would enter the device and vanish — a black
// hole wearing a tunnel's name — instead of receiving the untunnelable
// policy's honest forward-chain verdict (drop, or direct-out if the
// operator chose to disclose that much). That policy — or the
// untunnelable-egress marking below, whose receiving side is the kernel
// and has no such boundary — stays in charge of everything this ingress
// cannot carry.
if L3Enabled(m.Globals) && dnsIif != "" {
sb.WriteString(fmt.Sprintf("\t\t%s ip protocol icmp meta mark set 0x%x accept\n", dnsIif, L3Mark(m.Globals)))
if m.Globals.IPv6 {
// ND/RA hold the LAN's v6 plane together and mean nothing off-link.
// Most are already out of reach (multicast daddr and router-addressed
// unicast were accepted above), but a unicast NS/NA between global LAN
// addresses is not, and one tunnelled neighbour probe is enough to
// take v6 down — so accept them BEFORE the icmpv6 mark. MLD is
// multicast-addressed and already covered by the ff00::/8 accept.
sb.WriteString("\t\t" + dnsIif + " icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept\n")
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto ipv6-icmp meta mark set 0x%x accept\n", dnsIif, L3Mark(m.Globals)))
}
}
// --- untunnelable egress (Globals.UntunnelableEgress): the kernel carries the rest ---
// Everything the L3 ingress above leaves alone — ESP, AH, GRE, IGMP, SCTP,
// and ICMP too when l3_tunnel is off — is stamped with the named egress's
// OWN mark and accepted; the `ip rule fwmark -> table` pair that
// addEgressRouting already binds to that egress then routes it out the
// egress device. No new mark, no new table, no engine in the path.
//
// The filter is the WIDE negated set while the L3 block above matches ICMP
// only, and the asymmetry is the receiving side. The L3 mark delivers into
// the engine's TUN, where sing-tun's ForwardDispatcher classifies nothing
// beyond TCP/UDP/ICMP echo — wide there is a black hole. This mark
// delivers to the kernel's own forwarding path, which routes ANY IP
// protocol — narrow here would just amputate the feature.
//
// The order is load-bearing twice. AFTER the L3 block: when l3_tunnel is
// on, ICMP must be claimed by the mark that routes it through the engine's
// own rules, and this line may only sweep up what that block did not take
// (first match wins). AFTER the local-plane accepts (fib-local, RFC1918/
// link-local/multicast, ND/RA): pinging the gateway, LAN-to-LAN traffic
// and v6 neighbour discovery must never leave through an uplink. Nothing
// is re-accepted here: the per-egress mgmt bypass at the top of this chain
// and its forward-chain mirror already pass this mark.
if _, uMark, _, uOK := UntunnelableEgressBinding(m); uOK && dnsIif != "" {
if m.Globals.IPv6 {
if !L3Enabled(m.Globals) {
// The guard the L3 block plants before its ICMPv6 mark, needed
// here whenever that block is off: a unicast NS/NA between
// global LAN addresses is not covered by the daddr accepts
// above, and one neighbour probe sent out an uplink is enough
// to take LAN IPv6 down. With l3_tunnel on it already ran.
sb.WriteString("\t\t" + dnsIif + " icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept\n")
}
sb.WriteString(fmt.Sprintf("\t\t%s meta l4proto != { tcp, udp } meta mark set 0x%x accept\n", dnsIif, uMark))
} else {
// addEgressRouting installs the -6 rule/table pair only when
// Globals.IPv6 is on. Marked v6 without that pair does not go to
// the egress — it falls through to the MAIN v6 table carrying a
// mark the forward chain accepts, i.e. out the default WAN past
// the kill-switch. Scope the mark to v4 and leave v6 to the
// forward chain's existing verdicts instead.
sb.WriteString(fmt.Sprintf("\t\t%s meta nfproto ipv4 meta l4proto != { tcp, udp } meta mark set 0x%x accept\n", dnsIif, uMark))
}
}
// --- DNS anti-leak: only :853 is carved out; plain :53 is DIVERTED (D14) ---
if dnsIif != "" {
// Keep encrypted-DNS bypass ports (DoT :853 tcp, DoQ :853 udp) OUT of
@@ -838,11 +1112,52 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
// 1) mgmt / interface+tunnel egress marks: router-origin and egress
// traffic must never be dropped (mirror of the prerouting bypass).
// These same accepts are what let untunnelable-egress traffic
// (Globals.UntunnelableEgress) past the kill-switch: prerouting
// stamps the egress's own mark on non-TCP/UDP LAN ingress, and it
// is accepted HERE.
//
// THE MARK ALONE IS NOT ENOUGH, and believing it was is what made
// this the leakiest line in the file. The mark says where the
// packet was SENT; it does not say where it WENT. `ip rule fwmark
// 0x2100 lookup 8208` only diverts the lookup — when that table is
// empty the lookup FAILS OVER to the main table, silently, and the
// packet leaves out the default WAN still carrying the mark this
// line accepts, straight past a closed kill-switch. The table is
// empty for the most ordinary reason there is: the egress rides an
// interface, the interface goes down, and the kernel deletes every
// route through that device. Nothing in this ruleset changed, so
// the applier's idempotence check sees no work, and the leak is
// permanent.
//
// Ordinary egress traffic never had this hole because the engine
// binds those sockets to the device (SO_BINDTODEVICE, see
// generate/outbound.go) and a dead device fails the socket. The
// untunnelable-egress path has no such second opinion: a mark is
// all it is made of. So the accept carries the second opinion
// instead — `meta mark X oifname "dev"` is the conjunction of "we
// sent it there" AND "it is actually going there", which is
// strictly narrower than either half and true only when the
// routing did what the mark asked for.
//
// (An oifname accept ALONE would be the loose one — it would bless
// everything the routing table ever pushes at that device, whatever
// put it there — which is why the mark stays. The comment this
// replaces argued that point correctly and then drew the wrong
// conclusion from it: it justified dropping oifname INSTEAD OF the
// mark, which nobody proposed, rather than using both.)
//
// An egress with no device (EgressDevice == "") gets no accept at
// all: addEgressRouting installs no rule and no table for it and
// prerouting stamps nothing, so an accept for its mark could only
// ever wave through something that has no business here.
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", loopMark))
for i, eg := range m.Egresses {
if t := strings.ToLower(eg.Type); t == "interface" || t == "tunnel" {
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x accept\n", EgressMark(m.Globals, i)))
dev := EgressDevice(eg)
if dev == "" {
continue
}
sb.WriteString(fmt.Sprintf("\t\tmeta mark 0x%x oifname %q accept\n", EgressMark(m.Globals, i), dev))
}
// 2) LAN-to-LAN / link-local (v4): the LAN itself must keep working
// (same private-range bypass sets the prerouting chain uses).
@@ -860,11 +1175,28 @@ func renderNft(m *model.Model, plan *UntunnelablePlan, now time.Time) (string, [
sb.WriteString("\t\t" + dnsIif + " ip6 daddr fe80::/10 accept\n")
sb.WriteString("\t\t" + dnsIif + " ip6 daddr ff00::/8 accept\n")
}
// 4) Untunnelable protocols (never TCP/UDP, so never divertible): the
// 4) L3 ingress legs (Globals.L3Tunnel). The LAN->TUN leg still has a
// LAN iifname, so the fail-closed drop below would eat exactly the
// packets prerouting just marked and routed — accept by the TUN
// oifname first. The TUN->LAN leg (iifname) is the engine answering
// LAN clients. These accepts speak for OUR table only: fw4's
// forward chain runs independently and a drop there still wins —
// the shater-l3 firewall zone that keeps fw4 out of the way is
// provisioned in uci-defaults, not here.
if L3Enabled(m.Globals) {
sb.WriteString(fmt.Sprintf("\t\toifname \"%s\" accept\n", L3Device))
sb.WriteString(fmt.Sprintf("\t\tiifname \"%s\" accept\n", L3Device))
}
// 5) Untunnelable protocols (never TCP/UDP, so never divertible): the
// configured policy decides whether they leave directly or are dropped
// with everything else. Must precede the drops to have any effect.
// With an untunnelable egress bound, the traffic prerouting marked
// never reaches these lines (accepted by mark in 1); the policy
// keeps judging what the marking skipped — v6 while Globals.IPv6
// is off — and everything, as before, when the option is empty
// or does not resolve.
sb.WriteString(untunnelableRules(m.Globals, dnsIif, plan))
// 5) Fail-closed drops for public-bound LAN-ingress traffic.
// 6) Fail-closed drops for public-bound LAN-ingress traffic.
sb.WriteString("\t\t" + dnsIif + " meta nfproto ipv4 drop\n")
sb.WriteString("\t\t" + dnsIif + " meta nfproto ipv6 drop\n")
case ipv6Off:
+463 -3
View File
@@ -204,10 +204,20 @@ func TestRenderNftForwardFailClosed(t *testing.T) {
if !strings.Contains(fwd, "meta mark 0xff accept") {
t.Errorf("closed mode: forward chain missing the mgmt loop-guard accept:\n%s", fwd)
}
// Egress-marked traffic (interface egress at idx 0 => 0x2100) is bypassed too.
if !strings.Contains(fwd, "meta mark 0x2100 accept") {
// Egress-marked traffic (interface egress at idx 0 => 0x2100) is bypassed too,
// but ONLY in conjunction with the egress DEVICE. The mark records where the
// packet was sent; oifname records where it actually went. They disagree
// exactly when the egress's routing table has been emptied (ifdown, network
// restart) and the lookup fell through to main — i.e. when the packet is on
// its way out the default WAN wearing a mark this chain was told to trust.
if !strings.Contains(fwd, `meta mark 0x2100 oifname "wwan0" accept`) {
t.Errorf("closed mode: forward chain missing the egress-mark bypass:\n%s", fwd)
}
if strings.Contains(fwd, "meta mark 0x2100 accept") {
t.Errorf("closed mode: the egress-mark accept must be qualified by oifname — a mark-only "+
"accept passes the same packet after its table went empty and it fell back to the "+
"default WAN:\n%s", fwd)
}
// Open: no v4 fail-closed drop in the forward chain.
mo := realisticModel()
@@ -640,7 +650,7 @@ func TestRenderHoldNft(t *testing.T) {
{"atomic replace", "delete table inet shater"},
{"forward hook", "type filter hook forward priority filter;"},
{"router-origin survives", "meta mark 0xff accept"},
{"egress mark survives", "meta mark 0x2100 accept"},
{"egress mark survives", `meta mark 0x2100 oifname "wwan0" accept`},
{"LAN-to-LAN survives", "ip daddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16"},
{"ICMPv6 ND survives", "nd-neighbor-solicit"},
{"v4 blocked", "meta nfproto ipv4 drop"},
@@ -973,3 +983,453 @@ func TestUntunnelableInHoldingPlane(t *testing.T) {
t.Errorf("block policy must emit no relaxation in the holding plane:\n%s", rs2)
}
}
// --- L3 ingress: LAN ICMP into the engine's TUN (Globals.L3Tunnel) -----------
// preroutingChain extracts the `chain prerouting { ... }` block, mirroring
// forwardChain, so an ordering assert cannot be satisfied by a similar line
// living in the wrong hook.
func preroutingChain(t *testing.T, rs string) string {
t.Helper()
start := strings.Index(rs, "chain prerouting {")
if start < 0 {
t.Fatalf("no prerouting chain in ruleset:\n%s", rs)
}
rest := rs[start:]
end := strings.Index(rest, "\n\t}\n")
if end < 0 {
t.Fatalf("unterminated prerouting chain:\n%s", rest)
}
return rest[:end]
}
// TestL3IngressOptIn: l3_tunnel is opt-in, and with it unset the render must be
// byte-for-byte what it was before the feature existed — no TUN device, no L3
// mark, no extra ND accept. Every deployed router upgrades through this default;
// a single changed line here is an unreviewed behaviour change on all of them.
func TestL3IngressOptIn(t *testing.T) {
m := realisticModel() // L3Tunnel deliberately left false
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
// 0x2080 = FwmarkBase 0x2000 + the 0x80 L3 offset; it appears in NO other
// mark or table, so its absence covers every gated line that stamps or
// bypasses the mark.
for _, s := range []string{L3Device, "0x2080"} {
if strings.Contains(rs, s) {
t.Errorf("L3Tunnel=false must leave the render untouched, found %q:\n%s", s, rs)
}
}
// With IPv6 on and L3 off there is no ND/RA accept anywhere: the forward
// chain's ND line exists only in the IPv6-OFF branch. Finding one would mean
// the L3 prerouting block leaked past its gate.
m6 := realisticModel()
m6.Globals.IPv6 = true
rs6, err := RenderNft(m6)
if err != nil {
t.Fatalf("RenderNft(ipv6): %v", err)
}
if strings.Contains(rs6, "nd-router-solicit") {
t.Errorf("L3Tunnel=false leaked the prerouting ND/RA accept:\n%s", rs6)
}
}
// TestL3IngressDivertsICMP: with l3_tunnel=1, LAN ICMP is marked for the L3
// table in prerouting AND both TUN legs are let through the forward chain.
// Losing the mark means ping is back to the engine's faked echo replies; losing
// the forward legs means the marked packet is eaten by the very fail-closed
// drop the mark exists to route around.
func TestL3IngressDivertsICMP(t *testing.T) {
m := realisticModel()
m.Globals.L3Tunnel = true
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre := preroutingChain(t, rs)
mark := strings.Index(pre, "ip protocol icmp meta mark set 0x2080 accept")
if mark < 0 {
t.Fatalf("prerouting does not mark LAN ICMP for the L3 table:\n%s", pre)
}
// The marking must be ingress-scoped, or WAN-ingress ICMP gets marked too.
line := pre[strings.LastIndex(pre[:mark], "\n")+1 : mark]
if !strings.Contains(line, "iifname") {
t.Errorf("L3 mark is unscoped (would also mark WAN ingress): %q", strings.TrimSpace(line))
}
// Router-local and LAN-to-LAN ICMP must be accepted BEFORE the mark:
// pinging the gateway or a LAN neighbour must never enter the tunnel.
if fibLocal := strings.Index(pre, "fib daddr type local accept"); fibLocal < 0 || fibLocal > mark {
t.Errorf("router-local accept must precede the L3 mark or pinging the gateway breaks:\n%s", pre)
}
if lanV4 := strings.Index(pre, "ip daddr { 10.0.0.0/8"); lanV4 < 0 || lanV4 > mark {
t.Errorf("private-range accept must precede the L3 mark or LAN-to-LAN ping breaks:\n%s", pre)
}
// The mgmt bypass keeps an already-marked packet from being re-diverted.
if !strings.Contains(pre, "meta mark 0x2080 accept") {
t.Errorf("prerouting missing the L3-mark mgmt bypass:\n%s", pre)
}
fwd := forwardChain(t, rs)
oif := strings.Index(fwd, "oifname \"shater-l3\" accept")
if oif < 0 {
t.Fatalf("forward chain missing the LAN->TUN leg (oifname): the marked packet is dropped fail-closed:\n%s", fwd)
}
if !strings.Contains(fwd, "iifname \"shater-l3\" accept") {
t.Errorf("forward chain missing the TUN->LAN leg (iifname): the engine's replies are droppable:\n%s", fwd)
}
if drop := strings.Index(fwd, "meta nfproto ipv4 drop"); drop >= 0 && oif > drop {
t.Errorf("TUN accept must precede the fail-closed drop or it never matches:\n%s", fwd)
}
}
// TestL3IngressNDBeforeMark: the ND/RA accept must STRICTLY precede the ICMPv6
// mark. One tunnelled neighbour solicitation is enough to break address
// resolution, which takes IPv6 down for the whole LAN while every dashboard
// stays green.
func TestL3IngressNDBeforeMark(t *testing.T) {
m := realisticModel()
m.Globals.L3Tunnel = true
m.Globals.IPv6 = true
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre := preroutingChain(t, rs)
nd := strings.Index(pre, "icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept")
v6mark := strings.Index(pre, "meta l4proto ipv6-icmp meta mark set 0x2080 accept")
if v6mark < 0 {
t.Fatalf("IPv6 on: prerouting does not mark ICMPv6 for the L3 table:\n%s", pre)
}
if nd < 0 {
t.Fatalf("IPv6 on: prerouting has no ND/RA accept guarding the ICMPv6 mark:\n%s", pre)
}
if nd > v6mark {
t.Errorf("ND/RA accept sits AFTER the ICMPv6 mark: neighbour discovery gets tunnelled and LAN IPv6 dies:\n%s", pre)
}
}
// TestL3IngressLeavesUntunnelablePolicyInCharge: the L3 ingress carries ICMP and
// NOTHING else. ESP/AH/GRE/IGMP/SCTP cannot enter the engine — sing-tun's
// ForwardDispatcher classifies only TCP/UDP/ICMP echo and builds its NAT flows
// from port-like selectors, so other protocols are never dispatched — and
// marking them would trade the untunnelable policy's honest verdict (drop, or
// direct-out by explicit operator choice) for a silent black hole inside the
// TUN.
func TestL3IngressLeavesUntunnelablePolicyInCharge(t *testing.T) {
m := realisticModel()
m.Globals.L3Tunnel = true
m.Globals.IPv6 = true
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
// Every line that stamps the L3 mark must match ICMP explicitly — never the
// negated set `!= { tcp, udp }` that would sweep ESP/GRE along.
for _, line := range strings.Split(rs, "\n") {
if !strings.Contains(line, "meta mark set 0x2080") {
continue
}
if !strings.Contains(line, "ip protocol icmp") && !strings.Contains(line, "ipv6-icmp") {
t.Errorf("non-ICMP traffic marked into the TUN (black hole, not tunnel): %q", strings.TrimSpace(line))
}
}
// The policy machinery itself survives L3: with `direct` the negated-set
// accept is still emitted, and the TUN legs precede it, so ICMP tunnels
// while ESP/GRE still leave by the operator's explicit choice.
m.Globals.Untunnelable = "direct"
rsDirect, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft(direct): %v", err)
}
fwd := forwardChain(t, rsDirect)
policy := strings.Index(fwd, "meta l4proto != { tcp, udp } accept")
if policy < 0 {
t.Fatalf("untunnelable=direct lost its accept with L3 on:\n%s", fwd)
}
if oif := strings.Index(fwd, "oifname \"shater-l3\" accept"); oif < 0 || oif > policy {
t.Errorf("TUN legs must precede the untunnelable accept — L3 owns what it can carry:\n%s", fwd)
}
// Under the default `block`, ESP/GRE are still dropped by policy: no
// untunnelable accept appears, and the fail-closed drops remain.
m.Globals.Untunnelable = ""
rsBlock, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft(block): %v", err)
}
fwdBlock := forwardChain(t, rsBlock)
if strings.Contains(fwdBlock, "l4proto != { tcp, udp }") {
t.Errorf("block policy must emit no untunnelable accept, L3 on or off:\n%s", fwdBlock)
}
if !strings.Contains(fwdBlock, "meta nfproto ipv4 drop") {
t.Errorf("block policy lost the fail-closed drop with L3 on:\n%s", fwdBlock)
}
}
// TestL3IngressAbsentFromHoldingPlane: while the engine is down its TUN does not
// exist, so the holding plane must not send anything at shater-l3. Marking ICMP
// at a nonexistent device is not "ping keeps working" — it is a black hole that
// LOOKS like a feature while the engine is in exactly the state the holding
// plane exists to be honest about.
func TestL3IngressAbsentFromHoldingPlane(t *testing.T) {
m := realisticModel()
m.Globals.L3Tunnel = true
m.Globals.IPv6 = true
rs, err := RenderHoldNft(m)
if err != nil {
t.Fatalf("RenderHoldNft: %v", err)
}
for _, s := range []string{L3Device, "0x2080"} {
if strings.Contains(rs, s) {
t.Errorf("holding plane must not reference the L3 ingress (found %q) — the engine is down and the TUN is gone:\n%s", s, rs)
}
}
// The hold still fails closed for what the L3 ingress would have carried.
if !strings.Contains(rs, "meta nfproto ipv4 drop") {
t.Errorf("holding plane lost its fail-closed drop:\n%s", rs)
}
}
// --- untunnelable egress: the kernel carries what nothing else can ----------
// TestUntunnelableEgressOptIn: untunnelable_egress is opt-in, and with it unset
// the render must be byte-for-byte what it was before the feature existed.
// Every deployed router upgrades through this default; a marking line appearing
// here would silently re-route ESP/GRE/ICMP on all of them without anyone
// having asked for it.
func TestUntunnelableEgressOptIn(t *testing.T) {
m := realisticModel() // UntunnelableEgress deliberately left empty
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
// 0x2100 itself legitimately appears (the wwan egress's mgmt bypass); what
// must not exist is anything STAMPING it, or any negated-set match in
// prerouting — the policy's own negated set lives in the forward chain.
if strings.Contains(rs, "meta mark set 0x2100") {
t.Errorf("empty untunnelable_egress must not stamp the egress mark anywhere:\n%s", rs)
}
if pre := preroutingChain(t, rs); strings.Contains(pre, "l4proto != { tcp, udp }") {
t.Errorf("empty untunnelable_egress must leave prerouting free of the negated set:\n%s", pre)
}
// Same under a permissive policy: `direct` emits its negated-set accept in
// the FORWARD chain; prerouting marking must still not appear.
m.Globals.Untunnelable = "direct"
rsDirect, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft(direct): %v", err)
}
if pre := preroutingChain(t, rsDirect); strings.Contains(pre, "l4proto != { tcp, udp }") {
t.Errorf("the policy's accept belongs to forward; prerouting marking appeared without the option:\n%s", pre)
}
}
// TestUntunnelableEgressCarriesAllProtocols: with the option bound to an
// interface egress, prerouting stamps the egress's OWN mark on the whole
// non-TCP/UDP remainder and the forward chain accepts that mark ahead of the
// fail-closed drops. Losing the prerouting line quietly returns ESP/GRE to the
// untunnelable policy's verdict (typically: dropped); losing the forward accept
// means the kernel routes the marked packet at the egress and the kill-switch
// eats it on the way — the option reads configured and carries nothing.
func TestUntunnelableEgressCarriesAllProtocols(t *testing.T) {
m := realisticModel() // IPv6 off: the marking must stay v4-scoped
m.Globals.UntunnelableEgress = "wwan"
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre := preroutingChain(t, rs)
mark := strings.Index(pre, "meta nfproto ipv4 meta l4proto != { tcp, udp } meta mark set 0x2100 accept")
if mark < 0 {
t.Fatalf("prerouting does not mark the non-TCP/UDP remainder for the egress — or is not v4-scoped "+
"while IPv6 is off, in which case marked v6 finds no -6 rule, falls through to the main table "+
"and leaves over the default WAN:\n%s", pre)
}
// Ingress-scoped, or WAN-ingress ESP/GRE gets marked and re-routed too.
line := pre[strings.LastIndex(pre[:mark], "\n")+1 : mark]
if !strings.Contains(line, "iifname") {
t.Errorf("egress marking is unscoped (would also mark WAN ingress): %q", strings.TrimSpace(line))
}
// The mgmt bypass keeps an already-marked packet from being re-diverted.
if !strings.Contains(pre, "meta mark 0x2100 accept") {
t.Errorf("prerouting missing the egress-mark mgmt bypass:\n%s", pre)
}
fwd := forwardChain(t, rs)
acc := strings.Index(fwd, `meta mark 0x2100 oifname "wwan0" accept`)
drop := strings.Index(fwd, "meta nfproto ipv4 drop")
if acc < 0 {
t.Fatalf("forward chain does not accept the egress mark: the kernel routes marked ESP/GRE at the egress and the kill-switch drops it on the way:\n%s", fwd)
}
if drop >= 0 && acc > drop {
t.Errorf("the egress-mark accept must precede the fail-closed drop or it never matches:\n%s", fwd)
}
// IPv6 on: the -6 rule/table half exists too, so the marking widens to the
// family-neutral set.
m.Globals.IPv6 = true
rs6, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft(ipv6): %v", err)
}
pre6 := preroutingChain(t, rs6)
if !strings.Contains(pre6, " meta l4proto != { tcp, udp } meta mark set 0x2100 accept") ||
strings.Contains(pre6, "meta nfproto ipv4 meta l4proto != { tcp, udp } meta mark set 0x2100") {
t.Errorf("IPv6 on: the marking must cover both families — v6 ESP/GRE has a -6 rule and table to land in:\n%s", pre6)
}
// The holding plane stays marking-free either way: while the engine is down
// nothing guarantees the mark's `ip rule` was ever installed, and a mark
// with no rule is a main-table fallthrough that the hold's own egress-mark
// accept would wave out the default WAN.
hold, err := RenderHoldNft(m)
if err != nil {
t.Fatalf("RenderHoldNft: %v", err)
}
if strings.Contains(hold, "meta mark set") {
t.Errorf("holding plane must not stamp any mark — a mark without its routing is a kill-switch bypass:\n%s", hold)
}
}
// TestUntunnelableEgressFailClosedOnTypo: a name that resolves to no egress, or
// to one the kernel cannot route by mark (byedpi/direct egresses have no device
// and no mark routing), must change NOTHING — the untunnelable policy stays in
// sole charge. Rendering the marking anyway would stamp packets with a mark no
// `ip rule` serves: they fall through to the main table and leave over the
// default WAN — a typo turned kill-switch bypass.
func TestUntunnelableEgressFailClosedOnTypo(t *testing.T) {
build := func(name string) *model.Model {
m := realisticModel()
m.Egresses = append(m.Egresses,
model.Egress{Name: "dpi", Type: "byedpi"},
model.Egress{Name: "straight", Type: "direct"},
)
m.Globals.UntunnelableEgress = name
return m
}
baseline, err := RenderNft(build(""))
if err != nil {
t.Fatalf("RenderNft(baseline): %v", err)
}
for _, name := range []string{"wan2", "dpi", "straight"} {
rs, err := RenderNft(build(name))
if err != nil {
t.Fatalf("RenderNft(%q): %v", name, err)
}
if rs != baseline {
t.Errorf("untunnelable_egress=%q names no routable egress and must render byte-identical to the empty option:\n%s", name, rs)
}
}
if fwd := forwardChain(t, baseline); !strings.Contains(fwd, "meta nfproto ipv4 drop") {
t.Errorf("the fail-closed drop must survive an unresolvable option:\n%s", fwd)
}
}
// TestUntunnelableEgressEmptyInterfaceIsNotBrLan: an interface egress with NO
// interface must not resolve. IfaceDevice("") falls back to br-lan, so without
// the binding's own guard the LAN bridge would impersonate the egress device:
// the marking renders, addEgressRouting (which skips on the same emptiness
// check) never installs the mark's rule or table, marked ESP/GRE falls through
// to the main table, and the forward chain accepts it by mark — out the default
// WAN past a closed kill-switch.
func TestUntunnelableEgressEmptyInterfaceIsNotBrLan(t *testing.T) {
build := func(name string) *model.Model {
m := realisticModel()
m.Egresses = append(m.Egresses, model.Egress{Name: "hole", Type: "interface", Interface: ""})
m.Globals.UntunnelableEgress = name
return m
}
baseline, err := RenderNft(build(""))
if err != nil {
t.Fatalf("RenderNft(baseline): %v", err)
}
rs, err := RenderNft(build("hole"))
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
if rs != baseline {
t.Errorf("an egress without an interface has no device; rendering its marking routes ESP/GRE into the main table and out the default WAN:\n%s", rs)
}
}
// TestUntunnelableEgressYieldsICMPToL3: with both features on, the L3 ICMP mark
// must STRICTLY precede the egress marking — nft evaluates first-match, so the
// reverse order sends ICMP out the uplink instead of through the engine's
// tunnel, and the L3 ingress is silently dead while both options read enabled.
func TestUntunnelableEgressYieldsICMPToL3(t *testing.T) {
m := realisticModel()
m.Globals.L3Tunnel = true
m.Globals.UntunnelableEgress = "wwan"
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre := preroutingChain(t, rs)
icmp := strings.Index(pre, "ip protocol icmp meta mark set 0x2080 accept")
wide := strings.Index(pre, "meta l4proto != { tcp, udp } meta mark set 0x2100 accept")
if icmp < 0 {
t.Fatalf("L3 lost its ICMP mark with the egress option on:\n%s", pre)
}
if wide < 0 {
t.Fatalf("egress marking lost with L3 on:\n%s", pre)
}
if icmp > wide {
t.Errorf("egress marking precedes the L3 ICMP mark: ICMP rides the uplink and the L3 ingress silently dies:\n%s", pre)
}
// Same discipline for the ICMPv6 leg when IPv6 is on.
m.Globals.IPv6 = true
rs6, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft(ipv6): %v", err)
}
pre6 := preroutingChain(t, rs6)
v6 := strings.Index(pre6, "meta l4proto ipv6-icmp meta mark set 0x2080 accept")
wide6 := strings.Index(pre6, "meta l4proto != { tcp, udp } meta mark set 0x2100 accept")
if v6 < 0 || wide6 < 0 {
t.Fatalf("IPv6 on: missing the ICMPv6 mark (%d) or the egress marking (%d):\n%s", v6, wide6, pre6)
}
if v6 > wide6 {
t.Errorf("egress marking precedes the L3 ICMPv6 mark: ICMPv6 rides the uplink and the v6 half of the L3 ingress dies:\n%s", pre6)
}
}
// TestUntunnelableEgressExceptionsFirst: the local-plane accepts — router-local
// (fib), private/link-local ranges, and the v6 ND/RA guard — must all precede
// the egress marking. After it they are dead lines: pinging the gateway leaves
// through the uplink, LAN-to-LAN traffic is re-routed out of the LAN, and one
// re-routed neighbour solicitation takes LAN IPv6 down.
func TestUntunnelableEgressExceptionsFirst(t *testing.T) {
m := realisticModel()
m.Globals.IPv6 = true // L3 off: the ND/RA guard must come from the egress block itself
m.Globals.UntunnelableEgress = "wwan"
rs, err := RenderNft(m)
if err != nil {
t.Fatalf("RenderNft: %v", err)
}
pre := preroutingChain(t, rs)
wide := strings.Index(pre, "meta l4proto != { tcp, udp } meta mark set 0x2100 accept")
if wide < 0 {
t.Fatalf("egress marking missing:\n%s", pre)
}
for what, needle := range map[string]string{
"router-local (fib) accept": "fib daddr type local accept",
"private v4 ranges accept": "ip daddr { 10.0.0.0/8",
"v6 local ranges accept": "ip6 daddr { ::1, fc00::/7, fe80::/10, ff00::/8 } accept",
"ND/RA guard": "icmpv6 type { nd-router-solicit, nd-router-advert, nd-neighbor-solicit, nd-neighbor-advert } accept",
} {
at := strings.Index(pre, needle)
if at < 0 {
t.Errorf("%s is missing from prerouting, so that plane now leaves through the uplink:\n%s", what, pre)
continue
}
if at > wide {
t.Errorf("%s sits after the egress marking and is dead — that plane leaves through the uplink:\n%s", what, pre)
}
}
}
+65 -8
View File
@@ -379,11 +379,26 @@ var (
// the best-effort fallback when file logging is off but syslog is on. logread
// has no wall-clock retention worth promising, hence "ranges approximate" in
// the note the handler prepends.
logSyslogScrape = func(w io.Writer) error {
cmd := exec.Command("logread", "-e", "shater")
//
// The ctx is the REQUEST's, narrowed by logScrapeTimeout, and it is the only
// thing that ever ends this child process. Without it, a browser tab that was
// closed mid-download left `logread` running, its pipe open and the handler
// goroutine blocked writing into a socket nobody reads — one leaked process and
// two leaked descriptors per abandoned download, on a daemon that runs for
// months. CommandContext kills the child on cancel, which closes the pipe, which
// releases the copy goroutine Run is waiting on.
logSyslogScrape = func(ctx context.Context, w io.Writer) error {
cmd := exec.CommandContext(ctx, "logread", "-e", "shater")
cmd.Stdout = w
return cmd.Run()
}
// logScrapeTimeout hard-bounds the scrape even for a client that is still there.
//
// Why 30s: `logread` reads a ring buffer that is a few hundred KiB at most and
// exits; on a healthy box it completes in milliseconds. 30s is therefore not a
// budget anybody spends, it is the answer to "logread itself has wedged" — which
// is exactly the situation an operator downloads the log during.
logScrapeTimeout = 30 * time.Second
// logSyncTimeout bounds how long a download waits for the log sink's barrier
// (Server.SetLogSync). See syncLogSink for why the wait is bounded at all.
//
@@ -494,21 +509,30 @@ func (s *Server) handleLog(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Disposition",
`attachment; filename="shaterd-`+rng+`-`+now.Format("20060102")+`.log"`)
// The download is the one response that is legitimately big (the sink's file
// budget tops out at model.LogMaxKBMax = 8 MiB) and so the one that cannot live
// under a fixed WriteTimeout: a slow-but-real client would be cut off mid-file.
// out renews the write deadline on every chunk instead, so progress buys time and
// only a client that has genuinely STOPPED reading is dropped. See streamWriter.
out := newStreamWriter(w)
switch {
case !g.LogToFile && !g.LogToSyslog:
_, _ = io.WriteString(w, "# logging disabled\n")
_, _ = io.WriteString(out, "# logging disabled\n")
case !g.LogToFile:
_, _ = io.WriteString(w, "# note: file logging disabled; showing syslog ring only, ranges approximate\n")
if err := logSyslogScrape(w); err != nil {
_, _ = fmt.Fprintf(w, "# logread unavailable: %v\n", err)
_, _ = io.WriteString(out, "# note: file logging disabled; showing syslog ring only, ranges approximate\n")
ctx, cancel := context.WithTimeout(r.Context(), logScrapeTimeout)
defer cancel()
if err := logSyslogScrape(ctx, out); err != nil {
_, _ = fmt.Fprintf(out, "# logread unavailable: %v\n", err)
}
default:
segs := logsink.Segments(logSinkPath(g.LogPersist))
if len(segs) == 0 {
_, _ = io.WriteString(w, "# no log file yet\n")
_, _ = io.WriteString(out, "# no log file yet\n")
return
}
if err := logsink.CopyRange(w, segs, cutoff); err != nil {
if err := logsink.CopyRange(out, segs, cutoff); err != nil {
// Headers are long gone; nothing to do but note it server-side
// (usually the client hung up mid-download).
s.log.Debug("log download aborted: ", err)
@@ -516,6 +540,39 @@ func (s *Server) handleLog(w http.ResponseWriter, r *http.Request) {
}
}
// streamChunkTimeout is how long ONE write of a streamed response may block before
// the connection is torn down.
//
// Why 15s: a write blocks only while the socket send buffer is full, i.e. while the
// peer is not reading. A client that is reading at any rate at all — even a few KiB/s
// on a bad wireless link — drains the buffer far inside 15s and gets the deadline
// pushed forward again, so a slow download is never truncated. A client that has
// stopped (a suspended phone, a closed tab whose FIN was lost) is released in 15s
// instead of holding the goroutine, socket and fd until the process restarts.
const streamChunkTimeout = 15 * time.Second
// streamWriter renews the connection's write deadline before each chunk, turning a
// fixed server-wide WriteTimeout into a per-chunk idle timeout for one response. It
// is the only correct way to bound a large streamed body: the total is unknown, the
// per-write stall is not.
//
// SetWriteDeadline returning http.ErrNotSupported (httptest.NewRecorder, a wrapped
// ResponseWriter) is ignored on purpose — the writer then behaves exactly like the
// bare one and the server-wide WriteTimeout still applies.
type streamWriter struct {
w io.Writer
rc *http.ResponseController
}
func newStreamWriter(w http.ResponseWriter) *streamWriter {
return &streamWriter{w: w, rc: http.NewResponseController(w)}
}
func (s *streamWriter) Write(p []byte) (int, error) {
_ = s.rc.SetWriteDeadline(time.Now().Add(streamChunkTimeout))
return s.w.Write(p)
}
// handleStatsLog → GET /api/stats/log?limit=&before=&after=: the live DNS query log,
// ALWAYS newest first, each row carrying a monotonic `seq` cursor. With no cursor it
// returns the newest `limit` rows; before=<seq> returns the next OLDER page (seq<before);
+21
View File
@@ -109,9 +109,30 @@ type sessionRequest struct {
Token string `json:"token"`
}
// sessionBodyTimeout bounds how long the ONE unauthenticated route will wait for
// its (tiny) body.
//
// The server-wide DefaultReadTimeout has to accommodate a 4 MiB config PUT, so it is
// 60s. This route is different in kind: it is reachable with no credentials at all,
// and its entire legitimate body is a 64-hex-char token inside a JSON object — under
// 100 bytes, already delivered in the same TCP segment as the headers in every real
// client. Anything that needs longer than 5 seconds to produce it is not a panel.
//
// Without this, "Content-Length: 4096" plus one byte an hour bought an anonymous peer
// a goroutine, a socket and a file descriptor for as long as the read timeout allowed
// — repeatable until the daemon runs out of descriptors.
//
// A var, not a const, purely so the timeout tests can dial it down; production never
// assigns it.
var sessionBodyTimeout = 5 * time.Second
// handleSession validates a minted token and, on success, sets the session cookie.
// POST only. This is the panel's sole unauthenticated API route.
func (s *Server) handleSession(w http.ResponseWriter, r *http.Request) {
// Tighten the read deadline for this route only. ErrNotSupported is expected
// under httptest.NewRecorder (no real connection) and is not a failure: the
// server-wide ReadTimeout still applies.
_ = http.NewResponseController(w).SetReadDeadline(time.Now().Add(sessionBodyTimeout))
if r.Method != http.MethodPost {
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
return
+3 -2
View File
@@ -5,6 +5,7 @@ package panel
// are all package-var seams; each test swaps them and restores via t.Cleanup.
import (
"context"
"io"
"net/http"
"net/http/httptest"
@@ -131,7 +132,7 @@ func TestLogFileOffFallsBackToSyslogScrape(t *testing.T) {
g.LogToFile = false // syslog stays on
withLogSeams(t, g, "")
origScrape := logSyslogScrape
logSyslogScrape = func(w io.Writer) error {
logSyslogScrape = func(_ context.Context, w io.Writer) error {
_, err := io.WriteString(w, "ring line from logread\n")
return err
}
@@ -188,7 +189,7 @@ func TestLogFileOffIgnoresLeftoverSegments(t *testing.T) {
g.LogToFile = false // syslog stays on
withLogSeams(t, g, "2026-07-23T08:00:00Z INFO leftover line\n")
origScrape := logSyslogScrape
logSyslogScrape = func(w io.Writer) error {
logSyslogScrape = func(_ context.Context, w io.Writer) error {
_, err := io.WriteString(w, "ring line from logread\n")
return err
}
+70 -1
View File
@@ -39,6 +39,51 @@ const (
DefaultTokenTTL = 60 * time.Second
)
// Connection deadlines. Every one of these bounds a resource a REMOTE peer would
// otherwise hold indefinitely: a goroutine, a socket and a file descriptor each.
// On this box that matters more than on a server — a phone whose panel tab went to
// sleep mid-request is the normal case, not an attack, and the daemon runs for
// months without a restart, so "leaks one fd per stalled tab" is a countdown to the
// process-wide fd limit.
//
// Only ReadHeaderTimeout existed before. It bounds the request LINE and HEADERS and
// nothing else, so a peer that finished its headers, announced a Content-Length and
// then stopped sending was parked forever — reachable without any authentication at
// all through POST /api/session.
const (
// DefaultReadHeaderTimeout bounds the request line + headers. Unchanged value;
// named so the whole set is visible in one place.
DefaultReadHeaderTimeout = 10 * time.Second
// DefaultReadTimeout bounds reading an ENTIRE request, headers plus body.
//
// Why 60s: the largest legitimate body is a PUT /api/config at the 4 MiB cap
// (maxConfigBytes) — a real 380-node config is ~10x smaller, but the cap is what
// has to fit. 4 MiB in 60s is a 70 KB/s floor, which even a phone on a bad
// corner of the LAN clears by an order of magnitude. It is also short enough
// that a stalled connection is a bounded 60s cost instead of a permanent one.
// The unauthenticated route gets a much tighter budget of its own — see
// sessionBodyTimeout in auth.go.
DefaultReadTimeout = 60 * time.Second
// DefaultWriteTimeout bounds writing a response.
//
// Why 60s: every JSON endpoint answers in milliseconds; the one legitimately
// large response is the GET /api/log download, and that one renews its own
// deadline per chunk while the client keeps reading (see streamWriter in
// api.go), so this value is not the download's budget — it is the budget for a
// client that has stopped reading altogether.
DefaultWriteTimeout = 60 * time.Second
// DefaultIdleTimeout bounds an idle keep-alive connection between requests.
//
// Why 90s: the panel's own polling is the fastest legitimate reuse cadence and it
// is seconds, not minutes, so 90s never costs a live tab its connection, while a
// tab that was closed or suspended gives its fd back within a minute and a half
// instead of never.
DefaultIdleTimeout = 90 * time.Second
)
// Options configures a Server. The zero value is valid: every field falls back to
// the Default* above (or a safe built-in) in NewServer.
type Options struct {
@@ -56,6 +101,15 @@ type Options struct {
CookieSecure bool
// Logger is the daemon logger; nil => log.StdLogger().
Logger log.ContextLogger
// Connection deadlines. Zero => the Default* above; they exist as knobs mainly
// so tests can dial them down to milliseconds, but an embedder fronting the
// panel with a slow link can raise them. See the const block for the reasoning
// behind each default.
ReadHeaderTimeout time.Duration
ReadTimeout time.Duration
WriteTimeout time.Duration
IdleTimeout time.Duration
}
func (o Options) withDefaults() Options {
@@ -74,6 +128,18 @@ func (o Options) withDefaults() Options {
if o.Logger == nil {
o.Logger = log.StdLogger()
}
if o.ReadHeaderTimeout <= 0 {
o.ReadHeaderTimeout = DefaultReadHeaderTimeout
}
if o.ReadTimeout <= 0 {
o.ReadTimeout = DefaultReadTimeout
}
if o.WriteTimeout <= 0 {
o.WriteTimeout = DefaultWriteTimeout
}
if o.IdleTimeout <= 0 {
o.IdleTimeout = DefaultIdleTimeout
}
return o
}
@@ -119,7 +185,10 @@ func NewServer(a *apply.Applier, opts Options) *Server {
s.handler = s.buildRouter()
s.srv = &http.Server{
Handler: s.handler,
ReadHeaderTimeout: 10 * time.Second,
ReadHeaderTimeout: opts.ReadHeaderTimeout,
ReadTimeout: opts.ReadTimeout,
WriteTimeout: opts.WriteTimeout,
IdleTimeout: opts.IdleTimeout,
}
return s
}
+264
View File
@@ -0,0 +1,264 @@
// Timeout coverage for the panel's own http.Server (server.go) and for the two
// unbounded streaming paths in handleLog (api.go).
//
// These tests drive the REAL listener (Server.Start), not httptest over
// Server.Handler(): the whole point is the *http.Server field set, which
// httptest replaces with its own. Every deadline is dialled down through
// Options so the suite stays sub-second.
package panel
import (
"bufio"
"context"
"io"
"net"
"net/http"
"net/http/httptest"
"strconv"
"testing"
"time"
"github.com/sagernet/sing-box/shater/apply"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
)
// startLive brings up a Server on a real loopback port with the given Options
// overrides applied on top of the test defaults, and returns its address.
func startLive(t *testing.T, opts Options) (*Server, string) {
t.Helper()
t.Setenv("SHATER_SUBS_DIR", t.TempDir())
eng := engine.New()
a := apply.New(eng, nil)
opts.Addr = "127.0.0.1:0"
if opts.SessionTTL == 0 {
opts.SessionTTL = time.Hour
}
if opts.TokenTTL == 0 {
opts.TokenTTL = time.Minute
}
s := NewServer(a, opts)
ln, err := net.Listen("tcp", opts.Addr)
if err != nil {
t.Fatalf("listen: %v", err)
}
s.lnMu.Lock()
s.ln = ln
s.lnMu.Unlock()
go func() { _ = s.srv.Serve(ln) }()
t.Cleanup(func() { _ = s.Close() })
return s, ln.Addr().String()
}
// TestSlowBodyDoesNotHoldConnectionForever is the unauthenticated slow-loris:
// POST /api/session announcing a full 4 KiB body and then dribbling. Before the
// read deadline existed, the goroutine + socket + fd stayed parked until the
// client felt like finishing — i.e. forever.
func TestSlowBodyDoesNotHoldConnectionForever(t *testing.T) {
orig := sessionBodyTimeout
sessionBodyTimeout = 200 * time.Millisecond
t.Cleanup(func() { sessionBodyTimeout = orig })
_, addr := startLive(t, Options{
ReadTimeout: 30 * time.Second, // deliberately generous: the route's OWN budget must bite
WriteTimeout: 2 * time.Second,
})
conn, err := net.Dial("tcp", addr)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer conn.Close()
// Full headers (so ReadHeaderTimeout is satisfied), then one byte of a
// promised 4096-byte body and nothing more.
req := "POST /api/session HTTP/1.1\r\n" +
"Host: panel\r\n" +
"Content-Type: application/json\r\n" +
"Content-Length: 4096\r\n\r\n{"
if _, err := io.WriteString(conn, req); err != nil {
t.Fatalf("write request: %v", err)
}
// The server must give up on us. Either it closes (EOF/reset) or it answers
// 400 and closes; both end the read. What must NOT happen is a hang.
_ = conn.SetReadDeadline(time.Now().Add(5 * time.Second))
done := make(chan error, 1)
go func() {
_, err := io.Copy(io.Discard, conn)
done <- err
}()
select {
case err := <-done:
if ne, ok := err.(net.Error); ok && ne.Timeout() {
t.Fatalf("connection still open 5s into a stalled 4 KiB body — the read deadline is not set")
}
case <-time.After(5 * time.Second):
t.Fatalf("read never returned — the connection is held open by a stalled body")
}
}
// TestSlowBodyOnAuthenticatedRouteIsBounded covers the server-wide ReadTimeout —
// the backstop for every route that is NOT /api/session (which has its own,
// tighter budget). PUT /api/config accepts up to 4 MiB, so before ReadTimeout a
// stalled upload parked a goroutine and an fd indefinitely.
func TestSlowBodyOnAuthenticatedRouteIsBounded(t *testing.T) {
s, addr := startLive(t, Options{
ReadTimeout: 300 * time.Millisecond,
WriteTimeout: 2 * time.Second,
})
tok := s.MintToken()
conn, err := net.Dial("tcp", addr)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer conn.Close()
// Log in on this very connection, then start a PUT and stall it.
body := `{"token":"` + tok + `"}`
if _, err := io.WriteString(conn, "POST /api/session HTTP/1.1\r\nHost: panel\r\n"+
"Content-Type: application/json\r\n"+
"Content-Length: "+strconv.Itoa(len(body))+"\r\n\r\n"+body); err != nil {
t.Fatalf("write session: %v", err)
}
br := bufio.NewReader(conn)
resp, err := http.ReadResponse(br, nil)
if err != nil {
t.Fatalf("session response: %v", err)
}
_, _ = io.Copy(io.Discard, resp.Body)
resp.Body.Close()
var cookie string
for _, c := range resp.Cookies() {
if c.Name == DefaultCookieName {
cookie = c.Name + "=" + c.Value
}
}
if cookie == "" {
t.Fatalf("no session cookie")
}
if _, err := io.WriteString(conn, "PUT /api/config HTTP/1.1\r\nHost: panel\r\n"+
"Cookie: "+cookie+"\r\n"+
"Content-Type: application/json\r\n"+
"Content-Length: 65536\r\n\r\n{"); err != nil {
t.Fatalf("write PUT: %v", err)
}
_ = conn.SetReadDeadline(time.Now().Add(5 * time.Second))
if _, err := io.Copy(io.Discard, br); err != nil {
if ne, ok := err.(net.Error); ok && ne.Timeout() {
t.Fatalf("authenticated stalled body held the connection past ReadTimeout")
}
}
}
// TestSlowHeaderStillBounded pins the pre-existing ReadHeaderTimeout so a later
// refactor of the field set cannot drop it while adding the others.
func TestSlowHeaderStillBounded(t *testing.T) {
_, addr := startLive(t, Options{
ReadHeaderTimeout: 200 * time.Millisecond,
ReadTimeout: 5 * time.Second,
WriteTimeout: 2 * time.Second,
})
conn, err := net.Dial("tcp", addr)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer conn.Close()
if _, err := io.WriteString(conn, "GET /api/status HTTP/1.1\r\nHost: panel\r\n"); err != nil {
t.Fatalf("write: %v", err)
}
_ = conn.SetReadDeadline(time.Now().Add(3 * time.Second))
if _, err := io.Copy(io.Discard, conn); err != nil {
if ne, ok := err.(net.Error); ok && ne.Timeout() {
t.Fatalf("header read was never bounded")
}
}
}
// TestIdleKeepAliveIsBounded: a client that completes a request and then holds
// the keep-alive connection open without ever sending another one must be
// dropped. Without IdleTimeout the fd is that client's for as long as it wants.
func TestIdleKeepAliveIsBounded(t *testing.T) {
_, addr := startLive(t, Options{
ReadTimeout: 2 * time.Second,
WriteTimeout: 2 * time.Second,
IdleTimeout: 300 * time.Millisecond,
})
conn, err := net.Dial("tcp", addr)
if err != nil {
t.Fatalf("dial: %v", err)
}
defer conn.Close()
if _, err := io.WriteString(conn, "GET /api/status HTTP/1.1\r\nHost: panel\r\n\r\n"); err != nil {
t.Fatalf("write: %v", err)
}
br := bufio.NewReader(conn)
resp, err := http.ReadResponse(br, nil)
if err != nil {
t.Fatalf("read response: %v", err)
}
_, _ = io.Copy(io.Discard, resp.Body)
resp.Body.Close()
// Now go quiet. The server must close the idle connection.
_ = conn.SetReadDeadline(time.Now().Add(5 * time.Second))
if _, err := br.ReadByte(); err != nil {
if ne, ok := err.(net.Error); ok && ne.Timeout() {
t.Fatalf("idle keep-alive connection still open after 5s — IdleTimeout is not set")
}
return // EOF: the server closed it, which is the point
}
t.Fatalf("unexpected bytes on an idle connection")
}
// TestLogSyslogScrapeIsCancelledWithTheRequest: with file logging off, GET
// /api/log shells `logread`. The exec had no context, so a client that walked
// away left the child process, its pipe and the handler goroutine running.
func TestLogSyslogScrapeIsCancelledWithTheRequest(t *testing.T) {
g := model.DefaultGlobals()
g.LogToFile = false // syslog stays on -> the scrape path
withLogSeams(t, g, "")
entered := make(chan struct{})
released := make(chan struct{})
origScrape := logSyslogScrape
logSyslogScrape = func(ctx context.Context, w io.Writer) error {
close(entered)
<-ctx.Done() // in production this is what kills the `logread` child
close(released)
return ctx.Err()
}
t.Cleanup(func() { logSyslogScrape = origScrape })
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
ctx, cancel := context.WithCancel(context.Background())
req, _ := http.NewRequestWithContext(ctx, http.MethodGet, srv.URL+"/api/log?range=all", nil)
req.AddCookie(cookie)
go func() {
r, err := http.DefaultClient.Do(req)
if err == nil {
_, _ = io.Copy(io.Discard, r.Body)
r.Body.Close()
}
}()
select {
case <-entered:
case <-time.After(5 * time.Second):
t.Fatalf("scrape never started")
}
cancel()
select {
case <-released:
case <-time.After(5 * time.Second):
t.Fatalf("the syslog scrape outlived the request it was serving — no context is threaded into it")
}
}
+9 -1
View File
@@ -12,7 +12,13 @@
// CONTRACT: every type registered here is a type shater/generate emits
// (see generate.go's package comment for the authoritative list):
//
// - inbounds: tproxy, redirect, direct (dokodemo), socks, http, mixed
// - inbounds: tproxy, redirect, direct (dokodemo), socks, http, mixed, and
// tun — the synthetic "l3-in" L3 ingress generate emits under globals
// l3_tunnel, so non-TCP/UDP LAN traffic (ICMP echo) can reach an
// L3-capable outbound instead of being faked or dropped. tun keeps the
// slim promise: its gVisor stack is already linked because with_wireguard
// requires with_gvisor (scripts/router-tags.sh), so registering it adds no
// new build tag and no meaningful size.
// - outbounds: direct, block, selector, urltest, socks, http, shadowsocks,
// vmess, trojan, vless, shadowtls (+ hysteria2/tuic behind with_quic)
// - endpoints: wireguard/AWG (behind with_wireguard)
@@ -51,6 +57,7 @@ import (
"github.com/sagernet/sing-box/protocol/shadowtls"
"github.com/sagernet/sing-box/protocol/socks"
"github.com/sagernet/sing-box/protocol/trojan"
"github.com/sagernet/sing-box/protocol/tun"
"github.com/sagernet/sing-box/protocol/vless"
"github.com/sagernet/sing-box/protocol/vmess"
)
@@ -64,6 +71,7 @@ func Context(ctx context.Context) context.Context {
func InboundRegistry() *inbound.Registry {
registry := inbound.NewRegistry()
tun.RegisterInbound(registry)
redirect.RegisterRedirect(registry)
redirect.RegisterTProxy(registry)
direct.RegisterInbound(registry)
+314
View File
@@ -0,0 +1,314 @@
package stats
// Coverage for the retention bounds that were missing (the per-resolver and
// per-outbound DNS maps), for the visibility of every eviction, and for the size of
// Snapshot's critical section.
import (
"fmt"
"math/rand"
"net/netip"
"sort"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/common/dnstrack"
"github.com/sagernet/sing-box/log"
)
// warnLog records Warn lines so a test can assert an eviction was announced.
type warnLog struct {
log.ContextLogger
mu sync.Mutex
warns []string
}
func newWarnLog() *warnLog { return &warnLog{ContextLogger: log.StdLogger()} }
func (l *warnLog) Warn(args ...any) {
l.mu.Lock()
l.warns = append(l.warns, fmt.Sprint(args...))
l.mu.Unlock()
}
func (l *warnLog) lines() []string {
l.mu.Lock()
defer l.mu.Unlock()
return append([]string(nil), l.warns...)
}
// dnsEvent builds a LAN-sourced query event carrying a resolver tag and an outbound.
func dnsEvent(domain, server, outbound string) dnstrack.QueryEvent {
return dnstrack.QueryEvent{
Domain: domain,
QueryType: 1,
Client: netip.MustParseAddr("192.168.1.10"),
DNSServer: server,
Outbound: []string{outbound},
}
}
// TestOutboundMapHoldsAGenerationThenBoundsItself is the arithmetic behind
// maxOutboundKeys made executable. A live generation on this box is ~1000 outbound
// tags (≈380 nodes plus their per-group and per-chain copies); the map must carry
// that plus a rename-day overlap untouched, and must NOT carry an unbounded number
// of dead generations.
func TestOutboundMapHoldsAGenerationThenBoundsItself(t *testing.T) {
lg := newWarnLog()
a := New(nil, lg)
// Two whole generations: nothing may be evicted.
for gen := 0; gen < 2; gen++ {
for i := 0; i < 1000; i++ {
a.handleEvent(dnsEvent("a.example", "dns-local",
fmt.Sprintf("gen%d-node-%d", gen, i)))
}
}
if got := a.Snapshot().Dropped.Outbounds; got != 0 {
t.Fatalf("two live generations (2000 tags) already evicted %d entries — the cap is too low", got)
}
if len(lg.lines()) != 0 {
t.Fatalf("unexpected eviction notices while under the cap: %v", lg.lines())
}
// Now blow past it.
for i := 0; i < maxOutboundKeys; i++ {
a.handleEvent(dnsEvent("a.example", "dns-local", fmt.Sprintf("gen9-node-%d", i)))
}
a.mu.Lock()
size := len(a.outbounds)
a.mu.Unlock()
if size > maxOutboundKeys {
t.Fatalf("outbound map holds %d entries, above the %d cap — still unbounded", size, maxOutboundKeys)
}
snap := a.Snapshot()
if snap.Dropped.Outbounds == 0 {
t.Fatalf("map was capped but Snapshot reports nothing dropped — the eviction is invisible")
}
lines := lg.lines()
if len(lines) == 0 {
t.Fatalf("entries were evicted with no log line")
}
if !strings.Contains(lines[0], "outbounds") || !strings.Contains(lines[0], "sample, not a total") {
t.Fatalf("eviction notice does not name the map or its consequence: %q", lines[0])
}
}
// TestServerMapIsBounded: the resolver-tag map is small in practice, but "small in
// practice" was also true of every other map that turned out to grow forever.
func TestServerMapIsBounded(t *testing.T) {
a := New(nil, nil)
for i := 0; i < maxServerKeys*3; i++ {
a.handleEvent(dnsEvent("a.example", fmt.Sprintf("resolver-%d", i), ""))
}
a.mu.Lock()
size := len(a.servers)
a.mu.Unlock()
if size > maxServerKeys {
t.Fatalf("server map holds %d entries, above the %d cap", size, maxServerKeys)
}
if a.Snapshot().Dropped.Servers == 0 {
t.Fatalf("server map was capped but nothing was reported dropped")
}
}
// TestPruneKeepsTheBusiestEntries: eviction is by count, so the tags that actually
// carry traffic survive and the long tail goes. A cap that dropped the wrong ones
// would quietly falsify the panel's "top outbounds".
func TestPruneKeepsTheBusiestEntries(t *testing.T) {
a := New(nil, nil)
// One heavily-used tag...
for i := 0; i < 50; i++ {
a.handleEvent(dnsEvent("a.example", "dns-local", "busy-node"))
}
// ...and a long tail of one-hit wonders that overflows the cap.
for i := 0; i < maxOutboundKeys*2; i++ {
a.handleEvent(dnsEvent("a.example", "dns-local", fmt.Sprintf("tail-%d", i)))
}
a.mu.Lock()
_, kept := a.outbounds["busy-node"]
a.mu.Unlock()
if !kept {
t.Fatalf("the busiest outbound was evicted while single-hit entries survived")
}
}
// TestRetentionDisabledStillMeansUnlimited: the new caps must obey the existing
// master switch, or a deployment that deliberately asked for unlimited retention
// would silently get a bounded map instead.
func TestRetentionDisabledStillMeansUnlimited(t *testing.T) {
a := New(nil, nil, Config{RingSize: -1, TimelineMinutes: -1, MaxDomains: -1, RetentionDisabled: true})
for i := 0; i < maxOutboundKeys+50; i++ {
a.handleEvent(dnsEvent("a.example", fmt.Sprintf("resolver-%d", i), fmt.Sprintf("node-%d", i)))
}
a.mu.Lock()
outs, servers := len(a.outbounds), len(a.servers)
a.mu.Unlock()
if outs <= maxOutboundKeys {
t.Fatalf("RetentionDisabled still pruned the outbound map (%d entries)", outs)
}
if servers <= maxServerKeys {
t.Fatalf("RetentionDisabled still pruned the server map (%d entries)", servers)
}
if d := a.Snapshot().Dropped; d != (DropStats{}) {
t.Fatalf("RetentionDisabled reported drops: %+v", d)
}
}
// TestTopKMatchesAFullSort pins the refactor that shrank Snapshot's critical
// section: the top-K selections must produce exactly what the old
// materialise-everything-and-sort produced, order included. A projection that is
// merely "close" would reorder the panel's charts at random.
func TestTopKMatchesAFullSort(t *testing.T) {
rng := rand.New(rand.NewSource(20260726))
for trial := 0; trial < 200; trial++ {
n := rng.Intn(60)
m := make(map[string]uint64, n)
for i := 0; i < n; i++ {
m[fmt.Sprintf("name-%d", rng.Intn(80))] = uint64(rng.Intn(5))
}
for _, k := range []int{1, 3, 25, 1000} {
got := mapToServerStats(m, k)
want := make([]ServerStat, 0, len(m))
for name, c := range m {
want = append(want, ServerStat{Name: name, Count: c})
}
sort.Slice(want, func(i, j int) bool {
if want[i].Count != want[j].Count {
return want[i].Count > want[j].Count
}
return want[i].Name < want[j].Name
})
if k > 0 && len(want) > k {
want = want[:k]
}
if len(got) != len(want) {
t.Fatalf("trial %d k=%d: got %d rows, want %d", trial, k, len(got), len(want))
}
for i := range got {
if got[i] != want[i] {
t.Fatalf("trial %d k=%d row %d: got %+v, want %+v", trial, k, i, got[i], want[i])
}
}
}
}
}
// TestSnapshotResolvesNamesOutsideTheLock is item 4. Snapshot used to resolve every
// device's DHCP hostname while holding the aggregator mutex, and that resolution can
// read /tmp/dhcp.leases. While the mutex is held, handleEvent and handleConnEvent are
// parked — and the 64-slot buffers feeding them drop on overflow with no counter and
// no log, so the cost of a slow read is silently missing query-log rows.
func TestSnapshotResolvesNamesOutsideTheLock(t *testing.T) {
a := New(nil, nil)
// One real device row, so Snapshot has a name to resolve.
a.mu.Lock()
a.deviceDomains["192.168.1.50"] = map[string]uint64{"example.com": 3}
a.mu.Unlock()
var once sync.Once
entered := make(chan struct{})
release := make(chan struct{})
orig := parseLeases
parseLeases = func(path string) map[string]string {
once.Do(func() { close(entered) })
<-release
return map[string]string{}
}
t.Cleanup(func() { parseLeases = orig })
snapDone := make(chan struct{})
go func() {
a.Snapshot()
close(snapDone)
}()
select {
case <-entered:
case <-time.After(5 * time.Second):
close(release)
t.Fatalf("Snapshot never resolved a device name")
}
// Snapshot is now parked inside the lease read. The DNS hot path must be free.
// (A non-LAN client label needs no lease lookup of its own, so this measures the
// aggregator mutex and nothing else.)
folded := make(chan struct{})
go func() {
a.handleEvent(dnstrack.QueryEvent{
Domain: "hot.example", QueryType: 1,
Client: netip.MustParseAddr("198.51.100.7"),
})
close(folded)
}()
select {
case <-folded:
case <-time.After(5 * time.Second):
close(release)
t.Fatalf("handleEvent blocked behind Snapshot's lease-file read — the read is still under a.mu")
}
close(release)
<-snapDone
}
// TestSnapshotDeviceDomainsStillNamedAndOrdered guards the output of the same
// refactor: moving the work out of the lock must not change a single field.
func TestSnapshotDeviceDomainsStillNamedAndOrdered(t *testing.T) {
orig := parseLeases
parseLeases = func(path string) map[string]string {
return map[string]string{"192.168.1.50": "laptop"}
}
t.Cleanup(func() { parseLeases = orig })
a := New(nil, nil)
a.mu.Lock()
a.deviceDomains["192.168.1.50"] = map[string]uint64{"a.example": 1}
a.deviceDomains["192.168.1.60"] = map[string]uint64{"b.example": 9}
a.deviceDomains[routerDevice] = map[string]uint64{"c.example": 5}
a.mu.Unlock()
rows := a.Snapshot().DeviceDomains
if len(rows) != 3 {
t.Fatalf("got %d device rows, want 3", len(rows))
}
// Busiest first: .60 (9), router (5), .50 (1).
if rows[0].IP != "192.168.1.60" || rows[1].Name != routerDevice || rows[2].IP != "192.168.1.50" {
t.Fatalf("rows are not ordered by hit count: %+v", rows)
}
if rows[2].Name != "laptop" {
t.Fatalf("DHCP hostname was not resolved: %+v", rows[2])
}
if rows[1].IP != "" {
t.Fatalf("the router pseudo-row must carry no IP: %+v", rows[1])
}
}
// TestDeviceDomainDropsAreReported: the per-device caps existed already but evicted
// in silence. They are now counted, so an operator can tell "these are the top
// domains" from "these are the top domains we still had room for".
func TestDeviceDomainDropsAreReported(t *testing.T) {
a := New(nil, nil)
a.mu.Lock()
for i := 0; i < maxDomainsPerDevice+50; i++ {
a.foldDeviceDomainLocked("192.168.1.50", fmt.Sprintf("d%d.example", i))
}
// Overflow the client cap too.
for i := 0; i < maxDeviceClients+10; i++ {
a.foldDeviceDomainLocked(fmt.Sprintf("10.0.0.%d", i%256)+fmt.Sprintf("-%d", i), "x.example")
}
a.mu.Unlock()
d := a.Snapshot().Dropped
if d.DeviceDomains == 0 {
t.Fatalf("a device blew past the %d-domain cap and nothing was reported", maxDomainsPerDevice)
}
if d.DeviceClients == 0 {
t.Fatalf("clients were refused at the %d-client cap and nothing was reported", maxDeviceClients)
}
}
+40 -3
View File
@@ -21,12 +21,30 @@ var leasesPath = "/tmp/dhcp.leases"
// leaseCache resolves a client IP to its DHCP hostname, caching the parsed lease
// table and refreshing it at most once per ttl. Concurrency-safe.
//
// # Why the TTL refresh is asynchronous
//
// name() is called from paths that hold the AGGREGATOR mutex — handleEvent's
// deviceLabel (the DNS hot path), leaseName, and Snapshot's per-device projection.
// The cache has its own lock, so there is no deadlock, but a refresh is an
// os.ReadFile of /tmp/dhcp.leases, and doing that inline meant a blocking disk
// syscall ran while a.mu was held: every DNS event and every connection event in
// the daemon parked behind it, and the bounded 64-slot event buffers they feed
// from drop on overflow WITHOUT a counter. A file read on tmpfs is fast — until
// the one time it isn't, and then the cost is silently missing query-log rows.
//
// So only the COLD load is synchronous (once per process, and in production it
// happens on the poll loop's first tick, off any lock). Every later refresh runs
// in its own goroutine while the caller is served from the previous table, which
// is at most ttl stale — a DHCP hostname that is 30 seconds out of date is not a
// defect, a stalled DNS pipeline is.
type leaseCache struct {
ttl time.Duration
mu sync.Mutex
byIP map[string]string
fetched time.Time
loading bool // a refresh goroutine is in flight
}
func newLeaseCache() *leaseCache {
@@ -35,15 +53,20 @@ func newLeaseCache() *leaseCache {
// name returns a friendly device name for ip: the DHCP hostname when known, else
// the bare IP (never empty). Best-effort — a missing/uparseable lease file just
// yields the IP.
// yields the IP. Never blocks on disk except on the very first call.
func (c *leaseCache) name(ip string) string {
if ip == "" {
return ip
}
c.mu.Lock()
if c.byIP == nil || time.Since(c.fetched) > c.ttl {
switch {
case c.byIP == nil:
// Cold: nothing to serve, so this one read is unavoidable.
c.byIP = parseLeases(leasesPath)
c.fetched = time.Now()
case time.Since(c.fetched) > c.ttl && !c.loading:
c.loading = true
go c.refresh()
}
host := c.byIP[ip]
c.mu.Unlock()
@@ -53,6 +76,17 @@ func (c *leaseCache) name(ip string) string {
return ip
}
// refresh re-reads the lease table off any caller's lock and publishes it.
func (c *leaseCache) refresh() {
table := parseLeases(leasesPath)
now := time.Now()
c.mu.Lock()
c.byIP = table
c.fetched = now
c.loading = false
c.mu.Unlock()
}
// lanNetCache decides whether a connection SOURCE IP is a LAN client. It caches the
// LAN-side subnets parsed from /etc/config/network (via shater/devices.LANNets, which
// reuses the same UCI network parser the device-discovery page uses) and refreshes them
@@ -118,7 +152,10 @@ func (c *lanNetCache) isLAN(addr netip.Addr) bool {
// parseLeases reads the dnsmasq lease table into an ip->hostname map. Any read or
// parse failure yields an empty map (callers fall back to the raw IP).
func parseLeases(path string) map[string]string {
//
// A package var so a test can make the read block and prove that no caller performs
// it while holding the aggregator mutex; production never reassigns it.
var parseLeases = func(path string) map[string]string {
out := map[string]string{}
data, err := os.ReadFile(path)
if err != nil {
+106 -36
View File
@@ -75,67 +75,137 @@ func classifyCounter(name string) (display, kind string) {
}
// --- projection / sorting helpers ------------------------------------------
//
// Every projection below runs while the aggregator mutex is held, and while it is
// held the DNS and connection folding paths are parked. Their event buffers are 64
// slots deep and drop on overflow without a counter, so time spent here is paid in
// missing query-log rows and undercounted stats — invisibly.
//
// That is why these are top-K SELECTIONS, not "materialise everything, sort it,
// slice off 25". The domain and host maps hold up to 5000 entries and the answer is
// 25 of them; the old shape allocated a 5000-element slice and ran a 5000-element
// sort (~61k comparisons) under the lock to throw 99.5% of it away. topKInsert keeps
// only the K survivors and rejects a candidate that cannot beat the weakest of them
// in a single comparison, which is what almost every candidate does — so the work
// drops to about one comparison per map entry and the allocation to K elements.
//
// The ORDER is unchanged: same comparator, same "highest count first, name as
// tiebreak", so the output is byte-identical to what the sort produced.
// topKInsert inserts item into buf, a descending-ordered buffer holding at most k
// items under the strict weak ordering `before`. It returns the updated buffer.
// k <= 0 means unbounded, which degrades to insertion sort — callers with large
// inputs must pass a positive k.
func topKInsert[T any](buf []T, k int, item T, before func(a, b T) bool) []T {
if k > 0 && len(buf) >= k && !before(item, buf[len(buf)-1]) {
return buf // cannot displace even the weakest survivor: one comparison, done
}
i := sort.Search(len(buf), func(j int) bool { return before(item, buf[j]) })
if k > 0 && len(buf) >= k {
copy(buf[i+1:], buf[i:len(buf)-1]) // drop the weakest, shift right
buf[i] = item
return buf
}
buf = append(buf, item)
copy(buf[i+1:], buf[i:len(buf)-1])
buf[i] = item
return buf
}
// topKBuf allocates a buffer sized for a top-k selection over n candidates.
func topKBuf[T any](k, n int) []T {
if k > 0 && k < n {
n = k
}
return make([]T, 0, n)
}
// topDomainsLocked returns the n highest-count domains, most-popular first, with a
// stable tiebreak on the domain name. Caller holds the aggregator mutex.
func topDomainsLocked(m map[string]*domainAgg, n int) []DomainStat {
all := make([]DomainStat, 0, len(m))
for dom, ag := range m {
all = append(all, DomainStat{Domain: dom, Count: ag.count, Blocked: ag.blocked})
}
sort.Slice(all, func(i, j int) bool {
if all[i].Count != all[j].Count {
return all[i].Count > all[j].Count
before := func(a, b DomainStat) bool {
if a.Count != b.Count {
return a.Count > b.Count
}
return all[i].Domain < all[j].Domain
})
if n > 0 && len(all) > n {
all = all[:n]
return a.Domain < b.Domain
}
return all
out := topKBuf[DomainStat](n, len(m))
for dom, ag := range m {
out = topKInsert(out, n, DomainStat{Domain: dom, Count: ag.count, Blocked: ag.blocked}, before)
}
return out
}
// topHostsLocked returns the n highest-count destination hosts, most-popular first,
// with a stable tiebreak on the host key. Caller holds the aggregator mutex.
func topHostsLocked(m map[string]*hostAgg, n int) []HostStat {
all := make([]HostStat, 0, len(m))
before := func(a, b HostStat) bool {
if a.Count != b.Count {
return a.Count > b.Count
}
return a.Host < b.Host
}
out := topKBuf[HostStat](n, len(m))
for _, h := range m {
all = append(all, HostStat{
out = topKInsert(out, n, HostStat{
Host: h.host,
IP: h.ip,
Network: h.network,
Proto: h.proto,
Count: h.count,
Bytes: h.bytes,
})
}, before)
}
sort.Slice(all, func(i, j int) bool {
if all[i].Count != all[j].Count {
return all[i].Count > all[j].Count
}
return all[i].Host < all[j].Host
})
if n > 0 && len(all) > n {
all = all[:n]
}
return all
return out
}
// mapToServerStats projects a name->count map into a count-sorted slice.
func mapToServerStats(m map[string]uint64) []ServerStat {
out := make([]ServerStat, 0, len(m))
for name, c := range m {
out = append(out, ServerStat{Name: name, Count: c})
}
sort.Slice(out, func(i, j int) bool {
if out[i].Count != out[j].Count {
return out[i].Count > out[j].Count
// mapToServerStats projects a name->count map into a count-sorted slice, keeping the
// n highest counts (n <= 0 keeps everything).
//
// The limit is not cosmetic. a.outbounds is keyed by OUTBOUND TAG — a node name from
// the subscription — so it is one of the two maps whose key space a provider chooses,
// and this projection handed the whole thing to the panel on every /api/stats poll.
// The panel already renders only the top 10 (Insights.tsx), so nothing on screen
// changes; what changes is that a 2000-entry map stops being marshalled into JSON
// several times a minute.
func mapToServerStats(m map[string]uint64, n int) []ServerStat {
before := func(a, b ServerStat) bool {
if a.Count != b.Count {
return a.Count > b.Count
}
return out[i].Name < out[j].Name
})
return a.Name < b.Name
}
out := topKBuf[ServerStat](n, len(m))
for name, c := range m {
out = topKInsert(out, n, ServerStat{Name: name, Count: c}, before)
}
return out
}
// pruneCountMapLocked drops the lowest-count entries of a name->count map until keep
// remain, and reports how many it dropped. Same amortised strategy as
// pruneDomainsLocked: run only on overflow, so the sort is paid once per (max-keep)
// inserts. Caller holds the aggregator mutex.
func pruneCountMapLocked(m map[string]uint64, keep int) int {
if len(m) <= keep {
return 0
}
type kv struct {
name string
count uint64
}
all := make([]kv, 0, len(m))
for name, c := range m {
all = append(all, kv{name, c})
}
sort.Slice(all, func(i, j int) bool { return all[i].count < all[j].count })
drop := len(all) - keep
for i := 0; i < drop; i++ {
delete(m, all[i].name)
}
return drop
}
func sortDevicesByBytes(d []DeviceStat) {
sort.Slice(d, func(i, j int) bool {
if d[i].Bytes != d[j].Bytes {
+180 -43
View File
@@ -79,6 +79,40 @@ const (
maxDomainsPerDevice = 400 // per-client domain-map cap before it is pruned
keepDomainsPerDevice = 200 // per-client prune target (kept by count)
topDomainsPerDevice = 15 // domains Snapshot returns per device (highest count first)
// Caps for the two per-tag DNS maps. These were the last aggregates with no
// bound of any kind: every distinct DNSServer and every distinct outbound tag
// ever seen stayed in the map for the life of the process.
//
// a.servers is keyed by RESOLVER tag, which comes from the config — a handful
// on any real box. maxServerKeys is therefore a backstop against churn (a tag
// renamed on each of a few thousand applies), not against volume; 256 is ~25x
// the largest resolver set this generator emits, and at ~80 B/entry the full
// table is 20 KB.
maxServerKeys = 256
keepServerKeys = 128
// a.outbounds is keyed by OUTBOUND tag — on this box a node NAME chosen by the
// subscription provider, plus the derived per-group and per-chain copies. With
// ~380 nodes a live generation is on the order of 1000 distinct tags, and the
// provider renames nodes on the daily refresh, so this map grew by a whole
// generation per day forever. 2048 is two live generations, enough that a
// rename-day overlap never evicts a tag still in use; at ~120 B/entry (tag
// strings run 30-50 B) the full table is under 250 KB of a 512 MB box.
maxOutboundKeys = 2048
keepOutboundKeys = 1024
// topServersN is how many rows Snapshot returns for Servers / Outbounds. Mirrors
// topDomainsN; the panel renders the top 10 of either (Insights.tsx), so 25 is
// already generous.
topServersN = 25
// dropLogInterval throttles the "aggregate pruned" notices. Every prune is
// counted in Snapshot.Dropped and so is always visible in the API; the log line
// exists for an operator who is watching logread, and once per aggregate per
// five minutes is enough to make a standing condition obvious without becoming
// the flood the log sink then has to collapse.
dropLogInterval = 5 * time.Minute
)
// Config holds the memory-retention tuning knobs, sourced from model.Globals (Task B,
@@ -277,6 +311,35 @@ type ServerStat struct {
Count uint64 `json:"count"`
}
// DropStats reports what the retention caps have EVICTED since the daemon started —
// one running total per bounded aggregate.
//
// It exists because eviction that nobody can see is indistinguishable from data that
// never existed. Every one of these maps is bounded on purpose, and a bound that is
// working is fine; a bound that is CONSTANTLY working means the box is seeing more
// distinct domains / hosts / clients / outbound tags than the cap was sized for, and
// the numbers in TopDomains, TopHosts and DeviceDomains are then a sample rather than
// a total. That is a materially different claim, and the operator is entitled to
// know which one they are reading.
//
// A zero DropStats — the normal state — says every aggregate above is complete.
type DropStats struct {
// Domains / Hosts count entries dropped from the network-wide domain and
// destination-host maps by the maxDomains cap (lowest-count entries go first).
Domains uint64 `json:"domains"`
Hosts uint64 `json:"hosts"`
// DeviceDomains counts per-device domain entries dropped by maxDomainsPerDevice,
// summed over all devices. DeviceClients counts whole CLIENTS never tracked at
// all because maxDeviceClients was already reached (a spoofed-source flood, or a
// genuinely huge LAN).
DeviceDomains uint64 `json:"device_domains"`
DeviceClients uint64 `json:"device_clients"`
// Servers / Outbounds count entries dropped from the per-resolver and
// per-outbound-tag DNS maps (maxServerKeys / maxOutboundKeys).
Servers uint64 `json:"servers"`
Outbounds uint64 `json:"outbounds"`
}
// NodeHealthStat is one node's REAL health as measured by the engine's internal
// urltest probing (feedback #3 — replaces the panel's fake "enabled-count/total").
// A node's tag equals its node name (the generator emits every leaf node with
@@ -390,6 +453,11 @@ type Snapshot struct {
// Recent is a small tail of the query log embedded for convenience; the full
// log is served by GET /api/stats/log (RecentQueries).
Recent []LogEntry `json:"recent"`
// Dropped reports what the retention caps have evicted since start — see
// DropStats. All zeros (the normal state) means every aggregate above is a
// complete total rather than a sample. Additive field: a client that ignores it
// sees exactly the pre-existing contract.
Dropped DropStats `json:"dropped"`
}
// Normalized returns a copy of the snapshot in which EVERY slice field is non-nil,
@@ -521,6 +589,13 @@ type Aggregator struct {
keepDomains int
retentionDisabled bool
// dropped counts what the caps above have evicted, surfaced as Snapshot.Dropped.
// Guarded by mu like every aggregate it describes. dropLogged remembers when each
// kind last produced a log line so a standing overflow is reported at
// dropLogInterval instead of on every prune.
dropped DropStats
dropLogged map[string]time.Time
// which manager each consume loop is currently attached to (pointer identity is how
// Resubscribe detects a box-swap that replaced the manager).
curManager *dnstrack.Manager
@@ -603,6 +678,7 @@ func newAggregatorShared(eng *engine.Engine, logger log.ContextLogger, cfg ...Co
outbounds: make(map[string]uint64),
deviceDomains: make(map[string]map[string]uint64),
hosts: make(map[string]*hostAgg),
dropLogged: make(map[string]time.Time),
timelineMinutes: tlMin,
timelineUnlimited: tlUnlimited,
maxDomains: maxDom,
@@ -859,13 +935,26 @@ func (a *Aggregator) handleEvent(ev dnstrack.QueryEvent) {
}
}
// Per-server / per-outbound.
// Per-server / per-outbound. Both maps are keyed by tags that come from OUTSIDE
// this process (config resolver tags; subscription-chosen node names for the
// outbounds), so both are bounded the same insert-then-prune way as the domain
// map — and, like it, skipped when retention is deliberately disabled.
if ev.DNSServer != "" {
a.servers[ev.DNSServer]++
if !a.retentionDisabled && len(a.servers) > maxServerKeys {
n := pruneCountMapLocked(a.servers, keepServerKeys)
a.dropped.Servers += uint64(n)
a.noteDropLocked("servers", n, a.dropped.Servers, maxServerKeys)
}
}
for _, ob := range ev.Outbound {
if ob != "" {
a.outbounds[ob]++
if !a.retentionDisabled && len(a.outbounds) > maxOutboundKeys {
n := pruneCountMapLocked(a.outbounds, keepOutboundKeys)
a.dropped.Outbounds += uint64(n)
a.noteDropLocked("outbounds", n, a.dropped.Outbounds, maxOutboundKeys)
}
}
}
@@ -945,6 +1034,32 @@ func (a *Aggregator) pruneDomainsLocked() {
for i := 0; i < drop; i++ {
delete(a.domains, all[i].domain)
}
a.dropped.Domains += uint64(drop)
a.noteDropLocked("domains", drop, a.dropped.Domains, a.maxDomains)
}
// noteDropLocked records an eviction in the log, throttled per kind.
//
// Eviction is never silent here — the running totals go out on every /api/stats in
// Snapshot.Dropped — but the API is a pull, and an operator staring at logread while
// the numbers look wrong is exactly the person who needs to be told that TopDomains
// is a sample and not a total. Throttling to dropLogInterval per kind keeps a
// standing overflow visible without turning it into the flood logsink then has to
// collapse. Caller holds mu.
func (a *Aggregator) noteDropLocked(kind string, n int, total uint64, limit int) {
if n <= 0 || a.log == nil {
return
}
now := time.Now()
if last, ok := a.dropLogged[kind]; ok && now.Sub(last) < dropLogInterval {
return
}
a.dropLogged[kind] = now
// Deliberately "dropped", not "evicted": most kinds here evict the lowest-count
// entries already held, but device-clients drops the NEW arrival instead. One
// word that is true of both beats a precise one that is false half the time.
a.log.Warn("stats: ", kind, " at its ", limit, "-entry cap — dropped ", n,
" (", total, " total since start); this aggregate is a sample, not a total")
}
// --- connection consume loop (per-device domains) ---------------------------
@@ -1206,6 +1321,8 @@ func (a *Aggregator) pruneHostsLocked() {
for i := 0; i < drop; i++ {
delete(a.hosts, all[i].key)
}
a.dropped.Hosts += uint64(drop)
a.noteDropLocked("hosts", drop, a.dropped.Hosts, a.maxDomains)
}
// leaseName resolves a source IP to its DHCP hostname, returning "" (not the bare IP)
@@ -1255,6 +1372,8 @@ func (a *Aggregator) foldDeviceDomainLocked(ip, domain string) {
// the outer map unbounded). When full, existing clients keep updating but new
// ones are dropped. Skipped entirely when retention is disabled.
if !a.retentionDisabled && len(a.deviceDomains) >= maxDeviceClients {
a.dropped.DeviceClients++
a.noteDropLocked("device-clients", 1, a.dropped.DeviceClients, maxDeviceClients)
return
}
dm = make(map[string]uint64)
@@ -1262,27 +1381,19 @@ func (a *Aggregator) foldDeviceDomainLocked(ip, domain string) {
}
dm[domain]++
if !a.retentionDisabled && len(dm) > maxDomainsPerDevice {
pruneDeviceDomainsLocked(dm, keepDomainsPerDevice)
n := pruneDeviceDomainsLocked(dm, keepDomainsPerDevice)
a.dropped.DeviceDomains += uint64(n)
a.noteDropLocked("device-domains", n, a.dropped.DeviceDomains, maxDomainsPerDevice)
}
}
// pruneDeviceDomainsLocked trims one device's domain map back to keep entries, dropping
// the lowest-count domains (same amortised strategy as pruneDomainsLocked). Caller
// holds mu.
func pruneDeviceDomainsLocked(dm map[string]uint64, keep int) {
type kv struct {
domain string
count uint64
}
all := make([]kv, 0, len(dm))
for dom, c := range dm {
all = append(all, kv{dom, c})
}
sort.Slice(all, func(i, j int) bool { return all[i].count < all[j].count })
drop := len(all) - keep
for i := 0; i < drop; i++ {
delete(dm, all[i].domain)
}
// It returns how many entries it dropped, so the caller can account for them in
// Snapshot.Dropped.
func pruneDeviceDomainsLocked(dm map[string]uint64, keep int) int {
return pruneCountMapLocked(dm, keep)
}
// --- nft poll loop ----------------------------------------------------------
@@ -1403,8 +1514,6 @@ func (a *Aggregator) Snapshot() Snapshot {
nodeHealth := a.collectNodeHealth()
a.mu.Lock()
defer a.mu.Unlock()
snap := Snapshot{
Backend: a.backend,
EngineUp: a.eng != nil && a.eng.Running(),
@@ -1415,13 +1524,23 @@ func (a *Aggregator) Snapshot() Snapshot {
DeviceDomains: a.deviceDomainsLocked(),
TopHosts: topHostsLocked(a.hosts, topHostsN),
RuleTraffic: append([]RuleStat(nil), a.ruleTraffic...),
Servers: mapToServerStats(a.servers),
Outbounds: mapToServerStats(a.outbounds),
Servers: mapToServerStats(a.servers, topServersN),
Outbounds: mapToServerStats(a.outbounds, topServersN),
NodeHealth: nodeHealth,
Recent: a.recentLocked(20),
Dropped: a.dropped,
}
if !a.lastPoll.IsZero() {
snap.UpdatedAt = a.lastPoll.Format(time.RFC3339)
lastPoll := a.lastPoll
a.mu.Unlock()
// Naming and ordering happen OUTSIDE the lock. snap.DeviceDomains is already a
// private copy at this point, and a.leases.name may miss its cache — a disk read
// under a.mu parks the DNS and connection folding paths, whose 64-slot event
// buffers then drop silently. See finishDeviceDomains.
a.finishDeviceDomains(snap.DeviceDomains)
if !lastPoll.IsZero() {
snap.UpdatedAt = lastPoll.Format(time.RFC3339)
} else {
snap.UpdatedAt = time.Now().Format(time.RFC3339)
}
@@ -1432,10 +1551,14 @@ func (a *Aggregator) Snapshot() Snapshot {
}
// deviceDomainsLocked projects the per-device domain map into the serializable
// []DeviceDomainStat, resolving each source IP to a DHCP hostname and returning each
// device's top domains (highest count first). Devices are ordered by total hit count
// (busiest first), tie-broken by IP for stability. Caller holds mu. (leaseCache has its
// own lock, so name() here does not deadlock against a.mu.)
// []DeviceDomainStat: one row per tracked source IP carrying that device's top
// domains. Caller holds mu.
//
// It does the map walk and NOTHING else. Resolving DHCP hostnames and ordering the
// rows are deliberately left to finishDeviceDomains, which runs unlocked: the row
// slice is a private copy the moment this returns, so neither step needs the
// aggregator mutex, and both used to be paid with the DNS hot path stopped behind
// them — the name lookup potentially including an os.ReadFile of /tmp/dhcp.leases.
func (a *Aggregator) deviceDomainsLocked() []DeviceDomainStat {
out := make([]DeviceDomainStat, 0, len(a.deviceDomains))
for ip, dm := range a.deviceDomains {
@@ -1447,38 +1570,52 @@ func (a *Aggregator) deviceDomainsLocked() []DeviceDomainStat {
stat.Name = routerDevice
} else {
stat.IP = ip
stat.Name = a.leases.name(ip)
}
out = append(out, stat)
}
sort.Slice(out, func(i, j int) bool {
ti, tj := totalDomainCount(out[i].Domains), totalDomainCount(out[j].Domains)
return out
}
// finishDeviceDomains resolves each row's DHCP hostname and orders the rows by total
// hit count (busiest first, tie-broken by IP for stability) — the second half of
// deviceDomainsLocked, run WITHOUT the aggregator mutex. The output is identical to
// what the single locked pass produced.
func (a *Aggregator) finishDeviceDomains(rows []DeviceDomainStat) {
for i := range rows {
if rows[i].IP != "" {
rows[i].Name = a.leases.name(rows[i].IP)
}
}
sort.Slice(rows, func(i, j int) bool {
ti, tj := totalDomainCount(rows[i].Domains), totalDomainCount(rows[j].Domains)
if ti != tj {
return ti > tj
}
return out[i].IP < out[j].IP
return rows[i].IP < rows[j].IP
})
return out
}
// topDeviceDomains returns one device's n highest-count domains, most-popular first
// with a stable name tiebreak. Note the returned Count is the (possibly-capped) live
// count for domains that survived any prune.
// This is the single hottest projection in Snapshot: it runs once per tracked
// device, up to maxDeviceClients (512) of them, over maps of up to
// maxDomainsPerDevice (400) entries, all under the aggregator mutex. Materialising
// and fully sorting 400 entries per device to keep 15 was ~1.8M comparisons and
// ~200k slice elements of allocation with the DNS hot path stopped; the top-K
// selection rejects a candidate in one comparison and allocates 15 slots.
func topDeviceDomains(dm map[string]uint64, n int) []DomainCount {
all := make([]DomainCount, 0, len(dm))
for dom, c := range dm {
all = append(all, DomainCount{Domain: dom, Count: c})
}
sort.Slice(all, func(i, j int) bool {
if all[i].Count != all[j].Count {
return all[i].Count > all[j].Count
before := func(a, b DomainCount) bool {
if a.Count != b.Count {
return a.Count > b.Count
}
return all[i].Domain < all[j].Domain
})
if n > 0 && len(all) > n {
all = all[:n]
return a.Domain < b.Domain
}
return all
out := topKBuf[DomainCount](n, len(dm))
for dom, c := range dm {
out = topKInsert(out, n, DomainCount{Domain: dom, Count: c}, before)
}
return out
}
// totalDomainCount sums a device's returned domain counts (used only to order devices).
+5
View File
@@ -87,22 +87,27 @@ func (c *Client) DialContext(ctx context.Context) (net.Conn, error) {
request.Header.Set("Upgrade", "websocket")
err = request.Write(conn)
if err != nil {
conn.Close()
return nil, err
}
bufReader := std_bufio.NewReader(conn)
response, err := http.ReadResponse(bufReader, request)
if err != nil {
conn.Close()
return nil, err
}
if response.StatusCode != 101 ||
!strings.EqualFold(response.Header.Get("Connection"), "upgrade") ||
!strings.EqualFold(response.Header.Get("Upgrade"), "websocket") {
conn.Close()
response.Body.Close()
return nil, E.New("v2ray-http-upgrade: unexpected status: ", response.Status)
}
if bufReader.Buffered() > 0 {
buffer := buf.NewSize(bufReader.Buffered())
_, err = buffer.ReadFullFrom(bufReader, buffer.Len())
if err != nil {
conn.Close()
return nil, err
}
conn = bufio.NewCachedConn(conn, buffer)
@@ -0,0 +1,124 @@
package v2rayhttpupgrade
import (
"context"
"net"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/sagernet/sing-box/option"
M "github.com/sagernet/sing/common/metadata"
"github.com/stretchr/testify/require"
)
// A failed http-upgrade handshake used to return the error while leaving the
// dialed TCP connection open: nothing in the caller chain owns a conn that was
// never returned. On a router that keeps ~380 nodes under a health checker, one
// leaked descriptor per failed handshake is a slow death. Upstream 0f1763877.
type trackedConn struct {
net.Conn
closed atomic.Bool
}
func (c *trackedConn) Close() error {
c.closed.Store(true)
return c.Conn.Close()
}
// dialRecorder dials the real (test) listener and remembers every conn handed
// out, so the test can assert on the exact conn the client was given.
type dialRecorder struct {
access sync.Mutex
conns []*trackedConn
}
func (d *dialRecorder) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := new(net.Dialer).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *dialRecorder) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, net.ErrClosed
}
func (d *dialRecorder) only(t *testing.T) *trackedConn {
t.Helper()
d.access.Lock()
defer d.access.Unlock()
require.Len(t, d.conns, 1, "client must have dialed exactly once")
return d.conns[0]
}
// serveOnce accepts one connection, drains the request and writes raw response
// bytes back.
func serveOnce(t *testing.T, response string) net.Listener {
t.Helper()
listener, err := net.Listen("tcp", "127.0.0.1:0")
require.NoError(t, err)
t.Cleanup(func() { listener.Close() })
go func() {
conn, acceptErr := listener.Accept()
if acceptErr != nil {
return
}
defer conn.Close()
buffer := make([]byte, 4096)
conn.SetReadDeadline(time.Now().Add(5 * time.Second))
conn.Read(buffer)
if response != "" {
conn.Write([]byte(response))
}
}()
return listener
}
func newTestClient(t *testing.T, dialer *dialRecorder, listener net.Listener) *Client {
t.Helper()
client, err := NewClient(
context.Background(),
dialer,
M.ParseSocksaddr(listener.Addr().String()),
option.V2RayHTTPUpgradeOptions{Path: "/"},
nil,
)
require.NoError(t, err)
return client
}
func TestClientClosesConnOnUnexpectedStatus(t *testing.T) {
t.Parallel()
listener := serveOnce(t, "HTTP/1.1 403 Forbidden\r\nContent-Length: 0\r\n\r\n")
dialer := &dialRecorder{}
client := newTestClient(t, dialer, listener)
conn, err := client.DialContext(context.Background())
require.Error(t, err)
require.Nil(t, conn)
require.True(t, dialer.only(t).closed.Load(),
"the dialed conn must be closed when the upgrade is refused")
}
func TestClientClosesConnOnResponseFailure(t *testing.T) {
t.Parallel()
// The server hangs up without answering: http.ReadResponse fails.
listener := serveOnce(t, "")
dialer := &dialRecorder{}
client := newTestClient(t, dialer, listener)
conn, err := client.DialContext(context.Background())
require.Error(t, err)
require.Nil(t, conn)
require.True(t, dialer.only(t).closed.Load(),
"the dialed conn must be closed when the response cannot be read")
}
+6
View File
@@ -78,6 +78,12 @@ func (c *Client) offerNew() (*quic.Conn, error) {
packetConn.Close()
return nil, err
}
// quic-go does not take ownership of the packet conn passed to Dial:
// when the connection ends it only stops reading.
go func() {
<-quicConn.Context().Done()
packetConn.Close()
}()
c.conn.Store(quicConn)
c.rawConn = udpConn
return quicConn, nil
+203
View File
@@ -0,0 +1,203 @@
//go:build with_quic
package v2rayquic
import (
"context"
"net"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/sing-box/common/tls"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
qtls "github.com/sagernet/sing-quic"
M "github.com/sagernet/sing/common/metadata"
"github.com/stretchr/testify/require"
)
// quic-go does not take ownership of the packet conn handed to Dial: when the
// QUIC connection ends it only stops reading from it. offerNew() then dials a
// fresh one and overwrites c.rawConn, so the previous UDP socket was leaked for
// the lifetime of the process — one per reconnect, on a box with 512 MB and a
// health checker that reconnects constantly. Upstream 7067276170.
const testALPN = "shater-test"
type trackedConn struct {
net.Conn
closed atomic.Bool
}
func (c *trackedConn) Close() error {
c.closed.Store(true)
return c.Conn.Close()
}
type dialRecorder struct {
access sync.Mutex
conns []*trackedConn
}
func (d *dialRecorder) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := new(net.Dialer).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *dialRecorder) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, net.ErrClosed
}
func (d *dialRecorder) only(t *testing.T) *trackedConn {
t.Helper()
d.access.Lock()
defer d.access.Unlock()
require.Len(t, d.conns, 1, "client must have dialed exactly once")
return d.conns[0]
}
// serveQUIC brings up a real QUIC listener on localhost with a self-signed
// certificate, runs handler for every accepted connection, and returns its
// address.
func serveQUIC(t *testing.T, handler func(conn *quic.Conn)) M.Socksaddr {
t.Helper()
ctx := context.Background()
logger := log.NewNOPFactory().NewLogger("test")
keyPem, certificatePem, err := tls.GenerateCertificate(nil, nil, time.Now, "localhost", time.Now().Add(time.Hour))
require.NoError(t, err)
serverTLSConfig, err := tls.NewSTDServer(ctx, logger, option.InboundTLSOptions{
Enabled: true,
ServerName: "localhost",
ALPN: []string{testALPN},
Certificate: []string{string(certificatePem)},
Key: []string{string(keyPem)},
})
require.NoError(t, err)
require.NoError(t, serverTLSConfig.Start())
t.Cleanup(func() { serverTLSConfig.Close() })
packetConn, err := net.ListenPacket("udp", "127.0.0.1:0")
require.NoError(t, err)
t.Cleanup(func() { packetConn.Close() })
listener, err := qtls.Listen(packetConn, serverTLSConfig, &quic.Config{})
require.NoError(t, err)
t.Cleanup(func() { listener.Close() })
go func() {
for {
conn, acceptErr := listener.Accept(ctx)
if acceptErr != nil {
return
}
go handler(conn)
}
}()
return M.ParseSocksaddr(packetConn.LocalAddr().String())
}
func newTestClient(t *testing.T, dialer *dialRecorder, serverAddr M.Socksaddr) *Client {
t.Helper()
clientTLSConfig, err := tls.NewSTDClient(context.Background(), log.NewNOPFactory().NewLogger("test"), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
ALPN: []string{testALPN},
})
require.NoError(t, err)
transport, err := NewClient(context.Background(), dialer, serverAddr, option.V2RayQUICOptions{}, clientTLSConfig)
require.NoError(t, err)
client, isClient := transport.(*Client)
require.True(t, isClient)
return client
}
func TestClientClosesPacketConnWhenConnectionEnds(t *testing.T) {
// Drop the connection right after the handshake: this is the server-side
// reset / idle timeout the client must survive without leaking its socket.
serverAddr := serveQUIC(t, func(conn *quic.Conn) {
conn.CloseWithError(0, "bye")
})
dialer := &dialRecorder{}
client := newTestClient(t, dialer, serverAddr)
t.Cleanup(func() { client.Close() })
quicConn, err := client.offer()
require.NoError(t, err)
require.NotNil(t, quicConn)
// The server hangs up; the client keeps its Client alive (a health checker
// would simply dial again later).
select {
case <-quicConn.Context().Done():
case <-time.After(5 * time.Second):
t.Fatal("server never closed the QUIC connection")
}
tracked := dialer.only(t)
require.Eventually(t, tracked.closed.Load, 5*time.Second, 10*time.Millisecond,
"the UDP socket behind a dead QUIC connection must be closed, not leaked until Client.Close()")
}
// quic-go's Stream.Close() does not unblock a Write parked on flow control. The
// writer goroutine (for us: the copy loop of a proxied connection) then survives
// its own connection forever. Closing has to push the write deadline into the
// past as well.
func TestStreamCloseUnblocksBlockedWrite(t *testing.T) {
serverIsDone := make(chan struct{})
t.Cleanup(func() { close(serverIsDone) })
// Accept the stream but never read from it, so the client's writes fill the
// receive window and block.
serverAddr := serveQUIC(t, func(conn *quic.Conn) {
_, err := conn.AcceptStream(context.Background())
if err != nil {
return
}
<-serverIsDone
})
dialer := &dialRecorder{}
client := newTestClient(t, dialer, serverAddr)
t.Cleanup(func() { client.Close() })
stream, err := client.DialContext(context.Background())
require.NoError(t, err)
writeDone := make(chan error, 1)
go func() {
payload := make([]byte, 64*1024)
// 32 MiB is far past any quic-go receive window, so this must park.
for range 512 {
_, writeErr := stream.Write(payload)
if writeErr != nil {
writeDone <- writeErr
return
}
}
writeDone <- nil
}()
select {
case err = <-writeDone:
t.Fatal("the write never blocked, the test proves nothing: ", err)
case <-time.After(time.Second):
}
require.NoError(t, stream.Close())
select {
case <-writeDone:
case <-time.After(5 * time.Second):
t.Fatal("Write stayed blocked after Close: the writer goroutine is leaked")
}
}
+4
View File
@@ -2,6 +2,7 @@ package v2rayquic
import (
"net"
"time"
"github.com/sagernet/quic-go"
qtls "github.com/sagernet/sing-quic"
@@ -37,5 +38,8 @@ func (s *StreamWrapper) Upstream() any {
func (s *StreamWrapper) Close() error {
s.CancelRead(0)
s.Stream.Close()
// quic-go's Stream.Close does not unblock a Write blocked on flow control,
// but a past write deadline does; buffered data and the FIN are unaffected.
s.Stream.SetWriteDeadline(time.Now())
return nil
}
+2
View File
@@ -93,12 +93,14 @@ func (c *Client) dialContext(ctx context.Context, requestURL *url.URL, headers h
reader, _, err := ws.Dialer{Header: ws.HandshakeHeaderHTTP(headers), Protocols: protocols}.Upgrade(deadlineConn, requestURL)
deadlineConn.SetDeadline(time.Time{})
if err != nil {
conn.Close()
return nil, err
}
if reader != nil {
buffer := buf.NewSize(reader.Buffered())
_, err = buffer.ReadFullFrom(reader, buffer.Len())
if err != nil {
conn.Close()
return nil, err
}
conn = bufio.NewCachedConn(conn, buffer)
@@ -0,0 +1,122 @@
package v2raywebsocket
import (
"context"
"net"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/sagernet/sing-box/option"
M "github.com/sagernet/sing/common/metadata"
"github.com/stretchr/testify/require"
)
// A refused websocket upgrade used to return the error while leaving the dialed
// TCP connection open — the conn was never returned, so nobody else could close
// it. With ~380 nodes under a health checker every failing node leaks one
// descriptor per probe. Upstream 0f1763877.
type trackedConn struct {
net.Conn
closed atomic.Bool
}
func (c *trackedConn) Close() error {
c.closed.Store(true)
return c.Conn.Close()
}
type dialRecorder struct {
access sync.Mutex
conns []*trackedConn
}
func (d *dialRecorder) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := new(net.Dialer).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *dialRecorder) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, net.ErrClosed
}
func (d *dialRecorder) only(t *testing.T) *trackedConn {
t.Helper()
d.access.Lock()
defer d.access.Unlock()
require.Len(t, d.conns, 1, "client must have dialed exactly once")
return d.conns[0]
}
func serveOnce(t *testing.T, response string) net.Listener {
t.Helper()
listener, err := net.Listen("tcp", "127.0.0.1:0")
require.NoError(t, err)
t.Cleanup(func() { listener.Close() })
go func() {
conn, acceptErr := listener.Accept()
if acceptErr != nil {
return
}
defer conn.Close()
buffer := make([]byte, 4096)
conn.SetReadDeadline(time.Now().Add(5 * time.Second))
conn.Read(buffer)
if response != "" {
conn.Write([]byte(response))
}
}()
return listener
}
func newTestClient(t *testing.T, dialer *dialRecorder, listener net.Listener) *Client {
t.Helper()
transport, err := NewClient(
context.Background(),
dialer,
M.ParseSocksaddr(listener.Addr().String()),
option.V2RayWebsocketOptions{Path: "/"},
nil,
)
require.NoError(t, err)
client, isClient := transport.(*Client)
require.True(t, isClient)
return client
}
func TestClientClosesConnOnRefusedUpgrade(t *testing.T) {
t.Parallel()
listener := serveOnce(t, "HTTP/1.1 403 Forbidden\r\nContent-Length: 0\r\n\r\n")
dialer := &dialRecorder{}
client := newTestClient(t, dialer, listener)
conn, err := client.DialContext(context.Background())
require.Error(t, err)
require.Nil(t, conn)
require.True(t, dialer.only(t).closed.Load(),
"the dialed conn must be closed when the websocket upgrade is refused")
}
func TestClientClosesConnOnHandshakeEOF(t *testing.T) {
t.Parallel()
// The server hangs up mid-handshake: ws.Dialer.Upgrade fails on read.
listener := serveOnce(t, "")
dialer := &dialRecorder{}
client := newTestClient(t, dialer, listener)
conn, err := client.DialContext(context.Background())
require.Error(t, err)
require.Nil(t, conn)
require.True(t, dialer.only(t).closed.Load(),
"the dialed conn must be closed when the handshake cannot complete")
}
+1
View File
@@ -187,6 +187,7 @@ func (c *EarlyWebsocketConn) writeRequest(content []byte) error {
if len(lateData) > 0 {
_, err = conn.Write(lateData)
if err != nil {
conn.Close()
return err
}
}