Compare commits

..
Author SHA1 Message Date
omarandClaude Opus 5 1267d20fb8 docs: drop the L3 handoff note — it is merged, and it said to
test / go + panel tests (push) Successful in 8m33s
release / test gate (push) Successful in 8m8s
release / apk aarch64_cortex-a53 (push) Successful in 6m33s
release / apk x86_64 (push) Successful in 3m45s
release / release apk (push) Successful in 8s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 20:00:46 +03:00
omarandClaude Opus 5 35f697ed08 docs(openwrt): say why mtu_fix is inert instead of claiming an MTU we no longer set
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:52:39 +03:00
omarandClaude Opus 5 d0fb6befb1 fix(l3): the l3-in MTU is not a tunnel budget — 1420 was a forgery generator
shater-l3 was created at 1420, the WireGuard payload budget, copied one
layer too far out. It bought nothing: what actually goes into the tunnel
is sized by sing-tun's forwardToPort against Port.PortMTU(), which
already fragments to the outbound MTU without DF and answers a
well-formed `fragmentation needed` quoting it with DF. All 1420 did was
make the KERNEL split every packet above 1392 bytes of payload on its
way into the device -- and a fragment is the one thing sing-tun will not
judge. Dispatch returns on parsed.fragment before calling JudgeFlow, the
fragments reach the gVisor stack, it reassembles them, and the ICMP
forwarder's installFlow demands an unspecified port address that a
WireGuard endpoint never has. So it declined and answered the echo
itself. `ping -s 1392` honest, `ping -s 1393` a lie, and only for the
outbounds the feature exists for.

65535 rather than merely "large": no IP datagram can exceed it, so the
kernel cannot fragment at this device for any packet ever. Anything
smaller leaves a band open and re-opens the class. It is also sing-box's
own default TUN MTU on Linux.

Memory was measured, not argued. Three paired runs of the integration
test under -test.memprofilerate=1 allocate 5.41/5.48/5.47 MB at 65535
against 5.76/5.46/5.70 MB at 1420, and a -diff_base profile puts every
difference in netlink interface enumeration. Nothing in the read path
scales with the MTU: gVisor reads through fdbased.BufConfig, which
sing-tun pins to one 65535-byte view regardless. I predicted a ~1.8 MB
saving from GSO switching off above 49152 and was wrong -- protocol/tun
turns GSO back on at StartStateStart whenever a FlowOutbound exists, so
the GRO scaffolding is there at both values. The corrected reasoning is
in the constant's comment so the next reader does not redo the mistake.

The integration test now reads the MTU back off the real kernel device,
which is the assertion the value exists for: a kernel that clamped it
would restore the forgery without changing a generated byte.

D25's KNOWN HOLE block is replaced with what is genuinely left. Chiefly:
a big non-DF ping does not start WORKING, it starts failing HONESTLY --
classifyReturn declines fragments on the way back too, so the packet
really leaves, the far host really answers, and the reply is not NAT'd
home. And a client that fragments on the wire itself is still uncovered;
that is the nft carve-out's job, with a warning that conntrack defrag
may reassemble in prerouting and leave such a rule unable to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:50:08 +03:00
omarandClaude Opus 5 81c96019b5 fix(panel): let a routing rule say ICMP, instead of calling one broken
The Proto picker was a closed list of the two transports and the ten
sniffed L7 labels, and anything else drew "<value> — never matches".
The engine now routes ICMP by rule (Rule.Proto accepts icmp, icmpv4,
icmpv6), so a working ping rule was rendered as a dead one and could not
be created here at all — the operator had to hand-edit /etc/config/shater
and then watch the panel call the result broken.

Adds a third group, "Layer 3". All three spellings are offered: they are
not synonyms — icmpv4/icmpv6 pin the rule's ip_version — so hiding the
narrowing would both strand a capability outside the UI and silently
widen such a rule the first time someone edited it here.

The doc comment no longer claims the list IS generate/route.go's
sniffedProtocols; only the middle group is. ICMP goes to the emitted
rule's `network`, never to `protocol`, which is the whole reason it never
matched as a sniffed label.

An unknown value is still kept and offered as written, but the
never-matches flag is now judged on the lower-cased value, the way the
engine judges it — a hand-written `ICMP` is a live rule, not an inert one.

Verified: npm run build clean (tsc --noEmit + vite build); an icmp rule
added through the panel renders as a plain "PROTO icmp" chip; no
horizontal overflow at 360px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:26:07 +03:00
omarandClaude Opus 5 baed8ff8f2 fix(model): fwmark_base 0x7f routes the engine's own traffic into its own TUN
The panel offers fwmark_base and table_base as free hex fields under
"Advanced" and nothing has ever checked them. What makes that more than a
footgun is that the derived values are invisible from the number typed: the
L3 mark is base+0x80, so 0x7f lands it exactly on 0xff — the loop-guard mark
the engine stamps on its OWN traffic — and `ip rule fwmark 0xff lookup 8200`
then captures everything the engine sends and routes it into the engine's
TUN. The router loses the internet the moment l3_tunnel is switched on, for
a reason nothing on screen connects to a collapsed section. fwmark_base 0xff
had produced the same failure since long before the L3 offset existed.

table_base is worse and got the same treatment: its derived values can land
on the kernel's own table ids, and teardown does `ip route flush table <n>`.
It is count-sensitive (egress #i uses base+0x10+i), so the check takes the
egresses rather than living in ValidateGlobals.

Written as "derive every value this layout produces, then look for
duplicates and reserved ids" rather than as a blacklist, so a future offset
is covered by construction. The layout constants are duplicated from
netplane (the import only runs one way) and pinned by netplane's
TestMarkLayoutConstantsLockstep.

Warn-only, like every check in this file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 f80fb4dd1b fix(netplane): give every mark-driven table a floor, and check the L3 pair
Two halves of the same omission.

1. A fwmark lookup that finds an empty table does not fail — it falls
   through to main. Every mark-driven table now gets an `unreachable
   default` at the maximum metric: it loses to any real default route while
   one exists, it has no device so the kernel never garbage-collects it, and
   it turns "lookup failed, try main" into "lookup succeeded: unreachable".
   The fallthrough stops depending on somebody reading a warning at the
   moment an interface goes down. Deliberately not gated on the kill-switch:
   that switch decides whether traffic may escape the tunnel, while an egress
   binding is a statement about WHICH UPLINK, and silently substituting a
   different one is not what "fail open" was meant to permit.

   RoutingPresent's "does this table have a default route" test is tightened
   in the same breath, or the floor would answer it and turn the safety net
   into a blindfold.

2. RoutingPresent had never heard of addL3Routing. This is the same defect
   its own comment describes as already caught twice ("a presence check must
   cover everything its Apply counterpart installs"), committed a third time
   — and its trigger needs no interface to go down: editing a node URI
   restarts the engine, the kernel destroys shater-l3 and takes `default dev
   shater-l3 table 8200` with it, the rendered nft text is unchanged, so the
   fast-path skipped ApplyRouting forever and LAN ping stayed dead until
   someone restarted the daemon.

TestRoutingPresentSeesL3Table, TestEgressTableGetsFailClosedFloor and
TestEveryStampedMarkIsRoutedAndVerified all fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:54 +03:00
omarandClaude Opus 5 b71b793681 fix(netplane): a mark says where a packet was sent, not where it went
The forward chain let untunnelable-egress traffic past the kill-switch on
the strength of its fwmark alone. `ip rule fwmark X lookup N` does not
deliver the packet to table N, it delivers the LOOKUP there — and a lookup
that finds nothing falls through to main. So when the egress interface goes
down and the kernel garbage-collects its default route, every non-TCP/UDP
packet from the LAN is still stamped, still accepted here (above the
fail-closed drop), and leaves out the plain WAN with the router's real
address. Nothing we render changes, so no apply runs and nothing notices.

Ordinary egress traffic never had this hole: the engine binds those sockets
to the device, and a dead device fails the socket. The untunnelable-egress
path is made of nothing but a mark, so the accept now carries the second
opinion instead — `meta mark X oifname "dev"`, strictly narrower than either
half, true only when the routing did what the mark asked. The comment being
replaced argued correctly that oifname ALONE would be too loose, then drew
from that the conclusion that oifname should be dropped rather than added.

Same conjunction in the holding plane, where it is theory (that plane stamps
nothing) but where a bare mark accept has no business sitting.

Also folds the egress device resolution into one EgressDevice(), because the
binding and model.ValidateUntunnelableEgress had already drifted: the
validator trimmed the interface name and the binding did not, so `option
interface '   '` gave a panel saying "the option is ignored" over a data
plane that was marking packets for a table nobody built.

TestUntunnelableEgressAcceptIsBoundToItsDevice and
TestUntunnelableEgressResolutionLockstep fail on the code they replace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:23:32 +03:00
omarandClaude Opus 5 61c87ad1d9 fix(l3): guard the ICMP honest-drop at PreMatch, not inside the walk
The drop that keeps a ping from reading as tunnelled lived in
preMatchFlow, overriding the pre-declared continueResult. That covered
every exit of THAT function and none of the walk above it: the
prepareMatchMetadata error return (which arrived later, with the shared
metadata refactor), the sniff bail-outs, and the default: arm of the
rule-action switch all returned PreMatchContinue on their own.
adapter.JudgeFlow maps Continue to tun.ActionAccept, and sing-tun answers
Accept by rewriting Echo into EchoReply itself -- the exact forgery this
delta exists to remove. Narrow paths, but paths.

PreMatch is now a funnel over the renamed preMatch walk, so the guard
sits on the single return value and cannot be outgrown by a new exit.
PreMatchBypass joins the drop: sing-tun implements ActionBypass on the
nfqueue plane only, so on the TUN path it lands in the same default: arm
as Accept and forges too.

Every ICMP case has an explicit TCP/UDP twin; the JudgeFlow mapping
table is pinned outright, including the one fix that must NOT be made
there -- refusing ActionFlow for a port whose address is not unspecified
would drop every ping through WireGuard/AWG, because the forward
dispatcher and the ICMP forwarder share that function with identical
arguments and only the latter needs an unspecified address.

That leaves a real hole open, now named in D25 rather than papered over:
a FRAGMENTED echo to a WireGuard/AWG outbound is still answered by the
router. The dispatcher returns before asking for a verdict at all when
the packet is a fragment, and the reassembled packet reaches the ICMP
forwarder, whose installFlow demands the unspecified address a WireGuard
endpoint never has. The two fixes that would close it both live outside
pre-match and are written down; the Consequence paragraph is scoped
until one lands.

The stack comment in generate/inbound.go repeated the "only gvisor
really forwards ICMP" argument that D25 itself retracts -- both stacks
run the same ForwardDispatcher first. Brought in line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 19:21:29 +03:00
omar 4dee508e12 fix(route): let a rule say "icmp", and say when saying it is a lie
`icmp` fell through ruleMatchers' proto switch into RawDefaultRule.Protocol —
the SNIFFED-L7 field, compared against what the sniffers labelled a connection.
Nothing ever labels a flow "icmp" (PreMatch skips the sniff action for an ICMP
flow outright), so the rule was structurally valid and permanently dead. That
made the whole L3 ingress unusable on a real config: with no way to write "ICMP
goes here", every ping fell to the catch-all, which resolves to the chain's last
hop — a group of VLESS nodes that cannot carry layer 3 at all.

icmp is a NETWORK. NetworkItem.Match is a map lookup over metadata.Network, and
adapter.JudgeFlow sets that to N.NetworkICMP for BOTH ICMPv4 and ICMPv6 (one
case covers both protocol numbers), so there is exactly one network value and it
covers both families. `icmpv4`/`icmpv6` narrow that same network with an
ip_version item instead of inventing a second one: metadata.IPVersion comes from
the destination address, and an ICMPv6 packet always has an IPv6 destination —
no false positives, no false negatives.

An ICMP rule that cannot fire is not a dead setting: ICMP has no fall-through,
so route.preMatchFlow DROPS it. Four ways to get that silently are now reported:
l3_tunnel off (nothing enters the engine at all), icmpv6 with ipv6 off (neither
the nft mark nor the TUN address exists), a port matcher next to it (JudgeFlow
zeroes both ports), and a target that cannot carry layer 3 — decidable from the
model, because the capability is fixed by the outbound TYPE: only wireguard/AWG
endpoints and the direct outbound behind direct/interface egresses declare
N.NetworkICMP. A mixed group gets its own text (the answer follows group.Now()),
`block` gets none (dropping the ping IS the policy), and an unresolved target
gets none either (ruleKillFallback already said the louder thing).

Wording stays clear of shater/apply's criticalMarkers on purpose: a failed ping
is fail-CLOSED, and a cosmetic alarm is how the real one stops being read.
2026-07-26 19:18:06 +03:00
omarandClaude Opus 5 76da5134ef test(gate): the two tests that need a kernel may not skip in silence
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.

Two halves.

Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.

Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.

The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 18:50:46 +03:00
omar 4c630c9a13 docs: handoff note for the L3 branch
Transient, to be deleted when omp/work merges. Everything meant to outlive the
merge is already in D25/D26 and the lx changelog; this file is the part that is
only useful while the branch is still a branch — the verification commands, the
testbed recipe, what was proven on hardware and what was not, and the six files
that will conflict on rebase.
2026-07-26 18:31:59 +03:00
omar d8dbefcd07 docs: record the AWG site-to-site path as declined, not impossible
D26's "no port-like selector" line disposes of NAT-based forwarding and nothing
else, and read alone it says "impossible" — which is false and would be
re-derived at the cost of another research pass. The endpoint is protocol-blind
in both directions, so ESP could ride it untouched with the client's own source
address and no NAT whatsoever. That was declined for two reasons worth naming:
lx-owned code in the forward hot path, and a server-side AllowedIPs prerequisite
that turns a router option into a deployment contract.
2026-07-26 18:31:59 +03:00
omar 974208fc05 docs: record the kernel egress, and retract the reason D25 gave for the ceiling
D26 writes down where the engine's boundary actually is, because the intuitive
answer is wrong and someone will look for it again: the WG/AWG forward path
never consults gVisor in either direction, so the limit is sing-tun's
ForwardDispatcher — its parser and its port-shaped NAT — and the kernel egress
was chosen because it clears that limit without a line of new hot-path code, not
because userspace "cannot". Tailscale documents the same boundary for their
userspace mode and is quoted as corroboration, with the caveat that ours sits at
the dispatcher rather than the stack.

D25 said two things that do not survive checking, and both are corrected in
place rather than left for the next reader to trip over. It blamed the netstack
for the ICMP-echo ceiling; that was the dispatcher. And it called `stack: gvisor`
mandatory because the system stack fakes ping — the system stack runs the very
same dispatcher first and only forges an echo for packets the dispatcher
declined, so gvisor is a deliberate choice (already linked via with_wireguard,
and the combination the integration test exercises), not a necessity.

The operator note says what the option buys and refuses to call an egress a
tunnel on its own say-so: with a WireGuard device it is one, with a second WAN
the destination sees that uplink's address. It also says what the option does
not fix — multicast IPTV stays broken — and that IPsec through NAT-T is ordinary
UDP that never needed any of this.
2026-07-26 18:31:59 +03:00
omar 2eb71e8244 feat(netplane,model): hand the protocols the engine will not dispatch to the kernel
ESP, AH, GRE, IGMP and SCTP cannot enter the engine, and the reason is not the
one that looks obvious. A WireGuard or AmneziaWG endpoint forwards straight past
its gVisor stack — WritePackets reads the IP version and the destination address
and hands the raw bytes to the device, and the return path offers every
decrypted packet back before the stack sees it. WireGuard would carry ESP today
if anything handed it one. What refuses is sing-tun's ForwardDispatcher: its
parser recognises TCP, UDP and ICMP echo, and its NAT wants a port-shaped
selector that ESP, AH and GRE do not have. The retracted rationale is corrected
where it was written down, not quietly dropped.

So these protocols go to the kernel instead. untunnelable_egress names an
interface or tunnel egress; prerouting stamps that egress's OWN mark on
everything that is not TCP or UDP, and addEgressRouting has already bound that
mark to a table whose default route leaves via the device. Every protocol works
because nothing in the path has to understand any of them. No new mark, no new
table, no new code in the hot path.

Whether that is a tunnel depends on the device, and nothing here claims
otherwise: a WireGuard interface is one, a second WAN is a different uplink
whose real address the far end sees.

The wide `!= { tcp, udp }` filter is safe here and stays banned for the L3
ingress, for the same reason stated in both places: there the receiver is a
dispatcher that knows four protocols, here it is the kernel. ICMP is claimed by
the L3 ingress first when both are on. The local plane keeps its exclusions —
router-addressed traffic, private destinations, ICMPv6 ND/RA — and with IPv6 off
the marking is scoped to v4, because addEgressRouting installs no v6 rule then
and a marked v6 packet would fall into the main table.

An interface egress with an empty `interface` no longer resolves: IfaceDevice
defaults to br-lan, so it passed the binding while addEgressRouting skipped it —
mark set, no rule, straight past a closed kill switch and out the default WAN.
2026-07-26 18:31:59 +03:00
omar 668cccbf24 test(generate): the L3 device name is a singleton, so wait for the kernel to take it back
Both gated tests stand an engine up on shater-l3. Run together, the second met
`TUNSETIFF: device or resource busy` and failed for a reason that had nothing to
do with what it asserts — the first had closed its box and yielded while
unregister_netdevice was still catching up. Each passed alone, which is the
shape of a fixture bug that gets rediscovered rather than fixed.

The poll that already guarded the first test is now a shared helper both call.
It stays a poll rather than a sleep for the reason it always was: the removal is
usually immediate and a fixed wait would be either flaky or slow.
2026-07-26 18:31:59 +03:00
omar 4ea4585402 test(generate): pin that ping through an interface egress is real, and byedpi's is not
An interface egress is a direct outbound carrying BindInterface and a routing
mark, and direct builds its ICMP port from the very same dialer control — so
ping routed at that egress leaves through that device, marked, like every other
packet bound to it. Nothing said so. Both halves of that sentence are one
`common.Cast[*dialer.DefaultDialer]` away from being false: if the dialer ever
stops being a DefaultDialer, icmpPort is nil, PreMatchFlow declines, and ping
through the egress degrades to a drop without a single generated byte changing.
The gated test asserts the live outbound, not the config, because that is where
the cast happens.

The failure the codegen half guards is worse than a broken ping: losing
BindInterface or the mark does not stop the echo, it sends it out the main table
over the plain WAN with the real address, which is the one thing an egress
exists to prevent.

byedpi is a SOCKS outbound and cannot be a tun.Port, so ICMP aimed at it is
dropped. That is the honest end of l3-honest-drop and it is pinned too, because
the alternative the TUN stack offers is a forged reply.
2026-07-26 18:31:59 +03:00
omar dc6d102473 docs: put a number on the second netstack, and say what it does not bound
Measured on a throwaway harness in a container: peak RSS of a process that
brought the engine up went from ~26 MB to ~28 MB with l3_tunnel on, three
paired runs. It is x86_64, idle, with an empty ICMP NAT table, so it stays
listed as unverified for the router — an indicative figure is more useful than
silence only if it says loudly what it is not.
2026-07-26 18:31:58 +03:00
omar 683afc0a47 docs: record how ping got through the tunnel, and where it stops
D25 writes down the reasoning that is expensive to reconstruct: why a TUN rather
than TPROXY, why the interface is its own with auto_route off, why gvisor is
mandatory rather than preferred, and why the ceiling is ICMP echo — a boundary
in sing-tun's flow parser and gVisor's protocol set, not an unfinished edge of
ours. It also records what carries layer 3 and what does not, that masque could
and does not, and the two things still unproven: the live-router path end to
end, and what a second gVisor NIC costs in memory on the hardware.

D17 gains one line: its claim that TPROXY cannot carry ICMP is still true, and
is no longer the end of the story.
2026-07-26 18:31:58 +03:00
omar 2c3e20512e feat(openwrt): let fw4 know the L3 tunnel device before it exists
Both nft tables run and a drop in either one wins, so our forward accept for
shater-l3 decides nothing on its own: fw4 sees a device in no zone and drops the
forward, and the feature fails with exactly the symptom it was built to fix —
ping does not work, and nothing says why.

The zone names the device directly rather than a network. fw4 resolves a zone's
networks through netifd, and a proto-none interface for a device the daemon
creates is never up and contributes nothing, so list network would compile to an
empty device set. list device compiles to a plain iifname/oifname match that is
valid before the TUN exists and starts matching the moment shaterd creates it,
with no firewall reload at enable time.

It is seeded unconditionally, not gated on l3_tunnel: uci-defaults run once, and
a zone naming an absent device is inert. Gating it would mean the option could
be switched on and never take effect. The sections are named so a re-run is a
no-op instead of a second zone, and kmod-tun joins DEPENDS because /dev/net/tun
is not on a stock image.
2026-07-26 18:31:58 +03:00
omar 51b2f04672 feat(netplane,generate): carry LAN ping through the tunnel, on a TUN of its own
Kernel TPROXY needs a socket to hand a packet to, so it moves TCP and UDP and
nothing else. Everything else reached the forward chain and met the untunnelable
policy, whose best answer was "let it out with your real address" and whose
default was "drop it" — so on a stock install ping simply did not work, and the
setting that fixed it did so by leaking.

The engine has been able to do better for a while: sing-tun's ForwardDispatcher
does real ICMP forwarding with NAT on the echo id, and a WireGuard or AmneziaWG
endpoint is a tun.Port that carries the packet for real. What was missing was a
way in, because nothing on the router could hand it an IP packet.

l3_tunnel (opt-in, off by default) adds one: the generator emits an "l3-in" TUN
inbound and prerouting fwmarks LAN ICMP into it. The interface is its own and
auto_route is off, so the main routing table is never touched and the fwmark
plus addL3Routing's ip rule are the only entrance — the TPROXY plane is byte for
byte what it was. gvisor is not a preference: the system stack forges echo
replies locally, which is the very thing this is meant to end.

Only icmp and ipv6-icmp are ever marked, and only after the local plane is out
of the way — the router itself, private destinations, and ICMPv6 ND/RA, which
mean nothing off-link and take v6 down if one neighbour probe is tunnelled.
ESP, AH, GRE, IGMP and SCTP are deliberately left alone: sing-tun's parser and
gVisor's stack know no such protocol, so marking them would black-hole the
traffic while looking like a feature. They stay with the untunnelable policy,
which also keeps its say over what happens if the ip rule fails to install.

Ping and Windows tracert now cross the tunnel; IPv6 traceroute shows only the
destination, because the return path recognises TimeExceeded for v4 alone.
2026-07-26 18:31:58 +03:00
omar f190c8251e feat(lx): stop answering ping on behalf of a tunnel that never saw it
PreMatchContinue is not "fall back to the ordinary route" the way it is for TCP
and UDP. An ICMP flow has no ordinary route: the TUN stack takes the packet back
and answers the echo itself, swapping the addresses and writing a reply
(sing-tun stack_gvisor_icmp.go). So a ping routed to any outbound that cannot
carry layer 3 — every proxy protocol; only adapter.FlowOutbound can — came back
successful, and the operator read a working tunnel off a packet that was never
sent.

That is worse than the packet loss it replaced. Loss is a fault the operator can
see and chase; a forged reply is a fault that reports itself as health, and it
reports it on the one tool anyone reaches for first.

preMatchFlow now overrides continueResult once, at the top, for
N.NetworkICMP. One hunk covers every exit that used to fall through — no such
outbound, a group whose selection is gone, an outbound whose Network() omits
icmp, an outbound that is not a FlowOutbound — and keeps the diff to three lines
against a function upstream will keep editing. JudgeFlow carries the same
verdict in its !isPort branch, because FlowOutbound and tun.Port are separate
interfaces and drift between them must not reopen the forgery.

TCP and UDP are untouched, and the test pins that as hard as it pins the drop.
2026-07-26 18:31:58 +03:00
omarandClaude Opus 5 1945404eaa fix(armor): a reboot is not someone switching the product off
test / go + panel tests (push) Successful in 5m24s
release / test gate (push) Successful in 5m24s
release / apk aarch64_cortex-a53 (push) Successful in 3m9s
release / apk x86_64 (push) Successful in 3m9s
release / release apk (push) Successful in 8s
The boot armor never armed on the router it shipped to. procd runs the
K-links on the way down with the action `shutdown`, and stop_service
classified actions with an OPEN default:

    case $action in restart|reload) keep;; *) DISARM;; esac

`shutdown` matched nobody, fell into `*`, and deleted the arm token. The
mechanism erased itself at exactly the transition it exists for, so every
boot found nothing to load. Measured on the live router, one minute apart
across a reboot:

    13:28  /etc/shater/boot.nft present
    ----   reboot
    18s    at_S22: NO_TABLE  armor_file=NO_FILE

It did not fail every time, which is worse than failing always: on the way
down `rm` from this script raced a `SaveBootArmor` driven by the ifdown
hotplug storm, and whichever landed second won. Two reboots on the same box
an hour apart gave opposite outcomes.

Both lists are now positive and CLOSED. Only `stop` disarms; only
`restart`/`reload` hand off. An action nobody thought of changes nothing,
so the default now fails toward a boot that arms when it need not have --
recoverable in the second before the daemon applies, and still gated by
shater-armor's four state refusals. The old default failed toward the
plaintext window the feature was built to close.

Also closed, found while proving the above:

  * Every restart left the LAN in the clear for 80-90ms. The exit path was
    `Teardown(); armOnExit()`, and TeardownNft DELETES the table -- two nft
    transactions with no `inet shater` between them, leaving fw4's
    `lan -> wan ACCEPT` as the only policy. Every restart, every LuCI Save
    & Apply. TeardownExiting arms first under the apply lock and skips the
    delete iff a plane actually went in; RenderHoldNft is one `nft -f` that
    REPLACES the table, so the kernel never observes its absence.
    35k-sample instrument: 7 and 6 no-table hits before, 0 across three
    runs after.

  * SaveBootArmor fsynced the payload but not the directory, so a power cut
    could lose the rename that publishes it -- a boot with no armor and no
    error anywhere.

`stop` now also reads rc.d state, so a package transaction that stops the
service is not mistaken for a person switching it off. This one does not
reproduce on apk (it runs no pre-upgrade script and never calls prerm on an
upgrade; verified with apk adbdump and 245k samples across a real reinstall)
-- it is one returning opkg lane away from being live, and the removal case
is now stated rather than implicit.

Both new tests are mutation-checked: reverting the predicate fails naming
`shutdown`; reverting the teardown fails with `did [arm delete], want [arm]`.
initscript_test.go sources the SHIPPED shell and calls the real predicates
with every action procd uses -- a comment claiming `shutdown` was handled is
what shipped last time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 17:33:53 +03:00
omarandClaude Opus 5 6476722372 fix(panel): stop shipping a fabricated router in the binary
test / go + panel tests (push) Successful in 5m26s
release / test gate (push) Successful in 5m28s
release / apk aarch64_cortex-a53 (push) Successful in 6m7s
release / apk x86_64 (push) Successful in 3m5s
release / release apk (push) Successful in 7s
mock.ts was a static import and the mock switch was read from the query string at
runtime, so the bundle that ships inside the daemon carried a complete fictional
router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive,
without a single request to the daemon. The only tell was a line in the footer.
That is worse than any wrong number — there is no data at all and nothing says
so. It is out of the production bundle now, which is 21 kB smaller for it.

Unknown state stopped reading as good news in two more places. The kill-switch
tile treated an absent plane as armed, because the check was "not none" and
undefined satisfies it — the contract in the API types says the opposite. And the
apply page announced "daemon auto-rolled back" from its own timer, while the
daemon, seeing the state generation move, disarms and says it is NOT rolling back
in the log only.

Alerts moved to Settings. They are about the kill switch, apply failures, new
devices and subscription expiry, and they lived at the bottom of the DNS page,
while Settings mentioned them in prose with nothing to click.

Findings truncation is visible now: the notice that says how many were suppressed
arrives as info, and the attention list keeps only critical and warning, so past
fifty findings the operator saw forty-nine and no hint of the rest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:39 +03:00
omarandClaude Opus 5 4078334d85 fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.

They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.

Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.

The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.

And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.

The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:21 +03:00
omarandClaude Opus 5 0a6689b29e fix(quic,v2ray): close the sockets quic-go was never going to close
DialEarly with a packet conn the caller made sets a flag that means quic-go does
not own it: closing the transport only stops reading from the socket. Neither DNS
transport closed it. On the QUIC one it was closed on a failed handshake and
never on success, so every redial — idle timeout, retry error, engine reload —
left a UDP socket for the life of the process. On the HTTP/3 one the library
drives its own reconnects, so the leak compounds without anything in our code
looking wrong.

That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on
every reconnect without closing the previous one. Both are now owned by a watcher
tied to the connection's own context, so the socket lives exactly as long as the
connection does.

This matters more than it did last week: the shipped resolvers are DoH, and DNS
is intercepted by default now, so the whole network's query stream rides this
path on a router with 512 MB.

The same upstream commit fixes both halves. We had taken the v2ray half and not
the DNS one — the third time this session a paired fix arrived half-applied, and
the first of those cost a day of debugging. These two files are now byte-identical
to upstream so a rebase cannot reopen it.

Also from that family: websocket and httpupgrade leaked their conn on failed
handshakes, and a QUIC stream's Close did not release a blocked write.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:59 +03:00
omarandClaude Opus 5 ef22167b1a fix(apply): report the hold when the plane was armed by someone else
Booting with the armor loaded, or restarting through the handoff, left the status
saying the LAN was not being held while it was being dropped. Transient after a
successful apply, but permanent on the unreadable-config path — and there the
apply-failure alert words itself "traffic is NOT being blocked" at the exact
moment it is. That sends the operator to fix something that is not broken, past
the protection that is holding.

The table cannot be identified from here — netplane exposes no read-back and nft
does not keep comments — but identifying it is the wrong question. Holding does
not claim the holding plane is the object in the kernel; it claims the engine is
down and forwarded traffic is being dropped. A leftover full ruleset does that
too: with no engine socket the tproxy statement breaks its own rule before the
accept, so the packet reaches the forward chain unmarked and meets the primary
drop. What decides it is whether the last applied config was enabled and
fail-closed, which is exactly what the boot armor's presence already means.

So it is derived at read time rather than latched. A latch set from an inference
would have to be remembered in order to be cleared, which is the trap the active
flag already taught us. ArmHold also stops deferring to a table it cannot
inspect and installs its own render instead — the honest answer to "do not claim
a foreign table blindly" is to make it ours, and a fresh render beats a snapshot
that predates an interface rename.

Also closes the last of the detour transports: the subscription fetch took a
client and dropped it, and the exits that leak are the error ones, retried by
cron forever against a broken feed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:39 +03:00
omarandClaude Opus 5 cbda0fee0a fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has
already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its
predecessor, migrating the schema, building the engine and reading UCI — with a
UPX-compressed binary decompressing off flash first. Every boot therefore had a
window with no protection at all, landing exactly when Wi-Fi comes up and every
client reconnects. A restart, a reload or a package upgrade opened the same
window on purpose: Teardown does not consult the kill switch, and the init script
guarantees the interval is non-empty.

The holding plane is now persisted to /etc/shater/boot.nft on every apply and
loaded by a small service at 21, right after fw4 and netifd. Its presence is the
arm token: it exists only while the last applied config was enabled AND
fail-closed, and goes away the moment either stops being true. Writes are
content-gated — the cron reconcile runs a minute — and atomic, because the one
boot that reads this file is the boot after a power cut.

The service refuses to arm four ways so it can never brick a box, and its
enabled-check reads /etc/rc.d directly rather than asking rc.common, which would
take a blocking flock in the middle of boot. On exit the daemon re-arms only for
restart and reload, read from a snapshot of rc.common's action; anything else,
including an unknown one, degrades to a real stop that also disarms.

An unreadable config used to leave the router bare forever: the arm call sat in
the branch that requires a successful read, and nothing downstream could recover
it. It now arms from the same path.

A network nobody named was neither diverted nor blocked — the divert set is built
from inbounds and rule sources, and the same set scopes the fail-closed drops. It
is now enumerated from the interfaces whose firewall zone the operator forwards
to a WAN zone — their own statement that those clients reach the internet through
this box — and reported critically, by name, with both resolutions. Deliberately
not closed automatically: this router cannot know a guest SSID was meant to be
off the tunnel, and guessing is an outage. A device name that resolved to nothing
is reported the same way, for the same reason: there is no fail-closed action
available for a device we cannot name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:39:16 +03:00
omarandClaude Opus 5 7234817adb fix(panel): flush the log before serving it
test / go + panel tests (push) Successful in 4m59s
release / test gate (push) Successful in 4m58s
release / apk aarch64_cortex-a53 (push) Successful in 3m5s
release / apk x86_64 (push) Successful in 3m1s
release / release apk (push) Successful in 7s
Splitting the log sink made its writes asynchronous, so a download could miss
the last lines still in the queue — silently, with a successful response. Those
are the lines the operator came for: a log is downloaded to find out what just
happened.

The panel is handed a barrier, not the sink: a func() set once at startup, the
same shape as the reconfigure hook and the stats setter already in the tree. It
cannot write, reconfigure or close, so it stays a consumer, and nothing about
the sink's type reaches it.

The wait is bounded at the sink's own control budget and enforced on the panel
side, so a wedged writer cannot turn the download into the new place the daemon
gets stuck — the very thing the async split was for. Past the bound the handler
serves what is on disk. With no barrier installed the path behaves as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 07:27:42 +03:00
omarandClaude Opus 5 f1c36d6eea fix(panel): say when the engine is down, and ask before the irreversible
The header could not render "offline": it keyed on a field the daemon pinned to
true, so a dead engine behind a fail-closed plane showed a pulsing green lamp.
Health now needs both signals to agree before it reads as up, and a negative
from either is enough to say down — which is honest against the field that was
already honest, and stays honest now that the other one is too.

Deleting the last catch-all rule was described as "traffic will fall through to
the next rule" on the very row the page badges as the default route. What
happens instead is the kill switch: closed, the network loses the internet;
open, it leaves with the real address. The dialog now says which, by reading the
saved setting, and the toggle asks the same question — the generator only emits
enabled rules, so switching it off is the same event.

The master switch tore the whole plane down without a word, while deleting a
rule-set got a confirmation. Deleting a node or a resolver claimed to remove it
"from the config" without mentioning what still points at it, though the
reference finder was already there and used for renames.

Every Apply button armed the auto-rollback, and only one page said so. The
window is now recorded where all of them pass through, carried in a band under
the nav on every route, and persisted — so the countdown and the keep button
survive a reload, which is what made the window unconfirmable before. Overview's
Confirm button is gone rather than gated: Confirm cannot fail, so a permanently
live button could only ever report success.

Blocklists printed "filtering" from two config checkboxes without asking whether
the list had ever loaded — while the daemon grades a failed load critical. They
now show what the rule-set rows already showed, and say "not loaded — nothing
blocked" when that is the truth.

Also: the clock read UTC while every timestamp rendered in the browser's zone,
so the router appeared to have started in the future; the rule counter on
Overview counted saved rules rather than the ones in force, unlike the routing
page; and the hop badge counted the entry egress the rail below it does not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 07:27:28 +03:00
omarandClaude Opus 5 bb21ceb7f5 fix(lifecycle): a busy file is not a broken one, and three ways to lose state
Changing any stats knob on the persistent backend deleted months of history.
The replacement store was opened before the outgoing one was closed, so it hit
the first one's flock, timed out — and the open path treated ANY error as
corruption and unlinked the file. Unlink of an open file succeeds on Linux, so
the new ring opened an empty database while the panel was still told the
backend had not changed. The store now hands its resources over before asking
for them again, and deletion is gated on an allow-list of real corruption
signals; a busy, unreadable or read-only file degrades to the in-RAM ring and is
left alone.

The holder's reads were unguarded in a subtler way, caught only after the gate
failed twice: the accessor took the read lock, returned the pointer and released
it, so the call ran outside. A reader could hold a store the swap then closed and
be served its empty answer — an empty page presented as data. The accessor is
gone entirely, along with the possibility of handing out an unguarded reference.
Readers still do not block each other; the swap now waits out reads already in
flight, which is a page at most.

The urltest group published its chosen node through two plain fields written by
the prober and read on every dial and every panel poll — while the selector next
door does the same job atomically. They are one value now, so TCP and UDP can no
longer be read as a mismatched pair. Nothing had ever dialled through a group
while it was probing, which is why the detector had never seen it; a test now
does, and reproduces it deterministically against the old shape.

Close on a group whose ticker had already stopped returned before closing its
channel, and Touch would then arm a fresh loop nothing could stop. Reached by
pressing Test in the panel and applying a config within the next two minutes: the
orphan kept failing probes against a cancelled context and writing forged dead
verdicts into the board the live generation selects from. Close is now final.

The log sink held one mutex across a blocking write. Under procd stderr is a
pipe, so a reader that stopped draining wedged everything that logs — engine,
panel handlers, signal loop — while the process still answered a signal. It is
split: a front that assembles lines and a writer that owns the destinations,
joined by a bounded queue that drops and counts rather than blocking. Proven by
restoring the old shape: the package deadlocks for the full ten-minute timeout,
parked exactly where the field symptom said.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 07:27:06 +03:00
omarandClaude Opus 5 a8970b8ace fix(apply): stop the status from reporting a state the daemon is not in
Running was the constant true. The panel builds its header from it, so the
"offline" branch was unreachable code: with the engine dead and the LAN behind
a fail-closed hold, the operator saw a pulsing green lamp and, on the page
people open to fix things, "engine: running". The honest field sat beside it,
documented as the honest answer to are-we-proxying, and was read nowhere.

running now means shater is running: the daemon answered and its engine has a
started instance. active stays what it always was and is documented as such —
the "meant to be running" latch that gates hotplug and cron, not a health
signal. It is deliberately not cleared on hold, because the cron loop gates on
it and clearing it would switch off the reconcile that brings the engine back.

Two paths published nothing and so left the previous config's verdict standing
for as long as the fault lasted. A rollback with no snapshot re-applied the
engine and the plane and never touched the traffic verdict, so a router rolled
back to a direct default kept reporting the tunnel. And an apply that failed in
the netplane stage had already swapped the engine, then returned before every
publisher, so status described the config that was no longer running — and the
next reconcile, seeing an unchanged hash, failed the same way and published
nothing again. Both now publish, with an unknown verdict: after a no-snapshot
rollback the engine runs options this process does not hold, and guessing from
UCI would describe the config we rolled away from.

The severity classifier had drifted from the texts production emits. Markers
were compared case-sensitively against wording that had since changed, and the
entity pattern could not match a message beginning with an upper-case tag —
so a blocklist that failed to load graded as a warning while a typo in its URL
graded critical, and the panel's banner, which only lights for criticals, stayed
dark for the outage. RULESET-NOT-APPLIED and DNS-FILTER-NOT-APPLIED are now read
as the structural markers their producer documents them to be, so severity no
longer depends on wording at all. Five markers that matched no living text are
deleted; three protection-section texts drop to warning, because a blocklist
that is stale but still blocking lights the alarm on most reconciles behind a
flaky link, and an alarm that is always on is how the real one goes unread.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 07:26:40 +03:00
omarandClaude Opus 5 4996bc0984 fix(untunnelable): make block actually block
The block policy collapsed into direct whenever routing's final target was
direct — the common "tunnel only what is blocked, everything else direct"
shape. So ICMP, ESP, AH, GRE, IGMP and SCTP left with the client's real
address under the setting whose own field doc promises "nothing ever leaves
with the client's real IP", including a standing VPN on the real address,
which is exactly what the middle rung exists to separate out.

Both ends of the ladder now short-circuit before the plan is consulted and
neither may consult it: direct accepts everything, block emits no line at all
and lets the fail-closed drops the caller writes next do the work.

A rule scoped by source could also widen the other family: emit() skipped a
family whose destination list was empty but not one whose source list was, so
a rule carrying only IPv6 source prefixes rendered an IPv4 line with no
ip saddr clause — an accept for every IPv4 host on the LAN. The two halves now
read "scoped" the same way the catch-all collapse already did.

No destination plan is built for block at all now. It is the shipped default,
and a geoip-backed plan is ~159 000 prefixes pushed into kernel memory and the
ruleset text for a policy that cannot use them.

The operator-facing texts said IPTV works. It does not, on any of the three
rungs: inbound multicast is never matched by these rules and a client's
outbound multicast UDP dies at the fail-closed guard regardless. Saying
otherwise invited trading the ESP/GRE block away for nothing. What actually
stops working under block is stated instead, and precisely: raw ESP/AH and
GRE, but not IPsec through NAT or any UDP VPN, which are ordinary tunnelled
traffic.

TestOnlyPinnedAddressIsTunnelled is how this hid: it asserted, on the default
policy, that an exception line was emitted, and read that as the feature
working. It was block rendering direct. Its render assertions move to icmp,
where they mean something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 07:26:07 +03:00
omarandClaude Opus 5 a0de597d69 feat(dns): intercept by default, and bootstrap node addresses off the tunnel
test / go + panel tests (push) Successful in 4m56s
The posture was inverted. A client using the DHCP-supplied resolver — the router
itself — was NOT intercepted: dnsmasq answered and forwarded to the ISP in the
clear, so the filter, the blocklists, the per-device rules and BlockDoH were all
inert for exactly the clients that did nothing wrong. A client that hardcoded
8.8.8.8 to route around us WAS intercepted, by the catch-all. Meanwhile the
docs promised no DNS leaks. The default now matches the promise.

Turning it on crosses a threshold that was already dangerous for anyone with two
resolvers. Above one transport, a node's domain server address stops being
resolved by the transport directly and goes through the client DNS plane
instead — so a blocklist entry, a block_doh NXDOMAIN or any dns_rule can answer
your own node's hostname, and one sloppy line in an ad list stops being an ad
that got through and becomes a tunnel that never comes up.

So the fix is gated on having two or more transports, not on the intercept
toggle: resolver_default plus resolver_fallback always reached that threshold,
long before this change. When no endpoint_resolver is configured the plane now
carries a bootstrap server — the default resolver cloned with its detour
dropped, keeping its type, so a DoH default stays DoH and only the tunnel hop
goes. An explicit endpoint_resolver still wins.

This is not a restore of the previous behaviour and the comment says so: at one
transport the dialer used the default resolver WITH its detour, so a lone
DoH-through-the-tunnel resolver was already a bootstrap loop. It is strictly
better than what came before.

Existing installs keep whatever they set — the config file is a conffile and is
never replaced — and an explicit dns_intercept '0' survives the render-parse
round trip, which a default-true bool otherwise makes easy to lose.

The no-resolver warning stays, and no default resolver is shipped to silence it:
a placeholder would remove the sentence without moving a single query, and the
panel would then say a resolver was configured while nothing was filtered. Its
wording is corrected instead — .lan keeps working through the built-in local
transport, which the old text denied.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:54:15 +03:00
omarandClaude Opus 5 544da29863 fix(dns,netplane): close three paths that sent traffic out in the clear
A resolver whose detour no longer resolved fell back to "the default outbound",
which is not a default at all — it is a plain system socket. Every other place
in this generator fails such a reference closed, with an essay explaining why,
and wgdedup rewrites the very same field to block when it drops an endpoint. One
field, two opposite policies, and which one applied depended on whichever code
noticed the breakage first. A resolver detoured through a node the operator
switched off therefore handed the whole network's query stream to the ISP in the
clear, while the kill switch held the traffic itself.

It now fails closed, and the warning says what that means: the resolver answers
nothing, and if it is the default one, name resolution stops network-wide until
the target is restored. A dns_rule naming a missing resolver used to be dropped
whole, sending exactly the names the operator singled out to a resolver they did
not choose; it keeps its matchers and answers NXDOMAIN instead. Not a reject
action — one built in Go with an unset Method panics the engine at match time.

RoutingPresent never looked at per-egress rules or tables, and applyLocked skips
the whole routing stage on its word. So an egress table wiped by an ifdown was
never restored: the marked traffic fell through to main and left over the plain
WAN, permanently, with plane full and no warnings. It now verifies each binding
it installed, recording intent rather than outcome so a broken egress keeps the
plane reported absent and heals when the interface returns.

addEgressRouting discarded every ip error, so an egress that failed to install
reported success and the panel drew it green. Failures are now critical warnings
naming the egress, the device and what ip said — but still warnings, because
returning would abort the apply and punish the household for one bad uplink.

Also anchors the fwmark check: with a small fwmark_base the main mark is a
literal prefix of the first egress mark, so a substring match could answer "the
main rule is installed" while looking at an egress rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:53:53 +03:00
omarandClaude Opus 5 daaa0fda41 fix(wireguard): stop holding AmneziaWG down behind a WireGuard hop
The guard refused to start an AmneziaWG endpoint whose detour chain reached a
WireGuard one, and refused silently: not an error, just started=false, after
which every dial failed with "WireGuard is not ready yet". A selector hook went
further and suspended an already-working node the moment its group switched to a
WireGuard member.

It existed because AmneziaWG inside WireGuard hung the kernel on Android. We do
not ship Android, upstream dropped the guard once the cause was gone, and the
cure landed here yesterday — the ClientBind reserved-gate plus the submodule pin
that carries its twin. So the tree held both the cure and the prohibition on
using it, and the configuration simply did not come up while looking like a node
that "just does not work".

Also takes the two fixes that belong with it. ClientBind.conn was read on a
lock-free fast path and written under a mutex; upstream found that race with the
same end-to-end test we wrote yesterday, so we had taken one half of a pair
again. And the outer WireGuard UDP socket forced DF, unlike direct, hysteria and
tuic — with encapsulation the datagram regularly exceeds the path MTU and the
kernel drops it instead of fragmenting, a symptom indistinguishable from the bug
we spent yesterday on.

The race needed its own test: the existing e2e run did not flag it under -race
even at -count=15. Eight goroutines over both connect branches reproduce it
deterministically, naming the lock-free read and the guarded write.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:53:32 +03:00
omarandClaude Opus 5 56a276bcc1 ci: make the tests a gate instead of a decoration
The fork had a full suite and no CI that ran it. Upstream's test workflows
trigger on stable/testing/unstable; this repo only has main. And Gitea does not
read .github/workflows at all once .gitea/workflows exists, so those files were
decoration here. 115 of the 116 test files under shater/** had never executed in
CI even once, which is how TestDNSFilterRemoteBlocklistHTTPClient stayed red
across two published releases without anyone noticing.

The gate is a job inside release.yml that build-apk needs, because a separate
workflow cannot block another one. It runs the suite under the shipped tag set,
on Linux — 6 of 7 test files in transport/wireguard and 12 in shater/generate
compile only there or only under those tags, and those are exactly the files
covering AmneziaWG.

Three guards stop it from passing by running nothing, which is the failure this
whole change is about. The tag set may only ADD test files, never remove one.
Every package go list says has tests must appear as "ok <pkg>" in the output, so
a suite that collapses to "no test files" fails instead of passing. And the
panel run counts its test files first, because node --test exits 0 with "pass 0"
when the glob matches nothing.

The publish step used to exit 0 having published nothing: its assertions all
live inside a loop over artifacts, so an empty directory ran the body zero times
and reported success. It now counts what it published and fails on zero.

Verified by extracting the shipped step text and running it against stubs: empty
artifacts gives exit 0 before and exit 10 after; the rolling-release readback
still fires its own exit 14.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:53:14 +03:00
omarandClaude Opus 5 754bbcf1fa fix(submodule): point .gitmodules at the line the pin is actually on
The wireguard-go submodule is pinned to 7d15f33, which lives on lx-awg2-v005.
.gitmodules named `lx` — a separate line, 42 commits one way and 131 the other,
with no common recent history.

That is a loaded gun rather than a cosmetic mismatch. `lx` has no hasReserved()
gate in conn/bind_std.go at all, so a single `git submodule update --remote`
would move the pin there and silently restore the defect fixed yesterday: the
bind shreds the AmneziaWG magic header of every transport packet, handshakes
complete, no data moves, and no chain containing an AmneziaWG node carries
traffic. It would also drop the padding-overrun fix and the v0.0.5 re-graft.

Nothing about the checked-out tree changes — the pin is untouched. Only the
branch a --remote update would follow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 03:18:21 +03:00
omarandClaude Opus 5 63e6b709f8 fix(health): a chain blocked at a hop is dead on the board, unknown on the card
release / apk aarch64_cortex-a53 (push) Successful in 3m4s
release / apk x86_64 (push) Successful in 3m1s
release / release apk (push) Successful in 7s
The short-circuit left the chain's exit tag untested, because the exit is itself
a hop and every hop behind the break was rewritten that way. A stand run caught
it — the test lives in a file that does not compile on the dev host, so nothing
local could have.

That is not neutral silence. selectExcluding ranks untested ABOVE dead and says
so in its own comment: with no fresh-alive member, an untested one is a better
bet than a known-dead one. Leaving a provably broken path untested is therefore
a positive preference for it over a path we merely know is dead.

The two readings answer different questions and now differ on purpose. Is this
hop's own node alive — unknown behind a break, so the card keeps untested and
blocked_by. Can this chain carry traffic — known, no, because the hop in front
of it was probed and did not answer. The board carries that second answer, which
is the one selection, the freshness gate and the manual test all read.

The exit verdict is derived, not dialled: it records the consequence of a probe
that did happen one hop earlier, and it is re-derived every pass, so the moment
the blocker answers the walk reaches the exit again and the next verdict there is
a real measurement.

Also keeps a routed group warm. Its checker used to stop on the idle timeout and
nothing filled in behind it, so a rule that fires rarely would show untested
while being in force and pay a cold probe on the first real request. The gate
that adds this work answers false when it does not know — the mirror of the one
that withholds work, so plain sing-box keeps the lifecycle it always had.

And the tls-spoof suite now skips without tcpdump instead of failing sixteen
times: a missing tool is not measured, not broken. The same distinction this
commit is about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 02:48:26 +03:00
omarandClaude Opus 5 8612b0a9e9 fix(health): one dialler per target, and it is the group's own checker
A member of a chain hop wrapper was reachable by two probers: ours, from the
observatory plan, and sing-box's, from the urltest group the wrapper actually
is. Two independent readings of one node can disagree, and then neither can be
trusted — which is worse than the wasted dial.

The group's own checker is the right owner. A hop wrapper's members are the
per-chain copies, each carrying the previous hop as its detour, so that checker
already travels the chain prefix — the path the traffic takes. The plan now
records who dials each target and the observatory skips the ones a live checker
owns, keeping only what no group covers: node hops, the AmneziaWG endpoint,
selector members, and the members of groups that have been stood down.

The jobs stay in the plan rather than being deleted, and that is load-bearing:
the short-circuit reads the plan as the map of which tags measure which hop, so
deleting a urltest hop's members would erase that hop from the map and quietly
stop it blocking anything — on exactly the chains the feature exists for.

The short-circuit therefore moves to the group as well, through a ProbeGate the
engine implements: a scheduled check asks whether the path in front of it is up
before dialling, while an explicit check is never refused. Nothing is stored —
the gate recomputes from the live board every call — and Touch still arms the
ticker even while blocked, because a hop that refuses to tick has nothing left
to notice its own recovery. The gate answers yes whenever it does not know:
refusing on missing information is how a system talks itself into silence.

Two grounds now exist for a group not to probe and they must not be merged:
stood down means no rule reaches it at all, blocked means the path in front is
down right now. Both doc comments say so and name the chain hop wrapper as the
case where the difference bites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 02:27:00 +03:00
omarandClaude Opus 5 96d9cfaa63 fix(health): stop probing a chain below a hop that is already down
Hop probes were independent, so every hop was dialled whether or not the path
to it existed. A hop is dialled THROUGH the hops above it, so when hop 2 had no
live member left, the probe for hop 3 failed at hop 2 and hop 3 was recorded
dead. Dead means "we tested this and it did not work" — but nothing was learnt
about hop 3 at all. One broken hop painted the whole chain dead and pointed the
operator at the wrong place, and every one of those probes was a dial with a
timeout down a path already known to be broken.

Chain jobs now run in path order and the walk stops at the first hop that reads
dead. Hops below it are not dialled at all and are reported untested with
blocked_by naming the hop that stopped the walk — the honest answer, since
nothing was measured.

Nothing latches. There is no blocked flag: the gate is a fresh read of the
health board at every hop of every pass, and the cursor rewinds to the top each
cycle, so the first dead hop is never behind a break and is always retried. The
moment it answers, the rest of the chain runs in that same pass. Only a positive
dead blocks; untested never does, or a cold start would never open.

Blocked hops are rewritten rather than annotated, because board records do not
vanish when the prober stops dialling — they age out on their own TTL, and the
worst version of that is a stale dead pointing at a hop that may be fine.

The exit tag is exactly what stops being dialled, so the group test would have
waited out its full deadline and then reported "not reached yet" about a chain
it already knew was down. It now names the blocking hop immediately, gated on
the same freshness watermark so a break seen before the request cannot
short-circuit a pass that may be about to find that hop alive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 02:00:38 +03:00
omarandClaude Opus 5 3c7536dba0 fix(panel): show the chain hop by hop, and stop reading unused as broken
release / apk aarch64_cortex-a53 (push) Successful in 3m5s
release / apk x86_64 (push) Successful in 3m1s
release / release apk (push) Successful in 8s
The chain card gave a single verdict, so a dead hop was invisible: the operator
saw "the chain is unhealthy" and had to guess which of four hops to look at.
Meanwhile a group used only inside a chain showed "unused" next to a live
alive/dead count, which reads as a diagnosis when it only means nothing measures
it on that path.

Render the hops as a rail that severs below the first dead one, so which hop is
answered before a word is read, and split the two "not routed" messages into the
routing fact and the explicit non-fact. The group one names the case directly: a
group used only as a hop inside a chain reads unused here on purpose, and its
real health is on that chain's card.

Also fixes a bug this would otherwise have shipped: the readout painted every
ok:false in the critical colour, so "not routed" would have rendered as a fault
— the exact lie being removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 01:09:42 +03:00
omarandClaude Opus 5 3c92e1cbfd fix(health): one prober, on the path the rules actually use
A node reached only as a chain hop was being measured twice, and the reading
the panel showed was the wrong one. On a router in Russia that is not a cosmetic
difference: a node the chain carries fine behind a WireGuard hop is dead when
dialled straight out of the WAN, so the group card read "0 of 2 alive" while
that very group was carrying every packet.

Two dial paths existed outside the observatory plan. URLTestGroup.PostStart
warmed up every urltest group at box start whether or not any rule reached it,
and the panel's Test button reached URLTest.DialContext, whose first act is
Touch() — arming a ticker that re-swept those groups directly every probe
interval for the next thirty minutes. Both wrote under the BASE node tag, and
both dialled the base outbound, which carries no chain detour at all.

The observatory was never the liar: its plan roots come from the rules, and a
chain hop copy is stored only under its own tag, so no plan job could ever
write under a base tag. The fix is therefore to remove the other two paths, not
to touch the plan.

TestGroups now asks the observatory for an out-of-turn pass and reports what it
measured; a target no enabled rule routes to is not dialled at all and says so.
Unused urltest groups stand down their own self-check via a new SelfCheck option
(nil keeps today's behaviour, so every existing config is unchanged). The one
direct dial left is the exit-address lookup, which has no other possible source
— it now runs only for a target that is both routed and already read alive, so
it travels the routed path and never touches an unused group.

Chain hop wrappers are probed as measurements of their own and surfaced as
chains[].hops[], because "which hop is dead" is the question an operator has and
the chain-level verdict cannot answer it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 01:09:42 +03:00
omarandClaude Opus 5 bcc9df9282 test(wireguard): drive a real AmneziaWG tunnel through ClientBind
release / apk aarch64_cortex-a53 (push) Successful in 3m25s
release / apk x86_64 (push) Successful in 3m14s
release / release apk (push) Successful in 8s
The unit tests pin the reserved-byte gate on each side in isolation, which
would still pass if the two halves disagreed about when to apply it. This wires
two real wireguard-go devices together over loopback UDP through ClientBind on
both ends — the bind the detour path actually uses — configures ranged h1-h4
plus s4 and junk, and asserts an inner IP packet reaches the peer's TUN.

It is red against the unconditional clear and green with the gate, so it covers
the failure the field hit rather than the code we happened to write. Tagged
with_awg, so it runs under the shipped router tag set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:25:05 +03:00
omarandClaude Opus 5 ee3641fe45 fix(logsink): collapse interleaved floods, not just consecutive lines
The previous suppression compared each line with the one before it, which the
field never obliges. A dead chain makes the engine cycle the same message
across three outbound tags, so no two identical lines are adjacent: on the
router it produced 854 daemon lines in a ~760-line syslog ring and exactly one
summary, all while claiming "repeated 1 time". The rest of the system's log —
netifd, dnsmasq, the kernel — was evicted anyway.

Track a bounded table of open series keyed by the existing repeat key instead.
The first copy of a key prints; further copies inside its window are counted
whatever arrives in between; the window end emits one summary per key. The
summary now names its message, because several can close at once and "last
message" would simply be false under interleaving.

The table holds 256 keys and evicts the least recently seen, never silently: an
evicted series with a pending count prints its summary on the way out, marked
so the truncation is visible. Close, Reconfigure and any fatal flush every open
series first — a dying daemon may never reach Close.

TestRepeatAlternatingNotSuppressed asserted that A B A B must never be
collapsed. That assertion was the bug. It is replaced by a stronger one: the
messages get separate series, separate summaries and separate counts, so
distinct events still never fold into a single number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 439f62238f fix(engine): retire the superseded instance instead of leaving it running
Every config apply built a new box and left the old one alive. The engine's own
log gives it away: inside a single shaterd process, lines carried uptime
counters half an hour apart in the same second, and a live router was found
running four generations at once. A process restart cleared it, so the leak
accrued purely on re-apply.

That is not just wasted memory on a 512 MB box. Each surviving generation keeps
its WireGuard devices up, and two devices sharing one private key evict each
other at the peer — so the leak reproduced the duplicate-device defect between
generations, underneath the deduplication that only reasons about one config.

Retirement now has a hard budget: 5s, which is exactly sing-box's own
C.StopTimeout (past which upstream already calls a stop excessive) and stays
under C.FatalStopTimeout. It is paid after the replacement is serving and only
on an apply that changed something, so a no-op reconcile stays free.

A close that blows the budget is ABANDONED, not waited on, and the apply is
still reported as the success it is — the new box is built, started and
carrying traffic, and failing there would abort the netplane stage and leave a
stale ruleset over a healthy engine. The stuck instance is surfaced through
PendingCloses() into `shaterd status` and the panel, and clears itself if the
shutdown ever completes. Repeated applies over a stuck close no longer stack:
the abandoned generation is remembered, not re-created.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 d971eb85ee fix(wireguard): stop ClientBind from shredding the AmneziaWG magic header
An AmneziaWG node worked standalone and died the moment it was placed behind
an egress or a chain hop: the handshake completed, the peer answered, and then
not one byte of data ever arrived. The peer never confirmed the session, so it
re-handshook every 15 seconds, forever.

ClientBind cleared bytes 1-3 of every datagram on receive and stamped them on
send, unconditionally. Those bytes are Cloudflare's "reserved" field. They are
also where AmneziaWG puts the upper three bytes of its little-endian uint32
magic header, so zeroing them collapses the value to its low byte, which falls
outside every h1-h4 range and makes the peer classify the packet as an unknown
type and drop it silently.

Handshakes survived because s1/s2 padding pushes their magic past byte 3 — the
clear only scribbled on the random junk prefix. Transport packets have s4 = 0,
so their magic starts at byte 0 and took the hit. That asymmetry is the whole
signature: session up locally, zero data through.

Only the detour path was affected, because Endpoint.Start picks StdNetBind when
the dialer exposes WireGuardControl (no detour) and ClientBind otherwise. The
gate had already landed in StdNetBind; ClientBind was its untouched twin. The
two implement one contract and are now commented as the pair they are, so the
next fix cannot again land on one side only.

Measured on the box: h4 spans 0x60728123-0x60728155, so zeroing bytes 1-3
leaves 35..85 — the captured transport packet began with 56, while a node
without a detour carried a correct 0x6b039798 at the same moment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 a0f6083e28 fix(panel): show whether a rule is in force, not just what was saved
release / apk aarch64_cortex-a53 (push) Successful in 3m7s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 8s
With two catch-all rules both enabled in UCI and a WAN profile enabling one
and disabling the other, the panel drew BOTH switches on while the engine
ran only one chain. GET /api/config is right to return the raw model — that
is the desired state the panel PUTs back — but Routing.tsx read the row
state and the active count from it too, so the interface claimed a setting
was in force when it was not. Same defect class as the Protected badge.

/api/rules/reachability now carries the effective flag and, where the active
profile changed the outcome, its name and direction. The annotation is a
DIFF of ApplyProfileRuleOverrides output against desired state rather than a
second reading of the profiles name lists, so profile logic is not
duplicated and cannot drift — an unmigrated rule the profile is forbidden to
enable produces no diff and gets no badge, with nothing here needing to know
about LegacyDst.

In the UI the two states stay separate: the switch remains the only carrier
of desired state and still writes UCI, while the effective state drives the
dimmed row, the badge, the banner and the header count. Mirroring the
effective state into the switch would be worse than the original bug — the
operator would be toggling someone elses control, and the profiles decision
would be written back as their own choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 77369aedfe fix(logsink): collapse repeated lines instead of erasing the routers syslog
A broken outbound makes the engine repeat one line about once a second —
370 copies in six minutes. The routers syslog ring holds ~760 lines, so
within minutes it evicts the history of every other subsystem and our own
startup lines with it. Diagnosing the WireGuard duplication above required
restarting the service purely to catch the first seconds of a boot.

Collapse runs into "last message repeated N times". The comparison key is
level + text with the uptime field dropped: comparing whole lines would
suppress only same-second bursts, because that counter ticks. The per
connection "[id duration]" group is deliberately KEPT in the key — those ids
are distinct connections, and folding "50 connections failed" into one count
would be a worse lie than the flood. Window 5s, so a standing fault keeps
being reported instead of looking like a frozen log. fatal/panic are never
suppressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 515ae6d1b7 test(generate): make the remote-blocklist test exercise the remote path
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two
releases, which made the whole package exit non-zero no matter what the code
did — a real regression would have drowned in the familiar red.

The cause is not the packages no-network fetcher stub, as it first appears.
ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION
alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so
the fixture fell into the TEXT-list path: downloaded by generates own
fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which
every assertion below then contradicted. No stub content could fix that; the
stub decides the lists contents, not the rule-sets type.

Give the URL the .srs suffix the test always meant it to have, so the engine
fetches the compiled set itself through the direct outbound. No assertion is
weakened and the no-network stub stays in place.

Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags):
338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 a2ffbb1292 fix(generate): one WireGuard device per private key
A node may be copied freely by this package: a per-chain hop copy and a
per-group egress copy are rebuilt from the share-link so each can carry its
own Detour. For vless that is right — a copy is another TCP client. For
WireGuard it is not: each emitted endpoint is a real device holding the
nodes private key, and a peer keeps exactly ONE session per public key.
Two devices from one key evict each other continuously, and with keepalive
on both the loop never settles: NEITHER passes traffic.

buildOutboundsAndEndpoints emits the base endpoint for every enabled node
whether or not anything references it, so a WG node used only as a chain hop
always produced two devices. That is what any chain containing a WG node
looks like — every such chain was permanently dead.

Observed on the box: two UDP sockets from shaterd to the same peer port, the
servers peer endpoint flapping between them, +32 bytes/min through the
tunnel and every hop failing with "context deadline exceeded".

Deduplicate once on the assembled options, which catches all three producer
paths by construction. Duplicates are DELETED, not merely unreferenced:
box.New starts every endpoint regardless of reachability, so a leftover
would still bring its device up and still fight for the session. Dangling
references go to block, never to direct — a consumer whose tunnel just
disappeared must stop, not fall out onto the plain WAN.

Subscription fetch detours seed the reachability walk (they are direct
references like any rule), mirroring engine.ViaToTag exactly, with a
tripwire test against drift. A config that genuinely needs two devices for
one key keeps one and fail-closes the rest with a critical warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 1746d4d0ef fix: stop the panel and the shipped binary from lying about what works
release / apk aarch64_cortex-a53 (push) Successful in 9m13s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 7s
Four defects, all found by the owner on the live router, all of the same
family: something declared itself working while it was not.

WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave
"create WireGuard device: gVisor is not included in this build". The router
tag set carried with_wireguard and with_awg but not with_gvisor, so
sing-tun compiled its stub instead of the netstack every WireGuard device
needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving
requirement", so this was a broken promise, not a trim.

The tag itself was the small half. The tag set was the ONE build
configuration nothing in the repo tested: TestAmneziaWGEndpoint passes
because tests build with the full upstream tags. So the set now lives in
one file (scripts/router-tags.sh) and two guards hold it to the feature
list -- a static check that needs no tags, no Linux and no network (so the
next such gap fails on the developer's machine), and a behavioural one that
constructs every declared protocol through box.New UNDER THE SHIPPED TAGS,
where skipping is forbidden. Removing the tag now fails with the feature
name, the missing tag, and why: "Either add the tag back, or stop declaring
the feature -- those are the only two honest options." Cost: +2.8 MB raw,
+0.6-0.7 MB packed per arch. D23; D9 corrected.

THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came
from plane === 'full', which reports whether the data plane is installed --
nft table, policy routing, live engine -- and says nothing about where the
traffic goes. On a config with one `default -> direct` rule and no groups
the plane is fully installed and every packet leaves in the clear, so the
worst possible state rendered as the reassuring one.

The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the
moment they reach the engine, not from the model: buildRoute changes the
answer (a scheduled rule outside its window is never emitted, only the last
condition-less rule reaches Final, an unresolved target is rewritten by
ruleKillFallback), and re-deriving it anywhere else is a second
implementation that will drift -- model/reachability.go exists because two
already did. Four verdicts, not three: `blocked` is separate because under
a closed kill-switch with no catch-all nothing leaks, and calling that
"going out directly" is a lie in the alarm direction. Rider: Overview's
defaultTarget printed the highest-Order enabled rule as the default; a rule
becomes Final by having no conditions, whatever its Order.

"PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE
(B15). Once the browser suppresses dialogs, window.confirm returns false
immediately, so all 15 confirmations across 7 pages read as "cancelled" and
silently did nothing, with no way to recover from inside the panel. Replaced
with an in-app dialog the browser cannot mute: focus trapped and parked on
Cancel, Esc and veil cancel, focus returned to the opener, crit styling for
destructive commits. useConfirm() throws if the provider is missing rather
than falling back to a quiet false -- the failure mode being fixed.

HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed,
so a feed's nodes of those types vanished. The real landmine was one layer
up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so
without schemePrefixes the links were gone before any parser ran and the
fix would have looked complete. Undeliverable parameters are refused when
the node cannot work or would be less secure than the link asked (obfs,
pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it
survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot
produce a QUIC TLS config, and that fails at dial time, not at box.New.

Also: nodes added by hand can be named and renamed. The name is the
outbound tag, so a rename rewrites every reference in one PUT -- rule
targets, group members, chain hops, detours -- in the spelling each already
uses, and is refused outright when a group answers to the same bare name.
Subscription nodes state why they cannot be renamed instead of hiding the
control.

go build, go vet, go test ./shater/... (13 packages), panel npm run build
and npm test (13/13) all green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 20:08:58 +03:00
omarandClaude Opus 5 f86501bf77 ci!: drop the opkg lane — apk only, and fix the stale rolling release
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and
main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has
no `opkg` binary at all. The 24.10 lane was building and signing a feed no
device could consume.

Removed jobs `build` and `release` with the scripts only they called
(ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh)
and the usign trust anchor dist/shater-feed.pub. A committed public key is
an instruction: it invites the old install path for a feed that is no longer
produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD
still holds the secret half, and a usign secret contains its own public half,
so the identity is reconstructible if a 24.10 device ever needs serving.
D7 is marked SUPERSEDED by the new D22 rather than deleted.

Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from
2026-07-24 while every tag run published its versioned release correctly.
The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest
(workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run
never touched the rolling pointer. Asset replacement was never the problem;
ci/gitea-release.sh already deletes before recreating. A router pinned to
the rolling URL sat on 0.2.0 while `apk update` reported success: silent
staleness, the failure mode this repo keeps having to close.

The rolling pointer is now published on EVERY run, tag runs included, and a
new assert reads the release back over the API afterwards: our three
tag-versioned packages at the built version plus the index and the key must
be present (exit 13), and no package asset at any other version may survive
(exit 14). Same class of check as sdk-build-apk.sh's package-version assert,
added for the same reason -- the previous failure mode was silent.

KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references
it. Docs state plainly that mini_router is deliberately pinned to a
versioned URL and that the hand-edit per release is the price of pinning.

Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can
no longer install our packages. Its 25.12 rebuild is in flight separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:46:21 +03:00
omarandClaude Opus 5 eccfc6136c fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review of a8ef887c5 + 244b7c419 ("a rule's destination is a rule-set,
and nothing else") found that the change rested on a comment that was not
true. ParseUCIExport dropped dst_domain/dst_ip on the strength of "the
migration is re-run on every load"; model.Migrate() actually runs only from
`shaterd migrate`, i.e. the service init and uci-defaults. The daemon's run
path, the SIGHUP reconcile and the panel's config write never migrate.

So an uncommitted migration (a full /overlay is the documented way that
happens) turned `list dst_domain 'bank.ru'` + `target direct` into a rule
with NO matchers, which IS the spelling of a catch-all: generate points
route.Final at it and the LAST such rule wins. One failed `uci commit` sent
every packet on the router out the plain WAN, silently.

Rule.LegacyDst is the tripwire. It is non-empty exactly when the config
still carries the removed options, and three locks hang off it:
  - ParseUCIExport holds such a rule DISABLED. Chosen over "make IsCatchAll
    false" alone, which only covers matcher-less rules: `dst_domain` plus a
    `src` was never a catch-all, and routing it without its destination
    would still have sent a whole subnet direct.
  - IsCatchAll returns false for it, so it can never own route.Final even
    if something hands its Enabled bit back.
  - ApplyProfileRuleOverrides refuses to enable it (a profile with
    `list enable_rule` would otherwise have defeated the parser).
ValidateRules reports it through the existing warning channel, before the
Enabled gate, so the one message explaining the outage is not suppressed by
the fact that caused it. The init script logs a failed migration to syslog
instead of discarding its exit code and stderr.

The write path had none of this. PUT /api/config decodes a Model straight
from the request body and render.go wrote `enabled` from it, so a panel
save erased the operator's lists (as did the subscription cron, which
re-renders the whole package), and a crafted body with Enabled:true and no
LegacyDst put a live matcher-less rule on disk -- the same whole-router
leak, re-entered from the other side. WriteUCI now reads DISK state and
refuses a rule-changing write over an unmigrated config (409, not 500);
non-rule writers pass and legacyDstOpts carries the options across so cron
preserves them; withDiskLegacyDst takes the field from disk so a fabricated
one can never reach the renderer.

Migration hardening: an entry list that migrates to nothing no longer has
its legacy option deleted (that made "matches nothing" silently become
"matches everything"); a hand-written rule-set whose name collides is no
longer allowed to swallow the entries; delete failures propagate instead of
bumping schema_version past them forever; every error path reverts the
staged uci delta so another process's commit cannot flush a half-migration.

untunnelable.go follows the destination out of the rule: a rule whose
rule-sets are known to match by name is still skipped by the ping/IPTV/VPN
plan, as its v1 form was. D21 documents the AND->OR widening for the
engine's TCP/UDP path; it does not follow that a leak-guard should widen
itself during an upgrade, and with target=direct that meant previously
tunnelled ICMP leaving with the client's real address. Inline rule-sets are
now read from the options, so an engine that has not started yet no longer
costs the operator their ping.

Rule-set vocabulary: `full:`/`suffix:`/`keyword:`/`regexp:` in a text list
fetched by URL were dropped with no diagnostic at all (normaliseListDomain
rejects any token with a colon) -- not "reported as an unknown prefix".
Unifying was rejected: published filter lists are full of colon-bearing
syntax, and a third-party `regexp:` is compiled into the router's matcher
and run per query. The difference stands and is paid for in diagnostics,
per list, on every generate. D21 gains the source/vocabulary table.

Panel: the add form warns about a matcher-less rule exactly as the edit
form does, from one shared predicate; its isCatchAll matches the daemon's
new one; an unmigrated rule reads as held-off rather than merely switched
off. The comment promising a "New list" button that D21 rejected is gone.

go build ./..., go vet ./shater/..., go test ./shater/... (13 packages) and
panel `npm run build` are green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:04:21 +03:00
omarandClaude Opus 5 244b7c4199 feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.

The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.

The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.

isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.

Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
252 changed files with 36701 additions and 3913 deletions
+235 -287
View File
@@ -1,36 +1,45 @@
# Shater v0.2 — build the 4-package signed opkg feed and publish it as a rolling
# Gitea release consumable as an `src/gz` feed.
# Shater v0.2 — build the 4-package signed **apk** feed and publish it as
# per-arch Gitea releases consumable as an apk repository.
#
# WHAT CHANGED FROM v0.1
# v0.1 shipped 3 packages: xrayctl (SDK-compiled Go) + shater-core +
# luci-app-shater (hand-packed data .ipk). v0.2 collapses the runtime into ONE
# forked binary and ships 4 packages, all built the canonical SDK way:
# WHAT WE SHIP
# ONE forked binary plus its OpenWrt glue, 4 packages, all built the canonical
# SDK way:
# - shaterd PREBUILT static-musl + SPA-embedded + UPX binary. Built
# OUT OF TREE by scripts/build-shaterd.sh (Go + Node + UPX)
# and staged into openwrt/shaterd/files/ BEFORE the SDK
# build; the openwrt/shaterd package just $(INSTALL_BIN)s
# the arch-matched artifact. (arch-specific .ipk)
# the arch-matched artifact. (arch-specific .apk)
# - shater-core data glue, PKGARCH=all
# - luci-app-shater LuCI thin launcher, PKGARCH=all (uses feeds/luci/luci.mk)
# - byedpi ciadpi, C cross-compiled from source by the SDK (arch-specific)
#
# TARGET HARDWARE / ARCH MATRIX
# x86_64 -> the QEMU testbed VM (generic x86-64).
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 + BPI-R4, mediatek/filogic).
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 mini + BPI-R4,
# mediatek/filogic), both on 25.12 with apk-tools 3.
# Only shaterd + byedpi are arch-specific; shater-core + luci-app-shater are
# PKGARCH=all, so one build of each covers every device. opkg filters by
# Architecture at install time, so a single combined feed URL serves all.
# PKGARCH=all, so one build of each covers every device — but the RELEASES
# are still per-arch (see the release-apk job for why).
#
# FEED SIGNING (opkg / usign — OpenWrt 24.10 is opkg, not apk; apk lands at 25.12)
# The feed index (Packages) is usign-signed with the SECRET key in the Gitea
# repo secret KEY_BUILD; routers verify it with the committed public key
# dist/shater-feed.pub (fingerprint 5ac4b177689cb8e0). Do NOT regenerate the
# key — that invalidates every deployed router's trust.
# FORMAT: apk ONLY (25.12+)
# The fleet runs OpenWrt/ImmortalWrt 25.12, where opkg is replaced by Alpine
# apk (.apk files, binary packages.adb index, EC keys in /etc/apk/keys/). The
# old .ipk lane was removed in 2026-07 (docs-shater/DECISIONS.md D22): no
# device we serve has an opkg binary at all, so building and signing a second
# feed served nobody.
#
# FEED SIGNING (EC / apk)
# packages.adb is signed with the EC (prime256v1) SECRET key in the Gitea repo
# secret KEY_APK; routers verify it with the committed public key
# dist/shater-apk.pem (ci/gen-apk-key.sh). Do NOT regenerate the key — that
# invalidates every deployed router's trust.
#
# AUTO-RELEASE
# push a tag `vX.Y.Z` -> versioned release. workflow_dispatch / (optional) main
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
# API via curl (ci/gitea-release.sh) — no external action needed.
# push a tag `vX.Y.Z` -> versioned per-arch releases `apk-vX.Y.Z-<arch>`.
# workflow_dispatch -> rolling per-arch `apk-latest-<arch>` (always-fresh
# feed). Publish uses the Gitea API via curl (ci/gitea-release.sh) — no
# external action needed. NOTE: the apk release tags deliberately do NOT start
# with `v` so publishing them cannot re-trigger this workflow's `v*` filter.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
@@ -41,28 +50,13 @@
# exported via $GITHUB_ENV):
# tag `vX.Y.Z` -> X.Y.Z-r1
# anything else -> <nearest tag>-r<commits since it + 1>
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# and hands them to the SDK build as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
# built .ipk/.apk really carry that version, so the failure can never be
# silent again. This is also why both build jobs check out with fetch-depth: 0
# into the binary's constant.Version. ci/sdk-build-apk.sh then ASSERTS that the
# built .apk really carry that version, so the failure can never be silent
# again. This is also why the build job checks out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
#
# APK LANE (25.12+, ADDITIVE — T2)
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
# index, EC keys in /etc/apk/keys/). The `build-apk` + `release-apk` jobs
# below build the SAME 4 packages through the ImmortalWrt 25.12 apk-SDK and
# publish PER-ARCH apk repos as releases `apk-latest-<arch>` (rolling) /
# `apk-<tag>-<arch>` (versioned). Per-arch because apk filenames carry no
# architecture (shaterd-0.2.0-r1.apk would collide across arches in one flat
# release) and apk fetches packages relative to the packages.adb URL.
# Signed with the EC key in the Gitea secret KEY_APK; trust anchor
# dist/shater-apk.pem (ci/gen-apk-key.sh). The usign/opkg lane above is
# UNCHANGED and keeps serving the 24.10 fleet. NOTE: the apk release tags
# deliberately do NOT start with `v` so publishing them cannot re-trigger
# this workflow's `v*` tag filter.
# CACHING (T3 — fast CI)
# All caches use actions/cache pinned to v3.3.2: the LAST release speaking the
@@ -80,34 +74,33 @@
# (PKG_VERSION/PKG_HASH live there). Stale-safe: the buildroot verifies
# PKG_HASH on every dl/ file and re-downloads on mismatch, so restore-keys
# prefix fallback is allowed.
# - Go module + build cache — key = hash of go.sum; shared by all 4 build
# - Go module + build cache — key = hash of go.sum; shared by both build
# jobs (each builds both GOARCHes).
# - panel/node_modules — key = hash of panel/package-lock.json, exact-only
# (a lockfile change MUST miss); on hit build-shaterd.sh gets --fast.
# - apt .deb archives for the apk lane's debian:bookworm host-deps
# (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list is in it).
# - usign binary (.cache/tools) — static helper, fixed key.
# - apt .deb archives for the debian:bookworm host-deps of the apk SDK
# container (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list
# is in it).
# - SDK feeds/ git checkouts (.cache/feeds) — the single biggest recurring
# cost: `scripts/feeds update -a` cloned base+packages+luci+routing+
# telephony EVERY run (~7 min/job; github.com is ~1 MB/s from this
# runner — run 51 evidence). The feeds dir is symlinked into the SDK
# container from the workspace cache; `feeds update` on an existing clone
# is a fast fetch+checkout of the pinned revs. Correctness-safe: update
# always checks out feeds.conf's pins, and ci/sdk-build*.sh wipes the
# always checks out feeds.conf's pins, and ci/sdk-build-apk.sh wipes the
# cache + re-clones fresh if update ever fails on a cached checkout.
# Key = lane + SDK release (shared across the two arch jobs of a lane —
# same release pins identical feed revs; the sequential runner means the
# second arch restores what the first saved). restore-keys lets an SDK
# version bump start from the old clones (git fetch delta, not re-clone).
# Key = lane + SDK release (shared across the two arch jobs — the same
# release pins identical feed revs; the sequential runner means the second
# arch restores what the first saved). restore-keys lets an SDK version
# bump start from the old clones (git fetch delta, not re-clone).
# Act_runner facts this design leans on (verified in run 51 logs):
# - the cache backend works: restores/saves confirmed, hashFiles() works;
# - docker images (openwrt/sdk, debian:bookworm, runner-images) live on the
# PERSISTENT host daemon — "Image is up to date" each run, no re-download;
# - docker images (debian:bookworm, runner-images) live on the PERSISTENT
# host daemon — "Image is up to date" each run, no re-download;
# - each actions/cache SAVE is followed by an exact 3-minute act_runner
# stall (node process lingers; hit→no-save→no stall). Steady state saves
# nothing, so adding cache entries is fine, but keys that change every
# run (e.g. github.sha) would cost +3 min/entry/run — do NOT do that.
name: release
on:
@@ -125,54 +118,49 @@ concurrency:
cancel-in-progress: true
jobs:
build:
name: ${{ matrix.arch }}
# ---------------------------------------------------------------------------
# THE TEST GATE (2026-07-26). Everything below `needs:` this job, so a red test
# stops the release instead of shipping with it.
#
# WHY IT IS A JOB HERE AND NOT JUST .gitea/workflows/test.yml: a separate
# workflow cannot block another one — they run side by side and a red `test`
# workflow would have published anyway. Only a `needs:` edge inside THIS
# workflow is a gate. test.yml exists too, for fast feedback on `main`; both
# call the same scripts/run-tests.sh so they cannot drift.
#
# WHAT WAS BROKEN: the release tract ran two `go test` invocations in total —
# build-shaterd.sh's one-package buildtags check and check-router-tags.sh's
# three named tests. 115 of the 116 test files under shater/** had never run in
# CI (upstream's .github/workflows/test.yml triggers on branches this fork does
# not have, and Gitea ignores .github/workflows entirely once .gitea/workflows
# exists). TestDNSFilterRemoteBlocklistHTTPClient shipped red twice.
#
# WHAT IT COVERS: the whole suite under the SHIPPED build tags
# (scripts/router-tags.sh) on linux — the two dimensions that were missing.
# transport/wireguard compiles 1 test file without the tag set and 7 with it
# (the AmneziaWG ones); shater/generate has 44 test files on linux against 32
# elsewhere. Plus a -race pass and the panel's TypeScript tests. Details and
# the named, reasoned exclusions are in scripts/run-tests.sh.
test:
name: test gate
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
include:
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
steps:
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the engine via a go.mod
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
# must be present or `go build` dies with "no such file or directory".
# actions/checkout does not fetch submodules by default; init ONLY this one
# (clients/apple+android are large and unused here).
# go.mod `replace`s wireguard-go to ./submodules/wireguard-go, so without
# this even `go list` fails. Same step/reason as in build-apk below.
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod # pins Go 1.24.7 (go.mod `go` line)
cache: false # explicit actions/cache@v3.3.2 below (setup-go's
# built-in cache uses the new API act_runner lacks)
go-version-file: go.mod
cache: false # explicit actions/cache@v3.3.2 below
- name: Set up Node
uses: actions/setup-node@v4
with:
node-version: '20' # Vite 5 needs Node 18+; 20 LTS
# ---- caches (see the header comment for keys + version pin rationale) ----
# Same cache key as build-apk: this job runs first, so it warms the module
# + build cache the SDK-lane build then restores. (v3.3.2 pin: see header.)
- name: Cache Go modules + build cache
uses: actions/cache@v3.3.2
with:
@@ -183,92 +171,36 @@ jobs:
restore-keys: |
go-
# Node 24, NOT the 20 build-apk uses for the SPA: panel's tests are
# TypeScript run directly by `node --test`, and type stripping only exists
# from 22.6 — on node 20 `npm test` dies before running a single case.
- name: Set up Node
uses: actions/setup-node@v4
with:
node-version: '24'
- name: Cache panel node_modules
id: npm-cache
uses: actions/cache@v3.3.2
with:
path: panel/node_modules
key: npm-${{ hashFiles('panel/package-lock.json') }}
# NO restore-keys: node_modules must exactly match the lockfile;
# on any lockfile change this misses and `npm ci` runs fresh.
- name: Cache SDK dl/ (package sources)
uses: actions/cache@v3.3.2
with:
path: .cache/dl
key: dl-${{ hashFiles('openwrt/*/Makefile') }}
restore-keys: |
dl-
- name: Panel tests
run: bash scripts/run-panel-tests.sh
# feeds git checkouts (see header): both 24.10.4 arch jobs share one entry
# (same release = same feeds.conf.default pins), so derive the release
# from the matrix sdk tag (x86_64-24.10.4 -> 24.10.4).
- name: Compute feeds cache key
id: feedskey
run: echo "ver=$(echo '${{ matrix.sdk }}' | sed 's/.*-//')" >> "$GITHUB_OUTPUT"
- name: Cache SDK feeds checkouts
uses: actions/cache@v3.3.2
with:
path: .cache/feeds
key: feeds-opkg-${{ steps.feedskey.outputs.ver }}
restore-keys: |
feeds-opkg-
- name: Cache CI tools (usign)
uses: actions/cache@v3.3.2
with:
path: .cache/tools
key: tools-usign-v1
- name: Install UPX
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
# the SDK package build (the openwrt/shaterd package installs the staged
# artifact). $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
- name: Build signed feed (SDK)
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
run: bash ci/build-feed.sh "${{ matrix.arch }}" "${{ matrix.sdk }}" "out/${{ matrix.arch }}"
- name: Show feed
run: ls -l "out/${{ matrix.arch }}" && cat "out/${{ matrix.arch }}/Packages"
- name: Upload feed artifact
# v4 uses an artifact backend Gitea Actions does not implement
# (GHESNotSupportedError); v3 works on Gitea's act_runner.
uses: actions/upload-artifact@v3
with:
name: shater-${{ matrix.arch }}
path: out/${{ matrix.arch }}/*
if-no-files-found: error
- name: Go tests (shipped tags, linux, + race)
run: bash scripts/run-tests.sh
# ---------------------------------------------------------------------------
# APK lane (additive): the same 4 packages through the ImmortalWrt 25.12
# apk-SDK for the 25.12/apk fleet (BananaWRT 25.12-mtk-vendor routers + the
# future 25.12 VM). Produces a per-arch apk repo dir: *.apk + EC-signed
# packages.adb + shater-apk.pem. Artifact prefix `apkfeed-` (NOT `shater-`)
# so the opkg release job's `artifacts/shater-*` glob never picks these up.
# Build the 4 packages through the ImmortalWrt 25.12 apk-SDK for the 25.12/apk
# fleet (BPI-R3 mini on BananaWRT 25.12-mtk-vendor, BPI-R4 on OpenWrt 25.12,
# and the testbed VM). Produces a per-arch apk repo dir: *.apk + EC-signed
# packages.adb + shater-apk.pem, uploaded as the artifact `apkfeed-<arch>`.
build-apk:
name: apk ${{ matrix.arch }}
# THE GATE EDGE. A red test skips this job, which leaves no artifact, which
# (with the guards in release-apk) leaves nothing published.
needs: test
runs-on: ubuntu-latest
strategy:
fail-fast: false
@@ -281,8 +213,10 @@ jobs:
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
steps:
# fetch-depth: 0 — see the opkg lane: the package version comes from
# `git describe`, which needs tags + ancestry.
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
@@ -295,8 +229,10 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# Same single version computation as the opkg lane — both lanes MUST agree
# on the version, they package the identical tree.
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
@@ -373,11 +309,27 @@ jobs:
restore-keys: |
feeds-apk-
# D23 — the shipped tag set is a TRIMMED subset (scripts/router-tags.sh);
# everything else in CI builds with the full upstream set, so without this
# step the one combination we actually ship is never exercised. That is how
# `with_gvisor` was trimmed while `with_wireguard` stayed and every shipped
# binary answered a WireGuard node with "gVisor is not included in this
# build" (2026-07-25). The check runs the declared-feature/tag comparison
# and then constructs one node of every declared protocol through box.New
# UNDER THE SHIPPED TAGS. It runs before the artifact build so a tag trim
# that breaks a feature fails the release instead of shipping.
- name: Verify the shipped build-tag set (D23)
run: bash scripts/check-router-tags.sh
- name: Install UPX
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
# Same artifact-order contract as the opkg lane: the SPA-embedded shaterd
# binary is built OUT of the SDK and staged before the package build.
# Artifact-order contract: the SPA-embedded shaterd binary is built OUT of
# the SDK and staged into openwrt/shaterd/files/ BEFORE the package build
# (the openwrt/shaterd package only installs the staged artifact).
# $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
@@ -409,124 +361,34 @@ jobs:
if-no-files-found: error
# ---------------------------------------------------------------------------
# Publish once both arches are built. Rolling `latest` on dispatch, a versioned
# release on a `vX.Y.Z` tag. Self-contained (curl -> Gitea API).
release:
name: release
needs: build
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Download all arch feeds
uses: actions/download-artifact@v3
with:
path: artifacts
- name: Assemble release assets
id: assets
run: |
set -eu
mkdir -p release
# For each downloaded arch feed: one ready-to-serve tarball + loose ipks.
for d in artifacts/shater-*; do
[ -d "$d" ] || continue
arch="${d#artifacts/shater-}"
tar -C "$d" -czf "release/shater-feed-${arch}.tar.gz" .
# loose .ipk for direct `opkg install <url>` (dedupe shared _all ipks by name)
for ipk in "$d"/*.ipk; do
[ -e "$ipk" ] || continue
cp -n "$ipk" "release/$(basename "$ipk")"
done
done
# ship the feed's public key so routers can verify (see docs-shater/INSTALL.md)
cp -f dist/shater-feed.pub release/shater-feed.pub
ls -l release
echo "count=$(ls release | wc -l)" >> "$GITHUB_OUTPUT"
# restore the prebuilt usign binary (skips apt + cmake + clone + build)
- name: Cache CI tools (usign)
uses: actions/cache@v3.3.2
with:
path: .cache/tools
key: tools-usign-v1
- name: Install usign (feed signer)
run: bash ci/install-usign.sh
- name: Build & sign combined opkg feed index
# One Packages/Packages.gz over ALL loose .ipk (every arch + arch=all),
# with basename Filenames. opkg filters by Architecture, so a single
# release URL serves every device: BPI routers pick aarch64_cortex-a53 +
# all, the x86 testbed picks x86_64 + all. Signed with KEY_BUILD so
# routers keep check_signature on. This is what makes the release directly
# consumable as an `src/gz` feed (see docs-shater/INSTALL.md).
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
run: bash ci/make-index.sh release
- name: Determine release identity
id: rel
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
echo "tag=${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
echo "name=shater ${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
echo "prerelease=false" >> "$GITHUB_OUTPUT"
echo "rolling=false" >> "$GITHUB_OUTPUT"
else
echo "tag=latest" >> "$GITHUB_OUTPUT"
echo "name=shater latest (main)" >> "$GITHUB_OUTPUT"
echo "prerelease=true" >> "$GITHUB_OUTPUT"
echo "rolling=true" >> "$GITHUB_OUTPUT"
fi
- name: Publish Gitea release
env:
TOKEN: ${{ secrets.RELEASE_TOKEN != '' && secrets.RELEASE_TOKEN || github.token }}
TAG: ${{ steps.rel.outputs.tag }}
NAME: ${{ steps.rel.outputs.name }}
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
ROLLING: ${{ steps.rel.outputs.rolling }}
BODY: |
Automated build. Packages: shaterd + byedpi (per-arch), shater-core +
luci-app-shater (arch=all).
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
── Add as an opkg feed (recommended — then updating is one command) ──
This release is itself a SIGNED package feed; opkg filters by
architecture, so the same lines work on every device:
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" >> /etc/opkg/customfeeds.conf
opkg update
opkg install luci-app-shater # pulls shater-core + shaterd too
The public-key install is one-time; after it, `opkg update/upgrade`
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
── Update (name our packages — never a bare `opkg upgrade`) ──
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
── Or install the loose .ipk directly / from the tarball feed ──
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
opkg install /tmp/shater/luci-app-shater_*_all.ipk
run: bash ci/gitea-release.sh release/*
# ---------------------------------------------------------------------------
# Publish the apk lane: ONE release PER ARCH (apk package filenames carry no
# arch, and apk fetches `<name>-<ver>.apk` relative to the packages.adb URL —
# a flat multi-arch release would collide). Rolling `apk-latest-<arch>` on
# dispatch, `apk-<tag>-<arch>` on a version tag. The tags do NOT match the
# workflow's `v*` trigger, so publishing them cannot re-trigger the build.
# Publish: ONE release PER ARCH (apk package filenames carry no arch, and apk
# fetches `<name>-<ver>.apk` relative to the packages.adb URL — a flat
# multi-arch release would collide). Every run refreshes the ROLLING pointer
# `apk-latest-<arch>`; a `vX.Y.Z` tag run ALSO publishes the pinnable
# `apk-vX.Y.Z-<arch>`. The tags do NOT match the workflow's `v*` trigger, so
# publishing them cannot re-trigger the build.
#
# WHY THE ROLLING RELEASE IS PUBLISHED ON TAG RUNS TOO (fixed 2026-07-25):
# it used to be an either/or — `TAG=apk-latest-<arch>` on dispatch, ELSE
# `TAG=apk-<ver>-<arch>` — so once releases moved to tag pushes the rolling
# pointer was never written again. It froze at 0.2.0 (published 2026-07-24)
# while v0.2.9/v0.2.10 published fine, and every router whose
# /etc/apk/repositories.d/shater.list points at the rolling URL kept getting a
# successful, silent `apk update` with nothing new. Rolling is the whole point
# of that URL, so it is now written unconditionally and asserted afterwards.
release-apk:
name: release apk
needs: build-apk
needs: [test, build-apk]
# Publish whatever arch feeds succeeded — do NOT block the aarch64 release
# when an unrelated arch (e.g. x86_64) fails. download-artifact only fetches
# artifacts that exist, and the publish loop skips missing apkfeed-* dirs.
if: ${{ !cancelled() }}
#
# `needs.test.result == 'success'` is the second half of the gate. Without
# it, `!cancelled()` is true when the test job FAILS (build-apk is then
# skipped), this job runs with no artifacts at all, and — see the guard at
# the end of the publish step — used to exit 0 having published nothing. Red
# tests must SKIP this job, not "succeed" through it.
if: ${{ !cancelled() && needs.test.result == 'success' }}
runs-on: ubuntu-latest
steps:
- name: Checkout
@@ -537,6 +399,9 @@ jobs:
with:
path: artifacts
# Identity of the VERSIONED release only. The rolling pointer is published
# on every run with fixed prerelease=true/rolling=true, so it needs nothing
# from here.
- name: Determine release identity
id: rel
run: |
@@ -558,21 +423,47 @@ jobs:
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
ROLLING: ${{ steps.rel.outputs.rolling }}
run: |
set -eu
set -euo pipefail
# Counted, and asserted non-zero at the end. Until 2026-07-26 this loop
# was the step's whole body: with no artifacts the glob stayed
# unexpanded, `[ -d ... ]` was false, `continue` ran once, the loop
# ended and the step exited 0 — "release apk" went GREEN having
# published absolutely nothing. Any upstream failure (all arches
# failing to build, an artifact-name change, a download-artifact
# hiccup) therefore looked like a successful release.
published=0
for d in artifacts/apkfeed-*; do
[ -d "$d" ] || continue
arch="${d#artifacts/apkfeed-}"
if [ "$VER" = latest ]; then TAG="apk-latest-$arch"; else TAG="apk-$VER-$arch"; fi
ROLL="apk-latest-$arch"
# The version we just built, read straight off the artifact
# (`shaterd-<ver>-r<rel>.apk`). NOT recomputed with ci/version.sh:
# this job checks out shallow, so it has no tags to describe from.
pkg=""
for a in "$d"/shaterd-*.apk; do
if [ -f "$a" ]; then pkg="$(basename "$a")"; fi
done
[ -n "$pkg" ] || { echo "[release-apk] ERROR: no shaterd-*.apk in $d"; exit 11; }
want="${pkg#shaterd-}"; want="${want%.apk}"
echo "[release-apk] arch=$arch built version=$want"
BODY="Automated apk (OpenWrt/ImmortalWrt 25.12+) package repo for \`$arch\`.
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
This build: \`$want\`.
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
── Add as an apk repository ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
── Add as an apk repository (rolling — install once, then just update) ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$ROLL/shater-apk.pem
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
\`apk-latest-<arch>\` is a MOVING pointer: every release run replaces its
assets, so the same repo line keeps serving the newest build. To pin a
version instead, point the repo line at
\`.../download/apk-vX.Y.Z-\$(cat /etc/apk/arch)/packages.adb\` — then the
file must be edited by hand for each upgrade.
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
@@ -580,9 +471,66 @@ jobs:
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
those packages are upgraded along with needed dependencies\").
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
echo "[release-apk] publishing $TAG from $d"
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
Full guide: docs-shater/INSTALL.md §5."
# 1) the pinnable versioned release (tag runs only)
if [ "$VER" != latest ]; then
echo "[release-apk] publishing apk-$VER-$arch from $d"
TAG="apk-$VER-$arch" NAME="shater apk $VER ($arch)" BODY="$BODY" \
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
bash ci/gitea-release.sh "$d"/*
fi
# 2) the rolling pointer — ALWAYS, tag run included. ci/gitea-release.sh
# deletes the existing release before recreating it, so the old
# version's assets are REPLACED, never accumulated (two versions of
# one package in one index would let apk choose, not us).
echo "[release-apk] publishing $ROLL from $d"
TAG="$ROLL" NAME="shater apk latest ($arch)" BODY="$BODY" \
PRERELEASE=true ROLLING=true \
bash ci/gitea-release.sh "$d"/*
# 3) ASSERT the rolling release really serves THIS build — same class
# of check as ci/sdk-build-apk.sh's package-version assert, and for
# the same reason: the previous failure mode was silent. Reads the
# published release back over the API and requires our three
# tag-versioned packages at $want, the index, the key — and NO
# left-over package asset at any other version.
api="$GITHUB_SERVER_URL/api/v1/repos/$GITHUB_REPOSITORY/releases/tags/$ROLL"
got="$(curl -fsS -H "Authorization: token $TOKEN" "$api" \
| tr '{},' '\n\n\n' \
| sed -n 's/.*"name"[[:space:]]*:[[:space:]]*"\([^"]*\)".*/\1/p' | sort -u)" || {
echo "[release-apk] ERROR: cannot read back $ROLL from the API"; exit 12; }
echo "[release-apk] $ROLL assets: $(printf '%s ' $got)"
# here-string, NOT `printf | grep -q`: under `pipefail` the early
# exit of grep -q can SIGPIPE the writer and fail a passing check.
for f in "shaterd-$want.apk" "shater-core-$want.apk" \
"luci-app-shater-$want.apk" packages.adb shater-apk.pem; do
grep -qxF "$f" <<<"$got" || {
echo "[release-apk] ERROR: $ROLL does not contain '$f' after publish."
echo " A router pinned to the rolling URL would have silently"
echo " stayed on its old version with a successful apk update."
exit 13; }
done
stale="$(grep -E '^(shaterd|shater-core|luci-app-shater)-.*\.apk$' <<<"$got" \
| grep -vxF -e "shaterd-$want.apk" -e "shater-core-$want.apk" \
-e "luci-app-shater-$want.apk" || true)"
[ -z "$stale" ] || {
echo "[release-apk] ERROR: $ROLL still holds stale package assets:"
printf ' %s\n' $stale
echo " Two versions of one package in one feed = apk picks by its"
echo " own rules, not by our intent."
exit 14; }
echo "[release-apk] OK — $ROLL serves $want"
published=$((published + 1))
done
# The assert the loop above never had. Zero feeds published is a failed
# release, not a quiet success — say so with a non-zero exit.
if [ "$published" -eq 0 ]; then
echo "[release-apk] ERROR: no apkfeed-* artifact reached this job, so"
echo " NOTHING was published. Downloaded tree:"
ls -la artifacts 2>&1 | sed 's/^/ /' || echo " (no artifacts/ dir at all)"
exit 10
fi
echo "[release-apk] published $published arch feed(s)"
+85
View File
@@ -0,0 +1,85 @@
# Shater — the test gate, on every push to `main`.
#
# WHY THIS FILE EXISTS (2026-07-26)
# The fork had a full suite and no CI that ran it. Upstream's
# .github/workflows/test.yml triggers on `stable`/`testing`/`unstable`; this
# repo only has `main`. And Gitea does not read .github/workflows AT ALL once
# .gitea/workflows exists — so those files are decoration here. Result: 115 of
# the 116 test files under shater/** had never once executed in CI, and
# TestDNSFilterRemoteBlocklistHTTPClient stayed red across two published
# releases.
#
# RELATIONSHIP TO release.yml
# This workflow is the FAST FEEDBACK loop on `main`. It is NOT the release
# gate: a separate workflow cannot block another one. The gate is the `test`
# JOB inside .gitea/workflows/release.yml, which build-apk `needs:` — see the
# comment there. Both run the very same scripts/run-tests.sh, so they cannot
# drift apart.
name: test
on:
push:
branches: [main]
paths-ignore:
- '**.md'
- 'docs-shater/**'
pull_request:
branches: [main]
workflow_dispatch:
concurrency:
group: test-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
name: go + panel tests
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
# go.mod has `replace github.com/sagernet/wireguard-go => ./submodules/
# wireguard-go`, so WITHOUT this every `go list`/`go test` fails before it
# starts. Same step, same reason, as in release.yml's build job.
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: false # explicit actions/cache@v3.3.2 below
# v3.3.2 is the last release speaking the cache API act_runner implements
# (see the header of release.yml). Same key as the release build job, so
# whichever runs first warms the other.
- name: Cache Go modules + build cache
uses: actions/cache@v3.3.2
with:
path: |
~/go/pkg/mod
~/.cache/go-build
key: go-${{ hashFiles('go.sum') }}
restore-keys: |
go-
# Node 24, NOT the 20 the SPA build uses: panel's tests are TypeScript run
# through `node --test`, and type stripping only exists from 22.6. On
# node 20 `npm test` dies with a syntax error before running anything.
- name: Set up Node
uses: actions/setup-node@v4
with:
node-version: '24'
- name: Cache panel node_modules
uses: actions/cache@v3.3.2
with:
path: panel/node_modules
key: npm-${{ hashFiles('panel/package-lock.json') }}
- name: Panel tests
run: bash scripts/run-panel-tests.sh
- name: Go tests (shipped tags, linux, + race)
run: bash scripts/run-tests.sh
+1 -1
View File
@@ -63,7 +63,7 @@ nul
/venv/
/test/cache.db
# feed artifacts (tracked public key dist/shater-feed.pub is force-added)
# feed artifacts (the tracked apk trust anchor dist/shater-apk.pem is force-added)
/dist/
# local agent config (CLAUDE.md is deliberately tracked; .claude local settings are not)
+7 -1
View File
@@ -7,4 +7,10 @@
[submodule "submodules/wireguard-go"]
path = submodules/wireguard-go
url = https://github.com/Leadaxe/wireguard-go-awg2-lx
branch = lx
# The pin lives on lx-awg2-v005, NOT on lx: the two are separate lines (42
# commits apart one way, 131 the other). `lx` has no hasReserved() gate in
# conn/bind_std.go at all, so a `git submodule update --remote` against it
# would silently restore the bug where ClientBind/StdNetBind shred the
# AmneziaWG magic header and no chain carries traffic. Keep this pointing at
# the line the pin is actually on.
branch = lx-awg2-v005
+19 -23
View File
@@ -32,7 +32,9 @@ single-use token into the standalone SPA the daemon serves on its own port
## Highlights
- Transparent **TPROXY** data plane (TCP + UDP), SNI/Host/QUIC sniffing, no DNS leaks.
- Transparent **TPROXY** data plane (TCP + UDP), SNI/Host/QUIC sniffing, no DNS leaks
— `:53` interception is on by default and covers the queries a client sends to the
router itself, not just the ones aimed around it (`globals.dns_intercept`, D24).
- First-match routing by source / destination / list / geo / client → outbound /
selector / chain / direct / block; node groups with balancer/observatory;
multi-hop chains; per-rule egress.
@@ -49,29 +51,22 @@ Full list with MVP/T1/T2 tags — [`docs-shater/FEATURES.md`](docs-shater/FEATUR
## Install
Two signed feeds. Pick by the router's OpenWrt version. Verbatim commands and the
manual `.ipk`/`.apk` install are in [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
**opkg (OpenWrt 24.10):**
One signed **apk** feed (OpenWrt / ImmortalWrt / BananaWRT **25.12+**), one
release per arch. Verbatim commands, the manual `.apk` install and the
rolling-vs-pinned choice are in
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
```sh
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
opkg update && opkg install luci-app-shater # -> shater-core -> shaterd
```
**apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+):**
```sh
wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
wget -O /etc/apk/keys/shater-apk.pem "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" > /etc/apk/repositories.d/shater.list
apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
`apk-latest-<arch>` is a moving pointer refreshed by every release run — install
once and `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`
keeps the router current. Point the repo line at `apk-vX.Y.Z-<arch>` instead to
pin a build; that file then has to be edited by hand for every upgrade.
shater ships **inert** (globals off) so install never breaks connectivity. After
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
then `shaterd apply` and `shaterd confirm`.
@@ -91,16 +86,17 @@ into `openwrt/shaterd/files/`. Details in
| `panel/` | Admin SPA (Vite + React + TS) and its Go server |
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Product documentation |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, feed/release scripts, CI |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, apk feed/release scripts, CI |
| `SPECS/`, `docs-lx/` | Engine-fork constitution/specs and feature-config reference |
| `docs/`, `mkdocs.yml` | **Upstream** sing-box docs (mkdocs) — kept as-is |
| `adapter/ cmd/ dns/ route/ option/ protocol/ transport/ …` | sing-box-lx engine tree |
## CI, upstream & license
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes signed
feeds: opkg (usign, key `5ac4b177689cb8e0`) and apk (EC key `shater-apk.pem`). A
`vX.Y.Z` tag → versioned release; `workflow_dispatch` → rolling `latest`.
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes a signed
per-arch apk repo (EC key `shater-apk.pem`). A `vX.Y.Z` tag → the pinnable
`apk-vX.Y.Z-<arch>`; every run also refreshes the rolling `apk-latest-<arch>` and
asserts over the API that it really serves the version just built.
The engine is the **sing-box-lx** fork — a thin downstream of upstream sing-box that
lives by **rebase, never merge**; its constitution is
+31 -48
View File
@@ -10,7 +10,7 @@
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue.svg)](LICENSE)
![targets: x86_64 · aarch64_cortex-a53](https://img.shields.io/badge/targets-x86__64%20%C2%B7%20aarch64__cortex--a53-brightgreen.svg)
![feeds: opkg 24.10 · apk 25.12](https://img.shields.io/badge/feeds-opkg%2024.10%20%C2%B7%20apk%2025.12-orange.svg)
![feed: apk 25.12+](https://img.shields.io/badge/feed-apk%2025.12%2B-orange.svg)
---
@@ -128,42 +128,15 @@ data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-sh
## Установка
shater поставляется двумя подписанными фидами. Выберите по версии OpenWrt на роутере:
- **OpenWrt 24.10** → фид **opkg** (`.ipk`, `Packages.gz`, ключ usign).
- **OpenWrt / ImmortalWrt / BananaWRT 25.12+** → фид **apk** (`.apk`, `packages.adb`,
EC-ключ).
shater поставляется одним подписанным **apk-фидом** (OpenWrt / ImmortalWrt /
BananaWRT **25.12+**: `.apk`, индекс `packages.adb`, EC-ключ в `/etc/apk/keys/`).
Старый opkg-фид (`.ipk`, 24.10) снят — оба наших роутера на 25.12 с apk-tools 3,
бинаря `opkg` там просто нет (`docs-shater/DECISIONS.md` D22).
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`
(+ опциональный `byedpi`). `shaterd` подтягивается автоматически как зависимость.
### Путь A — фид opkg (OpenWrt 24.10)
```sh
# 1) доверяем ключу фида — ИМЯ файла обязано равняться отпечатку usign-ключа.
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
# 2) добавляем фид (один URL обслуживает все арки).
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
opkg update
opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # опционально: ByeDPI desync-egress
```
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
он тянет обновления и на системные пакеты, это классический способ окирпичить
роутер):
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
```
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
### Фид apk
`/etc/apk/arch` сам выбирает нужный per-arch релиз (apk-релизы раздельны по арке):
@@ -198,13 +171,22 @@ apk upgrade shaterd shater-core luci-app-shater byedpi
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
> **Роллинг или фиксация — это выбор URL в `shater.list`.** `apk-latest-<arch>`
> — движущийся указатель: каждый релизный прогон заменяет его ассеты, поэтому
> «поставил и забыл»: `apk update` сам видит новую сборку. `apk-vX.Y.Z-<arch>` —
> фиксация на конкретной сборке: роутер не получит ничего нового, пока
> `/etc/apk/repositories.d/shater.list` не отредактируют руками — на каждом
> роутере и на каждый релиз. На `mini_router` сознательно прописан
> версионированный URL, и ручная правка — его цена. Подробнее —
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) §5.1.
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
> `25.12-mtk-vendor` — в [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
> Полные инструкции — ручная установка из `.apk`, фиксация версии
> (`apk-vX.Y.Z-<arch>`), совместимость с BananaWRT `25.12-mtk-vendor` — в
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
### Включение
@@ -258,8 +240,8 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Документация продукта (см. таблицу ниже) |
| `scripts/` | `build-shaterd.sh` — сборка ship-артефакта |
| `ci/` | Скрипты сборки фидов и релизов (SDK, usign/EC, Gitea API) |
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанные фиды opkg/apk |
| `ci/` | Скрипты сборки apk-фида и релизов (SDK, EC-подпись, Gitea API) |
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанный apk-фид |
| `SPECS/` | Конституция форка движка и спеки (Spec Kit) |
| `docs-lx/` | Справочник конфигурации фич движка (`lx-config.md`, `.ru.md`) |
| `lx-test/`, `submodules/` | Примеры конфигов движка и submodule AmneziaWG-рантайма |
@@ -273,16 +255,17 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает все 4 пакета и
публикует **подписанные фиды**:
- **opkg (24.10):** один комбинированный релиз, подписан usign-ключом (публичный
`dist/shater-feed.pub`, отпечаток `5ac4b177689cb8e0`; секрет — в Gitea-secret
`KEY_BUILD`).
- **apk (25.12+):** параллельная линия, **по релизу на арку**, подписан EC-ключом
(`dist/shater-apk.pem`; секрет — `KEY_APK`).
- **apk (25.12+)** — единственный формат: **по релизу на арку**, индекс
`packages.adb` подписан EC-ключом (публичный `dist/shater-apk.pem`; секрет — в
Gitea-secret `KEY_APK`).
Триггеры: push тега **`vX.Y.Z`** → версионный релиз; `workflow_dispatch` →
плавающий `latest`/`apk-latest-<arch>` (всегда свежий фид). Публикация — через
Gitea API (`ci/gitea-release.sh`). Ключи **никогда не перегенерируются** — это
инвалидировало бы доверие на всех развёрнутых роутерах.
Триггеры: push тега **`vX.Y.Z`** → версионный релиз `apk-vX.Y.Z-<arch>`;
`workflow_dispatch` → только роллинг. Роллинг `apk-latest-<arch>` обновляется
**на каждом прогоне**, включая теговый, и после публикации проверяется через API:
в нём обязаны лежать наши три пакета ровно собранной версии и ни одного ассета
другой версии. Публикация — через Gitea API (`ci/gitea-release.sh`). Ключ
**никогда не перегенерируется** — это инвалидировало бы доверие на всех
развёрнутых роутерах.
---
@@ -307,7 +290,7 @@ build-тегами и живущий **ребейзом на каждый upstre
| Документ | О чём |
|----------|-------|
| [`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, testbed/инфра |
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка обоих фидов (opkg/apk) |
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка apk-фида (роллинг/фиксация) |
| [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md) | One-binary дизайн, auth-handoff, data/DNS/apply-потоки (диаграммы) |
| [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [`docs-shater/ROADMAP.md`](docs-shater/ROADMAP.md) | Фазовый план |
@@ -3,7 +3,33 @@
| Поле | Значение |
|------|----------|
| Тип | B (bug) |
| Статус | C (complete) |
| Статус | C (complete) — guard **снят** (см. баннер ниже) |
> ## ⛔️ Guard снят (2026-07-26) — первопричина к shater не относится
> **Оба guard'а (Start-guard в `protocol/wireguard/endpoint.go` и
> selector-guard в `protocol/group/awg_selector_guard.go`) удалены**, вместе с
> их adapter-хуками (`OutboundManager.ConsumersOf`, `AmneziaWGSuspendable`).
> Апстрим снял их коммитом `5fa3a0a17`; сюда снятие приехало отдельно.
>
> **Почему.** Зависание было **Android-специфичным** (`Libbox.newService` не
> возвращал управление). Android для shater не платформа и ей не станет —
> мы собираем роутерный бинарь под OpenWrt/aarch64. При этом лекарство для
> самой AWG-за-detour связки у нас уже есть: reserved-clear gate в
> `ClientBind` (`d971eb85e` + пин сабмодуля `7d15f33`), без которого AWG не
> поднимался вообще ни за каким detour'ом. Мы носили и лекарство, и запрет
> на его применение.
>
> **Чем это было плохо на практике.** Guard отказывал **молча**: не ошибкой,
> а `started=false`, после чего каждый дозвон падал с «WireGuard is not ready
> yet». Конфигурация «AmneziaWG за WireGuard-хопом» выглядела не как
> отклонённая, а как «нода почему-то не работает».
>
> **Регрессия:** `protocol/wireguard/awg_over_wireguard_start_lx_test.go`
> (`with_gvisor && with_awg`) — AWG-эндпоинт с `detour` на outbound типа
> `wireguard` доходит до PostStart и поднимает `started`. До снятия guard'а
> тест краснел.
>
> **Осталось:** сквозной прогон на железе (AWG поверх реального WG-хопа).
Отклонять (по образцу ядрового запрета «empty direct detour») конфигурацию, где
AmneziaWG-endpoint (источник с AWG-полями) имеет `detour` на **любой
+145
View File
@@ -0,0 +1,145 @@
// lx:begin l3-honest-drop
package adapter
import (
"net/netip"
"testing"
"github.com/sagernet/sing-tun"
"github.com/sagernet/sing-tun/gtcpip/header"
"github.com/stretchr/testify/require"
)
// judgeFlowRouter answers PreMatch with a canned verdict; JudgeFlow reads
// nothing else off the Router.
type judgeFlowRouter struct {
Router
result PreMatchResult
}
func (r *judgeFlowRouter) PreMatch(InboundContext, []byte) PreMatchResult { return r.result }
// judgeFlowPort is the tun.Port half of a FlowOutbound. inet4 is what
// PortAddresses reports for IPv4 — the one field the two ICMP consumers in
// sing-tun disagree about (see the comment on
// TestJudgeFlowICMPToBoundPortStaysAFlow).
type judgeFlowPort struct {
Outbound
inet4 netip.Addr
}
func (o *judgeFlowPort) Tag() string { return "wg-out" }
func (o *judgeFlowPort) Type() string { return "wireguard" }
func (o *judgeFlowPort) PortAddresses() (netip.Addr, netip.Addr) {
return o.inet4, netip.Addr{}
}
func (o *judgeFlowPort) PortMTU() uint32 { return 1420 }
func (o *judgeFlowPort) AttachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) DetachReturn(tun.Return) error { return nil }
func (o *judgeFlowPort) WritePackets(packets [][]byte) error { return nil }
// judgeFlowNonPort is a FlowOutbound-shaped result that is NOT a tun.Port — the
// interface drift the second line of defense in JudgeFlow exists for.
type judgeFlowNonPort struct {
Outbound
}
func (o *judgeFlowNonPort) Tag() string { return "drifted" }
func (o *judgeFlowNonPort) Type() string { return "drifted" }
func judgeFlow(t *testing.T, protocol uint8, result PreMatchResult) tun.FlowVerdict {
t.Helper()
return JudgeFlow(
&judgeFlowRouter{result: result},
"l3-in", "tun", protocol,
netip.MustParseAddrPort("192.168.1.2:1234"),
netip.MustParseAddrPort("1.1.1.1:1234"),
nil,
)
}
const (
judgeFlowICMP = uint8(header.ICMPv4ProtocolNumber)
judgeFlowTCP = uint8(header.TCPProtocolNumber)
)
// TestJudgeFlowICMPToBoundPortStaysAFlow is the guard on the ONE fix that must
// not be made here.
//
// sing-tun has two ICMP consumers with different requirements on the port:
//
// - ForwardDispatcher.createFlow (flow_dispatch.go) needs only a VALID port
// address — it NATs the echo identifier and rewrites the source to that
// address. This is the path every unfragmented LAN ping takes, and it is
// what makes ping-through-WireGuard/AWG work at all.
// - ICMPForwarder.installFlow (stack_gvisor_icmp.go) additionally requires the
// address to be UNSPECIFIED, because it writes the packet to the port
// unmodified. A WireGuard endpoint reports its concrete interface address
// (transport/wireguard/port.go), so installFlow declines and HandlePacket
// falls through to forging the echo reply.
//
// The tempting fix — "for ICMP, refuse ActionFlow when PortAddresses() is not
// unspecified, so the verdict becomes a drop and the forgery is unreachable" —
// is applied HERE, in the one function both consumers share, with byte-identical
// arguments from either. It would therefore kill the working path too: every
// ping through WireGuard/AWG, fragmented or not, would drop, and l3_tunnel would
// carry nothing but `direct`. Keep this test failing loudly if anyone tries.
func TestJudgeFlowICMPToBoundPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.MustParseAddr("10.2.0.2")}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action,
"ICMP to a WireGuard/AWG endpoint must stay a flow: the forward dispatcher NATs it by echo identifier and this is the whole point of l3_tunnel")
require.Same(t, tun.Port(port), verdict.Port)
}
// The `direct` shape: an unspecified port address. Both consumers accept it.
func TestJudgeFlowICMPToUnspecifiedPortStaysAFlow(t *testing.T) {
t.Parallel()
port := &judgeFlowPort{inet4: netip.IPv4Unspecified()}
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: port})
require.Equal(t, tun.ActionFlow, verdict.Action)
require.Same(t, tun.Port(port), verdict.Port)
}
// PreMatchDrop is the honest verdict and must arrive as ActionDrop: it is the
// only value (besides Reject) that stops ICMPForwarder.HandlePacket before the
// Echo -> EchoReply rewrite.
func TestJudgeFlowICMPDropReachesTheStackAsDrop(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchDrop})
require.Equal(t, tun.ActionDrop, verdict.Action)
}
// The second line of defense: a PreMatchFlow whose outbound is not a tun.Port
// must not degrade ICMP to ActionAccept, because Accept is the forged reply.
func TestJudgeFlowICMPNonPortOutboundDrops(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowICMP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionDrop, verdict.Action,
"FlowOutbound and tun.Port are distinct interfaces; a drift between them must not silently re-enable the echo forger")
}
func TestJudgeFlowTCPNonPortOutboundAccepts(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchFlow, Outbound: &judgeFlowNonPort{}})
require.Equal(t, tun.ActionAccept, verdict.Action,
"for TCP, falling back to Accept is upstream behaviour and must stay untouched")
}
// TCP keeps every mapping it had, including the Continue -> Accept default that
// is a forgery only for ICMP.
func TestJudgeFlowTCPContinueStaysAccept(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchContinue})
require.Equal(t, tun.ActionAccept, verdict.Action)
}
func TestJudgeFlowTCPBypassStaysBypass(t *testing.T) {
t.Parallel()
verdict := judgeFlow(t, judgeFlowTCP, PreMatchResult{Action: PreMatchBypass})
require.Equal(t, tun.ActionBypass, verdict.Action)
}
// lx:end l3-honest-drop
-22
View File
@@ -45,30 +45,8 @@ type OutboundManager interface {
Default() Outbound
Remove(tag string) error
Create(ctx context.Context, router Router, logger log.ContextLogger, tag string, outboundType string, options any) error
// lx:begin awg
// ConsumersOf returns the tags of outbounds that depend on (detour through)
// the given tag — the reverse of Dependencies(). Used by the selector guard to
// walk up to AmneziaWG consumers when a group switches to a WireGuard member.
ConsumersOf(tag string) []string
// lx:end awg
}
// lx:begin awg
// AmneziaWGSuspendable is implemented by an AmneziaWG endpoint so the selector
// guard can suspend it (bring its device down) when a group it detours through
// switches to a WireGuard member — AmneziaWG inside a WireGuard tunnel hangs the
// kernel on Android. The marker lives in adapter so protocol/group can act on it
// without importing protocol/wireguard.
type AmneziaWGSuspendable interface {
// IsAmneziaWG reports whether this endpoint runs AmneziaWG (has AWG params).
IsAmneziaWG() bool
// SuspendAmneziaWG brings the device down so no junk handshake is sent. It is
// idempotent and safe to call on a not-yet-started or already-suspended endpoint.
SuspendAmneziaWG()
}
// lx:end awg
// lx:begin idle-suspend
// IdleSuspendable is implemented by a WG/AWG endpoint so the router's idle tick
// (SPEC 020) can suspend it when it is idle and unreachable, without importing
-15
View File
@@ -208,21 +208,6 @@ func (m *Manager) Outbound(tag string) (adapter.Outbound, bool) {
return m.endpoint.Get(tag)
}
// lx:begin awg
// ConsumersOf returns a copy of the tags that detour through tag (reverse of
// Dependencies()), built from the dependByTag ledger populated at Create time.
func (m *Manager) ConsumersOf(tag string) []string {
m.access.RLock()
defer m.access.RUnlock()
consumers := m.dependByTag[tag]
if len(consumers) == 0 {
return nil
}
return append([]string(nil), consumers...)
}
// lx:end awg
func (m *Manager) Default() adapter.Outbound {
m.access.RLock()
defer m.access.RUnlock()
+11
View File
@@ -75,7 +75,18 @@ func JudgeFlow(router Router, inbound string, inboundType string, network uint8,
case PreMatchFlow:
port, isPort := result.Outbound.(tun.Port)
if !isPort {
// lx:begin l3-honest-drop
// Second line of defense behind route.(*Router).preMatchFlow: a
// PreMatchFlow result already implies the outbound is an
// adapter.FlowOutbound, but FlowOutbound and tun.Port are distinct
// interfaces, and a drift between them must not degrade ICMP to
// ActionAccept — the TUN stack would then forge the echo reply
// itself instead of admitting the tunnel cannot carry the packet.
if networkName == N.NetworkICMP {
return tun.FlowVerdict{Action: tun.ActionDrop}
}
return tun.FlowVerdict{Action: tun.ActionAccept}
// lx:end l3-honest-drop
}
verdict := tun.FlowVerdict{Action: tun.ActionFlow, Port: port, UDPTimeout: result.UDPTimeout, NewTracker: result.NewTracker}
if result.Destination.IsValid() {
+19 -20
View File
@@ -1,6 +1,6 @@
#!/bin/sh
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (the 25.12
# lane — additive next to ci/build-feed.sh, which stays the opkg/24.10 lane).
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (25.12+;
# the only packaging lane shater has — see docs-shater/DECISIONS.md D22).
#
# Usage: ci/build-feed-apk.sh <ARCH> <SDK_URL> <OUTDIR>
# e.g. ci/build-feed-apk.sh aarch64_cortex-a53 \
@@ -10,14 +10,14 @@
# This is the per-arch entrypoint the Gitea workflow's `build-apk` job calls.
# It runs on the CI RUNNER and:
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
# scripts/build-shaterd.sh (same artifact-order contract as the opkg lane);
# 2. drives a plain `debian:bookworm` container (workspace shared via
# `--volumes-from`, same trick as ci/build-feed.sh) that downloads the
# ImmortalWrt 25.12 apk-SDK tarball and runs ci/sdk-build-apk.sh in it:
# compile the 4 packages as .apk, then `apk mkndx --sign` the per-arch
# `packages.adb` index. Unlike the usign lane (index signed on the runner),
# apk indexing NEEDS the SDK's host `apk` tool, so index+sign happen inside
# the container.
# scripts/build-shaterd.sh (the artifact-order contract);
# 2. drives a plain `debian:bookworm` container (the job's workspace volume is
# shared into it with `--volumes-from $(hostname)`; a bare `-v $PWD:...`
# points at a host path that does not exist under act_runner's DinD) that
# downloads the ImmortalWrt 25.12 apk-SDK tarball and runs
# ci/sdk-build-apk.sh in it: compile the 4 packages as .apk, then
# `apk mkndx --sign` the per-arch `packages.adb` index. Indexing NEEDS the
# SDK's host `apk` tool, so index+sign happen inside the container.
#
# Why the ImmortalWrt SDK (not openwrt/sdk images): the 25.12 fleet runs
# BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base (target mediatek/filogic,
@@ -25,9 +25,9 @@
# mediatek-filogic 25.12 tag — hence the official SDK tarball.
#
# Env:
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret — the apk analog
# of KEY_BUILD). If set, packages.adb carries an embedded signature
# verifiable by dist/shater-apk.pem (routers: /etc/apk/keys/).
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret). If set,
# packages.adb carries an embedded signature verifiable by
# dist/shater-apk.pem (routers: /etc/apk/keys/).
# If unset, an UNSIGNED index is produced (warning; not shippable —
# apk signatures are effectively mandatory).
set -eu
@@ -56,10 +56,9 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
# standalone. Passed into the container below and re-exported to the
# unprivileged build user in ci/sdk-build-apk.sh.
# The workflow puts these in the job env via `ci/version.sh --env >>
# $GITHUB_ENV`; recompute here when run standalone. Passed into the container
# below and re-exported to the unprivileged build user in ci/sdk-build-apk.sh.
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
@@ -73,7 +72,7 @@ echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# SDK; PKG_HASH still verifies every file, so stale = re-downloaded.
# apt/ debian:bookworm .deb archives for the host-deps install.
# The nested container runs the build as an unprivileged user -> must be writable
# (same reason as the chmod 0777 "$OUT" in ci/build-feed.sh).
# (same reason as the chmod 0777 "$OUT" above).
CACHE="$REPO/.cache"
mkdir -p "$CACHE/sdk" "$CACHE/dl" "$CACHE/apt"
chmod -R a+rwX "$CACHE/dl" "$CACHE/apt" 2>/dev/null || true
@@ -96,8 +95,8 @@ sh "$REPO/ci/fetch-sdk.sh" "$SDK_URL" "$SDK_TAR"
# --- 1) SDK build + index + sign inside a debian container -------------------
# `--volumes-from $(hostname)` shares THIS job container's workspace volume into
# the nested container (see ci/build-feed.sh for why a bare -v does not work on
# the act_runner DinD setup).
# the nested container: a bare `-v $PWD:...` points at a host path that does not
# exist under the act_runner DinD setup.
echo "[apk-feed] SDK build arch=$ARCH (ImmortalWrt 25.12 apk-SDK)"
docker pull -q debian:bookworm
docker run --rm --volumes-from "$(hostname)" \
-106
View File
@@ -1,106 +0,0 @@
#!/bin/sh
# ci/build-feed.sh — build the signed opkg feed for ONE arch.
#
# Usage: ci/build-feed.sh <ARCH> <SDK_DOCKER_TAG> <OUTDIR>
# e.g. ci/build-feed.sh x86_64 x86_64-24.10.4 out/x86_64
# ci/build-feed.sh aarch64_cortex-a53 mediatek-filogic-24.10.4 out/aarch64_cortex-a53
#
# This is the reusable per-arch entrypoint the Gitea workflow calls. It runs on
# the CI RUNNER and:
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
# scripts/build-shaterd.sh (into openwrt/shaterd/files/) — proving artifact
# order: SPA+shaterd build BEFORE the SDK package build;
# 2. drives the arch-matched `openwrt/sdk` docker image to compile all 4
# packages (ci/sdk-build.sh) and collect their .ipk into OUTDIR;
# 3. builds + usign-signs the opkg `Packages` index over OUTDIR
# (ci/install-usign.sh + ci/make-index.sh; signs iff $KEY_BUILD is set).
#
# Env:
# KEY_BUILD usign SECRET key (Gitea repo secret). If set, the feed index is
# signed and verifiable by dist/shater-feed.pub (fp 5ac4b177689cb8e0).
# If unset, an UNSIGNED feed is produced (make-index warns).
set -eu
ARCH="${1:?arch required (x86_64 | aarch64_cortex-a53)}"
SDK_TAG="${2:?sdk docker tag required (e.g. x86_64-24.10.4)}"
OUT="${3:?output dir required}"
REPO="$(cd "$(dirname "$0")/.." && pwd)"
mkdir -p "$OUT"; OUT="$(cd "$OUT" && pwd)"
# $OUT is created here as ROOT on the runner, but the nested `openwrt/sdk`
# container runs as the unprivileged `buildbot` (uid 1000) — so it must be able
# to write the collected .ipk into $OUT. World-writable is set HERE (a chmod
# from inside the container, as buildbot, cannot fix a root-owned dir).
chmod 0777 "$OUT"
# --- 0) the prebuilt shaterd binary must already be staged for this arch ------
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
*) echo "[feed] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
esac
if [ ! -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" ]; then
echo "[feed] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
echo " Run scripts/build-shaterd.sh BEFORE ci/build-feed.sh." >&2
exit 3
fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# The workflow normally puts these in the job env (ci/version.sh --env >>
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
# 0.2.0-r3 and no router could ever see an update).
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
# Workspace dir restored/saved by actions/cache in the workflow and shared into
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
# there (ci/sdk-build.sh). PKG_HASH still verifies every file, so a stale cache
# can never produce a wrong build. Must be writable by the container's
# unprivileged buildbot user (same reason as the $OUT chmod above).
DL_DIR="$REPO/.cache/dl"
mkdir -p "$DL_DIR"
chmod -R a+rwX "$DL_DIR" 2>/dev/null || true
# --- 0.6) persistent feeds/ git checkouts -------------------------------------
# Workspace dir restored/saved by actions/cache (key: feeds-opkg-<release>) and
# symlinked over the SDK's feeds/ inside the container (ci/sdk-build.sh), so
# `scripts/feeds update -a` fetches deltas instead of re-cloning base+packages+
# luci from scratch (~7 min/run on this runner's slow github.com link).
# Top-level chmod only: the contents are created by the container's uid-1000
# build user and restored with the same ownership (tar-as-root preserves it).
FEEDS_CACHE="$REPO/.cache/feeds/opkg"
mkdir -p "$FEEDS_CACHE"
chmod a+rwX "$REPO/.cache" "$REPO/.cache/feeds" "$FEEDS_CACHE" 2>/dev/null || true
# --- 1) SDK package build (4 packages) in the arch-matched SDK image ----------
# We drive the `openwrt/sdk` docker image directly (not openwrt/gh-action-sdk):
# on a self-hosted Gitea act_runner the marketplace action fetch can be
# unavailable, and we need a CLEAN single-feed layout. `--volumes-from
# $(hostname)` shares THIS job container's workspace volume into the nested SDK
# container — a bare `-v $PWD:...` points at a host path that does not exist
# under the act_runner DinD setup. (Requires the job to run inside a container,
# which Gitea Actions does by default.)
echo "[feed] SDK build arch=$ARCH image=openwrt/sdk:$SDK_TAG"
docker pull "openwrt/sdk:$SDK_TAG"
docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
-e FEEDS_CACHE="$FEEDS_CACHE" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
"openwrt/sdk:$SDK_TAG" \
sh "$REPO/ci/sdk-build.sh"
# --- 2) index + sign the per-arch feed (usign, KEY_BUILD passed through) -------
sh "$REPO/ci/install-usign.sh"
KEY_BUILD="${KEY_BUILD:-}" bash "$REPO/ci/make-index.sh" "$OUT"
echo "[feed] done arch=$ARCH -> $OUT"
ls -l "$OUT"
+7 -10
View File
@@ -2,24 +2,21 @@
# ci/gen-apk-key.sh — generate the Shater **apk** feed signing keypair (25.12 lane).
#
# apk (OpenWrt/ImmortalWrt 25.12+) verifies package indexes with EC keys
# (prime256v1 PEM), NOT usign — the existing usign identity
# (dist/shater-feed.pub, fp 5ac4b177689cb8e0) keeps signing the opkg/24.10 feed
# and is NOT touched by this script. This generates a SEPARATE, second identity:
# (prime256v1 PEM). This is the ONLY feed identity shater has since the opkg
# lane was removed (D22) — the old usign key is history, not a second lane.
#
# dist/shater-apk.key EC PRIVATE key. NEVER commit (dist/ is gitignored).
# Paste its full PEM contents into the Gitea repo secret
# KEY_APK (the apk analog of the usign secret KEY_BUILD).
# Then delete the local file (or keep it in a password
# manager as the offline backup — losing it means every
# deployed router must re-trust a new key).
# dist/shater-apk.pem PUBLIC key. Commit it next to shater-feed.pub:
# KEY_APK. Then delete the local file (or keep it in a
# password manager as the offline backup — losing it
# means every deployed router must re-trust a new key).
# dist/shater-apk.pem PUBLIC key. Commit it:
# git add -f dist/shater-apk.pem
# (-f because /dist/ is gitignored). Routers install it
# as /etc/apk/keys/shater-apk.pem.
#
# Run ONCE. Refuses to overwrite: regenerating the key invalidates the trust of
# every router that already installed shater-apk.pem (same rule as D7 for the
# usign key).
# every router that already installed shater-apk.pem (see D22).
set -eu
REPO="$(cd "$(dirname "$0")/.." && pwd)"
-60
View File
@@ -1,60 +0,0 @@
#!/bin/bash
# Make `usign` available on the CI runner so ci/make-index.sh can sign the opkg
# feed index. The OpenWrt SDK ships usign, but the index/signing step runs on the
# bare runner (outside the SDK container), so we build the tiny standalone tool
# from source (no libubox — it is intentionally dependency-free so it can
# bootstrap a build system). No-op if usign is already on PATH.
#
# Ported unchanged from Shater v0.1 (ci/install-usign.sh): usign is
# format-agnostic and the signing story is identical for the v0.2 4-package feed.
#
# CI cache: a previously-built binary is reused from $USIGN_CACHE (default:
# <repo>/.cache/tools — a workspace dir the workflow persists via actions/cache),
# skipping the apt + cmake + clone + build (~1 min). After a fresh build the
# binary is copied there so the NEXT run hits the cache. usign is a tiny static
# helper with no versioned protocol — a stale cached binary cannot mis-sign.
set -eu
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
TOOLS="${USIGN_CACHE:-$REPO_ROOT/.cache/tools}"
# place <binary> — install onto PATH (system-wide if we can, else ~/bin)
place() {
local SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
if $SUDO install -m0755 "$1" /usr/local/bin/usign 2>/dev/null; then
:
else
mkdir -p "$HOME/bin"
install -m0755 "$1" "$HOME/bin/usign"
echo "$HOME/bin" >> "${GITHUB_PATH:-/dev/null}"
export PATH="$HOME/bin:$PATH"
fi
}
if command -v usign >/dev/null 2>&1; then
echo "[usign] already present: $(command -v usign)"
exit 0
fi
if [ -x "$TOOLS/usign" ]; then
place "$TOOLS/usign"
echo "[usign] restored from cache: $(command -v usign || echo "$HOME/bin/usign")"
exit 0
fi
SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
if ! command -v cmake >/dev/null 2>&1 || ! command -v cc >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y -qq cmake gcc git
fi
tmp="$(mktemp -d)"
# Canonical source; fall back to the GitHub mirror if git.openwrt.org is flaky.
git clone --depth 1 https://git.openwrt.org/project/usign.git "$tmp/usign" \
|| git clone --depth 1 https://github.com/openwrt/usign.git "$tmp/usign"
( cd "$tmp/usign" && cmake -DCMAKE_BUILD_TYPE=Release . >/dev/null && make >/dev/null )
place "$tmp/usign/usign"
# seed the cache for the next run (best-effort)
mkdir -p "$TOOLS" 2>/dev/null && install -m0755 "$tmp/usign/usign" "$TOOLS/usign" 2>/dev/null || true
echo "[usign] built: $(command -v usign || echo "$HOME/bin/usign")"
-39
View File
@@ -1,39 +0,0 @@
#!/bin/bash
# Build the opkg feed index (Packages + Packages.gz) with SHA256 for a dir of
# .ipk files, then optionally usign-sign it if $KEY_BUILD (the Gitea repo secret)
# is set and usign is present. Arg $1 = feed dir.
#
# Ported from Shater v0.1 (ci/make-index.sh), unchanged. It is package-count and
# package-name agnostic: it indexes whatever .ipk are in the dir, so it serves
# BOTH the per-arch feed built by ci/build-feed.sh AND the combined release feed
# assembled in the release job (shaterd + byedpi per-arch, shater-core +
# luci-app-shater = _all). opkg filters by Architecture at install time, so one
# combined URL serves every device.
#
# Feed format: opkg `src/gz` (.ipk + text Packages index, usign signature).
# OpenWrt 24.10 (our SDK) still uses opkg; apk arrives at 25.12. The committed
# trust anchor dist/shater-feed.pub is a usign (Ed25519) key, matching this.
set -euo pipefail
OUT="${1:?feed dir required}"; cd "$OUT"
: > Packages
for ipk in *.ipk; do
[ -e "$ipk" ] || continue
ctrl=$(tar -xzOf "$ipk" ./control.tar.gz | tar -xzO ./control)
sz=$(wc -c < "$ipk"); sha=$(sha256sum "$ipk" | cut -d' ' -f1)
printf '%s\n' "$ctrl" | sed '/^[[:space:]]*$/d' >> Packages
printf 'Filename: %s\nSize: %s\nSHA256sum: %s\n\n' "$ipk" "$sz" "$sha" >> Packages
done
gzip -kf Packages
if [ -n "${KEY_BUILD:-}" ]; then
# Signing was requested — a missing/broken signer must FAIL the build, not
# silently ship an unsigned feed that routers with check_signature on reject.
command -v usign >/dev/null 2>&1 || { echo "[index] ERROR: KEY_BUILD set but usign not found" >&2; exit 1; }
umask 077; printf '%s\n' "$KEY_BUILD" > /tmp/usign.sec
usign -S -m Packages -s /tmp/usign.sec || { rm -f /tmp/usign.sec; echo "[index] ERROR: usign signing failed" >&2; exit 1; }
rm -f /tmp/usign.sec
echo "[index] signed -> Packages.sig ($(head -1 Packages.sig))"
else
echo "[index] no KEY_BUILD -> UNSIGNED feed (opkg needs check_signature off, or set the secret)"
fi
echo "[index] contents:"; ls -l
+3 -4
View File
@@ -10,7 +10,6 @@
# the target fleet (BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base, its
# distfeeds even point at downloads.immortalwrt.org/releases/25.12-SNAPSHOT) is
# ImmortalWrt — so we extract the official ImmortalWrt SDK tarball ourselves.
# Same --volumes-from workspace-sharing pattern as ci/sdk-build.sh (opkg lane).
#
# The OpenWrt buildsystem refuses to run as root, so the SDK build itself runs
# as an unprivileged `build` user created here.
@@ -36,8 +35,8 @@ echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallba
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
# The prebuilt shaterd artifact must already be staged for this arch (same
# contract as the opkg lane — scripts/build-shaterd.sh runs first).
# The prebuilt shaterd artifact must already be staged for this arch
# (artifact-order contract — scripts/build-shaterd.sh runs first).
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
@@ -113,7 +112,7 @@ export HOME=/home/build
cd "$SDKDIR"
# Register this repo's openwrt/ as a src-link feed named `shater` (absolute
# path required) — identical to the opkg lane (ci/sdk-build.sh).
# path required).
cp -f feeds.conf.default feeds.conf
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
-143
View File
@@ -1,143 +0,0 @@
#!/bin/sh
# Runs INSIDE an `openwrt/sdk:<target>-<ver>` container (CWD = SDK root
# /builder). The job's workspace is shared into this container via
# `docker run --volumes-from`, so the repo is visible at $REPO and output goes
# to $OUT (a dir under the repo, hence also visible to the runner afterwards).
#
# Unlike Shater v0.1 (which compiled ONLY xrayctl in the SDK and hand-packed the
# pure-data packages with tar), v0.2 builds ALL FOUR packages the canonical way,
# via the SDK feed + `make package/<p>/compile`:
#
# shaterd prebuilt binary — Build/Compile only VALIDATES that
# openwrt/shaterd/files/shaterd-<amd64|arm64>.upx was staged
# by scripts/build-shaterd.sh on the runner BEFORE this ran.
# (arch-specific .ipk: RSTRIP/STRIP disabled — packed ELF.)
# shater-core PKGARCH=all data glue (procd init, sysctl, uci-defaults).
# luci-app-shater PKGARCH=all LuCI thin launcher — its Makefile does
# `include $(TOPDIR)/feeds/luci/luci.mk`, so the `luci` feed
# MUST be updated first (that is what creates feeds/luci/luci.mk).
# byedpi arch-specific C — the SDK cross-compiles ciadpi from the
# upstream tarball (needs network for PKG_SOURCE_URL).
#
# Env (required): ARCH, REPO, OUT.
set -eu
ARCH="${ARCH:?ARCH env required}"
REPO="${REPO:?REPO env required}"
OUT="${OUT:?OUT env required}"
mkdir -p "$OUT"
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
# Package version, derived from the git tag by ci/version.sh and handed in by
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
# imports every environment variable as a variable, and it propagates through
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
# byedpi deliberately keeps its own upstream version (see its Makefile).
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
# The prebuilt shaterd artifact must already be staged for this arch.
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
*) echo "[sdk] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
esac
test -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" || {
echo "[sdk] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
echo " scripts/build-shaterd.sh must run on the runner before the SDK build."; exit 3; }
# --- register this repo's openwrt/ as a src-link feed named `shater` ---------
# src-link REQUIRES an absolute path; $REPO/openwrt is exactly a feed root (it
# contains the 4 package dirs and nothing else that looks like a package).
cp -f feeds.conf.default feeds.conf
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
# Update metadata for ALL feeds: our `shater` feed + the SDK defaults (base,
# luci, packages, routing, telephony). We need `luci` for feeds/luci/luci.mk and
# `base`/`packages` for the runtime deps (kmod-nft-tproxy, kmod-nft-socket,
# ip-full, rpcd, luci-base) to resolve.
#
# Persistent feeds checkouts: $FEEDS_CACHE (a workspace dir the runner restores
# via actions/cache, shared into this container via --volumes-from) replaces
# the SDK's ephemeral feeds/ dir, so `feeds update` git-fetches deltas instead
# of re-cloning base+packages+luci every run (~7 min on the runner's slow
# github.com link). Correctness-safe: update always checks out feeds.conf's
# pinned revisions; if it ever fails on a cached checkout (e.g. a force-pushed
# upstream), the cache is wiped and the update retried with fresh clones.
if [ -n "${FEEDS_CACHE:-}" ] && mkdir -p "$FEEDS_CACHE" 2>/dev/null; then
rm -rf feeds
ln -s "$FEEDS_CACHE" feeds
echo "[sdk] feeds/ -> $FEEDS_CACHE (persistent cache)"
fi
echo "[sdk] feeds update -a"
if ! ./scripts/feeds update -a; then
[ -L feeds ] || { echo "[sdk] ERROR: feeds update failed"; exit 8; }
echo "[sdk] WARNING: feeds update failed on cached checkouts — wiping cache, cloning fresh"
find "$FEEDS_CACHE" -mindepth 1 -maxdepth 1 -exec rm -rf {} + 2>/dev/null || true
./scripts/feeds update -a
fi
echo "[sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
# Select our packages, then defconfig. `make package/<p>/compile` builds the
# explicit target regardless, but selecting first makes deps visible to defconfig.
for p in shaterd shater-core byedpi luci-app-shater; do
echo "CONFIG_PACKAGE_$p=m" >> .config
done
# Route source downloads through OpenWrt's fast CDN mirror FIRST — sourceware.org
# (elfutils) and other upstreams intermittently stall mid-transfer, and curl's
# --connect-timeout doesn't cover a stalled stream, so the SDK download hangs the
# build. LOCALMIRROR is tried before each package's own PKG_SOURCE_URL. (lx CI)
echo 'CONFIG_LOCALMIRROR="https://sources.cdn.openwrt.org"' >> .config
# Persistent dl/ across runs: $DL_DIR is a workspace dir the runner restores via
# actions/cache (see ci/build-feed.sh). Correctness-safe: the buildroot verifies
# PKG_HASH on every file already in dl/ and re-downloads on mismatch, so a stale
# cache can never leak a wrong source into the build.
if [ -n "${DL_DIR:-}" ]; then
echo "CONFIG_DOWNLOAD_FOLDER=\"$DL_DIR\"" >> .config
fi
echo "[sdk] defconfig"
make defconfig >/dev/null
# --- compile the 4 packages --------------------------------------------------
for p in shaterd shater-core byedpi luci-app-shater; do
echo "[sdk] === build $p ==="
make "package/$p/compile" V=s -j"$(nproc)"
done
# --- collect ONLY our 4 packages' .ipk (per-arch shaterd/byedpi + _all core/luci)
# NOT `find bin -name '*.ipk'`: the openwrt/sdk image ships HUNDREDS of prebuilt
# kmod/base .ipk under bin/, which a blanket copy would pull into the feed and
# get signed under OUR key. Match each package's own `<name>_<ver>_<arch>.ipk`.
found=0
for p in shaterd shater-core byedpi luci-app-shater; do
for ipk in $(find bin -type f -name "${p}_*.ipk"); do
cp -f "$ipk" "$OUT/"; found=$((found+1))
done
done
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
# --- assert the tag-derived version actually reached the packages -------------
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
# env -> make hand-off has several layers (docker -e, make's env import, the
# metadata dump), so verify the result instead of trusting it: every one of our
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
echo "[sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[sdk] version check OK — our 3 packages are $want"
fi
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
ls -l "$OUT"
+7 -9
View File
@@ -6,9 +6,9 @@
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
# feed's version string differs from the installed one, `apk update` saw nothing
# new and the routers could not be updated through the normal path at all.
# vs r2's 5 488 336 B). Since apk offers an upgrade only when the feed's version
# string differs from the installed one, `apk update` saw nothing new and the
# routers could not be updated through the normal path at all.
#
# So the version is now DERIVED, in CI, from the git tag, and the package
# Makefiles only carry a fallback for manual/offline builds.
@@ -21,14 +21,12 @@
# rolling `latest`)
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
#
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
# part first (numerically, component by component), the `r<rel>` only as a
# tie-break. Verified against the real tools, not from memory:
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# apk compares `<upstream>-r<rel>` as: the dotted upstream part first
# (numerically, component by component), the `r<rel>` only as a tie-break.
# Verified against the real tool, not from memory —
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
# identical results (opkg implements the Debian algorithm).
# That is exactly the ordering this scheme needs:
# * a release always outranks every rolling build that preceded it
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
@@ -0,0 +1,35 @@
//go:build darwin
package dialer
import (
"syscall"
"testing"
"golang.org/x/sys/unix"
)
// udpSocketDFSet reports whether the socket has "don't fragment" forced on
// (control.DisableUDPFragment sets IP_DONTFRAG=1 on darwin).
func udpSocketDFSet(t *testing.T, sysConn syscall.Conn) bool {
t.Helper()
rawConn, err := sysConn.SyscallConn()
if err != nil {
t.Fatal(err)
}
var (
value int
sockErr error
ctrlErr error
)
ctrlErr = rawConn.Control(func(fd uintptr) {
value, sockErr = unix.GetsockoptInt(int(fd), unix.IPPROTO_IP, unix.IP_DONTFRAG)
})
if ctrlErr != nil {
t.Fatal(ctrlErr)
}
if sockErr != nil {
t.Fatal(sockErr)
}
return value != 0
}
@@ -0,0 +1,36 @@
//go:build linux
package dialer
import (
"syscall"
"testing"
"golang.org/x/sys/unix"
)
// udpSocketDFSet reports whether the socket has "don't fragment" forced on
// (control.DisableUDPFragment sets IP_MTU_DISCOVER=IP_PMTUDISC_DO on linux,
// the same flag the user-visible failure was traced to on android).
func udpSocketDFSet(t *testing.T, sysConn syscall.Conn) bool {
t.Helper()
rawConn, err := sysConn.SyscallConn()
if err != nil {
t.Fatal(err)
}
var (
value int
sockErr error
ctrlErr error
)
ctrlErr = rawConn.Control(func(fd uintptr) {
value, sockErr = unix.GetsockoptInt(int(fd), unix.IPPROTO_IP, unix.IP_MTU_DISCOVER)
})
if ctrlErr != nil {
t.Fatal(ctrlErr)
}
if sockErr != nil {
t.Fatal(sockErr)
}
return value == unix.IP_PMTUDISC_DO
}
@@ -0,0 +1,14 @@
//go:build !darwin && !linux && !windows
package dialer
import (
"syscall"
"testing"
)
func udpSocketDFSet(t *testing.T, _ syscall.Conn) bool {
t.Helper()
t.Skip("DF socket-flag introspection implemented for darwin, linux and windows only")
return false
}
@@ -0,0 +1,43 @@
//go:build windows
package dialer
import (
"syscall"
"testing"
"golang.org/x/sys/windows"
)
// IP_MTU_DISCOVER on windows (ws2ipdef.h); control.DisableUDPFragment sets it to
// IP_PMTUDISC_DO, the same "don't fragment" state the linux helper checks.
const (
windowsIPMTUDiscover = 71
windowsPMTUDiscDo = 1
)
// udpSocketDFSet reports whether the socket has "don't fragment" forced on.
// shater addition: upstream ships linux + darwin only, so the whole suite
// skipped on the dev host — where it is the one platform we can actually run it
// on before the router build.
func udpSocketDFSet(t *testing.T, sysConn syscall.Conn) bool {
t.Helper()
rawConn, err := sysConn.SyscallConn()
if err != nil {
t.Fatal(err)
}
var (
value int
sockErr error
)
ctrlErr := rawConn.Control(func(fd uintptr) {
value, sockErr = windows.GetsockoptInt(windows.Handle(fd), windows.IPPROTO_IP, windowsIPMTUDiscover)
})
if ctrlErr != nil {
t.Fatal(ctrlErr)
}
if sockErr != nil {
t.Skip("IP_MTU_DISCOVER is not readable on this host: ", sockErr)
}
return value == windowsPMTUDiscDo
}
+99
View File
@@ -0,0 +1,99 @@
// lx: regression tests for the udp_fragment / UDPFragmentDefault
// plumbing. The WireGuard endpoint (and MASQUE outbound) rely on
// UDPFragmentDefault=true reaching the real UDP socket as "DF clear": with DF
// set, an outer datagram larger than the path MTU is silently dropped instead
// of fragmented, which blackholes nested tunnels (AWG-over-AWG, MASQUE-over-AWG)
// and AWG s4 transport junk. These tests assert the socket flag itself, on both
// paths a WireGuard bind can take: the dialer (ClientBind, detour case) and the
// listener control (StdNetBind via WireGuardControl, no-detour case).
package dialer
import (
"context"
"net"
"syscall"
"testing"
"github.com/sagernet/sing-box/option"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
)
func dialUDPForDF(t *testing.T, options option.DialerOptions) syscall.Conn {
t.Helper()
d, err := NewDefault(context.Background(), options)
if err != nil {
t.Fatal(err)
}
conn, err := d.DialContext(context.Background(), N.NetworkUDP, M.ParseSocksaddr("127.0.0.1:9"))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = conn.Close() })
sysConn, isSysConn := conn.(syscall.Conn)
if !isSysConn {
t.Fatalf("dialed UDP conn %T does not expose SyscallConn", conn)
}
return sysConn
}
func listenUDPForDF(t *testing.T, options option.DialerOptions) syscall.Conn {
t.Helper()
d, err := NewDefault(context.Background(), options)
if err != nil {
t.Fatal(err)
}
// WireGuardControl() is the listener control conn.StdNetBind installs on the
// socket a no-detour WireGuard endpoint sends its outer datagrams from — the
// exact socket the DF default decides the fate of.
listenConfig := net.ListenConfig{Control: d.WireGuardControl()}
packetConn, err := listenConfig.ListenPacket(context.Background(), "udp4", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = packetConn.Close() })
sysConn, isSysConn := packetConn.(syscall.Conn)
if !isSysConn {
t.Fatalf("listened UDP conn %T does not expose SyscallConn", packetConn)
}
return sysConn
}
// Upstream default: no UDPFragmentDefault, no udp_fragment → DF is set on both
// the dial and listener paths. Pins the baseline the endpoint fix opts out of.
func TestUDPFragmentDFByDefault_LX(t *testing.T) {
if !udpSocketDFSet(t, dialUDPForDF(t, option.DialerOptions{})) {
t.Fatal("default dialer must set DF on dialed UDP sockets")
}
if !udpSocketDFSet(t, listenUDPForDF(t, option.DialerOptions{})) {
t.Fatal("default dialer must set DF on listener-control UDP sockets")
}
}
// UDPFragmentDefault=true (what the WireGuard endpoint and MASQUE outbound now
// set) → DF clear on both paths, so oversize outer datagrams fragment instead
// of vanishing.
func TestUDPFragmentDefaultClearsDF_LX(t *testing.T) {
options := option.DialerOptions{UDPFragmentDefault: true}
if udpSocketDFSet(t, dialUDPForDF(t, options)) {
t.Fatal("UDPFragmentDefault=true must leave DF clear on dialed UDP sockets")
}
if udpSocketDFSet(t, listenUDPForDF(t, options)) {
t.Fatal("UDPFragmentDefault=true must leave DF clear on listener-control UDP sockets")
}
}
// Explicit user config always wins over the protocol default, in both
// directions.
func TestUDPFragmentExplicitOverride_LX(t *testing.T) {
fragmentOff := false
options := option.DialerOptions{UDPFragment: &fragmentOff, UDPFragmentDefault: true}
if !udpSocketDFSet(t, dialUDPForDF(t, options)) {
t.Fatal("udp_fragment=false must set DF even when the protocol default allows fragmentation")
}
fragmentOn := true
options = option.DialerOptions{UDPFragment: &fragmentOn}
if udpSocketDFSet(t, dialUDPForDF(t, options)) {
t.Fatal("udp_fragment=true must leave DF clear even without a protocol default")
}
}
+18
View File
@@ -25,6 +25,21 @@ func requireRoot(t *testing.T) {
}
}
// requireTCPDump skips when tcpdump is not installed.
//
// The same honesty this package's callers demand of a health reading: a missing
// INSTRUMENT is "not checked", never "broken". Without it every test in this
// file fails on `cmd.Start()` — sixteen red results that say nothing about the
// code and hide any real failure among them — on a machine where the only thing
// wrong is that a capture tool is absent. requireRoot has always drawn that line
// for privileges; this draws it for the tool.
func requireTCPDump(t *testing.T) {
t.Helper()
if _, err := exec.LookPath("tcpdump"); err != nil {
t.Skip("integration test requires tcpdump on PATH; install it to run this suite")
}
}
func tcpdumpObserver(t *testing.T, iface string, port uint16, needle string, do func(), wait time.Duration) bool {
t.Helper()
return tcpdumpObserverMulti(t, iface, port, []string{needle}, do, wait)[needle]
@@ -36,6 +51,9 @@ func tcpdumpObserver(t *testing.T, iface string, port uint16, needle string, do
// the wire.
func tcpdumpObserverMulti(t *testing.T, iface string, port uint16, needles []string, do func(), wait time.Duration) map[string]bool {
t.Helper()
// Every capture in this file funnels through here, so one guard covers the
// whole suite and no future test can forget it.
requireTCPDump(t)
ctx, cancel := context.WithTimeout(context.Background(), wait)
defer cancel()
cmd := exec.CommandContext(ctx, "tcpdump", "-i", iface, "-n", "-A", "-l",
+141
View File
@@ -0,0 +1,141 @@
// lx:begin health-board
package urltest
import (
"strconv"
"strings"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/adapter"
)
// captureEvictions swaps the eviction notice sink for the duration of a test and
// returns a func that reads back everything reported.
func captureEvictions(t *testing.T) func() []string {
t.Helper()
var (
mu sync.Mutex
msgs []string
)
orig := boardEvictionLog
boardEvictionLog = func(m string) {
mu.Lock()
msgs = append(msgs, m)
mu.Unlock()
}
t.Cleanup(func() { boardEvictionLog = orig })
return func() []string {
mu.Lock()
defer mu.Unlock()
return append([]string(nil), msgs...)
}
}
// TestBoardHoldsAGenerationWithoutEvicting is the "what it holds" half of the
// bound. A live generation on this box is ~1200 tags (≈380 nodes plus their
// per-group egress copies and chain hops); the board must carry that — and a
// second generation's worth of overlap during a subscription rename — with no
// eviction at all, or the ceiling would be silently degrading real health data.
func TestBoardHoldsAGenerationWithoutEvicting(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
const generation = 1200
for gen := 0; gen < 2; gen++ {
for i := 0; i < generation; i++ {
s.StoreURLTestHistory("gen"+strconv.Itoa(gen)+"-node-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 20})
}
}
if got := s.Evicted(); got != 0 {
t.Fatalf("two full generations (%d tags) evicted %d entries; the board must hold them",
2*generation, got)
}
if msgs := read(); len(msgs) != 0 {
t.Fatalf("unexpected eviction notices: %v", msgs)
}
// Everything is still readable.
if s.LoadURLTestHistory("gen0-node-0") == nil {
t.Fatalf("the first tag of the first generation was lost without an eviction")
}
}
// TestBoardEvictsOldestAndSaysSo is the "what happens when it overflows" half.
// Overflow must (a) actually bound the map, (b) drop the LEAST RECENTLY MEASURED
// tags — on this box, exactly the ones no config names any more — and (c) be
// audible: a silent eviction is a health board quietly forgetting nodes it is
// still being asked about.
func TestBoardEvictsOldestAndSaysSo(t *testing.T) {
read := captureEvictions(t)
s := NewHistoryStorage()
base := time.Now().Add(-24 * time.Hour)
// Stale generation first: measured a day ago, nothing since.
const stale = 1500
for i := 0; i < stale; i++ {
s.StoreURLTestHistory("stale-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: base.Add(time.Duration(i) * time.Millisecond), Delay: 30})
}
if s.Evicted() != 0 {
t.Fatalf("evicted before the ceiling was reached")
}
// Now push past the ceiling with fresh measurements.
for i := 0; i <= maxBoardEntries; i++ {
s.StoreURLTestHistory("fresh-"+strconv.Itoa(i),
&adapter.URLTestHistory{LastOK: time.Now(), Delay: 15})
}
if got := s.Evicted(); got == 0 {
t.Fatalf("board grew past %d entries without evicting anything — it is still unbounded", maxBoardEntries)
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("board holds %d entries, above the %d ceiling", size, maxBoardEntries)
}
// The day-old generation is what went, not the fresh one.
if s.LoadURLTestHistory("stale-0") != nil {
t.Fatalf("the oldest observation survived while newer ones were dropped")
}
if s.LoadURLTestHistory("fresh-"+strconv.Itoa(maxBoardEntries)) == nil {
t.Fatalf("the newest measurement was evicted")
}
msgs := read()
if len(msgs) == 0 {
t.Fatalf("entries were evicted with no notice — eviction must never be silent")
}
m := msgs[0]
for _, want := range []string{"health board full", "evicted", "re-probed"} {
if !strings.Contains(m, want) {
t.Fatalf("eviction notice %q does not say %q", m, want)
}
}
}
// TestBoardEvictionThroughMarkFailed pins the OTHER write path. MarkFailed is how
// a dead node is recorded, and a flood of dead renamed nodes is exactly the shape
// of the leak — so it has to prune too, not just the success path.
func TestBoardEvictionThroughMarkFailed(t *testing.T) {
captureEvictions(t)
s := NewHistoryStorage()
for i := 0; i <= maxBoardEntries; i++ {
s.MarkFailed("dead-" + strconv.Itoa(i))
}
s.access.RLock()
size := len(s.delayHistory)
s.access.RUnlock()
if size > maxBoardEntries {
t.Fatalf("MarkFailed grew the board to %d, above the %d ceiling", size, maxBoardEntries)
}
if s.Evicted() == 0 {
t.Fatalf("MarkFailed never prunes — the failure path is still unbounded")
}
}
// lx:end health-board
+118
View File
@@ -10,11 +10,128 @@
package urltest
import (
"sort"
"strconv"
"sync"
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/log"
)
// --- board capacity ---------------------------------------------------------
//
// The board is the one structure in the daemon whose key space is chosen by
// somebody else. Its keys are outbound TAGS, and on this box a tag is a node
// NAME straight out of the subscription — plus the derived per-group egress
// copies ("group-<g>-m<i>-<node>") and per-chain hop copies the probe planner
// creates for the same nodes. Providers rename their nodes freely, so a daily
// subscription refresh introduces a whole new generation of keys, while the
// store itself is pinned to the ENGINE's context (shater/engine.New) and so
// outlives every generation and every Apply — by design, so health survives a
// config change.
//
// Nothing ever removed a key. DeleteURLTestHistory exists but no shater path
// calls it (only daemon/ and clashapi/, which this fork does not run), so the
// map was strictly append-only for the life of the process — and the process is
// expected to live for months.
//
// The arithmetic: ~380 nodes, and a config with a couple of egress-bound groups
// plus a handful of chains puts a LIVE generation at roughly 380 base tags +
// 2x380 group copies + ~100 chain copies ≈ 1200 keys. One new generation per day
// is ~440k keys a year, at ~200 B per entry (map bucket + a tag string that is
// routinely 30-50 B with flag emoji, + a 56 B URLTestHistory) ≈ 88 MB of a
// 512 MB box — spent entirely on nodes that no longer exist.
const (
// maxBoardEntries is the hard ceiling. 4096 is ~3.4 live generations, so the
// board comfortably holds the current config plus the overlap while a
// subscription refresh swaps names, and still costs under a megabyte. A tighter
// bound would start evicting tags the running config actually uses; a looser one
// would stop being a bound in any useful sense.
maxBoardEntries = 4096
// keepBoardEntries is the prune target: drop a quarter at a time so the
// O(n log n) selection is amortised over ~1024 inserts instead of running on
// every probe once the board is full.
keepBoardEntries = 3072
)
// boardEvictionLog reports an eviction. A package var so tests can capture it;
// production leaves it writing to the process log, which under procd is the same
// syslog/logsink stream every other daemon line lands in.
//
// Eviction is NEVER silent. It is not free either: an evicted tag reverts to
// "untested" and its next probe re-measures it, so a board that evicts entries
// belonging to the LIVE config is a board whose ceiling is too low — and the only
// way anyone finds that out is this line.
var boardEvictionLog = func(msg string) { boardLogger().Warn(msg) }
// pruneLocked drops the least-recently-OBSERVED entries when the board exceeds
// maxBoardEntries. "Least recently observed" is max(LastOK, LastFail): the entry
// nothing has measured for the longest is, on this box, precisely a tag that no
// longer exists in any config — a renamed node, a removed group copy, a retired
// chain hop. Caller holds access.
func (s *HistoryStorage) pruneLocked() {
if len(s.delayHistory) <= maxBoardEntries {
return
}
type kv struct {
tag string
seen time.Time
}
all := make([]kv, 0, len(s.delayHistory))
for tag, h := range s.delayHistory {
seen := h.LastOK
if h.LastFail.After(seen) {
seen = h.LastFail
}
all = append(all, kv{tag, seen})
}
sort.Slice(all, func(i, j int) bool { return all[i].seen.Before(all[j].seen) })
drop := len(all) - keepBoardEntries
var oldest time.Time
for i := 0; i < drop; i++ {
if i == 0 {
oldest = all[i].seen
}
delete(s.delayHistory, all[i].tag)
}
s.evicted += uint64(drop)
msg := "urltest: health board full (" + strconv.Itoa(maxBoardEntries) + " tags) — evicted " +
strconv.Itoa(drop) + " least-recently-measured entries (" + strconv.FormatUint(s.evicted, 10) +
" total since start); they revert to untested and will be re-probed"
if !oldest.IsZero() {
msg += "; oldest observation was " + time.Since(oldest).Truncate(time.Second).String() + " ago"
}
boardEvictionLog(msg)
}
// Evicted reports how many entries the capacity bound has dropped since the store
// was created. Nonzero means the board reached maxBoardEntries at least once.
func (s *HistoryStorage) Evicted() uint64 {
if s == nil {
return 0
}
s.access.RLock()
defer s.access.RUnlock()
return s.evicted
}
// boardLogger is the process-wide fallback logger for eviction notices. The store
// is built from a plain constructor with no logger in sight (box.New, the daemon,
// shater/engine all call NewHistoryStorage()), so rather than change that
// signature everywhere the notice goes to the standard logger — which on the
// router is the daemon's own stderr, i.e. the same sink logsink owns.
var (
boardLogOnce sync.Once
boardLog log.ContextLogger
)
func boardLogger() log.ContextLogger {
boardLogOnce.Do(func() { boardLog = log.StdLogger() })
return boardLog
}
// HealthVerdict classifies a stored history entry at read time.
type HealthVerdict int
@@ -54,6 +171,7 @@ func (s *HistoryStorage) MarkFailed(tag string) {
updated.Delay = previous.Delay
}
s.delayHistory[tag] = updated
s.pruneLocked()
s.notifyUpdated()
s.access.Unlock()
}
+61
View File
@@ -0,0 +1,61 @@
package urltest
// lx: health board §5.C — the reachability half of "should this be probed".
//
// # Two different reasons not to probe, and why they cannot be one flag
//
// A group's OWN probing schedule is stood down for two unrelated reasons, and
// conflating them breaks one of the two:
//
// - NOT USED — no enabled routing rule reaches this group, so probing it
// measures a path nothing travels. That is a property of the CONFIG, it is
// decided once when the config is generated, and it travels in the config
// itself (option.URLTestOutboundOptions.SelfCheck). It cannot change while
// the box runs, because the rules cannot change while the box runs.
//
// - NOT REACHABLE RIGHT NOW — the group is a hop of a chain and a hop in
// FRONT of it is currently dead. Every member of this group dials through
// that hop, so every probe would fail inside it: the measurement would be
// about the broken hop, and would be recorded against this one. That is a
// property of the WORLD, it changes minute by minute, and it must be
// re-asked every time rather than baked into the config — a hop that comes
// back must resume probing on its own, with no reapply and nobody pressing
// anything.
//
// ProbeGate is the second one. It is deliberately a QUESTION asked at the
// moment of probing and never a stored answer: there is no flag to set, so
// there is no flag to forget to clear.
//
// The gate governs the group's own SCHEDULE only — the warm-up sweep and the
// ticker. An explicit check (a human, an API call) is a deliberate request and
// is never refused, exactly as with SelfCheck.
type ProbeGate interface {
// ProbeAllowed reports whether the outbound tagged tag may run its own
// scheduled probe right now.
//
// Implementations MUST answer true when they do not know: a gate that
// refuses on missing information would silence probing precisely when the
// system has the least idea what is going on, and nothing would ever
// measure its way out of that. A nil ProbeGate means "no gate" and every
// probe proceeds.
ProbeAllowed(tag string) bool
// ProbeWhenIdle reports whether the outbound tagged tag must keep measuring
// even when no traffic is passing through it.
//
// A urltest group normally probes only while it is in use: Touch arms the
// ticker on a dial, and the idle timeout stops it again. That is right for a
// group whose readings matter only while somebody is dialling it, and wrong
// for one the routing config REACHES: a rule that matches rarely — a narrow
// domain list, say — is in force the whole time, so the health of its target
// is a live question the whole time. Letting it go quiet means the panel
// reports "untested" about a rule that is armed, and the first real request
// pays a cold probe instead of picking an already-known-good member.
//
// Unlike ProbeAllowed, the safe answer here is FALSE when nothing is known.
// This one ADDS work, and a gate that claimed it on missing information would
// keep every group in the process probing forever — not a default anybody
// asked for. Absent gate, unknown tag, nothing configured yet: false, and the
// idle timeout behaves exactly as it always has.
ProbeWhenIdle(tag string) bool
}
+9
View File
@@ -21,6 +21,10 @@ type HistoryStorage struct {
access sync.RWMutex
delayHistory map[string]*adapter.URLTestHistory
updateHooks []*observable.Subscriber[struct{}]
// evicted counts entries dropped by the capacity bound (board_lx.go). The map
// is keyed by outbound tags chosen by a subscription provider, so it needs a
// ceiling; see the comment on maxBoardEntries.
evicted uint64
}
func NewHistoryStorage() *HistoryStorage {
@@ -71,6 +75,11 @@ func (s *HistoryStorage) StoreURLTestHistory(tag string, history *adapter.URLTes
}
// lx:end health-board
s.delayHistory[tag] = history
// lx:begin health-board — the map is keyed by provider-chosen tags and the
// store outlives every engine generation, so it must bound itself here: no
// shater path ever calls DeleteURLTestHistory. See maxBoardEntries.
s.pruneLocked()
// lx:end health-board
s.notifyUpdated()
s.access.Unlock()
}
-2
View File
@@ -1,2 +0,0 @@
untrusted comment: shater feed signing key
RWRaxLF3aJy44JbcxSFujtrFFEQ8lIsnTkd1K5TdjIhdlC2c0wa0fv4V
+6
View File
@@ -126,6 +126,12 @@ func (t *HTTP3Transport) newTransport() *http3.Transport {
conn.Close()
return nil, dialErr
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-quicConn.Context().Done()
conn.Close()
}()
return quicConn, nil
},
TLSClientConfig: t.tlsConfig,
+351
View File
@@ -0,0 +1,351 @@
package quic
import (
"context"
"crypto/tls"
"net"
"net/http"
"net/url"
"sync"
"testing"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/quic-go/http3"
sbTLS "github.com/sagernet/sing-box/common/tls"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/dns"
"github.com/sagernet/sing-box/dns/transport"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/logger"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
mDNS "github.com/miekg/dns"
)
var _ N.Dialer = (*trackingDialer)(nil)
// These tests pin down who owns the UDP socket handed to quic-go.
//
// quic-go's Dial/DialEarly take a net.PacketConn but do NOT take ownership of
// it: quic.setupTransport() builds a Transport with createdConn=false, and
// Transport.Close() then only calls conn.SetReadDeadline(time.Now()) instead of
// conn.Close(). So every QUIC connection torn down here — idle timeout, a
// retryable error, an engine reload calling Reset() — used to strand the UDP
// socket that carried it for the rest of the process's life. On a router that
// resolves through DoQ/DoH3 for months that is an unbounded fd leak.
//
// Both tests reconnect once and assert the socket from the FIRST connection is
// actually closed. Without the `<-conn.Context().Done() -> rawConn.Close()`
// watchdogs in quic.go / http3.go they fail on that assertion.
type trackedConn struct {
net.Conn
closeOnce sync.Once
closed chan struct{}
}
func (c *trackedConn) Close() error {
c.closeOnce.Do(func() { close(c.closed) })
return c.Conn.Close()
}
// trackingDialer hands out real UDP sockets and remembers every one of them.
type trackingDialer struct {
access sync.Mutex
conns []*trackedConn
}
func (d *trackingDialer) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
conn, err := (&net.Dialer{}).DialContext(ctx, network, destination.String())
if err != nil {
return nil, err
}
tracked := &trackedConn{Conn: conn, closed: make(chan struct{})}
d.access.Lock()
d.conns = append(d.conns, tracked)
d.access.Unlock()
return tracked, nil
}
func (d *trackingDialer) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return net.ListenUDP("udp", nil)
}
func (d *trackingDialer) count() int {
d.access.Lock()
defer d.access.Unlock()
return len(d.conns)
}
func (d *trackingDialer) at(index int) *trackedConn {
d.access.Lock()
defer d.access.Unlock()
return d.conns[index]
}
func (d *trackingDialer) closeAll() {
d.access.Lock()
defer d.access.Unlock()
for _, conn := range d.conns {
conn.Close()
}
}
func requireClosed(t *testing.T, conn *trackedConn, what string) {
t.Helper()
select {
case <-conn.closed:
case <-time.After(5 * time.Second):
t.Fatalf("%s: the UDP socket of the retired QUIC connection was never closed — quic-go does not own it, we must", what)
}
}
func requireDialed(t *testing.T, dialer *trackingDialer, want int) {
t.Helper()
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if dialer.count() >= want {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Fatalf("expected at least %d dial(s), got %d", want, dialer.count())
}
func testServerTLSConfig(t *testing.T, nextProtos []string) *tls.Config {
t.Helper()
certificate, err := sbTLS.GenerateKeyPair(nil, nil, nil, "localhost")
if err != nil {
t.Fatal(err)
}
return &tls.Config{
Certificates: []tls.Certificate{*certificate},
NextProtos: nextProtos,
MinVersion: tls.VersionTLS13,
}
}
func testClientTLSConfig(t *testing.T, nextProtos []string) sbTLS.Config {
t.Helper()
config, err := sbTLS.NewClient(context.Background(), logger.NOP(), "localhost", option.OutboundTLSOptions{
Enabled: true,
Insecure: true,
ServerName: "localhost",
})
if err != nil {
t.Fatal(err)
}
config.SetNextProtos(nextProtos)
return config
}
// startDoQServer serves a minimal DoQ responder and returns its address.
func startDoQServer(t *testing.T) M.Socksaddr {
t.Helper()
listener, err := quic.ListenAddr("127.0.0.1:0", testServerTLSConfig(t, []string{"doq"}), nil)
if err != nil {
t.Fatal(err)
}
ctx, cancel := context.WithCancel(context.Background())
t.Cleanup(func() {
cancel()
listener.Close()
})
go func() {
for {
conn, acceptErr := listener.Accept(ctx)
if acceptErr != nil {
return
}
go func(conn *quic.Conn) {
for {
stream, streamErr := conn.AcceptStream(ctx)
if streamErr != nil {
return
}
go func(stream *quic.Stream) {
defer stream.Close()
request, readErr := transport.ReadMessage(stream)
if readErr != nil {
return
}
response := new(mDNS.Msg)
response.SetReply(request)
transport.WriteMessage(stream, 0, response)
}(stream)
}
}(conn)
}
}()
return M.ParseSocksaddr(listener.Addr().String())
}
func testQuery() *mDNS.Msg {
message := new(mDNS.Msg)
message.SetQuestion("example.com.", mDNS.TypeA)
return message
}
func TestQUICTransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
serverAddr := startDoQServer(t)
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
dnsTransport := &Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeQUIC, "test-doq", nil),
dialer: dialer,
serverAddr: serverAddr,
tlsConfig: testClientTLSConfig(t, []string{"doq"}),
connection: transport.NewConnPool(transport.ConnPoolOptions[*quic.Conn]{
Mode: transport.ConnPoolSingle,
IsAlive: func(conn *quic.Conn) bool {
return conn != nil && !common.Done(conn.Context())
},
Close: func(conn *quic.Conn, _ error) {
conn.CloseWithError(0, "")
},
}),
}
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
// Retire the connection the way a retryable error or an engine reload does.
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
// The reconnect must still work, on a fresh socket.
if _, err := dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err := dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func TestHTTP3TransportClosesPacketConnOnReconnect(t *testing.T) {
t.Parallel()
mux := http.NewServeMux()
mux.HandleFunc("/dns-query", func(writer http.ResponseWriter, request *http.Request) {
message, err := readRequestMessage(request)
if err != nil {
writer.WriteHeader(http.StatusBadRequest)
return
}
response := new(mDNS.Msg)
response.SetReply(message)
rawResponse, err := response.Pack()
if err != nil {
writer.WriteHeader(http.StatusInternalServerError)
return
}
writer.Header().Set("Content-Type", transport.MimeType)
writer.Write(rawResponse)
})
listener, err := quic.ListenAddrEarly("127.0.0.1:0", testServerTLSConfig(t, []string{http3.NextProtoH3}), nil)
if err != nil {
t.Fatal(err)
}
server := &http3.Server{Handler: mux}
go server.ServeListener(listener)
t.Cleanup(func() {
server.Close()
listener.Close()
})
serverAddr := M.ParseSocksaddr(listener.Addr().String())
dialer := &trackingDialer{}
t.Cleanup(dialer.closeAll)
stdConfig := &tls.Config{
InsecureSkipVerify: true,
ServerName: "localhost",
NextProtos: []string{http3.NextProtoH3},
MinVersion: tls.VersionTLS13,
}
dnsTransport := &HTTP3Transport{
TransportAdapter: dns.NewTransportAdapter(C.DNSTypeHTTP3, "test-doh3", nil),
logger: logger.NOP(),
dialer: dialer,
destination: &url.URL{Scheme: "https", Host: "localhost", Path: "/dns-query"},
headers: http.Header{},
serverAddr: serverAddr,
tlsConfig: stdConfig,
}
dnsTransport.transport = dnsTransport.newTransport()
t.Cleanup(func() { dnsTransport.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("first exchange: ", err)
}
requireDialed(t, dialer, 1)
first := dialer.at(0)
dnsTransport.Reset()
requireClosed(t, first, "Reset()")
if _, err = dnsTransport.Exchange(ctx, testQuery()); err != nil {
t.Fatal("second exchange: ", err)
}
requireDialed(t, dialer, 2)
second := dialer.at(1)
if second == first {
t.Fatal("expected a new UDP socket for the reconnect")
}
if err = dnsTransport.Close(); err != nil {
t.Fatal(err)
}
requireClosed(t, second, "Close()")
}
func readRequestMessage(request *http.Request) (*mDNS.Msg, error) {
defer request.Body.Close()
rawMessage := make([]byte, 4096)
n, err := readFull(request.Body, rawMessage)
if err != nil {
return nil, err
}
var message mDNS.Msg
err = message.Unpack(rawMessage[:n])
if err != nil {
return nil, err
}
return &message, nil
}
func readFull(reader interface{ Read([]byte) (int, error) }, buffer []byte) (int, error) {
var total int
for total < len(buffer) {
n, err := reader.Read(buffer[total:])
total += n
if err != nil {
if total > 0 {
return total, nil
}
return total, err
}
}
return total, nil
}
+12
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"os"
"time"
"github.com/sagernet/quic-go"
"github.com/sagernet/sing-box/adapter"
@@ -117,6 +118,12 @@ func (t *Transport) Exchange(ctx context.Context, message *mDNS.Msg) (*mDNS.Msg,
rawConn.Close()
return nil, E.Cause(err, "establish QUIC connection")
}
// quic-go does not take ownership of the packet conn passed to
// DialEarly: when the connection ends it only stops reading.
go func() {
<-earlyConnection.Context().Done()
rawConn.Close()
}()
return earlyConnection, nil
})
if err != nil {
@@ -144,6 +151,11 @@ func (t *Transport) exchange(ctx context.Context, message *mDNS.Msg, conn *quic.
return nil, E.Cause(err, "open stream")
}
defer stream.CancelRead(0)
stopWatch := context.AfterFunc(ctx, func() {
stream.CancelRead(0)
_ = stream.SetWriteDeadline(time.Now())
})
defer stopWatch()
err = transport.WriteMessage(stream, 0, message)
if err != nil {
stream.Close()
+43
View File
@@ -12,6 +12,49 @@ as GitHub **pre-releases** and never become "Latest".
#### Unreleased (shater)
**`l3-honest-drop` — ICMP routed to an L4-only outbound is dropped, not
forged** — ships with `shaterd` (part of the shater L3 ingress,
`docs-shater/DECISIONS.md` D25), not as an lx release tag; recorded here because
it edits two upstream files. Without it the TUN stack answers an unroutable echo
ITSELF — sing-tun's `ICMPForwarder.HandlePacket` rewrites Echo→EchoReply
whenever the flow judgment comes back Accept (`stack_gvisor_icmp.go`) — so a
ping routed to vless/vmess/… would read as a working tunnel while the packet
never left the router.
* **`route/route.go` (`PreMatch`)** — the pre-match walk was renamed to
`preMatch` and the exported `PreMatch` became a thin FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for
`N.NetworkICMP`. An earlier version overrode `continueResult` inside
`preMatchFlow` instead; that covered only the exits reaching that function and
left three of the walk's own exits forging — the `prepareMatchMetadata` error
return, the sniff bail-outs, and the `default:` arm of the rule-action switch
(every action pre-match has no arm for: `hijack-dns`, `direct`, …). A guard on
the single return value cannot be outgrown by a new exit. `PreMatchBypass` is
folded in because sing-tun implements `ActionBypass` on the nfqueue plane only
— on the TUN path it lands in the same `default:` arm as Accept, i.e. forges.
* **`adapter/router.go` (`JudgeFlow`, the `!isPort` branch)** — ICMP returns
`ActionDrop` where it fell through to `ActionAccept`. Second line of defense:
`adapter.FlowOutbound` and `tun.Port` are distinct interfaces, and a drift
between them must not quietly re-enable the forged reply.
* **TCP/UDP behaviour is unchanged** — `PreMatchContinue` still means "take the
ordinary connection route" for both, `PreMatchBypass` still means bypass, and
the `!isPort` fallthrough still returns `ActionAccept` for them; pinned by
`route/prematch_icmp_lx_test.go` and `adapter/judgeflow_icmp_lx_test.go`
(both inside the marker), each ICMP case having an explicit TCP/UDP twin.
* **NOT covered: a FRAGMENTED echo to a WireGuard/AWG outbound is still
forged** — sing-tun's `ForwardDispatcher.Dispatch` returns before asking for a
verdict at all when `parsed.fragment`, and the reassembled packet reaches
`ICMPForwarder.HandlePacket`, whose `installFlow` demands an UNSPECIFIED port
address that a WireGuard endpoint never has. Fixing it inside `JudgeFlow`
is NOT possible — both consumers call it with identical arguments and the
working path needs the concrete address. Full chain, the two viable fixes and
the trap are in `docs-shater/DECISIONS.md` D25, "KNOWN HOLE".
* **Rebase cost: two small marked blocks** (`lx:begin/end l3-honest-drop`, a
wrapper function in `route/route.go` and one branch body in
`adapter/router.go`) plus the two self-contained test files — carried across
an upstream rebase by eye. Note that `PreMatch`'s own body now lives in
`preMatch`, so an upstream change to the walk applies to that function.
**Fork-layer + control-plane rework of proxy health** — ships with `shaterd`
(the shater router daemon), not as an lx release tag; recorded here because the
load-bearing half lives in fork zones (`common/urltest`, `protocol/group`).
+10 -11
View File
@@ -31,12 +31,11 @@ Do not delete it — we port proven pieces from it. What v0.1 has:
- **`luci-app-shater`** — a custom "instrument panel" LuCI app (client-side JS +
ucode/rpcd ubus backend): Overview with a live Signal Path, Simple/Advanced
toggle, quick-start wizard, Nodes/Subs/Rules/DNS/Live/Profiles/Settings pages.
- **CI + signed opkg feed** on Gitea: builds per-arch, signs the feed index with
usign, publishes a rolling `latest` Gitea release consumable as `src/gz`. **Feed
signing key fingerprint `5ac4b177689cb8e0`**; public key `dist/shater-feed.pub`,
secret in the Gitea repo secret `KEY_BUILD`.
- **CI + a signed package feed** on Gitea: builds per-arch, signs the feed index,
publishes a rolling `latest` Gitea release the router consumes as a feed.
(v0.1 shipped `.ipk` signed with a usign key — that lane is retired, D22.)
- Verified end-to-end on the VM: real LAN client proxied, DNS anti-leak, honest
fail-closed, opkg install/upgrade from the signed feed.
fail-closed, install/upgrade from the signed feed.
v0.1 is engine-locked to **xray-core**; its generator, share-link parser and
`run.json` are xray-shaped.
@@ -91,7 +90,7 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
- **`shater` branch `v0.1`** = the standalone xray-based version (frozen, ported
from).
- Until Phase 1 merges the engine in, `main` is the docs-first overlay seed you
are reading now (LICENSE, README, `docs-shater/`, `dist/shater-feed.pub`).
are reading now (LICENSE, README, `docs-shater/`, the feed signing key).
## What to port from v0.1 (don't rewrite these ideas)
@@ -105,8 +104,8 @@ overlay, don't redo:
- **Subscription fetch** (HAPP emulation, fingerprint reconcile, per-sub cache)
and the flexible **ruleset/list** model — though sing-box has its own share-link
parser and config schema we now target.
- **CI feed build + usign signing + Gitea release** (adapt to the single forked
binary; keep key `5ac4b177689cb8e0`).
- **CI feed build + index signing + Gitea release** (adapted to the single forked
binary; the format is apk, signed with the EC key — D22).
- The LuCI **design system** (the "instrument panel" identity) — reused for the
mini-dashboard and as the panel's visual language.
@@ -122,9 +121,9 @@ filter/stats engine wired into sing-box's DNS.
`https://github.com/SagerNet/sing-box`).
- **CI:** Gitea Actions (act_runner + Docker). v0.1's workflow was removed from
`main`; new CI is added when the v0.2 build exists.
- **Feed signing:** usign key `5ac4b177689cb8e0`; secret in repo secret
`KEY_BUILD`; public key `dist/shater-feed.pub` (kept so existing installs keep
verifying).
- **Feed signing:** EC (prime256v1) key for the apk index; secret in the repo
secret `KEY_APK`; public key `dist/shater-apk.pem`, installed on routers as
`/etc/apk/keys/shater-apk.pem`. Never regenerate it (D22).
- **Test VM:** OpenWrt 24.10.3 x86_64 in Docker (`docker ps --filter
name=openwrt-vm`). SSH via the ssh-manager MCP server `local_openwrt`
(localhost:2222, root/openwrt). LuCI at `http://127.0.0.1:8080` (root/openwrt),
+779 -1
View File
@@ -64,11 +64,17 @@ sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files m
stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is
GPL-3.0 for clarity.
## D7 — Keep the v0.1 feed signing identity
## D7 — Keep the v0.1 feed signing identity *(SUPERSEDED by D22)*
The usign feed key `5ac4b177689cb8e0` (public key in `dist/shater-feed.pub`,
secret in Gitea secret `KEY_BUILD`) carries over, so routers that already trust it
keep verifying v0.2 packages. Do not regenerate it without a documented rotation.
> **Superseded 2026-07-25 (D22).** The opkg feed this identity signed no longer
> exists, so there is nothing left for the key to verify. It was never rotated or
> compromised — it is simply unused. `dist/shater-feed.pub` was deleted from the
> tree; the reasoning, and how to resurrect the identity if it is ever needed
> again, is in D22.
## D8 — Preserve, don't destroy: v0.1 lives on its branch
The reset moved the full working xray-based project to the `v0.1` branch and
cleaned `main`. Nothing is lost; reusable logic (reliability layer, nft/routing,
@@ -95,6 +101,10 @@ runtime, forcing an ELF with `PT_INTERP=/lib64/ld-linux-x86-64.so.2` + `PT_DYNAM
plane is tproxy/redirect (netplane); generate never emits a tun inbound, so
the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever
appears it falls back to the system stack — re-add the tag then.
**REVERTED 2026-07-25 — that reasoning was wrong and shipped a dead feature.**
gVisor is not only the tun stack: it is the netstack of the **WireGuard
endpoint**, which we do emit and do declare [MVP]. See D23; the tag is back and
is now held there by a test.
- 2026-07-23: `with_clash_api` also dropped. The admin panel is shater's own
web server and generate never emits a `clash_api` service; the desktop/CLI
`LX_TAGS` keeps the tag for external dashboards.
@@ -309,6 +319,10 @@ Three values, not two, because the leaks differ in *kind*: an ICMP echo is ephem
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
*(Refined 2026-07-26 by D25: still true of TPROXY — but ICMP echo now has an
opt-in data plane of its own, the dedicated L3 TUN, so the policy no longer
speaks alone for ping; it keeps sole charge of ESP/GRE/IGMP and of the degraded
paths.)*
**Fail-open degradations must be visible in the panel, not only in `logread`.** The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
@@ -424,3 +438,767 @@ on the next render.
Consequence: all delay numbers are comparable (least ping ranks apples against
apples), and group settings lose two footgun fields while Settings keeps the
two that actually govern every check.
## D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). `config rule` carried THREE ways to say
where traffic is going: `dst_domain` (an inline domain list), `dst_ip` (an inline
CIDR list) and `dst_ruleset` (a reference to a `config ruleset`). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a `.srs`, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
**Decision: `dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is
the only destination matcher.** `Src`, `dst_port` and `proto` are untouched —
they are not lists of destinations and have no rule-set form.
- **Rejected: keep the inline lists as a shorthand.** "One obvious way" is the
whole point; a shorthand that quietly means something different from the long
form (see the bare-entry trap) is worse than no shorthand.
- **Rejected: promote inline lists to rule-sets lazily at generate time.** The
config on disk would then not say what the router does, and the panel would
have to render a list the user cannot find or edit.
### The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (`dst_domain`) | in a rule-set (`entry`) | migrated to |
|--------------------|--------------------------|-------------------------|--------------|
| `example.com` | **exact host** | **host + subdomains** | `full:example.com` |
| `full:example.com` | exact host | exact host | unchanged |
| `suffix:example.com` / `.example.com` | host + subdomains | host + subdomains | unchanged |
| `keyword:ads` | substring | substring | unchanged |
| `regexp:^ads\.` | pattern | pattern *(added here)* | unchanged |
| `geosite:x` / `geoip:x` | inert (engine field removed) | inert (unknown prefix) | unchanged |
`shaterd migrate` (schema v1→v2, `shater/model/migrate.go`) creates one inline
`config ruleset` per rule that still carries a legacy list — `rule-<rule name>`
for domains, `rule-<rule name>-ip` for addresses — moves the entries across with
the conversion above, appends the new name to `dst_ruleset`, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
`rule-<name>-2`). `regexp:` support was added to inline rule-sets in the same
change precisely so the move can be lossless.
`geosite:`/`geoip:` entries are copied VERBATIM rather than promoted to a
`source=geosite` rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
**One deliberate semantic change, called out:** a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because `rule_set: [a, b]` matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
That AND→OR change is about the ENGINE's TCP/UDP path, and it deliberately does
**not** extend to the untunnelable-protocol plane (`shater/apply/untunnelable.go`,
the ping / IPTV / VPN-passthrough policy in nftables). There, a v1
`dst_domain + dst_ip` rule could never claim a packet that carries no domain, so
the plan skipped it; reading the migrated form as OR would have made the address
half suddenly decisive, and with `target=direct` that means an upgrade quietly
sending previously-tunnelled ICMP out with the client's real source address. A
rule whose rule-sets are known to match by NAME is therefore still skipped by that
plan, and the skip is reported ("a routing rule matches by name … as well as by
address"). Split the rule in two if you want the addresses decided there.
### The vocabulary is about ENTRIES YOU TYPE, not about every list body
The table above is the vocabulary of an **inline** rule-set's `entry` values (and
of the DNS-filter/device lists, which share the classifier). The other two rule-set
sources are not other spellings of it:
| source | what it is | vocabulary |
|------------------------------|----------------------------------|------------|
| `inline` | entries you type | the table above |
| `url` → `.srs` / `.json` | a compiled rule-set, engine-owned | the engine's, not ours |
| `url` → anything else | a hosts / one-domain-per-line / AdBlock TEXT FILE | **none** — every line is a domain plus its subdomains |
| `file` | a local `.srs` / `.json` | the engine's, not ours |
**Rejected: run text lists through the entry classifier too.** A published
AdGuard/OISD list is full of colon-bearing tokens that are ordinary filter syntax
(`##…:has(…)`, `$domain=`, absolute URLs); classifying them would either mis-import
them or bury the operator under hundreds of "unrecognised prefix" warnings per
list. The formats also disagree structurally — a hosts line carries several names,
so the text parser works per token, while an entry is a whole line. And `regexp:`
arriving from a third-party URL is a pattern compiled into the router's matcher and
evaluated per query, which is a very different proposition from one the operator
typed.
So the difference stands and is paid for in diagnostics instead: a text list
containing `full:` / `suffix:` / `keyword:` / `regexp:` is reported per list, on
every generate, naming the entries and pointing at `source=inline` where they work
(`warnListEntryVocabulary`, `shater/generate/ruleset.go`). The check tests only
those four markers, never the general `word:` shape, so it fires on a human's
mistake and stays quiet on published filter syntax.
Consequence: one destination mechanism, one vocabulary, one place a list is
edited; every list is compiled once and reused. The panel's rule editor drops its
Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of
the rulesets that already exist, and nothing more. Creating and filling a list
stays in the Rulesets panel — **rejected: a "create a list from here" shortcut in
the rule editor**, because a second place to author a list is a second place for
its semantics and its duplicate-name rules to drift, and the whole point of this
decision was to stop having two.
## D22 — One packaging lane: apk. The opkg/`.ipk` lane is deleted, not disabled
Decided 2026-07-25 (product owner). CI built and published TWO signed feeds from
every run: opkg/usign (`.ipk` + `Packages.gz`, OpenWrt 24.10) and apk/EC (`.apk` +
`packages.adb`, OpenWrt/ImmortalWrt 25.12). The opkg half served nobody. Checked
on the actual hardware, not inferred:
| Device | Firmware | pkg arch | package manager |
|---|---|---|---|
| `mini_router` (BPi-R3 Mini) | ImmortalWrt 25.12.1 | `aarch64_cortex-a53` | apk-tools 3.0.5 |
| `main_router` (BPi-R4) | OpenWrt 25.12.0 | `aarch64_cortex-a53` | apk-tools 3.0.5 — **no `opkg` binary on the system at all** |
**Decision: delete the opkg lane outright.** Removed: the `build` + `release`
jobs from `.gitea/workflows/release.yml`; `ci/build-feed.sh`, `ci/sdk-build.sh`,
`ci/make-index.sh`, `ci/install-usign.sh`; and the trust anchor
`dist/shater-feed.pub`. The Gitea secret `KEY_BUILD` is now referenced by
nothing and can be deleted from the repo settings. `ci/version.sh`,
`ci/gitea-release.sh` and `ci/fetch-sdk.sh` are shared or apk-only and stay.
- **Rejected: keep the lane but stop triggering it** (comment it out / gate it on
a dispatch input). Dead code in CI is worse than no code: it keeps a second SDK
matrix, a second signing key and a second feed layout alive in everyone's head
and in every future edit, and it silently rots because nothing runs it. The
24.10 SDK images it pins are themselves a frozen dependency.
- **Rejected: keep `dist/shater-feed.pub` as a historical artifact.** A committed
trust anchor is an instruction — it invites someone to follow the old install
path for a feed that is no longer produced. Nothing is lost by removing it:
git history still holds the file, the SECRET half is untouched in `KEY_BUILD`,
and a usign secret key blob contains its own public half, so the identity can
be reconstructed if a 24.10 device ever has to be served again. Deleting the
file is reversible; a stale trust anchor pointing at an unmaintained feed is
the thing that quietly misleads.
- **Not done: revoking or rotating the usign key.** There is no incident. It is
retired, not burned (D7).
Consequence: one SDK, one key, one feed layout, one set of install instructions.
It also makes the rolling release `apk-latest-<arch>` the *only* install path
that does not require hand-editing a file per release — which is why the same
change fixed it: publishing was an either/or (`apk-latest-<arch>` on dispatch,
ELSE `apk-vX.Y.Z-<arch>` on a tag), so once releases moved to tag pushes the
rolling pointer stopped being written and froze at `0.2.0` while v0.2.9/v0.2.10
shipped — routers on the rolling URL got a successful, silent `apk update` with
nothing new. `release-apk` now writes the rolling pointer on every run and
asserts, by reading the published release back over the Gitea API, that it holds
our three tag-versioned packages at exactly the version just built and no asset
at any other version.
## D23 — The router tag set is a checked contract, not a string literal
`with_gvisor` was trimmed from the router set on 2026-07-23 (D9) as "unreachable
code: we never emit a tun inbound". True about tun — and irrelevant, because
gVisor is also the netstack of the **WireGuard endpoint**, which shater emits and
FEATURES.md declares [MVP] (AmneziaWG is called *"a driving requirement"*). Every
binary shipped between then and 2026-07-25 answered a configured WireGuard node
with:
```
create instance: initialize endpoint[0]: create WireGuard device:
gVisor is not included in this build, rebuild with -tags with_gvisor
```
`transport/wireguard/device_stack_stub.go` (`//go:build !with_gvisor`) returns
`tun.ErrGVisorNotIncluded` from **both** device constructors, so
`system_interface: true` is not an escape hatch either: WireGuard was 100% dead
in the shipped artifact while the panel offered it, the parser accepted `wg://`,
`awg://` and wg-quick `.conf` imports, and the owner had 7 WireGuard sections in
UCI on a production router.
- **Decision:** `with_gvisor` is part of the router tag set and stays there for
as long as we ship WireGuard. It costs **~2.8 MB raw / ~0.65 MB UPX per arch**
(measured 2026-07-25, both arches; `/overlay` on the production router is
6.9 GB with 205 MB used). A tag whose absence turns a declared feature into a
runtime error is not "dead weight" — it is the feature.
### Why the bug was invisible, and what now makes it visible
The defect was not a typo in a tag list. It was that **nothing connected the tag
list to the feature list**, and the shipped tag combination was the one build
configuration nothing exercised: the whole test suite compiles with the FULL
upstream set (`with_gvisor` included), so `TestAmneziaWGEndpoint` passed happily
while the artifact it was supposed to vouch for could not create a WireGuard
device. Tests proved the code was right; they never proved the *build* was.
Three pieces now hold it together:
1. **One definition of the set** — `scripts/router-tags.sh` (`SHATER_ROUTER_TAGS`
+ `SHATER_ROUTER_LDFLAGS`), sourced by `scripts/build-shaterd.sh` and by the
checker. The tag list used to live as a literal inside the build script, i.e.
in a file no test reads. A second copy is a second truth.
2. **A declared-feature table** — `shater/buildtags`: every tag-gated capability
we promise, with the exact tags it needs *to run* and why (the code anchor).
`TestRouterTagSetCoversDeclaredFeatures` parses the shell file and fails if a
declared feature lost a tag. It needs no build tags, no Linux, no network and
no privileges, so it runs in every plain `go test ./...` — including on the
Windows dev host, where nothing else can see the shipped configuration.
3. **A construction test under the shipped tags** —
`shater/generate.TestShippedTagSetConstructsDeclaredProtocols` drives one node
of every declared protocol (ss/vmess/trojan/vless ws-grpc-httpupgrade-quic-
xhttp/REALITY/uTLS-fp/hysteria2/tuic/**wg**/**awg**) through `box.New`+`Start`.
`scripts/check-router-tags.sh` runs it **with `SHATER_ROUTER_TAGS`**, and CI
runs that script (`.gitea/workflows/release.yml`) *before* the artifact is
built. In a router-tag-set run nothing may be skipped: a protocol that is not
compiled in fails the run instead of quietly disappearing from it.
(2) catches a trim the moment it is made and names the feature it kills; (3)
catches what a list comparison cannot — a tag that is present but insufficient.
Neither is a substitute for the other. A new protocol in `shater/parse` +
`shater/generate` means a new row in `buildtags.Features` and a new probe case;
`TestEveryTagGatedFeatureIsProbed` fails until both exist.
- **Rejected: "just add the tag".** The one-line fix restores WireGuard and
leaves the mechanism that hid it fully intact — the next size-driven trim is
equally invisible. The tag is the smallest part of this decision.
- **Rejected: run the WHOLE test suite with the router tag set in CI.** It is the
obvious move and it does not work: parts of the suite legitimately depend on
upstream-only tags, and the run costs a second full compile of a 25 MB binary's
worth of packages on every release. A focused, unprivileged construction test
buys the same evidence for ~10 s and, unlike a full run, can be *required* to
skip nothing.
- **Rejected: assert the tag set against upstream's `DEFAULT_BUILD_TAGS`.** That
makes any trim a failure, which turns the check into noise and re-litigates D9
on every upstream rebase. The contract is with our own feature list, not with
upstream's.
- **Not done: dropping `with_lx_command`.** It is inert for `shaterd` — nothing
under `shater/` imports `sing-box/daemon` or `experimental/libbox`, and
`go list -deps ./shater/cmd/shaterd` links neither, so it costs zero bytes. It
stays only so the router set remains a subset of the lx desktop set. Noted
because "a tag that buys nothing" is the mirror image of this bug and should be
removed deliberately, not silently.
## D24 — DNS interception is the DEFAULT (`dns_intercept=1`), not an opt-in
Decided 2026-07-26. `Globals.DNSIntercept` shipped as opt-in (`default false`, and
absent from both `DefaultGlobals` and the shipped `/etc/config/shater`). The result
was an **inverted** posture, which is the reason this is a decision and not a
preference:
- a client with **standard** settings — DNS = the router's address, exactly what
DHCP hands out — sent its queries to the router. The nft `:53` divert was behind
the flag (`netplane/nft.go`), and the rule right after it is an unconditional
`fib daddr type local accept`, so the query was delivered locally to dnsmasq and
forwarded to the ISP **in the clear**: no blocklists, no per-device DNS rules,
no Block-DoH, no resolver detour, nothing;
- a client that hard-coded `8.8.8.8` "to bypass the router" was addressing a
non-local IP and **was** caught by the ordinary tproxy catch-all.
The obedient client leaked; the evader did not. Meanwhile `FEATURES.md`, `README.md`
and D14 all promised "no DNS leaks" and "dnsmasq never sees LAN queries" — true only
for the traffic pattern the default did not cover. `dns_intercept` appeared nowhere
in `docs-shater/` at all.
**Decision: `DNSIntercept` is seeded ON in `model.DefaultGlobals`, and the shipped
`/etc/config/shater` carries an explicit `option dns_intercept '1'`.** Nothing about
the interception MECHANISM changed — only which side of the switch is the default.
**`.lan` and the private PTR zones keep working, and that is a pre-existing part of
the mechanism, not something bolted on for this flip.** `generate/dns.go` adds a
synthetic DNS server (`shater-local-dns`, plain UDP to `127.0.0.1:53`, detour
`direct`, so the daemon's own loop-mark keeps it out of the divert) and PREPENDS a
`domain_suffix` rule for `lan` + the RFC6303 private reverse zones, ahead of every
device/filter rule. Two honest limitations: it hardcodes `lan` (a router whose
dnsmasq domain was changed needs a `config dns_rule` for the new suffix), and it
only exists when the model has at least one `config resolver` — with none, buildDNS
emits no DNS plane at all and the engine falls back to its built-in `local`
transport, which reads `/etc/resolv.conf` (127.0.0.1 → dnsmasq), so local names
still resolve but nothing is filtered.
**A dead engine does NOT black out the LAN's DNS.** This was the first thing checked,
because "intercept everything" invites the reading "engine down = no DNS anywhere",
and that is not what happens:
- the fail-closed **holding plane** (D17, `RenderHoldNft`) hooks `forward` ONLY.
A query addressed to the router is INPUT-hook traffic, so dnsmasq answers it as
it always did — unfiltered and plaintext to the ISP. Deliberate: blocking it
would also cut the daemon's own name resolution and with it any chance of
self-recovery;
- with the FULL plane loaded and the engine's tproxy socket gone, the `tproxy`
statement returns `NFT_BREAK`, which aborts its own rule; the packet continues
down the chain into the same `fib daddr type local accept` and reaches dnsmasq.
So the failure mode is a DNS **fail-open** (working, unfiltered) while client
TRAFFIC stays fail-closed — and a query aimed at an EXTERNAL resolver is dropped
with the rest of the forwarded traffic. Operators must know this: "the tunnel is
down" does not mean "DNS is private".
**Existing installs.** `/etc/config/shater` is a conffile
(`openwrt/shater-core/Makefile`), so an upgrade never replaces it:
- a config that never mentioned the option (all of them, before this change) now
parses over the ON seed and **starts intercepting on the next apply**. That is the
intended behaviour change, and the only one this decision makes;
- an explicit `option dns_intercept '0'` keeps winning. It survives the
`WriteUCI→ReadUCI` round-trip because `render.go` emits booleans ALWAYS —
the trap a default-true bool has and a default-false one does not: a value
omitted at false would come back as the seed and silently re-enable itself.
`shater/model/dnsintercept_test.go` pins both directions, plus the shipped file.
**Not done: silencing the "no resolvers configured" warning by shipping a resolver.**
With interception on and no `config resolver`, generate warns — and it is right to:
every client query now lands in an engine that has no resolver plane, so it is
answered by the system resolver (dnsmasq → the ISP, in the clear) with filtering and
anti-leak inert. Shipping a `type local` resolver would make the warning disappear
while changing nothing about where the queries go: the panel would show a configured
resolver and the operator would believe DNS was handled. That is the inverted lie
this project keeps deleting. The warning stays; what it needs is the accurate
wording (it currently claims `.lan` breaks, which the fallback above disproves), not
a workaround. Note also that a fresh install ships INERT (`enabled '0'`) and
`Reconcile` tears down instead of generating, so the warning cannot appear before the
operator has enabled the stack — at which point it describes their live config.
**OPEN, and it gates shipping this default: the synthetic local server changes how
proxy-endpoint DOMAINS are resolved.** Found while landing D24, reproduced on Linux
with one resolver and a node addressed by a hostname:
- `common/dialer/dialer.go` resolves a domain server address through
`route.default_domain_resolver`; when that is unset it uses
`dnsTransport.Default()` — the engine's built-in `local` transport, i.e. a
bootstrap-DIRECT lookup — but **only while fewer than two DNS transports exist**.
With two or more and no default, it reports the `missing-domain-resolver`
deprecation and leaves the query transport nil, so `dns.Router.Lookup` falls back
to `lookupWithRules`: the CLIENT DNS plane.
- `dns_intercept` adds `shater-local-dns`, which takes a single-resolver config from
one transport to two. So a config whose only resolver is DoH-through-the-tunnel —
the recommended anti-leak setup — would start resolving its own node's hostname
through that same tunnel: a bootstrap loop where there was none.
- Evidence: the same model emits no deprecation notice with `dns_intercept=0` and
two `missing-domain-resolver` notices with `dns_intercept=1`;
`generate.TestDNSFilterRemoteBlocklistHTTPClient` (Linux-only) fails on exactly
that notice and is deliberately left failing rather than relaxed.
The fix belongs in `generate` (`route.go:160` already sets
`route.default_domain_resolver` from `endpointResolver()`, which is opt-in and unset
by default): when buildDNS emits the synthetic local server and no endpoint resolver
is configured, `default_domain_resolver` must be pointed at a bootstrap-direct
server, which restores exactly the pre-D24 behaviour and clears the notice. Until
that lands, an operator can get the same result by setting `endpoint_resolver` to a
direct resolver. Note the hazard is **not** created by D24 — any config with two
resolvers has it today; the default merely makes it universal.
## D25 — L3 ingress: LAN ICMP rides a dedicated TUN through the tunnel, not a policy verdict
Decided 2026-07-26. D17 made everything TPROXY cannot divert an explicit policy
(`Globals.Untunnelable` = block | icmp | direct) — and its premise still holds:
kernel TPROXY delivers a packet by handing it to a listening SOCKET, and sockets
exist for TCP and UDP only, so an ICMP echo has nothing to be handed to. But a
policy can only choose between losing the packet and leaking it with the
client's real source address; neither ever puts a ping THROUGH the tunnel. This
decision adds the data plane D17 could not have: **`globals.l3_tunnel` (opt-in,
default off; `model.Globals.L3Tunnel`) opens a second, dedicated ingress — a TUN
device — and LAN ICMP enters the engine as raw IP packets**, where the ordinary
route rules pick an outbound exactly as for any flow. The policy is refined, not
repealed: it keeps sole charge of the protocols the engine cannot ingest at all,
and of the degraded paths (both below).
**The whole mechanism is one mark, one rule, one device — the TPROXY plane is
untouched.** The nft prerouting chain stamps `L3Mark` (= fwmark_base + 0x80,
`netplane/nft.go` `l3MarkOffset`) on LAN `ip protocol icmp` / `meta l4proto
ipv6-icmp` ONLY, and only after every local plane was already accepted
(fib-local, RFC1918/link-local/multicast daddr sets) and — for v6 — after a
unicast ND/NA carve-out, because one tunnelled neighbour probe is enough to take
the LAN's v6 plane down (`renderNft`, the L3 block). `addL3Routing`
(`netplane/apply.go`) binds that mark to a table (= table_base + 0x08) whose
only content is `default dev shater-l3`; del-then-add idempotent, and a failed
rule or route is a NAMED operator warning, never an apply abort. `generate`
emits the synthetic `l3-in` TUN inbound bound to exactly `netplane.L3Device`,
MTU 65535 (the largest IP datagram there can be, so the KERNEL can never
fragment on the way in — see "the device MTU is not a tunnel budget" below),
point-to-point /30 + /126 addresses from private space,
the v6 one only when `globals.ipv6` is on — and only next to a tproxy inbound:
the ingress rides the same LAN divert plane, and without one the TUN would sit
dark while the config claims ICMP is tunnelled, so it is skipped with a warning
(`generate/inbound.go`, `appendL3TunInbound`). `shater/registry` registers the
`tun` inbound type; that costs no new build tag and no meaningful size because
`with_wireguard` already requires `with_gvisor` (D23, `scripts/router-tags.sh`).
**`auto_route: false` is load-bearing, not a default we happened to keep.**
sing-box's auto_route rewrites the router's MAIN routing table — it would drag
everything the router itself sends (WAN traffic, DNS, the tunnel's own underlay)
into this TUN. The fwmark rule + dedicated table above is deliberately the ONLY
entrance, and disabling the feature can never strand a stale default route in
main (`generate/inbound.go`; pinned by `TestL3TunnelEmitsTunInbound`).
**`stack: "gvisor"` is a deliberate choice, and the tempting reason for it is
wrong.** It is TRUE that sing-tun's system stack answers an ICMP echo LOCALLY —
`processIPv4ICMP` rewrites Echo→EchoReply in place and swaps the addresses
(sing-tun `stack_system.go:648`; the v6 twin sits right under it). It is FALSE
that this makes the system stack unusable here: `dispatchIPv4`
(`stack_system.go:355-372`) hands the packet to the SAME `ForwardDispatcher`
first and only falls through to that forger for packets addressed to the TUN
itself, exactly as the gVisor filter does (`stack_gvisor_filter.go:52-113`).
Both stacks would forward. gvisor is chosen because it is already linked —
`with_wireguard` requires `with_gvisor` (D23), so it costs no tag and no new
code path — and because it is the combination the integration test actually
exercises. Do not re-derive this as "the system stack fakes ping": it fakes ping
only where the dispatcher declined the packet.
**The ceiling is ICMP echo, and it is upstream's dispatcher — NOT the netstack.**
This distinction matters because the netstack answer is the intuitive one and it
is wrong. On the forward path a WireGuard/AWG endpoint never consults gVisor at
all: `Endpoint.WritePackets` (`transport/wireguard/port.go:21-58`) reads the IP
version and the destination address and hands the raw bytes to
`wgDevice.InputPackets` — the protocol byte is never examined — and
`returnDeviceWrapper.Write` (`:127-157`) offers every decrypted packet to
`returnPath.ReturnPackets` before the stack sees it. WireGuard would carry ESP
today if anything handed it one. What refuses is `ForwardDispatcher`: its parser
sets `hasFlow` for TCP, UDP and ICMP echo alone (`flow_parse.go`,
`parseTransport`, the echo identifier serving as the pseudo-port), and
`createFlow` NATs through a port-shaped selector (`flow_dispatch.go:325`,
`allocateSelector`) that ESP, AH and GRE do not have. So ESP/AH/GRE/IGMP/SCTP
cannot enter the engine in ANY configuration and REMAIN on the D17 policy —
or on the kernel egress of D26, which sidesteps the dispatcher entirely. The nft
plane encodes the same boundary on purpose: it marks `icmp`/`ipv6-icmp` only,
never `l4proto != { tcp, udp }`, because a marked ESP packet would enter the
device and vanish — a black hole wearing a tunnel's name — instead of receiving
the policy's honest verdict (`netplane/nft.go`, the prerouting L3 comment).
**What works and what does not, read off the upstream source.** ping v4/v6 —
yes. Windows `tracert` — yes: the gVisor return path recognises
`ICMPv4TimeExceeded` and `ICMPv4DstUnreachable` alongside EchoReply and NATs
them back to the LAN client (`stack_gvisor_icmp.go:341+`, `returnPacket`). IPv6
traceroute — intermediate hops stay invisible: the v6 branch of the same
function accepts EchoReply only, so just the final destination answers. Several
LAN clients behind the one tunnel address are already solved upstream:
`ForwardDispatcher` NATs by echo identifier and rewrites the source to the
outbound's port address (`flow_dispatch.go:325+`, `createFlow`; `icmpFlowKey`) —
we wrote no NAT of our own.
**Which outbounds can carry it.** The contract is `adapter.FlowOutbound`
(= `Outbound` + `tun.Port` + `PreMatchFlow`, `adapter/outbound.go`). In-tree
implementors: the WireGuard/AWG endpoint (`protocol/wireguard`), `direct`
(`protocol/direct`), `bridge` (`protocol/bridge`), `tailscale`
(`protocol/tailscale`). Of those, the shaterd registry can construct only
WireGuard/AWG and direct (`shater/registry/registry.go` — bridge and tailscale
are not registered). Every proxy protocol — vless/vmess/trojan/shadowsocks/
hysteria2/tuic/socks/http/shadowtls — is L4-only and cannot. Recorded as a known
gap: `masque` is L3 by nature (CONNECT-IP; it builds a userspace gVisor stack
per tunnel, `protocol/masque/outbound.go`) but implements no `tun.Port` and is
not in the shater registry, so today it cannot carry the ingress. Wiring it up
is possible future work, not a promise.
**ICMP to an L4-only outbound is DROPPED, and that took patching upstream files
(the `lx:l3-honest-drop` delta — see `docs-lx/lx-changelog.md`).** In the gVisor
stack the fallthrough verdict is a forgery: `ICMPForwarder.HandlePacket` answers
the echo ITSELF (Echo→EchoReply + address swap) whenever the flow judgment comes
back Accept (`stack_gvisor_icmp.go:120`), and upstream maps "no flow route" to
exactly that Accept — so a ping routed to vless would read as tunnelled while
the packet died on the router. Two small marked hunks make the truth observable:
`route/route.go` wraps the whole pre-match walk — the walk itself became
`preMatch`, and the exported `PreMatch` is now a FUNNEL that rewrites
`PreMatchContinue` and `PreMatchBypass` to `PreMatchDrop` for `N.NetworkICMP` —
and `adapter/router.go` (`JudgeFlow`, the `!isPort` branch) returns `ActionDrop`
for ICMP where it fell through to `ActionAccept` — the second line of defense,
because `FlowOutbound` and `tun.Port` are distinct interfaces and a drift
between them must not quietly re-enable the forger. TCP/UDP verdicts are
byte-identical; `route/prematch_icmp_lx_test.go` and
`adapter/judgeflow_icmp_lx_test.go` pin both directions. The operator-facing
text says the same out loud (`shater/apply/warnings.go`): proxy-routed addresses
"cannot be pinged at all — deliberately".
> **Why a funnel and not an override inside the walk.** The first version of
> this delta overrode the pre-declared `continueResult` inside `preMatchFlow`
> and claimed to cover "every exit point of the function at once". It covered
> every exit of THAT function; the walk above it has exits of its own that never
> reach it — the `prepareMatchMetadata` error return (which arrived later, with
> the shared-metadata refactor, upstream `b911fb078`), the sniff bail-outs, and
> the `default:` arm of the rule-action switch, which catches every action
> pre-match has no arm for (`hijack-dns`, `direct`, and whatever upstream adds
> next). Each of those returned `PreMatchContinue`, i.e. `tun.ActionAccept`,
> i.e. the forged reply. A guard on the single return value cannot be outgrown
> by a new exit. `PreMatchBypass` joined the drop for the same reason: sing-tun
> implements `ActionBypass` on the nfqueue plane only — the name appears nowhere
> in `flow_dispatch.go` or `stack_gvisor_icmp.go` — so on the TUN path it lands
> in the same `default:` arm as Accept and forges too. There is no honest bypass
> for a packet that is already inside the engine's TUN.
**The device MTU is NOT a tunnel budget, and pretending it was manufactured
forged replies.** `l3-in` is created with MTU **65535**, not the tunnel's 1420,
and the maximum is the whole argument. This MTU governs exactly one thing:
whether the KERNEL splits a packet on its way INTO the device. What the engine
then puts into the tunnel is sized separately and correctly, against the
OUTBOUND's MTU — `ForwardDispatcher.forwardToPort` (`flow_dispatch.go:445-481`)
measures every forwarded packet against `Port.PortMTU()` and either fragments to
it (no DF, `fragmentIPv4Packet`) or answers a well-formed `fragmentation needed`
quoting it (DF, `buildFragmentationNeeded`, source = the far host, so PMTU
discovery works end to end). That machinery was always there; it was simply
never handed a whole packet.
At 1420 it wasn't. Anything above 1392 bytes of payload was fragmented by the
kernel at this device, and a fragment is the one thing sing-tun will not judge:
`Dispatch` (`flow_dispatch.go:176-177`) returns on `parsed.fragment` BEFORE
calling `JudgeFlow` at all. The fragments fell through to the gVisor stack —
promiscuous and spoofing (`stack_gvisor.go:219-223`) — which reassembled them
and handed the echo to `ICMPForwarder.HandlePacket` (`stack_gvisor_icmp.go:105+`),
whose `installFlow` (`:233-244`) writes to the port UNMODIFIED and therefore
demands a port address that is valid **and UNSPECIFIED**. `direct` qualifies
(`IPv4Unspecified()`); a WireGuard/AWG endpoint reports its concrete interface
address (`transport/wireguard/port.go:13`) and does not. So it declined, and
`HandlePacket` fell past the switch and FORGED the reply: `SetType(EchoReply)` +
address swap. Net effect on the operator's bench: `ping -s 1392` honest,
`ping -s 1393` a lie told by the router — and the lie was, of course, only for
the outbounds this feature exists for. (Upstream applies the very same
unspecified test and answers it honestly in the cloudflared ICMP handler,
`protocol/cloudflare/inbound.go:163-167`: it drops. Only the TUN path forges.)
65535 rather than "big enough": no IP datagram can exceed it, so the kernel
CANNOT fragment at this device, for any packet, ever. Any smaller value leaves
a band of sizes open and re-opens the class. It is also sing-box's own default
TUN MTU on Linux. Pinned by `TestL3TunnelMTULeavesNothingForTheKernelToFragment`
and `TestL3TunnelMTUIsNotATunnelBudget` (`generate/l3mtu_test.go`), and — the
assertion that matters — by the integration test reading the MTU back off the
real kernel device, since a kernel that clamped it would restore the forgery
without changing a generated byte.
Memory was MEASURED, not reasoned about: three paired runs of
`TestIntegrationL3TunInboundStarts` under `-test.memprofilerate=1` (exact
accounting, not sampled) allocate 5.41 / 5.48 / 5.47 MB at 65535 against
5.76 / 5.46 / 5.70 MB at 1420, and a `-diff_base` profile attributes every
difference to netlink interface enumeration. Nothing in the read path scales
with the MTU: gVisor reads through `fdbased.BufConfig`, which sing-tun's `init`
pins to a single 65535-byte view regardless of MTU, and `fdbased` keeps `mtu`
only to return it from `MTU()`. Two adjacent facts, recorded because both are
easy to derive wrongly: (a) `protocol/tun` computes
`enableGSO = stack == gvisor && mtu < 49152`, so this MTU turns GSO off there —
and then `StartStateStart` turns it back ON unconditionally because an
`adapter.FlowOutbound` exists in the config, so the ~1.98 MB of TCP/UDP GRO
scaffolding is present at BOTH MTUs and is priced by the flow-capable outbound,
not by this number; (b) the `mtu_fix` on the `shater_l3` fw4 zone is now inert —
only ICMP is ever marked into the device — and its uci-defaults comment still
says "the tunnel MTU is 1420".
**What is still NOT covered, said plainly.**
1. **A big ping does not start WORKING — it starts FAILING HONESTLY.** Upstream's
ICMP NAT is unfragmented-only in BOTH directions: `classifyReturn`
(`flow_dispatch.go:703-710`) returns `returnPass` on `parsed.fragment` exactly
as the forward path does. So a non-DF `ping -s 2000` now genuinely leaves the
router (fragmented to the tunnel MTU by `forwardToPort`), the far host really
answers, and the reply — fragmented by the peer to fit the tunnel — is not
NAT'd back to the LAN client. The operator sees a timeout. That is the
feature's promise ("travels or fails honestly"), not a capability claim.
Carrying oversized ICMP end to end would need reassembly upstream does not
have; it is not planned.
2. **A client that puts fragments on the wire ITSELF.** The device MTU cannot
un-fragment what already arrived fragmented, so such packets still reach the
gVisor stack, still get reassembled there, and still receive a forged reply
when the outbound is WireGuard/AWG. This is the residue the planned
`ip frag-off & 0x3fff != 0` prerouting carve-out (`netplane/nft.go`) is for.
**Whoever writes that rule must first check whether it can ever match:** fw4's
ruleset uses conntrack, conntrack pulls in `nf_defrag_ipv4`/`nf_defrag_ipv6`,
and defrag REASSEMBLES in PREROUTING before our marking rules run. Where
defrag is active the case does not arise (the MTU covers it) and the rule is
dead; where it is not, the rule is the only cover. Verify on the bench with
`nft list ruleset | grep -c ct` and a fragment counter, do not assume.
3. **The DF path changed hands and is untested on hardware.** It used to be the
kernel that answered `fragmentation needed` (from the router's LAN address,
MTU 1420); it is now the engine (from the far host's address, quoting
`Port.PortMTU()`). Both are correct PMTUD; only the first has ever run on a
real router.
**fw4 has to be told about the device, and `list device` is the only spelling
that works.** nftables runs EVERY table on every packet and a drop in any one of
them wins — an accept in `inet shater` cannot override fw4, and fw4 WILL reject
this forward: netifd never learns about a device the daemon creates at runtime,
so `shater-l3` belongs to no zone and falls into fw4's zone-less defaults. Hence
a real fw4 zone `shater_l3` + a lan→shater_l3 forwarding, seeded idempotently
(NAMED sections) and unconditionally in uci-defaults
(`openwrt/shater-core/files/etc/uci-defaults/30_shater-core`, `seed_l3_zone`),
with `mtu_fix` set. That `mtu_fix` is now inert and should be read as such: it
clamps forwarded TCP MSS to the route MTU, the device MTU is 65535, and nothing
but ICMP is ever marked into this device — the uci-defaults comment still says
"the tunnel MTU is 1420" and is stale. The device is attached
via `list device`, deliberately NOT `list network`: fw4 resolves a zone's
networks through netifd, which yields an EMPTY device set for a runtime-created
TUN (a proto-none stub would have to be brought UP to contribute an l3_device,
and nothing ever brings it up), while `list device` compiles to a plain
iifname/oifname string match — valid before the TUN exists, matching from the
moment shaterd creates it. `kmod-tun` joined DEPENDS so a slimmed image cannot
lose `/dev/net/tun` (`openwrt/shater-core/Makefile`). Our own forward chain
accepts both TUN legs ahead of the fail-closed drops — accepts that speak for
OUR table only (`netplane/nft.go`, forward chain step 4).
**What the policy still owns, and the one combination that now warns.** With the
ingress on, the mark is stamped in prerouting and the ROUTING decision carries
echo into the TUN before the forward chain — where the policy's verdicts live —
is ever consulted; that holds under every `untunnelable` value. The policy
therefore governs exactly two things: the never-markable protocols above, and
the fallback when the L3 rule/route did not come up (engine down, partial apply)
— `block` turns that failure into an honest loss, `direct` into a silent leak
with the real address. That is why `l3_tunnel` + `untunnelable=direct` draws a
validation warning naming the safe choice (`model/validate.go`), why every
rule/route failure surfaces as a named panel warning rather than an abort
(`addL3Routing`), and why the D17 HOLDING plane never marks: the TUN is created
BY the engine, and the holding plane exists precisely because the engine is not
running — marking would dead-end ping in a device that does not exist
(`netplane/nft.go`, hold comment).
- **Rejected: `auto_route` / letting the engine own the routing.** It rewrites
the main table and intercepts the router's own WAN/DNS/underlay traffic; the
blast radius of a toggle meant for LAN ping would be the whole router.
- **Rejected: marking all `l4proto != { tcp, udp }` into the TUN.** ESP/AH/GRE/
IGMP/SCTP cannot be parsed into flows upstream; they would vanish inside the
device. A drop with a name (the policy's) beats a silent black hole.
- **Rejected: keeping upstream's accept-and-forge for unroutable ICMP.** A ping
that "works" without leaving the router is the inverted lie this project keeps
deleting (D17's fiction purge, D23's dead WireGuard, D24's obedient-client
leak).
**Proven, and not proven, said plainly.** The cold start is PROVEN, not assumed:
`TestIntegrationL3TunInboundStarts`
(`shater/generate/l3_integration_linux_test.go`, run as root with NET_ADMIN and
`/dev/net/tun`, PASS) drives an `l3_tunnel=1` config through the SLIM registry
(`registry.Context`, not upstream's `include.Context`) under the shipped router
tag set: `box.New` + `Start` accept it, the kernel really ends up with the
`shater-l3` device at the contract MTU 65535 — the assertion the value exists
for, since a kernel that clamped it would silently restore the forged-reply
band — and Close removes it; precisely
the "built with X, verified with Y" gap class D23 exists for (a lost
`tun.RegisterInbound` or a trimmed `with_gvisor` changes no generated byte and
would otherwise surface only on the operator's router). Each layer contract is
pinned besides (`generate` `TestL3Tunnel*`, `netplane` `TestL3Ingress*`,
`route/prematch_icmp_lx_test.go`). Exactly two things remain UNVERIFIED:
(a) the end-to-end path on live hardware — LAN client → prerouting mark →
ip rule → TUN → WireGuard peer → reply back to the client — has not been
exercised on a real router; (b) the steady-state memory cost of the second
gVisor netstack (the `l3-in` TUN beside the WireGuard endpoint's own) is
unmeasured on the target hardware. An indicative figure exists and is only
that: on x86_64 in a container, idle and carrying no flows, peak RSS of a
process that brought the same engine up went from ~26.0-26.8 MB without
`l3_tunnel` to ~28.3-28.7 MB with it over three paired runs — about +2.2 MB.
That was measured on a throwaway harness, not on aarch64, not under load, and
with an empty ICMP NAT table, so it bounds nothing on the router. Neither
item is folded into any claim above.
Consequence: a ping from the LAN either genuinely travels through the tunnel
(WireGuard/AWG, direct) or fails honestly, at every size the router itself can
put into the device — and a router that never opts in renders the pre-feature
plane byte-for-byte (`TestL3IngressOptIn` pins the off-state render). Read
"fails honestly" strictly: above the tunnel MTU a non-DF ping now leaves the
router for real and then times out, because upstream's ICMP NAT does not carry
fragments back either. The one qualifier left is item 2 above — a client that
puts fragments on the wire ITSELF, on a router where conntrack defrag is not
reassembling them first. This paragraph has been overclaimed twice already;
extend it only against a bench result, never against a reading.
## D26 — What the engine cannot carry, the kernel carries: `untunnelable_egress`
Decided 2026-07-26. D25 ended with ESP/AH/GRE/IGMP/SCTP still owned by the D17
policy — that is, with a choice between dropping them and leaking them out the
WAN, never a data plane. This decision gives them one, and deliberately NOT
ours: **`globals.untunnelable_egress` (default empty;
`model.Globals.UntunnelableEgress`) names an existing egress of type
interface/tunnel, and LAN traffic that is neither TCP nor UDP is stamped in
prerouting with that egress's own mark, so the KERNEL routes it out that
egress's device with the kernel's own NAT.** No proxy, no engine, no userspace
stack ever touches the packet — which is exactly why every protocol works.
**Where the engine's boundary actually is — recorded so nobody digs for it
twice.** It is NOT the gVisor stack, and it is not WireGuard: on the forward
path the WG/AWG endpoint never consults gVisor at all. `Endpoint.WritePackets`
(`transport/wireguard/port.go:21-58`) takes the raw IP packet bytes, reads
exactly the IP version and the destination address, and hands
`device.InputPacketRef`s to `wgDevice.InputPackets` — the protocol byte is
never read; on the way back (`port.go:127-157`) `returnDeviceWrapper.Write`
offers every decrypted packet to `returnPath.ReturnPackets` first and only the
unconsumed remainder falls through to the gVisor device. gVisor serves
`DialContext`/`ListenPacket` — traffic the ENGINE originates — while forwarded
traffic bypasses the stack in both directions, indifferent to protocol. The
real ceiling sits one step earlier, in sing-tun's `ForwardDispatcher`:
`parseTransport` (`flow_parse.go:106-153`) sets `hasFlow` for exactly TCP, UDP,
ICMPv4 Echo/EchoReply and ICMPv6 EchoRequest/EchoReply — a packet of any other
protocol is never dispatched as a flow — and `createFlow`
(`flow_dispatch.go:325`) builds its NAT through
`allocateSelector(packet.protocol, …, packet.source.Port())` (line 355), which
needs a port-like selector that ESP/AH/GRE simply do not have (SCTP has ports,
but the parser above never grants it a flow either). Tailscale documents the
same frontier for its own userspace mode — "Any IP protocol other than TCP or
UDP (such as SCTP) is not supported in userspace mode… All IP protocols are
supported" in kernel mode
(https://tailscale.com/docs/reference/kernel-vs-userspace-routers) — useful as
external corroboration of where userspace data planes generally end, though OUR
boundary is the dispatcher, not the stack. The kernel egress was therefore
chosen not because userspace "cannot" in principle, but because the kernel
delivers all protocols with zero new code on the hot path.
**The mechanism already existed; the feature is one binding and one marking
step.** `addEgressRouting` (`netplane/apply.go`) has always installed, for
every interface/tunnel egress, an `ip rule fwmark <EgressMark> lookup
<EgressTable>` plus a `default dev <device>` route in that table — per-rule
egress selection rides on it. The only missing piece was that nothing ever
marked non-TCP/UDP traffic: `untunnelable=direct` merely ACCEPTED it in the
forward chain, so it left over the main table, i.e. the WAN.
`UntunnelableEgressBinding` (`netplane/nft.go`) resolves the option to the
egress's index, its OWN mark and its OWN device — deliberately no third
mark/table pair to keep coherent — and the prerouting chain stamps that mark on
the untunnelable protocols. A name that does not resolve to an interface/tunnel
egress with a device renders nothing and is reported: the D17 policy stays in
sole charge, which is the fail-closed reading of a typo.
**Why `l4proto != { tcp, udp }` is safe here when D25 banned it.** D25 rejected
the broad filter because the receiving side was the `ForwardDispatcher`, which
classifies nothing beyond TCP/UDP/ICMP echo — a marked ESP packet would enter
the TUN and vanish, a black hole wearing a tunnel's name. Here the receiving
side is the kernel, which forwards ANY IP protocol and NATs what it has
machinery for: SCTP carries ports and NATs like TCP/UDP; GRE is NATed only
through the PPTP helper keyed on the call-id — the kernel's own comment calls
GRE "generally not very suited for NAT, as it has no protocol-specific part as
port numbers" (`net/netfilter/nf_conntrack_proto_gre.c`); ESP/AH pass as plain
routed IP. Nothing on this path can silently swallow a protocol it does not
understand, which was the entire objection.
**Order against D25: the L3 ingress claims ICMP first.** With `l3_tunnel` on,
LAN ICMP is marked into the engine's TUN before the egress carrier is consulted
— the engine path routes ping by the operator's rules, which a kernel egress
cannot do — and only the remaining protocols go to the egress. With `l3_tunnel`
off, ICMP goes to the egress with everything else. In both shapes marked
traffic is settled by ROUTING before the forward chain speaks, so the D17
policy now governs exactly the failure case — the rule or route that did not
come up — the same division D25 already established for the L3 mark.
**What the feature refuses to promise — and the operator text refuses with it
(`shater/apply/warnings.go`, the egress-carrier note).** (a) It is not a tunnel
per se: the option accepts any interface/tunnel egress, and on the target
routers a WireGuard device is the exception (`kmod-wireguard` is usually
absent) while a second WAN is routine. Through a WireGuard egress this
genuinely is a tunnel; through a second WAN it is simply another uplink, and
the destination sees that uplink's real address. No text, comment or doc line
may call it a tunnel unconditionally. (b) It does not revive IPTV: IGMP is
LAN-side multicast group management, WireGuard is L3 point-to-point and carries
no multicast, and multicast never crossed this router under any setting —
routing IGMP out an egress restores nothing, and no wording may hint otherwise.
(c) IPsec through NAT-T never needed it: RFC 3948 encapsulates ESP in UDP/4500,
so a modern IPsec client behind NAT is ordinary UDP that already follows the
routing rules; the raw-ESP case this feature carries is the no-NAT-T remainder.
- **Rejected: teaching the engine these protocols.** Extending `parseTransport`
and the selector NAT upstream would be new hot-path code in an
actively-maintained adversarial area, for protocols the kernel already
forwards for free — and for ESP/AH/GRE there is no port-like selector to NAT
by in the first place.
- **Rejected 2026-07-26: carrying them through the userspace AWG endpoint
site-to-site, with no NAT at all.** This is the alternative the "no port-like
selector" line above does NOT dispose of, and it is written down because the
obvious reading of that line — "impossible" — is wrong and would be
re-derived. The endpoint is already protocol-blind in both directions
(`transport/wireguard/port.go:21-58`, `:127-157`), so an ESP packet could be
forwarded UNTOUCHED, keeping the LAN client's own source address, and the
reply would come back addressed to that client and need only be written to the
TUN. No selector, no NAT, every protocol. It needs two things we declined to
take on: lx-owned code in the forward hot path, bypassing `ForwardDispatcher`
on both legs — precisely the surface CONSTITUTION §2 exists to keep small on
an actively-maintained upstream — and a SERVER-side prerequisite (our LAN
prefix in the peer's `AllowedIPs`, plus a route back), which turns a router
option into a deployment contract. The kernel egress above buys the same
protocols with zero hot-path code, so this stays a design on file, not a gap.
- **Rejected: a dedicated mark/table pair for the carrier.** `addEgressRouting`
already binds `EgressMark`/`EgressTable` to the device; a third pair would be
a second copy of the same route that could drift from the first.
**Not verified, said plainly.** The end-to-end path — LAN client → prerouting
mark → ip rule → egress device → far end and back — has not been exercised with
real ESP or GRE on live hardware. Nothing above claims it has.
Consequence: raw IPsec, PPTP/GRE, SCTP — and ICMP when the L3 ingress is off —
leave through an egress the operator explicitly named, under kernel routing and
kernel NAT, instead of being dropped or silently leaking out the WAN; and with
the option empty (the default) the plane renders byte-for-byte as before, with
the D17 policy in sole charge.
+42 -6
View File
@@ -12,9 +12,35 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
## Transparent proxying & routing
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** **L3 ingress for ICMP** (`globals.l3_tunnel`, opt-in, default off):
LAN ping travels THROUGH the tunnel instead of being dropped or answered by a
forged local reply. The engine opens a dedicated TUN (`shater-l3`, gVisor
stack, `auto_route` off); nft marks LAN icmp/icmpv6 only and a scoped
`ip rule` routes it in — the TPROXY plane and the main routing table stay
untouched (D25). Carried only by L3-capable egresses (WireGuard/AmneziaWG,
direct); ICMP routed to vless/vmess/… is honestly dropped, never faked.
Ceiling is upstream sing-tun's: ICMP echo only — Windows tracert works, IPv6
traceroute shows just the destination; ESP/AH/GRE/IGMP stay with the
`untunnelable` policy (D17) unless `untunnelable_egress` carries them (D26).
- **[MVP]** **Kernel egress for untunnelable protocols**
(`globals.untunnelable_egress`, opt-in, default empty): names an existing
interface/tunnel egress, and IPsec (ESP/AH), PPTP/GRE, SCTP — everything that
is neither TCP nor UDP, plus ICMP when the L3 ingress is off — is routed out
that egress's device by the KERNEL with kernel NAT, reusing the egress's own
fwmark/table from `addEgressRouting`; the proxy never sees a byte, which is
why every protocol works (D26). What that buys depends on the device: a
WireGuard interface really is a tunnel, a second WAN is just another uplink
whose real address the destination sees. It does not revive multicast IPTV,
and UDP-based VPNs (WireGuard, OpenVPN-UDP, IPsec NAT-T) never needed it —
they follow the routing rules as before. The `untunnelable` policy (D17)
keeps only the failure case: a route that did not come up.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination (domain/suffix/keyword/geosite), reusable domain/IP lists, port,
proto → target (outbound/selector/chain/direct/block) + egress.
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
(inline domains/CIDRs, a local or remote file, or a geosite/geoip category) that is
compiled once into a `.srs` and shared by every rule that references it. Domain
entries take `full:` (exact), `suffix:` / a leading dot (host + subdomains),
`keyword:` (substring) and `regexp:`; a bare entry means host + subdomains.
- **[MVP]** Node groups with balancer/observatory (least-ping/failover/round-robin).
- **[T1]** Multi-hop chains (L1→Ln); per-rule egress selection; egress via any
interface/tunnel (e.g. an AmneziaWG tunnel).
@@ -36,6 +62,16 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
type=fakeip + pool — there is no global "FakeIP mode"); no DNS leaks. Routing
is decided by in-engine rule-sets — the v0.1 dnsmasq→nftset population
mechanism does not exist in v0.2 (see generate/dns.go).
The hijack covers the queries a client sends **to the router itself** — the
address DHCP hands out — because `globals.dns_intercept` is **ON by default**
(D24). With it off, those queries go to dnsmasq and out to the ISP in the clear,
so the well-behaved client leaks while the one that hard-codes 8.8.8.8 does not.
`.lan` and the private PTR zones are preserved through dnsmasq either way. Two
things the promise does NOT cover, both by design: while the engine is DOWN the
holding plane hooks `forward` only, so dnsmasq still answers router-addressed
:53 unfiltered (client traffic and DNS to external resolvers stay blocked); and
with no `config resolver` at all there is no DNS plane to filter with — queries
fall through to the system resolver and generate says so.
- **[MVP]** Client DoT/DoH blocking (stop devices bypassing the filter).
- **[MVP]** **Blocklists** with **flexible sources**: `inline` (type your own) /
`file` / `url` (auto-update) / `geosite` category (only when geodata present).
@@ -89,8 +125,8 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
SIM uplink → different egress); backup/restore; i18n (EN + RU).
## Ops & distribution
- **[MVP]** Single signed binary; signed opkg feed on Gitea (reuse key
`5ac4b177689cb8e0`); one-line install; `opkg upgrade`.
- **[MVP]** Single signed binary; signed apk feed on Gitea (EC key
`dist/shater-apk.pem`); one-line install; named-package `apk upgrade`.
- **[T1]** Upstream-rebase cadence (track sing-box-lx tags) with a smoke suite.
- **[T2]** apk (OpenWrt 25.x) packaging; multi-router fleet management; REST/gRPC
external API; Telegram bot.
- **[T2]** Multi-router fleet management; REST/gRPC external API; Telegram bot.
(apk packaging landed and is now the only lane — D22.)
+122 -99
View File
@@ -1,7 +1,7 @@
# Shater v0.2 — Build & Install
How to build the ship artifact (the SPA-embedded `shaterd` binary) and install
the OpenWrt feed onto a router.
the signed apk repo onto a router.
## 1. Build the `shaterd` binary
@@ -28,7 +28,7 @@ Arg / env:
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
the **same** computation the package version comes from (§2.1), so the string
the panel shows always matches what `apk info shaterd` / `opkg status` report.
the panel shows always matches what `apk list -I shaterd` reports.
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
@@ -43,22 +43,44 @@ UPX="…/scratchpad/upx-4.2.4-win64/upx.exe" scripts/build-shaterd.sh v0.2.0 --f
The `dist/*` and `openwrt/shaterd/files/shaterd-*.upx` outputs are gitignored —
they are release artifacts, not source.
Tag set (D9 — keep in sync with `docs-shater/DECISIONS.md`):
Tag set (D9/D23) — defined in **one** place, `scripts/router-tags.sh`, which
documents every tag and is sourced by the build:
```
with_quic,with_wireguard,with_utls,
with_gvisor,with_quic,with_wireguard,with_utls,
badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command
```
We drop `with_purego,with_naive_outbound`: they pull cronet-go, which forces a
glibc `PT_INTERP` even under `CGO_ENABLED=0`, making the binary unusable on musl.
We drop `with_gvisor`: the shater data plane is tproxy/redirect and generate
never emits a tun inbound, so the userspace gvisor netstack is unreachable code.
We drop `with_clash_api`: the admin panel is shater's own web server and the
generator never emits a `clash_api` service, so the Clash server is dead code.
We drop `with_dhcp`: shater resolver types are `udp/tcp/doh/dot/local/fakeip`;
a `dhcp://` DNS transport is never generated or registered.
`with_gvisor` was dropped in 2026-07 as "unreachable — we emit no tun inbound"
and **put back on 2026-07-25**: gVisor is also the netstack of the WireGuard
endpoint, so without it every `wg://`/`awg://` node died at apply time with
*"gVisor is not included in this build"* while the panel still offered the
feature. It costs ~2.8 MB raw / ~0.65 MB UPX per arch. Full story: `DECISIONS.md`
D23.
### Changing the tag set
Run the guard — it is what stands between a size trim and a silently dead
feature, and CI runs it before the artifact is built:
```sh
scripts/check-router-tags.sh # from Windows/macOS it re-execs itself in golang:1.26
```
It (1) fails if a feature declared in `FEATURES.md` lost a build tag it needs to
run (`shater/buildtags`, no tags/OS/network required) and (2) constructs one node
of every declared protocol through `box.New` **compiled with the shipped tag
set** — nothing may be skipped in that run. Adding a protocol to
`shater/parse`+`shater/generate` means adding a row to `buildtags.Features` and a
probe case in `shater/generate/shipped_tags_linux_test.go`.
## 2. Packages
Four OpenWrt packages live under `openwrt/`:
@@ -91,9 +113,9 @@ See `openwrt-package-build-ci` for SDK/feed mechanics.
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
Both package managers offer an upgrade only when the feed's version string
differs from the installed one, so `apk update` saw nothing new and the routers
could not be updated through the normal path at all.
apk offers an upgrade only when the feed's version string differs from the
installed one, so `apk update` saw nothing new and the routers could not be
updated through the normal path at all.
`ci/version.sh` now derives them from `git describe`, once per CI job:
@@ -103,9 +125,8 @@ could not be updated through the normal path at all.
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
Ordering is what makes this safe, and both managers agree on it (checked with
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
Ordering is what makes this safe (checked with `apk version -t` on apk-tools
3.0.3): the dotted part decides first, `-rN` only breaks ties — so
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
outranks every rolling build before it, rolling builds between two releases grow
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
@@ -113,8 +134,9 @@ upgrade.
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
environment; the Makefiles read it with a literal fallback for manual/offline
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
so a lost variable fails the build instead of shipping a stale version.
builds. `ci/sdk-build-apk.sh` then **asserts** the produced `.apk` really carries
it, so a lost variable fails the build instead of shipping a stale version. The
release job asserts the same version again on the published rolling repo (§5.1).
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
@@ -124,21 +146,27 @@ when our packaging of it changes.
## 3. Install on a router
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
**The normal path is the signed apk repo — §5.** This section is the manual
fallback (a router with no route to the Gitea host, or a hand-carried build).
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`).
apk filenames carry no architecture, so make sure you copied the `.apk` built for
*this* router's arch (`cat /etc/apk/arch`):
```sh
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_<ver>_all.ipk
opkg install luci-app-shater_<ver>_all.ipk
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
# --allow-untrusted: our member .apk are unsigned by design — trust lives in the
# signed packages.adb index (§5), which a loose file install does not consult.
apk add --allow-untrusted ./shaterd-<ver>.apk
apk add --allow-untrusted ./shater-core-<ver>.apk
apk add --allow-untrusted ./luci-app-shater-<ver>.apk
apk add --allow-untrusted ./byedpi-0.17.3-r1.apk # optional: ByeDPI egress
```
Installing from a signed feed instead:
From the repo instead (§5 sets it up once), deps pull the rest in:
```sh
# add the feed (customfeeds.conf / apk repositories), then:
opkg update && opkg install shater-core luci-app-shater # shaterd pulled in as a dep
apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
## 4. Enable
@@ -158,92 +186,87 @@ daemon (`shaterd run`), which owns the engine, the `inet shater` data plane, pol
routing, in-process DNS, and the admin panel (default `:8088`). The LuCI app's
"Open panel" button mints a single-use token and hands the browser off to the panel.
## 5. Add the signed feed (recommended — then `opkg upgrade` just works)
### What enabling does to DNS
CI (`.gitea/workflows/release.yml`) publishes every build as a **rolling `latest`
Gitea release** that is itself a signed opkg `src/gz` feed: the release holds the
`.ipk` for all arches, a `Packages`/`Packages.gz` index, a usign `Packages.sig`,
and the public key `shater-feed.pub`. opkg filters by `Architecture`, so the **same
two lines work on every device** (x86 testbed picks `x86_64 + all`; the BPI routers
pick `aarch64_cortex-a53 + all`).
From the first apply, **every** LAN plaintext `:53` goes into the engine — including
the queries a client sends to the router's own address, which is what DHCP hands out.
That is `globals.dns_intercept`, and it is **on by default** (D24); without it those
queries reach dnsmasq and the ISP unfiltered, i.e. the client with default settings
leaks while the one that hard-coded `8.8.8.8` does not. What follows from it:
> **Format:** OpenWrt 24.10 (our SDK) uses **opkg** (`.ipk`, `Packages.gz`, usign),
> so the feed is `src/gz` and the trust anchor is the usign key
> `dist/shater-feed.pub` (fingerprint **`5ac4b177689cb8e0`**). apk only replaces
> opkg at OpenWrt **25.12** — see §6.
- `.lan` and private reverse (PTR) lookups still go to dnsmasq — the engine gets a
rule for those suffixes. If you renamed dnsmasq's domain away from `lan`, add a
`config dns_rule` for the new suffix.
- Configure at least one `config resolver`. With none, the engine has no resolver
plane: intercepted queries fall through to the system resolver (dnsmasq → your
ISP, in the clear), blocklists and per-device DNS rules are inert, and the apply
says so in its warnings.
- While the engine is DOWN, DNS is **not** blacked out: the fail-closed holding
plane hooks `forward` only, so dnsmasq keeps answering router-addressed `:53`
(unfiltered, plaintext) while client traffic and DNS to external resolvers stay
blocked. "The tunnel is down" is not "DNS is private".
One-time setup on the router:
To opt out, on the router:
```sh
# 1) trust the feed key — the FILENAME must equal the usign key fingerprint.
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
# 2) add the feed (one URL serves every arch).
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
# 3) refresh + install (shaterd is pulled in as a dependency).
opkg update
opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # optional: ByeDPI desync egress
uci set shater.globals.dns_intercept=0
uci commit shater
shaterd apply
```
With the key installed, opkg's default `check_signature 1` verifies the feed on
every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
(`vX.Y.Z`) publishes the identical layout at
`.../releases/download/vX.Y.Z` if you prefer to pin a version instead of tracking
`latest`.
Your `0` is kept: `/etc/config/shater` is a conffile (upgrades never replace it) and
the daemon always writes the option back explicitly, so it is never re-enabled by a
default.
### Updating
## 5. The signed apk repo (the normal install path)
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
tries to upgrade *every* installed package from *every* configured feed, which on
OpenWrt means base/system packages on the overlay and is a well-known way to
brick a router.
OpenWrt/ImmortalWrt **25.12** packages with Alpine's **apk**: `.apk` files, a
binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
effectively mandatory signatures (unsigned needs `--allow-untrusted`). This is
the only format shater publishes — the `.ipk`/opkg lane was removed in 2026-07
(`DECISIONS.md` D22); every device we serve is on 25.12 with apk-tools 3.
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
```
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
when the feed's `Version` differs from the installed one — that is exactly what
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
the version from the git tag on every build (§2.1), so there is nothing to bump
by hand any more; check with:
```sh
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
```
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
OpenWrt/ImmortalWrt **25.12** replaces opkg with Alpine's **apk**: `.apk` files,
a binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
effectively mandatory signatures (unsigned needs `--allow-untrusted`). The
package **Makefiles are unchanged** — the SDK release decides the format.
CI builds this lane **in parallel** with the opkg feed (same manual triggers:
`v*` tag push or `workflow_dispatch`): the `build-apk` jobs in
`.gitea/workflows/release.yml` compile the same 4 packages through the official
**ImmortalWrt 25.12 SDK** (tarballs from
CI (`v*` tag push or `workflow_dispatch`) compiles the 4 packages through the
official **ImmortalWrt 25.12 SDK** (tarballs from
`downloads.immortalwrt.org/releases/25.12.1/targets/{x86/64,mediatek/filogic}/`)
and publish **one release per arch** — rolling `apk-latest-x86_64` /
`apk-latest-aarch64_cortex-a53`, or `apk-vX.Y.Z-<arch>` for a tagged version.
Per-arch (unlike the combined opkg release) because apk filenames carry no
architecture and packages are fetched relative to the `packages.adb` URL.
and publishes **one release per arch**: the rolling `apk-latest-x86_64` /
`apk-latest-aarch64_cortex-a53`, plus `apk-vX.Y.Z-<arch>` on a tag. Per-arch
because apk filenames carry no architecture and packages are fetched *relative to
the `packages.adb` URL*, so one flat multi-arch release would collide.
> **Key:** apk cannot use the usign key. The apk trust anchor is the separate EC
> public key **`dist/shater-apk.pem`** (generated once by `ci/gen-apk-key.sh`;
> private half lives ONLY in the Gitea secret **`KEY_APK`**, the apk analog of
> `KEY_BUILD`). Never regenerate either key — that invalidates every deployed
> router's trust. The usign identity `shater-feed.pub` keeps signing the
> opkg/24.10 feed, untouched.
> **Key:** the trust anchor is the EC public key **`dist/shater-apk.pem`**
> (generated once by `ci/gen-apk-key.sh`; the private half lives ONLY in the
> Gitea secret **`KEY_APK`**). Never regenerate it — that invalidates every
> deployed router's trust.
One-time setup on a 25.12 router (BananaWRT `25.12-mtk-vendor` on the BPI-R3
mini, BPI-R4 on 25.12, or the future 25.12 VM — `/etc/apk/arch` picks the right
per-arch release automatically):
### 5.1 Rolling or pinned — pick the repo URL deliberately
The repo line names an **index file**, and which one you name is the whole
update policy:
| Repo line points at | Behaviour | Cost |
|---|---|---|
| `apk-latest-<arch>/packages.adb` (**rolling**) | Every release run REPLACES this release's assets, so `apk update && apk upgrade <our packages>` always sees the newest build. Install once, never touch the file again. | You get whatever CI published last; there is no per-router pin. |
| `apk-vX.Y.Z-<arch>/packages.adb` (**pinned**) | The router stays on exactly that build. `apk update` will never offer a newer shater. | `/etc/apk/repositories.d/shater.list` must be edited **by hand on every upgrade**, on every router. |
`mini_router` is deliberately on a **pinned** URL — a considered choice, and the
hand-edit per release is its price. Use rolling unless you specifically want to
freeze a device.
> The rolling release used to go stale silently: publishing was an either/or, so
> tag runs wrote only `apk-vX.Y.Z-<arch>` and `apk-latest-<arch>` was last
> refreshed on 2026-07-24 at `0.2.0` while v0.2.9/v0.2.10 shipped. A router on
> the rolling URL kept getting a successful `apk update` with nothing new. Fixed
> 2026-07-25: `release-apk` writes the rolling pointer on **every** run and then
> reads the release back over the Gitea API, asserting it holds our three
> tag-versioned packages at exactly the version just built and **no** leftover
> asset at another version (two versions of one package in one index would let
> apk choose instead of us).
### 5.2 One-time setup on the router
BananaWRT `25.12-mtk-vendor` on the BPI-R3 mini, OpenWrt 25.12 on the BPI-R4, or
the testbed VM — `/etc/apk/arch` picks the right per-arch release automatically:
```sh
# 1) trust the apk feed key (any *.pem filename under /etc/apk/keys works).
@@ -251,6 +274,7 @@ wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
# 2) add the repo — the line points at the packages.adb INDEX FILE itself.
# (rolling; for a pinned router put apk-vX.Y.Z-$(cat /etc/apk/arch) here — §5.1)
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
@@ -260,7 +284,7 @@ apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # optional: ByeDPI desync egress
```
### Updating
### 5.3 Updating
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
installed package against *every* configured repository at once; on a router
@@ -286,9 +310,8 @@ Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
rolling by pointing the repo line at
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
as `0.2.0-r3` and `apk update` offered nothing). Rolling vs pinned repo URL —
§5.1.
### BananaWRT `25.12-mtk-vendor` compatibility
+41 -4
View File
@@ -195,7 +195,7 @@ type Chain struct { Name string; Hops []string } // "group:<n>" | "node:<n>", L1
type Egress struct { Name,Type,Interface,Target string } // interface|proxy|direct|block
type Rule struct {
Name string; Enabled bool; Order int
Src []string; DstDomain,DstRuleset,DstIP []string; DstPort,Proto string
Src []string; DstRuleset []string; DstPort,Proto string // dst = ruleset only (v0.2 schema v2)
Target string // chain:|group:|node:|direct|block
Egress,Kill string
SchedEnabled bool; SchedDays []string; SchedStart,SchedEnd string; SchedUTCOffset int
@@ -251,14 +251,51 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
> v0.2: "restart engine only on change" → config-hash gate + Close+New box (no reload).
### uci.go — `/etc/config/shater` schema
- `config globals`: enabled, loglevel, kill_switch, dns_mode, ipv6, fwmark_base, table_base, confirm_timeout, resolver_default, resolver_fallback, probe_url, probe_interval, schema_version, active_profile.
- `config globals` — the full option set, with the value used when the option is
ABSENT (the `model.DefaultGlobals` seed). Booleans are always written back as
`'1'`/`'0'` by `render.go`, so an explicit value never decays into the seed:
| option | default | meaning |
|---|---|---|
| `enabled` | `0` as shipped | master switch; `0` ⇒ `Reconcile` tears the stack down instead of applying |
| `loglevel` (alias `log_level`) | `warning` | engine + daemon level; `none/off/silent/disabled` ⇒ log disabled, unknown ⇒ `warn` + a validation warning |
| `log_syslog` / `log_file` / `log_persist` | `1` / `1` / `0` | operational log (`shater/logsink`): syslog, rotated file, and whether that file lives on flash instead of tmpfs |
| `log_max_kb` | `2048` | size cap of the log file, clamped to 128…8192; `0` = "use the default", not "off" |
| `kill_switch` | `closed` | `closed` = fail-closed (block on engine loss, incl. a holding plane when the engine never started); `open` = plain routing |
| `ipv6` | `1` | `0` drops LAN IPv6 in the forward chain instead of leaving it unproxied |
| `fwmark_base` / `table_base` | `0x2000` | reserved fwmark / routing-table bases (must not collide with fw4 or other apps) |
| `confirm_timeout` | `0` | seconds before an unconfirmed apply auto-rolls back; `0` = commit-confirm off |
| `resolver_default` / `resolver_fallback` / `endpoint_resolver` | unset | `config resolver` names: the DNS catch-all, its failover chain, and the bootstrap-direct server that resolves proxy endpoint DOMAINS |
| `probe_url` / `probe_interval` | engine defaults | the ONE instrument all health probing uses (D20 — there are no per-group overrides) |
| `panel_port` | `0` ⇒ `8088` | admin-panel HTTP port |
| `dns_filter` | `0` | master enable of the blocklist/allowlist filter (D15); needs at least one `config resolver` |
| `dns_intercept` | **`1`** | force ALL LAN plaintext `:53` into the engine, INCLUDING queries addressed to the router itself. See D24 for why this is the default, what preserves `.lan`, and what happens while the engine is down |
| `block_doh` | `0` | NXDOMAIN the known public DoH hostnames + the Firefox canary and reject `:443` to their IPs, so clients fall back to `:53` (which the engine catches) |
| `group_health` | `1` | OUR background group probing (the observatory). Does not touch sing-box's own urltest inside a group |
| `untunnelable` | `block` | policy for what TPROXY cannot carry (ICMP/IGMP/ESP/AH/GRE/SCTP): `block` \| `icmp` (echo out, rest dropped) \| `direct` (all out, bypassing the tunnel) |
| `geo_provider` | unset = auto | `sagernet` \| `loyalsoldier` \| `metacubex` \| `custom`; auto = country codes from SagerNet, everything else from Loyalsoldier |
| `geosite_url` / `geoip_url` | unset | `{category}` templates, honoured only when `geo_provider=custom` |
| `geosite_index_url` / `geoip_index_url` | unset | git-trees URLs used to SUGGEST categories in the panel; empty = no suggestions |
| `stats_backend` | `memory` | `off` (no aggregation at all) \| `memory` (RAM, lost on restart) \| `sqlite` (aggregates in RAM + query/connection log on disk) |
| `stats_ring_size` / `stats_timeline_minutes` / `stats_max_domains` | `200` / `60` / `5000` | live-log length, sparkline minutes, domain-map cap. **`0` = UNLIMITED** (grows with traffic), which is why these three are always emitted |
| `stats_disk_limit_mb` | `64` | on-disk cap of `stats.db`; only meaningful for `stats_backend=sqlite`; `0` = unlimited |
| `stats_retention_disabled` | `0` | master switch that turns OFF all trimming/pruning — every aggregate then grows unbounded |
| `schema_version` | `0` = pre-versioned | UCI schema revision; `shaterd migrate` writes `2` |
| `active_profile` | unset | display bookkeeping: the last profile switched to |
Deleted options still parse (unknown keys are ignored) and drain out on the next
render: `dns_mode` (D17 — fake-IP is a resolver TYPE), `sweep_interval` (D19).
- `config inbound`: name, enabled, type, network, tproxy_port(12345), listen, port, auth, user, pass, target_addr, target_port, target_network, tcp, udp, sniff.
- `config subscription`: name, enabled, url, update_interval, fetch_via(direct|proxy), ua, hwid, device_os, ver_os, device_model, list header, format, list include/exclude/filter_proto/filter_country, dedup, expire_alert_days.
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
- `config chain`: name, list hop. `config egress`: name, type, interface, target.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url), url, path, format, update_interval, list entry.
- `config rule`: name, enabled, order, list src/dst_domain/dst_ruleset/dst_ip, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end/tz.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url|geosite|geoip), url, path, format, update_interval, list category, list entry.
- `config rule`: name, enabled, order, list src, list dst_ruleset, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end, sched_utc_offset.
v0.1 carried `dst_domain`/`dst_ip` on the rule itself; **schema v2 removed both** — a
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
+1 -1
View File
@@ -6,7 +6,7 @@ OpenWrt). Лицо репозитория и быстрый старт — в к
| Документ | О чём |
|----------|-------|
| [CONTEXT.md](CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, решения в кратце, testbed/инфра |
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка обоих фидов — opkg (24.10) и apk (25.12+) |
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка apk-фида (25.12+): роллинг или фиксация версии |
| [ARCHITECTURE.md](ARCHITECTURE.md) | One-binary дизайн, auth-handoff LuCI→панель, data/DNS/apply-потоки (диаграммы) |
| [FEATURES.md](FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [ROADMAP.md](ROADMAP.md) | Фазовый план |
+1 -1
View File
@@ -104,7 +104,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
## Phase 8 — Ship it ✅ DONE
- Adapt CI to build/sign the single forked binary for both arches; publish the
signed opkg feed (reuse key `5ac4b177689cb8e0`); install/upgrade docs.
signed feed (apk since D22, EC key `dist/shater-apk.pem`); install/upgrade docs.
- Set an upstream-rebase cadence (merge new sing-box-lx tags, run the smoke suite).
## Cross-cutting (every phase)
+2 -2
View File
@@ -21,8 +21,8 @@ PKG_NAME:=byedpi
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# regression: our tags are 0.2.x, and the version comparator (apk-tools 3,
# verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
+9 -1
View File
@@ -40,6 +40,10 @@ define Package/shater-core
# shaterd : the daemon our init supervises (`shaterd run`)
# kmod-nft-tproxy : kernel TPROXY (shaterd emits the `inet shater` rules)
# kmod-nft-socket : socket match used by the tproxy divert chain
# kmod-tun : /dev/net/tun — the daemon opens the `shater-l3` TUN
# for L3 ingress (globals.l3_tunnel); usually built-in
# on stock images, but a slimmed image without it would
# make the option fail with a cryptic open() error.
# ip-full : `ip rule`/`ip route`/rt_tables for policy routing
# nftables-json : shaterd shells out to `nft`, and netplane/stats.go
# parses `nft -j list ...` — the JSON output only exists
@@ -49,7 +53,7 @@ define Package/shater-core
# ca-bundle : the daemon is CGO_ENABLED=0, so crypto/x509 has no
# host cert fallback — without /etc/ssl/certs every
# HTTPS subscription / .srs ruleset fetch fails.
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +ip-full +nftables-json +ca-bundle
DEPENDS:=+shaterd +kmod-nft-tproxy +kmod-nft-socket +kmod-tun +ip-full +nftables-json +ca-bundle
PKGARCH:=all
endef
@@ -80,6 +84,10 @@ define Package/shater-core/install
$(INSTALL_DIR) $(1)/etc/init.d
$(INSTALL_BIN) ./files/etc/init.d/shater $(1)/etc/init.d/shater
$(INSTALL_BIN) ./files/etc/init.d/shater-cron $(1)/etc/init.d/shater-cron
# START=21 one-shot that loads the persisted fail-closed plane before fw4's
# `lan -> wan ACCEPT` can be the only thing on the box (the main init is
# START=99, i.e. seconds of plaintext forwarding on every boot).
$(INSTALL_BIN) ./files/etc/init.d/shater-armor $(1)/etc/init.d/shater-armor
$(INSTALL_DIR) $(1)/etc/hotplug.d/iface
$(INSTALL_BIN) ./files/etc/hotplug.d/iface/99-shater $(1)/etc/hotplug.d/iface/99-shater
+55 -2
View File
@@ -23,6 +23,32 @@ config globals 'globals'
option kill_switch 'closed'
# There is no dns_mode option: routing is decided by in-engine rule-sets and
# fake-IP is a resolver type (`config resolver` with type=fakeip + pool).
#
# Force ALL LAN plaintext DNS (:53) into the engine, INCLUDING queries the
# client sends to the router itself (the address DHCP hands out). ON by
# default: with it off, a client using the router as its resolver is answered
# by dnsmasq and forwarded to the ISP in the clear — no blocklists, no
# per-device DNS rules, no resolver detour — while a client that hard-codes
# 8.8.8.8 IS intercepted. The obedient client leaked; the evader did not.
#
# Set to '0' to opt out (dnsmasq answers router-addressed :53 again). Your
# explicit value is never overwritten: this file is a conffile, and the daemon
# always writes the option back as '1'/'0'.
#
# .lan and the private reverse (PTR) zones keep working: with at least one
# `config resolver` present the engine gets a synthetic server pointed at
# dnsmasq on 127.0.0.1:53 plus a rule that sends those suffixes to it; with no
# resolver at all the engine falls back to the system resolver, which is
# dnsmasq too. If you changed dnsmasq's domain away from `lan`, add a
# `config dns_rule` for it (only `lan` + RFC6303 reverse zones are built in).
#
# While the engine is DOWN the LAN is NOT left without DNS: the fail-closed
# holding plane hooks `forward` only, so dnsmasq still answers router-addressed
# :53 — unfiltered and in the clear, the documented trade-off (blocking it
# would also cut the daemon's own name resolution and its chance to recover).
# Queries aimed at an EXTERNAL resolver are dropped with the rest of the LAN's
# forwarded traffic.
option dns_intercept '1'
option ipv6 '1'
# Reserved fwmark base and routing-table base (do not overlap fw4/other apps).
option fwmark_base '0x2000'
@@ -62,7 +88,29 @@ config inbound
# list node 'my-node'
#
# A routing rule. target: chain:<n>|group:<n>|node:<n>|egress:<n>|direct|block.
# Match on src / dst_domain / dst_ruleset / dst_ip / dst_port / proto.
# Match on src / dst_ruleset / dst_port / proto. A rule with NO matcher at all is
# the default route for everything that reached it.
#
# WHERE the traffic is going is named ONLY by dst_ruleset — one or more
# `config ruleset` names; the rule matches when ANY of them matches. There is no
# inline domain or address list on a rule (`dst_domain`/`dst_ip` were removed in
# schema v2): a destination list is written once as a ruleset, compiled into a
# .srs and shared by every rule that references it. `shaterd migrate` converts
# older configs automatically, creating a `rule-<name>` ruleset per rule.
#config ruleset
# option name 'blocked-video'
# option type 'domain'
# option source 'inline'
# list entry 'youtube.com'
# list entry 'suffix:googlevideo.com'
#
#config rule
# option name 'video-via-main'
# option enabled '1'
# option order '50'
# list dst_ruleset 'blocked-video'
# option target 'group:main'
#
#config rule
# option name 'all-via-main'
# option enabled '1'
@@ -83,11 +131,16 @@ config inbound
# option type 'direct'
# option dpi 'fragment'
#
#config ruleset
# option name 'youtube'
# option source 'geosite'
# list category 'youtube'
#
#config rule
# option name 'youtube-fragment'
# option enabled '1'
# option order '50'
# list dst_domain 'geosite:youtube'
# list dst_ruleset 'youtube'
# option target 'egress:frag'
#
# A DNS resolver (type: doh|dot|plain|local|fakeip). `detour` routes its queries
+274 -2
View File
@@ -33,6 +33,30 @@
# be running. `start` raises ACTIVE_FLAG, `stop` clears it; hotplug/cron
# reconcile ONLY while the flag is up, so an admin `stop` STICKS — no
# background actor may resurrect interception behind a stopped daemon.
# * BEING REPLACED IS NOT BEING SWITCHED OFF. `restart` and `reload` (which is
# stop+start, i.e. every LuCI Save & Apply) both run through `stop`, and the
# daemon's SIGTERM teardown removes the fail-closed table unconditionally — it
# does not consult kill_switch at all. Between that teardown and the
# successor's first apply the init GUARANTEES a gap: it waits for the old
# process to exit (shater_wait_stopped), then runs `shaterd migrate`, then
# starts a daemon that still has to build an engine. So a restart is announced
# with RESTART_FLAG, which tells the outgoing daemon to leave the fail-closed
# holding plane STANDING — apply.TeardownExiting swaps it in with one nft
# transaction and then skips the delete, so the table is never absent, not even
# for the 80-90 ms the old arm-after-teardown order measured. A real `stop`
# raises no flag and therefore still means what it says.
# (A package UPGRADE does not come through here at all on apk v3: shater-core's
# script table is post-install / pre-deinstall / post-upgrade, with no
# pre-upgrade, so default_prerm — and its `stop` — runs only on REMOVAL.)
# * The FAIL-CLOSED PLANE MUST ALSO EXIST BEFORE THIS SCRIPT DOES. START=99 is
# after fw4 (19) and netifd (20), so at every boot the LAN forwards to the WAN
# in the clear for as long as it takes procd to decompress the daemon off
# flash and get an engine up. /etc/init.d/shater-armor (START=21) loads
# BOOT_ARMOR — a copy of the holding plane the daemon persists on every apply
# — to close that window. This script owns the DISARM half, and it owns it
# with a CLOSED LIST: an operator's `stop`, or a removal, and nothing else.
# Powering the box down must not — `shutdown` reaches stop_service too, and it
# is not a person switching the product off (see shater_stop_disarms).
# * The engine must never be permanently abandoned while interception stands:
# respawn retries are infinite (procd never gives up); a sustained-dead
# daemon is additionally escalated by the shater-cron watchdog.
@@ -50,11 +74,166 @@ ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Raised around a restart/reload, read by the OUTGOING `shaterd run` at SIGTERM:
# present => "you are being replaced, leave the fail-closed plane standing";
# absent => "you are being switched off, take everything down". tmpfs, so a
# power cut can never make the next boot look like a restart.
RESTART_FLAG=/var/run/shater.restarting
# The persisted fail-closed holding plane. Written by the daemon on every apply,
# loaded by /etc/init.d/shater-armor at boot. Its PRESENCE is the arm token, so
# removing it here is how a deliberate stop stops the next boot from blocking.
BOOT_ARMOR=/etc/shater/boot.nft
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# WHICH ACTION rc.common was invoked with, frozen at source time.
#
# rc.common does, in this order:
# initscript=$1; action=${2:-help}; shift 2; ...; . "$initscript"; $action "$@"
# so `action` is ALREADY assigned when this file is sourced, and every action then
# runs as a function in THAT SAME shell. MEASURED on the target (ImmortalWrt
# 25.12.1 r37978) with a throwaway probe init script, not read off documentation:
#
# /etc/init.d/X restart -> stop_service action=[restart], start_service [restart]
# /etc/init.d/X stop -> stop_service action=[stop]
# /etc/init.d/X reload -> reload_service action=[reload]
# `reboot` -> stop_service action=[SHUTDOWN] <-- see below
# the boot after it -> start_service action=[boot]
#
# A previous probe reported this variable EMPTY and the emptiness was written up as
# the defect. It was the probe: `sh -x /etc/init.d/shater restart` bypasses the
# `#!/bin/sh /etc/rc.common` shebang, so rc.common never runs, never assigns
# `action`, and the variable reads empty no matter what this file does.
#
# Frozen into our own variable because `action` is a short, generic name that other
# framework helpers also use as a local; a snapshot taken before any function runs
# cannot be shadowed later.
SHATER_RC_ACTION="$action"
# --- what an action MEANS --------------------------------------------------
#
# THE BUG THESE TWO PREDICATES REPLACE (v0.2.17, measured on the live router).
# The old stop_service was `case $action in restart|reload) keep;; *) DISARM;; esac`
# — an open default that swept up every action nobody had enumerated. `reboot` is
# one of them: procd runs the K-links with the action `shutdown`, so the shutdown
# path deleted the arm token on the way down and the next boot had nothing to load.
# The mechanism destroyed itself at exactly the moment it exists for. Instrument
# reading from the router, one minute apart across a reboot:
#
# 13:28 /etc/shater/boot.nft present
# ---- reboot (stop_service action=[shutdown] -> old `*` branch -> rm)
# 18s at_S22: NO_TABLE armor_file=NO_FILE
#
# So both lists below are POSITIVE and CLOSED. An action nobody thought about —
# `shutdown` above all, but also whatever a future procd invents — falls through
# both and changes nothing. The default now fails in the recoverable direction: at
# worst a boot arms when it need not have, which costs the second before the daemon
# applies and is still gated by shater-armor's own four state refusals. The old
# default failed in the direction of the plaintext window the feature was built to
# close.
#
# They are predicates rather than an inline `case` so the test gate can execute the
# real thing: it sources THIS FILE in /bin/sh and calls them with every action procd
# actually uses (shater/cmd/shaterd/initscript_test.go). A comment claiming
# `shutdown` is handled is what shipped last time.
# True only for the ONE action that means "the operator switched the product off".
# Deliberately not `shutdown`: powering a router down is not turning a feature off.
#
# NOT sufficient on its own — see shater_stop_disarms. `stop` is also how the
# package manager's plumbing reaches us, and a package manager is not a person.
shater_action_disarms() {
case "$1" in
stop) return 0 ;;
*) return 1 ;;
esac
}
# Is a package manager in the middle of a transaction RIGHT NOW?
#
# This is a state, read at the moment the decision is made, exactly like
# shater-armor's four refusals — not a record of an event. The same question is
# already asked (for the same reason: prerm/postinst plumbing is not a user
# action) by the detached bring-up in /etc/uci-defaults/30_shater-core.
shater_pkg_transaction() {
pidof apk >/dev/null 2>&1 && return 0
pidof opkg >/dev/null 2>&1 && return 0
return 1
}
# Is the main service still enabled at boot? Same glob, and for the same reason,
# as shater-armor's own check: `/etc/init.d/shater enabled` would source procd.sh
# and take a blocking flock, which is not something to do from inside a package
# manager's transaction.
shater_rc_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
# THE ACTUAL DISARM DECISION.
# $1 = action
# $2 = 1 when a package transaction is in flight
# $3 = 1 when the service is still enabled in rc.d
# All three are passed in rather than read inside, so the gate can drive every
# combination without a package manager or an /etc/rc.d.
#
# WHY IT IS NOT JUST THE ACTION. base-files' default_prerm runs, in this order:
#
# if [ "$PKG_UPGRADE" != "1" ]; then "$i" disable; fi
# "$i" stop
#
# so a package manager reaches stop_service wearing the operator's clothes. Two
# different intentions arrive as the same action, and the difference between them
# is readable at the moment of the decision:
#
# REMOVAL — prerm has ALREADY run `disable`, so S99shater is gone. The product
# is going away; the armor goes with it. (It is belt-and-braces even
# so: shater-armor refuses to arm without that symlink, and the whole
# init script is about to be deleted anyway.)
# REPLACED — the service is still enabled, so something intends to bring it
# back. That is not an operator switching anything off, and deleting
# the armor here would leave the next boot unprotected. "The next
# apply will rewrite it" is not an answer: the armor exists precisely
# to cover a reboot, and a reboot between an update and the first
# apply is how this product is deployed.
#
# MEASURED, because the paragraph above is about a path I got wrong once already.
# On THIS target (apk-tools 3.0.5, ImmortalWrt 25.12.1) shater-core's script table
# is post-install / pre-deinstall / post-upgrade, with NO pre-upgrade — so an apk
# UPGRADE never executes default_prerm and never calls `stop` at all. Verified with
# a real `apk fix --reinstall shater-core` while sampling the armor file: 245 625
# samples, zero disappearances, even with this guard mutated off. The upgrade half
# of this predicate is therefore defence-in-depth for a shape that is one
# `pre-upgrade` script (or a returning opkg lane) away, NOT a fix for an observed
# failure. The removal half is live today.
shater_stop_disarms() {
shater_action_disarms "$1" || return 1
# No package manager involved => a person typed it. The escape hatch must work.
[ "$2" = "1" ] || return 0
# A package transaction that has NOT disabled the service is replacing it.
[ "$3" = "1" ] && return 1
return 0
}
# True when a successor is coming, so the outgoing daemon should leave the
# fail-closed holding plane standing instead of removing it.
#
# `shutdown` is deliberately NOT a handoff either: nothing is coming, and the
# kernel that would hold the plane is going away with it. Leaving the flag down
# there also keeps the marker's meaning exact — it says "you are being replaced",
# and at shutdown nothing is.
shater_action_handoff() {
case "$1" in
restart|reload) return 0 ;;
*) return 1 ;;
esac
}
# --- helpers ---------------------------------------------------------------
# True only when the stack is explicitly enabled in UCI.
@@ -73,6 +252,25 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Announce/withdraw "this daemon is being replaced, not switched off". Read by
# `shaterd run` when it receives SIGTERM.
shater_mark_restart() {
mkdir -p "$(dirname "$RESTART_FLAG")" 2>/dev/null
: > "$RESTART_FLAG"
}
shater_clear_restart() { rm -f "$RESTART_FLAG"; }
# Remove the persisted boot armor, so the LAN is NOT blocked at the next boot
# before the daemon starts. Called from exactly two places, both of which are a
# statement about the PRODUCT rather than about this process: an operator typing
# `stop`, and a daemon binary that is no longer on the box. In neither case is
# anything going to come along and replace the armor with a real data plane, and a
# kill switch with nothing behind it is just a brick.
#
# NOT called on the shutdown path. That is the whole fix — see
# shater_action_disarms.
shater_disarm_boot() { rm -f "$BOOT_ARMOR"; }
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
@@ -136,9 +334,20 @@ start_service() {
# Guard: never claim to run without the daemon binary. A half-removed/failed
# shaterd upgrade must degrade to "plugin off", not to a box that thinks
# interception is live with nothing behind it.
#
# "Plugin off" now has to include DISARMING. With the boot armor in play, a
# missing binary is the one case where the fail-closed plane could stand
# forever with nothing able to replace it: the armor loads at START=21, the
# daemon never starts, and every later boot repeats it. The product being gone
# is not a security event — it is an uninstall — so the plane comes down and
# the LAN returns to plain routing, loudly.
if [ ! -x "$PROG" ]; then
shater_clear_restart
shater_disarm_boot
rm -f "$ACTIVE_FLAG"
nft delete table inet shater 2>/dev/null
_slog -p daemon.err \
"shaterd binary missing/not executable at $PROG — refusing to start (LAN stays on plain routing)"
"shaterd binary missing/not executable at $PROG — refusing to start; the fail-closed plane and its boot armor have been REMOVED (LAN back to plain routing, unprotected). Reinstall shaterd."
return 0
fi
@@ -150,9 +359,28 @@ start_service() {
# running, which is the boot case.
shater_wait_stopped
# The predecessor is gone and has already consumed the flag (it reads it in its
# SIGTERM handler). Withdraw it now, so a LATER `stop` is unambiguous even if
# this start fails further down.
shater_clear_restart
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
"$PROG" migrate >/dev/null 2>&1
#
# THE FAILURE IS LOGGED, NOT SWALLOWED. This is the only place the schema
# migration runs at boot (`shaterd run`, the SIGHUP reconcile and the panel's
# config write all read UCI directly), so if it fails here it does not get
# retried until the next start. And it CAN fail for a mundane reason — a full
# /overlay makes `uci commit` fail — after which the config still carries the
# schema-v1 `dst_domain`/`dst_ip` options. The daemon holds every rule that
# still has them DISABLED and reports it, so nothing is silently misrouted, but
# rules the operator wrote are then not in force and the reason has to be
# visible somewhere. Hence: log the binary's own stderr, and start anyway —
# refusing to start would take the admin panel down with it, and the panel is
# the only way to fix the box.
local migrate_out
migrate_out=$("$PROG" migrate 2>&1) || _slog -p daemon.err \
"UCI schema migration FAILED: ${migrate_out:-no output from $PROG migrate}. Starting anyway; routing rules that still carry the removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space on /overlay and re-run '$PROG migrate', or restart the service."
procd_open_instance shater
# shaterd runs in the FOREGROUND under procd (must never daemonize). `run` is
@@ -200,6 +428,45 @@ start_service() {
}
stop_service() {
# Say WHY we are stopping before procd sends the signal, because the daemon
# cannot tell from the signal alone and the answer changes what it leaves in
# the kernel. Two INDEPENDENT questions, and the old code conflated them into
# one two-armed `case` whose else-branch answered both wrongly for `shutdown`:
#
# 1. IS A SUCCESSOR COMING (this process only)? restart / reload.
# Raise RESTART_FLAG so the outgoing daemon replaces its data plane with
# the fail-closed HOLDING plane instead of removing it. The gap until the
# successor applies is not a moment: this script waits out the old
# process, runs `shaterd migrate`, then starts a daemon that must build an
# engine — all of it, before this flag existed, with `lan -> wan ACCEPT`
# and nothing else.
#
# 2. IS THE PRODUCT BEING SWITCHED OFF (across boots)? `stop` — and only
# `stop`, and only when a PERSON is behind it (shater_stop_disarms; the
# package manager reaches us through `stop` too). Then the boot armor goes
# with it, so the next boot does not quietly reinstate what the operator
# just switched off — the same rule ACTIVE_FLAG has always enforced for
# hotplug/cron.
#
# `shutdown` answers NO to both, which is the defect this replaced: a reboot is
# not a successor and it is certainly not an operator switching the product off.
# It is the boot the armor exists for. An upgrade answers NO to the second for
# the same kind of reason.
if shater_action_handoff "$SHATER_RC_ACTION"; then
shater_mark_restart
else
shater_clear_restart
fi
local in_pkg=0 rc_en=0
shater_pkg_transaction && in_pkg=1
shater_rc_enabled && rc_en=1
if shater_stop_disarms "$SHATER_RC_ACTION" "$in_pkg" "$rc_en"; then
shater_disarm_boot
elif [ "$in_pkg" = "1" ] && shater_action_disarms "$SHATER_RC_ACTION"; then
_slog -p daemon.info \
"stop came from a package transaction that left the service enabled — keeping the boot armor, so being replaced cannot leave the next boot unprotected"
fi
# Drop the live-flag FIRST so a concurrent hotplug/cron tick cannot rebuild
# what we are about to tear down. procd then sends SIGTERM to `shaterd run`,
# which runs its OWN honest teardown (engine.Close + netplane restore) — we
@@ -219,6 +486,11 @@ reload_service() {
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
#
# Marked EXPLICITLY as well as via SHATER_RC_ACTION: this is the path a routine
# Save & Apply takes, so it is the one that must not depend on reading an
# rc.common variable correctly. Belt and braces, one line.
shater_mark_restart
stop
start
}
@@ -0,0 +1,162 @@
#!/bin/sh /etc/rc.common
# /etc/init.d/shater-armor — the fail-closed plane, before the daemon exists.
#
# WHAT THIS CLOSES
#
# /etc/init.d/shater is START=99. By then fw4 (START=19) has long since loaded
# `lan -> wan ACCEPT` and netifd (START=20) has brought the LAN bridge up, so the
# router forwards LAN traffic to the WAN in the clear from the moment the link
# comes up until `shaterd run` has been decompressed off flash, has waited out any
# predecessor, has migrated UCI, has read the config and has installed its first
# table. On router-class hardware with a UPX-packed binary that is seconds — and
# they are exactly the seconds in which Wi-Fi finishes associating and every
# client on the network reconnects and starts talking. `kill_switch=closed` was
# configured the whole time and covered none of it.
#
# There was nothing in the package that could cover it either: no /etc/nftables.d
# include, no `nft -f` in uci-defaults. Protection existed only inside a Go
# process that had not started yet.
#
# HOW
#
# The daemon persists a copy of its fail-closed HOLDING plane (the same ruleset it
# installs when the engine is down: one forward chain, LAN-to-LAN and router
# traffic accepted, everything else from the diverted devices dropped) to
# $ARMOR on every apply. This script loads it early. When the daemon comes up it
# replaces the table atomically — the ruleset begins with `delete table` and adds
# its own in one netlink transaction — so there is never a moment with no table.
#
# `iifname` matches by NAME at packet time, not by ifindex at load time, so
# loading this before netifd has created br-lan is fine: the rules simply start
# matching when the device appears. That is why START can sit here rather than
# racing netifd.
#
# START=21: after fw4 (19) and netifd (20), because fw4's own start tears its
# table down and rebuilds it and we do not want to be in the middle of that, and
# because there is nothing to protect before the LAN device is being created. The
# residual exposure is the fraction of a second between netifd's `ifup` and this
# script, against seconds-to-a-minute before.
#
# THE ESCAPE HATCHES (a kill switch that cannot be switched off is a brick)
#
# These are STATE checks, evaluated here, at the moment of arming — not a record
# of something that happened on the way down. That distinction is the whole
# lesson of v0.2.17: the arm token was deleted by an EVENT on the shutdown path
# ("this looks like a stop"), and since `reboot` also runs the K-links, the
# mechanism reliably erased itself on the one transition it was built for. An
# event on the way down cannot be trusted to describe the world on the way up; a
# question asked on the way up can be.
#
# * $ARMOR only exists while the daemon's last applied config was BOTH enabled
# and fail-closed. `globals.enabled=0` and `kill_switch=open` each remove it
# at the next apply, and an operator typing `/etc/init.d/shater stop` removes
# it there and then. Powering the box off does NOT.
# * We refuse to arm when the main service is disabled in rc.d, or when the
# daemon binary is gone — in either case nothing would ever come along to
# replace the armor with a real data plane. These two are what makes a
# genuinely uninstalled/disabled product safe REGARDLESS of what the file
# says, which is why they are checked here rather than trusted to have been
# acted on earlier.
# * We refuse to arm when UCI can be read AND says the stack is disabled. A
# config that cannot be read is NOT a refusal: that case is precisely why the
# armor is a file rather than a query.
# * The chain hooks `forward` only, so SSH, LuCI and the admin panel (all input
# hook, to the router's own addresses) stay reachable. The operator can always
# get in and undo this.
#
# Note what a bare `/etc/init.d/shater stop` does NOT mean: it does not survive a
# reboot, because S99shater is still linked and procd starts the daemon again. So
# "stopped" is not a durable off-state and this script must not be designed as if
# it were — the durable ones are `disable` (no S??shater) and `globals.enabled=0`,
# and those are the two refusals above.
#
# busybox ash only — no bashisms.
START=21 # after firewall (19) and network (20), long before shater (99)
STOP=89
ARMOR=/etc/shater/boot.nft
PROG=/usr/bin/shaterd
# Syslog line that honors globals.log_syslog, like the other two inits. An
# unreadable UCI leaves the option empty => ON, which is what we want here: the
# one boot where the config cannot be read is the boot worth logging.
_slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater-armor "$@"
}
# Is the MAIN service enabled at boot? Answered by looking for its rc.d symlink
# rather than by running `/etc/init.d/shater enabled`: that is a USE_PROCD script,
# so every action of it sources procd.sh, which takes a blocking flock — and this
# runs at START=21, in the middle of boot, for a question a glob answers exactly
# as well. The START number is not hardcoded; any S<NN>shater counts.
shater_service_enabled() {
local f
for f in /etc/rc.d/S[0-9][0-9]shater; do
[ -e "$f" ] && return 0
done
return 1
}
start() {
# No saved plane => the stack has never applied an enabled, fail-closed config
# (or it was explicitly switched off). Nothing to do, and nothing to say.
[ -f "$ARMOR" ] || return 0
[ -s "$ARMOR" ] || {
_slog -p daemon.err "$ARMOR is empty — NOT arming; the LAN is unprotected until shaterd starts"
return 0
}
# Never arm something nothing can disarm.
[ -x "$PROG" ] || {
_slog -p daemon.err \
"$PROG is missing — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
shater_service_enabled || {
_slog -p daemon.warn \
"the shater service is disabled in rc.d — NOT arming (nothing would replace the block with a working data plane); the LAN stays on plain routing"
return 0
}
# A READABLE config that says "off" wins over the saved plane (it means the
# daemon was stopped before it could disarm). An UNREADABLE config does not:
# that is the case this whole mechanism exists for.
en=$(uci -q get shater.globals.enabled 2>/dev/null)
if [ -n "$en" ] && [ "$en" != "1" ]; then
rm -f "$ARMOR"
_slog -p daemon.info "globals.enabled=$en — boot armor removed, not arming"
return 0
fi
command -v nft >/dev/null 2>&1 || {
_slog -p daemon.err "nft is not installed — cannot arm; the LAN is unprotected until shaterd starts"
return 0
}
# Validate before loading: a truncated/incompatible snapshot must not leave a
# half-built table behind on the one boot it is needed.
if ! nft -c -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.err \
"$ARMOR did not validate (nft -c) — NOT arming; the LAN is unprotected until shaterd starts"
return 0
fi
if nft -f "$ARMOR" >/dev/null 2>&1; then
_slog -p daemon.warn \
"fail-closed plane armed from $ARMOR: LAN->WAN forwarding is BLOCKED until shaterd applies. SSH, LuCI and the admin panel stay reachable."
else
_slog -p daemon.err \
"could not load $ARMOR — the LAN is unprotected until shaterd starts"
fi
return 0
}
stop() {
# Deliberately a NO-OP. By the time anything stops this service the daemon owns
# `inet shater`, and deleting the table here would dismantle a LIVE data plane
# on the strength of a service that only ever ran for one second at boot. The
# disarm paths that matter live where the decision is actually made:
# /etc/init.d/shater stop (operator switched it off) and the daemon itself
# (globals.enabled=0 / kill_switch=open).
return 0
}
@@ -59,6 +59,68 @@ if uci -q get shater.globals >/dev/null 2>&1 || [ -f /etc/config/shater ]; then
uci -q commit shater
fi
# Introduce the daemon-created `shater-l3` TUN to fw4 (L3 ingress, D-L3). The
# daemon policy-routes LAN ICMP into that device from OUR nft table
# `inet shater`, but nftables runs EVERY table on every packet and a drop in
# any one of them wins — an accept in `inet shater` cannot override fw4. And
# fw4 WILL drop this forward: netifd knows nothing about a device the daemon
# creates at runtime, so it belongs to no zone and falls into fw4's zone-less
# defaults (REJECT). The device has to be declared to fw4 itself; it cannot be
# fixed from our own table.
#
# Seeded UNCONDITIONALLY (not gated on globals.l3_tunnel): uci-defaults run
# once, so gating on the option would require re-running this script when the
# option is flipped later — which never happens. An idle zone is harmless: its
# device match is a plain iifname/oifname STRING compare that simply never hits
# while the TUN does not exist.
#
# Idempotency: `config zone`/`config forwarding` are normally ANONYMOUS
# sections, and a naive `uci add firewall zone` would append a duplicate on
# every re-run (uci-defaults re-run on package upgrade/reinstall). All sections
# here are NAMED instead, guarded by an existence check — a re-run re-finds the
# section and touches nothing.
seed_l3_zone() {
# No fw4 on this image (bare nftables build) => nothing drops the forward
# on fw4's behalf and there is nothing to punch through.
[ -f /etc/config/firewall ] || return 0
if ! uci -q get firewall.shater_l3 >/dev/null; then
uci set firewall.shater_l3=zone
uci set firewall.shater_l3.name='shater_l3'
uci set firewall.shater_l3.input='REJECT'
uci set firewall.shater_l3.output='ACCEPT'
uci set firewall.shater_l3.forward='REJECT'
uci set firewall.shater_l3.masq='0'
# INERT TODAY, kept for the day it is not. mtu_fix clamps forwarded TCP
# MSS to the route MTU — but shater-l3 is 65535 (deliberately: at any
# smaller value the kernel fragments into the device, and the flow
# dispatcher refuses to judge a fragment and lets the stack forge the
# echo reply — see l3MTU in shater/generate/inbound.go), so the clamp has
# nothing to clamp to. And only ICMP is ever marked into this device, so
# no TCP rides here to be clamped in the first place. It earns its keep
# the moment either of those changes; removing it would make that day
# silent.
uci set firewall.shater_l3.mtu_fix='1'
# `list device`, deliberately NOT the usual `list network`: fw4
# resolves a zone's networks through netifd, and netifd never learns
# about a device the daemon creates at runtime — a stub interface
# (proto none) would need to be brought UP to contribute an l3_device,
# and nothing ever brings it up, so `list network` resolves to an
# EMPTY device set and fw4 keeps dropping the forward. `list device`
# instead compiles to an iifname/oifname STRING match, valid before
# the TUN exists and matching from the moment shaterd creates it —
# no netifd involvement and no firewall reload at enable time. Do not
# "normalize" this to `list network` in a refactor; it breaks silently.
uci add_list firewall.shater_l3.device='shater-l3'
fi
if ! uci -q get firewall.shater_l3_fwd >/dev/null; then
uci set firewall.shater_l3_fwd=forwarding
uci set firewall.shater_l3_fwd.src='lan'
uci set firewall.shater_l3_fwd.dest='shater_l3'
fi
uci -q commit firewall
}
seed_l3_zone
# Bring the UCI schema forward on upgrade (idempotent; refuses a newer schema).
[ -x /usr/bin/shaterd ] && /usr/bin/shaterd migrate >/dev/null 2>&1
@@ -110,8 +172,27 @@ SHATER_BRINGUP='
done
[ -x /etc/init.d/shater ] && /etc/init.d/shater enable
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron enable
# The boot-time fail-closed armor. `enable` only — it is a one-shot that loads
# the persisted holding plane at START=21, and running it NOW would install a
# block on a live box moments before the daemon replaces it anyway. It has to
# be enabled here regardless of whether the stack is on: the file it loads only
# exists while the daemon wants it to, so an enabled-but-unarmed service is a
# no-op, and enabling it later would mean the first boot after an upgrade is
# the one boot still exposed.
[ -x /etc/init.d/shater-armor ] && /etc/init.d/shater-armor enable
[ -x /etc/init.d/shater ] && /etc/init.d/shater restart
[ -x /etc/init.d/shater-cron ] && /etc/init.d/shater-cron restart
# Fold the seeded shater_l3 zone into the LIVE ruleset — matters on a live
# opkg/apk install only, where firewall started long before our commit and
# nothing else would re-read it until the next reboot. Gated on the fw4
# table actually being loaded: at FIRST boot this job can run before the
# S19 firewall start, and an early reload would install a ruleset built
# from a half-initialized netifd AND make the later start a no-op (fw4
# start skips when its table already exists). No table => the pending S19
# start reads the committed config by itself, no reload needed.
if nft list tables 2>/dev/null | grep -q "inet fw4"; then
[ -x /etc/init.d/firewall ] && /etc/init.d/firewall reload
fi
exit 0
'
SHATER_TMO=""
+5 -5
View File
@@ -38,9 +38,9 @@ PKG_NAME:=shaterd
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
# that version, so a lost env can never silently ship a stale one again.
# ci/build-feed-apk.sh exports them into the SDK build env. ci/sdk-build-apk.sh
# then ASSERTS that the produced .apk really carries that version, so a lost env
# can never silently ship a stale one again.
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
# are not "the release version"; releases are named by the tag.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
@@ -104,8 +104,8 @@ define Package/shaterd/install
$(INSTALL_BIN) $(CURDIR)/files/$(SHATERD_BIN) $(1)/usr/bin/shaterd
endef
# This package ships ONLY the binary — no init script — so opkg's default
# postinst never touches the running service. On `opkg upgrade shaterd` the new
# This package ships ONLY the binary — no init script — so the package manager's
# postinst never touches the running service. On `apk upgrade shaterd` the new
# ELF lands at /usr/bin/shaterd while the OLD image keeps running from its
# unlinked inode: the upgrade silently has no effect until the next reboot, and
# meanwhile the new CLI (`shaterd reconcile`, `status`, `mint-token` — invoked by
+26
View File
@@ -18,6 +18,32 @@ type URLTestOutboundOptions struct {
// lx: SPEC 019 v2 — load-balancing.
Mode string `json:"mode,omitempty"` // least_test (default) | round_robin
Balancer *URLTestBalancerOptions `json:"balancer,omitempty"`
// lx: health board §5.C — SelfCheck stands the group's OWN background
// health-check up or down. nil/absent == true, so every existing config keeps
// today's behaviour.
//
// Why this exists at all: a urltest group probes its members BY ITSELF — a
// warm-up sweep at PostStart and a ticker for as long as traffic keeps
// touching it — and it dials the members' outbounds DIRECTLY, from the
// router, over whatever the default WAN route is. For a group that traffic
// actually flows through, that is exactly right: the probe travels the same
// path the connections do. But for a group NO routing rule reaches, that
// same probe measures a path nothing uses — and it stores the result under
// the members' BASE tags, which every health consumer then reads as "the
// node's health". A node that is blocked on the direct WAN and perfectly
// alive behind a tunnel therefore reads "dead" the moment such a group
// probes it; the reading is not merely stale, it is FALSE, and it poisons
// the shared board for everyone (selection, the panel, the observatory's
// freshness gate). SelfCheck=false is how the control plane stands such a
// group's own schedule down: the shater engine computes which groups the
// applied rules actually reach (the observatory's used-set) and disables
// the self-check on the rest, so the ONLY prober left is the observatory —
// which probes along the real dial paths and nothing else.
//
// The flag suppresses only the group's own SCHEDULE (the PostStart warm-up
// and the Touch ticker). An EXPLICIT CheckOutbounds/URLTest call — the
// adapter interface a human or an API invokes on purpose — still works.
SelfCheck *bool `json:"self_check,omitempty"`
}
// URLTestBalancerOptions configures round_robin: a fixed-size pool of live nodes, lazily
+2 -1
View File
@@ -8,7 +8,8 @@
"dev": "vite",
"build": "tsc --noEmit && vite build",
"preview": "vite preview",
"typecheck": "tsc --noEmit"
"typecheck": "tsc --noEmit",
"test": "node --test src/*.test.ts"
},
"dependencies": {
"react": "^18.3.1",
+221
View File
@@ -262,6 +262,150 @@
}
}
/* ---- fixture band (dev builds only; see App.tsx MockBanner) ----
Deliberately outside the crit/amber vocabulary: nothing is wrong with the
router, there is no router. The hazard hatch is the service-sticker language a
piece of network hardware already uses for "this unit is not in service". */
.mock-band {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
margin-top: calc(var(--u, 8px) * 2);
padding: 10px 14px;
border: 1px dashed var(--faint);
border-radius: 9px;
background: repeating-linear-gradient(
-45deg,
var(--sink),
var(--sink) 9px,
var(--panel) 9px,
var(--panel) 18px
);
}
.mock-band-tag {
flex-shrink: 0;
align-self: flex-start;
padding: 3px 7px;
border: 1px solid var(--faint);
border-radius: 4px;
background: var(--raised);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 700;
letter-spacing: 0.14em;
color: var(--dim);
}
.mock-band-copy {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 2px;
}
.mock-band-headline {
font-family: var(--font-mono);
font-size: 12.5px;
font-weight: 700;
letter-spacing: 0.02em;
color: var(--ink);
}
.mock-band-detail {
font-size: 12.5px;
line-height: 1.5;
color: var(--dim);
max-width: 76ch;
}
.mock-band-detail code {
font-family: var(--font-mono);
font-size: 11.5px;
color: var(--ink);
}
/* ---- commit-confirm band (every page except Apply, which has the full panel) ----
Same plate as the protection banner so the two read as one family; the seconds
are the loud element because they are the only thing that is running out. */
.cfm-band {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
margin-top: calc(var(--u, 8px) * 2);
padding: 10px 14px;
border: 1px solid color-mix(in srgb, var(--amber) 50%, var(--groove));
border-radius: 9px;
background: linear-gradient(180deg, color-mix(in srgb, var(--amber) 10%, var(--raised)), var(--raised));
box-shadow: 0 1px 0 var(--edge) inset;
}
.cfm-band-count {
display: flex;
align-items: baseline;
gap: 2px;
flex-shrink: 0;
font-family: var(--font-mono);
color: var(--amber);
}
.cfm-band-num {
font-size: 22px;
font-weight: 700;
font-variant-numeric: tabular-nums;
line-height: 1;
}
.cfm-band-unit {
font-size: 11px;
letter-spacing: 0.06em;
}
.cfm-band-copy {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 3px;
}
.cfm-band-headline {
font-family: var(--font-mono);
font-size: 12.5px;
font-weight: 700;
letter-spacing: 0.02em;
color: var(--ink);
}
.cfm-band-detail {
font-size: 12.5px;
line-height: 1.5;
color: var(--dim);
max-width: 76ch;
}
.cfm-band-actions {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1);
flex-shrink: 0;
}
.cfm-band-link {
padding: 6px 11px;
border: 1px solid var(--groove);
border-radius: 6px;
font-family: var(--font-mono);
font-size: 11px;
letter-spacing: 0.06em;
text-transform: uppercase;
text-decoration: none;
color: var(--ink);
background: var(--raised);
}
.cfm-band-link:hover {
border-color: var(--accent);
color: var(--accent);
}
@media (max-width: 720px) {
.cfm-band {
flex-wrap: wrap;
}
.cfm-band-actions {
width: 100%;
justify-content: flex-end;
}
}
/* ---- last-apply findings (Overview) ----
Severity carries the colour; the accent is reserved for interactive controls. */
.findings {
@@ -312,6 +456,13 @@
.finding--warning {
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
}
/* The daemon's "the list is capped" disclosure. Dashed, because the row is about
what ISN'T here — it must not read as one more finding to work through. */
.finding--truncated {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 40%, var(--groove));
background: var(--panel);
}
.finding-copy {
flex: 1;
min-width: 0;
@@ -405,3 +556,73 @@
color: var(--dim);
max-width: 74ch;
}
/* ---- inline rename (shared) ----
The pencil-in-the-row interaction: click the ✎ beside a name, type over it,
Enter commits / Esc cancels / blur commits. Lifted out of Devices.css when
Nodes grew the same affordance — one interaction, one set of rules, so the two
pages can never drift apart. `--locked` is the same control with the action
withheld: it stays visible and focusable-looking so a missing rename reads as
a stated rule, not a dead button. */
.inline-rename {
flex: none;
display: inline-flex;
align-items: center;
justify-content: center;
width: 22px;
height: 22px;
padding: 0;
border: 1px solid transparent;
border-radius: 5px;
background: none;
color: var(--faint);
font-size: 12px;
line-height: 1;
cursor: pointer;
transition: color 0.15s, background 0.15s, border-color 0.15s;
}
.inline-rename:hover:not(:disabled) {
color: var(--accent);
background: color-mix(in srgb, var(--accent) 12%, transparent);
}
.inline-rename:focus-visible {
color: var(--accent);
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.inline-rename:disabled {
opacity: 0.5;
cursor: default;
}
/* Withheld, not broken: keep the glyph readable and let the cursor say "there is
a reason" rather than dimming it into invisibility. */
.inline-rename--locked {
opacity: 0.75;
cursor: help;
}
.inline-rename--locked:hover {
color: var(--dim);
background: none;
}
.inline-rename-input {
min-width: 0;
max-width: 24ch;
padding: 4px 8px;
border: 1px solid var(--accent);
border-radius: 6px;
background: var(--sink);
color: var(--ink);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
box-shadow: 0 1px 2px var(--shadow) inset;
}
.inline-rename-input:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.inline-rename-input:disabled {
opacity: 0.55;
}
+118 -7
View File
@@ -1,12 +1,14 @@
import './App.css'
import { useCallback, useEffect, useState } from 'react'
import { Faceplate, FaceplateHeader, Led, Module } from './components'
import { Button, Faceplate, FaceplateHeader, Led, Module } from './components'
import type { LedVariant } from './components'
import { ApiError, MOCK, getStatus } from './api'
import { ApiError, MOCK, confirm as apiConfirm, getStatus } from './api'
import type { Status } from './api'
import { usePendingConfirm } from './pendingConfirm'
import { bootstrapSession } from './session'
import { ROUTES, navigate, useRoute } from './router'
import { protectionState } from './planeState'
import { engineState, protectionState } from './planeState'
import { truncationNote } from './findings'
import type { Route } from './router'
import { Overview, Placeholder, Nodes, Routing, Apply, DNS, Devices, Targets, Settings, Profiles, Insights, Networks } from './pages'
@@ -99,12 +101,102 @@ export function App() {
footer={<StatusBar status={status} />}
>
<Nav route={route} />
<MockBanner />
<PlaneBanner status={status} route={route} />
<ConfirmBand route={route} onChanged={() => void refreshStatus()} />
<Page route={route} status={status} onStatusChange={() => void refreshStatus()} />
</Faceplate>
)
}
/**
* The commit-confirm countdown, on every page.
*
* The daemon arms an auto-rollback on EVERY apply, but only the Apply page ever
* said so: press Apply on Routing, read "Applied", walk away, and the router
* reverts a minute later with nothing on screen having mentioned it. This band
* carries that deadline — and the button that stops it — to wherever the operator
* actually is.
*
* Suppressed on Apply, which renders the full control room for the same window
* (and reads the same record, so a reload no longer loses the countdown there
* either).
*/
function ConfirmBand({ route, onChanged }: { route: Route; onChanged: () => void }) {
const armed = usePendingConfirm()
const [busy, setBusy] = useState(false)
const [error, setError] = useState<string | null>(null)
// Keeping the config is the only action offered here; rolling back early is a
// deliberate act with its own before/after readout, and that lives on Apply.
const keep = useCallback(async () => {
setBusy(true)
setError(null)
try {
const r = await apiConfirm()
if (r.error) setError(r.error)
} catch (e) {
setError(e instanceof Error ? e.message : 'request failed')
} finally {
setBusy(false)
onChanged()
}
}, [onChanged])
if (!armed || route === 'apply') return null
return (
<div className="cfm-band" role="alert">
<Led variant="amber" pulse />
<div className="cfm-band-count" role="timer" aria-label={`${armed.remaining} seconds until auto-rollback`}>
<span className="cfm-band-num">{armed.remaining}</span>
<span className="cfm-band-unit">s</span>
</div>
<div className="cfm-band-copy">
<span className="cfm-band-headline">This config is live but not kept</span>
<span className="cfm-band-detail">
{error
? `Couldn’t keep it — ${error}. Try again, or open Apply.`
: 'Every apply arms an auto-rollback. Keep this config before the timer runs out, or the router reverts to the last-good one.'}
</span>
</div>
<div className="cfm-band-actions">
<Button variant="primary" onClick={() => void keep()} disabled={busy}>
{busy ? 'Keeping…' : 'Keep this config'}
</Button>
<a className="cfm-band-link" href="#/apply" onClick={() => navigate('apply')}>
Apply page
</a>
</div>
</div>
)
}
/**
* Says, on every page, that nothing on screen came from a router.
*
* Only a DEV build can ever render this — the fixtures are not in a production
* bundle (api.ts initMockBackend), so an operator cannot reach this state at all.
* It is here for the person who CAN: a footer line reading "DEMO DATA" is easy to
* work past for an afternoon and then screenshot into a bug report, and every
* number above it is invented.
*/
function MockBanner() {
if (!MOCK) return null
return (
<div className="mock-band" role="status">
<span className="mock-band-tag">FIXTURES</span>
<div className="mock-band-copy">
<span className="mock-band-headline">No router is being read</span>
<span className="mock-band-detail">
Every reading on this page is invented by <code>src/mock.ts</code> for offline
development. Drop <code>?mock</code> from the address to talk to a daemon.
</span>
</div>
</div>
)
}
/**
* The protection state, pinned under the nav on every page EXCEPT Overview
* (which shows the same state as its own headline readout — see planeState.ts).
@@ -130,12 +222,15 @@ function PlaneBanner({ status, route }: { status: Status | null; route: Route })
const criticals = (status.warnings ?? []).filter((w) => w.severity === 'critical').length
const state = protectionState(status)
// The published list is capped at 50, so with a note attached the count is a
// floor. Say "at least" rather than quoting a total the daemon didn't send.
const atLeast = truncationNote(status.warnings) ? 'At least ' : ''
// Wording comes from the shared source of truth so the banner and Overview can
// never describe the same router differently.
const headline = state.alarm
? state.headline
: `${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
: `${atLeast}${criticals} protection ${criticals === 1 ? 'gap' : 'gaps'} from the last apply`
const detail = state.alarm
? state.detail
: 'Something you configured isn’t in effect. Review the findings before relying on it.'
@@ -225,14 +320,30 @@ function StatusBar({ status }: { status: Status | null }) {
)
}
/**
* The one lamp that is on screen no matter which page you are on.
*
* It used to read `status.running`, which the daemon hardcoded to `true` — so the
* "Offline" branch could never be reached and the plate said "Online" through an
* engine that had failed to start. It now asks {@link engineState}, whose whole
* job is to be able to answer "down", and refuses to guess when nothing has been
* reported: an unlit socket, not a green light.
*/
function masterIndicator(
phase: Phase,
status: Status | null,
): { label: string; variant: LedVariant; pulse?: boolean } {
if (phase === 'loading' || !status) return { label: 'Linking', variant: 'off' }
if (status.running && status.active) return { label: 'Online', variant: 'on', pulse: true }
if (status.running) return { label: 'Standby', variant: 'amber' }
return { label: 'Offline', variant: 'crit' }
switch (engineState(status)) {
case 'down':
return { label: 'Engine down', variant: 'crit' }
case 'up':
return status.active
? { label: 'Online', variant: 'on', pulse: true }
: { label: 'Standby', variant: 'amber' }
default:
return { label: 'Unknown', variant: 'off' }
}
}
function UnauthPlate() {
+336 -49
View File
@@ -13,8 +13,13 @@
// serves in-memory fixtures instead of hitting the network, so `npm run dev`
// and screenshot runs render without a live backend. A real backend in dev is
// reachable instead via the Vite proxy in vite.config.ts (no flag ⇒ real fetch).
//
// THE FIXTURES ARE A DEV-BUILD-ONLY ARTEFACT — see initMockBackend below. They
// used to be a plain static import, decided at RUNTIME off `location.search`, so
// the invented router shipped inside the binary that goes on real hardware and a
// link ending in `?dev` painted a healthy appliance without making one request.
import * as mock from './mock'
import { armPendingConfirm, clearPendingConfirm, noteConfirmTimeout } from './pendingConfirm'
// --- error type -------------------------------------------------------------
@@ -51,6 +56,38 @@ export class ApiError extends Error {
*/
export type Plane = 'full' | 'hold' | 'none'
/**
* Where the router's traffic actually ENDS UP, decided by the daemon from the
* engine config it is running (apply.Status.traffic ← generate.TrafficOf).
*
* tunnel — the default route goes into a tunnel: everything not matched by a
* more specific rule is proxied.
* split — the default leaves directly, but some rules do tunnel their traffic.
* direct — the default leaves directly and nothing is tunnelled at all.
* blocked — the default is the fail-closed backstop: unmatched traffic is
* dropped, not let out. Nothing leaks.
*
* `plane` DOES NOT ANSWER THIS and must never be read as if it did. `plane` says
* how much of the data plane is installed (nft table, policy routing, engine up);
* a router whose only rule is `default → direct` has all of it and sends the whole
* LAN out the plain WAN with its real address. That combination — plane "full",
* traffic "direct" — was live on a user's router under a green "Protected" LED.
*/
export type TrafficVerdict = 'tunnel' | 'split' | 'direct' | 'blocked'
export interface Traffic {
// '' or absent ⇒ not known (daemon that predates this field, nothing applied
// yet, or the plane is on hold). NEVER treat unknown as 'tunnel'.
verdict?: TrafficVerdict | ''
// The outbound tag the engine's default route names, in the engine's own
// vocabulary ("direct", "block", a node/group tag). Diagnostic — wording is
// driven by `verdict`, never by parsing this.
default?: string
// How many of the engine's route rules send their matched traffic into a tunnel.
// Separates "some of your traffic is protected" from "none of it is".
tunnel_rules?: number
}
/**
* One thing the last apply could not do. Deliberately fail-OPEN with a warning
* rather than refusing the whole config (the alternative was taking the network
@@ -88,6 +125,10 @@ export interface Status {
// How much of the data plane is installed. Absent on older daemons ⇒ unknown,
// in which case the UI shows nothing rather than guessing "full".
plane?: Plane
// Where the traffic actually goes under the running config. Absent on older
// daemons ⇒ unknown; see TrafficVerdict for why this is a separate question
// from `plane`.
traffic?: Traffic
// Findings from the last apply. ALWAYS an array from the daemon (never null);
// empty means the last apply was clean. Pre-sorted critical-first and capped at
// 50, where a truncated list ends with an `info` entry saying "suppressed".
@@ -354,19 +395,142 @@ export interface GroupHealth {
* any more and nothing to report here beyond the groups themselves.
*/
/**
* One hop of one chain, measured where that hop actually sits in the path.
*
* This is the reading the daemon always took and never showed. A chain is not a
* target with a single health — it is an ordered series of them, and the only
* question an operator ever asks about a broken chain is WHICH hop broke. The
* end-to-end exit reading cannot answer that: it says "the path is dead" for a
* four-hop chain and leaves the person to guess between four suspects.
*
* WIRE ORDER. `index` is 1-based and counts hops in the order the router dials
* them: hop 1 is the first physical hop, and each later hop is dialled THROUGH
* the ones before it. The hop carrying `exit: true` — always the largest index —
* is where traffic leaves for the internet. A leading `egress:` in the chain's
* configured Hops is NOT a numbered hop: the daemon lifts it into the entry
* detour of hop 1, so a chain written `egress:ewan → node:awgout → group:sub0`
* reports two hops, not three. Anything zipping this against the model's Hops
* must drop that leading egress first and give up on labelling entirely if the
* counts still disagree — a chain that splices sub-chains gets flattened here,
* and a confidently WRONG hop name is worse than no name.
*
* `tag` is the engine-side outbound (`chain-<name>-h2`). Debugging and tooltips
* only; it is never a label to put in front of a person.
*
* ORDERED WALK — THE READING STOPS AT THE FIRST DEAD HOP. Hops are NOT measured
* independently, and never were measurable that way: hop 3 is dialled THROUGH
* hop 2, so probing hop 3 while hop 2 is down measures hop 2 a second time and
* learns nothing about hop 3. The daemon therefore walks the path in wire order
* and stops at the first hop that does not answer. Every hop below that one is
* left undialled and reported `state: "untested"` — no measurement exists —
* carrying {@link ChainHopBlock} in `blocked_by` to name the hop that stopped the
* walk. So a chain never reports a dead hop with a live hop below it; that shape
* is not a rare case, it is unreachable.
*
* NODE HOP vs GROUP HOP. For `kind: "node"` the hop IS the measurement: `total`
* is 1, the counters follow its own state, and `selected` is ''. For
* `kind: "group"` the counters roll up that hop's per-hop member COPIES — the
* copies dialled through the hops in front of it, which is exactly why they can
* read alive here while the same group's standalone card reads dead. Both
* readings are true; they measure different dial paths. `selected` is the node
* NAME the hop routes through right now, and `delay_ms` / `age_seconds` belong
* to that selected member (or the freshest alive one).
*
* Invariants the daemon guarantees — never re-derive them, just read them:
* `tested === alive + dead` and `alive + dead + untested === total`.
*
* `state` is a closed set of THREE. `untested` is NEVER "dead" and never
* "healthy": it means nothing fresh enough is known. Without `blocked_by` that is
* a matter of timing — for a used chain it resolves on its own within seconds.
* With `blocked_by` it will not resolve until the named hop is fixed. There is no
* fourth state for that; the state stays `untested` because that is what it is.
* `age_seconds: -1` means the age is unknown.
*/
export interface ChainHopHealth {
/** 1-based WIRE order. Hop 1 is dialled first; see the note above. */
index: number
/** Engine outbound tag (`chain-<name>-h2`) — tooltips/debugging, never a label. */
tag: string
/** `node` ⇒ the hop is the measurement. `group` ⇒ the counters roll up members. */
kind: 'node' | 'group'
/** This hop is where traffic leaves for the internet. Always the largest index. */
exit: boolean
/** Closed set — switch on it exhaustively. `untested` is never "dead". */
state: 'alive' | 'dead' | 'untested'
/** RTT of the selected/freshest alive member; 0 (meaningless) when not alive. */
delay_ms: number
/** Age of that measurement in seconds; -1 when unknown. */
age_seconds: number
/** Node name this GROUP hop routes through right now; '' for a node hop. */
selected: string
total: number
tested: number
alive: number
dead: number
untested: number
/**
* PRESENT ONLY on a hop the ordered walk never reached — i.e. a hop sitting
* below one the prober found `dead`. The key is omitted otherwise; absent is
* the normal case and means "this hop was actually dialled".
*
* Its presence is the daemon's own statement that this hop has NO measurement,
* and it comes with the rest of that statement already filled in: `state` is
* `untested`, `delay_ms` is 0, `age_seconds` is -1, and the counters are
* `alive: 0, dead: 0, tested: 0, untested: total`. Read those; do not re-derive
* a verdict from them, and do not infer a block from zeroed counters either —
* an unprobed-yet hop has the same numbers and a very different meaning.
* `selected` MAY still be non-empty: the wrapper does have a pick, it simply
* was not measured, so it says which node the hop would use, not which node is
* carrying traffic.
*/
blocked_by?: ChainHopBlock
}
/**
* The hop that stopped the ordered walk, as reported on every hop below it.
*
* This exists because "no reading" and "no reading, and here is whose fault that
* is" are different answers to the operator's actual question. Without it a
* blocked hop is indistinguishable from one the observatory has not come round to
* yet, and the interface can only shrug.
*
* `index` is the 1-based WIRE index of the blocking hop and is ALWAYS smaller
* than the index of the hop carrying it, so it points at a hop already on screen.
* `tag` is that hop's engine outbound (`chain-<name>-h3`) — debugging and
* tooltips only, never a label to put in front of a person, exactly as on
* {@link ChainHopHealth}.tag.
*/
export interface ChainHopBlock {
/** 1-based wire index of the hop that did not answer. Always < this hop's index. */
index: number
/** That hop's engine outbound tag — tooltips/debugging, never a label. */
tag: string
}
/** Per-chain reachability, the chain analogue of {@link GroupHealth}.used (plan
* §5.E): a chain no enabled routing rule routes through is outside the
* observatory's plan, so its exit is never probed and the Targets card renders it
* "unused" instead of an exit-test readout. A chain has no membership counters —
* it is a fixed path, and its end-to-end health is the exit test's job. */
* observatory's plan, so nothing probes it and the Targets card says so instead
* of rendering a health reading. A chain has no membership counters of its own —
* it is a fixed path, and its health lives on its {@link ChainHopHealth} hops. */
export interface ChainHealth {
name: string
/** An enabled routing rule (the Final target, a DNS-resolver detour, a device
* target, …) reaches this chain, so the observatory probes its exit in the
* target, …) reaches this chain, so the observatory probes its hops in the
* background. false ⇒ nothing routes through the chain: it is skipped by the
* background probing and its end-to-end health stays untested. That is an
* "unused" note about the ROUTING CONFIG, never a health problem. */
* background probing and its health stays untested. That is an "unused" note
* about the ROUTING CONFIG, never a health problem. */
used: boolean
/**
* Per-hop health in wire order (see {@link ChainHopHealth}).
*
* MAY BE ABSENT, and absent does not mean "this chain has no hops". It means
* the engine never materialised per-hop outbounds for it: the chain is unused,
* or it collapses to a single hop and the daemon points traffic straight at
* that target instead of building a copy of it. Read a missing key as "nothing
* measured per hop", never as an empty path or as a fault.
*/
hops?: ChainHopHealth[]
}
export interface GroupsHealth {
@@ -809,9 +973,17 @@ export interface Rule {
Enabled: boolean
Order: number
Src?: string[] | null
DstDomain?: string[] | null
/**
* WHERE the traffic is going — the rule's only destination matcher. Each entry
* names a {@link Ruleset}; the rule matches when ANY of them matches.
*
* There is no inline domain or address list on a rule. `dst_domain`/`dst_ip`
* were removed in schema v2, and `shaterd migrate` folds every existing one
* into a generated `rule-<name>` ruleset, so a destination list is written and
* edited in exactly one place and compiled once into a .srs that every rule
* referencing it shares.
*/
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
/**
* Narrow the rule to one transport or one sniffed application protocol. A
@@ -952,12 +1124,59 @@ export interface Model {
// --- transport --------------------------------------------------------------
/** True when the URL asks for the offline fixture backend (?mock or ?dev). */
export const MOCK: boolean = (() => {
// --- the offline fixture backend (dev builds only) ---------------------------
/**
* True when the in-memory fixtures are serving this session instead of the
* daemon. ALWAYS false in a production build — see {@link initMockBackend}.
*
* A live binding, not a constant: it is decided once during boot, before the
* first render, and every importer sees the same value for the whole session.
*/
export let MOCK = false
/** The loaded fixture module. `null` unless a dev build was asked for `?mock`. */
let fixtures: typeof import('./mock') | null = null
/**
* Load the fixture backend, if this build has one and the URL asks for it.
* Call ONCE from the entry point and await it before the first render — the
* pages read {@link MOCK} while they render, so flipping it afterwards would
* leave a half-mocked screen.
*
* Two gates, and the order matters. `import.meta.env.DEV` is folded to a literal
* `false` by Vite at build time, so in a production build the whole body is
* unreachable, `import('./mock')` is tree-shaken out of the module graph, and the
* fixtures are not in the emitted bundle AT ALL — not lazily, not behind a flag.
* `vite.config.ts` fails the build if that ever stops being true.
*
* This is deliberately stronger than "hide the mock behind a query flag". The
* flag was the bug: `?dev` on a production URL rendered an invented healthy
* router — 119 of 122 nodes alive, "Protected" — with no request made and one
* line of small print in the footer to say so. A person cannot audit a bundle;
* the only honest guarantee is that the invented data is not in it.
*/
export async function initMockBackend(): Promise<boolean> {
if (import.meta.env.DEV && mockRequested()) {
fixtures = await import('./mock')
MOCK = true
}
return MOCK
}
/** Does the URL ask for the offline fixture backend (`?mock` or `?dev`)? */
function mockRequested(): boolean {
if (typeof location === 'undefined') return false
const q = new URLSearchParams(location.search)
return q.has('mock') || q.has('dev')
})()
}
/** The fixture backend, for the `MOCK ? … : …` branches below. Throws rather
* than inventing data if it is ever reached without having been loaded. */
function mock(): NonNullable<typeof fixtures> {
if (!fixtures) throw new Error('mock backend not loaded — call initMockBackend() first')
return fixtures
}
/** A decoded response plus the raw Headers, for endpoints whose contract puts
* pagination metadata outside the JSON body (see the stats log endpoints). */
@@ -1006,33 +1225,54 @@ async function req<T>(path: string, init?: RequestInit): Promise<T> {
// --- endpoints --------------------------------------------------------------
export function getStatus(): Promise<Status> {
return MOCK ? mock.getStatus() : req<Status>('api/status')
return MOCK ? mock().getStatus() : req<Status>('api/status')
}
export function getConfig(): Promise<Model> {
return MOCK ? mock.getConfig() : req<Model>('api/config')
export async function getConfig(): Promise<Model> {
const m = await (MOCK ? mock().getConfig() : req<Model>('api/config'))
// Every page reads the config, and the commit-confirm window's length is the
// only thing needed to arm a countdown — so it is captured here once instead of
// being threaded through eight pages. See pendingConfirm.ts.
noteConfirmTimeout(m.Globals?.ConfirmTimeout)
return m
}
export function putConfig(m: Model): Promise<{ ok: boolean; applied: boolean }> {
return MOCK
? mock.putConfig(m)
? mock().putConfig(m)
: req('api/config', { method: 'PUT', body: JSON.stringify(m) })
}
export function apply(): Promise<ApplyResult> {
return MOCK ? mock.apply() : req<ApplyResult>('api/apply', { method: 'POST' })
/**
* POST /api/apply.
*
* The daemon arms an auto-rollback on EVERY successful apply that changed
* something (panel/api.go handleApply → ArmRollback), whichever page's button was
* pressed. Recording it here — the one place every one of those buttons goes
* through — is what lets the countdown and the "Keep this config" control follow
* the operator around the panel instead of living in the Apply page's local
* state. See pendingConfirm.ts.
*/
export async function apply(): Promise<ApplyResult> {
const r = await (MOCK ? mock().apply() : req<ApplyResult>('api/apply', { method: 'POST' }))
if (!r.error && r.changed) armPendingConfirm()
return r
}
export function confirm(): Promise<ApplyResult> {
return MOCK ? mock.confirm() : req<ApplyResult>('api/confirm', { method: 'POST' })
export async function confirm(): Promise<ApplyResult> {
const r = await (MOCK ? mock().confirm() : req<ApplyResult>('api/confirm', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export function rollback(): Promise<ApplyResult> {
return MOCK ? mock.rollback() : req<ApplyResult>('api/rollback', { method: 'POST' })
export async function rollback(): Promise<ApplyResult> {
const r = await (MOCK ? mock().rollback() : req<ApplyResult>('api/rollback', { method: 'POST' }))
if (!r.error) clearPendingConfirm()
return r
}
export function getStats(): Promise<Stats> {
return MOCK ? mock.getStats() : req<Stats>('api/stats')
return MOCK ? mock().getStats() : req<Stats>('api/stats')
}
// --- daemon log download ------------------------------------------------------
@@ -1084,7 +1324,7 @@ function saveBlob(blob: Blob, filename: string): void {
*/
export async function downloadLog(range: LogRange): Promise<void> {
if (MOCK) {
saveBlob(new Blob([mock.getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
saveBlob(new Blob([mock().getLogText(range)], { type: 'text/plain' }), `shater-log-${range}.txt`)
return
}
let res: Response
@@ -1163,14 +1403,14 @@ function logPage<T>(env: { body: T[] | null; headers: Headers }): StatsLogPage<T
/** GET /api/stats/log — one page of the DNS query log with its cursor metadata. */
export function getStatsLogPage(q: StatsLogQuery = {}): Promise<StatsLogPage<QueryLogEntry>> {
return MOCK
? mock.getStatsLogPage(q)
? mock().getStatsLogPage(q)
: reqFull<QueryLogEntry[] | null>(`api/stats/log${statsLogQS(q)}`).then(logPage)
}
/** GET /api/stats/conns — one page of the connection log with its cursor metadata. */
export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<ConnLogEntry>> {
return MOCK
? mock.getStatsConnsPage(q)
? mock().getStatsConnsPage(q)
: reqFull<ConnLogEntry[] | null>(`api/stats/conns${statsLogQS(q)}`).then(logPage)
}
@@ -1178,13 +1418,13 @@ export function getStatsConnsPage(q: StatsLogQuery = {}): Promise<StatsLogPage<C
* Rows only; callers that tail the stream want {@link getStatsLogPage} instead. */
export function getStatsLog(q: number | StatsLogQuery = {}): Promise<QueryLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
return MOCK ? mock().getStatsLog(o) : req<QueryLogEntry[]>(`api/stats/log${statsLogQS(o)}`)
}
/** GET /api/stats/conns — the live connection-event log (device→dest), newest first. */
export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEntry[]> {
const o: StatsLogQuery = typeof q === 'number' ? { limit: q } : q
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
return MOCK ? mock().getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
@@ -1218,6 +1458,17 @@ export interface RuleReach {
shadowed_by_order?: number
/** Operator-facing sentence; absent when `unreachable` is false. */
reason?: string
/**
* Whether the rule is IN FORCE right now — `Rule.Enabled` after the active WAN
* profile's overrides. This is NOT `GET /api/config`'s `Enabled`: that one is
* the desired state the page PUTs back, and on a router with profiles the two
* legitimately disagree. Draw rows from this; keep the switch on the other.
*/
effective_enabled: boolean
/** The active profile that CHANGED this rule's state; absent when none did. */
overridden_by?: string
/** Which way it went. Absent together with `overridden_by`. */
override?: 'enabled' | 'disabled'
}
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
@@ -1228,12 +1479,12 @@ export interface RulesReachability {
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
return MOCK ? mock().getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
return MOCK ? mock().getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
}
/**
@@ -1244,7 +1495,7 @@ export function getRulesetStatus(): Promise<RulesetStatus[]> {
*/
export function updateRuleset(tag: string): Promise<RulesetStatus | { ok: boolean }> {
return MOCK
? mock.updateRuleset(tag)
? mock().updateRuleset(tag)
: req('api/ruleset/update', { method: 'POST', body: JSON.stringify({ tag }) })
}
@@ -1272,7 +1523,7 @@ export interface RulesetCheck {
*/
export function checkRulesetCategory(source: string, category: string): Promise<RulesetCheck> {
return MOCK
? mock.checkRulesetCategory(source, category)
? mock().checkRulesetCategory(source, category)
: req<RulesetCheck>('api/ruleset/check', {
method: 'POST',
body: JSON.stringify({ source, category }),
@@ -1301,18 +1552,18 @@ export interface RulesetCategories {
*/
export function getRulesetCategories(source: string): Promise<RulesetCategories> {
return MOCK
? mock.getRulesetCategories(source)
? mock().getRulesetCategories(source)
: req<RulesetCategories>(`api/ruleset/categories?source=${encodeURIComponent(source)}`)
}
/** GET /api/devices — discovered LAN clients merged with per-device config. */
export function getDevices(): Promise<DiscoveredDevice[]> {
return MOCK ? mock.getDevices() : req<DiscoveredDevice[]>('api/devices')
return MOCK ? mock().getDevices() : req<DiscoveredDevice[]>('api/devices')
}
/** GET /api/interfaces — the router's UCI network interfaces for the egress picker. */
export function getInterfaces(): Promise<Interface[]> {
return MOCK ? mock.getInterfaces() : req<Interface[]>('api/interfaces')
return MOCK ? mock().getInterfaces() : req<Interface[]>('api/interfaces')
}
/** POST /api/session — exchange a single-use handoff token for a session cookie. */
@@ -1348,7 +1599,7 @@ export function importWg(conf: string): Promise<{ uri: string; name: string }> {
*/
export function updateSubscription(name: string): Promise<{ added: number }> {
return MOCK
? mock.updateSubscription(name)
? mock().updateSubscription(name)
: req('api/subscription/update', { method: 'POST', body: JSON.stringify({ name }) })
}
@@ -1368,7 +1619,7 @@ export function updateSubscription(name: string): Promise<{ added: number }> {
export function getGroupsHealth(
opts: { group?: string; members?: boolean } = {},
): Promise<GroupsHealth> {
if (MOCK) return mock.getGroupsHealth(opts)
if (MOCK) return mock().getGroupsHealth(opts)
const p = new URLSearchParams()
if (opts.group) p.set('group', opts.group)
if (opts.members) p.set('members', '1')
@@ -1377,8 +1628,21 @@ export function getGroupsHealth(
}
/**
* One group's (or chain's) last test: which member the balancer picked, how fast
* it answered, and what the internet saw as the source address.
* What the OBSERVATORY measured for one group or chain — not a dial the panel
* made.
*
* This shape used to come from a fresh connection opened on demand, straight at
* the target. That was a lie on any router whose proxies are blocked when dialled
* directly and work only as a hop behind a tunnel: the card reported dead for a
* path that carries traffic all day. The daemon now has exactly one thing that
* measures — the background observatory, which probes along the REAL dial path,
* per-hop copies and all — and this endpoint reports what it found. There is no
* second measurement anywhere, and the panel never opens a connection of its own.
*
* So read the fields as a READ, not as a test run: `ok` and `delay_ms` are the
* observatory's verdict for the path traffic actually takes, and `tested_unix`
* (router clock, seconds) is when the OBSERVATORY took that measurement — which
* can be a few seconds before the refresh was asked for.
*
* `ok:true` with an EMPTY `exit_ip`/`exit_country` is a valid, successful result,
* not a partial failure: the delay was measured but the exit address could not be
@@ -1387,10 +1651,23 @@ export function getGroupsHealth(
*
* Chains ride the same endpoint. For a chain row, `group` carries the CHAIN's
* name and `selected` the node its last group hop picked ('' when the exit hop
* isn't a group). Everything else reads the same way.
* isn't a group). Per-hop detail is a different read: {@link ChainHopHealth}.
*
* `ok:false` ⇒ the test failed and `error` carries the human reason; every other
* field is meaningless. `tested_unix` is the router's clock, in seconds.
* `ok:false` ⇒ there is no usable measurement and `error` carries the human
* reason; every other field is meaningless. Four of those reasons are about the
* observatory rather than the path, and must not be rendered as "your target is
* broken":
*
* "not routed by any enabled rule, so nothing measures it — the observatory
* only probes paths the rules use"
* "the observatory has not reached this target yet — it refreshes on the
* global probe interval"
* "background probing is disabled, so there is nothing to measure this target
* with"
* "the observatory's probe through this path failed"
*
* Only the last one is a health finding. The first three say the measurement
* does not exist, which is a different thing and a different fix.
*/
export interface GroupTestResult {
group: string // group name — or a chain name for a chain row
@@ -1407,7 +1684,11 @@ export interface GroupTestResult {
* GET /api/groups/test — progress plus every result so far. `results` is ALWAYS
* an array (never null); `done`/`total` count finished vs targeted groups and
* chains while `running` is true. Idle reads `{running:false}` with the last
* run's results still attached, so a reload after a test still shows what it found.
* run's results still attached, so a reload still shows what was last read.
*
* "Running" means the observatory is working through an out-of-turn refresh pass
* over the named targets and this endpoint is collecting what it measures. It is
* not the panel dialling anything.
*/
export interface GroupTestStatus {
running: boolean
@@ -1440,18 +1721,24 @@ export interface GroupTestStart {
}
/**
* POST /api/groups/test — measure a target's delay and exit address. Pass a
* group or chain name to test one; pass nothing (or '') to test every group
* and every chain. Singleton: a second call while a run is in flight resolves
* to `{started:false, reason:'already running'}` rather than failing.
* POST /api/groups/test — ask the observatory for an out-of-turn refresh pass,
* then report what it measured. Pass a group or chain name to refresh one; pass
* nothing (or '') for every group and every chain.
*
* It does NOT dial. The observatory is the only thing in the daemon that
* measures anything, and it measures along the real dial path — so this is the
* "don't wait for the next probe interval" button, not a second opinion. The
* numbers it returns are the same numbers the cards are already showing, just
* fresher. Singleton: a second call while a pass is in flight resolves to
* `{started:false, reason:'already running'}` rather than failing.
*/
export function postGroupsTest(name = ''): Promise<GroupTestStart> {
return MOCK
? mock.postGroupsTest(name)
? mock().postGroupsTest(name)
: req<GroupTestStart>('api/groups/test', { method: 'POST', body: JSON.stringify({ name }) })
}
/** GET /api/groups/test — progress + results of the current/last group test. */
export function getGroupsTest(): Promise<GroupTestStatus> {
return MOCK ? mock.getGroupsTest() : req<GroupTestStatus>('api/groups/test')
return MOCK ? mock().getGroupsTest() : req<GroupTestStatus>('api/groups/test')
}
+15 -1
View File
@@ -1,4 +1,5 @@
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange. */
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange;
* .btn.crit is the solid-red destructive commit. */
.btn {
display: inline-block;
padding: 7px 12px;
@@ -30,3 +31,16 @@
color: #fff;
filter: brightness(1.05);
}
/* Destructive commit. The fill is crit stepped a little toward black so white
* label text clears 4.5:1 in BOTH themes — the raw --crit is bright enough in
* dark mode to fall under it. Red here always means "this removes something". */
.btn.crit {
border-color: transparent;
background: color-mix(in srgb, var(--crit) 88%, #000);
color: #fff;
}
.btn.crit:hover {
color: #fff;
filter: brightness(1.08);
}
+15 -7
View File
@@ -1,19 +1,27 @@
import './Button.css'
import { forwardRef } from 'react'
import type { ButtonHTMLAttributes } from 'react'
export interface ButtonProps extends ButtonHTMLAttributes<HTMLButtonElement> {
/** `primary` is the solid-orange call to action; `ghost` is the default. */
variant?: 'ghost' | 'primary'
/**
* `primary` is the solid-orange call to action; `crit` is the solid-red
* destructive commit (delete, remove) — semantic crit, never the accent;
* `ghost` is the default.
*/
variant?: 'ghost' | 'primary' | 'crit'
}
export function Button({ variant = 'ghost', className, type, ...rest }: ButtonProps) {
/** Ref-forwarding so a dialog can park focus on a specific button. */
export const Button = forwardRef<HTMLButtonElement, ButtonProps>(function Button(
{ variant = 'ghost', className, type, ...rest },
ref,
) {
return (
<button
ref={ref}
type={type ?? 'button'}
className={['btn', variant === 'primary' ? 'primary' : '', className]
.filter(Boolean)
.join(' ')}
className={['btn', variant === 'ghost' ? '' : variant, className].filter(Boolean).join(' ')}
{...rest}
/>
)
}
})
+49 -3
View File
@@ -1,11 +1,49 @@
import { useEffect, useState } from 'react'
/**
* The panel's wall clock, in the SAME timezone as every timestamp under it.
*
* It used to read `getUTCHours()` and print "UTC", while `format.ts` renders every
* log line, connection event and date through `toLocaleTimeString` — i.e. the
* browser's zone. In Moscow that put two clocks three hours apart on one plate,
* and the header was the one nobody could reconcile: the router's "started" time
* read later than the current time while the uptime said it had been up for hours.
*
* So the clock follows the rest of the panel — local, and it SAYS which offset
* that is, because a bare "12:41:07" beside a router in another zone is the
* ambiguity that started this. The zone label is the browser's UTC offset, not an
* abbreviation: "MSK"/"CEST" are not derivable everywhere, an offset always is.
*
* This is the BROWSER's clock, not the router's — the appliance has no RTC. Every
* router-sourced instant in the panel is converted to this clock before it is
* shown, which is what makes one label at the top honest for the whole page.
*/
function zoneLabel(d: Date): string {
// getTimezoneOffset() is minutes WEST of UTC, so the sign is inverted.
const min = -d.getTimezoneOffset()
if (min === 0) return 'UTC'
const sign = min < 0 ? '−' : '+'
const a = Math.abs(min)
const h = Math.floor(a / 60)
const m = a % 60
return `UTC${sign}${h}${m ? `:${String(m).padStart(2, '0')}` : ''}`
}
function format(d: Date): string {
const p = (n: number) => String(n).padStart(2, '0')
return `${p(d.getUTCHours())}:${p(d.getUTCMinutes())}:${p(d.getUTCSeconds())} UTC`
return `${p(d.getHours())}:${p(d.getMinutes())}:${p(d.getSeconds())} ${zoneLabel(d)}`
}
/** Live UTC readout, tabular digits, ticking once a second. */
/** The full zone name, for the title — "Europe/Moscow" says more than "+3" does. */
function zoneName(): string {
try {
return Intl.DateTimeFormat().resolvedOptions().timeZone || ''
} catch {
return ''
}
}
/** Live local readout, tabular digits, ticking once a second. */
export function Clock({ className }: { className?: string }) {
const [now, setNow] = useState(() => format(new Date()))
@@ -14,5 +52,13 @@ export function Clock({ className }: { className?: string }) {
return () => window.clearInterval(id)
}, [])
return <span className={['clock', className].filter(Boolean).join(' ')}>{now}</span>
const zone = zoneName()
return (
<span
className={['clock', className].filter(Boolean).join(' ')}
title={zone ? `Your device's clock — ${zone}. Every time in the panel is shown in this zone.` : undefined}
>
{now}
</span>
)
}
+162
View File
@@ -0,0 +1,162 @@
/* <ConfirmDialog> — the safety interlock plate.
*
* This replaces the browser's native confirm dialog, which a browser can mute for
* good ("prevent this page from creating additional dialogs"): after that it
* returns false with no dialog at all, so every delete button in the panel goes
* dead and silent with no way to recover short of a page reload. We draw the
* plate ourselves, so nothing can suppress it.
*
* Faceplate language: a small rack module lifted off the panel — corner screws
* (reused from Faceplate.css), an engraved label, a groove above the actions.
* Destructive intent is carried by the crit semantic, never by the orange accent:
* accent means "this control is active", crit means "this destroys something".
*/
/* The veil is a fixed dark wash in both themes — a light scrim over a light
* panel would not read as "the panel is out of reach". Follows the tokens.css
* pattern: light base, dark via media query, data-theme overrides win both ways. */
.cfm-scrim {
--cfm-veil: rgba(33, 29, 21, 0.52);
}
@media (prefers-color-scheme: dark) {
.cfm-scrim {
--cfm-veil: rgba(0, 0, 0, 0.66);
}
}
:root[data-theme='light'] .cfm-scrim {
--cfm-veil: rgba(33, 29, 21, 0.52);
}
:root[data-theme='dark'] .cfm-scrim {
--cfm-veil: rgba(0, 0, 0, 0.66);
}
.cfm-scrim {
position: fixed;
inset: 0;
z-index: 200;
display: flex;
align-items: center;
justify-content: center;
/* Short viewports: the plate scrolls with the veil instead of being clipped. */
overflow-y: auto;
padding: calc(var(--u, 8px) * 2);
background: var(--cfm-veil);
animation: cfm-veil-in 0.14s ease-out;
}
.cfm-card {
position: relative;
width: min(32rem, 100%);
max-height: calc(100dvh - var(--u, 8px) * 4);
overflow-y: auto;
padding: calc(var(--u, 8px) * 3.25);
border: 1px solid var(--groove);
border-radius: 12px;
/* same brushed plate as <Faceplate>, one step brighter so it reads as lifted */
background:
repeating-linear-gradient(
90deg,
transparent 0 2px,
color-mix(in srgb, var(--edge) 30%, transparent) 2px 3px
),
linear-gradient(180deg, var(--raised), color-mix(in srgb, var(--raised) 82%, var(--panel)));
box-shadow:
0 1px 0 var(--edge) inset,
0 30px 60px -22px var(--shadow),
0 4px 12px var(--shadow);
animation: cfm-card-in 0.18s cubic-bezier(0.2, 0.7, 0.3, 1);
}
.cfm-card:focus {
outline: none;
}
/* `still` is set from usePrefersReducedMotion — the plate appears, it never
* travels. (The global reduced-motion rule in tokens.css also neutralises the
* duration; this keeps the intent explicit at the component.) */
.cfm-scrim.still,
.cfm-scrim.still .cfm-card {
animation: none;
}
@keyframes cfm-veil-in {
from {
opacity: 0;
}
to {
opacity: 1;
}
}
@keyframes cfm-card-in {
from {
opacity: 0;
transform: translateY(6px) scale(0.99);
}
to {
opacity: 1;
transform: none;
}
}
/* ---- header: engraved label + state LED ---- */
.cfm-hd {
display: flex;
align-items: center;
gap: 10px;
margin-bottom: calc(var(--u, 8px) * 1.5);
}
.cfm-label {
flex: 1;
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: var(--track-label-wide, 0.24em);
color: var(--dim);
text-transform: uppercase;
}
/* ---- copy ---- */
.cfm-title {
margin: 0;
font-family: var(--font-mono);
font-weight: 700;
font-size: 17px;
line-height: 1.35;
color: var(--ink);
/* names can be long and unbroken — wrap rather than push the plate wide */
overflow-wrap: anywhere;
}
.cfm-body {
margin: calc(var(--u, 8px) * 1.5) 0 0;
max-width: 52ch;
font-family: var(--font-sans);
font-size: 13.5px;
line-height: 1.6;
color: var(--dim);
overflow-wrap: anywhere;
}
/* ---- action bar ---- */
.cfm-actions {
display: flex;
justify-content: flex-end;
gap: calc(var(--u, 8px));
margin-top: calc(var(--u, 8px) * 3);
padding-top: calc(var(--u, 8px) * 2);
border-top: 1px solid var(--groove);
}
@media (max-width: 420px) {
.cfm-card {
padding: calc(var(--u, 8px) * 2.5);
}
.cfm-actions {
flex-wrap: wrap;
}
.cfm-actions .btn {
flex: 1 1 auto;
text-align: center;
}
/* screws crowd a small plate — drop them rather than collide with the copy */
.cfm-card > .screw {
display: none;
}
}
+281
View File
@@ -0,0 +1,281 @@
import './ConfirmDialog.css'
import {
createContext,
useCallback,
useContext,
useEffect,
useId,
useRef,
useState,
} from 'react'
import type { ReactNode } from 'react'
import { createPortal } from 'react-dom'
import { Button } from './Button'
import { Led } from './Led'
import { usePrefersReducedMotion } from './usePrefersReducedMotion'
/**
* How the confirming button is painted.
*
* crit — the action destroys something. Semantic crit, never the accent.
* neutral — the action is a normal commit the operator should read first
* (a warning before saving); the accent's call-to-action is correct.
*/
export type ConfirmTone = 'crit' | 'neutral'
export interface ConfirmOptions {
/** Engraved eyebrow, e.g. "DELETE RULE". Names the operation, not the object. */
label?: string
/** The question. One line, ends in "?". */
title: string
/** The consequence — what changes on the router if this goes through. */
body?: ReactNode
/** Verb on the confirming button. Defaults to "Delete". */
confirmLabel?: string
/** Verb on the dismissing button. Defaults to "Cancel". */
cancelLabel?: string
/** Defaults to `crit` — the overwhelmingly common case is a delete. */
tone?: ConfirmTone
}
export interface ConfirmDialogProps extends ConfirmOptions {
open: boolean
/** Called exactly once per dialog, with the operator's answer. */
onResolve: (confirmed: boolean) => void
}
const FOCUSABLE =
'button:not([disabled]), [href], input:not([disabled]), select:not([disabled]), textarea:not([disabled]), [tabindex]:not([tabindex="-1"])'
/**
* The modal plate itself. Normally reached through `useConfirm()`; exported so a
* page that wants to own the open state can render it directly.
*
* Keyboard contract:
* - focus moves to Cancel on open, so a reflex Enter dismisses, never deletes;
* - Tab / Shift+Tab cycle inside the plate and cannot reach the page behind it;
* - Esc answers "no";
* - on close, focus returns to whatever opened the dialog.
*/
export function ConfirmDialog({
open,
onResolve,
label,
title,
body,
confirmLabel = 'Delete',
cancelLabel = 'Cancel',
tone = 'crit',
}: ConfirmDialogProps) {
const titleId = useId()
const bodyId = useId()
const cardRef = useRef<HTMLDivElement>(null)
const cancelRef = useRef<HTMLButtonElement>(null)
const openerRef = useRef<HTMLElement | null>(null)
const reduced = usePrefersReducedMotion()
// Take the page out of the tab order, park focus on Cancel, and hand focus
// back to the opener when the plate goes away.
useEffect(() => {
if (!open) return
const opener = document.activeElement
openerRef.current = opener instanceof HTMLElement ? opener : null
const prevOverflow = document.body.style.overflow
document.body.style.overflow = 'hidden'
// Cancel is the resting place: an Enter or a Space meant for the page lands
// on "no". The destructive button is one Tab away, deliberately.
;(cancelRef.current ?? cardRef.current)?.focus()
return () => {
document.body.style.overflow = prevOverflow
const back = openerRef.current
openerRef.current = null
if (back && document.contains(back)) back.focus()
}
}, [open])
// Esc answers no; Tab is caged. Capture phase so a page-level key handler
// never sees keys aimed at the dialog.
useEffect(() => {
if (!open) return
const onKey = (e: KeyboardEvent) => {
if (e.key === 'Escape') {
e.preventDefault()
e.stopPropagation()
onResolve(false)
return
}
if (e.key !== 'Tab') return
const card = cardRef.current
if (!card) return
const list = Array.from(card.querySelectorAll<HTMLElement>(FOCUSABLE))
if (list.length === 0) {
e.preventDefault()
card.focus()
return
}
const first = list[0]
const last = list[list.length - 1]
const active = document.activeElement as HTMLElement | null
if (!active || !card.contains(active)) {
e.preventDefault()
;(e.shiftKey ? last : first).focus()
} else if (e.shiftKey && active === first) {
e.preventDefault()
last.focus()
} else if (!e.shiftKey && active === last) {
e.preventDefault()
first.focus()
}
}
document.addEventListener('keydown', onKey, true)
return () => document.removeEventListener('keydown', onKey, true)
}, [open, onResolve])
if (!open) return null
return createPortal(
<div
className={['cfm-scrim', reduced ? 'still' : ''].filter(Boolean).join(' ')}
// A click on the field around the plate means "not now". Mousedown (not
// click) so a text selection dragged out of the plate can't dismiss it.
onMouseDown={(e) => {
if (e.target === e.currentTarget) onResolve(false)
}}
>
<div
className={`cfm-card tone-${tone}`}
ref={cardRef}
tabIndex={-1}
role="alertdialog"
aria-modal="true"
aria-labelledby={titleId}
aria-describedby={body != null ? bodyId : undefined}
>
<i className="screw tl" aria-hidden="true" />
<i className="screw tr" aria-hidden="true" />
<i className="screw bl" aria-hidden="true" />
<i className="screw br" aria-hidden="true" />
{/* Lamp first, then the engraved label — the way a real panel reads, and
it keeps the LED off the corner screw. */}
<div className="cfm-hd">
<Led variant={tone === 'crit' ? 'crit' : 'amber'} />
<span className="cfm-label">{label ?? (tone === 'crit' ? 'Confirm delete' : 'Confirm')}</span>
</div>
<h2 className="cfm-title" id={titleId}>
{title}
</h2>
{body != null && (
<p className="cfm-body" id={bodyId}>
{body}
</p>
)}
<div className="cfm-actions">
<Button ref={cancelRef} onClick={() => onResolve(false)}>
{cancelLabel}
</Button>
<Button variant={tone === 'crit' ? 'crit' : 'primary'} onClick={() => onResolve(true)}>
{confirmLabel}
</Button>
</div>
</div>
</div>,
document.body,
)
}
// ---- provider + hook --------------------------------------------------------
interface Request extends ConfirmOptions {
id: number
resolve: (v: boolean) => void
}
const ConfirmCtx = createContext<((o: ConfirmOptions) => Promise<boolean>) | null>(null)
/**
* Mount once at the app root. Everything below can then ask a question and await
* the answer.
*/
export function ConfirmProvider({ children }: { children: ReactNode }) {
const [req, setReq] = useState<Request | null>(null)
const pending = useRef<Request | null>(null)
const seq = useRef(0)
const confirm = useCallback(
(opts: ConfirmOptions) =>
new Promise<boolean>((resolve) => {
// A second question while one is open answers the first with "no" rather
// than leaving its promise — and its caller — hanging forever.
pending.current?.resolve(false)
seq.current += 1
const next: Request = { ...opts, id: seq.current, resolve }
pending.current = next
setReq(next)
}),
[],
)
const settle = useCallback((confirmed: boolean) => {
const open = pending.current
pending.current = null
setReq(null)
open?.resolve(confirmed)
}, [])
// Teardown must not strand a caller mid-await.
useEffect(
() => () => {
pending.current?.resolve(false)
pending.current = null
},
[],
)
// A question belongs to the page that asked it. The provider outlives the
// hash router, so a navigation would otherwise leave a stale plate floating
// over a page it has nothing to do with — answer it "no" and clear it.
useEffect(() => {
const onNav = () => {
if (pending.current) settle(false)
}
window.addEventListener('hashchange', onNav)
return () => window.removeEventListener('hashchange', onNav)
}, [settle])
return (
<ConfirmCtx.Provider value={confirm}>
{children}
{req !== null && <ConfirmDialog key={req.id} open onResolve={settle} {...req} />}
</ConfirmCtx.Provider>
)
}
/**
* Ask the operator, get a definite answer:
*
* const confirm = useConfirm()
* if (!(await confirm({ title: 'Delete rule "x"?', body: '…' }))) return
*
* The returned function is stable, so it is safe in a useCallback dep list. It
* always settles — cancel, Esc, click-outside and teardown all resolve `false`;
* only the confirming button resolves `true`.
*
* Name it `confirm` at the call site on purpose: the local binding shadows the
* global one inside that component, so an accidental bare `confirm(...)` cannot
* reach the suppressible native dialog.
*/
export function useConfirm(): (o: ConfirmOptions) => Promise<boolean> {
const ctx = useContext(ConfirmCtx)
if (!ctx) {
// Loud on purpose. A fallback that quietly resolved false would rebuild the
// exact bug this component exists to kill.
throw new Error('useConfirm() needs <ConfirmProvider> above it (mounted in main.tsx)')
}
return ctx
}
+2
View File
@@ -17,6 +17,8 @@ export { Button } from './Button'
export type { ButtonProps } from './Button'
export { Select } from './Select'
export type { SelectProps, SelectOption } from './Select'
export { ConfirmDialog, ConfirmProvider, useConfirm } from './ConfirmDialog'
export type { ConfirmDialogProps, ConfirmOptions, ConfirmTone } from './ConfirmDialog'
export { Clock } from './Clock'
export { CatSuggest } from './CatSuggest'
export { SrcPicker } from './SrcPicker'
+127
View File
@@ -0,0 +1,127 @@
// findings.ts — which apply-time finding is shown where.
//
// Run with `npm test` (node's built-in test runner + native TypeScript
// stripping; no test dependency is added to the SPA, which ships inside the
// daemon binary).
//
// Two defects are pinned here.
//
// 1. THE TRUNCATION NOTE WAS UNREACHABLE. The daemon caps Status.warnings at 50
// and overwrites the last slot with an `info` note counting what it dropped.
// Overview filtered `info` away wholesale, and the settings-page route keys on
// a section (`generate`) that no page owns — so the single line telling the
// operator "you are not seeing all of it" reached no screen at all.
//
// 2. FINDINGS ABOUT AN ENTITY NEVER REACHED THAT ENTITY'S PAGE. The generator
// drops a node it cannot build and names it; the Nodes page rendered that node
// as an ordinary row with a green toggle, because it never read the findings.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
attentionFindings,
entityFindings,
findingsByName,
sectionNotes,
truncationNote,
worstSeverity,
} from './findings.ts'
import type { StatusWarning } from './api.ts'
const crit = (section: string, name: string, message = 'broken'): StatusWarning => ({
severity: 'critical',
section,
name,
message,
})
const warn = (section: string, name: string, message = 'degraded'): StatusWarning => ({
severity: 'warning',
section,
name,
message,
})
const info = (section: string, name: string, message: string): StatusWarning => ({
severity: 'info',
section,
name,
message,
})
/** Verbatim from apply/warnings.go finalizeWarnings. */
const SUPPRESSED = info(
'generate',
'',
'7 further warning(s) suppressed; run `logread -e shater` for the full list',
)
// --- the truncation note ----------------------------------------------------
test('the truncation note is found, whatever else is in the list', () => {
const note = truncationNote([crit('rule', 'a'), warn('node', 'b'), SUPPRESSED])
assert.notEqual(note, null)
assert.match(note!.message, /7 further warning/)
})
test('a whole list has no truncation note', () => {
assert.equal(truncationNote([crit('rule', 'a'), warn('node', 'b')]), null)
assert.equal(truncationNote([]), null)
assert.equal(truncationNote(undefined), null)
})
test('an ordinary info note is not mistaken for the truncation note', () => {
const notes = [info('untunnelable', 'block', 'Ping and traceroute do not work…')]
assert.equal(truncationNote(notes), null)
})
test('the truncation note is kept out of the settings-page notes it would pollute', () => {
const all = [info('generate', '', 'cache: moved to /overlay'), SUPPRESSED]
const notes = sectionNotes(all, 'generate')
assert.equal(notes.length, 1)
assert.match(notes[0].message, /cache:/)
})
test('the attention list still carries only critical and warning', () => {
const all = [crit('rule', 'a'), warn('node', 'b'), info('untunnelable', 'block', 'x'), SUPPRESSED]
const attention = attentionFindings(all)
assert.equal(attention.length, 2)
assert.ok(attention.every((w) => w.severity !== 'info'))
})
// --- per-entity findings ----------------------------------------------------
test('a page takes only the sections it owns', () => {
const all = [
crit('node', 'tokyo-01', 'parse share-link: bad scheme (skipped)'),
warn('subscription', 'qomar', 'fetch failed'),
crit('rule', 'default', 'never applies'),
info('generate', '', 'cache: x'),
]
const mine = entityFindings(all, ['node', 'subscription'])
assert.deepEqual(
mine.map((w) => w.name),
['tokyo-01', 'qomar'],
)
})
test('entity findings never include info notes', () => {
const all = [info('node', 'tokyo-01', 'just a note'), SUPPRESSED]
assert.equal(entityFindings(all, ['node', 'generate']).length, 0)
})
test('findings index by name, and global (unnamed) ones are left out', () => {
const all = [
crit('node', 'tokyo-01', 'first'),
warn('node', 'tokyo-01', 'second'),
crit('node', '', 'global to the section'),
]
const byName = findingsByName(entityFindings(all, ['node']))
assert.equal(byName.size, 1)
assert.equal(byName.get('tokyo-01')!.length, 2)
})
test('one lamp per row takes the loudest severity', () => {
assert.equal(worstSeverity([warn('node', 'a'), crit('node', 'a')]), 'critical')
assert.equal(worstSeverity([warn('node', 'a')]), 'warning')
assert.equal(worstSeverity([]), null)
})
+91 -2
View File
@@ -6,7 +6,8 @@
//
// critical / warning — something needs attention: a protection promise is
// broken, or something you configured isn't in effect. These belong on
// Overview, where the operator looks first.
// Overview, where the operator looks first — and, when they name an entity,
// ALSO on the page that owns that entity (see `entityFindings`).
//
// info — a statement ABOUT the configuration, not a problem. It never clears,
// because nothing is wrong: it is simply describing a choice that was made.
@@ -16,9 +17,49 @@
// page that never goes away and never asks for anything trains people to skim
// the list — which is exactly how a real critical finding gets missed. Anything
// standing in the findings list should be something you could act on.
//
// The one exception is carved out below: the daemon's own note that it dropped
// findings to fit the cap. It is `info` by severity and unactionable by nature,
// and it is the single most important line in the list, because it is the list
// telling you it is not the whole list.
import type { StatusWarning } from './api'
/**
* The daemon's truncation disclosure, verbatim from apply/warnings.go
* finalizeWarnings:
*
* "%d further warning(s) suppressed; run `logread -e shater` for the full list"
*
* Matched on the stable clause rather than the whole sentence so a reworded tail
* still registers. If this ever stops matching, the failure mode is a list that
* silently claims to be complete — which is why `truncationNote` is tested.
*/
const SUPPRESSED_RE = /further warning\(s\) suppressed/
/**
* The daemon's "this list is incomplete" note, or null when the list is whole.
*
* Status.warnings is capped at 50, sorted critical-first, and the last slot is
* REPLACED by an `info` note counting what was dropped. That note therefore
* arrives on the one channel the panel filtered away wholesale: `info` never
* reached Overview, and the settings-page route (`sectionNotes`) keys on
* section `generate`, which no page owns. So the single line saying "there are
* findings you are not being shown" was the only one guaranteed to be invisible.
*
* Callers must render this WITH the attention list, not instead of it.
*/
export function truncationNote(warnings: StatusWarning[] | undefined): StatusWarning | null {
return (
(warnings ?? []).find((w) => w.severity === 'info' && SUPPRESSED_RE.test(w.message)) ?? null
)
}
/** Is this the truncation disclosure rather than an ordinary note? */
function isTruncationNote(w: StatusWarning): boolean {
return w.severity === 'info' && SUPPRESSED_RE.test(w.message)
}
/** Findings that need attention — the Overview list. Info notes are excluded. */
export function attentionFindings(warnings: StatusWarning[] | undefined): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'critical' || w.severity === 'warning')
@@ -29,10 +70,58 @@ export function attentionFindings(warnings: StatusWarning[] | undefined): Status
* (e.g. `untunnelable` → the Networks page's "Other traffic" section). Only info:
* a critical/warning is an attention item and stays on Overview, so it can't be
* quietly buried on a settings page instead.
*
* The truncation note is excluded: it is about the LIST, not about any section,
* and it has its own home beside the list ({@link truncationNote}).
*/
export function sectionNotes(
warnings: StatusWarning[] | undefined,
section: string,
): StatusWarning[] {
return (warnings ?? []).filter((w) => w.severity === 'info' && w.section === section)
return (warnings ?? []).filter(
(w) => w.severity === 'info' && w.section === section && !isTruncationNote(w),
)
}
/**
* The attention findings about entities ONE page owns — for that page to show
* beside the entities themselves.
*
* Overview is where you look when you already suspect something; a page like
* Nodes is where you look when you don't. The generator drops a node it cannot
* build — an unparseable share link, a WireGuard key materialised twice — and
* says so by name ("node \"x\": parse share-link: … (skipped)"), yet that node
* kept rendering as an ordinary row with a green toggle, because the page never
* read the findings at all. The switch says on; the engine has no such outbound.
*
* This does NOT move anything off Overview: the same finding appears in both
* places, which is correct — one list is "what is wrong with this router", the
* other is "what is wrong with this node".
*/
export function entityFindings(
warnings: StatusWarning[] | undefined,
sections: readonly string[],
): StatusWarning[] {
const want = new Set(sections)
return attentionFindings(warnings).filter((w) => want.has(w.section))
}
/** Index attention findings by entity name, for badging a row directly. Entries
* with an empty `name` are global to their section and are left out. */
export function findingsByName(findings: StatusWarning[]): Map<string, StatusWarning[]> {
const out = new Map<string, StatusWarning[]>()
for (const f of findings) {
if (!f.name) continue
const list = out.get(f.name)
if (list) list.push(f)
else out.set(f.name, [f])
}
return out
}
/** The loudest severity in a set — for a row badge that has room for one lamp. */
export function worstSeverity(findings: StatusWarning[]): 'critical' | 'warning' | null {
if (findings.some((f) => f.severity === 'critical')) return 'critical'
if (findings.length > 0) return 'warning'
return null
}
+33
View File
@@ -77,6 +77,39 @@ export function fmtDateTime(unix: number): string {
return d && t ? `${d}, ${t}` : d || t
}
// --- remote-list freshness ---------------------------------------------------
// A url/geo-sourced list re-fetches on a cadence and the engine reports when it
// last pulled (GET /api/ruleset/status). Routing shows this for rule-sets and DNS
// shows it for blocklists, so the two readings live here and cannot drift apart.
/** "updated 3h ago" / "never updated" for a remote list's last fetch (RFC3339). */
export function relFetch(iso: string): string {
if (!iso) return 'never updated'
const t = Date.parse(iso)
if (Number.isNaN(t)) return 'never updated'
const s = Math.max(0, Math.floor((Date.now() - t) / 1000))
if (s < 45) return 'updated just now'
const m = Math.floor(s / 60)
if (m < 60) return `updated ${m}m ago`
const h = Math.floor(m / 60)
if (h < 24) return `updated ${h}h ago`
const d = Math.floor(h / 24)
return `updated ${d}d ago`
}
/** "every 24h" for an auto-update cadence in seconds ("" when there is none). */
export function everyLabel(sec: number): string {
if (!sec || sec <= 0) return ''
if (sec % 3600 === 0) {
const h = sec / 3600
if (h < 48) return `every ${h}h`
if (sec % 86400 === 0) return `every ${sec / 86400}d`
return `every ${h}h`
}
if (sec % 60 === 0) return `every ${sec / 60}m`
return `every ${sec}s`
}
/**
* A coarse "how long until / since" reading for a unix deadline, relative to now.
*
+26 -5
View File
@@ -2,12 +2,33 @@ import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './tokens.css'
import { App } from './App'
import { initMockBackend } from './api'
import { ConfirmProvider } from './components'
const rootEl = document.getElementById('root')
if (!rootEl) throw new Error('#root not found')
createRoot(rootEl).render(
<StrictMode>
<App />
</StrictMode>,
)
// Settle the fixture question BEFORE the first render: pages read `MOCK` while
// they render, so a backend that arrives afterwards would paint half a screen
// from the daemon and half from fixtures. In a production build this resolves
// immediately and to `false` — the fixtures are not in the bundle to load (see
// api.ts initMockBackend and the assertNoMockFixtures plugin in vite.config.ts).
function mount() {
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl!).render(
<StrictMode>
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
}
// A fixture module that fails to load is a broken dev checkout, not a reason to
// hand the operator a blank plate — mount anyway and let the shell report that it
// cannot reach a daemon, which by then is the truth.
void initMockBackend().then(mount, (e) => {
console.error('mock backend failed to load; continuing against the real API', e)
mount()
})
+281 -40
View File
@@ -6,7 +6,7 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, Traffic } from './api'
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -129,9 +129,25 @@ const CONFIG: Model = {
{ Name: 'via-tunnel', Source: 'subscription', Subscription: 'primary', Strategy: 'leastping', Egress: 'awg' },
{ Name: 'fallback', Source: 'subscription', Subscription: 'backup', Strategy: 'roundrobin', Egress: '' },
],
// One multi-hop chain so `?mock` exercises the chain card's Test button and
// its result readout: enters through the awg tunnel, exits via the auto group.
Chains: [{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] }],
// Three chains, one per state the hop readout has to render.
Chains: [
// The owner's real production shape: leave through a WAN interface, cross an
// AmneziaWG node, then three subscription groups in series. The leading
// `egress:` is NOT a numbered hop — the daemon lifts it into hop 1's entry
// detour — so this reports FOUR hops, and hop 3 is dead while its neighbours
// answer. That single red notch in the middle of a live path is the entire
// reason per-hop health exists, so `?mock` must show it at a glance.
{
Name: 'ewan-wg-subs',
Hops: ['egress:wan', 'node:home-wg', 'group:auto', 'group:stealth', 'group:via-tunnel'],
},
// Used, but the observatory hasn't come round yet — every hop untested. Not
// dead and not healthy: the state the panel most easily renders as a fault.
{ Name: 'sub-fresh', Hops: ['node:home-wg', 'group:fallback'] },
// No enabled rule targets it, so the observatory skips it entirely and the
// daemon never materialises its hops: `used:false` and NO `hops` key.
{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] },
],
Egresses: [
{ Name: 'wan', Type: 'interface', Interface: 'wan' },
// An AmneziaWG tunnel — the whole point of a group-level egress binding.
@@ -145,6 +161,13 @@ const CONFIG: Model = {
Rules: [
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
// These two are what make the chains USED — the observatory probes only the
// paths an enabled rule can reach, so without them every chain card would
// read "not routed" and the hop rail would never appear in `?mock`. Kept
// ABOVE the condition-less rule at Order 40, which would otherwise swallow
// everything below it and mark them "never applies".
{ Name: 'media-via-chain', Enabled: true, Order: 22, DstRuleset: ['yt-geosite'], Target: 'chain:ewan-wg-subs' },
{ Name: 'spare-via-chain', Enabled: true, Order: 24, DstRuleset: ['ad-hosts'], Target: 'chain:sub-fresh' },
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
// A SECOND condition-less rule, above the real default. It reads like a working
// rule and does nothing: a rule with no conditions becomes the router's default,
@@ -171,6 +194,7 @@ const CONFIG: Model = {
// to an official remote list; the others are the usual url / inline lists.
Blocklists: [
{ Name: 'StevenBlack', Enabled: true, Source: 'url', URL: 'https://raw.githubusercontent.com/StevenBlack/hosts/master/hosts', Response: 'nxdomain', UpdateInterval: '24h' },
{ Name: 'oisd-basic', Enabled: true, Source: 'url', URL: 'https://big.oisd.nl/domainswild', Response: 'nxdomain', UpdateInterval: '24h' },
{ Name: 'telegram-block', Enabled: false, Source: 'geosite', Categories: ['telegram'], Response: 'nxdomain', UpdateInterval: '24h' },
],
Resolvers: [
@@ -300,34 +324,98 @@ const RULESET_STATUS: RulesetStatus[] = [
rule_count: 903,
},
{ tag: 'rs-ru-geoip-ru', name: 'ru-geoip', category: 'ru', kind: 'ruleset', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
// Blocklists report through the same endpoint under `bl-<name>`, which the DNS
// page never asked for — so a list that has NEVER been fetched still read
// "filtering". StevenBlack is that case here; oisd-basic is the healthy one, so
// both readings are exercisable offline.
{ tag: 'bl-StevenBlack', name: 'StevenBlack', category: '', kind: 'blocklist', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
{
tag: 'bl-oisd-basic',
name: 'oisd-basic',
category: '',
kind: 'blocklist',
remote: true,
last_updated: new Date(Date.now() - 6 * 3600_000).toISOString(),
interval_seconds: 86_400,
rule_count: 218_431,
},
// Disabled in CONFIG, so the row reads "off" whatever this says — it exists to
// prove the row does not start claiming things the moment a status appears.
{ tag: 'bl-telegram-block-telegram', name: 'telegram-block', category: 'telegram', kind: 'blocklist', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
]
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
* a rule with no conditions is the router's default, and the LAST such rule by
* Order wins — every earlier one can never apply. It reads the live CONFIG so
* edits made in `?mock` keep the badge honest. */
* edits made in `?mock` keep the badge honest.
*
* It also mirrors model.ResolveActiveProfile + ApplyProfileRuleOverrides, because
* `effective_enabled` is the whole point of the endpoint: CONFIG's `mobile-uplink`
* is active and both enables and disables rules, so `?mock` shows the same
* desired-vs-effective split the field config does. */
export async function getRulesReachability(): Promise<RulesReachability> {
await wait(60)
const rules = CONFIG.Rules ?? []
// A pin naming an existing, ENABLED profile wins outright. Otherwise auto-select:
// highest Priority among enabled profiles, ties by Name, skipping any with an
// iface condition (the WAN watcher owns those and expresses its verdict as the pin).
const profiles = CONFIG.Profiles ?? []
const pinned = String(CONFIG.Globals?.ActiveProfile ?? '').trim()
let prof: Profile | null = profiles.find((p) => p.Enabled && p.Name === pinned) ?? null
if (!prof) {
for (const p of profiles) {
if (!p.Enabled || (p.MatchIface ?? []).length > 0) continue
const pp = p.Priority ?? 0
const bp = prof?.Priority ?? 0
if (!prof || pp > bp || (pp === bp && p.Name < prof.Name)) prof = p
}
}
// Enable first, then Disable, so a name in both ends up disabled (Disable wins).
const effective = rules.map((r) => Boolean(r.Enabled))
if (prof) {
const force = (names: string[] | null | undefined, on: boolean) => {
for (const raw of names ?? []) {
const n = raw.trim()
rules.forEach((r, i) => {
if (r.Name === n) effective[i] = on
})
}
}
force(prof.EnableRules, true)
force(prof.DisableRules, false)
}
const activeProfile = prof
const out: RuleReach[] = rules.map((r, index) => ({
index,
name: String(r.Name ?? ''),
order: Number(r.Order ?? 0),
unreachable: false,
shadowed_by_index: -1,
effective_enabled: effective[index],
// Annotate only where the profile actually FLIPPED the outcome — a profile that
// disables an already-off rule has overridden nothing the operator can see.
...(activeProfile && effective[index] !== Boolean(r.Enabled)
? {
overridden_by: activeProfile.Name,
override: effective[index] ? ('enabled' as const) : ('disabled' as const),
}
: {}),
}))
const conditionless = (r: (typeof rules)[number]): boolean =>
!(r.Src ?? []).length &&
!(r.DstDomain ?? []).length &&
!(r.DstRuleset ?? []).length &&
!(r.DstIP ?? []).length &&
!String(r.DstPort ?? '').trim() &&
!String(r.Proto ?? '').trim()
const target = (r: (typeof rules)[number]): string =>
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
const defaults = rules
.map((r, index) => ({ r, index }))
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
// The EFFECTIVE flag, not the configured one: a rule the active profile
// switched off is not in force and cannot retire anything (model's
// RuleReachability runs over the effective set for the same reason).
.filter(({ r, index }) => effective[index] && conditionless(r) && target(r))
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
const winner = defaults[defaults.length - 1]
if (winner) {
@@ -407,7 +495,20 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
// ?mock&warn=1 → a full warning set (critical + warning + info) on top
// ?mock&ks=open → healthy plane but a FAIL-OPEN kill-switch, which is what
// makes the untunnelable policy inert (F8 case 4)
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
// ?mock&traffic=… → with the plane FULL, where the traffic actually ends up:
// split | direct | blocked | blackout | unknown. `direct` is
// the field case the readout used to call "Protected" (one
// rule, `default → direct`); `unknown` is a daemon too old to
// report. Default: tunnel.
// ?mock&plane=unreported → a daemon that sends NO `plane` field. The panel then
// knows nothing about what is installed, which is the state
// the Kill-switch module used to render as a green "ARMED"
// (`undefined !== 'none'` is true).
function mockPlane(): {
plane: 'full' | 'hold' | 'none' | undefined
engine: boolean
killSwitch: string
} {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
const killSwitch = params.get('ks') === 'open' ? 'open' : 'closed'
@@ -415,9 +516,34 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
if (p === 'hold') return { plane: 'hold', engine: false, killSwitch: 'closed' }
if (p === 'none') return { plane: 'none', engine: false, killSwitch }
if (p === 'open') return { plane: 'none', engine: false, killSwitch: 'open' }
if (p === 'unreported') return { plane: undefined, engine: true, killSwitch }
return { plane: 'full', engine: true, killSwitch }
}
// The daemon's verdict on where traffic goes (apply.Status.traffic). Only
// meaningful with the plane installed: with the engine down there is no running
// config to judge, and the daemon reports the unknown/zero value — so do the same
// here rather than leaving a stale "tunnel" behind a dead engine.
function mockTraffic(plane: 'full' | 'hold' | 'none' | undefined): Traffic | undefined {
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
switch (params.get('traffic')) {
case 'split':
return { verdict: 'split', default: 'direct', tunnel_rules: 3 }
case 'direct':
return { verdict: 'direct', default: 'direct', tunnel_rules: 0 }
case 'blocked':
return { verdict: 'blocked', default: 'block', tunnel_rules: 2 }
case 'blackout':
return { verdict: 'blocked', default: 'block', tunnel_rules: 0 }
case 'unknown':
// A daemon that predates the field sends no `traffic` at all.
return undefined
default:
return { verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }
}
}
const MOCK_WARNINGS: StatusWarning[] = [
{
severity: 'critical',
@@ -443,6 +569,29 @@ const MOCK_WARNINGS: StatusWarning[] = [
name: 'fakeip-pool',
message: 'fake-IP resolver cannot be used as a fallback; the failover chain was not built',
},
// Two findings the generator attributes to a NODE by name — the class that the
// Nodes page never showed, leaving a node the engine threw away rendered as an
// ordinary row with a green toggle. Both name real fixture nodes so the row
// badge, the collapsed-bucket "N flagged" count and the per-row strip all fire.
{
severity: 'warning',
section: 'node',
name: 'fi-trojan',
message: 'parse share-link: unsupported scheme "trojan+ws" (skipped)',
},
{
severity: 'warning',
section: 'node',
name: 'home-wg',
message:
'this WireGuard node is materialised twice in the engine config — as "home-wg" and as "group-stealth-m1-home-wg" — and traffic can reach both. A WireGuard peer keeps ONE session per public key, so two devices built from one private key evict each other continuously and NEITHER tunnel passes traffic. Only "home-wg" is kept; everything that routed through "group-stealth-m1-home-wg" is fail-closed (blocked) instead of leaving over the plain WAN',
},
{
severity: 'warning',
section: 'subscription',
name: 'backup',
message: 'fetch failed: dial tcp 203.0.113.9:443: i/o timeout — serving the nodes cached earlier',
},
{
severity: 'info',
section: 'generate',
@@ -451,6 +600,19 @@ const MOCK_WARNINGS: StatusWarning[] = [
},
]
/**
* The daemon's truncation disclosure, exactly as apply/warnings.go writes it when
* the published set overflows the 50-entry cap. Served under `?mock&trunc` so the
* "this list is incomplete" rendering is exercisable — it used to be dropped
* wholesale by the panel's `info` filter and reached no screen at all.
*/
const MOCK_TRUNCATION: StatusWarning = {
severity: 'info',
section: 'generate',
name: '',
message: '7 further warning(s) suppressed; run `logread -e shater` for the full list',
}
/**
* The standing `untunnelable` note the daemon reports. It is INFO, never a
* problem: it states a correct, chosen configuration. Two shapes, mirroring the
@@ -497,9 +659,12 @@ function mockWarnings(killSwitch: string): StatusWarning[] {
const params = new URLSearchParams(q)
const mode = (CONFIG.Globals as { Untunnelable?: string }).Untunnelable ?? 'block'
const notes = untunnelableNote(mode, killSwitch)
// `?trunc` adds the daemon's "the published list is capped" disclosure, which
// it appends IN PLACE OF the last entry it had room for.
const trunc = params.has('trunc') ? [{ ...MOCK_TRUNCATION }] : []
// A degraded plane always comes with the findings that explain it.
if (params.has('warn') || params.get('plane')) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes]
if (params.has('warn') || params.get('plane') || trunc.length > 0) {
return [...MOCK_WARNINGS.map((w) => ({ ...w })), ...notes, ...trunc]
}
return notes
}
@@ -520,6 +685,7 @@ export async function getStatus(): Promise<Status> {
can_rollback: armed || hasLastGood,
engine_running: engine,
plane,
traffic: mockTraffic(plane),
warnings: mockWarnings(killSwitch),
// Process uptime. Anchored to when this tab loaded plus a fixed head start, so
// the reading ticks forward across polls exactly like the real daemon's does.
@@ -1096,20 +1262,59 @@ function healthList(): GroupHealth[] {
return (CONFIG.Groups ?? []).map((g) => summarise(g.Name, GROUP_MEMBERS.get(g.Name) ?? []))
}
/** Per-chain reachability for the Targets page's "unused" badge (plan §5.E) — the
* chain analogue of healthList's `used` field. The mock's single chain `relay` is
* NOT referenced by any rule in CONFIG.Rules (they target group:auto / block /
* direct), so it reads used=false and its card renders "unused" — exactly the case
* the badge exists to surface. A stopped engine reports no chains. */
/**
* Per-hop health, keyed by chain name — what the observatory measured at each
* position of the path, in WIRE order.
*
* `ewan-wg-subs` is the fixture that matters, and it encodes the ORDERED WALK.
* Hop 1 is the WireGuard node and answers; hop 2 is a subscription group whose
* copies answer THROUGH it — 119 of 122 tested alive, which is the reading only a
* per-hop probe can produce, since the same members are dialled differently on
* their own card. Hop 3 is a group whose members all time out at that position,
* and the walk STOPS there: hop 4 is dialled through hop 3, so it was never
* dialled at all. It comes back `untested` with `blocked_by` naming hop 3, its
* counters zeroed, and `selected` still set — the wrapper has a pick, nothing
* crossed it to measure. A dead hop with a live hop under it is not in this
* fixture because the daemon can no longer produce one.
*
* `sub-fresh` is used but never yet reached: every hop untested, nothing dead,
* no block — the other reason a lamp is unlit, and the one that fixes itself.
* `relay` is absent from this map on purpose — an unused chain is never
* materialised, so the daemon sends no `hops` key at all, which is "nothing
* measured", not "no hops".
*/
const CHAIN_HOPS: Record<string, ChainHopHealth[]> = {
'ewan-wg-subs': [
{ index: 1, tag: 'chain-ewan-wg-subs-h1', kind: 'node', exit: false, state: 'alive', delay_ms: 41, age_seconds: 22, selected: '', total: 1, tested: 1, alive: 1, dead: 0, untested: 0 },
{ index: 2, tag: 'chain-ewan-wg-subs-h2', kind: 'group', exit: false, state: 'alive', delay_ms: 96, age_seconds: 18, selected: '🇳🇱 Amsterdam-01', total: 298, tested: 122, alive: 119, dead: 3, untested: 176 },
{ index: 3, tag: 'chain-ewan-wg-subs-h3', kind: 'group', exit: false, state: 'dead', delay_ms: 0, age_seconds: 15, selected: '', total: 2, tested: 2, alive: 0, dead: 2, untested: 0 },
{ index: 4, tag: 'chain-ewan-wg-subs-h4', kind: 'group', exit: true, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '🇸🇬 Singapore-09', total: 6, tested: 0, alive: 0, dead: 0, untested: 6, blocked_by: { index: 3, tag: 'chain-ewan-wg-subs-h3' } },
],
'sub-fresh': [
{ index: 1, tag: 'chain-sub-fresh-h1', kind: 'node', exit: false, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 1, tested: 0, alive: 0, dead: 0, untested: 1 },
{ index: 2, tag: 'chain-sub-fresh-h2', kind: 'group', exit: true, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 24, tested: 0, alive: 0, dead: 0, untested: 24 },
],
}
/** Per-chain reachability plus per-hop health for the Targets page. `used` is the
* chain analogue of healthList's field; `hops` is OMITTED (never null, never []),
* exactly like the daemon, for a chain the engine never materialised. A stopped
* engine reports no chains at all. */
function chainHealthList(): ChainHealth[] {
if (!mockPlane().engine) return []
return (CONFIG.Chains ?? []).map((c) => ({ name: c.Name, used: chainUsed(c.Name) }))
return (CONFIG.Chains ?? []).map((c) => {
const hops = CHAIN_HOPS[c.Name]
const h: ChainHealth = { name: c.Name, used: chainUsed(c.Name) }
if (hops) h.hops = hops.map((x) => ({ ...x }))
return h
})
}
/** A chain is "used" when some enabled routing rule (or Final, or a DNS detour)
* targets `chain:<name>` — the same reachability the daemon's observatory derives.
* The mock's rules never target a chain, so every chain reads used=false; a real
* config would mark the ones rules point at used=true. */
* Two of the mock's rules do (`media-via-chain` → ewan-wg-subs, `spare-via-chain`
* → sub-fresh), so those two chains read used=true and get a hop rail; `relay`
* is targeted by nothing and reads used=false, which is the unused note. */
function chainUsed(name: string): boolean {
const target = `chain:${name}`
return (CONFIG.Rules ?? []).some(
@@ -1161,15 +1366,23 @@ class ApiErrorLike extends Error {
}
}
// Mock group/chain test. Deliberately covers every state the UI has to render,
// one per target, so a single offline run exercises all of them:
// auto → ok WITH an exit address
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
// SUCCESS, and the case the UI most easily gets wrong
// relay → the chain: same wire shape, `group` carries the CHAIN's name and
// `selected` the node its exit group picked
// fallback → a failure carrying a human reason
// Results land one per GET poll, so the running/progress state is visible too.
// Mock refresh results. The endpoint no longer dials anything: it asks the
// observatory to measure out of turn and reports what the observatory found, so
// every row here is a READ of a background measurement. Deliberately covers every
// state the UI has to render, one per target, so a single offline run exercises
// all of them:
// auto → ok WITH an exit address
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
// SUCCESS, and the case the UI most easily gets wrong
// ewan-wg-subs → the chain: same wire shape, `group` carries the CHAIN's name
// and `selected` the node its exit hop picked
// via-tunnel → the one honest health FAILURE: a probe that ran and failed
// fallback,
// relay → not routed at all, so no measurement exists to report
// sub-fresh → routed, but the observatory hasn't come round yet
// The last three are absence of measurement, not a broken target, and the copy
// has to keep them apart. Results land one per GET poll, so the running/progress
// state is visible too.
const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_unix'>> = {
auto: {
selected: 'nl-reality-2',
@@ -1187,32 +1400,60 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
ok: true,
error: '',
},
// The chain — Selected is the node the chain's exit group (auto) picked.
relay: {
selected: 'nl-reality-2',
delay_ms: 61,
exit_ip: '185.12.34.56',
exit_country: 'NL',
ok: true,
error: '',
// The chain, and the pairing that makes the whole feature worth building. A
// chain is one series path, so with hop 3 dead the end-to-end probe is never
// even attempted — the daemon stops walking there. This row and the hop rail
// therefore have to tell one story, not two: both name hop 3, and neither
// offers hop 4 as a second suspect. Note the row does NOT say "the probe
// failed" — no probe of this chain's exit ran at all — which is why the daemon
// has a separate message for it.
'ewan-wg-subs': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'hop 3 of this chain was probed and did not answer, so nothing reaches the exit through it — fix that hop first',
},
// Dead through its tunnel, exactly as its membership health says — the exit
// test and the member health tell the same story about the same group.
// The one real health failure in the fixture: the observatory's probe ran along
// this path and did not come back.
'via-tunnel': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'no member answered through egress awg (6 of 6 timed out)',
error: 'the observatory’s probe through this path failed',
},
// Not a health verdict — nothing routes here, so no measurement of it exists.
fallback: {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'no reachable node in the group (all 3 members timed out)',
error:
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
},
relay: {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
},
// Routed, materialised, simply not reached yet. Untested is not dead.
'sub-fresh': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'the observatory has not reached this target yet — it refreshes on the global probe interval',
},
}
+443
View File
@@ -0,0 +1,443 @@
/* Alerts section (rendered on Settings) — inherits the Faceplate tokens and the
* shared page chrome from App.css (.toast, .mono). Every rule below is a
* one-to-one copy of the DNS.css rule the markup used before the section moved
* here, renamed `dns-*` → `alr-*` so nothing collides. Orange stays an accent. */
/* ---- section shell (matches the Settings group plates one-to-one) ---- */
.alr-section {
margin-top: calc(var(--u, 8px) * 3.5);
}
.alr-sec-hd {
display: flex;
align-items: baseline;
gap: 12px;
padding-bottom: 10px;
border-bottom: 1px solid var(--groove);
}
.alr-sec-title {
margin: 0;
font-family: var(--font-mono);
font-size: 13px;
font-weight: 700;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--dim);
}
.alr-sec-count {
font-size: 11px;
letter-spacing: 0.06em;
color: var(--faint);
}
.alr-sec-note {
margin: 10px 2px 0;
font-family: var(--font-sans);
font-size: 12.5px;
line-height: 1.55;
color: var(--dim);
max-width: 56ch;
}
/* ---- add form ---- */
.alr-add {
display: flex;
flex-direction: column;
gap: 10px;
margin-top: calc(var(--u, 8px) * 2);
}
.alr-add-top {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px;
}
.alr-input {
min-width: 0;
padding: 9px 12px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 12.5px;
letter-spacing: 0.02em;
box-shadow: 0 1px 2px var(--shadow) inset;
transition: border-color 0.15s, box-shadow 0.15s;
}
.alr-input::placeholder {
color: var(--faint);
}
.alr-input:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-input:disabled {
opacity: 0.55;
}
.alr-input--name {
flex: 0 1 14rem;
}
/* segmented type picker */
.alr-seg {
display: inline-flex;
border: 1px solid var(--groove);
border-radius: 7px;
overflow: hidden;
background: var(--sink);
}
.alr-seg-btn {
padding: 8px 14px;
border: 0;
background: transparent;
color: var(--dim);
font-family: var(--font-mono);
font-size: 11px;
letter-spacing: 0.08em;
text-transform: uppercase;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.alr-seg-btn + .alr-seg-btn {
border-left: 1px solid var(--groove);
}
.alr-seg-btn.on {
background: var(--accent);
color: #fff;
}
.alr-seg-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
.alr-resp {
display: inline-flex;
align-items: center;
gap: 8px;
}
.alr-resp-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-select {
padding: 8px 10px;
border: 1px solid var(--groove);
border-radius: 7px;
background: var(--sink);
color: var(--ink);
font-family: var(--font-mono);
font-size: 11.5px;
letter-spacing: 0.04em;
cursor: pointer;
}
.alr-select:focus-visible {
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.alr-add-actions {
display: flex;
align-items: center;
justify-content: flex-end;
gap: 14px;
flex-wrap: wrap;
}
.alr-field-err {
flex: 1;
min-width: 0;
margin: 0;
font-family: var(--font-mono);
font-size: 11.5px;
line-height: 1.5;
color: var(--crit);
}
/* ---- rows ---- */
.alr-rows {
list-style: none;
margin: calc(var(--u, 8px) * 2) 0 0;
padding: 0;
display: flex;
flex-direction: column;
gap: 8px;
}
.alr-row {
display: flex;
align-items: center;
gap: calc(var(--u, 8px) * 1.5);
padding: 12px 14px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(
180deg,
var(--raised),
color-mix(in srgb, var(--raised) 82%, var(--panel))
);
box-shadow: 0 1px 0 var(--edge) inset;
}
.alr-row-main {
flex: 1;
min-width: 0;
display: flex;
flex-direction: column;
gap: 4px;
}
.alr-row-l1 {
display: flex;
align-items: center;
gap: 8px;
flex-wrap: wrap;
}
.alr-row-name {
font-family: var(--font-mono);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
color: var(--ink);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 24ch;
}
.alr-row-l2 {
display: flex;
align-items: center;
gap: 10px;
flex-wrap: wrap;
font-size: 11.5px;
letter-spacing: 0.02em;
}
.alr-row-detail {
color: var(--dim);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 40ch;
}
/* badge — groove-bordered, not orange (accent stays reserved) */
.alr-badge {
display: inline-block;
padding: 2px 7px;
border: 1px solid var(--groove);
border-radius: 5px;
background: color-mix(in srgb, var(--sink) 60%, transparent);
font-family: var(--font-mono);
font-size: 10px;
font-weight: 600;
letter-spacing: 0.1em;
text-transform: uppercase;
color: var(--dim);
white-space: nowrap;
}
.alr-badge--accent {
border-color: color-mix(in srgb, var(--accent) 55%, var(--groove));
color: var(--accent);
}
.alr-masked {
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: 0.08em;
color: var(--faint);
text-transform: uppercase;
cursor: help;
}
.alr-del {
flex: none;
padding: 6px 12px;
font-size: 10.5px;
}
/* ---- empty plate ---- */
.alr-empty {
margin-top: calc(var(--u, 8px) * 2);
padding: calc(var(--u, 8px) * 3);
border: 1px dashed var(--groove);
border-radius: 9px;
background: color-mix(in srgb, var(--raised) 55%, transparent);
text-align: center;
}
.alr-empty-title {
display: block;
font-size: 13px;
font-weight: 700;
letter-spacing: 0.06em;
color: var(--dim);
}
.alr-empty-body {
margin: 8px auto 0;
max-width: 48ch;
font-family: var(--font-sans);
font-size: 13px;
line-height: 1.55;
color: var(--dim);
}
/* ---- loading skeleton ---- */
.alr-skel {
height: 62px;
border: 1px solid var(--groove);
border-radius: 8px;
background: linear-gradient(90deg, var(--raised), var(--sink), var(--raised));
background-size: 200% 100%;
animation: alr-skel-shift 1.4s ease-in-out infinite;
}
@keyframes alr-skel-shift {
from {
background-position: 200% 0;
}
to {
background-position: -200% 0;
}
}
/* the per-row delivery picker sits inline in the row */
.alr-detour {
flex: none;
display: flex;
flex-direction: column;
gap: 5px;
min-width: 0;
}
.alr-detour-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
.alr-detour-select {
max-width: 22rem;
}
/* current delivery-path readout on the row */
.alr-path {
color: var(--faint);
white-space: nowrap;
}
.alr-path[data-active='on'] {
color: var(--dim);
}
.alr-path-name {
color: var(--led-on);
font-weight: 600;
}
.alr-path[data-missing='y'] .alr-path-name {
color: var(--amber);
}
.alr-path-flag {
color: var(--amber);
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.alr-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.alr-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.alr-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-row delivery controls sit inline in the row (like .alr-detour) */
.alr-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.alr-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.alr-note--row {
margin-top: 2px;
}
/* alert event checkboxes */
.alr-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.alr-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.alr-event input {
accent-color: var(--fp-accent, currentColor);
}
/* ---- responsive ---- */
@media (max-width: 640px) {
.alr-row {
flex-wrap: wrap;
}
.alr-row-main {
flex-basis: calc(100% - 90px);
}
.alr-del {
margin-left: auto;
}
.alr-input--name {
flex-basis: 100%;
}
.alr-detour {
flex-basis: 100%;
order: 3;
flex-wrap: wrap;
}
/* A <select> won't shrink below its widest option unless it's allowed to:
without min-width:0 the long detour labels push the page into a horizontal
scroll at 390px. Let them fill the row and clip instead. */
.alr-detour-select,
.alr-resp .alr-select {
max-width: 100%;
width: 100%;
min-width: 0;
}
.alr-resp {
display: flex;
flex-wrap: wrap;
max-width: 100%;
}
.alr-ctl {
flex-basis: 100%;
order: 3;
}
}
@media (prefers-reduced-motion: reduce) {
.alr-skel {
animation: none;
}
.alr-input,
.alr-seg-btn {
transition: none;
}
}
+690
View File
@@ -0,0 +1,690 @@
import './Alerts.css'
import { useCallback, useMemo, useState } from 'react'
import { Button, Toggle, useConfirm } from '../components'
import type { Alert, Model } from '../api'
// The Alerts section — out-of-band notifications (Telegram bot / webhook) for
// kill-switch trips, apply failures, new devices and subscription expiry. It
// lived at the bottom of the DNS page, which is the last place an operator
// looking for "tell me when the tunnel dies" would think to look; it now renders
// as a group on Settings. The component owns no I/O: every mutation goes through
// the `onSave` prop so Settings keeps a single dirty banner and a single toast.
//
// NOTE on duplication: the detour helpers below (DetourCatalog, canonDetour,
// detourValues, describeDetour, DetourSelect) plus asArray / uniqueName /
// maskUrl / EmptyPlate are deliberate copies of the ones in DNS.tsx. DNS keeps
// its own for resolvers and DNS rules; extracting a shared module would couple
// two pages that otherwise share nothing, and that refactor is out of scope
// here. If a third consumer ever appears, promote them then.
// ---- local Model extension --------------------------------------------------
/** The Model with the Alerts slice surfaced (index-signature passthrough). */
type AlertsModel = Model & { Alerts?: Alert[] | null }
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers (copies of DNS.tsx — see the header note) ----------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
const HTTP_RE = /^https?:\/\//i
/** A remote URL often carries a token in its query/path — show host only. */
function maskUrl(url: string): { host: string; masked: boolean } {
try {
const u = new URL(url)
return { host: u.host, masked: u.search !== '' || u.pathname.replace(/\/+$/, '') !== '' }
} catch {
return { host: url || '—', masked: false }
}
}
function uniqueName(base: string, taken: Set<string>): string {
const seed = base.trim() || 'alert'
if (!taken.has(seed)) return seed
let i = 2
while (taken.has(`${seed}-${i}`)) i++
return `${seed}-${i}`
}
/** The live targets an alert's delivery can be pinned to (the picker). */
interface DetourCatalog {
groups: string[]
chains: string[]
egresses: { name: string; type: string }[]
nodes: string[]
}
/**
* Normalise a stored `Via` to a picker option value. Empty/`direct` ⇒
* `direct`; already-prefixed values (`group:`/`chain:`/`egress:`/`node:`) pass
* through; a bare legacy name is resolved against the catalog so a still-valid
* setup isn't mislabelled; anything unresolved is kept verbatim (shown stale).
*/
function canonDetour(raw: string | undefined, cat: DetourCatalog): string {
const d = (raw ?? '').trim()
if (!d || d.toLowerCase() === 'direct') return 'direct'
if (/^(node|group|chain|egress):/i.test(d)) return d
if (cat.egresses.some((e) => e.name === d)) return `egress:${d}`
if (cat.groups.includes(d)) return `group:${d}`
if (cat.chains.includes(d)) return `chain:${d}`
if (cat.nodes.includes(d)) return `node:${d}`
return d
}
/** Every valid option value for a catalog, including `direct`. */
function detourValues(cat: DetourCatalog): Set<string> {
const s = new Set<string>(['direct'])
for (const g of cat.groups) s.add(`group:${g}`)
for (const c of cat.chains) s.add(`chain:${c}`)
for (const e of cat.egresses) s.add(`egress:${e.name}`)
for (const n of cat.nodes) s.add(`node:${n}`)
return s
}
/** Describe a canonical detour value for the row readout. */
function describeDetour(
canon: string,
cat: DetourCatalog,
valid: Set<string>,
): { direct: boolean; prefix: string; name: string; missing: boolean } {
if (canon === 'direct') return { direct: true, prefix: '', name: '', missing: false }
const i = canon.indexOf(':')
const kind = i === -1 ? '' : canon.slice(0, i)
const name = i === -1 ? canon : canon.slice(i + 1)
const missing = !valid.has(canon)
let prefix = 'via'
if (kind === 'group') prefix = 'via group'
else if (kind === 'chain') prefix = 'via chain'
else if (kind === 'node') prefix = 'via node'
else if (kind === 'egress') {
const eg = cat.egresses.find((e) => e.name === name)
prefix = eg?.type === 'interface' ? 'via interface' : 'via egress'
}
return { direct: false, prefix, name, missing }
}
// ---- section ----------------------------------------------------------------
export function AlertsSection({
config,
busy,
loading,
onSave,
}: {
/** Full desired-state model; null until it has loaded. */
config: Model | null
/** A save/apply is in flight — controls lock. */
busy: boolean
/** The config is still loading — show a skeleton row. */
loading: boolean
/** Persist the whole next model; resolves true on success (Settings' `save`). */
onSave: (next: Model, okMsg: string) => Promise<boolean>
}): JSX.Element {
const confirm = useConfirm()
const model = config as AlertsModel | null
const alerts = useMemo<Alert[]>(() => asArray(model?.Alerts), [model])
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Built straight from the Model with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({
groups: asArray(config?.Groups).map((g) => g.Name),
chains: [],
egresses: asArray(config?.Egresses).map((e) => ({ name: e.Name, type: e.Type })),
nodes: asArray(config?.Nodes).map((n) => n.Name),
}),
[config],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
// ---- mutations — all writes go through onSave -----------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!model) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return onSave({ ...model, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[model, alerts, onSave],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void onSave({ ...model, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[model, alerts, onSave],
)
const removeAlert = useCallback(
async (idx: number) => {
if (!model) return
const target = alerts[idx]
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void onSave({ ...model, Alerts: next }, `Deleted ${target.Name}`)
},
[model, alerts, onSave, confirm],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!model) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void onSave(
{ ...model, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[model, alerts, onSave],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!model) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void onSave(
{ ...model, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[model, alerts, onSave],
)
return (
<div className="alr-section" aria-label="Alerts">
<header className="alr-sec-hd">
<h2 className="alr-sec-title">Alerts</h2>
<span className="alr-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="alr-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="alr-rows" aria-hidden="true">
<li className="alr-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="alr-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
)
}
// ---- alert add form + row ----------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="alr-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="alr-add-top">
<input
className="alr-input alr-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="alr-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'alr-seg-btn on' : 'alr-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="alr-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="alr-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="alr-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="alr-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="alr-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="alr-delivery">
<label className="alr-resp">
<span className="alr-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="alr-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="alr-add-actions">
{err && (
<p className="alr-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="alr-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="alr-row-main">
<div className="alr-row-l1">
<span className="alr-row-name">{alert.Name}</span>
<span className="alr-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="alr-badge alr-badge--accent">
{e}
</span>
))}
</div>
<div className="alr-row-l2 mono">
<span className="alr-row-detail">{detail.text}</span>
{detail.masked && (
<span className="alr-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="alr-path">direct</span>
) : (
<span className="alr-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="alr-path-name">{route.name}</strong>
{route.missing && <span className="alr-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="alr-note alr-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="alr-ctl">
<label className="alr-detour">
<span className="alr-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="alr-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="alr-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="alr-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
/** The live delivery picker: option list built from the Model's targets. */
function DetourSelect({
value,
catalog,
valid,
busy,
disabled,
ariaLabel,
onChange,
directLabel = 'Direct (no proxy)',
}: {
value: string // canonical value
catalog: DetourCatalog
valid: Set<string>
busy: boolean
disabled: boolean
ariaLabel: string
onChange: (v: string) => void
directLabel?: string
}) {
const missing = value !== 'direct' && !valid.has(value)
return (
<select
className="alr-select alr-detour-select"
value={value}
onChange={(e) => onChange(e.target.value)}
disabled={busy || disabled}
aria-label={ariaLabel}
>
<option value="direct">{directLabel}</option>
{catalog.groups.length > 0 && (
<optgroup label="Groups">
{catalog.groups.map((g) => (
<option key={g} value={`group:${g}`}>
Group {g} (balancer)
</option>
))}
</optgroup>
)}
{catalog.chains.length > 0 && (
<optgroup label="Chains">
{catalog.chains.map((c) => (
<option key={c} value={`chain:${c}`}>
Chain {c}
</option>
))}
</optgroup>
)}
{catalog.egresses.length > 0 && (
<optgroup label="Interfaces / egresses">
{catalog.egresses.map((e) => (
<option key={e.name} value={`egress:${e.name}`}>
Interface/egress {e.name}
{e.type ? ` (${e.type})` : ''}
</option>
))}
</optgroup>
)}
{catalog.nodes.length > 0 && (
<optgroup label="Nodes">
{catalog.nodes.map((n) => (
<option key={n} value={`node:${n}`}>
Node {n}
</option>
))}
</optgroup>
)}
{missing && <option value={value}>{value} (missing)</option>}
</select>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="alr-empty">
<span className="alr-empty-title mono">{title}</span>
<p className="alr-empty-body">{body}</p>
</div>
)
}
+108 -53
View File
@@ -11,6 +11,8 @@ import {
ApiError,
} from '../api'
import type { Globals, Status } from '../api'
import { engineReadout, killSwitchReadout } from '../planeState'
import { onPendingConfirmExpire, usePendingConfirm } from '../pendingConfirm'
// Short, readable config hash — drops the "sha256:" prefix like the footer does.
function short(hash: string): string {
@@ -25,13 +27,6 @@ function msg(e: unknown): string {
type Busy = 'apply' | 'confirm' | 'rollback' | null
/** A pending commit-confirm window: the daemon has armed an auto-rollback. */
interface Armed {
total: number // the ConfirmTimeout the window started with
remaining: number // seconds left before the daemon reverts
appliedHash: string // the hash that went live on apply (the "after" of apply)
}
type ActionKind = 'apply' | 'confirm' | 'rollback' | 'expire'
interface ActionResult {
kind: ActionKind
@@ -57,7 +52,11 @@ export default function Apply() {
const [configError, setConfigError] = useState<string | null>(null)
const [busy, setBusy] = useState<Busy>(null)
const [armed, setArmed] = useState<Armed | null>(null)
// The armed window is app-wide state, not this page's: it is recorded by the
// api layer on every apply and survives a reload. Keeping it local is what made
// refreshing this tab lose both the countdown and the only button that could
// stop it. See pendingConfirm.ts.
const armed = usePendingConfirm()
const [result, setResult] = useState<ActionResult | null>(null)
const [confirmingRollback, setConfirmingRollback] = useState(false)
@@ -104,32 +103,64 @@ export default function Apply() {
void loadConfig()
}, [loadConfig])
// ---- commit-confirm countdown: a calm 1s numeric tick, effect-scoped so the
// timer is always cleared on unmount / confirm / rollback (no leaked intervals) ----
// The window running out does NOT mean the daemon rolled back.
//
// apply.ArmRollback captures the data-plane generation when it arms, and on
// expiry it compares. If anything re-applied the plane in between — another
// panel apply, SIGHUP, a hotplug or the once-a-minute cron reconcile, the WAN
// profile auto-switch — it disarms and KEEPS the running config, logging "NOT
// rolling back" and nothing else. That is the common case on a production
// router, and this page used to print "daemon auto-rolled back to last-good
// config" for it: a confident report of an event that did not happen, with a
// hash pair underneath that quietly said "unchanged".
//
// The panel cannot see which branch ran — the daemon says so only in its log.
// So it reports the one thing it CAN observe, the live config hash, and waits
// for the revert to land before reading it (a rollback is a full re-apply and
// does not complete the instant the timer fires).
const liveHashRef = useRef('')
liveHashRef.current = status?.hash ?? ''
useEffect(() => {
if (!armed) return
if (armed.remaining <= 0) {
// Window elapsed — the daemon reverts to last-good on its own. Observe it.
const before = armed.appliedHash
setArmed(null)
flash('Auto-rolled back')
let cancelled = false
const off = onPendingConfirmExpire(() => {
const before = liveHashRef.current
flash('Confirm window elapsed')
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed. Reading what the daemon did…',
before,
after: before,
})
void (async () => {
const after = (await refreshStatus())?.hash ?? ''
let after = before
for (let i = 0; i < 4 && !cancelled; i++) {
await new Promise((r) => window.setTimeout(r, 1500))
if (cancelled) return
after = (await refreshStatus())?.hash ?? after
if (after !== before) break
}
if (cancelled) return
setResult({
kind: 'expire',
tone: 'warn',
text: 'Confirm window elapsed — daemon auto-rolled back to last-good config.',
text:
after !== before
? 'Confirm window elapsed and the live config changed — the daemon reverted to its last-good config.'
: 'Confirm window elapsed and the live config has not changed, so this config is still running. ' +
'The daemon only reverts if nothing else re-applied the data plane while the window was open; ' +
'otherwise it stands down and keeps what is live. Which one happened is in the daemon log — ' +
'download it from Settings, or run `logread -e shater`.',
before,
after,
})
})()
return
})
return () => {
cancelled = true
off()
}
const id = window.setTimeout(() => {
setArmed((a) => (a ? { ...a, remaining: a.remaining - 1 } : a))
}, 1000)
return () => window.clearTimeout(id)
}, [armed, flash, refreshStatus])
}, [flash, refreshStatus])
const confirmWindow = globals?.ConfirmTimeout ?? 0
@@ -146,8 +177,9 @@ export default function Apply() {
return
}
const after = (await refreshStatus())?.hash ?? before
// The window itself was recorded by api.apply(); this branch only writes the
// readout for it.
if (r.changed && confirmWindow > 0) {
setArmed({ total: confirmWindow, remaining: confirmWindow, appliedHash: after })
setResult({
kind: 'apply',
tone: 'good',
@@ -179,7 +211,8 @@ export default function Apply() {
const doConfirm = useCallback(async () => {
const before = status?.hash ?? ''
setBusy('confirm')
setArmed(null) // stop the countdown immediately; confirm cancels the auto-rollback
// api.confirm() clears the shared window on success — the countdown stops the
// moment the daemon agrees, not the moment we asked.
try {
const r = await apiConfirm()
if (r.error) {
@@ -208,7 +241,7 @@ export default function Apply() {
const before = status?.hash ?? ''
setConfirmingRollback(false)
setBusy('rollback')
setArmed(null) // rolling back also cancels any pending confirm window
// api.rollback() clears the shared window on success (rolling back ends it).
try {
const r = await apiRollback()
if (r.error) {
@@ -237,19 +270,37 @@ export default function Apply() {
}, [status, flash, refreshStatus])
// ---- derived display state (mirrors Overview's LED semantics) ----
const killArmed = globals ? globals.KillSwitch === 'closed' : false
const engineVariant: LedVariant = !status
? 'off'
: status.running && status.active
? 'on'
: status.running
? 'amber'
: 'crit'
const dataVariant: LedVariant = status?.table ? 'on' : status?.running ? 'amber' : 'off'
//
// The LIVE kill-switch wins over the saved one, exactly as on Overview: this row
// is a status readout, and the config on disk can already differ from what is
// installed. Falls back to the config only while /api/status is unread.
const killArmed = (status?.kill_switch ?? globals?.KillSwitch ?? 'closed') === 'closed'
// Whether that setting is actually installed — same three-plus-unknown reading
// as Overview, so the two pages cannot disagree about the same router.
const kill = killSwitchReadout(status, globals?.KillSwitch)
const killWord =
kill.state === 'open'
? 'open'
: kill.state === 'armed'
? 'fail-closed'
: kill.state === 'inert'
? 'closed · not in effect'
: 'closed · not reported'
// Every engine mark on this page comes from ONE reading, and that reading is
// able to say "stopped" — see planeState.engineState for why `status.running`
// could not. This page is where someone lands when the network is down; three
// green lamps here were the difference between "I broke it" and "nothing broke".
const engine = engineReadout(status)
const engineVariant: LedVariant = engine.variant
// No nft table means there is no data plane at all. Under a fail-closed switch
// that is a leak (crit); under an open one it is the documented choice (amber).
// It used to go amber whenever `running` was true — i.e. always — and unlit
// otherwise, so the one state worth shouting about had no colour of its own.
const dataVariant: LedVariant = status?.table ? 'on' : !status ? 'off' : killArmed ? 'crit' : 'amber'
const configVariant: LedVariant = status?.enabled ? 'on' : 'amber'
const liveHash = short(status?.hash ?? '')
const pct = armed ? Math.max(0, Math.round((armed.remaining / armed.total) * 100)) : 0
const pct = armed ? Math.max(0, Math.round((armed.remaining / armed.pending.total) * 100)) : 0
// Only offer rollback when the daemon says one would revert something: an armed
// commit-confirm snapshot, or an engine last-good predecessor. When false there
@@ -264,9 +315,7 @@ export default function Apply() {
label="Engine"
variant={engineVariant}
pulse={engineVariant === 'on'}
value={
!status ? 'checking…' : status.running ? (status.active ? 'active' : 'idle') : 'stopped'
}
value={engine.word}
/>
<StatusPip
label="Config"
@@ -278,11 +327,7 @@ export default function Apply() {
variant={dataVariant}
value={status?.table ? 'nft installed' : 'no table'}
/>
<StatusPip
label="Kill-switch"
variant={killArmed ? 'on' : 'amber'}
value={killArmed ? 'fail-closed' : 'open'}
/>
<StatusPip label="Kill-switch" variant={kill.variant} value={killWord} />
</div>
{statusError && (
@@ -310,9 +355,13 @@ export default function Apply() {
unit="· sha256"
led={{ variant: configVariant }}
rows={[
{ k: 'engine', v: status?.running ? 'running' : 'stopped', hot: !status?.running },
{ k: 'data plane', v: status?.table ? 'nft installed' : 'no table' },
{ k: 'kill-switch', v: killArmed ? 'fail-closed' : 'open', hot: !killArmed },
{ k: 'engine', v: engine.word, hot: engineVariant === 'crit' },
{
k: 'data plane',
v: status?.table ? 'nft installed' : 'no table',
hot: dataVariant === 'crit',
},
{ k: 'kill-switch', v: killWord, hot: kill.variant === 'crit' || !killArmed },
]}
/>
<Module
@@ -325,7 +374,7 @@ export default function Apply() {
}
led={{ variant: engineVariant }}
rows={[
{ k: 'state', v: !status ? 'checking…' : status.active ? 'active' : 'idle' },
{ k: 'state', v: engine.word, hot: engineVariant === 'crit' },
{ k: 'config', v: status?.enabled ? 'enabled' : 'disabled' },
{ k: 'schema', v: globals ? `v${globals.SchemaVersion}` : '—' },
]}
@@ -362,9 +411,13 @@ export default function Apply() {
</div>
<div className="cc-info">
<p className="cc-copy">
Applied config <span className="mono">{short(armed.appliedHash)}</span> is live but
not yet kept. Confirm to keep it — otherwise the daemon rolls back to the last-good
config when the timer hits zero.
{/* The live hash IS the applied one while a window is open — that
is what "live but not kept" means — so the readout survives a
reload instead of depending on what this tab remembers. */}
Applied config <span className="mono">{liveHash}</span> is live but not yet kept.
Confirm to keep it. At zero the daemon rolls back to the last-good config — unless
something else re-applies the data plane first, in which case it stands down and
keeps whatever is live.
</p>
<div className="cc-bar" aria-hidden="true">
<span className="cc-bar-fill" style={{ width: `${pct}%` }} />
@@ -481,7 +534,9 @@ function labelFor(kind: ActionKind): string {
case 'rollback':
return 'Rollback'
case 'expire':
return 'Auto-rollback'
// NOT "Auto-rollback": on expiry the daemon either reverts or stands down,
// and this page cannot tell which. Name the event it did observe.
return 'Window elapsed'
}
}
+80 -68
View File
@@ -387,6 +387,79 @@
.dns-row-state[data-active='on'] {
color: var(--led-on);
}
/* A list that is switched on but has nothing loaded is not "off" and is certainly
not "filtering" — warn semantics, the same amber the badges use. */
.dns-row-state[data-active='warn'] {
color: var(--amber);
}
/* ---- remote-list freshness (mirrors the rule-set rows on Routing) ---- */
.dns-row-sync {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 10px;
margin-top: 3px;
}
.dns-sync-fresh {
font-family: var(--font-mono);
font-size: 11px;
color: var(--dim);
}
.dns-sync-fresh[data-never='y'] {
color: var(--amber);
}
.dns-sync-every,
.dns-sync-rules {
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.02em;
color: var(--faint);
}
.dns-sync-every::before {
content: '↻ ';
}
.dns-sync-update {
display: inline-flex;
align-items: center;
gap: 6px;
padding: 3px 10px;
border: 1px solid var(--accent-soft);
border-radius: 5px;
background: var(--raised);
color: var(--accent);
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: var(--track-label);
text-transform: uppercase;
cursor: pointer;
transition: color 0.12s, border-color 0.12s, background 0.12s;
}
.dns-sync-update:hover:not(:disabled) {
border-color: var(--accent);
background: color-mix(in srgb, var(--accent) 12%, transparent);
}
.dns-sync-update:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 2px;
}
.dns-sync-update:disabled {
opacity: 0.6;
cursor: not-allowed;
}
.dns-sync-spin {
width: 10px;
height: 10px;
border: 2px solid color-mix(in srgb, var(--accent) 35%, transparent);
border-top-color: var(--accent);
border-radius: 50%;
animation: dns-sync-spin 0.7s linear infinite;
}
@keyframes dns-sync-spin {
to {
transform: rotate(360deg);
}
}
/* badge — groove-bordered, not orange (accent stays reserved) */
.dns-badge {
@@ -575,80 +648,19 @@
}
}
/* alert delivery: deliver-via picker + fallback toggle + caution note */
.dns-alert-delivery {
display: flex;
flex-wrap: wrap;
align-items: center;
gap: 10px 20px;
}
.dns-fallback {
display: inline-flex;
align-items: center;
gap: 8px;
cursor: pointer;
}
.dns-fallback-label {
font-size: 10px;
letter-spacing: var(--track-label, 0.18em);
text-transform: uppercase;
color: var(--faint);
}
/* the per-alert-row delivery controls sit inline in the row (like .dns-detour) */
.dns-alert-ctl {
flex: none;
display: flex;
flex-direction: column;
gap: 8px;
min-width: 0;
}
.dns-alert-note {
margin: 0;
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.5;
color: var(--amber);
max-width: 56ch;
}
.dns-alert-note--row {
margin-top: 2px;
}
@media (max-width: 640px) {
.dns-alert-ctl {
flex-basis: 100%;
order: 3;
}
}
/* alert event checkboxes */
.dns-events {
display: flex;
flex-wrap: wrap;
gap: 8px 16px;
margin: 0;
padding: 0;
border: 0;
}
.dns-event {
display: inline-flex;
align-items: center;
gap: 6px;
font-size: 13px;
color: var(--fp-text, inherit);
cursor: pointer;
}
.dns-event input {
accent-color: var(--fp-accent, currentColor);
}
@media (prefers-reduced-motion: reduce) {
.dns-skel {
animation: none;
}
/* No spin under reduced motion — the static ring + "Updating…" label carry it. */
.dns-sync-spin {
animation: none;
border-top-color: color-mix(in srgb, var(--accent) 35%, transparent);
}
.dns-chip,
.dns-input,
.dns-seg-btn {
.dns-seg-btn,
.dns-sync-update {
transition: none;
}
}
+237 -499
View File
@@ -1,8 +1,16 @@
import './DNS.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, CatSuggest, Led, SrcPicker, Toggle } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Alert, DNSRule, Model, Resolver } from '../api'
import { Button, CatSuggest, Led, SrcPicker, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
getRulesetStatus,
putConfig,
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { DNSRule, Model, Resolver, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
// The DNS / Blocklists page is a thin editor over the desired-state Model —
// exactly like Nodes.tsx. Every edit rewrites the relevant slice in-place, PUTs
@@ -52,28 +60,9 @@ type GlobalsX = Model['Globals'] & { DNSFilter?: boolean }
type DNSModel = Model & {
Blocklists?: Blocklist[] | null
Allowlists?: Allowlist[] | null
Alerts?: Alert[] | null
DNSRules?: DNSRule[] | null
}
/**
* Every event the daemon actually sends. A retired health-probe event was left
* out on purpose: nothing ever fired it, so a channel that subscribed to it would
* just stay quiet forever — the one failure mode an alert must not have. Only
* events with a live firing path are offered here.
*/
const ALERT_EVENTS: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'killswitch', label: 'Kill-switch' },
{ id: 'apply_fail', label: 'Apply failure' },
{ id: 'new_device', label: 'New device' },
{ id: 'sub_expiry', label: 'Subscription expiring' },
]
// Shown when an alert routes through a detour with no direct fallback — the exact
// case where a tunnel-down alert could fail to send. The user asked for this.
const VIA_NO_FALLBACK_NOTE =
'A kill-switch/tunnel-down alert may not send if it routes through the affected tunnel — enable fallback.'
// ---- helpers ---------------------------------------------------------------
const asArray = <T,>(a: T[] | null | undefined): T[] => (a ? a : [])
@@ -220,6 +209,7 @@ function describeDetour(
// ---- page ------------------------------------------------------------------
export default function DNS() {
const confirm = useConfirm()
const [config, setConfig] = useState<DNSModel | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -301,7 +291,6 @@ export default function DNS() {
const blocklists = useMemo(() => asArray(config?.Blocklists), [config])
const allowlists = useMemo(() => asArray(config?.Allowlists), [config])
const resolvers = useMemo<Resolver[]>(() => asArray(config?.Resolvers), [config])
const alerts = useMemo<Alert[]>(() => asArray(config?.Alerts), [config])
// Ascending Order — the engine evaluates DNS rules first-match, so the list is
// shown and edited in the order it actually runs.
const dnsRules = useMemo<DNSRule[]>(
@@ -309,6 +298,68 @@ export default function DNS() {
[config],
)
// ---- did the lists actually LOAD? -----------------------------------------
//
// A blocklist row said "filtering" whenever the list and the master switch were
// both on. Neither of those is evidence that anything is being blocked: a
// url/geosite list is fetched by the engine, the daemon treats a failed fetch as
// a CRITICAL apply finding, and the row went on saying "filtering" through it.
// The Routing page had already been given this reading for rule-sets — the same
// endpoint, the same tags (`bl-<name>` / `al-<name>`) — and the DNS page never
// asked. Slow poll: lists refresh on a ~24h cadence, so 15s only has to catch a
// manual Update-now. Grouped by NAME because a geosite list with N categories
// reports N records.
const [listStatus, setListStatus] = useState<Map<string, RulesetStatus[]>>(new Map())
const [updatingLists, setUpdatingLists] = useState<Set<string>>(new Set())
const loadListStatus = useCallback(async () => {
try {
const all = await getRulesetStatus()
const m = new Map<string, RulesetStatus[]>()
for (const s of all) {
if (s.kind !== 'blocklist' && s.kind !== 'allowlist') continue
const key = `${s.kind}:${s.name}`
const arr = m.get(key)
if (arr) arr.push(s)
else m.set(key, [s])
}
setListStatus(m)
} catch {
// Engine stopped or an older daemon — keep the last reading. The row falls
// back to "load not reported", which claims nothing either way.
}
}, [])
useEffect(() => {
void loadListStatus()
const id = window.setInterval(() => void loadListStatus(), 15000)
return () => window.clearInterval(id)
}, [loadListStatus])
const updateList = useCallback(
async (kind: 'blocklist' | 'allowlist', name: string) => {
const key = `${kind}:${name}`
setUpdatingLists((prev) => new Set(prev).add(key))
try {
// One geo list can hold several categories, each its own engine tag.
const recs = listStatus.get(key) ?? []
const tags = recs.length
? recs.map((r) => r.tag)
: [`${kind === 'blocklist' ? 'bl' : 'al'}-${name}`]
for (const tag of tags) await apiUpdateRuleset(tag)
await loadListStatus()
flash(`${name} refreshed`)
} catch (e) {
flash(`Refresh failed — ${errText(e)}`)
} finally {
setUpdatingLists((prev) => {
const next = new Set(prev)
next.delete(key)
return next
})
}
},
[listStatus, loadListStatus, flash],
)
const blOn = blocklists.filter((b) => b.Enabled).length
const alOn = allowlists.filter((a) => a.Enabled).length
const blNames = useMemo(() => new Set(blocklists.map((b) => b.Name)), [blocklists])
@@ -422,15 +473,19 @@ export default function DNS() {
)
const removeBlocklist = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = blocklists[idx]
if (!window.confirm(`Delete blocklist “${target.Name}”? This removes it from the config.`))
return
const ok = await confirm({
label: 'Delete blocklist',
title: `Delete blocklist “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = blocklists.filter((_, i) => i !== idx)
void save({ ...config, Blocklists: next }, `Deleted ${target.Name}`)
},
[config, blocklists, save],
[config, blocklists, save, confirm],
)
// ---- allowlist mutations --------------------------------------------------
@@ -457,15 +512,19 @@ export default function DNS() {
)
const removeAllowlist = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = allowlists[idx]
if (!window.confirm(`Delete allowlist “${target.Name}”? This removes it from the config.`))
return
const ok = await confirm({
label: 'Delete allowlist',
title: `Delete allowlist “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = allowlists.filter((_, i) => i !== idx)
void save({ ...config, Allowlists: next }, `Deleted ${target.Name}`)
},
[config, allowlists, save],
[config, allowlists, save, confirm],
)
// ---- resolver mutations ---------------------------------------------------
@@ -491,15 +550,58 @@ export default function DNS() {
[config, resolvers, save],
)
/**
* Delete a resolver, saying what it was still wired into.
*
* The three GLOBAL slots (default, fallback, endpoint) are cleared here, because
* a global pointing at nothing is never what anyone meant. The DNS RULES are a
* different matter: each one is a decision about which queries go where, and
* silently deleting or repointing them would change where a device's DNS goes
* without saying so. So they are named instead and left alone — the dialog is
* where the operator finds out they exist, which is precisely what this page
* used to skip: it cleared the two globals without a word and never mentioned
* the rules at all.
*/
const removeResolver = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = resolvers[idx]
if (!window.confirm(`Delete resolver “${target.Name}”? This removes it from the config.`))
return
const next = resolvers.filter((_, i) => i !== idx)
// Don't leave default/fallback pointing at a resolver that no longer exists.
const g = { ...config.Globals }
const slots: string[] = []
if (g.ResolverDefault === target.Name) slots.push('the default resolver')
if (g.ResolverFallback === target.Name) slots.push('the fallback resolver')
if (g.EndpointResolver === target.Name) slots.push('the endpoint resolver')
const usedBy = dnsRules.filter((r) => r.Resolver === target.Name)
const parts: string[] = []
if (slots.length > 0) {
parts.push(
`It is ${slots.join(' and ')} — ${
slots.length === 1 ? 'that slot is' : 'those slots are'
} cleared, so DNS falls back to the engine's built-in resolution.`,
)
}
if (usedBy.length === 1) {
parts.push(
`One DNS rule still sends queries to it (order ${usedBy[0].Order}). It is left as it is and will have nowhere to resolve — repoint it before you apply.`,
)
} else if (usedBy.length > 1) {
parts.push(
`${usedBy.length} DNS rules still send queries to it (orders ${usedBy
.map((r) => r.Order)
.join(', ')}). They are left as they are and will have nowhere to resolve — repoint them before you apply.`,
)
}
if (parts.length === 0) parts.push('Nothing else in the config points at it.')
const ok = await confirm({
label: 'Delete resolver',
title: `Delete resolver “${target.Name}”?`,
body: parts.join(' '),
})
if (!ok) return
const next = resolvers.filter((_, i) => i !== idx)
// Don't leave default/fallback/endpoint pointing at a resolver that's gone.
const cleared: string[] = []
if (g.ResolverDefault === target.Name) {
g.ResolverDefault = ''
@@ -509,12 +611,16 @@ export default function DNS() {
g.ResolverFallback = ''
cleared.push('fallback')
}
if (g.EndpointResolver === target.Name) {
g.EndpointResolver = ''
cleared.push('endpoint')
}
const msg = cleared.length
? `Deleted ${target.Name} — cleared ${cleared.join(' & ')}`
: `Deleted ${target.Name}`
void save({ ...config, Globals: g, Resolvers: next }, msg)
},
[config, resolvers, save],
[config, resolvers, dnsRules, save, confirm],
)
const setResolverDefault = useCallback(
@@ -578,84 +684,21 @@ export default function DNS() {
)
const removeDNSRule = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = dnsRules[idx]
if (!window.confirm(`Delete this DNS rule? Matching queries fall back to the default resolver.`))
return
const ok = await confirm({
label: 'Delete DNS rule',
title: 'Delete this DNS rule?',
body: 'Matching queries fall back to the default resolver.',
})
if (!ok) return
const next = dnsRules.filter((_, i) => i !== idx)
void save({ ...config, DNSRules: next }, `Deleted DNS rule → ${target.Resolver}`)
},
[config, dnsRules, save],
[config, dnsRules, save, confirm],
)
// ---- alert mutations ------------------------------------------------------
const addAlert = useCallback(
(draft: Alert): Promise<boolean> => {
if (!config) return Promise.resolve(false)
const taken = new Set(alerts.map((a) => a.Name))
const a: Alert = { ...draft, Name: uniqueName(draft.Name, taken) }
return save({ ...config, Alerts: [...alerts, a] }, `Added ${a.Name}`)
},
[config, alerts, save],
)
const toggleAlert = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Enabled: on } : a))
void save({ ...config, Alerts: next }, `${next[idx].Name} ${on ? 'enabled' : 'disabled'}`)
},
[config, alerts, save],
)
const removeAlert = useCallback(
(idx: number) => {
if (!config) return
const target = alerts[idx]
if (!window.confirm(`Delete alert “${target.Name}”? This removes it from the config.`)) return
const next = alerts.filter((_, i) => i !== idx)
void save({ ...config, Alerts: next }, `Deleted ${target.Name}`)
},
[config, alerts, save],
)
const setAlertVia = useCallback(
(idx: number, v: string) => {
if (!config) return
const via = v === 'direct' ? '' : v
const next = alerts.map((a, i) => (i === idx ? { ...a, Via: via || undefined } : a))
void save(
{ ...config, Alerts: next },
via ? `${next[idx].Name} delivers via ${via}` : `${next[idx].Name} delivers direct`,
)
},
[config, alerts, save],
)
const setAlertFallback = useCallback(
(idx: number, on: boolean) => {
if (!config) return
const next = alerts.map((a, i) => (i === idx ? { ...a, Fallback: on || undefined } : a))
void save(
{ ...config, Alerts: next },
`${next[idx].Name} direct fallback ${on ? 'on' : 'off'}`,
)
},
[config, alerts, save],
)
// Alerts route through Direct/group/node/egress only (no chains) — the contract
// vocabulary for Alert.Via. Reuse the resolver detour catalog with chains dropped.
const alertCatalog = useMemo<DetourCatalog>(
() => ({ ...detourCatalog, chains: [] }),
[detourCatalog],
)
const alertValid = useMemo(() => detourValues(alertCatalog), [alertCatalog])
const alertNames = useMemo(() => new Set(alerts.map((a) => a.Name)), [alerts])
const alertsOn = alerts.filter((a) => a.Enabled).length
const loading = config === null && loadError === null
return (
@@ -883,8 +926,11 @@ export default function DNS() {
categories={b.Categories}
response={b.Response}
filterOn={dnsFilterOn}
statuses={listStatus.get(`blocklist:${b.Name}`) ?? null}
updating={updatingLists.has(`blocklist:${b.Name}`)}
busy={busy}
onToggle={(on) => toggleBlocklist(i, on)}
onUpdateNow={() => void updateList('blocklist', b.Name)}
onDelete={() => removeBlocklist(i)}
/>
))}
@@ -942,9 +988,13 @@ export default function DNS() {
url={a.URL}
path={a.Path}
entries={a.Entries}
categories={a.Categories}
filterOn={dnsFilterOn}
statuses={listStatus.get(`allowlist:${a.Name}`) ?? null}
updating={updatingLists.has(`allowlist:${a.Name}`)}
busy={busy}
onToggle={(on) => toggleAllowlist(i, on)}
onUpdateNow={() => void updateList('allowlist', a.Name)}
onDelete={() => removeAllowlist(i)}
/>
))}
@@ -1069,57 +1119,6 @@ export default function DNS() {
)}
</div>
{/* ---- 5. ALERTS ---- */}
<div className="dns-section" aria-label="Alerts">
<header className="dns-sec-hd">
<h2 className="dns-sec-title">Alerts</h2>
<span className="dns-sec-count mono">
{alertsOn} / {alerts.length} on
</span>
</header>
<p className="dns-sec-note">
Out-of-band notifications. Delivered <strong>direct to the internet</strong> by default — so a
kill-switch or engine-down alert still reaches you when the proxy is down. You can route one
through a group, node or egress instead, with a direct fallback if that detour fails.
</p>
<AddAlertForm
busy={busy}
disabled={!config}
taken={alertNames}
catalog={alertCatalog}
valid={alertValid}
onAdd={addAlert}
/>
{loading ? (
<ul className="dns-rows" aria-hidden="true">
<li className="dns-skel" />
</ul>
) : alerts.length === 0 ? (
<EmptyPlate
title="No alerts"
body="Add a Telegram bot or a webhook above to get notified when the kill-switch trips, a new device joins, or an apply fails."
/>
) : (
<ul className="dns-rows">
{alerts.map((a, i) => (
<AlertRow
key={`${a.Name}-${i}`}
alert={a}
busy={busy}
catalog={alertCatalog}
valid={alertValid}
onToggle={(on) => toggleAlert(i, on)}
onVia={(v) => setAlertVia(i, v)}
onFallback={(on) => setAlertFallback(i, on)}
onDelete={() => removeAlert(i)}
/>
))}
</ul>
)}
</div>
{toast && (
<div className="toast" role="status">
{toast}
@@ -1129,342 +1128,6 @@ export default function DNS() {
)
}
// ---- alert add form + row --------------------------------------------------
function AddAlertForm({
busy,
disabled,
taken,
catalog,
valid,
onAdd,
}: {
busy: boolean
disabled: boolean
taken: Set<string>
catalog: DetourCatalog
valid: Set<string>
onAdd: (a: Alert) => Promise<boolean>
}) {
const [name, setName] = useState('')
const [type, setType] = useState<'telegram' | 'webhook'>('telegram')
const [token, setToken] = useState('')
const [chatId, setChatId] = useState('')
const [url, setUrl] = useState('')
const [events, setEvents] = useState<string[]>(['killswitch'])
const [via, setVia] = useState('direct')
const [fallback, setFallback] = useState(false)
const [err, setErr] = useState<string | null>(null)
const reset = () => {
setName('')
setType('telegram')
setToken('')
setChatId('')
setUrl('')
setEvents(['killswitch'])
setVia('direct')
setFallback(false)
}
const toggleEvent = (id: string) =>
setEvents((prev) => (prev.includes(id) ? prev.filter((e) => e !== id) : [...prev, id]))
const submit = async () => {
const nm = name.trim()
if (!nm) {
setErr('Give the alert a name.')
return
}
if (taken.has(nm)) {
setErr(`An alert named “${nm}” already exists.`)
return
}
if (type === 'telegram') {
if (!token.trim() || !chatId.trim()) {
setErr('Telegram needs a bot token and a chat ID.')
return
}
} else if (!HTTP_RE.test(url.trim())) {
setErr('Enter an http(s):// webhook URL.')
return
}
if (events.length === 0) {
setErr('Pick at least one event to notify on.')
return
}
setErr(null)
const routed = via !== 'direct'
const routing = { Via: routed ? via : undefined, Fallback: routed && fallback ? true : undefined }
const draft: Alert =
type === 'telegram'
? { Name: nm, Enabled: true, Type: 'telegram', Token: token.trim(), ChatID: chatId.trim(), Events: events, ...routing }
: { Name: nm, Enabled: true, Type: 'webhook', URL: url.trim(), Events: events, ...routing }
const ok = await onAdd(draft)
if (ok) reset()
}
const routed = via !== 'direct'
return (
<form
className="dns-add"
onSubmit={(e) => {
e.preventDefault()
void submit()
}}
>
<div className="dns-add-top">
<input
className="dns-input dns-input--name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Alert name"
aria-label="Alert name"
value={name}
onChange={(e) => {
setName(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<div className="dns-seg" role="group" aria-label="Alert type">
<button
type="button"
className={type === 'telegram' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'telegram'}
onClick={() => setType('telegram')}
disabled={busy || disabled}
>
Telegram
</button>
<button
type="button"
className={type === 'webhook' ? 'dns-seg-btn on' : 'dns-seg-btn'}
aria-pressed={type === 'webhook'}
onClick={() => setType('webhook')}
disabled={busy || disabled}
>
Webhook
</button>
</div>
</div>
{type === 'telegram' ? (
<>
<input
className="dns-input"
type="password"
spellCheck={false}
autoComplete="off"
placeholder="Bot token (kept secret)"
aria-label="Telegram bot token"
value={token}
onChange={(e) => {
setToken(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
<input
className="dns-input"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Chat ID (e.g. -1001234567890)"
aria-label="Telegram chat ID"
value={chatId}
onChange={(e) => {
setChatId(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
</>
) : (
<input
className="dns-input"
type="text"
inputMode="url"
spellCheck={false}
autoComplete="off"
placeholder="https://hooks.example.com/…"
aria-label="Webhook URL"
value={url}
onChange={(e) => {
setUrl(e.target.value)
if (err) setErr(null)
}}
disabled={busy || disabled}
/>
)}
<fieldset className="dns-events" aria-label="Events to notify on">
{ALERT_EVENTS.map((ev) => (
<label key={ev.id} className="dns-event">
<input
type="checkbox"
checked={events.includes(ev.id)}
onChange={() => toggleEvent(ev.id)}
disabled={busy || disabled}
/>
<span>{ev.label}</span>
</label>
))}
</fieldset>
<div className="dns-alert-delivery">
<label className="dns-resp">
<span className="dns-resp-label mono">Deliver via</span>
<DetourSelect
value={via}
catalog={catalog}
valid={valid}
busy={busy}
disabled={disabled}
ariaLabel="Deliver alert via"
onChange={setVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={setFallback}
label={fallback ? 'Disable direct fallback' : 'Enable direct fallback'}
disabled={busy || disabled || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
{routed && !fallback && (
<p className="dns-alert-note" role="note">
{VIA_NO_FALLBACK_NOTE}
</p>
)}
<div className="dns-add-actions">
{err && (
<p className="dns-field-err" role="alert">
{err}
</p>
)}
<Button type="submit" variant="primary" disabled={busy || disabled}>
{busy ? 'Saving…' : 'Add alert'}
</Button>
</div>
</form>
)
}
function AlertRow({
alert,
busy,
catalog,
valid,
onToggle,
onVia,
onFallback,
onDelete,
}: {
alert: Alert
busy: boolean
catalog: DetourCatalog
valid: Set<string>
onToggle: (on: boolean) => void
onVia: (v: string) => void
onFallback: (on: boolean) => void
onDelete: () => void
}) {
// Never render the token/URL in clear — show a masked descriptor only.
const detail = useMemo(() => {
if (alert.Type === 'telegram') {
return { text: `chat ${alert.ChatID || '—'}`, masked: !!alert.Token }
}
const { host, masked } = maskUrl(alert.URL ?? '')
return { text: host, masked: masked || !!alert.URL }
}, [alert.Type, alert.ChatID, alert.Token, alert.URL])
const events = asArray(alert.Events)
const canon = useMemo(() => canonDetour(alert.Via, catalog), [alert.Via, catalog])
const route = useMemo(() => describeDetour(canon, catalog, valid), [canon, catalog, valid])
const routed = canon !== 'direct'
const fallback = alert.Fallback ?? false
return (
<li className="dns-row">
<Toggle
pressed={alert.Enabled}
onChange={onToggle}
label={`${alert.Enabled ? 'Disable' : 'Enable'} alert ${alert.Name}`}
disabled={busy}
/>
<div className="dns-row-main">
<div className="dns-row-l1">
<span className="dns-row-name">{alert.Name}</span>
<span className="dns-badge">{alert.Type}</span>
{events.map((e) => (
<span key={e} className="dns-badge dns-badge--accent">
{e}
</span>
))}
</div>
<div className="dns-row-l2 mono">
<span className="dns-row-detail">{detail.text}</span>
{detail.masked && (
<span className="dns-masked" title="Secret is stored but hidden here">
secret hidden
</span>
)}
{route.direct ? (
<span className="dns-path">direct</span>
) : (
<span className="dns-path" data-active="on" data-missing={route.missing ? 'y' : undefined}>
{route.prefix} <strong className="dns-path-name">{route.name}</strong>
{route.missing && <span className="dns-path-flag"> (missing)</span>}
{fallback ? ' · +direct fallback' : ' · no fallback'}
</span>
)}
</div>
{routed && !fallback && <p className="dns-alert-note dns-alert-note--row">{VIA_NO_FALLBACK_NOTE}</p>}
</div>
<div className="dns-alert-ctl">
<label className="dns-detour">
<span className="dns-detour-label mono">Deliver via</span>
<DetourSelect
value={canon}
catalog={catalog}
valid={valid}
busy={busy}
disabled={false}
ariaLabel={`Deliver alert ${alert.Name} via`}
onChange={onVia}
directLabel="Direct (default)"
/>
</label>
<label className="dns-fallback">
<Toggle
pressed={fallback}
onChange={onFallback}
label={`${fallback ? 'Disable' : 'Enable'} direct fallback for ${alert.Name}`}
disabled={busy || !routed}
/>
<span className="dns-fallback-label mono">Fallback to direct</span>
</label>
</div>
<Button
className="dns-del"
onClick={onDelete}
disabled={busy}
aria-label={`Delete alert ${alert.Name}`}
>
Delete
</Button>
</li>
)
}
// ---- add form --------------------------------------------------------------
interface AddDraft {
@@ -1695,8 +1358,11 @@ function ListRow({
categories,
response,
filterOn,
statuses,
updating,
busy,
onToggle,
onUpdateNow,
onDelete,
}: {
name: string
@@ -1708,8 +1374,14 @@ function ListRow({
categories?: string[] | null
response?: BlockResponse
filterOn: boolean
/** What the running engine reports about this list, one record per geo category.
* null/[] ⇒ nothing reported: an older daemon, a stopped engine, or a list that
* has not been applied yet. The row then says so instead of guessing. */
statuses: RulesetStatus[] | null
updating: boolean
busy: boolean
onToggle: (on: boolean) => void
onUpdateNow: () => void
onDelete: () => void
}) {
const detail = useMemo<{ text: string; masked: boolean; title?: string }>(() => {
@@ -1733,8 +1405,45 @@ function ListRow({
}
}, [source, url, path, entries, categories])
// A list only actually filters when both it and the master switch are on.
const active = enabled && filterOn
// url and geosite lists are FETCHED by the engine; inline and file ones are read
// straight from the config and are loaded the moment they are applied.
const remote = source === 'url' || source === 'geosite'
const recs = statuses ?? []
const hasStatus = recs.length > 0
// A geo list with several categories: the OLDEST fetch (so a category that never
// arrived is never hidden behind a fresh sibling) and the SUM of the counts.
let ruleCount = 0
let neverAny = false
let oldestIso = ''
for (const s of recs) {
ruleCount += s.rule_count
if (!s.last_updated) neverAny = true
else if (!oldestIso || Date.parse(s.last_updated) < Date.parse(oldestIso)) oldestIso = s.last_updated
}
const interval = everyLabel(recs[0]?.interval_seconds ?? 0)
/**
* Whether this list is BLOCKING ANYTHING, which is a different question from
* whether it is switched on — and the one the row used to answer wrongly.
*
* "filtering" is now only said when the engine reports rules loaded for it. A
* remote list that has never been fetched (the daemon raises this as a critical
* apply finding) reads "not loaded", and one that fetched an empty list reads
* "empty". Nothing reported at all is "load not reported": unknown, not green.
*/
const state: { text: string; tone: 'on' | 'off' | 'warn' } = !enabled
? { text: 'off', tone: 'off' }
: !filterOn
? { text: 'inactive', tone: 'off' }
: !remote
? { text: 'filtering', tone: 'on' }
: !hasStatus
? { text: 'load not reported', tone: 'off' }
: neverAny
? { text: 'not loaded — nothing blocked', tone: 'warn' }
: ruleCount === 0
? { text: 'loaded empty — nothing blocked', tone: 'warn' }
: { text: 'filtering', tone: 'on' }
return (
<li className="dns-row">
@@ -1764,10 +1473,39 @@ function ListRow({
token hidden
</span>
)}
<span className="dns-row-state" data-active={active ? 'on' : 'off'}>
{active ? 'filtering' : 'inactive'}
<span className="dns-row-state" data-active={state.tone}>
{state.text}
</span>
</div>
{remote && (
<div className="dns-row-sync">
<span className="dns-sync-fresh" data-never={hasStatus && neverAny ? 'y' : undefined}>
{hasStatus ? relFetch(oldestIso) : 'status pending'}
</span>
{interval && <span className="dns-sync-every">{interval}</span>}
{ruleCount > 0 && (
<span className="dns-sync-rules">
{ruleCount.toLocaleString('en-US')} rule{ruleCount === 1 ? '' : 's'}
</span>
)}
<button
type="button"
className="dns-sync-update"
onClick={onUpdateNow}
disabled={busy || updating}
aria-label={`Update ${name} now`}
>
{updating ? (
<>
<span className="dns-sync-spin" aria-hidden="true" />
<span>Updating…</span>
</>
) : (
'Update now'
)}
</button>
</div>
)}
</div>
<Button
className="dns-del"
+3 -51
View File
@@ -138,57 +138,9 @@
/* inline rename: a quiet pencil affordance beside the name, and the mono input
it swaps to — in the same sink/groove tone as the domain editors. */
.dev-rename {
flex: none;
display: inline-flex;
align-items: center;
justify-content: center;
width: 22px;
height: 22px;
padding: 0;
border: 1px solid transparent;
border-radius: 5px;
background: none;
color: var(--faint);
font-size: 12px;
line-height: 1;
cursor: pointer;
transition: color 0.15s, background 0.15s, border-color 0.15s;
}
.dev-rename:hover:not(:disabled) {
color: var(--accent);
background: color-mix(in srgb, var(--accent) 12%, transparent);
}
.dev-rename:focus-visible {
color: var(--accent);
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.dev-rename:disabled {
opacity: 0.5;
cursor: default;
}
.dev-name-input {
min-width: 0;
max-width: 24ch;
padding: 4px 8px;
border: 1px solid var(--accent);
border-radius: 6px;
background: var(--sink);
color: var(--ink);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
box-shadow: 0 1px 2px var(--shadow) inset;
}
.dev-name-input:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.dev-name-input:disabled {
opacity: 0.55;
}
/* The pencil button and the name input now live in App.css as .inline-rename /
.inline-rename-input — Nodes grew the same affordance and the two pages must
not drift. */
.dev-id-l2 {
display: flex;
align-items: center;
+13 -7
View File
@@ -1,6 +1,6 @@
import './Devices.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Module, Toggle } from '../components'
import { Button, Led, Module, Toggle, useConfirm } from '../components'
import type { LedVariant } from '../components'
import { apply as apiApply, getConfig, getDevices, putConfig, ApiError } from '../api'
import type { Device, DiscoveredDevice, Model } from '../api'
@@ -81,6 +81,7 @@ function networkLabel(row: DeviceRow): string {
// ---- page ------------------------------------------------------------------
export default function Devices() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -251,17 +252,22 @@ export default function Devices() {
const nameOf = (row: DeviceRow) => row.cfg?.Name || row.hostname || row.ip || 'device'
const removeControl = useCallback(
(row: DeviceRow) => {
async (row: DeviceRow) => {
if (!config) return
const devs = asArray(config.Devices)
const idx = matchDevice(devs, row.mac, row.ip)
if (idx < 0) return
const nm = devs[idx].Name || nameOf(row)
if (!window.confirm(`Stop managing “${nm}”? Its per-device rules are removed; it falls back to network defaults.`))
return
const ok = await confirm({
label: 'Stop managing device',
title: `Stop managing “${nm}”?`,
body: 'Its per-device rules are removed; it falls back to network defaults.',
confirmLabel: 'Stop managing',
})
if (!ok) return
void save({ ...config, Devices: devs.filter((_, i) => i !== idx) }, `Removed control for ${nm}`)
},
[config, save],
[config, save, confirm],
)
const loading = config === null && loadError === null && devices === null && devError === null
@@ -471,7 +477,7 @@ function DeviceCard({
{renaming ? (
<input
ref={nameInput}
className="dev-name-input mono"
className="inline-rename-input mono"
type="text"
spellCheck={false}
autoComplete="off"
@@ -497,7 +503,7 @@ function DeviceCard({
</span>
<button
type="button"
className="dev-rename"
className="inline-rename"
onClick={beginRename}
disabled={busy}
aria-label={`Rename ${name}`}
+39 -19
View File
@@ -1,6 +1,6 @@
import './Networks.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Select, Toggle } from '../components'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Inbound, Interface, Model, Status } from '../api'
import { isLanNetwork, isWanNetwork, useInterfaces } from '../srcOptions'
@@ -85,6 +85,19 @@ const DEFAULT_TPROXY_PORT = 12345
* Collapsing those into one switch would make "I want ping to work" silently mean
* "I permit a parallel VPN bypass", so the middle option exists to remove that
* false choice — and the labels push anyone who wants diagnostics to `icmp`.
*
* WHY THIS COPY WAS REWRITTEN. `block` used to say "Nothing leaves except through
* the tunnel", and it was not true. The daemon let untunnelable traffic out toward
* every destination the ROUTING RULES send direct, on the argument that such a host
* already has your address from ordinary TCP. Under the commonest setup here —
* "tunnel what's blocked, send the rest direct" — the routing default IS direct, so
* that covered everything: `block` behaved exactly like `direct`, including ESP/GRE,
* i.e. the parallel-VPN case the middle rung exists to exclude. The daemon now drops
* unconditionally under `block`, and this copy states the price instead of hiding it
* (the owner's call: this router does not do ping and does not do IPTV).
*
* `icmp` still carries that destination-dependence for its NON-ping half, so its
* cost line says so rather than claiming "nothing else gets out".
*/
type Untunnelable = 'block' | 'icmp' | 'direct'
@@ -98,7 +111,10 @@ function normUntunnelable(raw: string | undefined): Untunnelable {
const UNTUNNELABLE_OPTIONS: ReadonlyArray<{ value: string; label: string }> = [
{ value: 'block', label: 'Block everything — most private' },
{ value: 'icmp', label: 'Allow ping only — for diagnostics' },
// Not "Allow ping only": the rung also lets the other untunnelable protocols
// out toward directly-routed addresses, and the cost line below says so. A
// label that promised "only" would be contradicted two lines under itself.
{ value: 'icmp', label: 'Allow ping — for diagnostics' },
{ value: 'direct', label: 'Allow everything — most compatible' },
]
@@ -110,13 +126,18 @@ interface PolicyCopy {
const UNTUNNELABLE_COPY: Record<Untunnelable, PolicyCopy> = {
block: {
works: 'Nothing leaves except through the tunnel.',
cost: 'Ping and traceroute won’t work from your devices, and neither will multicast IPTV or connecting to a VPN from a device on your network.',
// Scoped to "this traffic" on purpose. The old line — "Nothing leaves except
// through the tunnel" — was doubly loose: it was false (see the note above),
// and even read charitably it collides with directly-routed TCP, which does
// leave outside the tunnel by design.
works:
'None of this traffic leaves the router — it’s dropped, whatever your routing rules say. It’s the only setting whose promise doesn’t depend on how the rules are written.',
cost: 'Ping and traceroute stop working from your devices. So do IPsec and PPTP VPN connections made from a device on your network, multicast IPTV, and SCTP. VPNs that run over UDP — WireGuard, OpenVPN-UDP, and IPsec through NAT (IKEv2/NAT-T) — are unaffected: they go through the tunnel like everything else.',
tone: 'good',
},
icmp: {
works: 'Ping and traceroute work, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. Only for hosts you deliberately ping, and nothing else gets out — IPTV and VPN connections stay blocked.',
works: 'Ping and traceroute work everywhere, so you can check whether something is reachable.',
cost: 'Whatever you ping sees your real IP address instead of the tunnel’s. IPsec, PPTP and IPTV also get out — but only toward addresses your routing rules already send direct, so a VPN app on a device can still open its own connection beside this one if its server is one of those.',
tone: 'warn',
},
direct: {
@@ -242,6 +263,7 @@ function computeWarnings(inbounds: Inbound[], ifaces: Interface[]): Warning[] {
// ---- page ------------------------------------------------------------------
export default function Networks({ status }: { status?: Status | null }) {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
const ifaces = useInterfaces()
@@ -392,23 +414,21 @@ export default function Networks({ status }: { status?: Status | null }) {
)
const removeInbound = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = inbounds[idx]
if (
!window.confirm(
`Delete inbound “${target.Name}”?${
intercepts(target)
? ` ${target.Network || 'Its network'} stops going through the tunnel.`
: ''
}`,
)
)
return
const ok = await confirm({
label: 'Delete inbound',
title: `Delete inbound “${target.Name}”?`,
body: intercepts(target)
? `${target.Network || 'Its network'} stops going through the tunnel.`
: undefined,
})
if (!ok) return
const next = inbounds.filter((_, i) => i !== idx)
void save({ ...config, Inbounds: next }, `Deleted ${target.Name}`)
},
[config, inbounds, save],
[config, inbounds, save, confirm],
)
return (
@@ -552,7 +572,7 @@ export default function Networks({ status }: { status?: Status | null }) {
{untunnelable === 'block' && (
<p className="nw-sec-note nw-policy-hint">
If you just want to check whether a site is reachable, choose <strong>Allow ping only</strong>{' '}
If you just want to check whether a site is reachable, choose <strong>Allow ping</strong>{' '}
rather than allowing everything — it’s the narrower of the two.
</p>
)}
+110
View File
@@ -196,6 +196,52 @@
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
color: var(--amber);
}
/* "not built" — the saved switch says on and the engine has no such outbound. */
.badge--crit {
border-color: color-mix(in srgb, var(--crit) 55%, var(--groove));
color: var(--crit);
}
/* ---- last-apply findings, attached to the row they are about ----
Sits under the row's own two lines, inside the row plate, so a node the
generator threw away cannot read as an ordinary enabled node. Severity carries
the colour; the accent stays reserved for controls. */
.row-findings {
margin: 6px 0 0;
padding: 0;
list-style: none;
display: flex;
flex-direction: column;
gap: 5px;
}
.row-finding {
display: flex;
align-items: flex-start;
gap: 8px;
padding: 7px 9px;
border: 1px solid color-mix(in srgb, var(--amber) 40%, var(--groove));
border-radius: 6px;
background: color-mix(in srgb, var(--sink) 35%, transparent);
}
.row-finding--critical {
border-color: color-mix(in srgb, var(--crit) 45%, var(--groove));
}
.row-finding-msg {
flex: 1;
min-width: 0;
font-size: 12px;
line-height: 1.5;
color: var(--ink);
max-width: 82ch;
overflow-wrap: anywhere;
}
/* Findings that belong to no single row (see Nodes.tsx globalFindings). */
.node-findings {
margin-bottom: calc(var(--u, 8px) * 2);
}
.node-findings .row-findings {
margin-top: 0;
}
/* masked-credential marker */
.masked {
@@ -641,6 +687,20 @@ select.fp-input {
letter-spacing: 0.06em;
color: var(--faint);
}
/* A collapsed bucket has to carry its own bad news: a 300-node subscription is
closed by default, and the per-row findings inside it are otherwise unreachable
without knowing to look. */
.group-flagged {
flex: none;
display: inline-flex;
align-items: center;
gap: 6px;
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--amber);
}
.group-rows {
margin-top: 8px;
}
@@ -727,3 +787,53 @@ select.fp-input {
width: 9rem;
}
}
/* ---- inline node rename ----
The pencil / input pair itself is shared (.inline-rename[-input] in App.css);
only the row-local sizing and the refusal message live here. A node name is
longer than a device name (it carries a protocol and a host), so the field is
given more room than the shared 24ch default. */
.node-name-input {
max-width: 32ch;
font-family: var(--font-mono);
font-size: 12.5px;
}
/* Why a rename was refused, pinned under the row it was typed in. Semantic crit:
the name did not change, and that must not be mistaken for a saved edit. */
.row-err {
margin: 2px 0 0;
font-size: 11.5px;
line-height: 1.45;
color: var(--crit);
max-width: 68ch;
}
/* Stated once per subscription bucket: the same rule the locked control in every
row carries, so the absent rename is explained before it is looked for. */
.group-note {
margin: 0;
padding: 8px 12px;
border: 1px solid var(--groove);
border-top: 0;
background: color-mix(in srgb, var(--sink) 25%, transparent);
font-size: 11.5px;
line-height: 1.5;
color: var(--faint);
}
/* The optional name sits beside the link input on a wide row and drops onto its
own line when the row can no longer hold both. */
.add-name {
flex: 0 1 22ch;
min-width: 12ch;
}
.add-row--conf .add-name {
flex: none;
align-self: stretch;
}
@media (max-width: 640px) {
.add-row {
flex-wrap: wrap;
}
.add-name {
flex: 1 1 100%;
}
}
+638 -17
View File
@@ -2,16 +2,18 @@ import './Nodes.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import type { LedVariant } from '../components'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
getStatus,
putConfig,
importWg,
updateSubscription,
ApiError,
} from '../api'
import type { Model, Node as NodeCfg, Subscription } from '../api'
import type { Model, Node as NodeCfg, StatusWarning, Subscription } from '../api'
import { entityFindings, findingsByName } from '../findings'
import { fmtBytes, fmtDate, fmtUntil } from '../format'
// The whole page is a thin editor over the desired-state Model: every mutation
@@ -158,6 +160,219 @@ function uniqueName(base: string, taken: Set<string>): string {
return `${seed}-${i}`
}
// ---- node names are identity, not a caption --------------------------------
//
// A node's Name IS its sing-box outbound tag and the only thing every reference
// to it spells: a rule target `node:<name>`, a chain hop, a manual group's member
// list, a resolver detour, an alert delivery, a subscription fetch detour. Rename
// the node alone and every one of those points at nothing — and an unresolved
// target does NOT fall back to the default route, the daemon BLOCKS that traffic.
// So the rename either carries every reference with it, or it is refused.
/** Reserved outbound tags. A node called this is skipped by the generator entirely. */
const RESERVED_TAGS = ['direct', 'block']
/**
* Prefixes that `model.SplitTarget` reads as a KIND, not as part of a name. A
* name starting with one of them makes every bare reference to it ambiguous with
* a real `kind:name` reference, so it is refused rather than half-supported.
*/
const KIND_PREFIXES = ['node', 'group', 'egress', 'chain', 'direct', 'block']
/** Names are rendered into a line-oriented `uci export`; control chars are stripped there. */
function hasControlChar(s: string): boolean {
for (let i = 0; i < s.length; i++) {
const c = s.charCodeAt(i)
if (c < 0x20 || c === 0x7f) return true
}
return false
}
/**
* Why `name` cannot be a node name here, or null if it can.
*
* Every rule mirrors something the daemon actually does with the name, not a
* house style: reserved tags make generate skip the node; a duplicate makes two
* outbounds share a tag and the manager silently keeps the last one; a group of
* the same name is dropped by buildGroups ("rename the group"); an
* `egress-<name>` collision takes over a real egress outbound; and a control
* character is rewritten to a space by sanitizeUCIValue on write, so the saved
* name would not be the one you typed.
*/
function nodeNameError(
raw: string,
m: Model | null,
self: string | null,
): string | null {
const name = raw.trim()
if (!name) return 'A node needs a name.'
if (hasControlChar(name))
return 'Names can’t contain line breaks or control characters — they’re stripped when the config is written.'
if (RESERVED_TAGS.some((t) => t.toLowerCase() === name.toLowerCase()))
return `“${name}” is a reserved target name — a node called that is skipped by the engine. Pick another.`
const head = name.includes(':') ? name.slice(0, name.indexOf(':')).toLowerCase() : ''
if (head && KIND_PREFIXES.includes(head))
return `A name starting with “${head}:” reads as a ${head} reference everywhere it’s used. Pick another.`
if (!m) return null
const clash = asArray(m.Nodes).find((n) => n.Name === name && n.Name !== self)
if (clash)
return clash.FromSub
? `“${name}” is already a node from subscription “${clash.FromSub}”. Two nodes with one name share a single outbound — pick another.`
: `“${name}” is already another node. Pick another.`
if (asArray(m.Groups).some((g) => g.Name === name))
return `A group is already named “${name}”. The engine drops the group when a node takes its name — pick another.`
const egressClash = asArray(m.Egresses).find((e) => `egress-${e.Name}` === name)
if (egressClash)
return `“${name}” is the outbound tag of egress “${egressClash.Name}”. Pick another.`
return null
}
/** One place a node name is written, as a short label for the rename summary. */
interface NodeRefSite {
/** Which section — drives the "N rules, M groups" count. */
kind: 'rule' | 'group' | 'chain' | 'resolver' | 'alert' | 'subscription' | 'egress'
label: string
}
/** A target/detour string naming this node in its prefixed form (`node:<name>`). */
const isNodeRef = (v: string | undefined | null, name: string): boolean =>
(v ?? '') === `node:${name}`
/** …or in the bare form the engine also resolves (a group member, a bare hop/target). */
const isBareRef = (v: string | undefined | null, name: string): boolean => (v ?? '') === name
/**
* Every place `name` is written outside the node itself. Both spellings count:
* `resolveTarget` falls through to a bare node lookup, and a manual group's
* member list is bare by contract.
*/
function findNodeReferences(m: Model, name: string): NodeRefSite[] {
const out: NodeRefSite[] = []
for (const r of asArray(m.Rules)) {
if (isNodeRef(r.Target, name) || isBareRef(r.Target, name))
out.push({ kind: 'rule', label: `rule “${r.Name}” target` })
}
for (const g of asArray(m.Groups)) {
if (asArray(g.Nodes).some((n) => n === name))
out.push({ kind: 'group', label: `group “${g.Name}” member` })
}
for (const c of asArray(m.Chains)) {
if (asArray(c.Hops).some((h) => isNodeRef(h, name) || isBareRef(h, name)))
out.push({ kind: 'chain', label: `chain “${c.Name}” hop` })
}
for (const r of asArray(m.Resolvers)) {
if (isNodeRef(r.Detour, name)) out.push({ kind: 'resolver', label: `resolver “${r.Name}” DNS path` })
}
for (const a of asArray(m.Alerts)) {
if (isNodeRef(a.Via, name)) out.push({ kind: 'alert', label: `alert “${a.Name}” delivery` })
}
for (const s of asArray(m.Subscriptions)) {
if (isNodeRef(s.FetchDetour, name))
out.push({ kind: 'subscription', label: `subscription “${s.Name}” fetch` })
}
for (const e of asArray(m.Egresses)) {
if (isNodeRef(e.Target, name)) out.push({ kind: 'egress', label: `egress “${e.Name}” target` })
}
return out
}
/**
* What makes a rename impossible to carry rather than merely wide.
*
* A BARE reference is just a name; the engine resolves it node-first, then group.
* If something else already answers to the old name, we cannot tell which object
* a bare reference meant, and rewriting it would move a reference the operator
* never pointed at this node. That is a half-done cascade, so the rename is
* refused instead — with the collision named, so it can be fixed.
*/
function bareAmbiguity(m: Model, name: string): string | null {
const group = asArray(m.Groups).find((g) => g.Name === name)
if (!group) return null
const bare = [
...asArray(m.Rules)
.filter((r) => isBareRef(r.Target, name))
.map((r) => `rule “${r.Name}”`),
...asArray(m.Chains)
.filter((c) => asArray(c.Hops).some((h) => isBareRef(h, name)))
.map((c) => `chain “${c.Name}”`),
]
if (bare.length === 0) return null
return `A group is also named “${name}”, and ${bare.join(', ')} point${bare.length === 1 ? 's' : ''} at that bare name — there is no way to tell which of the two is meant. Rename the group first, then this node.`
}
/**
* Rewrite every reference from `from` to `to`. Returns a NEW Model with only the
* touched sections replaced; the Nodes section is the caller's business.
*
* Bare references are rewritten too — that is the whole point for a manual
* group's member list — which is safe only because `bareAmbiguity` has already
* refused the one case where a bare name could mean something else.
*/
function renameNodeReferences(m: Model, from: string, to: string): Model {
if (from === to) return m
/** Prefixed-only sites (a detour is never spelled bare). */
const pfx = (v: string | undefined) => (isNodeRef(v, from) ? `node:${to}` : v)
/** Sites that accept either spelling — each is rewritten in the spelling it already uses. */
const either = (v: string | undefined) => {
if (isNodeRef(v, from)) return `node:${to}`
if (isBareRef(v, from)) return to
return v
}
const next: Model = { ...m }
if (m.Rules) next.Rules = m.Rules.map((r) => ({ ...r, Target: either(r.Target) }))
if (m.Groups)
next.Groups = m.Groups.map((g) => ({
...g,
Nodes: g.Nodes ? g.Nodes.map((n) => (n === from ? to : n)) : g.Nodes,
}))
if (m.Chains)
next.Chains = m.Chains.map((c) => ({
...c,
Hops: c.Hops ? c.Hops.map((h) => either(h) ?? h) : c.Hops,
}))
if (m.Resolvers) next.Resolvers = m.Resolvers.map((r) => ({ ...r, Detour: pfx(r.Detour) }))
if (m.Alerts) next.Alerts = m.Alerts.map((a) => ({ ...a, Via: pfx(a.Via) }))
if (m.Subscriptions)
next.Subscriptions = m.Subscriptions.map((s) => ({ ...s, FetchDetour: pfx(s.FetchDetour) }))
if (m.Egresses) next.Egresses = m.Egresses.map((e) => ({ ...e, Target: pfx(e.Target) }))
return next
}
/** "3 rules, 1 group and 2 chains" — what the rename is about to rewrite. */
function refSummary(refs: NodeRefSite[]): string {
const plural: Record<NodeRefSite['kind'], [string, string]> = {
rule: ['rule', 'rules'],
group: ['group', 'groups'],
chain: ['chain', 'chains'],
resolver: ['resolver', 'resolvers'],
alert: ['alert', 'alerts'],
subscription: ['subscription', 'subscriptions'],
egress: ['egress', 'egresses'],
}
const order: NodeRefSite['kind'][] = [
'rule', 'group', 'chain', 'resolver', 'alert', 'subscription', 'egress',
]
const parts = order
.map((k) => [k, refs.filter((r) => r.kind === k).length] as const)
.filter(([, n]) => n > 0)
.map(([k, n]) => `${n} ${plural[k][n === 1 ? 0 : 1]}`)
if (parts.length === 1) return parts[0]
return `${parts.slice(0, -1).join(', ')} and ${parts[parts.length - 1]}`
}
/**
* The name a rename just committed to, waiting for its row to come back.
*
* A row is keyed by the node's NAME, so committing a rename unmounts the row and
* mounts a different one — carrying the focused element away with it. This baton
* survives that remount: the row that reappears under the new name claims it and
* puts the keyboard back on its own rename button, instead of dropping the user
* on <body> halfway down a list of 300 nodes.
*/
let pendingRenameFocus: string | null = null
// A subscription with more than this many nodes starts collapsed so the list
// doesn't become one endless scroll; an active search overrides it.
const LARGE_GROUP = 20
@@ -289,6 +504,7 @@ function DetourSelect({
// ---- page ------------------------------------------------------------------
export default function Nodes() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -305,6 +521,43 @@ export default function Nodes() {
void loadConfig()
}, [loadConfig])
// ---- what the last apply said about these nodes ---------------------------
//
// The generator drops a node it cannot build and names it: an unparseable
// share link (generate/outbound.go), a name colliding with a reserved tag, a
// WireGuard private key materialised twice (generate/wgdedup.go). Until now
// this page never read /api/status, so a node the engine had thrown away
// rendered as an ordinary row with a green toggle — the switch said on and
// there was no such outbound anywhere in the running config.
//
// Findings are attached to the ROWS, not summarised at the top: a 300-node
// subscription makes a list of names useless, and the row is where the false
// reassurance was.
const [findings, setFindings] = useState<StatusWarning[]>([])
const loadFindings = useCallback(async () => {
try {
const s = await getStatus()
setFindings(entityFindings(s.warnings, ['node', 'subscription']))
} catch {
// Status is a supplement here, not the page. Keep the last set rather than
// clearing it — a dropped poll is not the same as "the problem is fixed".
}
}, [])
useEffect(() => {
void loadFindings()
}, [loadFindings])
const nodeFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'node')),
[findings],
)
const subFindings = useMemo(
() => findingsByName(findings.filter((w) => w.section === 'subscription')),
[findings],
)
// Findings about nodes/subscriptions in general, which belong to no single row.
const globalFindings = useMemo(() => findings.filter((w) => !w.name), [findings])
// ---- toast + persistent apply banner --------------------------------------
const [toast, setToast] = useState<string | null>(null)
const toastTimer = useRef<number | undefined>(undefined)
@@ -359,8 +612,11 @@ export default function Nodes() {
flash(`Apply failed — ${errText(e)}`)
} finally {
setApplying(false)
// An apply is exactly what rewrites the findings — including clearing the
// ones the operator just fixed.
void loadFindings()
}
}, [flash, loadConfig])
}, [flash, loadConfig, loadFindings])
// ---- node mutations -------------------------------------------------------
const nodes = useMemo(() => asArray(config?.Nodes), [config])
@@ -385,6 +641,9 @@ export default function Nodes() {
const [nodeInput, setNodeInput] = useState('')
const [nodeErr, setNodeErr] = useState<string | null>(null)
const [addMode, setAddMode] = useState<'link' | 'conf'>('link')
// Optional. Empty keeps the old behaviour (a name derived from the server
// address), so "paste a link, press Add" stays a two-step path.
const [nodeName, setNodeName] = useState('')
const [importing, setImporting] = useState(false)
// ---- node search + collapsible grouping -----------------------------------
@@ -445,8 +704,12 @@ export default function Nodes() {
try {
const { uri, name } = await importWg(conf)
const taken = new Set(nodes.map((n) => n.Name))
// A typed name is used AS TYPED — uniqueName would silently turn a
// collision into "name-2", which is the confusion this field exists to
// end. It is validated instead, and a clash is refused out loud above.
const wanted = nodeName.trim()
const node: NodeCfg = {
Name: uniqueName(name || 'wireguard', taken),
Name: wanted || uniqueName(name || 'wireguard', taken),
Enabled: true,
URI: uri,
FromSub: '',
@@ -455,6 +718,7 @@ export default function Nodes() {
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
if (ok) {
setNodeInput('')
setNodeName('')
setAddMode('link')
}
} catch (e) {
@@ -463,11 +727,20 @@ export default function Nodes() {
setImporting(false)
}
},
[config, nodes, save, flash],
[config, nodes, nodeName, save, flash],
)
const addNode = useCallback(async () => {
if (!config) return
// The name is checked BEFORE the import round-trip, so a bad name costs
// nothing and the message lands in the form next to the field.
if (nodeName.trim()) {
const bad = nodeNameError(nodeName, config, null)
if (bad) {
setNodeErr(bad)
return
}
}
// Auto-detect a pasted config, whichever input it landed in.
if (nodeInput.includes(WG_MARKER)) {
await addWgConf(nodeInput)
@@ -486,10 +759,20 @@ export default function Nodes() {
const parsed = parseShareLink(uri)
const taken = new Set(nodes.map((n) => n.Name))
const base = parsed.suggested || `${parsed.proto.toLowerCase()}-${parsed.host}`.replace(/[^\w.:-]+/g, '-')
const node: NodeCfg = { Name: uniqueName(base, taken), Enabled: true, URI: uri, FromSub: '', Egress: '' }
const wanted = nodeName.trim()
const node: NodeCfg = {
Name: wanted || uniqueName(base, taken),
Enabled: true,
URI: uri,
FromSub: '',
Egress: '',
}
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
if (ok) setNodeInput('')
}, [config, nodeInput, nodes, save, addMode, addWgConf])
if (ok) {
setNodeInput('')
setNodeName('')
}
}, [config, nodeInput, nodeName, nodes, save, addMode, addWgConf])
const toggleNode = useCallback(
(idx: number, on: boolean) => {
@@ -500,15 +783,110 @@ export default function Nodes() {
[config, nodes, save],
)
/**
* Delete a node, naming everything that still points at it.
*
* `findNodeReferences` was already here and already right — it just wasn't asked
* on the one path where the answer matters. A RENAME carried its references and
* said so; a DELETE said "This removes it from the config", which is true of the
* node and silent about the rule, group member, chain hop or resolver detour
* left spelling a name nothing answers to. That is not a cosmetic dangle: an
* unresolved target does not fall through to the default route, so the traffic
* aimed at it is blocked.
*/
const removeNode = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = nodes[idx]
if (!window.confirm(`Delete node “${target.Name}”? This removes it from the config.`)) return
const refs = findNodeReferences(config, target.Name)
const shown = refs.slice(0, 4).map((r) => r.label)
const more = refs.length - shown.length
const ok = await confirm({
label: 'Delete node',
title: `Delete node “${target.Name}”?`,
body:
refs.length === 0
? 'Nothing else in the config points at it.'
: `${refSummary(refs)} still ${refs.length === 1 ? 'points' : 'point'} at it — ${shown.join(
', ',
)}${
more > 0 ? `, and ${more} more` : ''
}. Nothing rewrites them, and a target that no longer resolves does not fall through to the default route: the traffic aimed at it is blocked.`,
})
if (!ok) return
const next = nodes.filter((_, i) => i !== idx)
void save({ ...config, Nodes: next }, `Deleted ${target.Name}`)
},
[config, nodes, save],
[config, nodes, save, confirm],
)
/**
* Rename a manual node, carrying every reference with it.
*
* The name is this node's identity: its outbound tag, and the exact string a
* rule target, a chain hop, a manual group's member list, a resolver detour, an
* alert delivery and a subscription fetch detour all spell. So the rename is one
* atomic save of the Nodes section AND every referencing section, or it does not
* happen at all:
*
* - an invalid or colliding name is refused with the reason (`nodeNameError`);
* - a name a GROUP also answers to, with bare references pointing at it, is
* refused too — there is no way to know which object those meant, and
* guessing would move a reference the operator never pointed here;
* - anything else is shown exactly what it will rewrite, and only then saved.
*
* Errors surface through `onError` so they land in the row that was edited.
*/
const renameNode = useCallback(
async (idx: number, raw: string, onError: (msg: string) => void): Promise<boolean> => {
if (!config) return false
const target = nodes[idx]
const from = target.Name
const to = raw.trim()
if (to === from) return true
// Subscription names come back from the feed on the next update; renaming
// one would be undone without warning, so this path is manual-only.
if (target.FromSub) {
onError(`“${from}” is named by subscription “${target.FromSub}” — the feed rewrites it on the next update.`)
return false
}
const bad = nodeNameError(to, config, from)
if (bad) {
onError(bad)
return false
}
const blocked = bareAmbiguity(config, from)
if (blocked) {
onError(blocked)
return false
}
const refs = findNodeReferences(config, from)
if (refs.length > 0) {
const shown = refs.slice(0, 4).map((r) => r.label)
const more = refs.length - shown.length
const ok = await confirm({
tone: 'neutral',
label: 'Rename node',
title: `Rename “${from}” to “${to}”?`,
body: `This also updates ${refSummary(refs)} that point at it — ${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}. They are saved together, so nothing is left pointing at the old name.`,
confirmLabel: 'Rename',
})
if (!ok) return false
}
// One PUT: the node and every reference move in the same write, so no
// intermediate state exists where a reference dangles.
const carried = renameNodeReferences(config, from, to)
const next = asArray(carried.Nodes).map((n, i) => (i === idx ? { ...n, Name: to } : n))
return save(
{ ...carried, Nodes: next },
refs.length > 0
? `Renamed to ${to} — updated ${refs.length} reference${refs.length === 1 ? '' : 's'}`
: `Renamed to ${to}`,
)
},
[config, nodes, save, confirm],
)
// Pin (or clear) one node's dial egress. Same optimistic save→apply path as
@@ -563,16 +941,20 @@ export default function Nodes() {
)
const removeSub = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = subs[idx]
const hasCache = nodes.some((n) => n.FromSub === target.Name)
const extra = hasCache ? ' Its cached nodes stay until you next apply.' : ''
if (!window.confirm(`Delete subscription “${target.Name}”?${extra}`)) return
const ok = await confirm({
label: 'Delete subscription',
title: `Delete subscription “${target.Name}”?`,
body: hasCache ? 'Its cached nodes stay until you next apply.' : undefined,
})
if (!ok) return
const next = subs.filter((_, i) => i !== idx)
void save({ ...config, Subscriptions: next }, `Deleted ${target.Name}`)
},
[config, subs, nodes, save],
[config, subs, nodes, save, confirm],
)
// Commit an options edit for one subscription. The editor hands back a fully
@@ -654,6 +1036,14 @@ export default function Nodes() {
</div>
)}
{/* Findings about nodes in general — no single row owns them, so they sit
above the lists rather than being dropped for having no name. */}
{globalFindings.length > 0 && (
<div className="node-findings">
<RowFindings findings={globalFindings} />
</div>
)}
{/* ---- NODES ---- */}
<div className="node-section" aria-label="Nodes">
<header className="sec-hd">
@@ -736,6 +1126,20 @@ export default function Nodes() {
disabled={busy || importing || !config}
/>
)}
<input
className="fp-input add-name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Name (optional)"
aria-label="Node name — optional"
value={nodeName}
onChange={(e) => {
setNodeName(e.target.value)
if (nodeErr) setNodeErr(null)
}}
disabled={busy || importing || !config}
/>
<Button type="submit" variant="primary" disabled={busy || importing || !config}>
{importing ? 'Importing…' : saving ? 'Saving…' : addMode === 'conf' ? 'Import' : 'Add node'}
</Button>
@@ -743,7 +1147,8 @@ export default function Nodes() {
<p className="add-hint">
{addMode === 'conf'
? 'Paste a wg-quick / AmneziaWG .conf — it starts with [Interface].'
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}{' '}
Leave the name empty and it’s taken from the server address; you can rename it later.
</p>
</div>
{nodeErr && (
@@ -794,12 +1199,14 @@ export default function Nodes() {
<NodeGroup
key={g.key || '__manual__'}
group={g}
findings={nodeFindings}
open={isGroupOpen(g)}
busy={busy}
egressNames={egressNames}
onToggle={() => toggleGroup(g)}
onToggleNode={toggleNode}
onRemoveNode={removeNode}
onRenameNode={renameNode}
onSetEgress={setNodeEgress}
/>
))}
@@ -882,6 +1289,7 @@ export default function Nodes() {
busy={busy}
catalog={detourCatalog}
valid={detourValid}
findings={subFindings.get(s.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => toggleSub(i, on)}
onDelete={() => removeSub(i)}
onEdit={(patch) => editSub(i, patch)}
@@ -908,23 +1316,35 @@ function NodeGroup({
open,
busy,
egressNames,
findings,
onToggle,
onToggleNode,
onRemoveNode,
onRenameNode,
onSetEgress,
}: {
group: NodeGroupData
open: boolean
busy: boolean
egressNames: string[]
/** Last-apply findings per node name (findings.ts findingsByName). */
findings: Map<string, StatusWarning[]>
onToggle: () => void
onToggleNode: (idx: number, on: boolean) => void
onRemoveNode: (idx: number) => void
onRenameNode: (idx: number, name: string, onError: (msg: string) => void) => Promise<boolean>
onSetEgress: (idx: number, egress: string) => Promise<boolean>
}) {
const panelId = `node-group-${group.key || 'manual'}`
// The same inventory count as the section header, scoped to this bucket.
const count = useMemo(() => fmtEnabled(group.items.map((i) => i.node)), [group.items])
// How many nodes in this bucket the last apply had something to say about —
// shown on the COLLAPSED header, because a subscription of 300 nodes is
// collapsed by default and the row badge below would never be seen otherwise.
const flagged = useMemo(
() => group.items.filter(({ node }) => findings.has(node.Name)).length,
[group.items, findings],
)
return (
<section className={`node-group${open ? ' node-group--open' : ''}`}>
<h3 className="group-hd-wrap">
@@ -938,8 +1358,20 @@ function NodeGroup({
<span className="group-caret" aria-hidden="true" />
<span className="group-name">{group.label}</span>
<span className="group-count mono">{count}</span>
{flagged > 0 && (
<span className="group-flagged" title="Findings from the last apply">
<Led variant="amber" />
{flagged} flagged
</span>
)}
</button>
</h3>
{open && group.key !== '' && (
<p className="group-note">
Names come from the subscription feed and are rewritten on every update, so nodes in this
list can’t be renamed here.
</p>
)}
{open && (
<ul id={panelId} className="rows-list group-rows">
{group.items.map(({ node, idx }) => (
@@ -948,8 +1380,10 @@ function NodeGroup({
node={node}
busy={busy}
egressNames={egressNames}
findings={findings.get(node.Name) ?? EMPTY_FINDINGS}
onToggle={(on) => onToggleNode(idx, on)}
onDelete={() => onRemoveNode(idx)}
onRename={(name, onError) => onRenameNode(idx, name, onError)}
onSetEgress={(egress) => onSetEgress(idx, egress)}
/>
))}
@@ -959,19 +1393,42 @@ function NodeGroup({
)
}
/** One shared empty array, so a clean row doesn't get a fresh identity per render. */
const EMPTY_FINDINGS: StatusWarning[] = []
/**
* Did the generator say it left this entity OUT of the engine config?
*
* The producers all end the sentence with the same word — "(skipped)" for an
* unparseable share link or a bad WireGuard endpoint (generate/outbound.go),
* "skipped" for a name colliding with a reserved tag — and wgdedup says only one
* of the duplicates "is kept". Read the daemon's word rather than inventing a
* verdict: a finding that does NOT say this may well be about a node that is
* running perfectly, and badging it "not built" would be a new lie in place of
* the old one.
*/
function skipped(findings: StatusWarning[]): boolean {
return findings.some((f) => /\bskipped\b|\bis kept\b/i.test(f.message))
}
function NodeRow({
node,
busy,
egressNames,
findings,
onToggle,
onDelete,
onRename,
onSetEgress,
}: {
node: NodeCfg
busy: boolean
egressNames: string[]
/** What the last apply said about THIS node; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onRename: (name: string, onError: (msg: string) => void) => Promise<boolean>
onSetEgress: (egress: string) => Promise<boolean>
}) {
const { proto, host, hasCreds } = useMemo(() => parseShareLink(node.URI), [node.URI])
@@ -982,6 +1439,68 @@ function NodeRow({
const [open, setOpen] = useState(false)
const panelId = `node-egress-${node.FromSub || 'manual'}-${node.Name}`
// ---- inline rename (same interaction as a device row) ---------------------
// Enter commits, Esc cancels, blur commits; a ref-guard keeps Esc-then-blur
// from committing twice. Unlike a device, the commit can be REFUSED (a name
// collision, or references that can't be carried), so the input stays open
// with the reason under it instead of closing on a change that never happened.
const [renaming, setRenaming] = useState(false)
const [draft, setDraft] = useState(node.Name)
const [renameErr, setRenameErr] = useState<string | null>(null)
const nameInput = useRef<HTMLInputElement>(null)
const renameBtn = useRef<HTMLButtonElement>(null)
const finished = useRef(false)
const beginRename = () => {
setDraft(node.Name)
setRenameErr(null)
finished.current = false
setRenaming(true)
}
const finishRename = async (commit: boolean) => {
if (finished.current) return
finished.current = true
const nm = draft.trim()
if (!commit || !nm || nm === node.Name) {
setRenaming(false)
setRenameErr(null)
return
}
// Armed BEFORE the save: the renamed row remounts the moment the config
// state lands, which is before this await resolves. Arming afterwards would
// always miss it.
pendingRenameFocus = nm
const ok = await onRename(nm, (msg) => setRenameErr(msg))
if (ok) {
setRenaming(false)
setRenameErr(null)
} else {
if (pendingRenameFocus === nm) pendingRenameFocus = null
// Refused — hold the field open on the rejected text so it can be fixed.
finished.current = false
nameInput.current?.focus()
}
}
useEffect(() => {
if (renaming) {
nameInput.current?.focus()
nameInput.current?.select()
}
}, [renaming])
// Claim the baton if this row is the one the rename produced. The row remounts
// while the PUT is still in flight, so on that first pass the button is still
// disabled and focus() would be a silent no-op — the baton is held until the
// save settles and this effect re-runs with a focusable button.
useEffect(() => {
if (pendingRenameFocus !== node.Name) return
const btn = renameBtn.current
if (!btn || btn.disabled) return
pendingRenameFocus = null
btn.focus()
}, [node.Name, busy])
return (
<li className={`row-item node-row${open ? ' node-row--open' : ''}`}>
<div className="row-head">
@@ -993,10 +1512,83 @@ function NodeRow({
/>
<div className="row-main">
<div className="row-line1">
<span className="row-name">{node.Name}</span>
{renaming ? (
<input
ref={nameInput}
className="inline-rename-input node-name-input mono"
type="text"
spellCheck={false}
autoComplete="off"
value={draft}
aria-label={`Rename node ${node.Name}`}
aria-invalid={renameErr ? true : undefined}
onChange={(e) => {
setDraft(e.target.value)
if (renameErr) setRenameErr(null)
}}
onBlur={() => void finishRename(true)}
onKeyDown={(e) => {
if (e.key === 'Enter') {
e.preventDefault()
void finishRename(true)
} else if (e.key === 'Escape') {
e.preventDefault()
void finishRename(false)
}
}}
disabled={busy}
/>
) : (
<>
<span className="row-name" title={node.Name}>
{node.Name}
</span>
{managed ? (
// Not hidden — withheld, with the reason attached. A control
// that quietly isn't there reads as a bug; this one states the
// rule, and the same sentence is on the group header above.
<button
type="button"
className="inline-rename inline-rename--locked"
disabled
aria-label={`Can’t rename ${node.Name} — its name comes from subscription “${node.FromSub}” and is rewritten on the next update`}
title={`Named by subscription “${node.FromSub}” — the feed rewrites this name on the next update. Rename it in the subscription, or add the node manually.`}
>
🔒
</button>
) : (
<button
ref={renameBtn}
type="button"
className="inline-rename"
onClick={beginRename}
disabled={busy}
aria-label={`Rename node ${node.Name}`}
title="Rename"
>
✎
</button>
)}
</>
)}
<span className="badge">{proto}</span>
{node.Stale && <span className="badge badge--warn">stale</span>}
{/* The toggle above is the SAVED state. When the last apply couldn't
build this node the engine has no such outbound, and the two
disagree — so the row says which, rather than leaving a green
switch to imply the node is carrying traffic. The word is the
daemon's own where it used one. */}
{findings.length > 0 && (
<span className={`badge badge--${skipped(findings) ? 'crit' : 'warn'}`}>
{skipped(findings) ? 'not built' : 'flagged'}
</span>
)}
</div>
{renameErr && (
<p className="row-err" role="alert">
{renameErr}
</p>
)}
<div className="row-line2 mono">
<span className="row-host">{host}</span>
{hasCreds && (
@@ -1014,6 +1606,7 @@ function NodeRow({
</span>
)}
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
{canPin && (
@@ -1133,6 +1726,7 @@ function SubRow({
busy,
catalog,
valid,
findings,
onToggle,
onDelete,
onEdit,
@@ -1143,6 +1737,8 @@ function SubRow({
busy: boolean
catalog: DetourCatalog
valid: Set<string>
/** What the last apply said about THIS subscription; empty when it said nothing. */
findings: StatusWarning[]
onToggle: (on: boolean) => void
onDelete: () => void
onEdit: (patch: Subscription) => Promise<boolean>
@@ -1167,6 +1763,7 @@ function SubRow({
<span className="row-name">{sub.Name}</span>
{sub.Format && sub.Format !== 'auto' && <span className="badge">{sub.Format}</span>}
{sub.FetchVia === 'proxy' && <span className="badge">via proxy</span>}
{findings.length > 0 && <span className="badge badge--warn">flagged</span>}
</div>
<div className="row-line2 mono">
<span className="row-host">{host}</span>
@@ -1179,6 +1776,7 @@ function SubRow({
every {interval} · {count} node{count === 1 ? '' : 's'}
</span>
</div>
<RowFindings findings={findings} />
</div>
<div className="row-actions">
<Button
@@ -1650,6 +2248,29 @@ function HeaderRows({
)
}
/**
* What the last apply said about THIS row, under the row it is about.
*
* Deliberately inside the row rather than in a list at the top of the page: the
* failure being fixed is a node that looks fine, and a name in a summary three
* screens up does not fix that. The wording is the daemon's own — these messages
* already name the entity and say what was done about it ("(skipped)", "only X
* is kept"), so paraphrasing them here would only invent a second vocabulary.
*/
function RowFindings({ findings }: { findings: StatusWarning[] }) {
if (findings.length === 0) return null
return (
<ul className="row-findings" aria-label="Findings from the last apply">
{findings.map((f, i) => (
<li key={i} className={`row-finding row-finding--${f.severity}`}>
<Led variant={f.severity === 'critical' ? 'crit' : 'amber'} />
<span className="row-finding-msg">{f.message}</span>
</li>
))}
</ul>
)
}
function EmptyPlate({ title, body }: { title: string; body: string }) {
return (
<div className="empty-plate">
+171 -54
View File
@@ -4,17 +4,18 @@ import type { LedVariant } from '../components'
import { fmtDateTime, fmtDuration } from '../format'
import {
apply as apiApply,
confirm as apiConfirm,
rollback as apiRollback,
getConfig,
getRulesReachability,
getStats,
ApiError,
} from '../api'
import type { Model, Stats, Status, StatusWarning } from '../api'
import { confirmTimeout } from '../pendingConfirm'
import { navigate } from '../router'
import type { Route } from '../router'
import { attentionFindings } from '../findings'
import { protectionState } from '../planeState'
import { attentionFindings, truncationNote } from '../findings'
import { engineReadout, killSwitchReadout, protectionState } from '../planeState'
// null-safe length for a Go slice that may arrive as null.
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
@@ -27,7 +28,8 @@ function short(hash: string): string {
return h.length > 12 ? h.slice(0, 12) : h
}
type ControlKind = 'apply' | 'confirm' | 'rollback'
// Confirm is no longer one of them — see the note beside the controls row.
type ControlKind = 'apply' | 'rollback'
/**
* Live service uptime in seconds, ticking between status polls.
@@ -42,17 +44,32 @@ type ControlKind = 'apply' | 'confirm' | 'rollback'
* Returns null when the daemon doesn't report uptime (older builds) — the caller
* then renders nothing rather than inventing a number.
*/
function useUptime(status: Status | null): number | null {
function useUptime(status: Status | null): { seconds: number; startedUnix: number } | null {
const base = useRef<{ uptime: number; at: number } | null>(null)
// The instant the daemon came up, ON THE BROWSER'S CLOCK.
//
// `status.started_unix` is the router's own clock, and the router has no RTC —
// it runs on UTC with no tzdata. Rendering it through the browser's timezone
// printed a start time three hours in the FUTURE for a Moscow operator, beside
// an uptime of "2 h 41 min". Deriving it instead as now-minus-uptime is a
// difference of two client timestamps, so it is skew-proof and can never land
// ahead of the clock in the header.
const started = useRef<number | null>(null)
const [, forceTick] = useState(0)
const reported = status?.uptime_seconds
useEffect(() => {
if (typeof reported !== 'number' || !Number.isFinite(reported)) {
base.current = null
started.current = null
return
}
base.current = { uptime: reported, at: Date.now() }
const now = Date.now()
base.current = { uptime: reported, at: now }
// Re-baselining every poll would jitter the displayed second back and forth;
// only move it when the estimate has genuinely drifted (a daemon restart).
const est = Math.round(now / 1000 - reported)
if (started.current === null || Math.abs(started.current - est) > 5) started.current = est
forceTick((n) => n + 1)
}, [reported])
@@ -64,8 +81,11 @@ function useUptime(status: Status | null): number | null {
return () => window.clearInterval(id)
}, [])
if (!base.current) return null
return base.current.uptime + Math.max(0, (Date.now() - base.current.at) / 1000)
if (!base.current || started.current === null) return null
return {
seconds: base.current.uptime + Math.max(0, (Date.now() - base.current.at) / 1000),
startedUnix: started.current,
}
}
export function Overview({
@@ -93,6 +113,29 @@ export function Overview({
void loadConfig()
}, [loadConfig])
// ---- how many rules are actually IN FORCE ----------------------------------
//
// `Rule.Enabled` from /api/config is the DESIRED state; the active WAN profile
// overrides it in either direction, and the daemon reports the result as
// `effective_enabled`. Counting the saved switches told a router running one
// chain that it had "2 / 2" — the Routing page had already been fixed to read
// the verdicts, and the home page kept summing the config beside it.
//
// null ⇒ no verdicts (older daemon, engine stopped, endpoint unreachable). The
// module then says so rather than passing the saved count off as the live one.
const [inForce, setInForce] = useState<number | null>(null)
const loadReach = useCallback(async () => {
try {
const { rules } = await getRulesReachability()
setInForce(rules.filter((r) => r.effective_enabled).length)
} catch {
setInForce(null)
}
}, [])
useEffect(() => {
void loadReach()
}, [loadReach])
// ---- live filter stats: poll the aggregate snapshot, degrade to honest empty states ----
const [stats, setStats] = useState<Stats | null>(null)
useEffect(() => {
@@ -121,7 +164,7 @@ export function Overview({
}
}, [])
// ---- apply / confirm / rollback ----
// ---- apply / rollback ----
const [busy, setBusy] = useState<ControlKind | null>(null)
const [result, setResult] = useState<{ ok: boolean; msg: string } | null>(null)
const [toast, setToast] = useState<string | null>(null)
@@ -139,22 +182,25 @@ export function Overview({
setBusy(kind)
setResult(null)
try {
const fn = kind === 'apply' ? apiApply : kind === 'confirm' ? apiConfirm : apiRollback
const r = await fn()
const r = kind === 'apply' ? await apiApply() : await apiRollback()
if (r.error) {
setResult({ ok: false, msg: r.error })
flash(`${kind} failed`)
} else {
// An apply that changed something armed an auto-rollback, and saying
// "data plane reconciled" while a timer runs is how someone walks away
// from a config that then reverts. Name the window when there is one.
const window = confirmTimeout()
const msg =
kind === 'apply'
? r.changed
? 'Applied — data plane reconciled'
? window > 0
? `Applied — keep this config within ${window}s or it rolls back`
: 'Applied — data plane reconciled'
: 'Applied — already up to date'
: kind === 'confirm'
? 'Confirmed — auto-rollback cancelled'
: 'Rolled back to last-good config'
: 'Rolled back to last-good config'
setResult({ ok: true, msg })
flash(kind === 'apply' ? 'Applied' : kind === 'confirm' ? 'Confirmed' : 'Rolled back')
flash(kind === 'apply' ? 'Applied' : 'Rolled back')
}
} catch (e) {
const msg = e instanceof Error ? e.message : 'request failed'
@@ -163,16 +209,20 @@ export function Overview({
} finally {
setBusy(null)
onStatusChange()
if (kind !== 'confirm') void loadConfig()
void loadConfig()
// An apply or a rollback is exactly what changes which rules are in force.
void loadReach()
}
},
[flash, loadConfig, onStatusChange],
[flash, loadConfig, loadReach, onStatusChange],
)
// ---- service uptime (PROCESS uptime, not "time since the last apply") ----
const uptime = useUptime(status)
const uptimeText = uptime === null ? '' : fmtDuration(uptime)
const startedAt = status?.started_unix ? fmtDateTime(status.started_unix) : ''
const uptimeText = uptime === null ? '' : fmtDuration(uptime.seconds)
// On YOUR clock, derived from the uptime — never `status.started_unix`, which is
// the router's clock and has no timezone to convert from. See useUptime.
const startedAt = uptime === null ? '' : fmtDateTime(uptime.startedUnix)
// ---- derived display state ----
const g = config?.Globals
@@ -251,24 +301,27 @@ export function Overview({
? `${worstGroup.group} — no answer`
: `${worstGroup.group} — ${worstGroup.dead} down`
const engineVariant: LedVariant = !status
? 'off'
: status.running && status.active
? 'on'
: status.running
? 'amber'
: 'crit'
// One reading for the engine, and it is able to say "stopped": `status.running`
// was a constant `true` on the daemon, so this LED could never go crit and the
// Engine module was green through a process that had failed to start. See
// planeState.engineState.
const engine = engineReadout(status)
const engineVariant: LedVariant = engine.variant
const protection = protectionState(status)
// Configured fail-closed AND actually enforcing it. `none` means nothing is
// installed, so the setting is inert no matter what it says.
const killInEffect = killArmed && status?.plane !== 'none'
// Configured fail-closed, actually enforcing it, or not known — three answers,
// and the third is not folded into the first. See planeState.killSwitchReadout.
const kill = killSwitchReadout(status, g?.KillSwitch)
// Findings that need attention. `info` notes are statements about the config,
// not problems, so they live beside the setting they describe (see findings.ts)
// — keeping this list to things someone could actually act on.
const warnings = attentionFindings(status?.warnings)
const criticalCount = warnings.filter((w) => w.severity === 'critical').length
// The daemon caps the published list at 50 and says so in an `info` note — the
// one channel this page filters away. Carried separately so the list can admit
// it is not the whole list. See findings.ts truncationNote.
const truncated = truncationNote(status?.warnings)
return (
<section className="page" aria-label="Overview">
@@ -291,7 +344,7 @@ export function Overview({
</p>
)}
<Findings warnings={warnings} criticalCount={criticalCount} />
<Findings warnings={warnings} criticalCount={criticalCount} truncated={truncated} />
<div className="grid">
{/* Groups, not nodes: a group is where a dial path is defined, so it is the
@@ -334,14 +387,31 @@ export function Overview({
/>
)}
{/* "N / M in force", the same reading the Routing page shows — never the
count of saved switches. The lamp follows the same rule: a table of
rules none of which are in force routes exactly nothing, and it used
to sit under a green light saying "0 / 7". */}
<Module
name="Routing"
value={String(enabledCount(config?.Rules))}
unit={`/ ${len(config?.Rules)} rules`}
led={{ variant: len(config?.Rules) ? 'on' : 'amber' }}
value={inForce === null ? String(enabledCount(config?.Rules)) : String(inForce)}
unit={
inForce === null
? `/ ${len(config?.Rules)} rules saved`
: `/ ${len(config?.Rules)} in force`
}
led={{
variant:
len(config?.Rules) === 0
? 'amber'
: inForce === null
? 'off'
: inForce === 0
? 'amber'
: 'on',
}}
rows={[
{ k: 'egresses', v: String(len(config?.Egresses)) },
{ k: 'default', v: defaultTarget(config), hot: true },
{ k: 'default', v: defaultTarget(status, config), hot: true },
]}
/>
@@ -381,17 +451,17 @@ export function Overview({
/>
{/* A kill-switch set to fail-closed is only ARMED if something is actually
installed to enforce it. With no plane it is configured but inert, and
saying "ARMED" there would be a false reassurance next to a readout
that says nothing is protected. */}
installed to enforce it, and "we haven't been told" is neither. With no
plane it is configured but inert; with no reading the lamp stays unlit
rather than joining the healthy branch by default. */}
<Module
name="Kill-switch"
value={killInEffect ? 'ARMED' : killArmed ? 'NOT IN EFFECT' : 'OPEN'}
led={{ variant: killInEffect ? 'on' : killArmed ? 'crit' : 'amber' }}
value={kill.value}
led={{ variant: kill.variant }}
rows={[
{ k: 'setting', v: killArmed ? 'fail-closed' : 'fail-open', hot: !killArmed },
...(killArmed && !killInEffect
? [{ k: 'blocking now', v: 'no — nothing installed', hot: true }]
...(kill.blockingNow
? [{ k: 'blocking now', v: kill.blockingNow, hot: kill.hot }]
: [{ k: 'ipv6', v: g?.IPv6 ? 'covered' : 'off' }]),
{ k: 'confirm', v: g?.ConfirmTimeout ? `${g.ConfirmTimeout}s window` : 'no auto-rollback' },
]}
@@ -414,6 +484,7 @@ export function Overview({
unit={status?.version?.includes('-') ? '· ' + status.version.split('-').slice(1).join('-') : ''}
led={{ variant: engineVariant }}
rows={[
{ k: 'process', v: engine.word, hot: engineVariant === 'crit' },
{ k: 'config hash', v: <span className="mono">{short(status?.hash ?? '')}</span> },
// Uptime of the daemon PROCESS. "started" is the moment it came up,
// by the router's clock — not the moment a config was applied.
@@ -431,9 +502,14 @@ export function Overview({
<Button variant="primary" onClick={() => void run('apply')} disabled={busy !== null}>
{busy === 'apply' ? 'Applying…' : 'Apply config'}
</Button>
<Button onClick={() => void run('confirm')} disabled={busy !== null}>
{busy === 'confirm' ? 'Confirming…' : 'Confirm'}
</Button>
{/* A "Confirm" button used to sit here permanently, and pressing it
always printed "Confirmed — auto-rollback cancelled": `apply.Confirm()`
returns nil whether or not a window was ever armed, so the message was
a success report for an event that usually had not happened.
Keeping a config is now offered only while a window is actually open,
and that is announced by the app-wide band directly above this page —
which is where the button lives, beside the countdown it belongs to,
rather than duplicated here. */}
{canRollback && (
<Button onClick={() => void run('rollback')} disabled={busy !== null}>
{busy === 'rollback' ? 'Rolling back…' : 'Rollback'}
@@ -464,11 +540,23 @@ const SECTION_ROUTE: Record<string, Route> = {
rule: 'routing',
ruleset: 'routing',
blocklist: 'dns',
allowlist: 'dns',
resolver: 'dns',
dns_rule: 'dns',
device: 'devices',
chain: 'targets',
group: 'targets',
// A node the generator dropped (unparseable share link, duplicate WireGuard
// key, name colliding with a reserved tag) is reported under `node` — and had
// nowhere to jump to, so the one page that could show it a green toggle was
// also the one page the finding could not reach.
node: 'nodes',
subscription: 'nodes',
egress: 'targets',
inbound: 'networks',
interface: 'networks',
profile: 'profiles',
alert: 'settings',
// The standing note about non-TCP/UDP traffic — its control lives on Networks.
untunnelable: 'networks',
}
@@ -486,11 +574,14 @@ const SECTION_ROUTE: Record<string, Route> = {
function Findings({
warnings,
criticalCount,
truncated,
}: {
warnings: StatusWarning[]
criticalCount: number
/** The daemon's "N further suppressed" note, when the list was capped. */
truncated: StatusWarning | null
}) {
if (warnings.length === 0) return null
if (warnings.length === 0 && !truncated) return null
const rank = { critical: 0, warning: 1, info: 2 } as const
const sorted = [...warnings].sort((a, b) => rank[a.severity] - rank[b.severity])
@@ -500,6 +591,10 @@ function Findings({
<header className="findings-hd">
<h2 className="findings-title">Last apply</h2>
<span className="findings-count mono">
{/* "at least" whenever the list was capped: the counts below it are a
floor, not a total, and the cap drops the least severe FIRST — so
on a config with fifty criticals the thing it drops is a critical. */}
{truncated ? 'at least ' : ''}
{criticalCount > 0
? `${criticalCount} critical · ${warnings.length} total`
: `${warnings.length} note${warnings.length === 1 ? '' : 's'}`}
@@ -538,6 +633,20 @@ function Findings({
</li>
)
})}
{/* The list saying it is not the whole list. Last, because it is about
everything above it — and never filtered out with the other `info`
notes, which is where it used to disappear. */}
{truncated && (
<li className="finding finding--truncated">
<Led variant="amber" />
<div className="finding-copy">
<span className="finding-where mono">list truncated</span>
<span className="finding-msg">
Some findings are missing from this list. {truncated.message}
</span>
</div>
</li>
)}
</ul>
</section>
)
@@ -563,12 +672,20 @@ const NAV_LABEL: Record<Route, string> = {
// for the apply/rollback flow, where the individual flags are the actual
// subject of the page.)
function defaultTarget(config: Model | null): string {
const rules = config?.Rules ?? []
if (rules.length === 0) return '—'
// The highest Order enabled rule is the effective catch-all.
const enabled = rules.filter((r) => r.Enabled)
if (enabled.length === 0) return 'none'
const last = enabled.reduce((a, b) => (b.Order >= a.Order ? b : a))
return last.Target || last.Egress || last.Name
/** Where everything not matched by a rule goes — the engine's route `final`.
*
* Taken from the daemon (status.traffic.default), which reads it off the config
* it is running. The guess this replaced was "the highest-Order enabled rule",
* and that is not what the default is: a rule only becomes the default by having
* NO conditions at all, whatever its Order (model.IsCatchAll), so a specific
* high-Order rule was routinely printed here as the router's default. It also
* described the config on disk rather than the one running, and could not see a
* target that failed to resolve and fell back.
*
* Falls back to the rule count only when the daemon has not reported — never to
* a guess about where traffic goes. */
function defaultTarget(status: Status | null, config: Model | null): string {
const d = status?.traffic?.default
if (d) return d
return len(config?.Rules) === 0 ? '—' : 'not reported'
}
+10 -4
View File
@@ -1,6 +1,6 @@
import './Profiles.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, getInterfaces, putConfig, ApiError } from '../api'
import type { Interface, Model, Profile } from '../api'
@@ -36,6 +36,7 @@ function namesOf(v: unknown): string[] {
// ---- page ------------------------------------------------------------------
export default function Profiles() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -192,9 +193,14 @@ export default function Profiles() {
)
const deleteProfile = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
if (!window.confirm(`Delete profile “${name}”? Its overrides stop applying.`)) return
const ok = await confirm({
label: 'Delete profile',
title: `Delete profile “${name}”?`,
body: 'Its overrides stop applying.',
})
if (!ok) return
const next = profiles.filter((p) => p.Name !== name)
const g =
config.Globals.ActiveProfile === name
@@ -202,7 +208,7 @@ export default function Profiles() {
: config.Globals
void save({ ...config, Profiles: next, Globals: g }, `Deleted ${name}`)
},
[config, profiles, save],
[config, profiles, save, confirm],
)
// ---- expansion (only one profile editor open at a time) -------------------
+86 -6
View File
@@ -58,6 +58,45 @@
color: var(--ink);
}
/* ---- active-profile banner ----
*
* Deliberately NOT the accent plate the save→apply bar wears above. Orange is
* "there is something for you to do" on this faceplate, and an active WAN profile
* is a standing condition, not a pending action. A quiet plate with an amber tag
* reads as "note the state" — and it is the SAME amber the overridden rows below
* carry, so the banner and its rows are visibly one story rather than two
* unrelated oddities. */
.rt-prof-banner {
display: flex;
align-items: flex-start;
gap: 10px;
margin: 0 0 calc(var(--u, 8px) * 2.5);
padding: 10px 14px;
border: 1px solid color-mix(in srgb, var(--amber) 35%, var(--groove));
border-radius: 8px;
background: color-mix(in srgb, var(--amber) 7%, transparent);
font-family: var(--font-sans);
font-size: 12px;
line-height: 1.55;
color: var(--dim);
}
.rt-prof-banner strong {
color: var(--ink);
font-weight: 600;
}
.rt-prof-tag {
flex: none;
margin-top: 1px;
padding: 2px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
font-size: 9px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--amber);
}
/* ---- empty state ---- */
.rt-empty {
padding: calc(var(--u, 8px) * 4) 0 calc(var(--u, 8px) * 3);
@@ -289,6 +328,51 @@
opacity: 0.62;
}
/* ---- a rule the active WAN profile overrides ----
*
* The row itself needs no new paint: an overridden-off rule already wears `.off`
* (it is off, whatever its switch says) and an overridden-on rule wears nothing
* (it is on). What was missing was never colour — it was the sentence naming who
* decided. So this is the per-row twin of the banner and borrows .rt-dead-note's
* type wholesale: same voice, same size, one <p> margin to reset. */
/* The same pill as .rt-badge.dead, so the two override states read as one pair,
* but in accent — a rule the profile forces ON is active, and active is orange on
* this faceplate. The pill is also what keeps it from running into the plain
* "default route · final" badge beside it, where "final on · by profile" read as
* one phrase. */
.rt-badge.prof-on {
padding: 1px 7px;
border: 1px solid var(--accent-soft);
border-radius: 999px;
background: color-mix(in srgb, var(--accent) 10%, transparent);
}
.rt-prof-note {
margin: 0;
}
.rt-prof-note strong {
color: var(--ink);
font-weight: 600;
}
/* Switch + its legend. The caption shows ONLY while a profile overrides the rule,
* and it is what keeps the control honest: the plate says what the router is
* doing, this says the switch is about the saved setting. A legend under the
* control it names is the faceplate's own idiom. */
.rt-switch {
display: inline-flex;
flex-direction: column;
align-items: center;
gap: 3px;
}
.rt-switch-note {
font-family: var(--font-mono);
font-size: 8.5px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--faint);
}
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
.rt-target {
display: inline-flex;
@@ -396,9 +480,6 @@
gap: 5px;
min-width: 0;
}
.rt-field-wide {
grid-column: span 2;
}
.rt-flabel {
font-family: var(--font-mono);
font-size: 9px;
@@ -515,6 +596,8 @@ select.rt-input {
border-color: var(--accent);
box-shadow: 0 1px 0 var(--edge) inset, 0 0 0 1px var(--accent-soft);
}
/* "no matchers" flag in the plate foot — shared by BOTH rule forms (add and
* edit), so the same non-blocking warning reads identically in either. */
.rt-edit-warn {
font-family: var(--font-mono);
font-size: 11.5px;
@@ -837,9 +920,6 @@ select.rt-input {
justify-content: flex-start;
align-self: start;
}
.rt-field-wide {
grid-column: auto;
}
.rt-rs-row {
grid-template-columns: 1fr;
row-gap: 10px;
+487 -187
View File
@@ -1,7 +1,7 @@
import './Routing.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import type { FormEvent, ReactNode } from 'react'
import { Button, CatSuggest, SrcPicker, Toggle } from '../components'
import { Button, CatSuggest, SrcPicker, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
@@ -12,6 +12,7 @@ import {
ApiError,
} from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
import { everyLabel, relFetch } from '../format'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -22,9 +23,7 @@ import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
type RRule = Rule & {
Src?: string[] | null
DstDomain?: string[] | null
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
Proto?: string
Kill?: string
@@ -35,19 +34,57 @@ type RRule = Rule & {
// Minutes east of UTC anchoring the schedule's wall-clock times; captured
// from the editing browser on save (the router has no tzdata). 0 ⇒ UTC.
SchedUTCOffset?: number
// The daemon's unmigrated-rule tripwire (model.Rule.LegacyDst), read-only here.
// Non-empty ⇒ the config STILL carries the schema-v1 `dst_domain`/`dst_ip` that
// schema v2 removed, i.e. `shaterd migrate` never ran or could not commit. Each
// element is the raw `<option>=<value>` text so the panel can quote what was
// found. The daemon holds such a rule disabled; the panel only reports it (it
// is never rendered back to UCI, so a config write from here drains it out).
// Field name is the Go one: model.Rule has no json tags.
LegacyDst?: string[] | null
}
/**
* Everything `Proto` can match, and nothing else. The engine understands two
* transports and exactly ten application protocols its sniffers can name
* (generate/route.go sniffedProtocols); a value outside this set builds a rule
* that is perfectly valid and can never fire — so its traffic quietly falls
* Whether a rule is in force, kept strictly apart from whether it is switched on.
*
* `on` is the EFFECTIVE state — what the router is actually doing — and every mark
* on the row is drawn from it. `profile`/`dir` are set only when the active WAN
* profile is the reason the two differ, so the row can name who overrode the
* saved setting instead of leaving the operator to guess why a switch that reads
* "on" routes nothing.
*/
type RuleForce = {
on: boolean
profile: string | null
dir: 'enabled' | 'disabled' | null
}
/**
* Everything `Proto` can match, and nothing else. A value outside this set builds
* a rule that is perfectly valid and can never fire — so its traffic quietly falls
* through to whatever rule sits below it. That is why this is a closed list and
* not a text box.
*
* Split into two groups because they answer different questions: the transport is
* known the moment a packet arrives, while an app protocol is only known once the
* first bytes have been read and labelled.
* Three groups, because the engine reads them through three different matchers
* (generate/route.go, ruleMatchers) and they answer different questions:
*
* - Transport — the L4 network. Known the moment a packet arrives.
* - Detected protocol — the L7 label a sniffer puts on a connection once its
* first bytes have been read. This group, and ONLY this group, is the engine's
* `sniffedProtocols` set; anything else routed into that matcher is inert.
* - Layer 3 — ICMP. Not a sniffed label: it lands in the emitted rule's
* `network`, never in `protocol` (the sniffers are skipped outright for an
* ICMP flow, so they never report "icmp"). All three spellings are the SAME
* one network; `icmpv4`/`icmpv6` additionally pin `ip_version`, which the
* engine derives from the destination address.
*
* ICMP carries caveats the picker deliberately does not try to enforce, because
* the daemon reports each one against the whole config on apply: it reaches the
* engine only while globals l3_tunnel is on, it has no ports (a port matcher
* beside it can never be satisfied), `icmpv6` also needs globals ipv6 on, and it
* is DROPPED rather than falling through when routed at a target that cannot
* carry layer 3 — i.e. every proxy protocol. Only wireguard/AmneziaWG nodes and
* direct/interface egresses can carry a ping.
*/
const PROTO_TRANSPORT: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'tcp', label: 'TCP' },
@@ -65,10 +102,27 @@ const PROTO_APP: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'rdp', label: 'RDP' },
{ id: 'ntp', label: 'NTP' },
]
const PROTO_VALUES = new Set([...PROTO_TRANSPORT, ...PROTO_APP].map((p) => p.id))
/**
* The family-qualified spellings are offered next to plain `icmp` rather than
* hidden behind it: the engine treats them as first-class and the difference is
* observable (an `ip_version` item on the same rule), so hiding them would leave a
* capability reachable only by hand-editing /etc/config/shater — and would mean
* that anyone who edited such a rule here lost the narrowing on the next save.
*/
const PROTO_L3: ReadonlyArray<{ id: string; label: string }> = [
{ id: 'icmp', label: 'ICMP (ping)' },
{ id: 'icmpv4', label: 'ICMP — IPv4 only' },
{ id: 'icmpv6', label: 'ICMP — IPv6 only' },
]
const PROTO_VALUES = new Set(
[...PROTO_TRANSPORT, ...PROTO_APP, ...PROTO_L3].map((p) => p.id),
)
/** The Proto picker's option list — shared by the inline add row and the editor. */
function ProtoOptions({ value }: { value: string }) {
// The engine lower-cases `Proto` before matching it, so a hand-written `ICMP`
// is a working rule; judge it the same way and flag only what really is inert.
const matches = PROTO_VALUES.has(value.trim().toLowerCase())
return (
<>
<option value="">any</option>
@@ -86,10 +140,18 @@ function ProtoOptions({ value }: { value: string }) {
</option>
))}
</optgroup>
{/* A stored value the engine can't detect is kept and flagged, never
silently rewritten — the rule it belongs to is live right now. */}
<optgroup label="Layer 3">
{PROTO_L3.map((p) => (
<option key={p.id} value={p.id}>
{p.label}
</option>
))}
</optgroup>
{/* A stored value none of the groups spells verbatim is kept and offered as
written, never silently rewritten — the rule it belongs to is live right
now. It is flagged only when the engine cannot match it either. */}
{value !== '' && !PROTO_VALUES.has(value) && (
<option value={value}>{value} — never matches</option>
<option value={value}>{matches ? value : `${value} — never matches`}</option>
)}
</>
)
@@ -97,7 +159,6 @@ function ProtoOptions({ value }: { value: string }) {
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
const byOrder = (a: RRule, b: RRule): number => a.Order - b.Order
const csv = (s: string): string[] => s.split(',').map((x) => x.trim()).filter(Boolean)
// --- ruleset helpers --------------------------------------------------------
// A `config ruleset` (api.ts Ruleset) is a named domain/ipcidr list a rule
@@ -183,48 +244,113 @@ function errMsg(e: unknown): string {
// --- remote-list freshness (feedback #9) ------------------------------------
// A url-source ruleset re-fetches on a cadence; the engine reports when it last
// pulled and how many rules the list holds. Match a ruleset to its status by the
// engine tag `rs-<name>`.
// engine tag `rs-<name>`. The two readings (relFetch / everyLabel) live in
// format.ts because the DNS page shows the same ones for blocklists.
/** "updated 3h ago" / "never updated" for a remote list's last fetch. */
function relFetch(iso: string): string {
if (!iso) return 'never updated'
const t = Date.parse(iso)
if (Number.isNaN(t)) return 'never updated'
const s = Math.max(0, Math.floor((Date.now() - t) / 1000))
if (s < 45) return 'updated just now'
const m = Math.floor(s / 60)
if (m < 60) return `updated ${m}m ago`
const h = Math.floor(m / 60)
if (h < 24) return `updated ${h}h ago`
const d = Math.floor(h / 24)
return `updated ${d}d ago`
}
/** "every 24h" for an auto-update cadence in seconds ("" when there is none). */
function everyLabel(sec: number): string {
if (!sec || sec <= 0) return ''
if (sec % 3600 === 0) {
const h = sec / 3600
if (h < 48) return `every ${h}h`
if (sec % 86400 === 0) return `every ${sec / 86400}d`
return `every ${h}h`
}
if (sec % 60 === 0) return `every ${sec / 60}m`
return `every ${sec}s`
}
/** A rule with no matcher of any kind is the effective catch-all (route Final). */
/** A rule with no matcher of any kind is the effective catch-all (route Final).
* Mirrors model.IsCatchAll on the daemon side — the two must agree or the
* "never applies" badge lands on a different row than the apply warning.
*
* AN UNMIGRATED RULE IS NEVER A CATCH-ALL, and that is the first thing checked
* here, exactly as on the Go side (model/reachability.go). When LegacyDst is
* non-empty the rule's destination is still written in the schema-v1 options
* the parser no longer reads, so its lack of matchers means "the destination is
* unreadable", not "matches everything" — reading it the other way is precisely
* what turned an uncommitted `shaterd migrate` into route Final for the whole
* router. The daemon holds such a rule disabled and reports false here; if this
* copy disagreed, the panel would paint the row "default route · final" while
* the daemon routes nothing through it. */
function isCatchAll(r: RRule): boolean {
if (len(r.LegacyDst) > 0) return false
return (
len(r.Src) === 0 &&
len(r.DstDomain) === 0 &&
len(r.DstRuleset) === 0 &&
len(r.DstIP) === 0 &&
!(r.DstPort && r.DstPort.trim()) &&
!(r.Proto && r.Proto.trim())
)
}
/** isCatchAll's twin for a form still being edited: the live fields of either
* rule form with no matcher left in them.
*
* Such a rule is not "matches all" in the ordinary sense — the engine emits it
* as route.Final, and among several the LAST one in rule order owns it. So what
* saving one actually does depends on what is last right now, and there are
* three cases:
* - no rules at all, or the last rule is a conditional one → the new rule
* lands last and TAKES the default, silently retargeting every otherwise
* unmatched flow (e.g. the whole LAN to `direct`, past the tunnel);
* - the last rule is already a catch-all → nextOrder() deliberately inserts
* the new one BEFORE it (and bumps the old one up), so the existing default
* keeps route.Final and the new rule is dead on arrival — the daemon
* reports it as shadowed and the row renders as such.
* Both outcomes are worth a warning, and neither form knows which it will be
* (the add form has no rule list), so the shared text says only what is certain:
* the rule has no matchers. It is a legal configuration either way, so neither
* form blocks it — they warn, from this one predicate, so the flag cannot drift
* out of sync between add and edit. */
function formHasNoMatchers(f: {
src: string[]
port: string
rulesets: string[]
proto: string
}): boolean {
return (
f.src.length === 0 && f.port.trim() === '' && f.rulesets.length === 0 && f.proto.trim() === ''
)
}
/**
* What happens to the network when the default route stops being emitted —
* whether it is deleted or merely switched off (the engine emits neither).
*
* The old text was one sentence for every rule: "Traffic it matched will fall
* through to the next rule." For an ordinary rule that is true. For the catch-all
* there IS no next rule, and what happens instead is decided by the kill-switch:
* generate/route.go sets `final := tagBlock` and only `kill_switch=open` swaps
* that for `direct`. So removing the default either takes the whole network
* offline or puts the whole network on the naked WAN — and the page said "falls
* through to the next rule" for both, on a row it had already badged
* "DEFAULT ROUTE · FINAL".
*
* `successor` is the rule that would inherit route.Final instead (a config can
* carry more than one conditionless rule; the last one wins). When there is one,
* nothing is lost — the honest warning is that the destination changes.
*/
function defaultRouteConsequence(
killSwitch: string,
successor: { name: string; order: number; target: string } | null,
): ReactNode {
if (successor) {
return (
<>
This is the router’s <strong>default route</strong> — everything no other rule matches
follows it. Remove it and “{successor.name}” (order {successor.order}) has no conditions
either, so it takes over: unmatched traffic goes to{' '}
<strong className="mono">{successor.target}</strong> instead.
</>
)
}
if (killSwitch === 'open') {
return (
<>
This is the router’s <strong>default route</strong> — everything no other rule matches
follows it, and no other rule matches everything. With the kill-switch set to{' '}
<strong>fail-open</strong>, unmatched traffic then leaves through your normal internet
connection with your real address — unproxied and unfiltered.
</>
)
}
return (
<>
This is the router’s <strong>default route</strong> — everything no other rule matches
follows it, and no other rule matches everything. With the kill-switch set to{' '}
<strong>fail-closed</strong>, unmatched traffic is then <strong>blocked</strong>: devices on
your network lose the internet until you add a default back.
</>
)
}
/** Effective routing target for a rule (Target wins; a bare Egress is a target too). */
function effectiveTarget(r: RRule): string {
if (r.Target && r.Target.trim()) return r.Target.trim()
@@ -263,16 +389,19 @@ interface TargetGroups {
nodes: TargetOpt[] // node:<n> (huge — rendered last)
}
// Free-text destination matchers offered by the ADD form. Domains are NOT one of
// them (the user's call): domain matching goes through named rulesets — that's
// what they exist for. 'none' = the rule matches by rulesets/source/proto alone.
// (Legacy rules that already carry DstDomain stay editable in the edit form.)
type MatchKind = 'none' | 'ip' | 'port'
// The add form's fields. WHERE traffic is going is a ruleset choice and nothing
// else — a rule has no inline domain or address list any more, so the old
// Match-kind picker (rulesets / ip / port) collapsed into a plain Port field
// beside the ruleset picker. The cost is real and accepted: routing a single
// domain is no longer done here — you leave for the Rulesets panel, create the
// list, fill it, and come back to check it. A "create a list from here" shortcut
// was proposed and rejected (DECISIONS.md D21): a second place to author a list
// is a second place for its entry semantics and duplicate-name rules to drift,
// which is the exact thing D21 removed.
interface AddForm {
name: string
src: string[]
matchKind: MatchKind
matchValue: string
port: string
rulesets: string[]
proto: string
target: string
@@ -284,8 +413,7 @@ interface AddForm {
const EMPTY_FORM: AddForm = {
name: '',
src: [],
matchKind: 'none',
matchValue: '',
port: '',
rulesets: [],
proto: '',
target: 'direct',
@@ -318,6 +446,7 @@ const browserTZName = (): string => {
}
export default function Routing() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
const [actionError, setActionError] = useState<string | null>(null)
@@ -435,7 +564,7 @@ export default function Routing() {
}, [config])
/**
* The verdict for one rule, or null when it can fire.
* The daemon's verdict for one rule, or null when we have none that describes it.
*
* Verdicts are fetched separately from the config, so between an optimistic edit
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
@@ -443,18 +572,53 @@ export default function Routing() {
* badge on a working rule: a mismatch means the verdict is not about this row,
* and no badge is the honest answer.
*/
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const verdictOf = useCallback(
(r: RRule): RuleReach | null => {
const i = modelIndex.get(r)
if (i === undefined) return null
const v = reach.get(i)
if (!v || !v.unreachable || !v.shadowed_by) return null
if (v.name !== r.Name || v.order !== r.Order) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
if (!v || v.name !== r.Name || v.order !== r.Order) return null
return v
},
[modelIndex, reach],
)
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const v = verdictOf(r)
if (!v || !v.unreachable || !v.shadowed_by) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
},
[verdictOf],
)
/**
* Whether a rule is IN FORCE, and who decided that — the two states this page
* used to conflate.
*
* `Rule.Enabled` from /api/config is the DESIRED state: what the operator saved,
* what the switch edits, what gets PUT back. The active WAN profile can override
* it in either direction, and then the desired state is no longer what the router
* is doing. Drawing the row from `Enabled` is what let a config with two rules
* `enabled '1'` show two live switches while the engine ran one chain.
*
* With no verdict — an older daemon, a stopped one, or one still describing the
* previous config — the desired state is all we know, so the row falls back to it
* and claims no profile rather than inventing one. The `typeof` guard is for the
* older daemon specifically: it answers without `effective_enabled` at all, and
* reading `undefined` as false would gray out every rule on the page.
*/
const forceOf = useCallback(
(r: RRule): RuleForce => {
const v = verdictOf(r)
if (!v || typeof v.effective_enabled !== 'boolean') {
return { on: !!r.Enabled, profile: null, dir: null }
}
return { on: v.effective_enabled, profile: v.overridden_by ?? null, dir: v.override ?? null }
},
[verdictOf],
)
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
const rulesets = useMemo<Ruleset[]>(
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
@@ -555,13 +719,17 @@ export default function Routing() {
// Delete the ruleset AND strip its name from any rule that referenced it, so no
// rule is left pointing at a matcher that no longer exists (one atomic persist).
const deleteRuleset = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const used = rulesetUsage.get(name) ?? 0
const warn = used
? `Delete ruleset "${name}"? It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
: `Delete ruleset "${name}"?`
if (!window.confirm(warn)) return
const ok = await confirm({
label: 'Delete ruleset',
title: `Delete ruleset "${name}"?`,
body: used
? `It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
: undefined,
})
if (!ok) return
const nextRulesets = rulesets.filter((r) => r.Name !== name)
const nextRules = rules.map((r) => {
const cur = r.DstRuleset ?? []
@@ -572,20 +740,64 @@ export default function Routing() {
`ruleset ${name} deleted`,
)
},
[config, rules, rulesets, rulesetUsage, persist],
[config, rules, rulesets, rulesetUsage, persist, confirm],
)
/**
* Is this rule the one the ENGINE uses as route.Final right now, and in force?
*
* Both halves matter. A conditionless rule the daemon reports as shadowed owns
* nothing (the row already says "never applies"), and one that is switched off
* is not being emitted either — removing either changes no traffic, so neither
* earns a warning.
*/
const isLiveDefault = useCallback(
(r: RRule): boolean => isCatchAll(r) && shadowOf(r) === null && forceOf(r).on,
[shadowOf, forceOf],
)
/** Which rule would inherit route.Final if `name` stopped being emitted: the
* LAST remaining conditionless, switched-on rule. null when there is none. */
const successorDefault = useCallback(
(name: string): { name: string; order: number; target: string } | null => {
for (let i = rules.length - 1; i >= 0; i--) {
const r = rules[i]
if (r.Name === name) continue
if (r.Enabled && isCatchAll(r)) {
return { name: r.Name, order: r.Order, target: effectiveTarget(r) }
}
}
return null
},
[rules],
)
const killSwitch = (config?.Globals?.KillSwitch ?? 'closed') === 'open' ? 'open' : 'closed'
const onToggle = useCallback(
(name: string) => {
async (name: string) => {
const target = rules.find((r) => r.Name === name)
if (!target) return
const nextState = !target.Enabled
// Switching the default route OFF is the same event as deleting it — the
// generator emits only enabled rules — so it asks the same question. It used
// to ask nothing at all, which made the least reversible control on the page
// the only one with no confirmation.
if (!nextState && isLiveDefault(target)) {
const ok = await confirm({
label: 'Turn off default route',
title: `Turn off the default route “${name}”?`,
body: defaultRouteConsequence(killSwitch, successorDefault(name)),
confirmLabel: 'Turn it off',
})
if (!ok) return
}
commitRules(
rules.map((r) => (r.Name === name ? { ...r, Enabled: nextState } : r)),
`${name} ${nextState ? 'enabled' : 'disabled'}`,
)
},
[rules, commitRules],
[rules, commitRules, confirm, isLiveDefault, killSwitch, successorDefault],
)
const onMove = useCallback(
@@ -612,17 +824,33 @@ export default function Routing() {
)
const onDelete = useCallback(
(name: string) => {
if (!window.confirm(`Delete rule "${name}"? Traffic it matched will fall through to the next rule.`)) return
async (name: string) => {
const target = rules.find((r) => r.Name === name)
if (!target) return
// The catch-all has no "next rule" to fall through to — see
// defaultRouteConsequence. Every other rule keeps the plain sentence.
const isDefault = isLiveDefault(target)
const ok = await confirm({
label: isDefault ? 'Delete default route' : 'Delete rule',
title: isDefault ? `Delete the default route “${name}”?` : `Delete rule “${name}”?`,
body: isDefault
? defaultRouteConsequence(killSwitch, successorDefault(name))
: 'Traffic it matched will fall through to the next rule.',
})
if (!ok) return
commitRules(
rules.filter((r) => r.Name !== name),
`${name} deleted`,
)
},
[rules, commitRules],
[rules, commitRules, confirm, isLiveDefault, killSwitch, successorDefault],
)
// Insert a new rule just above the catch-all (so a specific rule can actually match).
// isCatchAll() is false for an unmigrated rule, which is the right answer here too:
// such a rule is held disabled and owns no default route, so there is nothing to
// insert ahead of — the new rule simply goes last, where a rule with no matchers
// does become the default.
const nextOrder = useCallback((): { order: number; bumpCatchAll?: { name: string; order: number } } => {
if (rules.length === 0) return { order: 10 }
const last = rules[rules.length - 1]
@@ -674,18 +902,15 @@ export default function Routing() {
return
}
setFormError(null)
const mv = form.matchValue.trim()
const rule: RRule = {
Name: name,
Enabled: true,
Order: 0,
Src: form.src,
// Domains are matched via rulesets only — the add form has no free-text
// domain matcher by design.
DstDomain: [],
// Destination = rulesets, always. Domains and addresses live in a
// `config ruleset` so one list serves every rule that needs it.
DstRuleset: form.rulesets,
DstIP: form.matchKind === 'ip' ? csv(mv) : [],
DstPort: form.matchKind === 'port' ? mv : '',
DstPort: form.port.trim(),
Proto: form.proto,
Target: form.target,
Egress: '',
@@ -756,7 +981,14 @@ export default function Routing() {
)
}
const enabledCount = rules.filter((r) => r.Enabled).length
// One force verdict per displayed rule, computed once and handed down — the row,
// the counter and the banner must all be reading the SAME answer.
const force = rules.map((r) => forceOf(r))
// EFFECTIVE, not configured. A counter that added up saved switches said "2 / 2
// active" for a config the router was running one rule of.
const enabledCount = force.filter((f) => f.on).length
const overridden = force.filter((f) => f.profile !== null)
const overrideProfile = overridden[0]?.profile ?? null
return (
<section className="page" aria-label="Routing rules">
@@ -765,12 +997,28 @@ export default function Routing() {
Rules run top to bottom on the bus — the <strong>first match wins</strong>. Traffic that
reaches the bottom follows the default route.
</p>
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules active`}>
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules in force`}>
{enabledCount}
<small> / {rules.length} active</small>
<small> / {rules.length} in force</small>
</span>
</div>
{/* Said once at the top, so the per-row badges below read as consequences of
one thing rather than as N unrelated oddities. Only shown when a profile
actually changed something: a router that uses no profiles, or one whose
profile agrees with every saved switch, gets no banner at all. */}
{overrideProfile && (
<p className="rt-prof-banner" role="status">
<span className="rt-prof-tag mono">profile</span>
<span>
<strong className="mono">{overrideProfile}</strong> is the active WAN profile and is
overriding {overridden.length === 1 ? '1 rule' : `${overridden.length} rules`} below. The
switches keep showing what you saved; the rows show what the router is running. Change
which rules a profile forces on the <strong>Profiles</strong> page.
</span>
</p>
)}
{actionError && (
<p className="page-error" role="alert">
{actionError}
@@ -789,7 +1037,7 @@ export default function Routing() {
{rules.length === 0 ? (
<div className="rt-empty">
<p>No rules — all traffic follows the default route.</p>
<p className="rt-empty-sub">Add a rule below to steer a domain, address, or port.</p>
<p className="rt-empty-sub">Add a rule below to steer a destination list, source, or port.</p>
</div>
) : (
<ol className="rt-list" aria-label="Routing rules in first-match order">
@@ -818,6 +1066,7 @@ export default function Routing() {
busy={saving}
editingOther={editingRule !== null}
shadow={shadowOf(r)}
force={force[i]}
onEdit={onEditRule}
onToggle={onToggle}
onMove={onMove}
@@ -868,6 +1117,7 @@ function RuleRow({
busy,
editingOther,
shadow,
force,
onEdit,
onToggle,
onMove,
@@ -880,6 +1130,9 @@ function RuleRow({
editingOther: boolean
/** Set when the daemon reports this rule can never fire; null when it can. */
shadow: { by: string; byOrder: number; reason: string } | null
/** Whether the rule is IN FORCE, and which profile decided that (see RuleForce).
* Every mark on this row comes from here; `rule.Enabled` drives only the switch. */
force: RuleForce
onEdit: (name: string) => void
onToggle: (name: string) => void
onMove: (name: string, dir: 'up' | 'down') => void
@@ -895,17 +1148,26 @@ function RuleRow({
// matched above" line): those are the claim that made two `default` rules
// indistinguishable in the first place.
const dead = shadow !== null
// The unmigrated-rule tripwire (see RRule.LegacyDst / model.IsCatchAll): this
// rule's destination is still in the removed schema-v1 options, so the daemon
// holds it disabled. It is NOT an ordinary disabled rule — nobody switched it
// off — so it gets the same "wired but not connected" amber treatment as a
// shadowed rule, plus a badge and a line saying what to run. isCatchAll()
// already refuses to call it the default, so `final` marks cannot land here.
const legacyDst = (rule.LegacyDst ?? []).filter(Boolean)
const unmigrated = legacyDst.length > 0
const inert = dead || unmigrated
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
const cls = [
'rt-rule',
rule.Enabled ? '' : 'off',
isDefault ? 'final' : '',
dead ? 'dead' : '',
]
// Dimmed by the EFFECTIVE state, never by the saved one. A rule the active
// profile switched off is not in force, and the row has to read that way even
// though its switch — which edits the saved setting — is still on.
const cls = ['rt-rule', force.on ? '' : 'off', isDefault ? 'final' : '', inert ? 'dead' : '']
.filter(Boolean)
.join(' ')
// What the switch says, spelled out, for the moment the two disagree.
const savedState = rule.Enabled ? 'on' : 'off'
return (
<li className={cls}>
@@ -919,7 +1181,7 @@ function RuleRow({
>
▲
</button>
<span className={isDefault ? 'rt-ord final' : dead ? 'rt-ord dead' : 'rt-ord'}>
<span className={isDefault ? 'rt-ord final' : inert ? 'rt-ord dead' : 'rt-ord'}>
{isDefault ? '·' : rule.Order}
</span>
<button
@@ -937,10 +1199,26 @@ function RuleRow({
<div className="rt-head">
<span className="rt-name">{rule.Name}</span>
{isDefault && <span className="rt-badge">default route · final</span>}
{unmigrated && <span className="rt-badge dead">held off · not migrated</span>}
{dead && <span className="rt-badge dead">never applies</span>}
{/* Amber for the rule the profile switched OFF (warn semantics: wired but
not connected), accent for the one it switched ON — orange is the
faceplate's active state, and a force-enabled rule is exactly that. */}
{force.dir === 'disabled' && <span className="rt-badge dead">off · by profile</span>}
{force.dir === 'enabled' && <span className="rt-badge prof-on">on · by profile</span>}
</div>
<div className="rt-match">
{dead ? (
{unmigrated ? (
// Why the rule is off and what fixes it. Same voice as the shadow note:
// state, cause, one command. The daemon says the same thing through the
// apply warnings (model.ValidateRules); this puts it on the row it is about.
<span className="rt-dead-note">
Destination still written the old way (<span className="mono">{legacyDst.join(', ')}</span>
) — this config was never migrated, so shaterd cannot read where this rule sends
traffic and holds it disabled. Run <span className="mono">shaterd migrate</span> on the
router to turn those entries into a ruleset, then enable the rule again.
</span>
) : dead ? (
// The badge says it never fires; this line says what beat it and what to
// do. Visible text, not a tooltip — the operator has to be able to find
// the other rule, and two rows can carry the same name.
@@ -955,6 +1233,22 @@ function RuleRow({
<Matchers rule={rule} />
)}
</div>
{/* Added BELOW the matchers, not instead of them: the rule's conditions are
still worth reading — the operator is deciding whether to change the
profile or the rule.
Two clauses only. The banner at the top of the page already carries the
general explanation and the way to change it, and a profile that
overrides several rules would otherwise repeat that paragraph on every
one of them. What is left is the part only this row can say: whether it
is in force, and what its own switch is showing instead. */}
{force.profile && (
<p className="rt-dead-note rt-prof-note">
{force.dir === 'disabled' ? 'Not in force' : 'In force'} — profile{' '}
<strong className="mono">{force.profile}</strong> switches this rule{' '}
{force.dir === 'disabled' ? 'off' : 'on'}. The switch still reads{' '}
<strong>{savedState}</strong>: that is the saved setting.
</p>
)}
</div>
<div className={`rt-target ${tone}`} title={`target: ${target}`}>
@@ -974,12 +1268,40 @@ function RuleRow({
>
Edit
</button>
<Toggle
pressed={rule.Enabled}
onChange={() => onToggle(rule.Name)}
label={`${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`}
disabled={frozen}
/>
{/* An unmigrated rule cannot be switched on from here, and the switch says
so rather than pretending: the daemon holds it disabled, but a config
write from the panel DROPS the unreadable legacy options (render.go
emits neither), so enabling it here would save a live rule with no
destination left at all — the catch-all this tripwire exists to
prevent. `shaterd migrate` clears LegacyDst and the switch comes back. */}
{/* The switch edits the SAVED setting and nothing else, so it keeps showing
rule.Enabled even while the active profile forces the opposite. Mirroring
the effective state here would be worse than the bug it replaces: the
operator would flip a switch that was never theirs, and the PUT would
write the profile's decision into UCI as if they had chosen it. The row
above says what the router is doing; the "saved" caption says what this
control is for. */}
<span className="rt-switch">
<Toggle
pressed={rule.Enabled}
onChange={() => onToggle(rule.Name)}
label={
unmigrated
? `Rule ${rule.Name} is held disabled until the config is migrated`
: force.profile
? `Saved setting for rule ${rule.Name} is ${savedState}; profile ${force.profile} is forcing it ${
force.dir === 'disabled' ? 'off' : 'on'
}. This switch changes the saved setting only.`
: `${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`
}
disabled={frozen || unmigrated}
/>
{force.profile && (
<span className="rt-switch-note" aria-hidden="true">
saved
</span>
)}
</span>
<button
type="button"
className="rt-del"
@@ -1012,9 +1334,7 @@ function Matchers({ rule }: { rule: RRule }): ReactNode {
)
}
listChip('src', rule.Src, 'src')
listChip('dns', rule.DstDomain, 'dom')
listChip('ruleset', rule.DstRuleset, 'rs')
listChip('ip', rule.DstIP, 'ip')
if (rule.DstPort && rule.DstPort.trim()) {
chips.push(
<span className="rt-chip" key="port">
@@ -1130,7 +1450,16 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
)
}
/** The dst_ruleset checkbox group. Renders nothing when no rulesets exist. */
/**
* The destination picker: which rulesets this rule matches (dst_ruleset).
*
* Checkboxes and nothing else. This is the ONLY way a rule names a destination,
* so it renders even when the config has no lists yet — an empty picker that says
* where lists come from is the honest answer, and hiding it would leave the rule
* form with no destination control at all. Building and filling a list is the
* Rulesets panel's job, deliberately kept out of the rule editor so a list is
* created in exactly one place.
*/
function RulesetPicker({
options,
selected,
@@ -1142,24 +1471,26 @@ function RulesetPicker({
busy: boolean
onToggle: (name: string) => void
}): ReactNode {
if (options.length === 0) return null
return (
<div className="rt-rsel">
<span className="rt-flabel">Match rulesets — dst_ruleset</span>
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
<span className="rt-flabel">Destination — dst_ruleset</span>
{options.length > 0 && (
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
)}
<p className="rt-rsel-hint">
The rule also matches any traffic in the checked list(s). Combine with a domain, address, or
port, or use a ruleset on its own.
{options.length === 0
? 'No rulesets yet. Add one under Rulesets below, then come back and check it here — a rule matches a destination through a ruleset only.'
: 'The rule matches traffic in ANY checked list. Narrow it further with a source, port or protocol.'}
</p>
</div>
)
@@ -1285,7 +1616,14 @@ function AddRule({
set('rulesets', form.rulesets.includes(n) ? form.rulesets.filter((x) => x !== n) : [...form.rulesets, n])
const toggleDay = (d: string) =>
set('schedDays', form.schedDays.includes(d) ? form.schedDays.filter((x) => x !== d) : [...form.schedDays, d])
const matchPlaceholder = form.matchKind === 'ip' ? '10.0.0.0/8, 100.64.0.0/10' : '443, 8080-8090'
// Same flag as the edit form, from the same predicate: no matcher = this rule
// asks to be route.Final. Whether it wins depends on what is last — it takes
// the default when there is no catch-all yet (or the last rule is conditional),
// and lands dead when there is one, because nextOrder() inserts ahead of it.
// Both are worth flagging, so the text states the fact and not the outcome.
// Shown, never blocking: a default route is a legal thing to write.
const noMatchers = formHasNoMatchers(form)
return (
<form className="rt-add" onSubmit={onSubmit} aria-label="Add a routing rule">
@@ -1318,32 +1656,17 @@ function AddRule({
</label>
<label className="rt-field">
<span className="rt-flabel">Match</span>
<select
<span className="rt-flabel">Port(s)</span>
<input
className="rt-input mono"
value={form.matchKind}
onChange={(e) => set('matchKind', e.target.value as MatchKind)}
>
<option value="none">rulesets only</option>
<option value="ip">ip / cidr</option>
<option value="port">port</option>
</select>
value={form.port}
onChange={(e) => set('port', e.target.value)}
placeholder="443, 8080-8090"
autoComplete="off"
spellCheck={false}
/>
</label>
{form.matchKind !== 'none' && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">{form.matchKind === 'port' ? 'Port(s)' : 'Address(es)'}</span>
<input
className="rt-input mono"
value={form.matchValue}
onChange={(e) => set('matchValue', e.target.value)}
placeholder={matchPlaceholder}
autoComplete="off"
spellCheck={false}
/>
</label>
)}
<label className="rt-field">
<span className="rt-flabel">Proto</span>
<select
@@ -1390,6 +1713,9 @@ function AddRule({
<Button type="submit" variant="primary" disabled={busy}>
{busy ? 'Saving…' : 'Add rule'}
</Button>
{noMatchers && !error && (
<span className="rt-edit-warn">no matchers — matches everything</span>
)}
{error && (
<span className="rt-add-error" role="alert">
{error}
@@ -1401,11 +1727,11 @@ function AddRule({
}
// --- edit-a-rule plate (inline, replaces the row it edits) ------------------
// Unlike AddRule, editing exposes all three destination matchers at once
// (Domain(s) / IP-CIDR(s) / Port) rather than a single Match picker — a real rule
// can carry several matcher kinds simultaneously and none may be silently dropped.
// The full original rule is spread into the result on save, so Order / Enabled /
// Kill / Egress (and anything else off-form) survive untouched.
// Same fields as AddRule, on purpose: a rule carries exactly one destination
// mechanism (rulesets) plus port/proto/source, so there is nothing an edit can
// reveal that the add form hides. The full original rule is spread into the
// result on save, so Order / Enabled / Kill / Egress (and anything else off-form)
// survive untouched.
function RuleEditForm({
initial,
names,
@@ -1425,8 +1751,6 @@ function RuleEditForm({
}) {
const [name, setName] = useState(initial.Name)
const [src, setSrc] = useState<string[]>([...(initial.Src ?? [])])
const [domain, setDomain] = useState((initial.DstDomain ?? []).join(', '))
const [ip, setIp] = useState((initial.DstIP ?? []).join(', '))
const [port, setPort] = useState(initial.DstPort ?? '')
const [proto, setProto] = useState(initial.Proto ?? '')
const [target, setTarget] = useState(effectiveTarget(initial))
@@ -1442,14 +1766,15 @@ function RuleEditForm({
const toggleDay = (d: string) =>
setSchedDays((cur) => (cur.includes(d) ? cur.filter((x) => x !== d) : [...cur, d]))
// This rule's destination may still be in the schema-v1 options (RRule.LegacyDst).
// The form cannot show or edit them — they are not fields any more — and saving
// writes the model back through render.go, which does not emit them. So a save
// here silently DISCARDS that destination. Say so before it happens; the fix is
// `shaterd migrate`, which converts them into a ruleset this form can check.
const legacyDst = (initial.LegacyDst ?? []).filter(Boolean)
// A rule with no matcher of any kind is a catch-all — legal, but worth flagging.
const noMatchers =
src.length === 0 &&
csv(domain).length === 0 &&
csv(ip).length === 0 &&
port.trim() === '' &&
rulesets.length === 0 &&
proto.trim() === ''
const noMatchers = formHasNoMatchers({ src, port, rulesets, proto })
const submit = (e: FormEvent) => {
e.preventDefault()
@@ -1471,8 +1796,6 @@ function RuleEditForm({
...initial,
Name: nm,
Src: src,
DstDomain: csv(domain),
DstIP: csv(ip),
DstPort: port.trim(),
DstRuleset: rulesets,
Proto: proto,
@@ -1491,7 +1814,15 @@ function RuleEditForm({
<form className="rt-add rt-edit-form" onSubmit={submit} aria-label={`Edit rule ${initial.Name}`}>
<div className="rt-add-hd">
<span className="rt-add-title">Edit {initial.Name}</span>
<span className="rt-add-sub">Empty a field to drop that matcher. Save, then Apply.</span>
<span className="rt-add-sub">
Uncheck a list or clear a field to drop that matcher. Save, then Apply.
</span>
{legacyDst.length > 0 && (
<span className="rt-edit-warn">
not migrated — saving discards its old destination ({legacyDst.join(', ')}); run
shaterd migrate first
</span>
)}
</div>
<div className="rt-fields">
@@ -1521,37 +1852,6 @@ function RuleEditForm({
/>
</label>
{/* Domains are matched via rulesets by design — this legacy field only
appears when the rule ALREADY carries free-text domains, so they
stay visible and clearable rather than silently preserved. */}
{(initial.DstDomain ?? []).length > 0 && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">Domain(s) — legacy</span>
<input
className="rt-input mono"
value={domain}
onChange={(e) => setDomain(e.target.value)}
placeholder="youtube.com, *.googlevideo.com"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
)}
<label className="rt-field rt-field-wide">
<span className="rt-flabel">IP / CIDR(s)</span>
<input
className="rt-input mono"
value={ip}
onChange={(e) => setIp(e.target.value)}
placeholder="10.0.0.0/8, 100.64.0.0/10"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
<label className="rt-field">
<span className="rt-flabel">Port(s)</span>
<input
+58 -4
View File
@@ -1,7 +1,8 @@
import './Settings.css'
import { useCallback, useEffect, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import { Button, Led, Select, Toggle } from '../components'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { AlertsSection } from './Alerts'
import { apply as apiApply, downloadLog, getConfig, putConfig, ApiError } from '../api'
import type { Globals, LogRange, Model } from '../api'
@@ -125,6 +126,7 @@ const STATS_BACKENDS: ReadonlyArray<{ value: string; label: string }> = [
// ---- page ------------------------------------------------------------------
export default function Settings() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -257,6 +259,48 @@ export default function Settings() {
const groupHealthOn = globals?.GroupHealth !== false
const killSwitch = globals?.KillSwitch === 'open' ? 'open' : 'closed'
/**
* The master switch, which is the most destructive control in the panel and was
* the only one that asked nothing.
*
* Turning it off is not "pausing the proxy": apply.go runs Teardown() — the nft
* table goes, the policy routing goes, `plane` becomes `none`. The kill-switch
* does not save you, because a kill-switch is a rule in a table that no longer
* exists. Everything on the LAN then leaves through the plain WAN, unproxied and
* unfiltered. Deleting a rule-set asked for confirmation; this did not.
*
* Turning it back ON is not destructive and is not gated.
*/
const toggleService = useCallback(
async (on: boolean) => {
if (!on) {
const ok = await confirm({
label: 'Turn off the service',
title: 'Turn the proxy engine off?',
body: (
<>
This tears the whole data plane down — the firewall table, the policy routing and the
DNS interception are removed, not paused. Nothing is proxied, filtered or blocked, and
every device leaves through your normal internet connection with its real address.{' '}
{killSwitch === 'closed' ? (
<>
The kill-switch does not hold here: with nothing installed there is nothing left
to block with.
</>
) : (
<>The kill-switch is already open, so nothing changes about that.</>
)}
</>
),
confirmLabel: 'Turn it off',
})
if (!ok) return
}
setGlobal('Enabled', on, on ? 'Engine enabled' : 'Engine disabled')
},
[confirm, killSwitch, setGlobal],
)
const killNote =
killSwitch === 'open'
? 'Fail-open — if the engine stops, traffic falls back to the direct WAN. Stays online, but unprotected.'
@@ -296,10 +340,13 @@ export default function Settings() {
<div className="set-groups">
{/* ---- SERVICE ---- */}
<Group title="Service" count={globals?.Enabled ? 'enabled' : 'disabled'}>
<Field label="Proxy engine" note="Master on/off for the whole appliance.">
<Field
label="Proxy engine"
note="Master on/off for the whole appliance. Off removes the firewall table and the policy routing — every device goes out directly, with no kill-switch to catch it."
>
<Toggle
pressed={globals?.Enabled ?? false}
onChange={(on) => setGlobal('Enabled', on, on ? 'Engine enabled' : 'Engine disabled')}
onChange={(on) => void toggleService(on)}
label={globals?.Enabled ? 'Disable proxy engine' : 'Enable proxy engine'}
size="md"
disabled={busy || !ready}
@@ -360,7 +407,7 @@ export default function Settings() {
<Field
label="Log level"
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level."
note="Verbosity of the daemon log. “none” silences the engine and drops the control-plane to panic-only — a turn-down, not a true off: even warnings and errors are hidden. The toggles below decide where whatever is emitted gets written; turning both off is the only full silence. Failures still raise alerts regardless of this level — set up where they go in the Alerts section below."
>
<Select
value={globals?.LogLevel || 'warning'}
@@ -574,6 +621,13 @@ export default function Settings() {
</Field>
</Group>
{/* ---- ALERTS ---- */}
{/* Extracted from the DNS page — out-of-band notifications belong with
the appliance-wide knobs, next to the log level whose note points
here. Renders its own section header (same plate as a Group); all
writes go through `save`, so the dirty banner and toast stay one. */}
<AlertsSection config={config} busy={busy} loading={loading} onSave={save} />
{/* ---- STATISTICS & LOGGING ---- */}
<Group
title="Statistics &amp; logging"
+250
View File
@@ -704,6 +704,13 @@
.tg-test--bad .tg-test-msg {
color: var(--crit);
}
/* "Nothing measured this" is not a failure and must never be dressed as one: an
unlit lamp and the faintest text on the card, the same register the group
readout uses for its unmeasured state. */
.tg-test--none .tg-test-msg {
font-family: var(--font-sans);
color: var(--faint);
}
.tg-test--wait .tg-test-msg {
color: var(--amber);
}
@@ -1038,6 +1045,235 @@
}
/* ---- responsive ---- */
/* ---- chain hop rail ----
* The chain section's signature, and the one place this card spends any
* boldness: the path is drawn as a CONDUCTOR with a numbered lamp at each hop,
* and the conductor is SEVERED below the first hop that was probed and did not
* answer. A chain is a single series path, so the question is never "how many
* hops are green", it is "where does my traffic stop" — and a broken line answers
* that before a word has been read.
*
* The two marks carry two different facts and must not be conflated:
* - the LAMP is that hop's own measurement (good / warn / crit / unlit). Below
* the break there is no measurement to draw: the daemon stops walking at the
* first dead hop, so those lamps are UNLIT and the row says which hop stopped
* the walk. Unlit is never a shade of red — it claims nothing, which is the
* truth about a hop nobody dialled;
* - the CONDUCTOR is reachability through the path, which really does stop.
*
* Orange is untouched here. Semantics carry every colour, and everything that is
* not a lamp is groove-grey. No transitions and no animation anywhere in the
* rail, so there is nothing for reduced-motion to switch off. */
.ch-rail {
gap: 8px;
}
.ch-eyebrow {
font-family: var(--font-mono);
font-size: 9px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--faint);
}
.ch-hops {
--ch-num: 1.8ch; /* the engraved hop number's gutter */
--ch-gap: 8px;
--ch-led: 10px; /* must match .led's width */
--ch-lampy: 14px; /* row top → lamp centre; the conductor's anchor */
/* x of the conductor: the number gutter, one gap, then the lamp's centre */
--ch-spine: calc(var(--ch-num) + var(--ch-gap) + var(--ch-led) / 2);
list-style: none;
margin: 0;
padding: 0;
}
.ch-hop {
position: relative;
display: grid;
grid-template-columns: var(--ch-num) var(--ch-led) minmax(0, 1fr);
column-gap: var(--ch-gap);
align-items: start;
}
/* the conductor — two halves per row, so a break lands on one link only */
.ch-hop::before,
.ch-hop::after {
content: '';
position: absolute;
left: var(--ch-spine);
width: 2px;
margin-left: -1px;
/* Brighter than a plain groove: this line IS the readout, and at groove
strength it disappeared into the panel and took the whole idea with it. */
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
}
.ch-hop::before {
top: 0;
height: calc(var(--ch-lampy) - var(--ch-led) / 2 - 3px);
}
.ch-hop::after {
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 3px);
bottom: 0;
}
/* Nothing feeds hop 1 from above, and nothing leaves the exit downward — the
path starts and ends inside this rail. */
.ch-hop--first::before {
display: none;
}
.ch-hop--exit::after {
bottom: auto;
height: 9px;
}
/* …the exit ends on a crossbar instead of trailing off: end of line. */
.ch-hop--exit .ch-socket {
position: relative;
}
.ch-hop--exit .ch-socket::after {
content: '';
position: absolute;
left: 50%;
transform: translateX(-50%);
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 12px);
width: 11px;
height: 2px;
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
}
/* THE SEVER. Everything from the dead hop's outgoing link downward is drawn as a
broken conductor: unmistakably not-a-line at a glance, and unmistakably not a
colour, because a colour here would compete with the lamps that carry health. */
.ch-hop--dead::after,
.ch-hop--severed::before,
.ch-hop--severed::after {
background: repeating-linear-gradient(
to bottom,
color-mix(in srgb, var(--dim) 45%, var(--groove)) 0 3px,
transparent 3px 7px
);
}
.ch-num {
font-size: 10px;
line-height: calc(var(--ch-lampy) * 2);
text-align: right;
color: var(--faint);
}
.ch-socket {
display: flex;
align-items: center;
height: calc(var(--ch-lampy) * 2);
}
.ch-body {
min-width: 0;
/* Separates one hop from the next. The conductor runs through this space, so
too little of it and two hops read as one wrapped row. */
padding-bottom: 8px;
}
.ch-l1 {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 8px;
min-height: calc(var(--ch-lampy) * 2);
}
.ch-name {
font-size: 12px;
color: var(--ink);
overflow-wrap: anywhere;
}
.ch-hop--untested .ch-name {
color: var(--dim);
}
/* The exit marker is NEUTRAL on purpose. The config path above this rail tags its
exit green, which is free there — but in here green means "answering", and a
green badge on the last hop would read as a health claim about it. */
.ch-tag {
font-family: var(--font-mono);
font-size: 8.5px;
letter-spacing: 0.14em;
text-transform: uppercase;
color: var(--faint);
}
.ch-delay {
font-size: 11.5px;
font-weight: 700;
color: var(--ink);
}
.ch-quiet {
font-family: var(--font-sans);
font-size: 12px;
color: var(--faint);
}
/* A blocked hop's phrase carries a tooltip with the blocking hop's engine
outbound, so it takes the same help cursor as .ch-dead. No colour of its own:
the finding is red once, on the hop that actually failed. */
.ch-blocked {
cursor: help;
}
.ch-age {
margin-left: auto;
font-size: 10.5px;
color: var(--faint);
white-space: nowrap;
}
/* The counters read exactly as they do on a group card — alive out of TESTED,
with the untested remainder as a quiet aside only when there is one. Same
register, same weights, deliberately not a second dialect. */
.ch-l2 {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 8px;
margin-top: 1px;
font-size: 11.5px;
letter-spacing: 0.02em;
color: var(--dim);
}
.ch-count {
font-size: 12px;
color: var(--dim);
white-space: nowrap;
}
.ch-count b {
font-size: 14px;
font-weight: 700;
color: var(--ink);
}
.ch-hop--dead .ch-count b {
color: var(--crit);
}
.ch-word {
font-size: 10.5px;
letter-spacing: 0.12em;
text-transform: uppercase;
color: var(--faint);
}
.ch-dead {
padding: 1px 6px;
border-radius: 4px;
background: color-mix(in srgb, var(--crit) 12%, transparent);
font-size: 10.5px;
color: var(--crit);
white-space: nowrap;
cursor: help;
}
.ch-rest {
font-size: 10.5px;
color: var(--faint);
}
/* On a blocked hop this chip says "set to", not "now": a pick nothing crossed.
It steps back to faint so it can't be mistaken for a live reading. */
.ch-hop--blocked .ch-now {
color: var(--faint);
}
.ch-now {
max-width: 28ch;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
font-size: 10.5px;
color: var(--dim);
}
@media (max-width: 640px) {
.tg-sec-hd {
flex-wrap: wrap;
@@ -1060,6 +1296,20 @@
.gh-now {
max-width: 100%;
}
/* On a phone the age stamp stops being pushed to a lonely right edge and just
joins the end of the hop's line; the selected node gets the full width
instead of an ellipsis it doesn't need. */
.ch-age {
margin-left: 0;
}
.ch-now {
max-width: 100%;
}
/* Every field of a hop wraps onto its own line at this width, so the gap
between hops has to grow with them or the rail reads as one block of text. */
.ch-body {
padding-bottom: 12px;
}
.gh-mems {
max-height: 260px;
}
+458 -100
View File
@@ -1,6 +1,6 @@
import './Targets.css'
import { Fragment, useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import type { LedVariant } from '../components'
import {
apply as apiApply,
@@ -22,6 +22,8 @@ import type {
GroupTestResult,
GroupTestStatus,
Chain,
ChainHealth,
ChainHopHealth,
Egress,
Interface,
Node,
@@ -161,18 +163,18 @@ function renameReferences(m: Model, kind: RefKind, from: string, to: string): Mo
}
/**
* The sentence a delete confirmation appends: what still points at this target,
* and what happens to it. Empty list ⇒ an explicit "nothing references it", so
* the operator can delete a stray with confidence instead of guessing.
* The body of a delete confirmation: what still points at this target, and what
* happens to it. Empty list ⇒ an explicit "nothing references it", so the
* operator can delete a stray with confidence instead of guessing.
*/
function refWarning(refs: RefSite[]): string {
if (refs.length === 0) return ' Nothing references it.'
if (refs.length === 0) return 'Nothing references it.'
const shown = refs.slice(0, 4).map((r) => r.label)
const more = refs.length - shown.length
const list = `${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}`
return refs.length === 1
? ` It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
: ` It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
? `It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
: `It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
}
/**
@@ -436,14 +438,14 @@ const normalizeTest = (st: GroupTestStatus): GroupTestStatus => ({
})
/**
* How the header names the reach of a running exit test. The name matters more
* How the header names the reach of a running refresh pass. The name matters more
* than the number when there is only one: "auto" tells the operator which button
* they pressed; "1 target" tells them nothing they didn't already know.
*/
function scopeLabel(scope: string[], targetCount: number): string {
if (scope.length === 1) return scope[0]
if (scope.length === 0) return 'exits' // pre-scope daemon — say nothing false
return scope.length >= targetCount ? 'every exit' : `${scope.length} exits`
if (scope.length === 0) return 'targets' // pre-scope daemon — say nothing false
return scope.length >= targetCount ? 'every target' : `${scope.length} targets`
}
/** Which editor (add or edit-by-name) is open within a section. */
@@ -457,6 +459,7 @@ interface Opt {
// ---- page ------------------------------------------------------------------
export default function Targets() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -589,10 +592,12 @@ export default function Targets() {
[health],
)
// ---- group/chain exit test: how fast, through which node, out which address --
// The POST only kicks a run off, and a 2 s poll of the GET carries progress
// plus every result so far. One endpoint covers groups and chains alike:
// POST with a group or chain name tests that one; an empty name tests them all.
// ---- out-of-turn refresh: how fast, through which node, out which address ----
// The POST does NOT dial. It asks the observatory — the only thing in the daemon
// that measures anything, and it measures along the real dial path — to come
// round out of turn; a 2 s poll of the GET carries progress plus every reading
// so far. One endpoint covers groups and chains alike: POST with a name refreshes
// that one, an empty name refreshes them all.
const [gtest, setGtest] = useState<GroupTestStatus>(IDLE_TEST)
const [gtestErr, setGtestErr] = useState<string | null>(null)
const [polling, setPolling] = useState(false)
@@ -635,7 +640,7 @@ export default function Targets() {
void readTest().then((st) => {
if (!alive || !st || st.running) return
setPolling(false)
flash('Group test complete')
flash('Readings refreshed')
})
}, 2000)
return () => {
@@ -651,16 +656,16 @@ export default function Targets() {
if (r.started) {
setGtestErr(null)
setPolling(true)
flash(name ? `Testing ${name}…` : 'Testing every exit…')
flash(name ? `Refreshing ${name}…` : 'Refreshing every reading…')
void readTest()
} else if (r.reason === 'already running') {
setPolling(true) // pick up the run someone else started
flash('A group test is already running')
setPolling(true) // pick up the pass someone else started
flash('The prober is already refreshing')
} else {
flash(`Couldn’t start the test — ${r.reason || 'the daemon refused it'}`)
flash(`Couldn’t ask for a refresh — ${r.reason || 'the daemon refused it'}`)
}
} catch (e) {
flash(`Couldn’t start the test — ${errText(e)}`)
flash(`Couldn’t ask for a refresh — ${errText(e)}`)
}
},
[flash, readTest],
@@ -787,13 +792,18 @@ export default function Targets() {
)
const removeGroup = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'group', name)
if (!window.confirm(`Delete group “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete group',
title: `Delete group “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Groups: groups.filter((g) => g.Name !== name) }, `Deleted ${name}`)
},
[config, groups, save],
[config, groups, save, confirm],
)
// ---- chain mutations ------------------------------------------------------
@@ -820,13 +830,18 @@ export default function Targets() {
)
const removeChain = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'chain', name)
if (!window.confirm(`Delete chain “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete chain',
title: `Delete chain “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Chains: chains.filter((c) => c.Name !== name) }, `Deleted ${name}`)
},
[config, chains, save],
[config, chains, save, confirm],
)
// ---- egress mutations -----------------------------------------------------
@@ -853,13 +868,18 @@ export default function Targets() {
)
const removeEgress = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'egress', name)
if (!window.confirm(`Delete egress “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete egress',
title: `Delete egress “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Egresses: egresses.filter((e) => e.Name !== name) }, `Deleted ${name}`)
},
[config, egresses, save],
[config, egresses, save, confirm],
)
const busy = saving || applying
@@ -894,10 +914,11 @@ export default function Targets() {
<h2 className="tg-sec-title">Groups</h2>
<span className="tg-sec-count mono">{groups.length} configured</span>
{/* The observatory's background probing is invisible by design — it
keeps every used group's and chain's numbers fresh on its own. The
one manual run left is the exit test: it is scoped to the groups
and chains it names, so its progress says WHICH, and its badge
lands only on those cards. */}
keeps every used group's and chain's numbers fresh on its own, along
the path traffic actually takes. The one manual control left does
not measure anything itself: it asks that prober to come round out
of turn. It is scoped to the groups and chains it names, so its
progress says WHICH, and its badge lands only on those cards. */}
<div className="tg-sec-ctl">
{groupHealthOn && (
<>
@@ -905,11 +926,11 @@ export default function Targets() {
<span
className="tg-run tg-run--exit"
role="status"
title="An exit test sends one connection through each group or chain it covers and reports the delay and the address the internet sees."
title="The background prober is measuring the targets this refresh covers, along the path each one's traffic really takes."
>
<Led variant="amber" pulse />
<span className="tg-run-what">
exit test · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
refreshing · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
</span>
<span className="tg-run-n mono">
{gtest.done}/{gtest.total}
@@ -919,9 +940,9 @@ export default function Targets() {
<Button
onClick={() => void runTest()}
disabled={busy || !config || (groups.length === 0 && chains.length === 0) || gtest.running}
title="Send one connection through each group and chain and report the delay and the exit address the internet sees"
title="Ask the background prober to measure every group and chain out of turn, then show what it measured. The panel opens no connection of its own."
>
{gtest.running ? 'Testing…' : 'Test every exit'}
{gtest.running ? 'Refreshing…' : 'Refresh every reading'}
</Button>
</>
)}
@@ -941,6 +962,13 @@ export default function Targets() {
dials out through a tunnel measures them through that tunnel, so the same node can be alive
in one group and dead in another.
</p>
<p className="tg-sec-note">
One thing measures, and the panel is not it. A background prober walks every path your
rules use — hop by hop, exactly as traffic goes — and every number on this page is a read
of what it found. <strong>Refresh every reading</strong> asks it to come round out of turn
instead of waiting for the next pass; it opens no connection of its own, so a target no
rule routes through has nothing to report and says so.
</p>
{groupHealthOn && healthErr && (
<p className="tg-test-err" role="alert">
@@ -951,7 +979,7 @@ export default function Targets() {
{groupHealthOn && gtestErr && (
<p className="tg-test-err" role="alert">
Couldn’t read the test results — {gtestErr}.{' '}
Couldn’t read the refreshed numbers — {gtestErr}.{' '}
<button className="linkish" onClick={() => void readTest()}>
Retry
</button>
@@ -1094,7 +1122,9 @@ export default function Targets() {
chain={c}
busy={busy}
showHealth={groupHealthOn}
used={healthByChain.get(c.Name)?.used}
// The whole chain health record, not just `.used` — the card
// renders the observatory's per-hop measurements from it.
health={healthByChain.get(c.Name)}
test={testByGroup.get(c.Name)}
// The badge is this card's business only when the run names it.
testing={gtest.running && testScope.has(c.Name)}
@@ -1216,7 +1246,7 @@ function GroupRow({
group: Group
busy: boolean
/** Group health checks are on (Settings). When false, the card drops its health
* readout, its exit-test readout and its Test button — it is config only. */
* readout, its end-to-end reading and its Refresh button — it is config only. */
showHealth: boolean
/** This group's membership health, or undefined when the engine hasn't built
* it (not applied yet, or dropped for having no usable members). */
@@ -1226,14 +1256,14 @@ function GroupRow({
healthKnown: boolean
test?: GroupTestResult
/**
* A group exit test covering THIS group is in flight.
* A refresh pass covering THIS group is in flight.
*
* Deliberately not "a test is running": the caller resolves it against the run's
* scope. There is no per-card equivalent for the health run — that one measures
* every group at once and is reported once, in the section header.
*/
testing: boolean
/** Any exit test is in flight; the daemon runs one at a time. */
/** Any refresh pass is in flight; the daemon runs one at a time. */
testBusy: boolean
onTest: () => void
onEdit: () => void
@@ -1293,7 +1323,11 @@ function GroupRow({
health={health}
healthKnown={healthKnown}
/>
<GroupTestReadout test={test} pending={testing && !test} />
<GroupTestReadout
test={test}
pending={testing && !test}
hideAbsence={health?.used === false}
/>
</>
)}
</div>
@@ -1304,7 +1338,7 @@ function GroupRow({
editLabel={`Edit group ${group.Name}`}
deleteLabel={`Delete group ${group.Name}`}
onTest={showHealth ? onTest : undefined}
testLabel={showHealth ? `Test the exit of group ${group.Name}` : undefined}
testLabel={showHealth ? `Refresh the reading for group ${group.Name}` : undefined}
testDisabled={testBusy}
/>
</li>
@@ -1363,21 +1397,7 @@ function GroupHealthReadout({
// its members would stay "untested" forever. That is a fact about the ROUTING
// CONFIG, not about the members — so instead of counters that could only ever
// read as a permanent unknown, the card says so, quietly: unused, not unwell.
if (!health.used) {
return (
<div className="gh gh--unused">
<div className="gh-line">
<span
className="gh-unused"
title="No enabled rule routes through this group, so its members are not probed. Add it to a rule to see health."
>
unused
</span>
<span className="gh-quiet">not probed — no enabled rule routes through this group</span>
</div>
</div>
)
}
if (!health.used) return <NotRoutedNote kind="group" />
const v = verdictOf(health)
const { total, tested, alive, dead, untested } = health
@@ -1540,6 +1560,56 @@ function GroupHealthReadout({
)
}
/**
* The card's answer when NOTHING ROUTES THROUGH THIS TARGET. Shared by the group
* card and the chain card, because it is the same misunderstanding on both.
*
* It has to carry two statements, and the old one-liner ("not probed — no enabled
* rule routes through this group") only carried the first. Read fast it still
* landed as a verdict: a card that normally shows health and today shows a grey
* pill reads as "the health is bad". So the two meanings are now separated, on
* purpose and in this order:
*
* 1. the ROUTING FACT — nothing routes here, so nothing measures it;
* 2. the NON-FACT — this is not a health reading at all. Absent numbers are
* absence of measurement, never failure.
*
* For a GROUP there is a third line, and it is the confusion this whole change
* exists to end: a group used only as a hop inside a chain is never routed to
* DIRECTLY, so it correctly reads unused here while carrying real traffic as a
* hop. Its health is measured at that hop, on the chain's card.
*
* Unused is neutral — groove-grey, never amber, never crit. It is a state of the
* config, and the config is not sick.
*/
function NotRoutedNote({ kind }: { kind: 'group' | 'chain' }) {
return (
<div className="gh gh--unused">
<div className="gh-line">
<span className="gh-unused">unused</span>
<span className="gh-quiet">
No enabled rule routes through this {kind}, so the observatory never probes it.
</span>
</div>
<p className="gh-say">
That is a routing fact, not a health reading. There are no numbers here because nothing
measured this {kind} — not because it failed.
</p>
{kind === 'group' ? (
<p className="gh-say">
A group used only as a hop inside a chain reads unused here on purpose: the rules point at
the chain, not at the group. Its members are measured at that hop, so its real health is on
that chain’s card, hop by hop.
</p>
) : (
<p className="gh-say">
Point a rule at this chain and the observatory starts measuring every hop within seconds.
</p>
)}
</div>
)
}
/**
* One group's member rows, fetched on demand.
*
@@ -1645,7 +1715,9 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
}
/**
* What a group test found, in the four states it actually has.
* What the OBSERVATORY measured for this target end to end, in the four states it
* actually has. Nothing here was dialled by the panel — it is a read of the
* background prober's own measurement along the real path.
*
* The one worth spelling out: `ok` with an EMPTY `exit_ip` is a SUCCESS. The
* delay was measured; only the address lookup came back empty. Rendering that as
@@ -1653,15 +1725,45 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
* traffic, so it reads as a result with the address slot marked unknown — dim,
* not red, and the LED stays green.
*/
function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending: boolean }) {
/**
* Errors that mean NO MEASUREMENT EXISTS, as opposed to "this target is broken".
*
* Three of the observatory's four failure reasons are about the observatory, not
* about the path: nothing routes here, nothing has reached it yet, or background
* probing is switched off. Painting those crit-red — which is what `ok:false`
* used to buy you — reports a fault that nobody has found, on a target that may
* be carrying traffic perfectly. Only "the observatory's probe through this path
* failed" is a health finding, and it is deliberately NOT in this list.
*
* Matched on a stable fragment rather than the whole sentence, so a daemon that
* rewords the tail still classifies. An error we don't recognise stays red: an
* unknown failure is likelier to be real than not, and that is the safe default.
*/
const NO_MEASUREMENT = [
'not routed by any enabled rule',
'has not reached this target yet',
'background probing is disabled',
]
const isAbsence = (err: string): boolean => NO_MEASUREMENT.some((frag) => err.includes(frag))
function GroupTestReadout({
test,
pending,
hideAbsence,
}: {
test?: GroupTestResult
pending: boolean
/** The card already explains why nothing measures this target (the unused
* note), so an absence error here would just say it a second time. */
hideAbsence?: boolean
}) {
if (pending) {
// "testing", never "measuring": the health run owns that word and covers every
// group at once. Two runs that read the same on a card is how one group's test
// came to look like all four were busy.
// Names who is working and on what: the prober, on this target. The badge is
// scoped to the cards the run covers, so it can say "this one" honestly.
return (
<div className="tg-test tg-test--wait" role="status">
<Led variant="amber" pulse />
<span className="tg-test-msg">testing this exit…</span>
<span className="tg-test-msg">waiting for the prober to measure this…</span>
</div>
)
}
@@ -1670,10 +1772,22 @@ function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending:
const at = test.tested_unix ? fmtClock(test.tested_unix) : ''
if (!test.ok) {
// No measurement exists. Unlit lamp, quiet text: this panel's way of saying
// "no verdict", which is precisely the state — never a red one.
if (isAbsence(test.error)) {
if (hideAbsence) return null
return (
<div className="tg-test tg-test--none" role="status">
<Led variant="off" />
<span className="tg-test-msg">{test.error}</span>
{at && <span className="tg-test-at mono">{at}</span>}
</div>
)
}
return (
<div className="tg-test tg-test--bad" role="status">
<Led variant="crit" />
<span className="tg-test-msg">{test.error || 'the test failed'}</span>
<span className="tg-test-msg">{test.error || 'the probe failed'}</span>
{at && <span className="tg-test-at mono">{at}</span>}
</div>
)
@@ -2098,7 +2212,7 @@ function ChainRow({
chain,
busy,
showHealth,
used,
health,
test,
testing,
testBusy,
@@ -2109,31 +2223,41 @@ function ChainRow({
chain: Chain
busy: boolean
/** Group health checks are on (Settings). When false, the card drops its
* exit-test readout and Test button — it is config only. */
* health readout and Refresh button — it is config only. */
showHealth: boolean
/** This chain's reachability (GroupHealth.Used's chain analogue, plan §5.E).
* undefined ⇒ the health endpoint hasn't reported this chain (not applied yet, or
* a daemon version without chains): no badge. false ⇒ no enabled rule routes
* through the chain, so the observatory never probes it and the card renders
* "unused" instead of an exit-test readout. */
used?: boolean
/** Everything the observatory knows about this chain: whether any enabled rule
* routes through it, and the per-hop measurements along it.
* undefined ⇒ the health endpoint hasn't reported this chain at all (not
* applied yet, or a daemon version without chains): the card says nothing
* rather than guessing. */
health?: ChainHealth
test?: GroupTestResult
/** An exit test covering THIS chain is in flight (the caller resolves it
/** A refresh pass covering THIS chain is in flight (the caller resolves it
* against the run's scope, exactly as for a group card). */
testing: boolean
/** Any exit test is in flight; the daemon runs one at a time. */
/** Any refresh pass is in flight; the daemon runs one at a time. */
testBusy: boolean
onTest: () => void
onEdit: () => void
onDelete: () => void
}) {
const hops = asArray(chain.Hops)
// A LEADING `egress:` is not a hop and the rail below already knows it: the
// daemon lifts it into hop 1's entry detour (see hopLabels), so it is tagged
// "entry" and never numbered. The badge counted it anyway, which is how a chain
// drawn with four hops came to be labelled "5 hops" directly above them.
const entryEgress = hops.length > 0 && hops[0].startsWith('egress:')
const numbered = entryEgress ? hops.length - 1 : hops.length
return (
<li className="tg-row">
<div className="tg-row-main">
<div className="tg-row-l1">
<span className="tg-row-name">{chain.Name}</span>
<span className="tg-badge">{hops.length} hop{hops.length === 1 ? '' : 's'}</span>
<span className="tg-badge">
{numbered === 0 && entryEgress
? 'entry only · no exit'
: `${numbered} hop${numbered === 1 ? '' : 's'}`}
</span>
</div>
<div className="tg-row-l2">
{hops.length === 0 ? (
@@ -2165,25 +2289,21 @@ function ChainRow({
{showHealth && (
<>
{/* A chain no enabled rule routes through is never probed (the
observatory walks only reachable paths), so instead of an exit-test
readout the card says so, quietly — the same "unused" pattern the
group card uses (GroupHealthReadout), not a new design. `used` is
undefined until the health endpoint reports this chain (or from a
daemon version without chains): no badge then. */}
{used === false && (
<div className="gh gh--unused">
<div className="gh-line">
<span
className="gh-unused"
title="No enabled rule routes through this chain, so its exit is not probed. Add it to a rule to see health."
>
unused
</span>
<span className="gh-quiet">not probed — no enabled rule routes through this chain</span>
</div>
</div>
)}
<GroupTestReadout test={test} pending={testing && !test} />
observatory walks only reachable paths), so instead of a health
readout the card says so — the same "unused" note the group card
uses, not a new design. `health` is undefined until the endpoint
reports this chain (or on a daemon without chains): say nothing
then rather than guess. */}
{health?.used === false ? (
<NotRoutedNote kind="chain" />
) : health?.used ? (
<ChainHopRail chain={chain.Name} defs={hops} hops={health.hops} />
) : null}
<GroupTestReadout
test={test}
pending={testing && !test}
hideAbsence={health?.used === false}
/>
</>
)}
</div>
@@ -2194,13 +2314,250 @@ function ChainRow({
editLabel={`Edit chain ${chain.Name}`}
deleteLabel={`Delete chain ${chain.Name}`}
onTest={showHealth ? onTest : undefined}
testLabel={showHealth ? `Test the exit of chain ${chain.Name}` : undefined}
testLabel={showHealth ? `Refresh the reading for chain ${chain.Name}` : undefined}
testDisabled={testBusy}
/>
</li>
)
}
// ---- chain hop rail --------------------------------------------------------
/**
* The API gives hops an index and no name. The page already knows the names — the
* model's own `Hops` strings ("egress:ewan", "node:awgout", "group:sub0") — so
* zip the two by POSITION.
*
* Two things make that safe rather than clever. A LEADING `egress:` is not a
* numbered hop: the daemon lifts it into hop 1's entry detour, so it is dropped
* before counting. And if the counts still disagree — a chain that splices
* sub-chains gets FLATTENED by the daemon, producing more wire hops than the
* config lists — every label is dropped. A hop labelled with its neighbour's name
* is worse than a hop with no name at all: it would send someone to fix the wrong
* target.
*/
function hopLabels(defs: string[], hops: ChainHopHealth[]): (string | undefined)[] {
const numbered = defs.length > 0 && defs[0].startsWith('egress:') ? defs.slice(1) : defs
if (numbered.length !== hops.length) return hops.map(() => undefined)
return hops.map((h) => (h.index >= 1 && h.index <= numbered.length ? numbered[h.index - 1] : undefined))
}
/**
* One hop's lamp.
*
* `dead` is crit and `untested` is an UNLIT socket — never red, because nothing
* has been measured and an unlit lamp is this panel's way of saying "no verdict".
* That covers a hop the walk never reached (`blocked_by`) too: it is neither
* healthy nor broken, and unlit is the only mark that claims neither.
* The fourth case is the page's existing house reading, applied here for
* consistency rather than invented: a group hop that is carrying traffic but has
* confirmed failures on its board is amber. `state` stays the daemon's word for
* "can this hop carry traffic"; the amber only qualifies HOW WELL.
*/
function hopLed(h: ChainHopHealth): LedVariant {
if (h.state === 'dead') return 'crit'
if (h.state === 'untested') return 'off'
return h.dead > 0 ? 'amber' : 'on'
}
/**
* What the observatory measured at each position of a chain — the reading the
* daemon always took and the panel never showed.
*
* THE DESIGN RISK, and the one place this card spends any boldness: the rail
* draws the CONDUCTOR as well as the lamps, and severs it below the first dead
* hop. A chain is a single series path, so the operator's real question is never
* "how many hops are green" — it is "where does my traffic stop". Four lamps in a
* column answer the first question and leave the second to arithmetic. A broken
* conductor answers the second one before you have read a single word, which is
* the whole reason this feature exists.
*
* It stays honest by keeping two different facts on two different marks. The
* CONDUCTOR is reachability through the path, and that genuinely does stop at the
* break. The LAMPS are measurements — and there are none below the break to show:
* the daemon walks the path in order and stops at the first hop that does not
* answer, because every later hop is dialled THROUGH that one. Those hops arrive
* `untested` with `blocked_by` naming the hop that stopped the walk, so their
* lamps stay UNLIT: not a soft red, not a pale green, just this panel's way of
* saying no verdict exists about a hop nobody reached. The row says so in words
* too, naming that hop, because "why is this row empty" is the question the shape
* alone cannot answer.
*
* Nothing here is re-derived from the daemon's counters — the block, the zeroed
* numbers and the state all come off the wire. The only thing the panel adds is
* what a chain structurally is.
*
* Everything around the rail is deliberately quiet: no colour but the semantic
* lamps, no motion at all, the orange accent untouched.
*/
function ChainHopRail({
chain,
defs,
hops,
}: {
chain: string
/** The chain's configured hops, straight off the model — the only source of names. */
defs: string[]
/** Absent ⇒ the engine never materialised per-hop outbounds. NOT "no hops". */
hops?: ChainHopHealth[]
}) {
const ordered = useMemo(() => [...asArray(hops)].sort((a, b) => a.index - b.index), [hops])
const labels = useMemo(() => hopLabels(defs, ordered), [defs, ordered])
// The first hop that was probed and did not answer. Everything after it is
// unreachable THROUGH THIS CHAIN, whatever its own lamp says. `untested` is
// never a break: nothing was measured, so nothing is known to be severed.
const breakAt = ordered.findIndex((h) => h.state === 'dead')
if (ordered.length === 0) {
// Say why, in one line, instead of an empty rail. The daemon collapses a
// single-target chain into a plain alias and never builds copies to measure,
// so we can tell the two absences apart from the config alone.
const numbered = defs.filter((d, i) => !(i === 0 && d.startsWith('egress:')))
return (
<div className="gh gh--absent">
<span className="gh-absent-msg">
{numbered.length <= 1
? 'This chain has a single hop, so the engine points traffic straight at that target instead of building a path to measure. Its health is on that target’s own card.'
: 'The engine hasn’t built this chain’s hops yet, so there is nothing measured per hop. They appear once it is running with this config applied.'}
</span>
</div>
)
}
return (
<div className="gh ch-rail">
<span className="ch-eyebrow">measured, hop by hop</span>
<ol className="ch-hops">
{ordered.map((h, i) => {
const label = labels[i]
const severed = breakAt >= 0 && i > breakAt
// Straight off the wire: present ⇒ the walk never reached this hop, so
// there is nothing measured here and the daemon has already named the
// hop that stopped it. Never inferred from the counters.
const blocked = h.blocked_by
const cls = [
'ch-hop',
`ch-hop--${h.state}`,
blocked ? 'ch-hop--blocked' : '',
severed ? 'ch-hop--severed' : '',
h.exit ? 'ch-hop--exit' : '',
i === 0 ? 'ch-hop--first' : '',
]
.filter(Boolean)
.join(' ')
const age = fmtAge(h.age_seconds)
return (
<li key={h.tag || h.index} className={cls}>
<span className="ch-num mono" aria-hidden="true">
{h.index}
</span>
<span className="ch-socket">
<Led variant={hopLed(h)} />
</span>
<div className="ch-body">
<div className="ch-l1">
<span className="ch-name mono" title={`engine outbound ${h.tag}`}>
{label ?? (h.kind === 'group' ? 'a group hop' : 'a node hop')}
</span>
{h.exit && <span className="ch-tag">exit</span>}
{h.state === 'alive' && h.delay_ms > 0 && (
<span className="ch-delay mono">{h.delay_ms} ms</span>
)}
{/* Two different silences. A plain untested hop is a timing
gap that fills in by itself; a BLOCKED one never will,
because the walk stopped above it — so it says which hop
stopped it instead of implying someone should wait. */}
{blocked ? (
<span
className="ch-quiet ch-blocked"
title={`hop ${blocked.index} did not answer, so nothing was dialled through it (engine outbound ${blocked.tag})`}
>
no reading — the probe stopped at hop {blocked.index}
</span>
) : (
h.state === 'untested' && <span className="ch-quiet">not measured yet</span>
)}
{age && <span className="ch-age mono">{age}</span>}
</div>
{/* A node hop IS its own measurement (total 1), so counters would
only restate the lamp. A group hop rolls up its per-hop member
copies, and those read exactly as they do everywhere else in
this app: alive out of TESTED, with the untested remainder as a
quiet aside only when there is one. A blocked hop has those
counters zeroed by the daemon, so it lands in the tested === 0
branch — and there it must say the members were never REACHED,
not that they are still waiting their turn. */}
{h.kind === 'group' && h.total > 0 && (
<div className="ch-l2">
{h.tested === 0 ? (
<span className="ch-rest mono">
{h.total} member{h.total === 1 ? '' : 's'},{' '}
{blocked ? 'none of them reached' : 'none measured'}
</span>
) : (
<>
<span className="ch-count mono">
<b>{h.alive}</b> / {h.tested}
</span>
<span className="ch-word">alive</span>
{h.dead > 0 && (
<span
className="ch-dead mono"
title={`${h.dead} member${h.dead === 1 ? '' : 's'} were probed at this hop and did not answer`}
>
{h.dead} not answering
</span>
)}
{h.untested > 0 && (
<span className="ch-rest mono">
tested {h.tested} of {h.total}
</span>
)}
</>
)}
{/* The wrapper keeps its pick even when nothing crossed it,
so on a blocked hop this is the node it WOULD use — say
that, rather than "now", which claims live traffic. */}
{h.selected && (
<span
className="ch-now mono"
title={
blocked
? `Hop ${h.index} of “${chain}” is set to ${h.selected}; nothing crossed it to measure`
: `Traffic crossing hop ${h.index} of “${chain}” is on ${h.selected}`
}
>
{blocked ? 'set to' : 'now'} → {h.selected}
</span>
)}
</div>
)}
</div>
</li>
)
})}
</ol>
{/* The sentence the rail's shape implies, written out — because the break is
the answer someone came here for, and a graphic alone should never be the
only place a finding exists. It names the dead hop, since that is the one
thing here anybody can act on. */}
{breakAt >= 0 && (
<p className="gh-say gh-say--bad">
Hop {ordered[breakAt].index}
{labels[breakAt] ? ` (${labels[breakAt]})` : ''} was probed and did not answer, so traffic
stops there
{breakAt < ordered.length - 1
? ' — and the hops below it are dialled through it, so nothing reached them and nothing is known about them.'
: '.'}
</p>
)}
</div>
)
}
function ChainEditor({
initial,
hopOptions,
@@ -2675,8 +3032,8 @@ function RowActions({
busy: boolean
editLabel: string
deleteLabel: string
// Only groups and chains can be tested, so the control is optional and absent
// everywhere else rather than a disabled stub on every row.
// Only groups and chains are probed, so the refresh control is optional and
// absent everywhere else rather than a disabled stub on every row.
onTest?: () => void
testLabel?: string
testDisabled?: boolean
@@ -2689,8 +3046,9 @@ function RowActions({
onClick={onTest}
disabled={busy || testDisabled}
aria-label={testLabel}
title="Ask the background prober to measure this target out of turn. It does not open a connection from the panel."
>
Test
Refresh
</Button>
)}
<Button className="tg-act" onClick={onEdit} disabled={busy} aria-label={editLabel}>
+91
View File
@@ -0,0 +1,91 @@
// pendingConfirm — the record of an armed auto-rollback, shared by the whole panel.
//
// Run with `npm test`. The module imports React only for its hook; the plain
// functions exercised here touch neither React nor the DOM, and `localStorage` is
// absent under node, which is itself one of the cases worth pinning (the panel
// must still work, it just forgets on reload).
//
// What these protect:
// - arming when commit-confirm is OFF must record nothing. The daemon does not
// arm a window then, and a countdown for a rollback that will never happen is
// the same class of lie as the "Confirmed" message this module replaced.
// - a window that has elapsed reads as gone, so nothing renders "0 s left".
// - expiry notifies exactly once even though several components watch it.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import {
armPendingConfirm,
clearPendingConfirm,
confirmTimeout,
noteConfirmTimeout,
onPendingConfirmExpire,
readPendingConfirm,
} from './pendingConfirm.ts'
test('commit-confirm off ⇒ arming records nothing', () => {
noteConfirmTimeout(0)
assert.equal(confirmTimeout(), 0)
armPendingConfirm()
assert.equal(readPendingConfirm(), null)
})
test('a window is recorded with the timeout the config reported', () => {
noteConfirmTimeout(90)
armPendingConfirm()
const p = readPendingConfirm()
assert.notEqual(p, null)
assert.equal(p!.total, 90)
// Deadline is in the future and within a second of now + the window.
const left = (p!.until - Date.now()) / 1000
assert.ok(left > 89 && left <= 90, `expected ~90s left, got ${left}`)
clearPendingConfirm()
assert.equal(readPendingConfirm(), null)
})
test('switching commit-confirm off drops a window that was already armed', () => {
noteConfirmTimeout(60)
armPendingConfirm()
assert.notEqual(readPendingConfirm(), null)
noteConfirmTimeout(0)
assert.equal(readPendingConfirm(), null)
})
test('an elapsed window reads as gone, never as a countdown at zero', () => {
noteConfirmTimeout(1)
armPendingConfirm()
const p = readPendingConfirm()
assert.notEqual(p, null)
// Wind the clock forward rather than sleeping through the window.
const realNow = Date.now
Date.now = () => realNow() + 5000
try {
assert.equal(readPendingConfirm(), null)
} finally {
Date.now = realNow
}
clearPendingConfirm()
})
test('confirming does NOT fire the expiry listeners', () => {
noteConfirmTimeout(30)
let fired = 0
const off = onPendingConfirmExpire(() => {
fired++
})
armPendingConfirm()
clearPendingConfirm()
off()
assert.equal(fired, 0)
})
test('a nonsense timeout is ignored rather than taken as "off"', () => {
noteConfirmTimeout(45)
noteConfirmTimeout(Number.NaN)
noteConfirmTimeout(-1)
noteConfirmTimeout(undefined)
assert.equal(confirmTimeout(), 45)
clearPendingConfirm()
noteConfirmTimeout(0)
})
+186
View File
@@ -0,0 +1,186 @@
// The commit-confirm window, as ONE fact the whole panel can see.
//
// WHY THIS EXISTS. `POST /api/apply` always arms an auto-rollback for
// `Globals.ConfirmTimeout` seconds (shater/panel/api.go handleApply →
// ArmRollback) — EVERY Apply button does that, not just the one on the Apply
// page. But the countdown, and the button that stops it, lived in one component's
// local state. So:
//
// - pressing Apply on Routing/DNS/Nodes/Settings/Devices/Profiles/Targets said
// "Applied" and nothing else; the operator walked away and the router quietly
// reverted a minute later;
// - reloading the tab wiped the countdown AND the "Keep this config" button, so
// there was no way left to confirm from the panel at all.
//
// The daemon does not report a deadline, so this module records the one the panel
// itself armed, in `localStorage`. That is a deliberately modest claim — it knows
// about windows THIS BROWSER opened and says nothing about one opened elsewhere —
// but it survives a reload, a new tab and a navigation, which is what the two
// failures above needed.
//
// It is also the answer to "is there anything to confirm?". `apply.Confirm()`
// returns nil unconditionally, so a Confirm button that is always live can only
// ever report success. Gating it on a record here means the panel offers the
// action when it knows a window is open, and then reports an outcome it knows.
import { useEffect, useState } from 'react'
/** A live commit-confirm window the panel armed. */
export interface PendingConfirm {
/** Epoch ms at which the daemon auto-rolls back if nobody confirms. */
until: number
/** The window it started with, in seconds — the progress bar's denominator. */
total: number
}
const KEY = 'shater.pendingConfirm'
type Listener = () => void
const listeners = new Set<Listener>()
const expiryListeners = new Set<Listener>()
/** localStorage is absent under SSR/tests and throws in some privacy modes. A
* panel that cannot remember a window must still work — it just forgets on
* reload, which is exactly the old behaviour and no worse. */
function store(): Storage | null {
try {
return typeof localStorage === 'undefined' ? null : localStorage
} catch {
return null
}
}
function load(): PendingConfirm | null {
const s = store()
if (!s) return null
try {
const raw = s.getItem(KEY)
if (!raw) return null
const v = JSON.parse(raw) as Partial<PendingConfirm>
if (typeof v.until !== 'number' || typeof v.total !== 'number') return null
if (!Number.isFinite(v.until)) return null
return { until: v.until, total: v.total }
} catch {
return null
}
}
// The single in-process copy. Storage is the durable mirror, not the source of
// truth for a running tab: a `storage` event re-hydrates it when another tab
// writes.
let armed: PendingConfirm | null = load()
function emit() {
for (const l of [...listeners]) l()
}
function write(v: PendingConfirm | null) {
armed = v
const s = store()
if (s) {
try {
if (v) s.setItem(KEY, JSON.stringify(v))
else s.removeItem(KEY)
} catch {
// Storage full or blocked — the in-process copy still drives this tab.
}
}
emit()
}
/** The armed window, or null when there is none or it has already elapsed. */
export function readPendingConfirm(): PendingConfirm | null {
if (!armed) return null
return armed.until > Date.now() ? armed : null
}
// ---- the window's length ----------------------------------------------------
// `Globals.ConfirmTimeout` is all that is needed to arm a window, and every page
// reads the config anyway — so api.getConfig() feeds it here rather than each
// caller threading it through. 0 (or never seen) means commit-confirm is off, and
// arming then does nothing: an apply on such a router really is immediate.
let timeout = 0
export function noteConfirmTimeout(seconds: number | undefined) {
if (typeof seconds !== 'number' || !Number.isFinite(seconds) || seconds < 0) return
timeout = Math.floor(seconds)
// A window armed before commit-confirm was switched off is no longer real —
// drop it rather than count down to an event that will not happen.
if (timeout === 0 && armed) write(null)
}
export function confirmTimeout(): number {
return timeout
}
/** Record the window an apply just opened. Call only when the apply CHANGED
* something: an unchanged apply reconciles nothing and arms nothing. */
export function armPendingConfirm() {
if (timeout <= 0) return
write({ until: Date.now() + timeout * 1000, total: timeout })
}
/** Confirm and rollback both end the window. */
export function clearPendingConfirm() {
if (armed) write(null)
}
/** Fires when a window ran out on its own — i.e. the daemon has reverted — and
* NOT when it was confirmed or rolled back. Returns an unsubscribe. */
export function onPendingConfirmExpire(fn: Listener): () => void {
expiryListeners.add(fn)
return () => {
expiryListeners.delete(fn)
}
}
/** Idempotent: several mounted countdowns race to notice the same deadline, and
* only the first one gets to announce it. */
function expire() {
if (!armed) return
write(null)
for (const l of [...expiryListeners]) l()
}
// Another tab confirming, rolling back or applying is the same event as this one
// doing it.
if (typeof window !== 'undefined') {
window.addEventListener('storage', (e) => {
if (e.key !== KEY && e.key !== null) return
armed = load()
emit()
})
}
/**
* The armed window and its remaining seconds, ticking once a second.
*
* Returns null when nothing is armed. While non-null `remaining` is at least 1:
* reaching zero clears the record and notifies {@link onPendingConfirmExpire}, so
* no component ever renders "0 s left" for a window that is already over.
*/
export function usePendingConfirm(): { pending: PendingConfirm; remaining: number } | null {
const [, tick] = useState(0)
useEffect(() => {
const sync = () => tick((n) => n + 1)
listeners.add(sync)
sync()
return () => {
listeners.delete(sync)
}
}, [])
useEffect(() => {
const id = window.setInterval(() => {
if (armed && armed.until <= Date.now()) expire()
else if (armed) tick((n) => n + 1)
}, 1000)
return () => window.clearInterval(id)
}, [])
const pending = readPendingConfirm()
if (!pending) return null
return { pending, remaining: Math.max(1, Math.ceil((pending.until - Date.now()) / 1000)) }
}
+253
View File
@@ -0,0 +1,253 @@
// protectionState — the one sentence the whole panel shows about "am I protected".
//
// Run with `npm test` (node's built-in test runner + native TypeScript stripping;
// no test dependency is added to the SPA, which ships inside the daemon binary).
//
// The case this file was written for is "plane full, traffic direct": the exact
// state of a live router — one enabled rule, `default → direct`, no groups, no
// rule-sets — where every part of the data plane was installed and the readout
// therefore said "Protected — traffic from your network is going through the
// tunnel", under a green LED, while the whole LAN went out the plain WAN.
//
// planeState.ts has no runtime imports (both of its imports are `import type`),
// so this runs against the real module with nothing stubbed.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { engineReadout, engineState, killSwitchReadout, protectionState } from './planeState.ts'
import type { Status, Traffic } from './api.ts'
/** A healthy, fully-installed router; `traffic` is what each case varies. */
function status(over: Partial<Status> = {}): Status {
return {
running: true,
enabled: true,
active: true,
table: true,
hash: 'abc',
version: '1.11.0-shater',
kill_switch: 'closed',
engine_running: true,
plane: 'full',
warnings: [],
...over,
}
}
function withTraffic(traffic: Traffic | undefined): Status {
return status({ traffic })
}
// --- the field case ---------------------------------------------------------
test('plane full + default direct is NOT reported as protected', () => {
const s = protectionState(withTraffic({ verdict: 'direct', default: 'direct', tunnel_rules: 0 }))
assert.notEqual(s.headline, 'Protected')
assert.equal(s.variant, 'crit')
assert.equal(s.alarm, true)
// The claim that was false must not survive anywhere in the copy.
assert.doesNotMatch(s.detail, /going through the tunnel/)
// ...and the honest consequence must be stated, not implied.
assert.match(s.detail, /real address/)
})
// --- the other verdicts under a full plane ----------------------------------
test('plane full + default into a tunnel is protected', () => {
const s = protectionState(withTraffic({ verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }))
assert.equal(s.variant, 'on')
assert.equal(s.headline, 'Protected')
assert.equal(s.alarm, false)
})
test('plane full + direct default with tunnelling rules is split, not protected', () => {
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 3 }))
assert.equal(s.variant, 'amber')
assert.notEqual(s.headline, 'Protected')
// Says how much is protected, and that the default is not.
assert.match(s.detail, /3 rules/)
assert.match(s.detail, /normal internet connection/)
// A working selective setup must not raise a banner on every other page.
assert.equal(s.alarm, false)
})
test('split names a single rule in the singular', () => {
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 1 }))
assert.match(s.detail, /^One rule sends traffic/)
})
test('plane full + blocked default with rules leaks nothing and is never crit', () => {
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 2 }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, false)
assert.match(s.detail, /nothing is leaving unprotected/)
})
test('plane full + blocked default with no rules says the network has no way out', () => {
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 0 }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
assert.doesNotMatch(s.detail, /going through the tunnel/)
})
test('plane full with no verdict claims nothing either way', () => {
for (const t of [undefined, { verdict: '' as const }]) {
const s = protectionState(withTraffic(t))
assert.notEqual(s.headline, 'Protected')
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, false)
}
})
// --- the branches that were already correct ---------------------------------
test('no status yet', () => {
const s = protectionState(null)
assert.equal(s.variant, 'off')
assert.equal(s.alarm, false)
})
test('service switched off is a deliberate state, not a fault', () => {
const s = protectionState(status({ enabled: false }))
assert.equal(s.variant, 'amber')
assert.equal(s.headline, 'Turned off')
assert.equal(s.alarm, false)
})
test('hold: the kill-switch caught it — protected, offline', () => {
const s = protectionState(status({ plane: 'hold', engine_running: false, active: false }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
assert.match(s.headline, /blocked/)
})
test('none + fail-closed is the leak, and it is crit', () => {
const s = protectionState(status({ plane: 'none', table: false, engine_running: false }))
assert.equal(s.variant, 'crit')
assert.equal(s.alarm, true)
})
test('none + fail-open is the operator’s documented choice, stated not alarmed at', () => {
const s = protectionState(
status({ plane: 'none', table: false, engine_running: false, kill_switch: 'open' }),
)
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
})
test('daemon too old to send `plane` keeps its own fallback', () => {
// Nothing here may depend on `traffic`: a daemon with no `plane` has no
// `traffic` either, and this branch reads what it can observe instead.
const { plane, ...noPlane } = status()
void plane
assert.equal(protectionState(noPlane as Status).headline, 'Protected')
assert.equal(protectionState({ ...noPlane, running: false } as Status).headline, 'Service stopped')
assert.equal(
protectionState({ ...noPlane, active: false } as Status).headline,
'Starting up',
)
})
// --- engineState: the reading that could not say "down" ----------------------
//
// `apply.Status.running` was a hardcoded `true` on the daemon, so every panel LED
// derived from it was lit before it was read: App's master indicator could not
// reach its "Offline" branch, and Apply's "engine: running / stopped" row had one
// reachable value. These pin the three answers, and that "up" needs agreement.
test('engine_running:false is down even while the daemon claims it is running', () => {
assert.equal(engineState(status({ running: true, engine_running: false })), 'down')
assert.equal(engineReadout(status({ running: true, engine_running: false })).variant, 'crit')
assert.equal(engineReadout(status({ running: true, engine_running: false })).word, 'stopped')
})
test('a daemon that reports itself stopped is down whatever engine_running says', () => {
assert.equal(engineState(status({ running: false, engine_running: true })), 'down')
})
test('up needs both, and then active/idle splits the lamp', () => {
assert.equal(engineState(status({ running: true, engine_running: true })), 'up')
assert.equal(engineReadout(status({ active: true })).variant, 'on')
assert.equal(engineReadout(status({ active: true })).word, 'active')
assert.equal(engineReadout(status({ active: false })).variant, 'amber')
assert.equal(engineReadout(status({ active: false })).word, 'idle')
})
test('an older daemon with no engine_running is unknown — an unlit lamp, never green', () => {
const { engine_running, ...old } = status()
void engine_running
assert.equal(engineState(old as Status), 'unknown')
const r = engineReadout(old as Status)
assert.equal(r.variant, 'off')
assert.notEqual(r.variant, 'on')
assert.equal(r.word, 'not reported')
})
test('no status at all is unknown, not down', () => {
assert.equal(engineState(null), 'unknown')
assert.equal(engineReadout(null).variant, 'off')
assert.equal(engineReadout(null).word, 'checking…')
})
// --- killSwitchReadout: "I don't know" is not "it's armed" -------------------
//
// The Overview module read `killArmed && status?.plane !== 'none'`, and
// `undefined !== 'none'` is true — so a daemon that never reported `plane`, and
// the seconds before the first status arrives, both lit a green lamp over the
// word ARMED. These pin the fourth answer that expression could not express.
test('a daemon that does not report `plane` reads as not reported, never ARMED', () => {
const { plane, ...noPlane } = status()
void plane
const k = killSwitchReadout(noPlane as Status)
assert.equal(k.state, 'unknown')
assert.notEqual(k.value, 'ARMED')
assert.equal(k.variant, 'off')
assert.notEqual(k.variant, 'on')
assert.equal(k.blockingNow, 'not known')
})
test('no status at all is unknown too, and says there is no reading', () => {
const k = killSwitchReadout(null)
assert.equal(k.state, 'unknown')
assert.equal(k.variant, 'off')
assert.equal(k.blockingNow, 'no reading yet')
})
test('fail-closed with a plane installed is armed', () => {
for (const plane of ['full', 'hold'] as const) {
const k = killSwitchReadout(status({ plane }))
assert.equal(k.state, 'armed')
assert.equal(k.value, 'ARMED')
assert.equal(k.variant, 'on')
assert.equal(k.blockingNow, null)
}
})
test('fail-closed with no plane is configured but blocking nothing', () => {
const k = killSwitchReadout(status({ plane: 'none', table: false }))
assert.equal(k.state, 'inert')
assert.equal(k.value, 'NOT IN EFFECT')
assert.equal(k.variant, 'crit')
assert.equal(k.hot, true)
})
test('fail-open is the operator’s choice — amber, and never a plane question', () => {
for (const plane of ['full', 'none', undefined] as const) {
const k = killSwitchReadout(status({ kill_switch: 'open', plane }))
assert.equal(k.state, 'open')
assert.equal(k.value, 'OPEN')
assert.equal(k.variant, 'amber')
}
})
test('the live kill_switch wins over the saved one; the saved one only fills a gap', () => {
const { kill_switch, ...noKill } = status()
void kill_switch
// Live says open, config says closed → live wins.
assert.equal(killSwitchReadout(status({ kill_switch: 'open' }), 'closed').state, 'open')
// Nothing live → fall back to the saved policy.
assert.equal(killSwitchReadout(noKill as Status, 'open').state, 'open')
assert.equal(killSwitchReadout(noKill as Status, 'closed').state, 'armed')
})
+245 -17
View File
@@ -11,7 +11,7 @@
// same router differently.
import type { LedVariant } from './components'
import type { Status } from './api'
import type { Status, Traffic } from './api'
export interface ProtectionState {
variant: LedVariant
@@ -21,6 +21,135 @@ export interface ProtectionState {
alarm: boolean
}
// ---------------------------------------------------------------------------
// Is the engine actually up?
// ---------------------------------------------------------------------------
/**
* Three answers, and "unknown" is one of them.
*
* up — the sing-box process is running.
* down — it is not. Nothing is being proxied or filtered.
* unknown — nobody has told us. Never paint this green.
*
* THIS EXISTS BECAUSE `status.running` COULD NOT SAY "down". It is the DAEMON's
* own liveness, and on the daemons this panel shipped against it was a hardcoded
* `true` (apply.go) — so `running ? 'running' : 'stopped'` had exactly one
* reachable branch, and every LED derived from it was lit before it was read. A
* router whose engine failed to start, with a fail-closed holding plan installed
* and no internet on the LAN, showed three green lamps on the page people go to
* when they are trying to fix it.
*
* `engine_running` is the field that answers the question honestly, so it decides
* `up`. Either field may still prove a NEGATIVE — a daemon that reports itself
* stopped cannot be running an engine — and a negative always wins, so "up" needs
* both to agree. Neither field asserting anything leaves `unknown`.
*/
export type EngineState = 'up' | 'down' | 'unknown'
export function engineState(status: Status | null): EngineState {
if (!status) return 'unknown'
if (!status.running) return 'down'
if (typeof status.engine_running === 'boolean') return status.engine_running ? 'up' : 'down'
return 'unknown'
}
/** How the engine's lamp is painted and what the readout beside it says.
*
* `unknown` is an UNLIT socket, never amber and never green: amber is this
* panel's "degraded", and there is nothing to be degraded about when no reading
* has arrived. `down` is crit even when the kill-switch caught it — the engine
* being dead is the fault; whether traffic leaks is a separate lamp. */
export function engineReadout(status: Status | null): { variant: LedVariant; word: string } {
switch (engineState(status)) {
case 'down':
return { variant: 'crit', word: 'stopped' }
case 'up':
return status?.active ? { variant: 'on', word: 'active' } : { variant: 'amber', word: 'idle' }
default:
return { variant: 'off', word: status ? 'not reported' : 'checking…' }
}
}
// ---------------------------------------------------------------------------
// Is the kill-switch actually blocking anything?
// ---------------------------------------------------------------------------
/**
* Four answers, and "unknown" is one of them.
*
* armed — configured fail-closed AND a data plane is installed to enforce it.
* inert — configured fail-closed, but there is no plane. Nothing is blocking.
* unknown — the daemon has not said how much plane is installed, so whether the
* setting is in force is not known. NEVER paint this green.
* open — configured fail-open. Nothing is meant to be blocked.
*/
export type KillSwitchState = 'armed' | 'inert' | 'unknown' | 'open'
export interface KillSwitchReadout {
state: KillSwitchState
/** The word the module puts in its readout. */
value: string
variant: LedVariant
/** Is it blocking right now — the row under the readout. `null` ⇒ nothing to add. */
blockingNow: string | null
/** True when `blockingNow` is bad news and should be drawn hot. */
hot: boolean
}
/**
* THE UNKNOWN BRANCH IS THE WHOLE POINT. This used to be
*
* killArmed && status?.plane !== 'none'
*
* and `undefined !== 'none'` is true — so a daemon that had not reported `plane`
* at all, and a panel that had not yet received its first status, both landed in
* the "ARMED" branch under a green lamp. Every other unknown in this file is an
* unlit socket for exactly this reason (see engineReadout): the kill-switch is
* the last thing standing between the LAN and the plain WAN, and "I don't know
* whether it is installed" must never be dressed as "it is".
*
* `configured` is the SAVED policy from /api/config, used only while
* /api/status has not reported one. The live value wins wherever it exists, as
* everywhere else in the panel: this is a status readout, and the config on disk
* can already differ from what is installed.
*/
export function killSwitchReadout(
status: Status | null,
configured?: string,
): KillSwitchReadout {
const closed = (status?.kill_switch ?? configured ?? 'closed') === 'closed'
if (!closed) {
return { state: 'open', value: 'OPEN', variant: 'amber', blockingNow: null, hot: false }
}
switch (status?.plane) {
case 'full':
case 'hold':
// Something is installed, so the fail-closed guard is really in the path.
return { state: 'armed', value: 'ARMED', variant: 'on', blockingNow: null, hot: false }
case 'none':
return {
state: 'inert',
value: 'NOT IN EFFECT',
variant: 'crit',
blockingNow: 'no — nothing installed',
hot: true,
}
default:
return {
state: 'unknown',
value: 'NOT REPORTED',
variant: 'off',
// Terse on purpose: this is a two-column readout row, and the long form
// wrapped onto three lines beside a one-word key.
blockingNow: status ? 'not known' : 'no reading yet',
hot: false,
}
}
}
/**
* `plane` + `engine_running` express the state more precisely than the three
* booleans the old status strip exposed (engine active / config enabled / nft
@@ -31,6 +160,24 @@ export interface ProtectionState {
* hold — the kill-switch caught it. Protected, but offline.
* none (fail-closed) — there is no protection at all. Online, and exposed.
* Collapsing them would erase the only difference that matters.
*
* PLANE IS NOT THE WHOLE ANSWER, AND THAT USED TO BE A LIE. `plane: 'full'`
* returned "Protected — traffic from your network is going through the tunnel",
* which is a claim `plane` cannot support: it only says the nft table, the policy
* routing and the engine are all installed. Where the diverted packets go once the
* engine has them is decided by the engine's default route, and a router in the
* field ran with one rule — `default → direct`, no groups, no rule-sets. Fully
* installed plane, zero tunnel, whole LAN out the plain WAN with its real address,
* green LED, "Protected".
*
* So `full` now branches on `status.traffic`, the daemon's verdict on the config
* it is actually running (see api.ts TrafficVerdict). It is computed on the daemon
* because only the daemon knows what was GENERATED and STARTED: the panel's
* /api/config is desired state, which diverges from the running one whenever edits
* are unapplied or a rollback is pending, and re-deriving the default route from it
* would mean a second implementation of the generator's rule loop — schedules,
* shadowed catch-alls, targets that failed to resolve and fell back to direct —
* drifting against the first.
*/
export function protectionState(status: Status | null): ProtectionState {
if (!status) {
@@ -56,12 +203,7 @@ export function protectionState(status: Status | null): ProtectionState {
switch (status.plane) {
case 'full':
return {
variant: 'on',
headline: 'Protected',
detail: 'Traffic from your network is going through the tunnel.',
alarm: false,
}
return fullPlaneState(status.traffic)
case 'hold':
return {
variant: 'amber',
@@ -88,8 +230,18 @@ export function protectionState(status: Status | null): ProtectionState {
}
}
// Older daemon with no `plane` field: fall back to what we can observe.
if (status.running && status.active && status.table) {
// Older daemon with no `plane` field: fall back to what we can observe. The
// engine's own state is asked FIRST — "stopped" is the answer that matters, and
// reading it off `running` is what used to make it unreachable (see engineState).
if (engineState(status) === 'down') {
return {
variant: 'crit',
headline: 'Service stopped',
detail: 'The engine isn’t running, so traffic isn’t being proxied or filtered.',
alarm: true,
}
}
if (status.active && status.table) {
return {
variant: 'on',
headline: 'Protected',
@@ -97,14 +249,6 @@ export function protectionState(status: Status | null): ProtectionState {
alarm: false,
}
}
if (!status.running) {
return {
variant: 'crit',
headline: 'Service stopped',
detail: 'The service isn’t running, so traffic isn’t being handled.',
alarm: true,
}
}
return {
variant: 'amber',
headline: 'Starting up',
@@ -112,3 +256,87 @@ export function protectionState(status: Status | null): ProtectionState {
alarm: false,
}
}
/**
* The plane is fully installed — now say where the traffic it carries ends up.
*
* Only `tunnel` earns "Protected". The other verdicts each describe a real router
* someone can be sitting in front of, and they are kept apart because the thing to
* DO about them differs:
*
* split — deliberate for most people who reach it, accidental for the rest
* (a default rule that was never pointed anywhere). Amber, but no
* alarm: raising a banner on every page of a working selective setup
* is how a banner stops being read.
* direct — the engine is running and forwarding every connection out the plain
* WAN. The traffic outcome is identical to `plane: 'none'` under a
* closed kill-switch, so it gets the same weight: crit, and it
* interrupts. The wording differs because the fix does — nothing
* failed here, the routing simply says "direct".
* blocked — the fail-closed default. Nothing is leaking, so this is never crit;
* with no tunnelling rules at all it means the network has no way out
* and someone should be told why.
* unknown — an older daemon, or the seconds between this daemon starting and its
* first apply. We do not know, so we do not claim. Saying "Protected"
* here is the exact bug being removed.
*/
function fullPlaneState(traffic: Traffic | undefined): ProtectionState {
const tunnelRules = traffic?.tunnel_rules ?? 0
switch (traffic?.verdict) {
case 'tunnel':
return {
variant: 'on',
headline: 'Protected',
detail: 'Traffic from your network is going through the tunnel.',
alarm: false,
}
case 'split':
return {
variant: 'amber',
headline: 'Partly protected — the rest goes out directly',
detail: `${ruleCount(tunnelRules)} through the tunnel. Everything they don’t match leaves through your normal internet connection, with your real address.`,
alarm: false,
}
case 'direct':
return {
variant: 'crit',
headline: 'Not protected — nothing is going through the tunnel',
detail:
'The service is running, but your routing sends every connection straight out your normal internet connection, with your real address. On the Routing page, point the default rule at a group or a node.',
alarm: true,
}
case 'blocked':
return tunnelRules > 0
? {
variant: 'amber',
headline: 'Partly protected — everything else is blocked',
detail: `${ruleCount(tunnelRules)} through the tunnel. Anything they don’t match is blocked instead of being let out, so nothing is leaving unprotected.`,
alarm: false,
}
: {
variant: 'amber',
headline: 'Nothing is getting out',
detail:
'No rule sends traffic anywhere, so every connection from your network is being blocked rather than let out unprotected. Add a default rule on the Routing page.',
alarm: true,
}
default:
return {
variant: 'amber',
headline: 'Checking where traffic goes',
detail:
'The router is up and handling your traffic. It hasn’t reported yet whether that traffic is going through the tunnel.',
alarm: false,
}
}
}
/** "One rule sends traffic" / "4 rules send traffic", so the detail lines above
* can name a number the operator can go and count on the Routing page. Falls
* back to the vague form only if the daemon sent a verdict without a count. */
function ruleCount(n: number): string {
if (n <= 0) return 'Some traffic goes'
if (n === 1) return 'One rule sends traffic'
return `${n} rules send traffic`
}
+6 -1
View File
@@ -18,5 +18,10 @@
"noFallthroughCasesInSwitch": true,
"forceConsistentCasingInFileNames": true
},
"include": ["src", "vite.config.ts"]
"include": ["src", "vite.config.ts"],
// *.test.ts runs under node's built-in test runner (`npm test`), which strips
// types rather than checking them. They are excluded here because they import
// node:test / node:assert, and the SPA deliberately carries no @types/node — it
// is embedded in the daemon binary, so every devDependency is weight on a router.
"exclude": ["src/**/*.test.ts"]
}
+48 -1
View File
@@ -1,15 +1,62 @@
import { defineConfig } from 'vite'
import type { Plugin } from 'vite'
import react from '@vitejs/plugin-react'
// Minimal ambient for the dev-proxy target override — avoids pulling in @types/node
// just for one env read. Vite runs this file under Node where `process` exists.
declare const process: { env: Record<string, string | undefined> }
/** `src/mock.ts`, as the module graph spells it (POSIX-normalised for Windows). */
const MOCK_MODULE = 'src/mock.ts'
/**
* Refuse to emit a production bundle that contains the offline fixture backend.
*
* `src/mock.ts` describes an invented, healthy router: a full config, 122 nodes
* with 119 of them alive, "Protected". It exists so `npm run dev` renders without
* a daemon. It shipped inside the binary that goes on real hardware, switched on
* by nothing more than a `?dev` on the end of the URL — so a link someone was
* sent, or a bookmark they saved, showed an appliance in perfect health while
* making no request to the appliance at all.
*
* api.ts now loads it behind `import.meta.env.DEV`, which Vite folds to a literal
* `false` for a build, so Rollup drops the dynamic import and the module never
* enters the graph. That is a property of a build tool's optimiser, and an
* optimiser is not a promise: one refactor that makes the condition non-static
* silently puts the fixtures back. So the property is CHECKED rather than
* trusted — if `src/mock.ts` reaches any emitted chunk, the build fails here
* instead of shipping.
*/
function assertNoMockFixtures(): Plugin {
return {
name: 'shater:assert-no-mock-fixtures',
apply: 'build',
generateBundle(_options, bundle) {
const guilty: string[] = []
for (const [file, output] of Object.entries(bundle)) {
if (output.type !== 'chunk') continue
for (const id of output.moduleIds) {
if (id.replace(/\\/g, '/').endsWith(MOCK_MODULE)) guilty.push(`${file} ← ${id}`)
}
}
if (guilty.length > 0) {
this.error(
`the offline fixture backend (${MOCK_MODULE}) reached the production bundle:\n ` +
guilty.join('\n ') +
`\nFixtures describe a router that does not exist. Keep every path to them behind ` +
`\`import.meta.env.DEV\` so Rollup can drop them, and never gate them on a runtime ` +
`flag such as a query parameter.`,
)
}
},
}
}
// The SPA is embedded in the forked sing-box binary and served by the daemon on
// its own port. Relative base so it works under any mount path; single small
// bundle (no code-splitting) keeps the embed simple and the flash budget low.
export default defineConfig({
plugins: [react()],
plugins: [react(), assertNoMockFixtures()],
base: './',
build: {
outDir: 'dist',
-93
View File
@@ -1,93 +0,0 @@
// lx:begin awg
package group
import (
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
)
// suspendAmneziaWGConsumersOnWireGuardSwitch is called from Selector.SelectOutbound
// BEFORE the switch is committed. If the member about to be selected is — or chains
// down via detour to — a WireGuard-based endpoint (type "wireguard", covering plain
// WG and AmneziaWG), it walks UP from this group to every AmneziaWG endpoint that
// detours through it and suspends each one (brings its device down). Rationale:
// AmneziaWG traffic encapsulated inside a WireGuard tunnel hangs the kernel on
// Android; the static Start-guard cannot cover this because a selector's chosen
// member is only known at runtime.
//
// Called before s.selected is updated, so the race is closed: once the group
// points at the WireGuard member, the AmneziaWG consumers are already suspended
// (started=false) and a concurrent reconnect fails with "not ready" instead of
// sending a junk handshake into WireGuard.
func suspendAmneziaWGConsumersOnWireGuardSwitch(outboundManager adapter.OutboundManager, groupTag string, selected adapter.Outbound) {
if outboundManager == nil || groupTag == "" {
return
}
if !chainReachesWireGuard(outboundManager, selected, make(map[string]bool)) {
return
}
suspendAmneziaWGConsumers(outboundManager, groupTag, make(map[string]bool))
}
// chainReachesWireGuard reports whether outbound is — or transitively detours
// down to, or (being a group) contains a member that is — a WireGuard-based
// endpoint. visited guards against cycles.
func chainReachesWireGuard(outboundManager adapter.OutboundManager, outbound adapter.Outbound, visited map[string]bool) bool {
if outbound == nil {
return false
}
tag := outbound.Tag()
if tag != "" {
if visited[tag] {
return false
}
visited[tag] = true
}
if outbound.Type() == C.TypeWireGuard {
return true
}
// Down the detour chain (vless -> ... -> wireguard).
for _, dependency := range outbound.Dependencies() {
if member, loaded := outboundManager.Outbound(dependency); loaded {
if chainReachesWireGuard(outboundManager, member, visited) {
return true
}
}
}
// A nested group: any member reaching WireGuard counts.
if group, isGroup := outbound.(adapter.OutboundGroup); isGroup {
for _, memberTag := range group.All() {
if member, loaded := outboundManager.Outbound(memberTag); loaded {
if chainReachesWireGuard(outboundManager, member, visited) {
return true
}
}
}
}
return false
}
// suspendAmneziaWGConsumers walks UP from tag via the reverse-dependency ledger
// (ConsumersOf) and suspends every AmneziaWG endpoint that detours through it,
// directly or transitively (e.g. AWG -> vless -> group). visited guards cycles.
func suspendAmneziaWGConsumers(outboundManager adapter.OutboundManager, tag string, visited map[string]bool) {
for _, consumerTag := range outboundManager.ConsumersOf(tag) {
if visited[consumerTag] {
continue
}
visited[consumerTag] = true
consumer, loaded := outboundManager.Outbound(consumerTag)
if !loaded {
continue
}
if awg, isAWG := consumer.(adapter.AmneziaWGSuspendable); isAWG && awg.IsAmneziaWG() {
awg.SuspendAmneziaWG()
}
// Keep walking up: a non-AWG hop (vless) or a parent group may itself have
// an AmneziaWG consumer above it.
suspendAmneziaWGConsumers(outboundManager, consumerTag, visited)
}
}
// lx:end awg
-120
View File
@@ -1,120 +0,0 @@
// lx:begin awg
package group
import (
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
)
// fakeOutbound is a minimal adapter.Outbound; only Type/Tag/Dependencies are read.
type fakeOutbound struct {
adapter.Outbound
tag string
outboundTyp string
detour string
}
func (o *fakeOutbound) Type() string { return o.outboundTyp }
func (o *fakeOutbound) Tag() string { return o.tag }
func (o *fakeOutbound) Dependencies() []string {
if o.detour == "" {
return nil
}
return []string{o.detour}
}
// fakeAWG implements adapter.AmneziaWGSuspendable and records suspension.
type fakeAWG struct {
fakeOutbound
awg bool
suspended bool
}
func (a *fakeAWG) IsAmneziaWG() bool { return a.awg }
func (a *fakeAWG) SuspendAmneziaWG() { a.suspended = true }
// fakeManager resolves tags and reverse-deps (ConsumersOf) from fixed maps.
type fakeManager struct {
adapter.OutboundManager
byTag map[string]adapter.Outbound
consumers map[string][]string
}
func (m *fakeManager) Outbound(tag string) (adapter.Outbound, bool) {
ob, ok := m.byTag[tag]
return ob, ok
}
func (m *fakeManager) ConsumersOf(tag string) []string { return m.consumers[tag] }
func TestChainReachesWireGuard(t *testing.T) {
wg := &fakeOutbound{tag: "wg", outboundTyp: C.TypeWireGuard}
vlessToWG := &fakeOutbound{tag: "v2wg", outboundTyp: C.TypeVLESS, detour: "wg"}
vlessLeaf := &fakeOutbound{tag: "vleaf", outboundTyp: C.TypeVLESS}
mgr := &fakeManager{byTag: map[string]adapter.Outbound{
"wg": wg, "v2wg": vlessToWG, "vleaf": vlessLeaf,
}}
if !chainReachesWireGuard(mgr, wg, map[string]bool{}) {
t.Fatal("direct wireguard member must reach wireguard")
}
if !chainReachesWireGuard(mgr, vlessToWG, map[string]bool{}) {
t.Fatal("vless detouring to wireguard must reach wireguard")
}
if chainReachesWireGuard(mgr, vlessLeaf, map[string]bool{}) {
t.Fatal("plain vless must not reach wireguard")
}
}
func TestSuspendAmneziaWGConsumers(t *testing.T) {
// awg-direct detours through the group "sel"
awgDirect := &fakeAWG{fakeOutbound: fakeOutbound{tag: "awg-direct", outboundTyp: C.TypeWireGuard, detour: "sel"}, awg: true}
// awg-via-hop -> vless-hop -> sel
awgViaHop := &fakeAWG{fakeOutbound: fakeOutbound{tag: "awg-hop", outboundTyp: C.TypeWireGuard, detour: "vless-hop"}, awg: true}
vlessHop := &fakeOutbound{tag: "vless-hop", outboundTyp: C.TypeVLESS, detour: "sel"}
// plain-wg detours through sel but is NOT amneziawg — must stay untouched
plainWG := &fakeAWG{fakeOutbound: fakeOutbound{tag: "plain-wg", outboundTyp: C.TypeWireGuard, detour: "sel"}, awg: false}
mgr := &fakeManager{
byTag: map[string]adapter.Outbound{
"awg-direct": awgDirect, "awg-hop": awgViaHop,
"vless-hop": vlessHop, "plain-wg": plainWG,
},
consumers: map[string][]string{
"sel": {"awg-direct", "vless-hop", "plain-wg"},
"vless-hop": {"awg-hop"},
},
}
suspendAmneziaWGConsumers(mgr, "sel", map[string]bool{})
if !awgDirect.suspended {
t.Error("direct AmneziaWG consumer must be suspended")
}
if !awgViaHop.suspended {
t.Error("transitive AmneziaWG consumer (via vless hop) must be suspended")
}
if plainWG.suspended {
t.Error("plain (non-AmneziaWG) wireguard consumer must NOT be suspended")
}
}
// A switch to a non-wireguard member must suspend nothing.
func TestSuspendSkippedForNonWireGuardSwitch(t *testing.T) {
awg := &fakeAWG{fakeOutbound: fakeOutbound{tag: "awg", outboundTyp: C.TypeWireGuard, detour: "sel"}, awg: true}
vlessLeaf := &fakeOutbound{tag: "vleaf", outboundTyp: C.TypeVLESS}
mgr := &fakeManager{
byTag: map[string]adapter.Outbound{"awg": awg, "vleaf": vlessLeaf},
consumers: map[string][]string{"sel": {"awg"}},
}
// selected member is plain vless (does not reach wireguard) → no suspension
suspendAmneziaWGConsumersOnWireGuardSwitch(mgr, "sel", vlessLeaf)
if awg.suspended {
t.Error("must not suspend when the selected member does not reach wireguard")
}
}
// lx:end awg
-10
View File
@@ -128,16 +128,6 @@ func (s *Selector) SelectOutbound(tag string) bool {
if s.selected.Load() == detour {
return true
}
// lx:begin awg
// Suspend AmneziaWG consumers BEFORE switching: if the new member is (or chains
// to) a WireGuard endpoint, any AmneziaWG endpoint that detours through this
// group would tunnel AWG inside WireGuard and hang the kernel on Android. Doing
// this before s.selected.Swap closes the race — by the time the group points at
// the WireGuard member, those consumers are already down (started=false), so a
// concurrent reconnect fails with "not ready" instead of sending a junk
// handshake into WireGuard.
suspendAmneziaWGConsumersOnWireGuardSwitch(s.outbound, s.Tag(), detour)
// lx:end awg
s.selected.Store(detour)
invalidateReachability(s.ctx) // lx: SPEC 020 — active selection changed
if s.Tag() != "" {
+292 -45
View File
@@ -44,6 +44,20 @@ type URLTest struct {
group *URLTestGroup
interruptExternalConnections bool
balancer *balancer // lx: SPEC 019 — nil for least_test (default)
// lx: health board §5.C — true when options.SelfCheck == false: the group's
// OWN probing schedule (PostStart warm-up + Touch ticker) is stood down and
// the observatory is the only thing that measures its members. Stored
// INVERTED so the zero value keeps today's behaviour for every construction
// path that does not go through NewURLTest (hand-built groups in tests).
// See option.URLTestOutboundOptions.SelfCheck for the full reasoning.
selfCheckDisabled bool
// lx: health board §5.C — the RUNTIME half of the same question, read from
// the context registry at construction. selfCheckDisabled above says "this
// group is not used by the config"; this says "this group cannot be reached
// right now", which changes while the box runs and is therefore asked
// afresh at every scheduled check rather than stored. nil = no gate.
// See urltest.ProbeGate for why the two must stay separate.
probeGate urltest.ProbeGate
}
func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLogger, tag string, options option.URLTestOutboundOptions) (adapter.Outbound, error) {
@@ -71,6 +85,12 @@ func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLo
idleTimeout: time.Duration(options.IdleTimeout),
interruptExternalConnections: options.InterruptExistConnections,
balancer: balancer,
// nil/absent means true (self-check on) — the documented default, so a
// config written before the flag existed behaves exactly as it always has.
selfCheckDisabled: options.SelfCheck != nil && !*options.SelfCheck,
// Absent from the registry (plain sing-box, tests) yields nil, which
// means "no gate" — every scheduled probe proceeds, as before.
probeGate: service.FromContext[urltest.ProbeGate](ctx),
}
if len(outbound.tags) == 0 {
return nil, E.New("missing tags")
@@ -92,6 +112,13 @@ func (s *URLTest) Start() error {
return err
}
group.balancer = s.balancer // lx: SPEC 019 v2 — health-check drives the pool through it
// lx: health board §5.C — carry the stand-down flag onto the group the same
// way the balancer travels: set after construction, immutable from then on.
group.selfCheckDisabled = s.selfCheckDisabled
// The gate and the tag to ask it about travel together: the gate answers
// per-outbound, and the group is the thing whose schedule is being gated.
group.probeGate = s.probeGate
group.tag = s.Tag()
if s.balancer != nil {
// lx: health board §5.B — slot liveness reads through the board verdict, so a
// death recorded by any prober or a failed dial takes effect on the next pick,
@@ -122,12 +149,14 @@ func (s *URLTest) Now() string {
if s.balancer != nil {
return s.group.lastSelected.Load()
}
if s.group.selectedOutboundTCP != nil {
return s.group.selectedOutboundTCP.Tag()
} else if s.group.selectedOutboundUDP != nil {
return s.group.selectedOutboundUDP.Tag()
// One load, so the two halves reported here are the SAME decision.
selected := s.group.selected.Load()
if selected.tcp != nil {
return selected.tcp.Tag()
} else if selected.udp != nil {
return selected.udp.Tag()
}
// lx: SPEC 019 — cold start: before the first URL-test, selectedOutbound* is nil but
// lx: SPEC 019 — cold start: before the first URL-test the pair is empty but
// traffic already flows via the Select() fallback (outbounds[0] when no history yet).
// Mirror exactly what the next DialContext would pick, so the UI shows the real node
// instead of blank. Select() is the same source of truth DialContext uses.
@@ -296,29 +325,155 @@ func (s *URLTest) NewPacketConnection(ctx context.Context, conn N.PacketConn, me
}
type URLTestGroup struct {
ctx context.Context
outbound adapter.OutboundManager
pause pause.Manager
pauseCallback *list.Element[pause.Callback]
logger log.Logger
outbounds []adapter.Outbound
link string
interval time.Duration
tolerance uint16
idleTimeout time.Duration
history *urltest.HistoryStorage
checking atomic.Bool
selectedOutboundTCP adapter.Outbound
selectedOutboundUDP adapter.Outbound
ctx context.Context
outbound adapter.OutboundManager
pause pause.Manager
pauseCallback *list.Element[pause.Callback]
logger log.Logger
outbounds []adapter.Outbound
link string
interval time.Duration
tolerance uint16
idleTimeout time.Duration
history *urltest.HistoryStorage
checking atomic.Bool
// selected is the least_test cache: the member this group currently prefers, per
// network. It is written by the probing goroutine and read on EVERY dial through
// the group (selectExcluding / dialSelect) and by the panel (Now), so it is an
// atomic value rather than two plain fields — the same thing Selector does one file
// over (selector.go, common.TypedValue[adapter.Outbound]). An interface field is two
// words; a torn read of one hands a dial a type descriptor with the wrong data
// pointer, which is not a wrong node but a corrupt one.
//
// The TCP and UDP halves live in ONE value on purpose. They are decided together, by
// one pass over one board reading, and publishing them separately let a reader pick
// up the new TCP choice against the previous UDP choice — the group's hysteresis
// silently applied to a decision that was never made.
selected common.TypedValue[selectedPair]
interruptGroup *interrupt.Group
interruptExternalConnections bool
access sync.Mutex
ticker *time.Ticker
close chan struct{}
started bool
lastActive common.TypedValue[time.Time]
lastSelected common.TypedValue[string] // lx: SPEC 019 — Now() in balanced modes
balancer *balancer // lx: SPEC 019 v2 — round_robin pool; nil for least_test
// started is read by Touch on every dial, outside g.access, and written by
// PostStart under it — an atomic because that is what it always was in effect.
started atomic.Bool
// closed latches in Close and is what makes Close FINAL. Guarded by access.
//
// It exists because "has a ticker" is not the same question as "is shut down", and
// Close used to ask the first one: with no ticker armed it returned before closing
// g.close, leaving the group indistinguishable from a running one. A Touch arriving
// afterwards — an outbound snapshot taken before an Apply is still dialable for up
// to two minutes, see shater/engine/grouptest.go — then armed a fresh ticker whose
// loopCheck waits on a channel nobody will ever close, in a box whose context is
// already cancelled. Every tick of it fails instantly and files a "dead" verdict on
// the SHARED health board that the live generation selects nodes from. One retired
// group can go on declaring the whole node set dead for the uptime of the daemon.
closed bool
lastActive common.TypedValue[time.Time]
lastSelected common.TypedValue[string] // lx: SPEC 019 — Now() in balanced modes
balancer *balancer // lx: SPEC 019 v2 — round_robin pool; nil for least_test
// lx: health board §5.C — mirrors URLTest.selfCheckDisabled (set by Start,
// immutable afterwards, zero value = probing on). Guards ONLY the group's
// own schedule: the PostStart warm-up sweep and the Touch ticker. An
// explicit CheckOutbounds/URLTest call is untouched — the flag stands down
// the schedule, not the capability.
selfCheckDisabled bool
// lx: health board §5.C — the runtime gate and the tag it is asked about.
// Both mirror URLTest's fields (set by Start, immutable afterwards); nil
// gate or empty tag means every scheduled check proceeds. Consulted only
// through selfCheckAllowed, and only on the SCHEDULE.
probeGate urltest.ProbeGate
tag string
}
// selectedPair is one published least_test decision: the member chosen for TCP and the
// member chosen for UDP, as of the same probing round. Either half may be nil (nothing
// picked yet for that network).
type selectedPair struct {
tcp adapter.Outbound
udp adapter.Outbound
}
// selectedFor returns the cached choice for one network (nil when there is none, or when
// network is neither TCP nor UDP — the caller then falls through to a fresh selection).
func (g *URLTestGroup) selectedFor(network string) adapter.Outbound {
pair := g.selected.Load()
switch network {
case N.NetworkTCP:
return pair.tcp
case N.NetworkUDP:
return pair.udp
}
return nil
}
// setSelected publishes a decision. It is the ONLY writer of g.selected, and it writes
// the pair whole — see the field comment for why the two halves may not be split.
func (g *URLTestGroup) setSelected(tcp, udp adapter.Outbound) {
g.selected.Store(selectedPair{tcp: tcp, udp: udp})
}
// selfCheckAllowed reports whether the group's OWN probing schedule may dial
// right now. lx: health board §5.C.
//
// Two independent refusals, in the order they can be answered cheapest first:
//
// selfCheckDisabled — the config says no rule reaches this group. Fixed for
// the life of the box; see standDownUnusedSelfCheck.
// probeGate — the world says this group cannot be reached right now,
// typically a chain hop sitting behind a dead hop. Asked
// fresh EVERY time, which is the entire mechanism by which
// a recovered hop resumes probing: there is no state here
// to reset, so there is none to get stuck.
//
// Neither refusal touches an explicit CheckOutbounds/URLTest — a deliberate
// request is never a scheduled one.
func (g *URLTestGroup) selfCheckAllowed() bool {
if g.selfCheckDisabled {
return false
}
if g.probeGate == nil || g.tag == "" {
return true
}
return g.probeGate.ProbeAllowed(g.tag)
}
// scheduledCheck is one firing of the group's own schedule — the warm-up sweep
// and every ticker tick go through here, and nothing else does. Having exactly
// one gated entry point is what keeps the two callers from drifting apart, and
// it is the seam the tests drive to assert that a gated group makes no dial
// attempt at all.
func (g *URLTestGroup) scheduledCheck() {
if !g.selfCheckAllowed() {
return
}
g.CheckOutbounds(false)
}
// keepWarm reports whether this group must keep measuring with no traffic
// flowing through it. lx: health board §5.C — see urltest.ProbeGate.ProbeWhenIdle.
//
// The default is NO, in every direction: no gate, no tag, or a group whose
// self-check is stood down anyway. Only a gate that positively says "the routing
// config reaches this group" turns the idle timeout off, so plain sing-box and
// every hand-built group keep the lifecycle they have always had.
func (g *URLTestGroup) keepWarm() bool {
if g.selfCheckDisabled || g.probeGate == nil || g.tag == "" {
return false
}
return g.probeGate.ProbeWhenIdle(g.tag)
}
// startTickerLocked arms the group's own probing ticker. g.access MUST be held
// and g.ticker MUST be nil. Extracted so PostStart and Touch arm it identically
// — two ways in, one construction, no chance of one of them forgetting the pause
// registration.
func (g *URLTestGroup) startTickerLocked() {
ticker := time.NewTicker(g.interval)
g.ticker = ticker
g.pauseCallback = pause.RegisterTicker(g.pause, ticker, g.interval, nil)
go g.loopCheck(ticker, g.close)
}
func NewURLTestGroup(ctx context.Context, outboundManager adapter.OutboundManager, logger log.Logger, outbounds []adapter.Outbound, link string, interval time.Duration, tolerance uint16, idleTimeout time.Duration, interruptExternalConnections bool) (*URLTestGroup, error) {
@@ -358,41 +513,112 @@ func NewURLTestGroup(ctx context.Context, outboundManager adapter.OutboundManage
func (g *URLTestGroup) PostStart() {
g.access.Lock()
defer g.access.Unlock()
g.started = true
if g.closed {
return
}
g.started.Store(true)
g.lastActive.Store(time.Now())
// lx: SPEC 019 v2 — seed the pool so round_robin can route from the first connection,
// before the first health-check completes (history-warm nodes first, else config order).
// The seed only READS the board, so it runs even with the self-check stood down.
g.seedPool()
go g.CheckOutbounds(false)
// lx: health board §5.C — the warm-up sweep is the first half of the group's
// own probing schedule, and it fires for EVERY group at box start, including
// groups no routing rule reaches. For those, the sweep dials every member
// directly from the router — a path nothing uses — and records the outcome
// under the members' base tags, forging the board reading the observatory
// exists to keep honest. A stood-down group therefore skips it entirely; the
// observatory (or nothing, for a truly unused group) is what measures its
// members.
//
// The same call is now also where a chain hop behind a DEAD hop declines to
// sweep: every member of such a group dials through the broken hop, so the
// sweep would measure that hop once per member and file the result against
// this one. selfCheckAllowed keeps both refusals in one place.
go g.scheduledCheck()
// A group the routing config REACHES keeps measuring whether or not anybody
// dials it, so its ticker is armed here instead of waiting for a Touch that
// may never come. Without this, a used group with no traffic gets this one
// warm-up sweep and then nothing: its members age past the verdict TTL and
// the panel reports "untested" about a rule that is in force, while the first
// real request pays a cold probe. Nothing else would fill the gap — the
// observatory stands off a urltest group's members entirely (probeplan.go
// SelfChecked), which is the whole point of one dialler per target.
//
// lastActive was stored a moment ago, so loopCheck's opening "idle longer
// than the interval" check does not fire and this cannot double up with the
// sweep above.
if g.keepWarm() && g.ticker == nil {
g.startTickerLocked()
}
}
func (g *URLTestGroup) Touch() {
if !g.started {
if !g.started.Load() {
return
}
// lx: health board §5.C — Touch's only job is to keep the group's OWN
// probing ticker alive while traffic flows. With the self-check stood down
// there is deliberately no ticker to start or feed: the observatory owns the
// schedule, and a stray dial through an unused group (a stale rule cache, a
// manual pin) must not arm 30 minutes of direct probing under the members'
// base tags. Checked before the lock because the flag is immutable after
// Start, exactly like the started fast-path above.
//
// The runtime gate is deliberately NOT consulted here. Touch only arms the
// ticker; refusing to arm it would mean a hop that recovers has no ticker
// left to notice — the block would outlive the failure, which is the one
// outcome this must never have. The ticker runs and each tick re-asks the
// gate (loopCheck -> scheduledCheck), so a blocked hop costs a predicate
// call per interval and resumes the moment the hop in front answers.
if g.selfCheckDisabled {
return
}
g.access.Lock()
defer g.access.Unlock()
// A closed group arms nothing. Touch is reachable long after Close — a caller
// holding an outbound from a snapshot taken before an Apply keeps dialling it (up
// to the 120s budget of shater/engine/grouptest.go) — and the ticker it would arm
// has no way left to stop: see the `closed` field for what that costs.
if g.closed {
return
}
if g.ticker != nil {
g.lastActive.Store(time.Now())
return
}
ticker := time.NewTicker(g.interval)
g.ticker = ticker
g.pauseCallback = pause.RegisterTicker(g.pause, ticker, g.interval, nil)
go g.loopCheck(ticker, g.close)
g.startTickerLocked()
}
// Close shuts the group down for good. It is idempotent, and it is FINAL: no later Touch
// can bring the probing schedule back.
//
// It used to return early when no ticker happened to be armed, without ever closing
// g.close — so a group that was closed while idle stayed, from the point of view of every
// other method, a perfectly live group. That is the whole defect: the close channel is the
// only way a loopCheck goroutine ever exits (its idle-timeout escape does not fire for a
// group the routing config reaches, keepWarm), so a ticker armed after such a Close is
// immortal, and every one of its ticks writes a failure to the shared health board on
// behalf of a box that no longer exists.
func (g *URLTestGroup) Close() error {
g.access.Lock()
defer g.access.Unlock()
if g.ticker == nil {
if g.closed {
return nil
}
g.ticker.Stop()
g.ticker = nil
g.pause.UnregisterCallback(g.pauseCallback)
g.pauseCallback = nil
close(g.close)
g.closed = true
// Unconditionally, BEFORE looking at the ticker: this is the signal every loopCheck
// waits on, including any that a Touch armed after the last one was retired by the
// idle timeout.
if g.close != nil {
close(g.close)
}
if g.ticker != nil {
g.ticker.Stop()
g.ticker = nil
g.pause.UnregisterCallback(g.pauseCallback)
g.pauseCallback = nil
}
return nil
}
@@ -405,10 +631,15 @@ func (g *URLTestGroup) Select(network string) (adapter.Outbound, bool) {
return g.selectExcluding(network, nil)
}
// loopCheck is the group's own schedule. lx: health board §5.C — every probe it
// fires goes through scheduledCheck, so a stood-down or currently-unreachable
// group ticks without dialling. The ticker's LIFECYCLE (the idle timeout below)
// is deliberately left alone: a gated group keeps its ticker exactly as long as
// an ungated one would, because the ticker is what will notice the recovery.
func (g *URLTestGroup) loopCheck(ticker *time.Ticker, closeChan <-chan struct{}) {
if time.Since(g.lastActive.Load()) > g.interval {
g.lastActive.Store(time.Now())
g.CheckOutbounds(false)
g.scheduledCheck()
}
for {
select {
@@ -416,7 +647,13 @@ func (g *URLTestGroup) loopCheck(ticker *time.Ticker, closeChan <-chan struct{})
return
case <-ticker.C:
}
if time.Since(g.lastActive.Load()) > g.idleTimeout {
// The idle timeout retires the ticker of a group nobody is dialling —
// unless the routing config reaches it, in which case its health is a
// live question whether or not traffic is flowing and the ticker must
// outlive the silence. Asked here rather than remembered from PostStart
// so it tracks the running config, and asked OUTSIDE g.access because the
// answer comes from the engine, which has locks of its own.
if !g.keepWarm() && time.Since(g.lastActive.Load()) > g.idleTimeout {
g.access.Lock()
if g.ticker == ticker {
g.ticker.Stop()
@@ -427,7 +664,7 @@ func (g *URLTestGroup) loopCheck(ticker *time.Ticker, closeChan <-chan struct{})
g.access.Unlock()
return
}
g.CheckOutbounds(false)
g.scheduledCheck()
}
}
@@ -525,19 +762,29 @@ func (g *URLTestGroup) testNodes(ctx context.Context, outbounds []adapter.Outbou
return result
}
// performUpdateCheck re-ranks the members after a probing round and publishes the result.
// It is the only writer of g.selected: it reads the current pair ONCE, decides both
// networks against that one snapshot, and stores the outcome as a single value, so no
// reader can ever observe a half-applied decision. Callers are serialised by g.checking
// (urlTest), which is what makes the read-decide-store sequence safe without a lock.
func (g *URLTestGroup) performUpdateCheck() {
current := g.selected.Load()
next := current
var updated bool
if outbound, exists := g.Select(N.NetworkTCP); outbound != nil && (g.selectedOutboundTCP == nil || (exists && outbound != g.selectedOutboundTCP)) {
if g.selectedOutboundTCP != nil {
if outbound, exists := g.Select(N.NetworkTCP); outbound != nil && (current.tcp == nil || (exists && outbound != current.tcp)) {
if current.tcp != nil {
updated = true
}
g.selectedOutboundTCP = outbound
next.tcp = outbound
}
if outbound, exists := g.Select(N.NetworkUDP); outbound != nil && (g.selectedOutboundUDP == nil || (exists && outbound != g.selectedOutboundUDP)) {
if g.selectedOutboundUDP != nil {
if outbound, exists := g.Select(N.NetworkUDP); outbound != nil && (current.udp == nil || (exists && outbound != current.udp)) {
if current.udp != nil {
updated = true
}
g.selectedOutboundUDP = outbound
next.udp = outbound
}
if next != current {
g.setSelected(next.tcp, next.udp)
}
if updated {
g.interruptGroup.Interrupt(g.interruptExternalConnections)
+2 -14
View File
@@ -74,13 +74,7 @@ func (g *URLTestGroup) selectExcluding(network string, exclude map[string]bool)
var minOutbound adapter.Outbound
// Keep the upstream hysteresis: the currently selected outbound only yields to a
// member faster by more than tolerance — but only while it is still alive itself.
var current adapter.Outbound
switch network {
case N.NetworkTCP:
current = g.selectedOutboundTCP
case N.NetworkUDP:
current = g.selectedOutboundUDP
}
current := g.selectedFor(network)
if current != nil {
currentTag := RealTag(current)
if !exclude[currentTag] && g.history.Verdict(currentTag, ttl) == urltest.VerdictAlive {
@@ -144,13 +138,7 @@ func (s *URLTest) dialSelect(ctx context.Context, network string, destination M.
if s.balancer != nil {
return s.selectBalanced(ctx, network, destination, tried)
}
var outbound adapter.Outbound
switch N.NetworkName(network) {
case N.NetworkTCP:
outbound = s.group.selectedOutboundTCP
case N.NetworkUDP:
outbound = s.group.selectedOutboundUDP
}
outbound := s.group.selectedFor(N.NetworkName(network))
if outbound != nil {
realTag := RealTag(outbound)
if !tried[realTag] && s.group.history.Verdict(realTag, s.group.healthTTL()) != urltest.VerdictDead {
+33 -10
View File
@@ -7,6 +7,7 @@ import (
"context"
"errors"
"net"
"sync"
"testing"
"time"
@@ -32,10 +33,31 @@ type healthNode struct {
func (n *healthNode) Tag() string { return n.tag }
func (n *healthNode) Network() []string { return []string{N.NetworkTCP, N.NetworkUDP} }
func (n *healthNode) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
if n.dialed != nil {
*n.dialed = append(*n.dialed, n.tag)
// healthDialMu guards the shared dialed slice. testNodes probes a group's
// members CONCURRENTLY (a batch of 10), so two nodes pointing at one slice
// append from two goroutines; the append is what the -race build trips on, not
// the code under test.
var healthDialMu sync.Mutex
func (n *healthNode) record() {
if n.dialed == nil {
return
}
healthDialMu.Lock()
*n.dialed = append(*n.dialed, n.tag)
healthDialMu.Unlock()
}
// dialsOf reads a dial log under the same lock. Every assertion on a log a
// concurrent sweep may still be writing must go through it.
func dialsOf(dialed *[]string) []string {
healthDialMu.Lock()
defer healthDialMu.Unlock()
return append([]string(nil), *dialed...)
}
func (n *healthNode) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
n.record()
if n.fail {
return nil, errors.New("dial refused")
}
@@ -45,9 +67,7 @@ func (n *healthNode) DialContext(ctx context.Context, network string, destinatio
}
func (n *healthNode) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
if n.dialed != nil {
*n.dialed = append(*n.dialed, n.tag)
}
n.record()
if n.fail {
return nil, errors.New("listen refused")
}
@@ -85,6 +105,9 @@ func healthTestGroup(hist *urltest.HistoryStorage, manager adapter.OutboundManag
tolerance: 50,
logger: log.NewNOPFactory().Logger(),
interruptGroup: interrupt.NewGroup(),
// Real groups always have this channel (NewURLTestGroup); it is what Close
// signals every loopCheck through, so a hand-built group needs it too.
close: make(chan struct{}),
}
}
@@ -161,7 +184,7 @@ func TestSelectHysteresisKeepsAliveCurrent(t *testing.T) {
hist := urltest.NewHistoryStorage()
a, b := &balNode{tag: "a"}, &balNode{tag: "b"}
g := healthTestGroup(hist, nil, a, b)
g.selectedOutboundTCP = a
g.setSelected(a, nil)
storeAlive(hist, "a", 100)
storeAlive(hist, "b", 60) // within tolerance (100 ≤ 60+50) → keep a
if selected, _ := g.Select(N.NetworkTCP); selected != adapter.Outbound(a) {
@@ -178,7 +201,7 @@ func TestSelectDeadCurrentLosesToAlive(t *testing.T) {
hist := urltest.NewHistoryStorage()
a, b := &balNode{tag: "a"}, &balNode{tag: "b"}
g := healthTestGroup(hist, nil, a, b)
g.selectedOutboundTCP = a
g.setSelected(a, nil)
storeAlive(hist, "a", 10)
hist.MarkFailed("a")
storeAlive(hist, "b", 500)
@@ -331,7 +354,7 @@ func TestDialContextRetriesThroughNextAlive(t *testing.T) {
s := healthURLTest(g, nil, nil)
storeAlive(hist, "a", 10)
storeAlive(hist, "b", 100)
g.selectedOutboundTCP = a // the checker had picked a; it dies between ticks
g.setSelected(a, nil) // the checker had picked a; it dies between ticks
conn, err := s.DialContext(context.Background(), N.NetworkTCP, destDomain("example.com"))
if err != nil {
t.Fatalf("DialContext failed despite live member b: %v", err)
@@ -427,7 +450,7 @@ func TestListenPacketRetriesBeforeFirstSend(t *testing.T) {
s := healthURLTest(g, nil, nil)
storeAlive(hist, "a", 10)
storeAlive(hist, "b", 100)
g.selectedOutboundUDP = a
g.setSelected(nil, a)
conn, err := s.ListenPacket(context.Background(), destDomain("example.com"))
if err != nil {
t.Fatalf("ListenPacket failed despite live member b: %v", err)
+330
View File
@@ -0,0 +1,330 @@
package group
// Concurrency tests for the urltest group: the group's own probing schedule running at
// the same time as traffic going through it.
//
// This combination had no coverage at all. Every existing test either probes OR dials,
// never both at once, so the race detector had nothing to detect: the cached least_test
// choice was written by the prober goroutine and read on every single dial, with no
// synchronisation whatsoever, and the suite stayed green for as long as those two things
// never happened in the same test.
//
// An unsynchronised interface field is not a "usually fine" race. It is two words — type
// descriptor and data pointer — and a reader that catches the store half way holds a
// descriptor addressing the wrong value. What comes out is not a suboptimal node, it is a
// corrupt one, on the path of every connection the group carries.
import (
"context"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/common/urltest"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/service/pause"
)
// TestURLTestDialRacesProbeTicker drives real dials through a least_test group while the
// group's own probing schedule keeps re-deciding which member to use — the situation on
// every router where a urltest group carries traffic, since the ticker fires on its own
// interval regardless of what the connections are doing.
//
// It is a -race test first and an assertion test second: the failure it was written for
// is reported by the detector, not by a wrong value. Run it under -race or it proves
// almost nothing (the gate's [4/4] pass does).
func TestURLTestDialRacesProbeTicker(t *testing.T) {
hist := urltest.NewHistoryStorage()
a, b := &healthNode{tag: "a"}, &healthNode{tag: "b"}
manager := managerOf(a, b)
g := healthTestGroup(hist, manager, a, b)
s := healthURLTest(g, nil, manager)
storeAlive(hist, "a", 20)
storeAlive(hist, "b", 500)
var proberWG, dialWG sync.WaitGroup
stop := make(chan struct{})
// The prober: one full turn of the group's own schedule per iteration. CheckOutbounds
// is the real tick (probe every member, then publish); the probes cannot reach
// anything from a unit test, so the board is then re-armed with a winner that MOVES
// and the publish step is run again — otherwise the cached choice is written once and
// the window in which a reader can catch a torn write is a few nanoseconds wide.
proberWG.Add(1)
go func() {
defer proberWG.Done()
for i := 0; ; i++ {
select {
case <-stop:
return
default:
}
g.CheckOutbounds(true)
fast, slow := "a", "b"
if i%2 == 1 {
fast, slow = "b", "a"
}
storeAlive(hist, fast, 20)
storeAlive(hist, slow, 500)
g.performUpdateCheck()
}
}()
// The traffic: every dial reads the cached choice (dialSelect), and so does the panel
// (Now). Both are the read side of the race.
var dials, nows atomic.Int64
for range 4 {
dialWG.Add(1)
go func() {
defer dialWG.Done()
for range 300 {
if conn, err := s.DialContext(context.Background(), N.NetworkTCP, M.Socksaddr{}); err == nil {
_ = conn.Close()
dials.Add(1)
}
if pc, err := s.ListenPacket(context.Background(), M.Socksaddr{}); err == nil {
_ = pc.Close()
}
// The panel polls this while everything above is happening.
if tag := s.Now(); tag != "" && tag != "a" && tag != "b" {
t.Errorf("Now() = %q, which is not a member of the group — a torn read of the cached choice", tag)
}
nows.Add(1)
}
}()
}
// The dialers are the bounded side; the prober runs until they are done.
dialersDone := make(chan struct{})
go func() { dialWG.Wait(); close(dialersDone) }()
select {
case <-dialersDone:
case <-time.After(60 * time.Second):
close(stop)
proberWG.Wait()
t.Fatal("dialers did not finish — the group deadlocked against its own prober")
}
close(stop)
proberWG.Wait()
if dials.Load() == 0 {
t.Fatal("no dial succeeded — the test never exercised the read side it exists to race")
}
if nows.Load() == 0 {
t.Fatal("Now() was never polled")
}
}
// TestSelectedPairPublishedTogether pins the pairing half of the same defect: the TCP and
// UDP choices are one decision, taken from one board reading, and they become visible
// together. They used to be two separate field writes, so a reader could take the new TCP
// choice against the previous UDP one — a combination no probing round ever decided, and
// the group's hysteresis silently applied to it.
func TestSelectedPairPublishedTogether(t *testing.T) {
hist := urltest.NewHistoryStorage()
a, b := &healthNode{tag: "a"}, &healthNode{tag: "b"}
g := healthTestGroup(hist, managerOf(a, b), a, b)
// Round 1: a wins both networks.
storeAlive(hist, "a", 20)
storeAlive(hist, "b", 500)
g.performUpdateCheck()
if got := g.selected.Load(); got.tcp != adapter.Outbound(a) || got.udp != adapter.Outbound(a) {
t.Fatalf("after round 1 the pair is (%v, %v), want (a, a)", tagOrNil(got.tcp), tagOrNil(got.udp))
}
// Round 2: b wins both, by more than the tolerance.
storeAlive(hist, "a", 500)
storeAlive(hist, "b", 20)
g.performUpdateCheck()
got := g.selected.Load()
if got.tcp != adapter.Outbound(b) || got.udp != adapter.Outbound(b) {
t.Fatalf("after round 2 the pair is (%v, %v), want (b, b) — both halves move together",
tagOrNil(got.tcp), tagOrNil(got.udp))
}
// And what the dial path reads per network agrees with the published pair.
if g.selectedFor(N.NetworkTCP) != got.tcp || g.selectedFor(N.NetworkUDP) != got.udp {
t.Fatal("selectedFor disagrees with the published pair")
}
if g.selectedFor("icmp") != nil {
t.Fatal("selectedFor on an unknown network must yield nothing, not a TCP choice")
}
}
// TestURLTestGroupProbeRacesPanelRead is the narrower of the pair: the panel's Now() poll
// against the prober, with no dialling at all. shater/engine/grouphealth.go,
// shater/stats/stats.go and shater/engine/grouptest.go all call Now() from their own
// goroutines while the group's ticker runs.
func TestURLTestGroupProbeRacesPanelRead(t *testing.T) {
hist := urltest.NewHistoryStorage()
a, b := &healthNode{tag: "a"}, &healthNode{tag: "b"}
manager := managerOf(a, b)
g := healthTestGroup(hist, manager, a, b)
s := healthURLTest(g, nil, manager)
var wg sync.WaitGroup
stop := make(chan struct{})
wg.Add(1)
go func() {
defer wg.Done()
for i := 0; ; i++ {
select {
case <-stop:
return
default:
}
fast, slow := "a", "b"
if i%2 == 1 {
fast, slow = "b", "a"
}
storeAlive(hist, fast, 20)
storeAlive(hist, slow, 500)
g.performUpdateCheck()
}
}()
for range 3 {
wg.Add(1)
go func() {
defer wg.Done()
for range 2000 {
_ = s.Now()
}
}()
}
time.Sleep(50 * time.Millisecond)
close(stop)
wg.Wait()
}
// tagOrNil renders a possibly-nil outbound for a failure message.
func tagOrNil(o adapter.Outbound) string {
if o == nil {
return "<nil>"
}
return o.Tag()
}
// --- Close is final ---------------------------------------------------------
// TestGroupCloseIsFinalForALaterTouch is the immortal-ticker regression.
//
// Close used to return early whenever no ticker happened to be armed — which is the
// normal state of a group nobody is dialling — WITHOUT closing g.close. Nothing else
// records that a group was shut down (started is never cleared), so a Touch arriving
// afterwards armed a fresh ticker and a fresh loopCheck goroutine waiting on a channel
// that would never be closed. Its only other exit, the idle timeout, does not fire for a
// group the routing config reaches.
//
// A later Touch is not hypothetical: shater/engine/grouptest.go dials through outbounds
// taken from a snapshot at the start of a run and keeps doing so for up to 120s, so
// "press Test in the panel, then apply a config within two minutes" is enough. The
// retired group then probes forever through a cancelled context — every probe fails
// instantly — and files "dead" for its members on the SHARED health board that the LIVE
// generation picks nodes from.
func TestGroupCloseIsFinalForALaterTouch(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
var dialed []string
a := &healthNode{tag: "a", fail: true, dialed: &dialed}
g := healthTestGroup(hist, managerOf(a), a)
g.pause = pause.ManagerFromContext(pause.WithDefaultManager(context.Background()))
// Fast enough that a surviving ticker proves itself within the test's patience.
g.interval = 10 * time.Millisecond
g.idleTimeout = time.Hour
// Started, but idle: no ticker armed. This is the state Close mishandled.
g.started.Store(true)
if err := g.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
g.Touch()
g.access.Lock()
ticker := g.ticker
g.access.Unlock()
if ticker != nil {
t.Fatal("Touch armed a probing ticker on a CLOSED group — nothing can stop it: " +
"its loopCheck waits on a channel that will never be closed")
}
// The consequence, stated in the terms that actually hurt: no probe, so no forged
// verdict on the shared board.
time.Sleep(150 * time.Millisecond)
if got := dialsOf(&dialed); len(got) != 0 {
t.Fatalf("a closed group dialled %v — a retired generation is writing to the live health board", got)
}
if v := hist.Verdict("a", 10*time.Minute); v != urltest.VerdictUntested {
t.Fatalf("verdict(a) = %v after closing the group, want untested — the dead marks are forged", v)
}
}
// TestGroupCloseStopsAnArmedTicker keeps the original behaviour honest: when a ticker IS
// armed, Close still stops it, unregisters the pause callback and signals loopCheck.
func TestGroupCloseStopsAnArmedTicker(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
a := &healthNode{tag: "a", fail: true}
g := healthTestGroup(hist, managerOf(a), a)
g.pause = pause.ManagerFromContext(pause.WithDefaultManager(context.Background()))
g.interval = 10 * time.Millisecond
g.idleTimeout = time.Hour
g.started.Store(true)
g.lastActive.Store(time.Now())
g.Touch()
g.access.Lock()
armed := g.ticker != nil
g.access.Unlock()
if !armed {
t.Fatal("Touch did not arm the ticker on a live group")
}
if err := g.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
g.access.Lock()
stillArmed := g.ticker != nil
g.access.Unlock()
if stillArmed {
t.Fatal("Close left the ticker armed")
}
select {
case <-g.close:
default:
t.Fatal("Close did not signal loopCheck")
}
// Idempotent: a second Close must not close an already-closed channel (panic) or
// undo anything.
if err := g.Close(); err != nil {
t.Fatalf("second Close: %v", err)
}
}
// TestGroupPostStartAfterCloseDoesNothing: the other way a retired group can be woken.
// PostStart is called on every member of a box at start-up; a group closed by a racing
// shutdown must not be brought back by it.
func TestGroupPostStartAfterCloseDoesNothing(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
var dialed []string
a := &healthNode{tag: "a", fail: true, dialed: &dialed}
g := healthTestGroup(hist, managerOf(a), a)
g.pause = pause.ManagerFromContext(pause.WithDefaultManager(context.Background()))
g.interval = 10 * time.Millisecond
g.idleTimeout = time.Hour
_ = g.Close()
g.PostStart()
time.Sleep(150 * time.Millisecond)
if g.started.Load() {
t.Fatal("PostStart marked a closed group as started")
}
if got := dialsOf(&dialed); len(got) != 0 {
t.Fatalf("PostStart on a closed group ran the warm-up sweep: dialled %v", got)
}
}
+272
View File
@@ -0,0 +1,272 @@
package group
// lx: health board §5.C tests — SelfCheck stands the group's OWN probing
// schedule down: no PostStart warm-up sweep, no Touch ticker. The explicit
// CheckOutbounds path stays available, and the nil default keeps probing.
import (
"context"
"sync"
"testing"
"time"
"github.com/sagernet/sing-box/common/urltest"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/service/pause"
)
// waitForHistory polls until the store holds an entry for tag or the deadline
// passes; reports whether it appeared. PostStart's sweep runs on its own
// goroutine, so both directions of the assertion need a bounded wait.
func waitForHistory(hist *urltest.HistoryStorage, tag string, deadline time.Duration) bool {
stop := time.Now().Add(deadline)
for time.Now().Before(stop) {
if hist.LoadURLTestHistory(tag) != nil {
return true
}
time.Sleep(5 * time.Millisecond)
}
return false
}
// A group with the self-check stood down writes NOTHING to the history storage
// on PostStart: the warm-up sweep — which would dial the member directly from
// the router and mark the failure under its base tag — must not fire. And
// Touch, the other half of the schedule, must not start a ticker either.
func TestSelfCheckDisabledPostStartWritesNothing(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.selfCheckDisabled = true
g.PostStart()
// The absence of a write is the assertion, so give the (non-existent) sweep
// real time to have happened before declaring victory.
if waitForHistory(hist, "a", 150*time.Millisecond) {
t.Fatal("a stood-down group's PostStart wrote to the board; the warm-up sweep must not fire")
}
g.Touch()
g.access.Lock()
ticker := g.ticker
g.access.Unlock()
if ticker != nil {
t.Fatal("Touch armed the probing ticker on a stood-down group")
}
}
// The default (SelfCheck nil, i.e. the zero-value field on a hand-built group)
// keeps today's behaviour: PostStart's warm-up sweep runs and records the
// failing member on the board.
func TestSelfCheckDefaultStillProbesOnPostStart(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.PostStart()
if !waitForHistory(hist, "a", 5*time.Second) {
t.Fatal("default group's PostStart never probed; the self-check must stay on unless stood down")
}
if v := hist.Verdict("a", 10*time.Minute); v != urltest.VerdictDead {
t.Fatalf("verdict(a) = %v, want dead from the warm-up sweep", v)
}
}
// An EXPLICIT CheckOutbounds still probes a stood-down group: the flag
// suppresses the group's own schedule, never a deliberate request (the adapter
// interface a human or an API invokes on purpose).
func TestSelfCheckDisabledExplicitCheckStillProbes(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.selfCheckDisabled = true
g.CheckOutbounds(true)
if hist.LoadURLTestHistory("a") == nil {
t.Fatal("an explicit CheckOutbounds(true) did not probe; the flag must only stand down the schedule")
}
}
// The option → outbound plumbing: nil/absent means on, an explicit false means
// stood down, an explicit true means on. NewURLTest is the only place the
// option is read, so this is where a plumbing regression would hide.
func TestSelfCheckOptionPlumbing(t *testing.T) {
build := func(selfCheck *bool) *URLTest {
t.Helper()
opts := option.URLTestOutboundOptions{Outbounds: []string{"a"}}
opts.SelfCheck = selfCheck
ob, err := NewURLTest(context.Background(), nil, log.NewNOPFactory().Logger(), "t", opts)
if err != nil {
t.Fatalf("NewURLTest: %v", err)
}
return ob.(*URLTest)
}
if build(nil).selfCheckDisabled {
t.Fatal("nil SelfCheck must keep the self-check ON (the compatibility default)")
}
on, off := true, false
if build(&on).selfCheckDisabled {
t.Fatal("SelfCheck=true must keep the self-check on")
}
if !build(&off).selfCheckDisabled {
t.Fatal("SelfCheck=false must stand the self-check down")
}
}
// lx: health board §5.C — the RUNTIME gate (urltest.ProbeGate). The self-check
// flag above says "the config reaches nothing here"; the gate says "the path in
// front of this group is down right now". Both stand the SCHEDULE down; neither
// touches an explicit check; and only the gate is allowed to change its mind
// while the box runs.
// fakeGate answers from a mutable set of blocked tags, so one test can watch a
// group stop dialling and start again without rebuilding anything.
type fakeGate struct {
mu sync.Mutex
blocked map[string]bool
warm map[string]bool
asked int
}
func (g *fakeGate) ProbeAllowed(tag string) bool {
g.mu.Lock()
defer g.mu.Unlock()
g.asked++
return !g.blocked[tag]
}
func (g *fakeGate) ProbeWhenIdle(tag string) bool {
g.mu.Lock()
defer g.mu.Unlock()
return g.warm[tag]
}
func (g *fakeGate) set(tag string, blocked bool) {
g.mu.Lock()
defer g.mu.Unlock()
g.blocked[tag] = blocked
}
func (g *fakeGate) asks() int {
g.mu.Lock()
defer g.mu.Unlock()
return g.asked
}
// A gated group makes NO DIAL ATTEMPT on its own schedule — the assertion is on
// the attempt log, not on the board, because a probe that ran and failed leaves
// the same "nothing useful known" as one that never ran, and only the attempt
// log tells them apart. This is the waste half of the chain-hop fix: a hop
// sitting behind a dead hop would otherwise spend one probe timeout per member
// rediscovering the same broken hop.
func TestProbeGateBlocksScheduledDials(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
var dialed []string
a := &healthNode{tag: "chain-c-h3-a", fail: true, dialed: &dialed}
b := &healthNode{tag: "chain-c-h3-b", fail: true, dialed: &dialed}
gate := &fakeGate{blocked: map[string]bool{"chain-c-h3": true}}
g := healthTestGroup(hist, managerOf(a, b), a, b)
g.tag = "chain-c-h3"
g.probeGate = gate
// Touch arms a real ticker, so this group needs the two things
// healthTestGroup leaves out because nothing else in that suite starts one:
// a pause manager to register the ticker with, and the close channel Close
// shuts the loop down through.
g.pause = pause.ManagerFromContext(pause.WithDefaultManager(context.Background()))
g.close = make(chan struct{})
// The warm-up sweep: gated, so nothing is dialled. Give the (non-existent)
// sweep real time to have happened — the absence is the assertion.
g.PostStart()
if waitForHistory(hist, "chain-c-h3-a", 150*time.Millisecond) {
t.Fatal("a gated group's PostStart wrote to the board")
}
// A ticker tick, driven directly: this is the exact call loopCheck makes.
g.scheduledCheck()
if got := dialsOf(&dialed); len(got) != 0 {
t.Fatalf("gated group dialled %v; a hop behind a dead hop must not dial at all", got)
}
// Touch still arms the ticker. Refusing to arm it would leave a recovered
// hop with nothing to notice — the block would outlive the failure.
g.Touch()
g.access.Lock()
ticker := g.ticker
g.access.Unlock()
if ticker == nil {
t.Fatal("Touch did not arm the ticker on a gated group; nothing would be left to spot the recovery")
}
_ = g.Close()
// The hop in front comes back. Nothing is reset, nothing is reapplied — the
// next scheduled check simply asks again and gets a different answer.
gate.set("chain-c-h3", false)
before := gate.asks()
g.scheduledCheck()
if gate.asks() <= before {
t.Error("scheduledCheck did not re-ask the gate; a cached answer is a block that outlives its cause")
}
if len(dialsOf(&dialed)) == 0 {
t.Fatal("the group did not resume dialling after the hop in front recovered")
}
}
// An EXPLICIT check is a deliberate request and is never gated — the same rule
// SelfCheck already follows. The gate stands down the schedule, not the
// capability.
func TestProbeGateDoesNotBlockExplicitCheck(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
var dialed []string
a := &healthNode{tag: "chain-c-h3-a", fail: true, dialed: &dialed}
g := healthTestGroup(hist, managerOf(a), a)
g.tag = "chain-c-h3"
g.probeGate = &fakeGate{blocked: map[string]bool{"chain-c-h3": true}}
g.CheckOutbounds(true)
if len(dialsOf(&dialed)) == 0 {
t.Fatal("an explicit CheckOutbounds was refused by the gate")
}
}
// The two refusals are independent and compose the obvious way; and the absent
// cases (no gate at all, an ungated tag) leave today's behaviour untouched,
// which is what every plain sing-box config and every hand-built group relies
// on.
func TestSelfCheckAllowedCombinations(t *testing.T) {
hist := urltest.NewHistoryStorage()
defer hist.Close()
a := &healthNode{tag: "a"}
base := func() *URLTestGroup { return healthTestGroup(hist, managerOf(a), a) }
if g := base(); !g.selfCheckAllowed() {
t.Error("a plain group with no gate must probe (the zero value is the compatibility default)")
}
g := base()
g.selfCheckDisabled = true
g.tag, g.probeGate = "chain-c-h3", &fakeGate{blocked: map[string]bool{}}
if g.selfCheckAllowed() {
t.Error("an UNUSED group must stay down even when the path in front is fine")
}
g = base()
g.tag, g.probeGate = "chain-c-h3", &fakeGate{blocked: map[string]bool{"chain-c-h3": true}}
if g.selfCheckAllowed() {
t.Error("a group behind a dead hop must not run its schedule")
}
g = base()
g.tag, g.probeGate = "auto", &fakeGate{blocked: map[string]bool{"chain-c-h3": true}}
if !g.selfCheckAllowed() {
t.Error("an unrelated group was gated by another tag's block")
}
g = base()
g.probeGate = &fakeGate{blocked: map[string]bool{"": true}}
if !g.selfCheckAllowed() {
t.Error("a group with no tag must not be gated; there is nothing to ask about")
}
}
@@ -0,0 +1,129 @@
//go:build with_gvisor && with_awg
// lx: regression for the removal of the AmneziaWG-over-WireGuard start guard.
//
// The guard refused to bring up an AmneziaWG endpoint whose detour chain reached
// a WireGuard-based endpoint — and refused *silently*: Start returned nil with
// started=false, so the endpoint looked configured but every dial through it
// failed with "WireGuard is not ready yet". The root cause it protected against
// (a kernel hang on Android) is gone on this graft (ClientBind reserved-gate),
// and Android is not a supported platform here at all.
//
// This test builds a real AmneziaWG endpoint (junk + ranged magic headers) whose
// detour points at an outbound of type "wireguard", drives both start stages,
// and asserts the endpoint reports itself started. With the guard in place the
// first stage short-circuits and started stays false — this test fails.
package wireguard
import (
"context"
"crypto/rand"
"encoding/base64"
"net"
"net/netip"
"os"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
"github.com/sagernet/sing/service"
"github.com/sagernet/sing/service/pause"
)
// wgTypedOutbound is an adapter.Outbound that reports type "wireguard" — the hop
// the guard used to refuse to start behind. Dialling through it always fails:
// the point of the test is that the upper endpoint comes UP, not that it carries
// traffic (that is the job of the transport-level e2e stand).
type wgTypedOutbound struct {
adapter.Outbound
tag string
}
func (o *wgTypedOutbound) Type() string { return C.TypeWireGuard }
func (o *wgTypedOutbound) Tag() string { return o.tag }
func (o *wgTypedOutbound) Dependencies() []string { return nil }
func (o *wgTypedOutbound) DialContext(ctx context.Context, network string, destination M.Socksaddr) (net.Conn, error) {
return nil, os.ErrClosed
}
func (o *wgTypedOutbound) ListenPacket(ctx context.Context, destination M.Socksaddr) (net.PacketConn, error) {
return nil, os.ErrClosed
}
// startChainManager resolves tags from a fixed map. adapter.OutboundManager is
// embedded so this compiles against either shape of the interface.
type startChainManager struct {
adapter.OutboundManager
byTag map[string]adapter.Outbound
}
func (m *startChainManager) Outbound(tag string) (adapter.Outbound, bool) {
ob, loaded := m.byTag[tag]
return ob, loaded
}
func randomKey(t *testing.T) string {
t.Helper()
var key [32]byte
if _, err := rand.Read(key[:]); err != nil {
t.Fatal(err)
}
// Clamp so wireguard-go accepts it as a curve25519 private key.
key[0] &= 248
key[31] = (key[31] & 127) | 64
return base64.StdEncoding.EncodeToString(key[:])
}
// TestAmneziaWGOverWireGuardDetourStarts pins the invariant: an AmneziaWG
// endpoint detouring through a WireGuard hop must come up like any other.
func TestAmneziaWGOverWireGuardDetourStarts(t *testing.T) {
ctx := pause.WithDefaultManager(context.Background())
ctx = service.ContextWith[adapter.OutboundManager](ctx, &startChainManager{
byTag: map[string]adapter.Outbound{
"wg-hop": &wgTypedOutbound{tag: "wg-hop"},
},
})
options := option.WireGuardEndpointOptions{
MTU: 1280,
Address: badoption.Listable[netip.Prefix]{netip.MustParsePrefix("10.7.0.2/32")},
PrivateKey: randomKey(t),
Peers: []option.WireGuardPeer{{
Address: "10.9.9.9",
Port: 51820,
PublicKey: randomKey(t),
AllowedIPs: badoption.Listable[netip.Prefix]{netip.MustParsePrefix("0.0.0.0/0")},
}},
AmneziaWGOptions: option.AmneziaWGOptions{
Jc: 3,
Jmin: 8,
Jmax: 80,
S4: 16,
H1: "10-20",
H2: "30-40",
H3: "50-60",
H4: "70-80",
},
}
options.Detour = "wg-hop"
ep, err := NewEndpoint(ctx, nil, log.NewNOPFactory().NewLogger("wg-awg"), "wg-awg", options)
if err != nil {
t.Fatal("create amneziawg endpoint over a wireguard detour: ", err)
}
defer ep.Close()
if err = ep.Start(adapter.StartStateStart); err != nil {
t.Fatal("start stage: ", err)
}
if err = ep.Start(adapter.StartStatePostStart); err != nil {
t.Fatal("post-start stage: ", err)
}
if !ep.(*Endpoint).started.Load() {
t.Fatal("an amneziawg endpoint behind a wireguard hop must start; it is silently held down")
}
}
-100
View File
@@ -1,100 +0,0 @@
// lx:begin awg
package wireguard
import (
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
)
// fakeOutbound is a minimal adapter.Outbound for the start-guard chain walk:
// only Type() and Dependencies() (the detour) are consulted. The embedded
// interface is nil — any other method would panic, which never happens here.
type fakeOutbound struct {
adapter.Outbound
tag string
outboundTyp string
detour string
}
func (o *fakeOutbound) Type() string { return o.outboundTyp }
func (o *fakeOutbound) Tag() string { return o.tag }
func (o *fakeOutbound) Dependencies() []string {
if o.detour == "" {
return nil
}
return []string{o.detour}
}
// fakeGroup is an adapter.OutboundGroup (selector/urltest stand-in); the chain
// walk must stop at it without expanding All().
type fakeGroup struct {
fakeOutbound
members []string
}
func (g *fakeGroup) Now() string { return "" }
func (g *fakeGroup) All() []string { return g.members }
type fakeOutboundManager struct {
adapter.OutboundManager
byTag map[string]adapter.Outbound
}
func (m *fakeOutboundManager) Outbound(tag string) (adapter.Outbound, bool) {
ob, loaded := m.byTag[tag]
return ob, loaded
}
func TestAwgDetourChainReachesWireGuard(t *testing.T) {
mgr := &fakeOutboundManager{byTag: map[string]adapter.Outbound{
// AWG -> wg-out (direct)
"wg-out": &fakeOutbound{tag: "wg-out", outboundTyp: C.TypeWireGuard},
// AWG -> vless-hop -> wg-deep (transitive)
"vless-hop": &fakeOutbound{tag: "vless-hop", outboundTyp: C.TypeVLESS, detour: "wg-deep"},
"wg-deep": &fakeOutbound{tag: "wg-deep", outboundTyp: C.TypeWireGuard},
// AWG -> vless-leaf -> direct-leaf (no wireguard anywhere)
"vless-leaf": &fakeOutbound{tag: "vless-leaf", outboundTyp: C.TypeVLESS, detour: "direct-leaf"},
"direct-leaf": &fakeOutbound{tag: "direct-leaf", outboundTyp: C.TypeDirect},
// AWG -> sel (selector hiding a wireguard member) — walk must stop, return ""
"sel": &fakeGroup{
fakeOutbound: fakeOutbound{tag: "sel", outboundTyp: C.TypeSelector},
members: []string{"wg-out"},
},
// cyclic detour: a -> b -> a, no wireguard
"cyc-a": &fakeOutbound{tag: "cyc-a", outboundTyp: C.TypeVLESS, detour: "cyc-b"},
"cyc-b": &fakeOutbound{tag: "cyc-b", outboundTyp: C.TypeVLESS, detour: "cyc-a"},
}}
cases := []struct {
name string
start string
wantEmpty bool
wantTag string
}{
{"direct wireguard", "wg-out", false, "wg-out"},
{"transitive via vless", "vless-hop", false, "wg-deep"},
{"no wireguard in chain", "vless-leaf", true, ""},
{"selector in the middle is skipped", "sel", true, ""},
{"cyclic chain terminates", "cyc-a", true, ""},
{"unknown tag", "nope", true, ""},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := awgDetourChainReachesWireGuard(mgr, tc.start, make(map[string]bool))
if tc.wantEmpty {
if got != "" {
t.Fatalf("expected no wireguard in chain, got %q", got)
}
return
}
if got != tc.wantTag {
t.Fatalf("expected blocked-by %q, got %q", tc.wantTag, got)
}
})
}
}
// lx:end awg
+20 -118
View File
@@ -4,7 +4,6 @@ import (
"context"
"net"
"net/netip"
"strconv"
"sync"
"sync/atomic"
"time"
@@ -45,28 +44,14 @@ type Endpoint struct {
localAddresses []netip.Prefix
endpoint *wireguard.Endpoint
started atomic.Bool
// lx:begin awg
// awgActive marks this endpoint as running AmneziaWG (AmneziaWGOptions.IsSet());
// detour is its configured upstream tag. Start uses them to refuse to bring up
// an AmneziaWG-over-WireGuard chain, which hangs the kernel on Android — see
// awgDetourChainReachesWireGuard. The ledger lives here (not just in the dialer
// guard) because the hang happens synchronously in Start, before any dial.
awgActive bool
detour string
// awgChainBlocked is set by Start when the AmneziaWG-over-WireGuard guard
// fires: the device is left unstarted (started stays false) so no junk
// handshake runs and the kernel cannot hang, while the rest of the instance
// comes up. PostStart then skips this endpoint too.
awgChainBlocked bool
// lx:end awg
// lx:begin idle-suspend
// SPEC 020 idle-suspend state. lastActivity is the unix-nano timestamp of the
// last dial through this endpoint, stamped at PostStart and on every dial entry.
// idleAsleep is true while the endpoint is Down due to idle-suspend (distinct
// from a guard-suspend, which sets started=false and clears idleAsleep, so a
// guard-suspended endpoint fast-paths out of resumeOnDial and is never
// idle-woken). resumeMu serialises the idle tick's suspend decision, a dial's
// wake, and the AmneziaWG guard-suspend against one another.
// from a deliberately-stopped endpoint, which has started=false and
// idleAsleep=false, so it fast-paths out of resumeOnDial and is never
// idle-woken). resumeMu serialises the idle tick's suspend decision against a
// dial's wake.
lastActivity atomic.Int64
idleAsleep atomic.Bool
resumeMu sync.Mutex
@@ -74,6 +59,16 @@ type Endpoint struct {
}
func NewEndpoint(ctx context.Context, router adapter.Router, logger log.ContextLogger, tag string, options option.WireGuardEndpointOptions) (adapter.Endpoint, error) {
// lx: allow OS-level fragmentation of the OUTER UDP socket by default, the
// same opt-out direct/hysteria/hysteria2/tuic already take. Without it the
// dialer sets DF (IP_MTU_DISCOVER=IP_PMTUDISC_DO on linux), and an outer
// datagram over the path MTU — routine once anything is encapsulated: WG's
// own ~32 B header, AmneziaWG s4 transport junk, or this endpoint carrying a
// nested tunnel — is dropped by the kernel ("message too long") instead of
// fragmented, so the tunnel comes up and then carries nothing. An explicit
// `udp_fragment: false` on the node still restores DF (UDPFragment wins over
// UDPFragmentDefault in common/dialer).
options.UDPFragmentDefault = true
ep := &Endpoint{
Adapter: endpoint.NewAdapterWithDialerOptions(C.TypeWireGuard, tag, []string{N.NetworkTCP, N.NetworkUDP, N.NetworkICMP}, options.DialerOptions),
ctx: ctx,
@@ -81,10 +76,6 @@ func NewEndpoint(ctx context.Context, router adapter.Router, logger log.ContextL
dnsRouter: service.FromContext[adapter.DNSRouter](ctx),
logger: logger,
localAddresses: options.Address,
// lx:begin awg
awgActive: options.AmneziaWGOptions.IsSet(),
detour: options.Detour,
// lx:end awg
}
if options.Detour != "" && options.ListenPort != 0 {
return nil, E.New("`listen_port` is conflict with `detour`")
@@ -116,7 +107,8 @@ func NewEndpoint(ctx context.Context, router adapter.Router, logger log.ContextL
Dialer: outboundDialer,
CreateDialer: func(interfaceName string) N.Dialer {
return common.Must1(dialer.NewDefault(ctx, option.DialerOptions{
BindInterface: interfaceName,
BindInterface: interfaceName,
UDPFragmentDefault: true, // lx: same reason as above — this is the bind-to-interface twin of the outer socket
}))
},
Name: options.Name,
@@ -157,33 +149,6 @@ func NewEndpoint(ctx context.Context, router adapter.Router, logger log.ContextL
}
func (w *Endpoint) Start(stage adapter.StartStage) error {
// lx:begin awg
// Refuse to bring up an AmneziaWG endpoint whose detour chain reaches a
// WireGuard-based endpoint: encapsulating AWG (junk handshake) inside a
// WireGuard tunnel hangs the kernel on Android. The hang happens here, in the
// synchronous Start path (peer-domain resolution over the detour, then the
// device's junk handshake) — before any dial — so the lazy DetourDialer guard
// never gets a chance to fire. We must catch it at Start instead.
//
// Behaviour is "variant B": do NOT return an error (that would abort the whole
// instance start). Instead log, skip device startup, and leave started=false
// so the rest of the config comes up and every dial through this endpoint
// fails cleanly with "WireGuard is not ready yet". A selector/urltest in the
// middle hides the real target at start time, so the chain walk stops at a
// group and that case is left to the lazy DetourDialer guard at dial time.
if stage == adapter.StartStateStart && w.awgActive && w.detour != "" {
if outboundManager := service.FromContext[adapter.OutboundManager](w.ctx); outboundManager != nil {
if blockedBy := awgDetourChainReachesWireGuard(outboundManager, w.detour, make(map[string]bool)); blockedBy != "" {
w.awgChainBlocked = true
w.logger.Error("amneziawg endpoint will not start: its detour chain reaches wireguard-based endpoint ", strconv.Quote(blockedBy), " — amneziawg over wireguard is not supported. Use a non-wireguard detour (e.g. vless).")
return nil
}
}
}
if w.awgChainBlocked {
return nil
}
// lx:end awg
switch stage {
case adapter.StartStateStart:
return w.endpoint.Start(false)
@@ -200,69 +165,6 @@ func (w *Endpoint) Start(stage adapter.StartStage) error {
return nil
}
// lx:begin awg
// awgDetourChainReachesWireGuard walks the transitive detour chain starting at
// tag and returns the tag of the first WireGuard-based outbound it reaches
// (type "wireguard", covering plain WireGuard and AmneziaWG), or "" if none. It
// follows each outbound's detour dependency; it deliberately does NOT expand
// selector/urltest groups, whose chosen member is only known at runtime — that
// case is handled lazily by the DetourDialer guard. visited guards against cyclic
// detour configs. All outbounds are registered before any Start, so every tag in
// the chain is resolvable here even though some may not have started yet.
func awgDetourChainReachesWireGuard(outboundManager adapter.OutboundManager, tag string, visited map[string]bool) string {
if tag == "" || visited[tag] {
return ""
}
visited[tag] = true
outbound, loaded := outboundManager.Outbound(tag)
if !loaded {
return ""
}
if outbound.Type() == C.TypeWireGuard {
return tag
}
if _, isGroup := outbound.(adapter.OutboundGroup); isGroup {
// Runtime-resolved target — leave it to the lazy DetourDialer guard.
return ""
}
for _, dependency := range outbound.Dependencies() {
if blockedBy := awgDetourChainReachesWireGuard(outboundManager, dependency, visited); blockedBy != "" {
return blockedBy
}
}
return ""
}
// IsAmneziaWG reports whether this endpoint runs AmneziaWG. Implements
// adapter.AmneziaWGSuspendable.
func (w *Endpoint) IsAmneziaWG() bool {
return w.awgActive
}
// SuspendAmneziaWG brings the device down and marks the endpoint not-ready, so a
// junk handshake is never sent and every dial fails with "WireGuard is not ready
// yet". Called by the selector guard when a group this endpoint detours through
// switches to a WireGuard member (AmneziaWG over WireGuard hangs the kernel on
// Android). Idempotent. Implements adapter.AmneziaWGSuspendable.
func (w *Endpoint) SuspendAmneziaWG() {
// Take resumeMu so this is ordered against resumeOnDial/SuspendIfIdle: without
// it, a dial that already passed resumeOnDial's idleAsleep checks could wake
// the endpoint back up right after we clear the flag, defeating the guard.
w.resumeMu.Lock()
defer w.resumeMu.Unlock()
if w.started.CompareAndSwap(true, false) {
w.logger.Error("amneziawg endpoint suspended: a selector in its detour chain switched to a wireguard-based member — amneziawg over wireguard is not supported")
}
// Clear any idle-suspend state so resumeOnDial does not resurrect a
// guard-suspended endpoint: if it was idle-asleep first, idleAsleep would still
// be true and the next dial would wake it (SPEC 022 #2). With idleAsleep=false
// resumeOnDial's fast path returns started (now false) and the endpoint stays down.
w.idleAsleep.Store(false)
w.endpoint.Suspend()
}
// lx:end awg
// lx:begin idle-suspend
// stampActivity records the current time as the last dial through this endpoint.
@@ -286,8 +188,8 @@ func (w *Endpoint) IdleSince() time.Duration {
// holder — when it is unreachable from the active routing tree AND has been idle
// past the threshold. Silent on every non-transition (edge-triggered logging).
//
// It never touches a guard-suspended endpoint: that one already has
// started==false but idleAsleep==false, and the `!started` guard below short-
// It never touches a deliberately-stopped endpoint: that one already has
// started==false but idleAsleep==false, and the `!started` check below short-
// circuits before the CAS. resumeMu mutually excludes this against resumeOnDial.
func (w *Endpoint) SuspendIfIdle(reachable bool, threshold time.Duration) {
w.resumeMu.Lock()
@@ -296,7 +198,7 @@ func (w *Endpoint) SuspendIfIdle(reachable bool, threshold time.Duration) {
return
}
if !w.started.Load() {
// Already down some other way (guard-suspend, awg-chain-blocked, closed).
// Already down some other way (deliberately stopped, closed).
return
}
if w.idleAsleep.CompareAndSwap(false, true) {
@@ -313,7 +215,7 @@ func (w *Endpoint) SuspendIfIdle(reachable bool, threshold time.Duration) {
// session); that cost is on the first packet, as for any cold WG dial.
//
// Returns true if the endpoint is dialable (awake), false if it must stay down
// (guard-suspend / chain-blocked — not an idle-suspend, so we do not resurrect it).
// (deliberately stopped / closed — not an idle-suspend, so we do not resurrect it).
func (w *Endpoint) resumeOnDial() bool {
w.stampActivity()
if !w.idleAsleep.Load() {
+15 -15
View File
@@ -103,21 +103,21 @@ func TestSuspendIfIdle_idempotentCAS(t *testing.T) {
}
}
// TestSuspendIfIdle_guardSuspendedNotTouched is the §8 invariant verified live on
// an AWG-over-WG endpoint (wg-3 in the prod run): a guard-suspended endpoint has
// started=false WITHOUT idleAsleep. The idle tick must early-return on !started and
// NOT flip idleAsleep — otherwise a later resumeOnDial would idle-wake it and
// re-trigger the AWG-over-WG kernel hang the guard exists to prevent.
func TestSuspendIfIdle_guardSuspendedNotTouched(t *testing.T) {
// TestSuspendIfIdle_stoppedNotTouched is the §8 invariant: a deliberately-stopped
// endpoint (Close, or a start that never completed) has started=false WITHOUT
// idleAsleep. The idle tick must early-return on !started and NOT flip idleAsleep
// — otherwise a later resumeOnDial would idle-wake a device that was
// intentionally down.
func TestSuspendIfIdle_stoppedNotTouched(t *testing.T) {
w := newIdleTestEndpoint()
w.started.Store(false) // guard-suspend (device.Down at Start), idleAsleep stays false
w.started.Store(false) // stopped, idleAsleep stays false
w.lastActivity.Store(time.Now().Add(-time.Hour).UnixNano())
w.SuspendIfIdle(false, 30*time.Second)
if w.idleAsleep.Load() {
t.Fatal("a guard-suspended endpoint must NOT be flagged idleAsleep by the tick")
t.Fatal("a stopped endpoint must NOT be flagged idleAsleep by the tick")
}
if w.started.Load() {
t.Fatal("the tick must not change started for a guard-suspended endpoint")
t.Fatal("the tick must not change started for a stopped endpoint")
}
}
@@ -152,17 +152,17 @@ func TestResumeOnDial_dialBeforeTickRace(t *testing.T) {
}
}
func TestResumeOnDial_guardSuspendedNotWoken(t *testing.T) {
// A guard-suspended endpoint has started=false but idleAsleep=false.
// resumeOnDial must NOT wake it (returns started, i.e. false).
func TestResumeOnDial_stoppedNotWoken(t *testing.T) {
// A deliberately-stopped endpoint (Close / failed start) has started=false but
// idleAsleep=false. resumeOnDial must NOT wake it (returns started, i.e. false).
w := newIdleTestEndpoint()
w.started.Store(false) // simulate guard/awg-chain suspend (not idle)
w.started.Store(false) // stopped, not idle-suspended
ok := w.resumeOnDial()
if ok {
t.Fatal("resumeOnDial must not resurrect a guard-suspended (non-idle) endpoint")
t.Fatal("resumeOnDial must not resurrect a stopped (non-idle) endpoint")
}
if w.idleAsleep.Load() {
t.Fatal("guard-suspended endpoint must not be flagged idleAsleep")
t.Fatal("stopped endpoint must not be flagged idleAsleep")
}
}
+268
View File
@@ -0,0 +1,268 @@
// lx:begin l3-honest-drop
package route
import (
"context"
"net/netip"
"testing"
"github.com/sagernet/sing-box/adapter"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
R "github.com/sagernet/sing-box/route/rule"
"github.com/sagernet/sing/common/json/badoption"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/stretchr/testify/require"
)
// The contract under test: PreMatch never answers "continue" (nor "bypass") for
// an ICMP flow. adapter.JudgeFlow maps both to tun.ActionAccept, and the TUN
// stack answers Accept by FORGING the echo reply itself
// (sing-tun stack_gvisor_icmp.go ICMPForwarder.HandlePacket, the fallthrough
// under the Flow/Reject/Drop switch). A verdict of "continue" therefore reads to
// the operator as a working ping off a tunnel that never carried the packet.
//
// Every test below has a TCP/UDP twin: the honest drop must not leak into the
// protocols where "continue" really does mean "take the ordinary connection
// route".
// icmpL4Outbound is a minimal L4-only outbound (the vless/vmess/... shape): it
// does NOT implement adapter.FlowOutbound, and Network() lists only TCP/UDP.
// Unused Outbound methods come from the embedded nil interface and are never
// called on the pre-match paths under test.
type icmpL4Outbound struct {
adapter.Outbound
tag string
}
func (o *icmpL4Outbound) Tag() string { return o.tag }
func (o *icmpL4Outbound) Type() string { return "vless" }
func (o *icmpL4Outbound) Network() []string { return []string{N.NetworkTCP, N.NetworkUDP} }
// icmpOutboundManager resolves tags from a fixed map and hands the same L4-only
// outbound out as the default; the rest of the OutboundManager surface is never
// touched by the pre-match walk.
type icmpOutboundManager struct {
adapter.OutboundManager
defaultOutbound adapter.Outbound
outbounds map[string]adapter.Outbound
}
func (m *icmpOutboundManager) Default() adapter.Outbound { return m.defaultOutbound }
func (m *icmpOutboundManager) Outbound(tag string) (adapter.Outbound, bool) {
outbound, loaded := m.outbounds[tag]
return outbound, loaded
}
// icmpDNSRouter / icmpDNSTransportManager implement only what
// prepareMatchMetadata reaches. FakeIP returns nil unless a transport is
// installed, which is how the "fakeip lookup failed" exit is driven below.
type icmpDNSRouter struct {
adapter.DNSRouter
}
func (s *icmpDNSRouter) LookupReverseMapping(netip.Addr) (string, bool) { return "", false }
type icmpDNSTransportManager struct {
adapter.DNSTransportManager
fakeIP adapter.FakeIPTransport
}
func (s *icmpDNSTransportManager) FakeIP() adapter.FakeIPTransport {
if s.fakeIP == nil {
return nil
}
return s.fakeIP
}
// icmpMissingFakeIPTransport claims every address and then fails to look any of
// them up — exactly the "missing fakeip record, try enable
// `experimental.cache_file`" error prepareMatchMetadata returns.
type icmpMissingFakeIPTransport struct {
adapter.FakeIPTransport
}
func (t *icmpMissingFakeIPTransport) Store() adapter.FakeIPStore {
return &icmpMissingFakeIPStore{}
}
type icmpMissingFakeIPStore struct {
adapter.FakeIPStore
}
func (s *icmpMissingFakeIPStore) Contains(netip.Addr) bool { return true }
func (s *icmpMissingFakeIPStore) Lookup(netip.Addr) (string, bool) { return "", false }
type icmpRouterOptions struct {
fakeIP adapter.FakeIPTransport
rules []option.Rule
}
func icmpTestRouter(t *testing.T, options icmpRouterOptions) *Router {
t.Helper()
logger := log.NewNOPFactory().NewLogger("test")
defaultOutbound := &icmpL4Outbound{tag: "proxy-out"}
router := &Router{
ctx: context.Background(),
logger: logger,
dns: &icmpDNSRouter{},
dnsTransport: &icmpDNSTransportManager{fakeIP: options.fakeIP},
outbound: &icmpOutboundManager{
defaultOutbound: defaultOutbound,
outbounds: map[string]adapter.Outbound{defaultOutbound.Tag(): defaultOutbound},
},
}
for i, ruleOptions := range options.rules {
rule, err := R.NewRule(router.ctx, logger, ruleOptions, false)
require.NoError(t, err, "build rule[%d]", i)
router.rules = append(router.rules, rule)
}
return router
}
func icmpTestMetadata(network string) adapter.InboundContext {
return adapter.InboundContext{
Inbound: "l3-in",
InboundType: C.TypeTun,
Network: network,
Source: M.SocksaddrFrom(netip.MustParseAddr("192.168.1.2"), 0),
Destination: M.SocksaddrFrom(netip.MustParseAddr("1.1.1.1"), 0),
}
}
// lanRuleWithAction matches every packet from the test source, so the action is
// what the test is actually about.
func lanRuleWithAction(action option.RuleAction) option.Rule {
return option.Rule{
Type: C.RuleTypeDefault,
DefaultOptions: option.DefaultRule{
RawDefaultRule: option.RawDefaultRule{
SourceIPCIDR: badoption.Listable[string]{"192.168.1.0/24"},
},
RuleAction: action,
},
}
}
// --- exit 1: an outbound that cannot carry layer 3 --------------------------
func TestPreMatchICMPToL4OutboundDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"ICMP to an L4-only outbound fell through to the ordinary pre-match path: the TUN stack will forge the echo reply and ping will lie about a tunnel that never saw the packet")
}
func TestPreMatchTCPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"TCP to an L4-only outbound must keep taking the ordinary connection route; the ICMP honest-drop must not leak into TCP/UDP pre-match")
}
func TestPreMatchUDPToL4OutboundContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{})
result := router.PreMatch(icmpTestMetadata(N.NetworkUDP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"UDP to an L4-only outbound must keep taking the ordinary connection route")
}
// --- exit 2: prepareMatchMetadata failed before any rule was walked ---------
// This exit arrived with the shared prepareMatchMetadata refactor (upstream
// b911fb078): it returns before the rule walk, so it never reaches preMatchFlow
// where the ICMP override used to live.
func TestPreMatchICMPMetadataErrorDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"a fakeip record that cannot be resolved must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPMetadataErrorContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{fakeIP: &icmpMissingFakeIPTransport{}})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"for TCP the metadata-error exit must keep meaning `take the ordinary connection route`")
}
// --- exit 3: a rule action the pre-match walk does not handle ---------------
// hijack-dns is one of the actions PreMatch's switch has no arm for, so it lands
// in the default arm. Any future unhandled action lands there too — that is why
// the guard is a funnel on the return value and not a per-arm override.
func TestPreMatchICMPUnhandledRuleActionDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"an unhandled rule action must not degrade ICMP to continue: continue is tun.ActionAccept, and Accept is a forged echo reply")
}
func TestPreMatchTCPUnhandledRuleActionContinues(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeHijackDNS})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchContinue, result.Action,
"the unhandled-action exit must stay a continue for TCP")
}
// --- exit 4: an explicit bypass ---------------------------------------------
// sing-tun implements ActionBypass on the nfqueue plane only; on the TUN path it
// falls into the same default arm as Accept (flow_dispatch.go judgeAndInstall,
// and the ICMP forwarder's switch has no Bypass case either), i.e. into the same
// forgery. There is no honest bypass for a packet already inside the engine's
// TUN.
func TestPreMatchICMPBypassDrops(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchDrop, result.Action,
"bypass degrades to tun.ActionAccept on the TUN path, which is the forged echo reply again")
}
func TestPreMatchTCPBypassIsStillBypass(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{Action: C.RuleActionTypeBypass})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkTCP), nil)
require.Equal(t, adapter.PreMatchBypass, result.Action,
"the ICMP honest-drop must not turn a TCP bypass rule into a drop")
}
// --- the verdicts that must pass through untouched ---------------------------
// A reject rule already carries its own honest verdict; the funnel must not
// rewrite it (a Reject sends an ICMP unreachable, which is information, not a
// forged liveness signal).
func TestPreMatchICMPRejectIsNotRewritten(t *testing.T) {
t.Parallel()
router := icmpTestRouter(t, icmpRouterOptions{
rules: []option.Rule{lanRuleWithAction(option.RuleAction{
Action: C.RuleActionTypeReject,
RejectOptions: option.RejectActionOptions{Method: C.RuleActionRejectMethodDefault},
})},
})
result := router.PreMatch(icmpTestMetadata(N.NetworkICMP), nil)
require.Equal(t, adapter.PreMatchReject, result.Action,
"the ICMP funnel must only rewrite continue/bypass, never an explicit reject")
}
// lx:end l3-honest-drop

Some files were not shown because too many files have changed in this diff Show More