Compare commits

..
15 Commits
Author SHA1 Message Date
omarandClaude Opus 5 3c7536dba0 fix(panel): show the chain hop by hop, and stop reading unused as broken
release / apk aarch64_cortex-a53 (push) Successful in 3m5s
release / apk x86_64 (push) Successful in 3m1s
release / release apk (push) Successful in 8s
The chain card gave a single verdict, so a dead hop was invisible: the operator
saw "the chain is unhealthy" and had to guess which of four hops to look at.
Meanwhile a group used only inside a chain showed "unused" next to a live
alive/dead count, which reads as a diagnosis when it only means nothing measures
it on that path.

Render the hops as a rail that severs below the first dead one, so which hop is
answered before a word is read, and split the two "not routed" messages into the
routing fact and the explicit non-fact. The group one names the case directly: a
group used only as a hop inside a chain reads unused here on purpose, and its
real health is on that chain's card.

Also fixes a bug this would otherwise have shipped: the readout painted every
ok:false in the critical colour, so "not routed" would have rendered as a fault
— the exact lie being removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 01:09:42 +03:00
omarandClaude Opus 5 3c92e1cbfd fix(health): one prober, on the path the rules actually use
A node reached only as a chain hop was being measured twice, and the reading
the panel showed was the wrong one. On a router in Russia that is not a cosmetic
difference: a node the chain carries fine behind a WireGuard hop is dead when
dialled straight out of the WAN, so the group card read "0 of 2 alive" while
that very group was carrying every packet.

Two dial paths existed outside the observatory plan. URLTestGroup.PostStart
warmed up every urltest group at box start whether or not any rule reached it,
and the panel's Test button reached URLTest.DialContext, whose first act is
Touch() — arming a ticker that re-swept those groups directly every probe
interval for the next thirty minutes. Both wrote under the BASE node tag, and
both dialled the base outbound, which carries no chain detour at all.

The observatory was never the liar: its plan roots come from the rules, and a
chain hop copy is stored only under its own tag, so no plan job could ever
write under a base tag. The fix is therefore to remove the other two paths, not
to touch the plan.

TestGroups now asks the observatory for an out-of-turn pass and reports what it
measured; a target no enabled rule routes to is not dialled at all and says so.
Unused urltest groups stand down their own self-check via a new SelfCheck option
(nil keeps today's behaviour, so every existing config is unchanged). The one
direct dial left is the exit-address lookup, which has no other possible source
— it now runs only for a target that is both routed and already read alive, so
it travels the routed path and never touches an unused group.

Chain hop wrappers are probed as measurements of their own and surfaced as
chains[].hops[], because "which hop is dead" is the question an operator has and
the chain-level verdict cannot answer it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 01:09:42 +03:00
omarandClaude Opus 5 bcc9df9282 test(wireguard): drive a real AmneziaWG tunnel through ClientBind
release / apk aarch64_cortex-a53 (push) Successful in 3m25s
release / apk x86_64 (push) Successful in 3m14s
release / release apk (push) Successful in 8s
The unit tests pin the reserved-byte gate on each side in isolation, which
would still pass if the two halves disagreed about when to apply it. This wires
two real wireguard-go devices together over loopback UDP through ClientBind on
both ends — the bind the detour path actually uses — configures ranged h1-h4
plus s4 and junk, and asserts an inner IP packet reaches the peer's TUN.

It is red against the unconditional clear and green with the gate, so it covers
the failure the field hit rather than the code we happened to write. Tagged
with_awg, so it runs under the shipped router tag set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:25:05 +03:00
omarandClaude Opus 5 ee3641fe45 fix(logsink): collapse interleaved floods, not just consecutive lines
The previous suppression compared each line with the one before it, which the
field never obliges. A dead chain makes the engine cycle the same message
across three outbound tags, so no two identical lines are adjacent: on the
router it produced 854 daemon lines in a ~760-line syslog ring and exactly one
summary, all while claiming "repeated 1 time". The rest of the system's log —
netifd, dnsmasq, the kernel — was evicted anyway.

Track a bounded table of open series keyed by the existing repeat key instead.
The first copy of a key prints; further copies inside its window are counted
whatever arrives in between; the window end emits one summary per key. The
summary now names its message, because several can close at once and "last
message" would simply be false under interleaving.

The table holds 256 keys and evicts the least recently seen, never silently: an
evicted series with a pending count prints its summary on the way out, marked
so the truncation is visible. Close, Reconfigure and any fatal flush every open
series first — a dying daemon may never reach Close.

TestRepeatAlternatingNotSuppressed asserted that A B A B must never be
collapsed. That assertion was the bug. It is replaced by a stronger one: the
messages get separate series, separate summaries and separate counts, so
distinct events still never fold into a single number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 439f62238f fix(engine): retire the superseded instance instead of leaving it running
Every config apply built a new box and left the old one alive. The engine's own
log gives it away: inside a single shaterd process, lines carried uptime
counters half an hour apart in the same second, and a live router was found
running four generations at once. A process restart cleared it, so the leak
accrued purely on re-apply.

That is not just wasted memory on a 512 MB box. Each surviving generation keeps
its WireGuard devices up, and two devices sharing one private key evict each
other at the peer — so the leak reproduced the duplicate-device defect between
generations, underneath the deduplication that only reasons about one config.

Retirement now has a hard budget: 5s, which is exactly sing-box's own
C.StopTimeout (past which upstream already calls a stop excessive) and stays
under C.FatalStopTimeout. It is paid after the replacement is serving and only
on an apply that changed something, so a no-op reconcile stays free.

A close that blows the budget is ABANDONED, not waited on, and the apply is
still reported as the success it is — the new box is built, started and
carrying traffic, and failing there would abort the netplane stage and leave a
stale ruleset over a healthy engine. The stuck instance is surfaced through
PendingCloses() into `shaterd status` and the panel, and clears itself if the
shutdown ever completes. Repeated applies over a stuck close no longer stack:
the abandoned generation is remembered, not re-created.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 d971eb85ee fix(wireguard): stop ClientBind from shredding the AmneziaWG magic header
An AmneziaWG node worked standalone and died the moment it was placed behind
an egress or a chain hop: the handshake completed, the peer answered, and then
not one byte of data ever arrived. The peer never confirmed the session, so it
re-handshook every 15 seconds, forever.

ClientBind cleared bytes 1-3 of every datagram on receive and stamped them on
send, unconditionally. Those bytes are Cloudflare's "reserved" field. They are
also where AmneziaWG puts the upper three bytes of its little-endian uint32
magic header, so zeroing them collapses the value to its low byte, which falls
outside every h1-h4 range and makes the peer classify the packet as an unknown
type and drop it silently.

Handshakes survived because s1/s2 padding pushes their magic past byte 3 — the
clear only scribbled on the random junk prefix. Transport packets have s4 = 0,
so their magic starts at byte 0 and took the hit. That asymmetry is the whole
signature: session up locally, zero data through.

Only the detour path was affected, because Endpoint.Start picks StdNetBind when
the dialer exposes WireGuardControl (no detour) and ClientBind otherwise. The
gate had already landed in StdNetBind; ClientBind was its untouched twin. The
two implement one contract and are now commented as the pair they are, so the
next fix cannot again land on one side only.

Measured on the box: h4 spans 0x60728123-0x60728155, so zeroing bytes 1-3
leaves 35..85 — the captured transport packet began with 56, while a node
without a detour carried a correct 0x6b039798 at the same moment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:23:24 +03:00
omarandClaude Opus 5 a0f6083e28 fix(panel): show whether a rule is in force, not just what was saved
release / apk aarch64_cortex-a53 (push) Successful in 3m7s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 8s
With two catch-all rules both enabled in UCI and a WAN profile enabling one
and disabling the other, the panel drew BOTH switches on while the engine
ran only one chain. GET /api/config is right to return the raw model — that
is the desired state the panel PUTs back — but Routing.tsx read the row
state and the active count from it too, so the interface claimed a setting
was in force when it was not. Same defect class as the Protected badge.

/api/rules/reachability now carries the effective flag and, where the active
profile changed the outcome, its name and direction. The annotation is a
DIFF of ApplyProfileRuleOverrides output against desired state rather than a
second reading of the profiles name lists, so profile logic is not
duplicated and cannot drift — an unmigrated rule the profile is forbidden to
enable produces no diff and gets no badge, with nothing here needing to know
about LegacyDst.

In the UI the two states stay separate: the switch remains the only carrier
of desired state and still writes UCI, while the effective state drives the
dimmed row, the badge, the banner and the header count. Mirroring the
effective state into the switch would be worse than the original bug — the
operator would be toggling someone elses control, and the profiles decision
would be written back as their own choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 77369aedfe fix(logsink): collapse repeated lines instead of erasing the routers syslog
A broken outbound makes the engine repeat one line about once a second —
370 copies in six minutes. The routers syslog ring holds ~760 lines, so
within minutes it evicts the history of every other subsystem and our own
startup lines with it. Diagnosing the WireGuard duplication above required
restarting the service purely to catch the first seconds of a boot.

Collapse runs into "last message repeated N times". The comparison key is
level + text with the uptime field dropped: comparing whole lines would
suppress only same-second bursts, because that counter ticks. The per
connection "[id duration]" group is deliberately KEPT in the key — those ids
are distinct connections, and folding "50 connections failed" into one count
would be a worse lie than the flood. Window 5s, so a standing fault keeps
being reported instead of looking like a frozen log. fatal/panic are never
suppressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 515ae6d1b7 test(generate): make the remote-blocklist test exercise the remote path
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two
releases, which made the whole package exit non-zero no matter what the code
did — a real regression would have drowned in the familiar red.

The cause is not the packages no-network fetcher stub, as it first appears.
ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION
alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so
the fixture fell into the TEXT-list path: downloaded by generates own
fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which
every assertion below then contradicted. No stub content could fix that; the
stub decides the lists contents, not the rule-sets type.

Give the URL the .srs suffix the test always meant it to have, so the engine
fetches the compiled set itself through the direct outbound. No assertion is
weakened and the no-network stub stays in place.

Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags):
338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 a2ffbb1292 fix(generate): one WireGuard device per private key
A node may be copied freely by this package: a per-chain hop copy and a
per-group egress copy are rebuilt from the share-link so each can carry its
own Detour. For vless that is right — a copy is another TCP client. For
WireGuard it is not: each emitted endpoint is a real device holding the
nodes private key, and a peer keeps exactly ONE session per public key.
Two devices from one key evict each other continuously, and with keepalive
on both the loop never settles: NEITHER passes traffic.

buildOutboundsAndEndpoints emits the base endpoint for every enabled node
whether or not anything references it, so a WG node used only as a chain hop
always produced two devices. That is what any chain containing a WG node
looks like — every such chain was permanently dead.

Observed on the box: two UDP sockets from shaterd to the same peer port, the
servers peer endpoint flapping between them, +32 bytes/min through the
tunnel and every hop failing with "context deadline exceeded".

Deduplicate once on the assembled options, which catches all three producer
paths by construction. Duplicates are DELETED, not merely unreferenced:
box.New starts every endpoint regardless of reachability, so a leftover
would still bring its device up and still fight for the session. Dangling
references go to block, never to direct — a consumer whose tunnel just
disappeared must stop, not fall out onto the plain WAN.

Subscription fetch detours seed the reachability walk (they are direct
references like any rule), mirroring engine.ViaToTag exactly, with a
tripwire test against drift. A config that genuinely needs two devices for
one key keeps one and fail-closes the rest with a critical warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 1746d4d0ef fix: stop the panel and the shipped binary from lying about what works
release / apk aarch64_cortex-a53 (push) Successful in 9m13s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 7s
Four defects, all found by the owner on the live router, all of the same
family: something declared itself working while it was not.

WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave
"create WireGuard device: gVisor is not included in this build". The router
tag set carried with_wireguard and with_awg but not with_gvisor, so
sing-tun compiled its stub instead of the netstack every WireGuard device
needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving
requirement", so this was a broken promise, not a trim.

The tag itself was the small half. The tag set was the ONE build
configuration nothing in the repo tested: TestAmneziaWGEndpoint passes
because tests build with the full upstream tags. So the set now lives in
one file (scripts/router-tags.sh) and two guards hold it to the feature
list -- a static check that needs no tags, no Linux and no network (so the
next such gap fails on the developer's machine), and a behavioural one that
constructs every declared protocol through box.New UNDER THE SHIPPED TAGS,
where skipping is forbidden. Removing the tag now fails with the feature
name, the missing tag, and why: "Either add the tag back, or stop declaring
the feature -- those are the only two honest options." Cost: +2.8 MB raw,
+0.6-0.7 MB packed per arch. D23; D9 corrected.

THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came
from plane === 'full', which reports whether the data plane is installed --
nft table, policy routing, live engine -- and says nothing about where the
traffic goes. On a config with one `default -> direct` rule and no groups
the plane is fully installed and every packet leaves in the clear, so the
worst possible state rendered as the reassuring one.

The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the
moment they reach the engine, not from the model: buildRoute changes the
answer (a scheduled rule outside its window is never emitted, only the last
condition-less rule reaches Final, an unresolved target is rewritten by
ruleKillFallback), and re-deriving it anywhere else is a second
implementation that will drift -- model/reachability.go exists because two
already did. Four verdicts, not three: `blocked` is separate because under
a closed kill-switch with no catch-all nothing leaks, and calling that
"going out directly" is a lie in the alarm direction. Rider: Overview's
defaultTarget printed the highest-Order enabled rule as the default; a rule
becomes Final by having no conditions, whatever its Order.

"PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE
(B15). Once the browser suppresses dialogs, window.confirm returns false
immediately, so all 15 confirmations across 7 pages read as "cancelled" and
silently did nothing, with no way to recover from inside the panel. Replaced
with an in-app dialog the browser cannot mute: focus trapped and parked on
Cancel, Esc and veil cancel, focus returned to the opener, crit styling for
destructive commits. useConfirm() throws if the provider is missing rather
than falling back to a quiet false -- the failure mode being fixed.

HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed,
so a feed's nodes of those types vanished. The real landmine was one layer
up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so
without schemePrefixes the links were gone before any parser ran and the
fix would have looked complete. Undeliverable parameters are refused when
the node cannot work or would be less secure than the link asked (obfs,
pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it
survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot
produce a QUIC TLS config, and that fails at dial time, not at box.New.

Also: nodes added by hand can be named and renamed. The name is the
outbound tag, so a rename rewrites every reference in one PUT -- rule
targets, group members, chain hops, detours -- in the spelling each already
uses, and is refused outright when a group answers to the same bare name.
Subscription nodes state why they cannot be renamed instead of hiding the
control.

go build, go vet, go test ./shater/... (13 packages), panel npm run build
and npm test (13/13) all green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 20:08:58 +03:00
omarandClaude Opus 5 f86501bf77 ci!: drop the opkg lane — apk only, and fix the stale rolling release
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and
main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has
no `opkg` binary at all. The 24.10 lane was building and signing a feed no
device could consume.

Removed jobs `build` and `release` with the scripts only they called
(ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh)
and the usign trust anchor dist/shater-feed.pub. A committed public key is
an instruction: it invites the old install path for a feed that is no longer
produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD
still holds the secret half, and a usign secret contains its own public half,
so the identity is reconstructible if a 24.10 device ever needs serving.
D7 is marked SUPERSEDED by the new D22 rather than deleted.

Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from
2026-07-24 while every tag run published its versioned release correctly.
The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest
(workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run
never touched the rolling pointer. Asset replacement was never the problem;
ci/gitea-release.sh already deletes before recreating. A router pinned to
the rolling URL sat on 0.2.0 while `apk update` reported success: silent
staleness, the failure mode this repo keeps having to close.

The rolling pointer is now published on EVERY run, tag runs included, and a
new assert reads the release back over the API afterwards: our three
tag-versioned packages at the built version plus the index and the key must
be present (exit 13), and no package asset at any other version may survive
(exit 14). Same class of check as sdk-build-apk.sh's package-version assert,
added for the same reason -- the previous failure mode was silent.

KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references
it. Docs state plainly that mini_router is deliberately pinned to a
versioned URL and that the hand-edit per release is the price of pinning.

Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can
no longer install our packages. Its 25.12 rebuild is in flight separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:46:21 +03:00
omarandClaude Opus 5 eccfc6136c fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review of a8ef887c5 + 244b7c419 ("a rule's destination is a rule-set,
and nothing else") found that the change rested on a comment that was not
true. ParseUCIExport dropped dst_domain/dst_ip on the strength of "the
migration is re-run on every load"; model.Migrate() actually runs only from
`shaterd migrate`, i.e. the service init and uci-defaults. The daemon's run
path, the SIGHUP reconcile and the panel's config write never migrate.

So an uncommitted migration (a full /overlay is the documented way that
happens) turned `list dst_domain 'bank.ru'` + `target direct` into a rule
with NO matchers, which IS the spelling of a catch-all: generate points
route.Final at it and the LAST such rule wins. One failed `uci commit` sent
every packet on the router out the plain WAN, silently.

Rule.LegacyDst is the tripwire. It is non-empty exactly when the config
still carries the removed options, and three locks hang off it:
  - ParseUCIExport holds such a rule DISABLED. Chosen over "make IsCatchAll
    false" alone, which only covers matcher-less rules: `dst_domain` plus a
    `src` was never a catch-all, and routing it without its destination
    would still have sent a whole subnet direct.
  - IsCatchAll returns false for it, so it can never own route.Final even
    if something hands its Enabled bit back.
  - ApplyProfileRuleOverrides refuses to enable it (a profile with
    `list enable_rule` would otherwise have defeated the parser).
ValidateRules reports it through the existing warning channel, before the
Enabled gate, so the one message explaining the outage is not suppressed by
the fact that caused it. The init script logs a failed migration to syslog
instead of discarding its exit code and stderr.

The write path had none of this. PUT /api/config decodes a Model straight
from the request body and render.go wrote `enabled` from it, so a panel
save erased the operator's lists (as did the subscription cron, which
re-renders the whole package), and a crafted body with Enabled:true and no
LegacyDst put a live matcher-less rule on disk -- the same whole-router
leak, re-entered from the other side. WriteUCI now reads DISK state and
refuses a rule-changing write over an unmigrated config (409, not 500);
non-rule writers pass and legacyDstOpts carries the options across so cron
preserves them; withDiskLegacyDst takes the field from disk so a fabricated
one can never reach the renderer.

Migration hardening: an entry list that migrates to nothing no longer has
its legacy option deleted (that made "matches nothing" silently become
"matches everything"); a hand-written rule-set whose name collides is no
longer allowed to swallow the entries; delete failures propagate instead of
bumping schema_version past them forever; every error path reverts the
staged uci delta so another process's commit cannot flush a half-migration.

untunnelable.go follows the destination out of the rule: a rule whose
rule-sets are known to match by name is still skipped by the ping/IPTV/VPN
plan, as its v1 form was. D21 documents the AND->OR widening for the
engine's TCP/UDP path; it does not follow that a leak-guard should widen
itself during an upgrade, and with target=direct that meant previously
tunnelled ICMP leaving with the client's real address. Inline rule-sets are
now read from the options, so an engine that has not started yet no longer
costs the operator their ping.

Rule-set vocabulary: `full:`/`suffix:`/`keyword:`/`regexp:` in a text list
fetched by URL were dropped with no diagnostic at all (normaliseListDomain
rejects any token with a colon) -- not "reported as an unknown prefix".
Unifying was rejected: published filter lists are full of colon-bearing
syntax, and a third-party `regexp:` is compiled into the router's matcher
and run per query. The difference stands and is paid for in diagnostics,
per list, on every generate. D21 gains the source/vocabulary table.

Panel: the add form warns about a matcher-less rule exactly as the edit
form does, from one shared predicate; its isCatchAll matches the daemon's
new one; an unmigrated rule reads as held-off rather than merely switched
off. The comment promising a "New list" button that D21 rejected is gone.

go build ./..., go vet ./shater/..., go test ./shater/... (13 packages) and
panel `npm run build` are green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:04:21 +03:00
omarandClaude Opus 5 244b7c4199 feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.

The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.

The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.

isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.

Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
132 changed files with 14513 additions and 2020 deletions
+162 -315
View File
@@ -1,36 +1,45 @@
# Shater v0.2 — build the 4-package signed opkg feed and publish it as a rolling
# Gitea release consumable as an `src/gz` feed.
# Shater v0.2 — build the 4-package signed **apk** feed and publish it as
# per-arch Gitea releases consumable as an apk repository.
#
# WHAT CHANGED FROM v0.1
# v0.1 shipped 3 packages: xrayctl (SDK-compiled Go) + shater-core +
# luci-app-shater (hand-packed data .ipk). v0.2 collapses the runtime into ONE
# forked binary and ships 4 packages, all built the canonical SDK way:
# WHAT WE SHIP
# ONE forked binary plus its OpenWrt glue, 4 packages, all built the canonical
# SDK way:
# - shaterd PREBUILT static-musl + SPA-embedded + UPX binary. Built
# OUT OF TREE by scripts/build-shaterd.sh (Go + Node + UPX)
# and staged into openwrt/shaterd/files/ BEFORE the SDK
# build; the openwrt/shaterd package just $(INSTALL_BIN)s
# the arch-matched artifact. (arch-specific .ipk)
# the arch-matched artifact. (arch-specific .apk)
# - shater-core data glue, PKGARCH=all
# - luci-app-shater LuCI thin launcher, PKGARCH=all (uses feeds/luci/luci.mk)
# - byedpi ciadpi, C cross-compiled from source by the SDK (arch-specific)
#
# TARGET HARDWARE / ARCH MATRIX
# x86_64 -> the QEMU testbed VM (generic x86-64).
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 + BPI-R4, mediatek/filogic).
# aarch64_cortex-a53 -> BOTH production routers (BPI-R3 mini + BPI-R4,
# mediatek/filogic), both on 25.12 with apk-tools 3.
# Only shaterd + byedpi are arch-specific; shater-core + luci-app-shater are
# PKGARCH=all, so one build of each covers every device. opkg filters by
# Architecture at install time, so a single combined feed URL serves all.
# PKGARCH=all, so one build of each covers every device — but the RELEASES
# are still per-arch (see the release-apk job for why).
#
# FEED SIGNING (opkg / usign — OpenWrt 24.10 is opkg, not apk; apk lands at 25.12)
# The feed index (Packages) is usign-signed with the SECRET key in the Gitea
# repo secret KEY_BUILD; routers verify it with the committed public key
# dist/shater-feed.pub (fingerprint 5ac4b177689cb8e0). Do NOT regenerate the
# key — that invalidates every deployed router's trust.
# FORMAT: apk ONLY (25.12+)
# The fleet runs OpenWrt/ImmortalWrt 25.12, where opkg is replaced by Alpine
# apk (.apk files, binary packages.adb index, EC keys in /etc/apk/keys/). The
# old .ipk lane was removed in 2026-07 (docs-shater/DECISIONS.md D22): no
# device we serve has an opkg binary at all, so building and signing a second
# feed served nobody.
#
# FEED SIGNING (EC / apk)
# packages.adb is signed with the EC (prime256v1) SECRET key in the Gitea repo
# secret KEY_APK; routers verify it with the committed public key
# dist/shater-apk.pem (ci/gen-apk-key.sh). Do NOT regenerate the key — that
# invalidates every deployed router's trust.
#
# AUTO-RELEASE
# push a tag `vX.Y.Z` -> versioned release. workflow_dispatch / (optional) main
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
# API via curl (ci/gitea-release.sh) — no external action needed.
# push a tag `vX.Y.Z` -> versioned per-arch releases `apk-vX.Y.Z-<arch>`.
# workflow_dispatch -> rolling per-arch `apk-latest-<arch>` (always-fresh
# feed). Publish uses the Gitea API via curl (ci/gitea-release.sh) — no
# external action needed. NOTE: the apk release tags deliberately do NOT start
# with `v` so publishing them cannot re-trigger this workflow's `v*` filter.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
@@ -41,28 +50,13 @@
# exported via $GITHUB_ENV):
# tag `vX.Y.Z` -> X.Y.Z-r1
# anything else -> <nearest tag>-r<commits since it + 1>
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# and hands them to the SDK build as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
# built .ipk/.apk really carry that version, so the failure can never be
# silent again. This is also why both build jobs check out with fetch-depth: 0
# into the binary's constant.Version. ci/sdk-build-apk.sh then ASSERTS that the
# built .apk really carry that version, so the failure can never be silent
# again. This is also why the build job checks out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
#
# APK LANE (25.12+, ADDITIVE — T2)
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
# index, EC keys in /etc/apk/keys/). The `build-apk` + `release-apk` jobs
# below build the SAME 4 packages through the ImmortalWrt 25.12 apk-SDK and
# publish PER-ARCH apk repos as releases `apk-latest-<arch>` (rolling) /
# `apk-<tag>-<arch>` (versioned). Per-arch because apk filenames carry no
# architecture (shaterd-0.2.0-r1.apk would collide across arches in one flat
# release) and apk fetches packages relative to the packages.adb URL.
# Signed with the EC key in the Gitea secret KEY_APK; trust anchor
# dist/shater-apk.pem (ci/gen-apk-key.sh). The usign/opkg lane above is
# UNCHANGED and keeps serving the 24.10 fleet. NOTE: the apk release tags
# deliberately do NOT start with `v` so publishing them cannot re-trigger
# this workflow's `v*` tag filter.
# CACHING (T3 — fast CI)
# All caches use actions/cache pinned to v3.3.2: the LAST release speaking the
@@ -80,34 +74,33 @@
# (PKG_VERSION/PKG_HASH live there). Stale-safe: the buildroot verifies
# PKG_HASH on every dl/ file and re-downloads on mismatch, so restore-keys
# prefix fallback is allowed.
# - Go module + build cache — key = hash of go.sum; shared by all 4 build
# - Go module + build cache — key = hash of go.sum; shared by both build
# jobs (each builds both GOARCHes).
# - panel/node_modules — key = hash of panel/package-lock.json, exact-only
# (a lockfile change MUST miss); on hit build-shaterd.sh gets --fast.
# - apt .deb archives for the apk lane's debian:bookworm host-deps
# (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list is in it).
# - usign binary (.cache/tools) — static helper, fixed key.
# - apt .deb archives for the debian:bookworm host-deps of the apk SDK
# container (.cache/apt) — key = hash of ci/sdk-build-apk.sh (the apt list
# is in it).
# - SDK feeds/ git checkouts (.cache/feeds) — the single biggest recurring
# cost: `scripts/feeds update -a` cloned base+packages+luci+routing+
# telephony EVERY run (~7 min/job; github.com is ~1 MB/s from this
# runner — run 51 evidence). The feeds dir is symlinked into the SDK
# container from the workspace cache; `feeds update` on an existing clone
# is a fast fetch+checkout of the pinned revs. Correctness-safe: update
# always checks out feeds.conf's pins, and ci/sdk-build*.sh wipes the
# always checks out feeds.conf's pins, and ci/sdk-build-apk.sh wipes the
# cache + re-clones fresh if update ever fails on a cached checkout.
# Key = lane + SDK release (shared across the two arch jobs of a lane —
# same release pins identical feed revs; the sequential runner means the
# second arch restores what the first saved). restore-keys lets an SDK
# version bump start from the old clones (git fetch delta, not re-clone).
# Key = lane + SDK release (shared across the two arch jobs — the same
# release pins identical feed revs; the sequential runner means the second
# arch restores what the first saved). restore-keys lets an SDK version
# bump start from the old clones (git fetch delta, not re-clone).
# Act_runner facts this design leans on (verified in run 51 logs):
# - the cache backend works: restores/saves confirmed, hashFiles() works;
# - docker images (openwrt/sdk, debian:bookworm, runner-images) live on the
# PERSISTENT host daemon — "Image is up to date" each run, no re-download;
# - docker images (debian:bookworm, runner-images) live on the PERSISTENT
# host daemon — "Image is up to date" each run, no re-download;
# - each actions/cache SAVE is followed by an exact 3-minute act_runner
# stall (node process lingers; hit→no-save→no stall). Steady state saves
# nothing, so adding cache entries is fine, but keys that change every
# run (e.g. github.sha) would cost +3 min/entry/run — do NOT do that.
name: release
on:
@@ -125,148 +118,11 @@ concurrency:
cancel-in-progress: true
jobs:
build:
name: ${{ matrix.arch }}
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
include:
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
steps:
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the engine via a go.mod
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
# must be present or `go build` dies with "no such file or directory".
# actions/checkout does not fetch submodules by default; init ONLY this one
# (clients/apple+android are large and unused here).
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod # pins Go 1.24.7 (go.mod `go` line)
cache: false # explicit actions/cache@v3.3.2 below (setup-go's
# built-in cache uses the new API act_runner lacks)
- name: Set up Node
uses: actions/setup-node@v4
with:
node-version: '20' # Vite 5 needs Node 18+; 20 LTS
# ---- caches (see the header comment for keys + version pin rationale) ----
- name: Cache Go modules + build cache
uses: actions/cache@v3.3.2
with:
path: |
~/go/pkg/mod
~/.cache/go-build
key: go-${{ hashFiles('go.sum') }}
restore-keys: |
go-
- name: Cache panel node_modules
id: npm-cache
uses: actions/cache@v3.3.2
with:
path: panel/node_modules
key: npm-${{ hashFiles('panel/package-lock.json') }}
# NO restore-keys: node_modules must exactly match the lockfile;
# on any lockfile change this misses and `npm ci` runs fresh.
- name: Cache SDK dl/ (package sources)
uses: actions/cache@v3.3.2
with:
path: .cache/dl
key: dl-${{ hashFiles('openwrt/*/Makefile') }}
restore-keys: |
dl-
# feeds git checkouts (see header): both 24.10.4 arch jobs share one entry
# (same release = same feeds.conf.default pins), so derive the release
# from the matrix sdk tag (x86_64-24.10.4 -> 24.10.4).
- name: Compute feeds cache key
id: feedskey
run: echo "ver=$(echo '${{ matrix.sdk }}' | sed 's/.*-//')" >> "$GITHUB_OUTPUT"
- name: Cache SDK feeds checkouts
uses: actions/cache@v3.3.2
with:
path: .cache/feeds
key: feeds-opkg-${{ steps.feedskey.outputs.ver }}
restore-keys: |
feeds-opkg-
- name: Cache CI tools (usign)
uses: actions/cache@v3.3.2
with:
path: .cache/tools
key: tools-usign-v1
- name: Install UPX
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
# the SDK package build (the openwrt/shaterd package installs the staged
# artifact). $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
- name: Build signed feed (SDK)
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
run: bash ci/build-feed.sh "${{ matrix.arch }}" "${{ matrix.sdk }}" "out/${{ matrix.arch }}"
- name: Show feed
run: ls -l "out/${{ matrix.arch }}" && cat "out/${{ matrix.arch }}/Packages"
- name: Upload feed artifact
# v4 uses an artifact backend Gitea Actions does not implement
# (GHESNotSupportedError); v3 works on Gitea's act_runner.
uses: actions/upload-artifact@v3
with:
name: shater-${{ matrix.arch }}
path: out/${{ matrix.arch }}/*
if-no-files-found: error
# ---------------------------------------------------------------------------
# APK lane (additive): the same 4 packages through the ImmortalWrt 25.12
# apk-SDK for the 25.12/apk fleet (BananaWRT 25.12-mtk-vendor routers + the
# future 25.12 VM). Produces a per-arch apk repo dir: *.apk + EC-signed
# packages.adb + shater-apk.pem. Artifact prefix `apkfeed-` (NOT `shater-`)
# so the opkg release job's `artifacts/shater-*` glob never picks these up.
# Build the 4 packages through the ImmortalWrt 25.12 apk-SDK for the 25.12/apk
# fleet (BPI-R3 mini on BananaWRT 25.12-mtk-vendor, BPI-R4 on OpenWrt 25.12,
# and the testbed VM). Produces a per-arch apk repo dir: *.apk + EC-signed
# packages.adb + shater-apk.pem, uploaded as the artifact `apkfeed-<arch>`.
build-apk:
name: apk ${{ matrix.arch }}
runs-on: ubuntu-latest
@@ -281,8 +137,10 @@ jobs:
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
steps:
# fetch-depth: 0 — see the opkg lane: the package version comes from
# `git describe`, which needs tags + ancestry.
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
@@ -295,8 +153,10 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# Same single version computation as the opkg lane — both lanes MUST agree
# on the version, they package the identical tree.
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
@@ -373,11 +233,27 @@ jobs:
restore-keys: |
feeds-apk-
# D23 — the shipped tag set is a TRIMMED subset (scripts/router-tags.sh);
# everything else in CI builds with the full upstream set, so without this
# step the one combination we actually ship is never exercised. That is how
# `with_gvisor` was trimmed while `with_wireguard` stayed and every shipped
# binary answered a WireGuard node with "gVisor is not included in this
# build" (2026-07-25). The check runs the declared-feature/tag comparison
# and then constructs one node of every declared protocol through box.New
# UNDER THE SHIPPED TAGS. It runs before the artifact build so a tag trim
# that breaks a feature fails the release instead of shipping.
- name: Verify the shipped build-tag set (D23)
run: bash scripts/check-router-tags.sh
- name: Install UPX
run: sudo apt-get update -qq && sudo apt-get install -y -qq upx-ucl
# Same artifact-order contract as the opkg lane: the SPA-embedded shaterd
# binary is built OUT of the SDK and staged before the package build.
# Artifact-order contract: the SPA-embedded shaterd binary is built OUT of
# the SDK and staged into openwrt/shaterd/files/ BEFORE the package build
# (the openwrt/shaterd package only installs the staged artifact).
# $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
@@ -409,117 +285,21 @@ jobs:
if-no-files-found: error
# ---------------------------------------------------------------------------
# Publish once both arches are built. Rolling `latest` on dispatch, a versioned
# release on a `vX.Y.Z` tag. Self-contained (curl -> Gitea API).
release:
name: release
needs: build
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Download all arch feeds
uses: actions/download-artifact@v3
with:
path: artifacts
- name: Assemble release assets
id: assets
run: |
set -eu
mkdir -p release
# For each downloaded arch feed: one ready-to-serve tarball + loose ipks.
for d in artifacts/shater-*; do
[ -d "$d" ] || continue
arch="${d#artifacts/shater-}"
tar -C "$d" -czf "release/shater-feed-${arch}.tar.gz" .
# loose .ipk for direct `opkg install <url>` (dedupe shared _all ipks by name)
for ipk in "$d"/*.ipk; do
[ -e "$ipk" ] || continue
cp -n "$ipk" "release/$(basename "$ipk")"
done
done
# ship the feed's public key so routers can verify (see docs-shater/INSTALL.md)
cp -f dist/shater-feed.pub release/shater-feed.pub
ls -l release
echo "count=$(ls release | wc -l)" >> "$GITHUB_OUTPUT"
# restore the prebuilt usign binary (skips apt + cmake + clone + build)
- name: Cache CI tools (usign)
uses: actions/cache@v3.3.2
with:
path: .cache/tools
key: tools-usign-v1
- name: Install usign (feed signer)
run: bash ci/install-usign.sh
- name: Build & sign combined opkg feed index
# One Packages/Packages.gz over ALL loose .ipk (every arch + arch=all),
# with basename Filenames. opkg filters by Architecture, so a single
# release URL serves every device: BPI routers pick aarch64_cortex-a53 +
# all, the x86 testbed picks x86_64 + all. Signed with KEY_BUILD so
# routers keep check_signature on. This is what makes the release directly
# consumable as an `src/gz` feed (see docs-shater/INSTALL.md).
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
run: bash ci/make-index.sh release
- name: Determine release identity
id: rel
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
echo "tag=${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
echo "name=shater ${GITHUB_REF#refs/tags/}" >> "$GITHUB_OUTPUT"
echo "prerelease=false" >> "$GITHUB_OUTPUT"
echo "rolling=false" >> "$GITHUB_OUTPUT"
else
echo "tag=latest" >> "$GITHUB_OUTPUT"
echo "name=shater latest (main)" >> "$GITHUB_OUTPUT"
echo "prerelease=true" >> "$GITHUB_OUTPUT"
echo "rolling=true" >> "$GITHUB_OUTPUT"
fi
- name: Publish Gitea release
env:
TOKEN: ${{ secrets.RELEASE_TOKEN != '' && secrets.RELEASE_TOKEN || github.token }}
TAG: ${{ steps.rel.outputs.tag }}
NAME: ${{ steps.rel.outputs.name }}
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
ROLLING: ${{ steps.rel.outputs.rolling }}
BODY: |
Automated build. Packages: shaterd + byedpi (per-arch), shater-core +
luci-app-shater (arch=all).
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
── Add as an opkg feed (recommended — then updating is one command) ──
This release is itself a SIGNED package feed; opkg filters by
architecture, so the same lines work on every device:
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" >> /etc/opkg/customfeeds.conf
opkg update
opkg install luci-app-shater # pulls shater-core + shaterd too
The public-key install is one-time; after it, `opkg update/upgrade`
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
── Update (name our packages — never a bare `opkg upgrade`) ──
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
── Or install the loose .ipk directly / from the tarball feed ──
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
opkg install /tmp/shater/luci-app-shater_*_all.ipk
run: bash ci/gitea-release.sh release/*
# ---------------------------------------------------------------------------
# Publish the apk lane: ONE release PER ARCH (apk package filenames carry no
# arch, and apk fetches `<name>-<ver>.apk` relative to the packages.adb URL —
# a flat multi-arch release would collide). Rolling `apk-latest-<arch>` on
# dispatch, `apk-<tag>-<arch>` on a version tag. The tags do NOT match the
# workflow's `v*` trigger, so publishing them cannot re-trigger the build.
# Publish: ONE release PER ARCH (apk package filenames carry no arch, and apk
# fetches `<name>-<ver>.apk` relative to the packages.adb URL — a flat
# multi-arch release would collide). Every run refreshes the ROLLING pointer
# `apk-latest-<arch>`; a `vX.Y.Z` tag run ALSO publishes the pinnable
# `apk-vX.Y.Z-<arch>`. The tags do NOT match the workflow's `v*` trigger, so
# publishing them cannot re-trigger the build.
#
# WHY THE ROLLING RELEASE IS PUBLISHED ON TAG RUNS TOO (fixed 2026-07-25):
# it used to be an either/or — `TAG=apk-latest-<arch>` on dispatch, ELSE
# `TAG=apk-<ver>-<arch>` — so once releases moved to tag pushes the rolling
# pointer was never written again. It froze at 0.2.0 (published 2026-07-24)
# while v0.2.9/v0.2.10 published fine, and every router whose
# /etc/apk/repositories.d/shater.list points at the rolling URL kept getting a
# successful, silent `apk update` with nothing new. Rolling is the whole point
# of that URL, so it is now written unconditionally and asserted afterwards.
release-apk:
name: release apk
needs: build-apk
@@ -537,6 +317,9 @@ jobs:
with:
path: artifacts
# Identity of the VERSIONED release only. The rolling pointer is published
# on every run with fixed prerelease=true/rolling=true, so it needs nothing
# from here.
- name: Determine release identity
id: rel
run: |
@@ -558,21 +341,39 @@ jobs:
PRERELEASE: ${{ steps.rel.outputs.prerelease }}
ROLLING: ${{ steps.rel.outputs.rolling }}
run: |
set -eu
set -euo pipefail
for d in artifacts/apkfeed-*; do
[ -d "$d" ] || continue
arch="${d#artifacts/apkfeed-}"
if [ "$VER" = latest ]; then TAG="apk-latest-$arch"; else TAG="apk-$VER-$arch"; fi
ROLL="apk-latest-$arch"
# The version we just built, read straight off the artifact
# (`shaterd-<ver>-r<rel>.apk`). NOT recomputed with ci/version.sh:
# this job checks out shallow, so it has no tags to describe from.
pkg=""
for a in "$d"/shaterd-*.apk; do
if [ -f "$a" ]; then pkg="$(basename "$a")"; fi
done
[ -n "$pkg" ] || { echo "[release-apk] ERROR: no shaterd-*.apk in $d"; exit 11; }
want="${pkg#shaterd-}"; want="${want%.apk}"
echo "[release-apk] arch=$arch built version=$want"
BODY="Automated apk (OpenWrt/ImmortalWrt 25.12+) package repo for \`$arch\`.
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
This build: \`$want\`.
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
── Add as an apk repository ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
── Add as an apk repository (rolling — install once, then just update) ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$ROLL/shater-apk.pem
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
\`apk-latest-<arch>\` is a MOVING pointer: every release run replaces its
assets, so the same repo line keeps serving the newest build. To pin a
version instead, point the repo line at
\`.../download/apk-vX.Y.Z-\$(cat /etc/apk/arch)/packages.adb\` — then the
file must be edited by hand for each upgrade.
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
@@ -580,9 +381,55 @@ jobs:
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
those packages are upgraded along with needed dependencies\").
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
echo "[release-apk] publishing $TAG from $d"
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
Full guide: docs-shater/INSTALL.md §5."
# 1) the pinnable versioned release (tag runs only)
if [ "$VER" != latest ]; then
echo "[release-apk] publishing apk-$VER-$arch from $d"
TAG="apk-$VER-$arch" NAME="shater apk $VER ($arch)" BODY="$BODY" \
PRERELEASE="$PRERELEASE" ROLLING="$ROLLING" \
bash ci/gitea-release.sh "$d"/*
fi
# 2) the rolling pointer — ALWAYS, tag run included. ci/gitea-release.sh
# deletes the existing release before recreating it, so the old
# version's assets are REPLACED, never accumulated (two versions of
# one package in one index would let apk choose, not us).
echo "[release-apk] publishing $ROLL from $d"
TAG="$ROLL" NAME="shater apk latest ($arch)" BODY="$BODY" \
PRERELEASE=true ROLLING=true \
bash ci/gitea-release.sh "$d"/*
# 3) ASSERT the rolling release really serves THIS build — same class
# of check as ci/sdk-build-apk.sh's package-version assert, and for
# the same reason: the previous failure mode was silent. Reads the
# published release back over the API and requires our three
# tag-versioned packages at $want, the index, the key — and NO
# left-over package asset at any other version.
api="$GITHUB_SERVER_URL/api/v1/repos/$GITHUB_REPOSITORY/releases/tags/$ROLL"
got="$(curl -fsS -H "Authorization: token $TOKEN" "$api" \
| tr '{},' '\n\n\n' \
| sed -n 's/.*"name"[[:space:]]*:[[:space:]]*"\([^"]*\)".*/\1/p' | sort -u)" || {
echo "[release-apk] ERROR: cannot read back $ROLL from the API"; exit 12; }
echo "[release-apk] $ROLL assets: $(printf '%s ' $got)"
# here-string, NOT `printf | grep -q`: under `pipefail` the early
# exit of grep -q can SIGPIPE the writer and fail a passing check.
for f in "shaterd-$want.apk" "shater-core-$want.apk" \
"luci-app-shater-$want.apk" packages.adb shater-apk.pem; do
grep -qxF "$f" <<<"$got" || {
echo "[release-apk] ERROR: $ROLL does not contain '$f' after publish."
echo " A router pinned to the rolling URL would have silently"
echo " stayed on its old version with a successful apk update."
exit 13; }
done
stale="$(grep -E '^(shaterd|shater-core|luci-app-shater)-.*\.apk$' <<<"$got" \
| grep -vxF -e "shaterd-$want.apk" -e "shater-core-$want.apk" \
-e "luci-app-shater-$want.apk" || true)"
[ -z "$stale" ] || {
echo "[release-apk] ERROR: $ROLL still holds stale package assets:"
printf ' %s\n' $stale
echo " Two versions of one package in one feed = apk picks by its"
echo " own rules, not by our intent."
exit 14; }
echo "[release-apk] OK — $ROLL serves $want"
done
+1 -1
View File
@@ -63,7 +63,7 @@ nul
/venv/
/test/cache.db
# feed artifacts (tracked public key dist/shater-feed.pub is force-added)
# feed artifacts (the tracked apk trust anchor dist/shater-apk.pem is force-added)
/dist/
# local agent config (CLAUDE.md is deliberately tracked; .claude local settings are not)
+16 -22
View File
@@ -49,29 +49,22 @@ Full list with MVP/T1/T2 tags — [`docs-shater/FEATURES.md`](docs-shater/FEATUR
## Install
Two signed feeds. Pick by the router's OpenWrt version. Verbatim commands and the
manual `.ipk`/`.apk` install are in [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
**opkg (OpenWrt 24.10):**
One signed **apk** feed (OpenWrt / ImmortalWrt / BananaWRT **25.12+**), one
release per arch. Verbatim commands, the manual `.apk` install and the
rolling-vs-pinned choice are in
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
```sh
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
opkg update && opkg install luci-app-shater # -> shater-core -> shaterd
```
**apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+):**
```sh
wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
wget -O /etc/apk/keys/shater-apk.pem "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" > /etc/apk/repositories.d/shater.list
apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
`apk-latest-<arch>` is a moving pointer refreshed by every release run — install
once and `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`
keeps the router current. Point the repo line at `apk-vX.Y.Z-<arch>` instead to
pin a build; that file then has to be edited by hand for every upgrade.
shater ships **inert** (globals off) so install never breaks connectivity. After
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
then `shaterd apply` and `shaterd confirm`.
@@ -91,16 +84,17 @@ into `openwrt/shaterd/files/`. Details in
| `panel/` | Admin SPA (Vite + React + TS) and its Go server |
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Product documentation |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, feed/release scripts, CI |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, apk feed/release scripts, CI |
| `SPECS/`, `docs-lx/` | Engine-fork constitution/specs and feature-config reference |
| `docs/`, `mkdocs.yml` | **Upstream** sing-box docs (mkdocs) — kept as-is |
| `adapter/ cmd/ dns/ route/ option/ protocol/ transport/ …` | sing-box-lx engine tree |
## CI, upstream & license
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes signed
feeds: opkg (usign, key `5ac4b177689cb8e0`) and apk (EC key `shater-apk.pem`). A
`vX.Y.Z` tag → versioned release; `workflow_dispatch` → rolling `latest`.
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes a signed
per-arch apk repo (EC key `shater-apk.pem`). A `vX.Y.Z` tag → the pinnable
`apk-vX.Y.Z-<arch>`; every run also refreshes the rolling `apk-latest-<arch>` and
asserts over the API that it really serves the version just built.
The engine is the **sing-box-lx** fork — a thin downstream of upstream sing-box that
lives by **rebase, never merge**; its constitution is
+31 -48
View File
@@ -10,7 +10,7 @@
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue.svg)](LICENSE)
![targets: x86_64 · aarch64_cortex-a53](https://img.shields.io/badge/targets-x86__64%20%C2%B7%20aarch64__cortex--a53-brightgreen.svg)
![feeds: opkg 24.10 · apk 25.12](https://img.shields.io/badge/feeds-opkg%2024.10%20%C2%B7%20apk%2025.12-orange.svg)
![feed: apk 25.12+](https://img.shields.io/badge/feed-apk%2025.12%2B-orange.svg)
---
@@ -128,42 +128,15 @@ data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-sh
## Установка
shater поставляется двумя подписанными фидами. Выберите по версии OpenWrt на роутере:
- **OpenWrt 24.10** → фид **opkg** (`.ipk`, `Packages.gz`, ключ usign).
- **OpenWrt / ImmortalWrt / BananaWRT 25.12+** → фид **apk** (`.apk`, `packages.adb`,
EC-ключ).
shater поставляется одним подписанным **apk-фидом** (OpenWrt / ImmortalWrt /
BananaWRT **25.12+**: `.apk`, индекс `packages.adb`, EC-ключ в `/etc/apk/keys/`).
Старый opkg-фид (`.ipk`, 24.10) снят — оба наших роутера на 25.12 с apk-tools 3,
бинаря `opkg` там просто нет (`docs-shater/DECISIONS.md` D22).
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`
(+ опциональный `byedpi`). `shaterd` подтягивается автоматически как зависимость.
### Путь A — фид opkg (OpenWrt 24.10)
```sh
# 1) доверяем ключу фида — ИМЯ файла обязано равняться отпечатку usign-ключа.
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
# 2) добавляем фид (один URL обслуживает все арки).
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
opkg update
opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # опционально: ByeDPI desync-egress
```
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
он тянет обновления и на системные пакеты, это классический способ окирпичить
роутер):
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
```
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
### Фид apk
`/etc/apk/arch` сам выбирает нужный per-arch релиз (apk-релизы раздельны по арке):
@@ -198,13 +171,22 @@ apk upgrade shaterd shater-core luci-app-shater byedpi
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
> **Роллинг или фиксация — это выбор URL в `shater.list`.** `apk-latest-<arch>`
> — движущийся указатель: каждый релизный прогон заменяет его ассеты, поэтому
> «поставил и забыл»: `apk update` сам видит новую сборку. `apk-vX.Y.Z-<arch>` —
> фиксация на конкретной сборке: роутер не получит ничего нового, пока
> `/etc/apk/repositories.d/shater.list` не отредактируют руками — на каждом
> роутере и на каждый релиз. На `mini_router` сознательно прописан
> версионированный URL, и ручная правка — его цена. Подробнее —
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) §5.1.
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
> `25.12-mtk-vendor` — в [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
> Полные инструкции — ручная установка из `.apk`, фиксация версии
> (`apk-vX.Y.Z-<arch>`), совместимость с BananaWRT `25.12-mtk-vendor` — в
> [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
### Включение
@@ -258,8 +240,8 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Документация продукта (см. таблицу ниже) |
| `scripts/` | `build-shaterd.sh` — сборка ship-артефакта |
| `ci/` | Скрипты сборки фидов и релизов (SDK, usign/EC, Gitea API) |
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанные фиды opkg/apk |
| `ci/` | Скрипты сборки apk-фида и релизов (SDK, EC-подпись, Gitea API) |
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанный apk-фид |
| `SPECS/` | Конституция форка движка и спеки (Spec Kit) |
| `docs-lx/` | Справочник конфигурации фич движка (`lx-config.md`, `.ru.md`) |
| `lx-test/`, `submodules/` | Примеры конфигов движка и submodule AmneziaWG-рантайма |
@@ -273,16 +255,17 @@ arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), s
CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает все 4 пакета и
публикует **подписанные фиды**:
- **opkg (24.10):** один комбинированный релиз, подписан usign-ключом (публичный
`dist/shater-feed.pub`, отпечаток `5ac4b177689cb8e0`; секрет — в Gitea-secret
`KEY_BUILD`).
- **apk (25.12+):** параллельная линия, **по релизу на арку**, подписан EC-ключом
(`dist/shater-apk.pem`; секрет — `KEY_APK`).
- **apk (25.12+)** — единственный формат: **по релизу на арку**, индекс
`packages.adb` подписан EC-ключом (публичный `dist/shater-apk.pem`; секрет — в
Gitea-secret `KEY_APK`).
Триггеры: push тега **`vX.Y.Z`** → версионный релиз; `workflow_dispatch` →
плавающий `latest`/`apk-latest-<arch>` (всегда свежий фид). Публикация — через
Gitea API (`ci/gitea-release.sh`). Ключи **никогда не перегенерируются** — это
инвалидировало бы доверие на всех развёрнутых роутерах.
Триггеры: push тега **`vX.Y.Z`** → версионный релиз `apk-vX.Y.Z-<arch>`;
`workflow_dispatch` → только роллинг. Роллинг `apk-latest-<arch>` обновляется
**на каждом прогоне**, включая теговый, и после публикации проверяется через API:
в нём обязаны лежать наши три пакета ровно собранной версии и ни одного ассета
другой версии. Публикация — через Gitea API (`ci/gitea-release.sh`). Ключ
**никогда не перегенерируется** — это инвалидировало бы доверие на всех
развёрнутых роутерах.
---
@@ -307,7 +290,7 @@ build-тегами и живущий **ребейзом на каждый upstre
| Документ | О чём |
|----------|-------|
| [`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, testbed/инфра |
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка обоих фидов (opkg/apk) |
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка apk-фида (роллинг/фиксация) |
| [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md) | One-binary дизайн, auth-handoff, data/DNS/apply-потоки (диаграммы) |
| [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [`docs-shater/ROADMAP.md`](docs-shater/ROADMAP.md) | Фазовый план |
+19 -20
View File
@@ -1,6 +1,6 @@
#!/bin/sh
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (the 25.12
# lane — additive next to ci/build-feed.sh, which stays the opkg/24.10 lane).
# ci/build-feed-apk.sh — build the signed **apk** feed for ONE arch (25.12+;
# the only packaging lane shater has — see docs-shater/DECISIONS.md D22).
#
# Usage: ci/build-feed-apk.sh <ARCH> <SDK_URL> <OUTDIR>
# e.g. ci/build-feed-apk.sh aarch64_cortex-a53 \
@@ -10,14 +10,14 @@
# This is the per-arch entrypoint the Gitea workflow's `build-apk` job calls.
# It runs on the CI RUNNER and:
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
# scripts/build-shaterd.sh (same artifact-order contract as the opkg lane);
# 2. drives a plain `debian:bookworm` container (workspace shared via
# `--volumes-from`, same trick as ci/build-feed.sh) that downloads the
# ImmortalWrt 25.12 apk-SDK tarball and runs ci/sdk-build-apk.sh in it:
# compile the 4 packages as .apk, then `apk mkndx --sign` the per-arch
# `packages.adb` index. Unlike the usign lane (index signed on the runner),
# apk indexing NEEDS the SDK's host `apk` tool, so index+sign happen inside
# the container.
# scripts/build-shaterd.sh (the artifact-order contract);
# 2. drives a plain `debian:bookworm` container (the job's workspace volume is
# shared into it with `--volumes-from $(hostname)`; a bare `-v $PWD:...`
# points at a host path that does not exist under act_runner's DinD) that
# downloads the ImmortalWrt 25.12 apk-SDK tarball and runs
# ci/sdk-build-apk.sh in it: compile the 4 packages as .apk, then
# `apk mkndx --sign` the per-arch `packages.adb` index. Indexing NEEDS the
# SDK's host `apk` tool, so index+sign happen inside the container.
#
# Why the ImmortalWrt SDK (not openwrt/sdk images): the 25.12 fleet runs
# BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base (target mediatek/filogic,
@@ -25,9 +25,9 @@
# mediatek-filogic 25.12 tag — hence the official SDK tarball.
#
# Env:
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret — the apk analog
# of KEY_BUILD). If set, packages.adb carries an embedded signature
# verifiable by dist/shater-apk.pem (routers: /etc/apk/keys/).
# KEY_APK EC (prime256v1) PRIVATE key PEM (Gitea repo secret). If set,
# packages.adb carries an embedded signature verifiable by
# dist/shater-apk.pem (routers: /etc/apk/keys/).
# If unset, an UNSIGNED index is produced (warning; not shippable —
# apk signatures are effectively mandatory).
set -eu
@@ -56,10 +56,9 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
# standalone. Passed into the container below and re-exported to the
# unprivileged build user in ci/sdk-build-apk.sh.
# The workflow puts these in the job env via `ci/version.sh --env >>
# $GITHUB_ENV`; recompute here when run standalone. Passed into the container
# below and re-exported to the unprivileged build user in ci/sdk-build-apk.sh.
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
@@ -73,7 +72,7 @@ echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# SDK; PKG_HASH still verifies every file, so stale = re-downloaded.
# apt/ debian:bookworm .deb archives for the host-deps install.
# The nested container runs the build as an unprivileged user -> must be writable
# (same reason as the chmod 0777 "$OUT" in ci/build-feed.sh).
# (same reason as the chmod 0777 "$OUT" above).
CACHE="$REPO/.cache"
mkdir -p "$CACHE/sdk" "$CACHE/dl" "$CACHE/apt"
chmod -R a+rwX "$CACHE/dl" "$CACHE/apt" 2>/dev/null || true
@@ -96,8 +95,8 @@ sh "$REPO/ci/fetch-sdk.sh" "$SDK_URL" "$SDK_TAR"
# --- 1) SDK build + index + sign inside a debian container -------------------
# `--volumes-from $(hostname)` shares THIS job container's workspace volume into
# the nested container (see ci/build-feed.sh for why a bare -v does not work on
# the act_runner DinD setup).
# the nested container: a bare `-v $PWD:...` points at a host path that does not
# exist under the act_runner DinD setup.
echo "[apk-feed] SDK build arch=$ARCH (ImmortalWrt 25.12 apk-SDK)"
docker pull -q debian:bookworm
docker run --rm --volumes-from "$(hostname)" \
-106
View File
@@ -1,106 +0,0 @@
#!/bin/sh
# ci/build-feed.sh — build the signed opkg feed for ONE arch.
#
# Usage: ci/build-feed.sh <ARCH> <SDK_DOCKER_TAG> <OUTDIR>
# e.g. ci/build-feed.sh x86_64 x86_64-24.10.4 out/x86_64
# ci/build-feed.sh aarch64_cortex-a53 mediatek-filogic-24.10.4 out/aarch64_cortex-a53
#
# This is the reusable per-arch entrypoint the Gitea workflow calls. It runs on
# the CI RUNNER and:
# 1. asserts the prebuilt shaterd binary for this arch was already staged by
# scripts/build-shaterd.sh (into openwrt/shaterd/files/) — proving artifact
# order: SPA+shaterd build BEFORE the SDK package build;
# 2. drives the arch-matched `openwrt/sdk` docker image to compile all 4
# packages (ci/sdk-build.sh) and collect their .ipk into OUTDIR;
# 3. builds + usign-signs the opkg `Packages` index over OUTDIR
# (ci/install-usign.sh + ci/make-index.sh; signs iff $KEY_BUILD is set).
#
# Env:
# KEY_BUILD usign SECRET key (Gitea repo secret). If set, the feed index is
# signed and verifiable by dist/shater-feed.pub (fp 5ac4b177689cb8e0).
# If unset, an UNSIGNED feed is produced (make-index warns).
set -eu
ARCH="${1:?arch required (x86_64 | aarch64_cortex-a53)}"
SDK_TAG="${2:?sdk docker tag required (e.g. x86_64-24.10.4)}"
OUT="${3:?output dir required}"
REPO="$(cd "$(dirname "$0")/.." && pwd)"
mkdir -p "$OUT"; OUT="$(cd "$OUT" && pwd)"
# $OUT is created here as ROOT on the runner, but the nested `openwrt/sdk`
# container runs as the unprivileged `buildbot` (uid 1000) — so it must be able
# to write the collected .ipk into $OUT. World-writable is set HERE (a chmod
# from inside the container, as buildbot, cannot fix a root-owned dir).
chmod 0777 "$OUT"
# --- 0) the prebuilt shaterd binary must already be staged for this arch ------
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
*) echo "[feed] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
esac
if [ ! -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" ]; then
echo "[feed] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
echo " Run scripts/build-shaterd.sh BEFORE ci/build-feed.sh." >&2
exit 3
fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# The workflow normally puts these in the job env (ci/version.sh --env >>
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
# 0.2.0-r3 and no router could ever see an update).
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
# Workspace dir restored/saved by actions/cache in the workflow and shared into
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
# there (ci/sdk-build.sh). PKG_HASH still verifies every file, so a stale cache
# can never produce a wrong build. Must be writable by the container's
# unprivileged buildbot user (same reason as the $OUT chmod above).
DL_DIR="$REPO/.cache/dl"
mkdir -p "$DL_DIR"
chmod -R a+rwX "$DL_DIR" 2>/dev/null || true
# --- 0.6) persistent feeds/ git checkouts -------------------------------------
# Workspace dir restored/saved by actions/cache (key: feeds-opkg-<release>) and
# symlinked over the SDK's feeds/ inside the container (ci/sdk-build.sh), so
# `scripts/feeds update -a` fetches deltas instead of re-cloning base+packages+
# luci from scratch (~7 min/run on this runner's slow github.com link).
# Top-level chmod only: the contents are created by the container's uid-1000
# build user and restored with the same ownership (tar-as-root preserves it).
FEEDS_CACHE="$REPO/.cache/feeds/opkg"
mkdir -p "$FEEDS_CACHE"
chmod a+rwX "$REPO/.cache" "$REPO/.cache/feeds" "$FEEDS_CACHE" 2>/dev/null || true
# --- 1) SDK package build (4 packages) in the arch-matched SDK image ----------
# We drive the `openwrt/sdk` docker image directly (not openwrt/gh-action-sdk):
# on a self-hosted Gitea act_runner the marketplace action fetch can be
# unavailable, and we need a CLEAN single-feed layout. `--volumes-from
# $(hostname)` shares THIS job container's workspace volume into the nested SDK
# container — a bare `-v $PWD:...` points at a host path that does not exist
# under the act_runner DinD setup. (Requires the job to run inside a container,
# which Gitea Actions does by default.)
echo "[feed] SDK build arch=$ARCH image=openwrt/sdk:$SDK_TAG"
docker pull "openwrt/sdk:$SDK_TAG"
docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
-e FEEDS_CACHE="$FEEDS_CACHE" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
"openwrt/sdk:$SDK_TAG" \
sh "$REPO/ci/sdk-build.sh"
# --- 2) index + sign the per-arch feed (usign, KEY_BUILD passed through) -------
sh "$REPO/ci/install-usign.sh"
KEY_BUILD="${KEY_BUILD:-}" bash "$REPO/ci/make-index.sh" "$OUT"
echo "[feed] done arch=$ARCH -> $OUT"
ls -l "$OUT"
+7 -10
View File
@@ -2,24 +2,21 @@
# ci/gen-apk-key.sh — generate the Shater **apk** feed signing keypair (25.12 lane).
#
# apk (OpenWrt/ImmortalWrt 25.12+) verifies package indexes with EC keys
# (prime256v1 PEM), NOT usign — the existing usign identity
# (dist/shater-feed.pub, fp 5ac4b177689cb8e0) keeps signing the opkg/24.10 feed
# and is NOT touched by this script. This generates a SEPARATE, second identity:
# (prime256v1 PEM). This is the ONLY feed identity shater has since the opkg
# lane was removed (D22) — the old usign key is history, not a second lane.
#
# dist/shater-apk.key EC PRIVATE key. NEVER commit (dist/ is gitignored).
# Paste its full PEM contents into the Gitea repo secret
# KEY_APK (the apk analog of the usign secret KEY_BUILD).
# Then delete the local file (or keep it in a password
# manager as the offline backup — losing it means every
# deployed router must re-trust a new key).
# dist/shater-apk.pem PUBLIC key. Commit it next to shater-feed.pub:
# KEY_APK. Then delete the local file (or keep it in a
# password manager as the offline backup — losing it
# means every deployed router must re-trust a new key).
# dist/shater-apk.pem PUBLIC key. Commit it:
# git add -f dist/shater-apk.pem
# (-f because /dist/ is gitignored). Routers install it
# as /etc/apk/keys/shater-apk.pem.
#
# Run ONCE. Refuses to overwrite: regenerating the key invalidates the trust of
# every router that already installed shater-apk.pem (same rule as D7 for the
# usign key).
# every router that already installed shater-apk.pem (see D22).
set -eu
REPO="$(cd "$(dirname "$0")/.." && pwd)"
-60
View File
@@ -1,60 +0,0 @@
#!/bin/bash
# Make `usign` available on the CI runner so ci/make-index.sh can sign the opkg
# feed index. The OpenWrt SDK ships usign, but the index/signing step runs on the
# bare runner (outside the SDK container), so we build the tiny standalone tool
# from source (no libubox — it is intentionally dependency-free so it can
# bootstrap a build system). No-op if usign is already on PATH.
#
# Ported unchanged from Shater v0.1 (ci/install-usign.sh): usign is
# format-agnostic and the signing story is identical for the v0.2 4-package feed.
#
# CI cache: a previously-built binary is reused from $USIGN_CACHE (default:
# <repo>/.cache/tools — a workspace dir the workflow persists via actions/cache),
# skipping the apt + cmake + clone + build (~1 min). After a fresh build the
# binary is copied there so the NEXT run hits the cache. usign is a tiny static
# helper with no versioned protocol — a stale cached binary cannot mis-sign.
set -eu
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
TOOLS="${USIGN_CACHE:-$REPO_ROOT/.cache/tools}"
# place <binary> — install onto PATH (system-wide if we can, else ~/bin)
place() {
local SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
if $SUDO install -m0755 "$1" /usr/local/bin/usign 2>/dev/null; then
:
else
mkdir -p "$HOME/bin"
install -m0755 "$1" "$HOME/bin/usign"
echo "$HOME/bin" >> "${GITHUB_PATH:-/dev/null}"
export PATH="$HOME/bin:$PATH"
fi
}
if command -v usign >/dev/null 2>&1; then
echo "[usign] already present: $(command -v usign)"
exit 0
fi
if [ -x "$TOOLS/usign" ]; then
place "$TOOLS/usign"
echo "[usign] restored from cache: $(command -v usign || echo "$HOME/bin/usign")"
exit 0
fi
SUDO=""; [ "$(id -u)" = 0 ] || SUDO="sudo"
if ! command -v cmake >/dev/null 2>&1 || ! command -v cc >/dev/null 2>&1; then
$SUDO apt-get update -qq
$SUDO apt-get install -y -qq cmake gcc git
fi
tmp="$(mktemp -d)"
# Canonical source; fall back to the GitHub mirror if git.openwrt.org is flaky.
git clone --depth 1 https://git.openwrt.org/project/usign.git "$tmp/usign" \
|| git clone --depth 1 https://github.com/openwrt/usign.git "$tmp/usign"
( cd "$tmp/usign" && cmake -DCMAKE_BUILD_TYPE=Release . >/dev/null && make >/dev/null )
place "$tmp/usign/usign"
# seed the cache for the next run (best-effort)
mkdir -p "$TOOLS" 2>/dev/null && install -m0755 "$tmp/usign/usign" "$TOOLS/usign" 2>/dev/null || true
echo "[usign] built: $(command -v usign || echo "$HOME/bin/usign")"
-39
View File
@@ -1,39 +0,0 @@
#!/bin/bash
# Build the opkg feed index (Packages + Packages.gz) with SHA256 for a dir of
# .ipk files, then optionally usign-sign it if $KEY_BUILD (the Gitea repo secret)
# is set and usign is present. Arg $1 = feed dir.
#
# Ported from Shater v0.1 (ci/make-index.sh), unchanged. It is package-count and
# package-name agnostic: it indexes whatever .ipk are in the dir, so it serves
# BOTH the per-arch feed built by ci/build-feed.sh AND the combined release feed
# assembled in the release job (shaterd + byedpi per-arch, shater-core +
# luci-app-shater = _all). opkg filters by Architecture at install time, so one
# combined URL serves every device.
#
# Feed format: opkg `src/gz` (.ipk + text Packages index, usign signature).
# OpenWrt 24.10 (our SDK) still uses opkg; apk arrives at 25.12. The committed
# trust anchor dist/shater-feed.pub is a usign (Ed25519) key, matching this.
set -euo pipefail
OUT="${1:?feed dir required}"; cd "$OUT"
: > Packages
for ipk in *.ipk; do
[ -e "$ipk" ] || continue
ctrl=$(tar -xzOf "$ipk" ./control.tar.gz | tar -xzO ./control)
sz=$(wc -c < "$ipk"); sha=$(sha256sum "$ipk" | cut -d' ' -f1)
printf '%s\n' "$ctrl" | sed '/^[[:space:]]*$/d' >> Packages
printf 'Filename: %s\nSize: %s\nSHA256sum: %s\n\n' "$ipk" "$sz" "$sha" >> Packages
done
gzip -kf Packages
if [ -n "${KEY_BUILD:-}" ]; then
# Signing was requested — a missing/broken signer must FAIL the build, not
# silently ship an unsigned feed that routers with check_signature on reject.
command -v usign >/dev/null 2>&1 || { echo "[index] ERROR: KEY_BUILD set but usign not found" >&2; exit 1; }
umask 077; printf '%s\n' "$KEY_BUILD" > /tmp/usign.sec
usign -S -m Packages -s /tmp/usign.sec || { rm -f /tmp/usign.sec; echo "[index] ERROR: usign signing failed" >&2; exit 1; }
rm -f /tmp/usign.sec
echo "[index] signed -> Packages.sig ($(head -1 Packages.sig))"
else
echo "[index] no KEY_BUILD -> UNSIGNED feed (opkg needs check_signature off, or set the secret)"
fi
echo "[index] contents:"; ls -l
+3 -4
View File
@@ -10,7 +10,6 @@
# the target fleet (BananaWRT 25.12-mtk-vendor = ImmortalWrt 25.12 base, its
# distfeeds even point at downloads.immortalwrt.org/releases/25.12-SNAPSHOT) is
# ImmortalWrt — so we extract the official ImmortalWrt SDK tarball ourselves.
# Same --volumes-from workspace-sharing pattern as ci/sdk-build.sh (opkg lane).
#
# The OpenWrt buildsystem refuses to run as root, so the SDK build itself runs
# as an unprivileged `build` user created here.
@@ -36,8 +35,8 @@ echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallba
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
# The prebuilt shaterd artifact must already be staged for this arch (same
# contract as the opkg lane — scripts/build-shaterd.sh runs first).
# The prebuilt shaterd artifact must already be staged for this arch
# (artifact-order contract — scripts/build-shaterd.sh runs first).
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
@@ -113,7 +112,7 @@ export HOME=/home/build
cd "$SDKDIR"
# Register this repo's openwrt/ as a src-link feed named `shater` (absolute
# path required) — identical to the opkg lane (ci/sdk-build.sh).
# path required).
cp -f feeds.conf.default feeds.conf
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
-143
View File
@@ -1,143 +0,0 @@
#!/bin/sh
# Runs INSIDE an `openwrt/sdk:<target>-<ver>` container (CWD = SDK root
# /builder). The job's workspace is shared into this container via
# `docker run --volumes-from`, so the repo is visible at $REPO and output goes
# to $OUT (a dir under the repo, hence also visible to the runner afterwards).
#
# Unlike Shater v0.1 (which compiled ONLY xrayctl in the SDK and hand-packed the
# pure-data packages with tar), v0.2 builds ALL FOUR packages the canonical way,
# via the SDK feed + `make package/<p>/compile`:
#
# shaterd prebuilt binary — Build/Compile only VALIDATES that
# openwrt/shaterd/files/shaterd-<amd64|arm64>.upx was staged
# by scripts/build-shaterd.sh on the runner BEFORE this ran.
# (arch-specific .ipk: RSTRIP/STRIP disabled — packed ELF.)
# shater-core PKGARCH=all data glue (procd init, sysctl, uci-defaults).
# luci-app-shater PKGARCH=all LuCI thin launcher — its Makefile does
# `include $(TOPDIR)/feeds/luci/luci.mk`, so the `luci` feed
# MUST be updated first (that is what creates feeds/luci/luci.mk).
# byedpi arch-specific C — the SDK cross-compiles ciadpi from the
# upstream tarball (needs network for PKG_SOURCE_URL).
#
# Env (required): ARCH, REPO, OUT.
set -eu
ARCH="${ARCH:?ARCH env required}"
REPO="${REPO:?REPO env required}"
OUT="${OUT:?OUT env required}"
mkdir -p "$OUT"
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
# Package version, derived from the git tag by ci/version.sh and handed in by
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
# imports every environment variable as a variable, and it propagates through
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
# byedpi deliberately keeps its own upstream version (see its Makefile).
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
# The prebuilt shaterd artifact must already be staged for this arch.
case "$ARCH" in
x86_64) sfx=amd64 ;;
aarch64_cortex-a53) sfx=arm64 ;;
*) echo "[sdk] ERROR: unsupported ARCH '$ARCH'"; exit 2 ;;
esac
test -f "$REPO/openwrt/shaterd/files/shaterd-$sfx.upx" || {
echo "[sdk] ERROR: openwrt/shaterd/files/shaterd-$sfx.upx not staged."
echo " scripts/build-shaterd.sh must run on the runner before the SDK build."; exit 3; }
# --- register this repo's openwrt/ as a src-link feed named `shater` ---------
# src-link REQUIRES an absolute path; $REPO/openwrt is exactly a feed root (it
# contains the 4 package dirs and nothing else that looks like a package).
cp -f feeds.conf.default feeds.conf
grep -q '^src-link shater ' feeds.conf || echo "src-link shater $REPO/openwrt" >> feeds.conf
# Update metadata for ALL feeds: our `shater` feed + the SDK defaults (base,
# luci, packages, routing, telephony). We need `luci` for feeds/luci/luci.mk and
# `base`/`packages` for the runtime deps (kmod-nft-tproxy, kmod-nft-socket,
# ip-full, rpcd, luci-base) to resolve.
#
# Persistent feeds checkouts: $FEEDS_CACHE (a workspace dir the runner restores
# via actions/cache, shared into this container via --volumes-from) replaces
# the SDK's ephemeral feeds/ dir, so `feeds update` git-fetches deltas instead
# of re-cloning base+packages+luci every run (~7 min on the runner's slow
# github.com link). Correctness-safe: update always checks out feeds.conf's
# pinned revisions; if it ever fails on a cached checkout (e.g. a force-pushed
# upstream), the cache is wiped and the update retried with fresh clones.
if [ -n "${FEEDS_CACHE:-}" ] && mkdir -p "$FEEDS_CACHE" 2>/dev/null; then
rm -rf feeds
ln -s "$FEEDS_CACHE" feeds
echo "[sdk] feeds/ -> $FEEDS_CACHE (persistent cache)"
fi
echo "[sdk] feeds update -a"
if ! ./scripts/feeds update -a; then
[ -L feeds ] || { echo "[sdk] ERROR: feeds update failed"; exit 8; }
echo "[sdk] WARNING: feeds update failed on cached checkouts — wiping cache, cloning fresh"
find "$FEEDS_CACHE" -mindepth 1 -maxdepth 1 -exec rm -rf {} + 2>/dev/null || true
./scripts/feeds update -a
fi
echo "[sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
# Select our packages, then defconfig. `make package/<p>/compile` builds the
# explicit target regardless, but selecting first makes deps visible to defconfig.
for p in shaterd shater-core byedpi luci-app-shater; do
echo "CONFIG_PACKAGE_$p=m" >> .config
done
# Route source downloads through OpenWrt's fast CDN mirror FIRST — sourceware.org
# (elfutils) and other upstreams intermittently stall mid-transfer, and curl's
# --connect-timeout doesn't cover a stalled stream, so the SDK download hangs the
# build. LOCALMIRROR is tried before each package's own PKG_SOURCE_URL. (lx CI)
echo 'CONFIG_LOCALMIRROR="https://sources.cdn.openwrt.org"' >> .config
# Persistent dl/ across runs: $DL_DIR is a workspace dir the runner restores via
# actions/cache (see ci/build-feed.sh). Correctness-safe: the buildroot verifies
# PKG_HASH on every file already in dl/ and re-downloads on mismatch, so a stale
# cache can never leak a wrong source into the build.
if [ -n "${DL_DIR:-}" ]; then
echo "CONFIG_DOWNLOAD_FOLDER=\"$DL_DIR\"" >> .config
fi
echo "[sdk] defconfig"
make defconfig >/dev/null
# --- compile the 4 packages --------------------------------------------------
for p in shaterd shater-core byedpi luci-app-shater; do
echo "[sdk] === build $p ==="
make "package/$p/compile" V=s -j"$(nproc)"
done
# --- collect ONLY our 4 packages' .ipk (per-arch shaterd/byedpi + _all core/luci)
# NOT `find bin -name '*.ipk'`: the openwrt/sdk image ships HUNDREDS of prebuilt
# kmod/base .ipk under bin/, which a blanket copy would pull into the feed and
# get signed under OUR key. Match each package's own `<name>_<ver>_<arch>.ipk`.
found=0
for p in shaterd shater-core byedpi luci-app-shater; do
for ipk in $(find bin -type f -name "${p}_*.ipk"); do
cp -f "$ipk" "$OUT/"; found=$((found+1))
done
done
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
# --- assert the tag-derived version actually reached the packages -------------
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
# env -> make hand-off has several layers (docker -e, make's env import, the
# metadata dump), so verify the result instead of trusting it: every one of our
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
echo "[sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[sdk] version check OK — our 3 packages are $want"
fi
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
ls -l "$OUT"
+7 -9
View File
@@ -6,9 +6,9 @@
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
# feed's version string differs from the installed one, `apk update` saw nothing
# new and the routers could not be updated through the normal path at all.
# vs r2's 5 488 336 B). Since apk offers an upgrade only when the feed's version
# string differs from the installed one, `apk update` saw nothing new and the
# routers could not be updated through the normal path at all.
#
# So the version is now DERIVED, in CI, from the git tag, and the package
# Makefiles only carry a fallback for manual/offline builds.
@@ -21,14 +21,12 @@
# rolling `latest`)
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
#
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
# part first (numerically, component by component), the `r<rel>` only as a
# tie-break. Verified against the real tools, not from memory:
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# apk compares `<upstream>-r<rel>` as: the dotted upstream part first
# (numerically, component by component), the `r<rel>` only as a tie-break.
# Verified against the real tool, not from memory —
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
# identical results (opkg implements the Debian algorithm).
# That is exactly the ordering this scheme needs:
# * a release always outranks every rolling build that preceded it
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
-2
View File
@@ -1,2 +0,0 @@
untrusted comment: shater feed signing key
RWRaxLF3aJy44JbcxSFujtrFFEQ8lIsnTkd1K5TdjIhdlC2c0wa0fv4V
+10 -11
View File
@@ -31,12 +31,11 @@ Do not delete it — we port proven pieces from it. What v0.1 has:
- **`luci-app-shater`** — a custom "instrument panel" LuCI app (client-side JS +
ucode/rpcd ubus backend): Overview with a live Signal Path, Simple/Advanced
toggle, quick-start wizard, Nodes/Subs/Rules/DNS/Live/Profiles/Settings pages.
- **CI + signed opkg feed** on Gitea: builds per-arch, signs the feed index with
usign, publishes a rolling `latest` Gitea release consumable as `src/gz`. **Feed
signing key fingerprint `5ac4b177689cb8e0`**; public key `dist/shater-feed.pub`,
secret in the Gitea repo secret `KEY_BUILD`.
- **CI + a signed package feed** on Gitea: builds per-arch, signs the feed index,
publishes a rolling `latest` Gitea release the router consumes as a feed.
(v0.1 shipped `.ipk` signed with a usign key — that lane is retired, D22.)
- Verified end-to-end on the VM: real LAN client proxied, DNS anti-leak, honest
fail-closed, opkg install/upgrade from the signed feed.
fail-closed, install/upgrade from the signed feed.
v0.1 is engine-locked to **xray-core**; its generator, share-link parser and
`run.json` are xray-shaped.
@@ -91,7 +90,7 @@ We are rebasing onto a new engine and a new UI architecture. Full rationale in
- **`shater` branch `v0.1`** = the standalone xray-based version (frozen, ported
from).
- Until Phase 1 merges the engine in, `main` is the docs-first overlay seed you
are reading now (LICENSE, README, `docs-shater/`, `dist/shater-feed.pub`).
are reading now (LICENSE, README, `docs-shater/`, the feed signing key).
## What to port from v0.1 (don't rewrite these ideas)
@@ -105,8 +104,8 @@ overlay, don't redo:
- **Subscription fetch** (HAPP emulation, fingerprint reconcile, per-sub cache)
and the flexible **ruleset/list** model — though sing-box has its own share-link
parser and config schema we now target.
- **CI feed build + usign signing + Gitea release** (adapt to the single forked
binary; keep key `5ac4b177689cb8e0`).
- **CI feed build + index signing + Gitea release** (adapted to the single forked
binary; the format is apk, signed with the EC key — D22).
- The LuCI **design system** (the "instrument panel" identity) — reused for the
mini-dashboard and as the panel's visual language.
@@ -122,9 +121,9 @@ filter/stats engine wired into sing-box's DNS.
`https://github.com/SagerNet/sing-box`).
- **CI:** Gitea Actions (act_runner + Docker). v0.1's workflow was removed from
`main`; new CI is added when the v0.2 build exists.
- **Feed signing:** usign key `5ac4b177689cb8e0`; secret in repo secret
`KEY_BUILD`; public key `dist/shater-feed.pub` (kept so existing installs keep
verifying).
- **Feed signing:** EC (prime256v1) key for the apk index; secret in the repo
secret `KEY_APK`; public key `dist/shater-apk.pem`, installed on routers as
`/etc/apk/keys/shater-apk.pem`. Never regenerate it (D22).
- **Test VM:** OpenWrt 24.10.3 x86_64 in Docker (`docker ps --filter
name=openwrt-vm`). SSH via the ssh-manager MCP server `local_openwrt`
(localhost:2222, root/openwrt). LuCI at `http://127.0.0.1:8080` (root/openwrt),
+243 -1
View File
@@ -64,11 +64,17 @@ sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files m
stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is
GPL-3.0 for clarity.
## D7 — Keep the v0.1 feed signing identity
## D7 — Keep the v0.1 feed signing identity *(SUPERSEDED by D22)*
The usign feed key `5ac4b177689cb8e0` (public key in `dist/shater-feed.pub`,
secret in Gitea secret `KEY_BUILD`) carries over, so routers that already trust it
keep verifying v0.2 packages. Do not regenerate it without a documented rotation.
> **Superseded 2026-07-25 (D22).** The opkg feed this identity signed no longer
> exists, so there is nothing left for the key to verify. It was never rotated or
> compromised — it is simply unused. `dist/shater-feed.pub` was deleted from the
> tree; the reasoning, and how to resurrect the identity if it is ever needed
> again, is in D22.
## D8 — Preserve, don't destroy: v0.1 lives on its branch
The reset moved the full working xray-based project to the `v0.1` branch and
cleaned `main`. Nothing is lost; reusable logic (reliability layer, nft/routing,
@@ -95,6 +101,10 @@ runtime, forcing an ELF with `PT_INTERP=/lib64/ld-linux-x86-64.so.2` + `PT_DYNAM
plane is tproxy/redirect (netplane); generate never emits a tun inbound, so
the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever
appears it falls back to the system stack — re-add the tag then.
**REVERTED 2026-07-25 — that reasoning was wrong and shipped a dead feature.**
gVisor is not only the tun stack: it is the netstack of the **WireGuard
endpoint**, which we do emit and do declare [MVP]. See D23; the tag is back and
is now held there by a test.
- 2026-07-23: `with_clash_api` also dropped. The admin panel is shater's own
web server and generate never emits a `clash_api` service; the desktop/CLI
`LX_TAGS` keeps the tag for external dashboards.
@@ -424,3 +434,235 @@ on the next render.
Consequence: all delay numbers are comparable (least ping ranks apples against
apples), and group settings lose two footgun fields while Settings keeps the
two that actually govern every check.
## D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). `config rule` carried THREE ways to say
where traffic is going: `dst_domain` (an inline domain list), `dst_ip` (an inline
CIDR list) and `dst_ruleset` (a reference to a `config ruleset`). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a `.srs`, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
**Decision: `dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is
the only destination matcher.** `Src`, `dst_port` and `proto` are untouched —
they are not lists of destinations and have no rule-set form.
- **Rejected: keep the inline lists as a shorthand.** "One obvious way" is the
whole point; a shorthand that quietly means something different from the long
form (see the bare-entry trap) is worse than no shorthand.
- **Rejected: promote inline lists to rule-sets lazily at generate time.** The
config on disk would then not say what the router does, and the panel would
have to render a list the user cannot find or edit.
### The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (`dst_domain`) | in a rule-set (`entry`) | migrated to |
|--------------------|--------------------------|-------------------------|--------------|
| `example.com` | **exact host** | **host + subdomains** | `full:example.com` |
| `full:example.com` | exact host | exact host | unchanged |
| `suffix:example.com` / `.example.com` | host + subdomains | host + subdomains | unchanged |
| `keyword:ads` | substring | substring | unchanged |
| `regexp:^ads\.` | pattern | pattern *(added here)* | unchanged |
| `geosite:x` / `geoip:x` | inert (engine field removed) | inert (unknown prefix) | unchanged |
`shaterd migrate` (schema v1→v2, `shater/model/migrate.go`) creates one inline
`config ruleset` per rule that still carries a legacy list — `rule-<rule name>`
for domains, `rule-<rule name>-ip` for addresses — moves the entries across with
the conversion above, appends the new name to `dst_ruleset`, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
`rule-<name>-2`). `regexp:` support was added to inline rule-sets in the same
change precisely so the move can be lossless.
`geosite:`/`geoip:` entries are copied VERBATIM rather than promoted to a
`source=geosite` rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
**One deliberate semantic change, called out:** a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because `rule_set: [a, b]` matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
That AND→OR change is about the ENGINE's TCP/UDP path, and it deliberately does
**not** extend to the untunnelable-protocol plane (`shater/apply/untunnelable.go`,
the ping / IPTV / VPN-passthrough policy in nftables). There, a v1
`dst_domain + dst_ip` rule could never claim a packet that carries no domain, so
the plan skipped it; reading the migrated form as OR would have made the address
half suddenly decisive, and with `target=direct` that means an upgrade quietly
sending previously-tunnelled ICMP out with the client's real source address. A
rule whose rule-sets are known to match by NAME is therefore still skipped by that
plan, and the skip is reported ("a routing rule matches by name … as well as by
address"). Split the rule in two if you want the addresses decided there.
### The vocabulary is about ENTRIES YOU TYPE, not about every list body
The table above is the vocabulary of an **inline** rule-set's `entry` values (and
of the DNS-filter/device lists, which share the classifier). The other two rule-set
sources are not other spellings of it:
| source | what it is | vocabulary |
|------------------------------|----------------------------------|------------|
| `inline` | entries you type | the table above |
| `url` → `.srs` / `.json` | a compiled rule-set, engine-owned | the engine's, not ours |
| `url` → anything else | a hosts / one-domain-per-line / AdBlock TEXT FILE | **none** — every line is a domain plus its subdomains |
| `file` | a local `.srs` / `.json` | the engine's, not ours |
**Rejected: run text lists through the entry classifier too.** A published
AdGuard/OISD list is full of colon-bearing tokens that are ordinary filter syntax
(`##…:has(…)`, `$domain=`, absolute URLs); classifying them would either mis-import
them or bury the operator under hundreds of "unrecognised prefix" warnings per
list. The formats also disagree structurally — a hosts line carries several names,
so the text parser works per token, while an entry is a whole line. And `regexp:`
arriving from a third-party URL is a pattern compiled into the router's matcher and
evaluated per query, which is a very different proposition from one the operator
typed.
So the difference stands and is paid for in diagnostics instead: a text list
containing `full:` / `suffix:` / `keyword:` / `regexp:` is reported per list, on
every generate, naming the entries and pointing at `source=inline` where they work
(`warnListEntryVocabulary`, `shater/generate/ruleset.go`). The check tests only
those four markers, never the general `word:` shape, so it fires on a human's
mistake and stays quiet on published filter syntax.
Consequence: one destination mechanism, one vocabulary, one place a list is
edited; every list is compiled once and reused. The panel's rule editor drops its
Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of
the rulesets that already exist, and nothing more. Creating and filling a list
stays in the Rulesets panel — **rejected: a "create a list from here" shortcut in
the rule editor**, because a second place to author a list is a second place for
its semantics and its duplicate-name rules to drift, and the whole point of this
decision was to stop having two.
## D22 — One packaging lane: apk. The opkg/`.ipk` lane is deleted, not disabled
Decided 2026-07-25 (product owner). CI built and published TWO signed feeds from
every run: opkg/usign (`.ipk` + `Packages.gz`, OpenWrt 24.10) and apk/EC (`.apk` +
`packages.adb`, OpenWrt/ImmortalWrt 25.12). The opkg half served nobody. Checked
on the actual hardware, not inferred:
| Device | Firmware | pkg arch | package manager |
|---|---|---|---|
| `mini_router` (BPi-R3 Mini) | ImmortalWrt 25.12.1 | `aarch64_cortex-a53` | apk-tools 3.0.5 |
| `main_router` (BPi-R4) | OpenWrt 25.12.0 | `aarch64_cortex-a53` | apk-tools 3.0.5 — **no `opkg` binary on the system at all** |
**Decision: delete the opkg lane outright.** Removed: the `build` + `release`
jobs from `.gitea/workflows/release.yml`; `ci/build-feed.sh`, `ci/sdk-build.sh`,
`ci/make-index.sh`, `ci/install-usign.sh`; and the trust anchor
`dist/shater-feed.pub`. The Gitea secret `KEY_BUILD` is now referenced by
nothing and can be deleted from the repo settings. `ci/version.sh`,
`ci/gitea-release.sh` and `ci/fetch-sdk.sh` are shared or apk-only and stay.
- **Rejected: keep the lane but stop triggering it** (comment it out / gate it on
a dispatch input). Dead code in CI is worse than no code: it keeps a second SDK
matrix, a second signing key and a second feed layout alive in everyone's head
and in every future edit, and it silently rots because nothing runs it. The
24.10 SDK images it pins are themselves a frozen dependency.
- **Rejected: keep `dist/shater-feed.pub` as a historical artifact.** A committed
trust anchor is an instruction — it invites someone to follow the old install
path for a feed that is no longer produced. Nothing is lost by removing it:
git history still holds the file, the SECRET half is untouched in `KEY_BUILD`,
and a usign secret key blob contains its own public half, so the identity can
be reconstructed if a 24.10 device ever has to be served again. Deleting the
file is reversible; a stale trust anchor pointing at an unmaintained feed is
the thing that quietly misleads.
- **Not done: revoking or rotating the usign key.** There is no incident. It is
retired, not burned (D7).
Consequence: one SDK, one key, one feed layout, one set of install instructions.
It also makes the rolling release `apk-latest-<arch>` the *only* install path
that does not require hand-editing a file per release — which is why the same
change fixed it: publishing was an either/or (`apk-latest-<arch>` on dispatch,
ELSE `apk-vX.Y.Z-<arch>` on a tag), so once releases moved to tag pushes the
rolling pointer stopped being written and froze at `0.2.0` while v0.2.9/v0.2.10
shipped — routers on the rolling URL got a successful, silent `apk update` with
nothing new. `release-apk` now writes the rolling pointer on every run and
asserts, by reading the published release back over the Gitea API, that it holds
our three tag-versioned packages at exactly the version just built and no asset
at any other version.
## D23 — The router tag set is a checked contract, not a string literal
`with_gvisor` was trimmed from the router set on 2026-07-23 (D9) as "unreachable
code: we never emit a tun inbound". True about tun — and irrelevant, because
gVisor is also the netstack of the **WireGuard endpoint**, which shater emits and
FEATURES.md declares [MVP] (AmneziaWG is called *"a driving requirement"*). Every
binary shipped between then and 2026-07-25 answered a configured WireGuard node
with:
```
create instance: initialize endpoint[0]: create WireGuard device:
gVisor is not included in this build, rebuild with -tags with_gvisor
```
`transport/wireguard/device_stack_stub.go` (`//go:build !with_gvisor`) returns
`tun.ErrGVisorNotIncluded` from **both** device constructors, so
`system_interface: true` is not an escape hatch either: WireGuard was 100% dead
in the shipped artifact while the panel offered it, the parser accepted `wg://`,
`awg://` and wg-quick `.conf` imports, and the owner had 7 WireGuard sections in
UCI on a production router.
- **Decision:** `with_gvisor` is part of the router tag set and stays there for
as long as we ship WireGuard. It costs **~2.8 MB raw / ~0.65 MB UPX per arch**
(measured 2026-07-25, both arches; `/overlay` on the production router is
6.9 GB with 205 MB used). A tag whose absence turns a declared feature into a
runtime error is not "dead weight" — it is the feature.
### Why the bug was invisible, and what now makes it visible
The defect was not a typo in a tag list. It was that **nothing connected the tag
list to the feature list**, and the shipped tag combination was the one build
configuration nothing exercised: the whole test suite compiles with the FULL
upstream set (`with_gvisor` included), so `TestAmneziaWGEndpoint` passed happily
while the artifact it was supposed to vouch for could not create a WireGuard
device. Tests proved the code was right; they never proved the *build* was.
Three pieces now hold it together:
1. **One definition of the set** — `scripts/router-tags.sh` (`SHATER_ROUTER_TAGS`
+ `SHATER_ROUTER_LDFLAGS`), sourced by `scripts/build-shaterd.sh` and by the
checker. The tag list used to live as a literal inside the build script, i.e.
in a file no test reads. A second copy is a second truth.
2. **A declared-feature table** — `shater/buildtags`: every tag-gated capability
we promise, with the exact tags it needs *to run* and why (the code anchor).
`TestRouterTagSetCoversDeclaredFeatures` parses the shell file and fails if a
declared feature lost a tag. It needs no build tags, no Linux, no network and
no privileges, so it runs in every plain `go test ./...` — including on the
Windows dev host, where nothing else can see the shipped configuration.
3. **A construction test under the shipped tags** —
`shater/generate.TestShippedTagSetConstructsDeclaredProtocols` drives one node
of every declared protocol (ss/vmess/trojan/vless ws-grpc-httpupgrade-quic-
xhttp/REALITY/uTLS-fp/hysteria2/tuic/**wg**/**awg**) through `box.New`+`Start`.
`scripts/check-router-tags.sh` runs it **with `SHATER_ROUTER_TAGS`**, and CI
runs that script (`.gitea/workflows/release.yml`) *before* the artifact is
built. In a router-tag-set run nothing may be skipped: a protocol that is not
compiled in fails the run instead of quietly disappearing from it.
(2) catches a trim the moment it is made and names the feature it kills; (3)
catches what a list comparison cannot — a tag that is present but insufficient.
Neither is a substitute for the other. A new protocol in `shater/parse` +
`shater/generate` means a new row in `buildtags.Features` and a new probe case;
`TestEveryTagGatedFeatureIsProbed` fails until both exist.
- **Rejected: "just add the tag".** The one-line fix restores WireGuard and
leaves the mechanism that hid it fully intact — the next size-driven trim is
equally invisible. The tag is the smallest part of this decision.
- **Rejected: run the WHOLE test suite with the router tag set in CI.** It is the
obvious move and it does not work: parts of the suite legitimately depend on
upstream-only tags, and the run costs a second full compile of a 25 MB binary's
worth of packages on every release. A focused, unprivileged construction test
buys the same evidence for ~10 s and, unlike a full run, can be *required* to
skip nothing.
- **Rejected: assert the tag set against upstream's `DEFAULT_BUILD_TAGS`.** That
makes any trim a failure, which turns the check into noise and re-litigates D9
on every upstream rebase. The contract is with our own feature list, not with
upstream's.
- **Not done: dropping `with_lx_command`.** It is inert for `shaterd` — nothing
under `shater/` imports `sing-box/daemon` or `experimental/libbox`, and
`go list -deps ./shater/cmd/shaterd` links neither, so it costs zero bytes. It
stays only so the router set remains a subset of the lx desktop set. Noted
because "a tag that buys nothing" is the mirror image of this bug and should be
removed deliberately, not silently.
+10 -6
View File
@@ -13,8 +13,12 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination (domain/suffix/keyword/geosite), reusable domain/IP lists, port,
proto → target (outbound/selector/chain/direct/block) + egress.
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
(inline domains/CIDRs, a local or remote file, or a geosite/geoip category) that is
compiled once into a `.srs` and shared by every rule that references it. Domain
entries take `full:` (exact), `suffix:` / a leading dot (host + subdomains),
`keyword:` (substring) and `regexp:`; a bare entry means host + subdomains.
- **[MVP]** Node groups with balancer/observatory (least-ping/failover/round-robin).
- **[T1]** Multi-hop chains (L1→Ln); per-rule egress selection; egress via any
interface/tunnel (e.g. an AmneziaWG tunnel).
@@ -89,8 +93,8 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
SIM uplink → different egress); backup/restore; i18n (EN + RU).
## Ops & distribution
- **[MVP]** Single signed binary; signed opkg feed on Gitea (reuse key
`5ac4b177689cb8e0`); one-line install; `opkg upgrade`.
- **[MVP]** Single signed binary; signed apk feed on Gitea (EC key
`dist/shater-apk.pem`); one-line install; named-package `apk upgrade`.
- **[T1]** Upstream-rebase cadence (track sing-box-lx tags) with a smoke suite.
- **[T2]** apk (OpenWrt 25.x) packaging; multi-router fleet management; REST/gRPC
external API; Telegram bot.
- **[T2]** Multi-router fleet management; REST/gRPC external API; Telegram bot.
(apk packaging landed and is now the only lane — D22.)
+98 -107
View File
@@ -1,7 +1,7 @@
# Shater v0.2 — Build & Install
How to build the ship artifact (the SPA-embedded `shaterd` binary) and install
the OpenWrt feed onto a router.
the signed apk repo onto a router.
## 1. Build the `shaterd` binary
@@ -28,7 +28,7 @@ Arg / env:
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
the **same** computation the package version comes from (§2.1), so the string
the panel shows always matches what `apk info shaterd` / `opkg status` report.
the panel shows always matches what `apk list -I shaterd` reports.
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
@@ -43,22 +43,44 @@ UPX="…/scratchpad/upx-4.2.4-win64/upx.exe" scripts/build-shaterd.sh v0.2.0 --f
The `dist/*` and `openwrt/shaterd/files/shaterd-*.upx` outputs are gitignored —
they are release artifacts, not source.
Tag set (D9 — keep in sync with `docs-shater/DECISIONS.md`):
Tag set (D9/D23) — defined in **one** place, `scripts/router-tags.sh`, which
documents every tag and is sourced by the build:
```
with_quic,with_wireguard,with_utls,
with_gvisor,with_quic,with_wireguard,with_utls,
badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command
```
We drop `with_purego,with_naive_outbound`: they pull cronet-go, which forces a
glibc `PT_INTERP` even under `CGO_ENABLED=0`, making the binary unusable on musl.
We drop `with_gvisor`: the shater data plane is tproxy/redirect and generate
never emits a tun inbound, so the userspace gvisor netstack is unreachable code.
We drop `with_clash_api`: the admin panel is shater's own web server and the
generator never emits a `clash_api` service, so the Clash server is dead code.
We drop `with_dhcp`: shater resolver types are `udp/tcp/doh/dot/local/fakeip`;
a `dhcp://` DNS transport is never generated or registered.
`with_gvisor` was dropped in 2026-07 as "unreachable — we emit no tun inbound"
and **put back on 2026-07-25**: gVisor is also the netstack of the WireGuard
endpoint, so without it every `wg://`/`awg://` node died at apply time with
*"gVisor is not included in this build"* while the panel still offered the
feature. It costs ~2.8 MB raw / ~0.65 MB UPX per arch. Full story: `DECISIONS.md`
D23.
### Changing the tag set
Run the guard — it is what stands between a size trim and a silently dead
feature, and CI runs it before the artifact is built:
```sh
scripts/check-router-tags.sh # from Windows/macOS it re-execs itself in golang:1.26
```
It (1) fails if a feature declared in `FEATURES.md` lost a build tag it needs to
run (`shater/buildtags`, no tags/OS/network required) and (2) constructs one node
of every declared protocol through `box.New` **compiled with the shipped tag
set** — nothing may be skipped in that run. Adding a protocol to
`shater/parse`+`shater/generate` means adding a row to `buildtags.Features` and a
probe case in `shater/generate/shipped_tags_linux_test.go`.
## 2. Packages
Four OpenWrt packages live under `openwrt/`:
@@ -91,9 +113,9 @@ See `openwrt-package-build-ci` for SDK/feed mechanics.
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
Both package managers offer an upgrade only when the feed's version string
differs from the installed one, so `apk update` saw nothing new and the routers
could not be updated through the normal path at all.
apk offers an upgrade only when the feed's version string differs from the
installed one, so `apk update` saw nothing new and the routers could not be
updated through the normal path at all.
`ci/version.sh` now derives them from `git describe`, once per CI job:
@@ -103,9 +125,8 @@ could not be updated through the normal path at all.
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
Ordering is what makes this safe, and both managers agree on it (checked with
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
Ordering is what makes this safe (checked with `apk version -t` on apk-tools
3.0.3): the dotted part decides first, `-rN` only breaks ties — so
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
outranks every rolling build before it, rolling builds between two releases grow
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
@@ -113,8 +134,9 @@ upgrade.
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
environment; the Makefiles read it with a literal fallback for manual/offline
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
so a lost variable fails the build instead of shipping a stale version.
builds. `ci/sdk-build-apk.sh` then **asserts** the produced `.apk` really carries
it, so a lost variable fails the build instead of shipping a stale version. The
release job asserts the same version again on the published rolling repo (§5.1).
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
@@ -124,21 +146,27 @@ when our packaging of it changes.
## 3. Install on a router
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
**The normal path is the signed apk repo — §5.** This section is the manual
fallback (a router with no route to the Gitea host, or a hand-carried build).
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`).
apk filenames carry no architecture, so make sure you copied the `.apk` built for
*this* router's arch (`cat /etc/apk/arch`):
```sh
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_<ver>_all.ipk
opkg install luci-app-shater_<ver>_all.ipk
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
# --allow-untrusted: our member .apk are unsigned by design — trust lives in the
# signed packages.adb index (§5), which a loose file install does not consult.
apk add --allow-untrusted ./shaterd-<ver>.apk
apk add --allow-untrusted ./shater-core-<ver>.apk
apk add --allow-untrusted ./luci-app-shater-<ver>.apk
apk add --allow-untrusted ./byedpi-0.17.3-r1.apk # optional: ByeDPI egress
```
Installing from a signed feed instead:
From the repo instead (§5 sets it up once), deps pull the rest in:
```sh
# add the feed (customfeeds.conf / apk repositories), then:
opkg update && opkg install shater-core luci-app-shater # shaterd pulled in as a dep
apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
## 4. Enable
@@ -158,92 +186,55 @@ daemon (`shaterd run`), which owns the engine, the `inet shater` data plane, pol
routing, in-process DNS, and the admin panel (default `:8088`). The LuCI app's
"Open panel" button mints a single-use token and hands the browser off to the panel.
## 5. Add the signed feed (recommended — then `opkg upgrade` just works)
## 5. The signed apk repo (the normal install path)
CI (`.gitea/workflows/release.yml`) publishes every build as a **rolling `latest`
Gitea release** that is itself a signed opkg `src/gz` feed: the release holds the
`.ipk` for all arches, a `Packages`/`Packages.gz` index, a usign `Packages.sig`,
and the public key `shater-feed.pub`. opkg filters by `Architecture`, so the **same
two lines work on every device** (x86 testbed picks `x86_64 + all`; the BPI routers
pick `aarch64_cortex-a53 + all`).
OpenWrt/ImmortalWrt **25.12** packages with Alpine's **apk**: `.apk` files, a
binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
effectively mandatory signatures (unsigned needs `--allow-untrusted`). This is
the only format shater publishes — the `.ipk`/opkg lane was removed in 2026-07
(`DECISIONS.md` D22); every device we serve is on 25.12 with apk-tools 3.
> **Format:** OpenWrt 24.10 (our SDK) uses **opkg** (`.ipk`, `Packages.gz`, usign),
> so the feed is `src/gz` and the trust anchor is the usign key
> `dist/shater-feed.pub` (fingerprint **`5ac4b177689cb8e0`**). apk only replaces
> opkg at OpenWrt **25.12** — see §6.
One-time setup on the router:
```sh
# 1) trust the feed key — the FILENAME must equal the usign key fingerprint.
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
# 2) add the feed (one URL serves every arch).
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
# 3) refresh + install (shaterd is pulled in as a dependency).
opkg update
opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # optional: ByeDPI desync egress
```
With the key installed, opkg's default `check_signature 1` verifies the feed on
every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
(`vX.Y.Z`) publishes the identical layout at
`.../releases/download/vX.Y.Z` if you prefer to pin a version instead of tracking
`latest`.
### Updating
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
tries to upgrade *every* installed package from *every* configured feed, which on
OpenWrt means base/system packages on the overlay and is a well-known way to
brick a router.
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
```
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
when the feed's `Version` differs from the installed one — that is exactly what
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
the version from the git tag on every build (§2.1), so there is nothing to bump
by hand any more; check with:
```sh
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
```
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
OpenWrt/ImmortalWrt **25.12** replaces opkg with Alpine's **apk**: `.apk` files,
a binary `packages.adb` index, EC (prime256v1) keys in `/etc/apk/keys/`, and
effectively mandatory signatures (unsigned needs `--allow-untrusted`). The
package **Makefiles are unchanged** — the SDK release decides the format.
CI builds this lane **in parallel** with the opkg feed (same manual triggers:
`v*` tag push or `workflow_dispatch`): the `build-apk` jobs in
`.gitea/workflows/release.yml` compile the same 4 packages through the official
**ImmortalWrt 25.12 SDK** (tarballs from
CI (`v*` tag push or `workflow_dispatch`) compiles the 4 packages through the
official **ImmortalWrt 25.12 SDK** (tarballs from
`downloads.immortalwrt.org/releases/25.12.1/targets/{x86/64,mediatek/filogic}/`)
and publish **one release per arch** — rolling `apk-latest-x86_64` /
`apk-latest-aarch64_cortex-a53`, or `apk-vX.Y.Z-<arch>` for a tagged version.
Per-arch (unlike the combined opkg release) because apk filenames carry no
architecture and packages are fetched relative to the `packages.adb` URL.
and publishes **one release per arch**: the rolling `apk-latest-x86_64` /
`apk-latest-aarch64_cortex-a53`, plus `apk-vX.Y.Z-<arch>` on a tag. Per-arch
because apk filenames carry no architecture and packages are fetched *relative to
the `packages.adb` URL*, so one flat multi-arch release would collide.
> **Key:** apk cannot use the usign key. The apk trust anchor is the separate EC
> public key **`dist/shater-apk.pem`** (generated once by `ci/gen-apk-key.sh`;
> private half lives ONLY in the Gitea secret **`KEY_APK`**, the apk analog of
> `KEY_BUILD`). Never regenerate either key — that invalidates every deployed
> router's trust. The usign identity `shater-feed.pub` keeps signing the
> opkg/24.10 feed, untouched.
> **Key:** the trust anchor is the EC public key **`dist/shater-apk.pem`**
> (generated once by `ci/gen-apk-key.sh`; the private half lives ONLY in the
> Gitea secret **`KEY_APK`**). Never regenerate it — that invalidates every
> deployed router's trust.
One-time setup on a 25.12 router (BananaWRT `25.12-mtk-vendor` on the BPI-R3
mini, BPI-R4 on 25.12, or the future 25.12 VM — `/etc/apk/arch` picks the right
per-arch release automatically):
### 5.1 Rolling or pinned — pick the repo URL deliberately
The repo line names an **index file**, and which one you name is the whole
update policy:
| Repo line points at | Behaviour | Cost |
|---|---|---|
| `apk-latest-<arch>/packages.adb` (**rolling**) | Every release run REPLACES this release's assets, so `apk update && apk upgrade <our packages>` always sees the newest build. Install once, never touch the file again. | You get whatever CI published last; there is no per-router pin. |
| `apk-vX.Y.Z-<arch>/packages.adb` (**pinned**) | The router stays on exactly that build. `apk update` will never offer a newer shater. | `/etc/apk/repositories.d/shater.list` must be edited **by hand on every upgrade**, on every router. |
`mini_router` is deliberately on a **pinned** URL — a considered choice, and the
hand-edit per release is its price. Use rolling unless you specifically want to
freeze a device.
> The rolling release used to go stale silently: publishing was an either/or, so
> tag runs wrote only `apk-vX.Y.Z-<arch>` and `apk-latest-<arch>` was last
> refreshed on 2026-07-24 at `0.2.0` while v0.2.9/v0.2.10 shipped. A router on
> the rolling URL kept getting a successful `apk update` with nothing new. Fixed
> 2026-07-25: `release-apk` writes the rolling pointer on **every** run and then
> reads the release back over the Gitea API, asserting it holds our three
> tag-versioned packages at exactly the version just built and **no** leftover
> asset at another version (two versions of one package in one index would let
> apk choose instead of us).
### 5.2 One-time setup on the router
BananaWRT `25.12-mtk-vendor` on the BPI-R3 mini, OpenWrt 25.12 on the BPI-R4, or
the testbed VM — `/etc/apk/arch` picks the right per-arch release automatically:
```sh
# 1) trust the apk feed key (any *.pem filename under /etc/apk/keys works).
@@ -251,6 +242,7 @@ wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
# 2) add the repo — the line points at the packages.adb INDEX FILE itself.
# (rolling; for a pinned router put apk-vX.Y.Z-$(cat /etc/apk/arch) here — §5.1)
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
@@ -260,7 +252,7 @@ apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # optional: ByeDPI desync egress
```
### Updating
### 5.3 Updating
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
installed package against *every* configured repository at once; on a router
@@ -286,9 +278,8 @@ Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
rolling by pointing the repo line at
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
as `0.2.0-r3` and `apk update` offered nothing). Rolling vs pinned repo URL —
§5.1.
### BananaWRT `25.12-mtk-vendor` compatibility
+7 -3
View File
@@ -195,7 +195,7 @@ type Chain struct { Name string; Hops []string } // "group:<n>" | "node:<n>", L1
type Egress struct { Name,Type,Interface,Target string } // interface|proxy|direct|block
type Rule struct {
Name string; Enabled bool; Order int
Src []string; DstDomain,DstRuleset,DstIP []string; DstPort,Proto string
Src []string; DstRuleset []string; DstPort,Proto string // dst = ruleset only (v0.2 schema v2)
Target string // chain:|group:|node:|direct|block
Egress,Kill string
SchedEnabled bool; SchedDays []string; SchedStart,SchedEnd string; SchedUTCOffset int
@@ -257,8 +257,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
- `config chain`: name, list hop. `config egress`: name, type, interface, target.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url), url, path, format, update_interval, list entry.
- `config rule`: name, enabled, order, list src/dst_domain/dst_ruleset/dst_ip, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end/tz.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url|geosite|geoip), url, path, format, update_interval, list category, list entry.
- `config rule`: name, enabled, order, list src, list dst_ruleset, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end, sched_utc_offset.
v0.1 carried `dst_domain`/`dst_ip` on the rule itself; **schema v2 removed both** — a
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
+1 -1
View File
@@ -6,7 +6,7 @@ OpenWrt). Лицо репозитория и быстрый старт — в к
| Документ | О чём |
|----------|-------|
| [CONTEXT.md](CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, решения в кратце, testbed/инфра |
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка обоих фидов — opkg (24.10) и apk (25.12+) |
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка apk-фида (25.12+): роллинг или фиксация версии |
| [ARCHITECTURE.md](ARCHITECTURE.md) | One-binary дизайн, auth-handoff LuCI→панель, data/DNS/apply-потоки (диаграммы) |
| [FEATURES.md](FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [ROADMAP.md](ROADMAP.md) | Фазовый план |
+1 -1
View File
@@ -104,7 +104,7 @@ build new logic in the `shater/`, `panel/`, `openwrt/` overlay.
## Phase 8 — Ship it ✅ DONE
- Adapt CI to build/sign the single forked binary for both arches; publish the
signed opkg feed (reuse key `5ac4b177689cb8e0`); install/upgrade docs.
signed feed (apk since D22, EC key `dist/shater-apk.pem`); install/upgrade docs.
- Set an upstream-rebase cadence (merge new sing-box-lx tags, run the smoke suite).
## Cross-cutting (every phase)
+2 -2
View File
@@ -21,8 +21,8 @@ PKG_NAME:=byedpi
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# regression: our tags are 0.2.x, and the version comparator (apk-tools 3,
# verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
+29 -2
View File
@@ -62,7 +62,29 @@ config inbound
# list node 'my-node'
#
# A routing rule. target: chain:<n>|group:<n>|node:<n>|egress:<n>|direct|block.
# Match on src / dst_domain / dst_ruleset / dst_ip / dst_port / proto.
# Match on src / dst_ruleset / dst_port / proto. A rule with NO matcher at all is
# the default route for everything that reached it.
#
# WHERE the traffic is going is named ONLY by dst_ruleset — one or more
# `config ruleset` names; the rule matches when ANY of them matches. There is no
# inline domain or address list on a rule (`dst_domain`/`dst_ip` were removed in
# schema v2): a destination list is written once as a ruleset, compiled into a
# .srs and shared by every rule that references it. `shaterd migrate` converts
# older configs automatically, creating a `rule-<name>` ruleset per rule.
#config ruleset
# option name 'blocked-video'
# option type 'domain'
# option source 'inline'
# list entry 'youtube.com'
# list entry 'suffix:googlevideo.com'
#
#config rule
# option name 'video-via-main'
# option enabled '1'
# option order '50'
# list dst_ruleset 'blocked-video'
# option target 'group:main'
#
#config rule
# option name 'all-via-main'
# option enabled '1'
@@ -83,11 +105,16 @@ config inbound
# option type 'direct'
# option dpi 'fragment'
#
#config ruleset
# option name 'youtube'
# option source 'geosite'
# list category 'youtube'
#
#config rule
# option name 'youtube-fragment'
# option enabled '1'
# option order '50'
# list dst_domain 'geosite:youtube'
# list dst_ruleset 'youtube'
# option target 'egress:frag'
#
# A DNS resolver (type: doh|dot|plain|local|fakeip). `detour` routes its queries
+15 -1
View File
@@ -152,7 +152,21 @@ start_service() {
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
"$PROG" migrate >/dev/null 2>&1
#
# THE FAILURE IS LOGGED, NOT SWALLOWED. This is the only place the schema
# migration runs at boot (`shaterd run`, the SIGHUP reconcile and the panel's
# config write all read UCI directly), so if it fails here it does not get
# retried until the next start. And it CAN fail for a mundane reason — a full
# /overlay makes `uci commit` fail — after which the config still carries the
# schema-v1 `dst_domain`/`dst_ip` options. The daemon holds every rule that
# still has them DISABLED and reports it, so nothing is silently misrouted, but
# rules the operator wrote are then not in force and the reason has to be
# visible somewhere. Hence: log the binary's own stderr, and start anyway —
# refusing to start would take the admin panel down with it, and the panel is
# the only way to fix the box.
local migrate_out
migrate_out=$("$PROG" migrate 2>&1) || _slog -p daemon.err \
"UCI schema migration FAILED: ${migrate_out:-no output from $PROG migrate}. Starting anyway; routing rules that still carry the removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space on /overlay and re-run '$PROG migrate', or restart the service."
procd_open_instance shater
# shaterd runs in the FOREGROUND under procd (must never daemonize). `run` is
+5 -5
View File
@@ -38,9 +38,9 @@ PKG_NAME:=shaterd
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
# that version, so a lost env can never silently ship a stale one again.
# ci/build-feed-apk.sh exports them into the SDK build env. ci/sdk-build-apk.sh
# then ASSERTS that the produced .apk really carries that version, so a lost env
# can never silently ship a stale one again.
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
# are not "the release version"; releases are named by the tag.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
@@ -104,8 +104,8 @@ define Package/shaterd/install
$(INSTALL_BIN) $(CURDIR)/files/$(SHATERD_BIN) $(1)/usr/bin/shaterd
endef
# This package ships ONLY the binary — no init script — so opkg's default
# postinst never touches the running service. On `opkg upgrade shaterd` the new
# This package ships ONLY the binary — no init script — so the package manager's
# postinst never touches the running service. On `apk upgrade shaterd` the new
# ELF lands at /usr/bin/shaterd while the OLD image keeps running from its
# unlinked inode: the upgrade silently has no effect until the next reboot, and
# meanwhile the new CLI (`shaterd reconcile`, `status`, `mint-token` — invoked by
+26
View File
@@ -18,6 +18,32 @@ type URLTestOutboundOptions struct {
// lx: SPEC 019 v2 — load-balancing.
Mode string `json:"mode,omitempty"` // least_test (default) | round_robin
Balancer *URLTestBalancerOptions `json:"balancer,omitempty"`
// lx: health board §5.C — SelfCheck stands the group's OWN background
// health-check up or down. nil/absent == true, so every existing config keeps
// today's behaviour.
//
// Why this exists at all: a urltest group probes its members BY ITSELF — a
// warm-up sweep at PostStart and a ticker for as long as traffic keeps
// touching it — and it dials the members' outbounds DIRECTLY, from the
// router, over whatever the default WAN route is. For a group that traffic
// actually flows through, that is exactly right: the probe travels the same
// path the connections do. But for a group NO routing rule reaches, that
// same probe measures a path nothing uses — and it stores the result under
// the members' BASE tags, which every health consumer then reads as "the
// node's health". A node that is blocked on the direct WAN and perfectly
// alive behind a tunnel therefore reads "dead" the moment such a group
// probes it; the reading is not merely stale, it is FALSE, and it poisons
// the shared board for everyone (selection, the panel, the observatory's
// freshness gate). SelfCheck=false is how the control plane stands such a
// group's own schedule down: the shater engine computes which groups the
// applied rules actually reach (the observatory's used-set) and disables
// the self-check on the rest, so the ONLY prober left is the observatory —
// which probes along the real dial paths and nothing else.
//
// The flag suppresses only the group's own SCHEDULE (the PostStart warm-up
// and the Touch ticker). An EXPLICIT CheckOutbounds/URLTest call — the
// adapter interface a human or an API invokes on purpose — still works.
SelfCheck *bool `json:"self_check,omitempty"`
}
// URLTestBalancerOptions configures round_robin: a fixed-size pool of live nodes, lazily
+2 -1
View File
@@ -8,7 +8,8 @@
"dev": "vite",
"build": "tsc --noEmit && vite build",
"preview": "vite preview",
"typecheck": "tsc --noEmit"
"typecheck": "tsc --noEmit",
"test": "node --test src/*.test.ts"
},
"dependencies": {
"react": "^18.3.1",
+70
View File
@@ -405,3 +405,73 @@
color: var(--dim);
max-width: 74ch;
}
/* ---- inline rename (shared) ----
The pencil-in-the-row interaction: click the ✎ beside a name, type over it,
Enter commits / Esc cancels / blur commits. Lifted out of Devices.css when
Nodes grew the same affordance — one interaction, one set of rules, so the two
pages can never drift apart. `--locked` is the same control with the action
withheld: it stays visible and focusable-looking so a missing rename reads as
a stated rule, not a dead button. */
.inline-rename {
flex: none;
display: inline-flex;
align-items: center;
justify-content: center;
width: 22px;
height: 22px;
padding: 0;
border: 1px solid transparent;
border-radius: 5px;
background: none;
color: var(--faint);
font-size: 12px;
line-height: 1;
cursor: pointer;
transition: color 0.15s, background 0.15s, border-color 0.15s;
}
.inline-rename:hover:not(:disabled) {
color: var(--accent);
background: color-mix(in srgb, var(--accent) 12%, transparent);
}
.inline-rename:focus-visible {
color: var(--accent);
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.inline-rename:disabled {
opacity: 0.5;
cursor: default;
}
/* Withheld, not broken: keep the glyph readable and let the cursor say "there is
a reason" rather than dimming it into invisibility. */
.inline-rename--locked {
opacity: 0.75;
cursor: help;
}
.inline-rename--locked:hover {
color: var(--dim);
background: none;
}
.inline-rename-input {
min-width: 0;
max-width: 24ch;
padding: 4px 8px;
border: 1px solid var(--accent);
border-radius: 6px;
background: var(--sink);
color: var(--ink);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
box-shadow: 0 1px 2px var(--shadow) inset;
}
.inline-rename-input:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.inline-rename-input:disabled {
opacity: 0.55;
}
+182 -18
View File
@@ -51,6 +51,38 @@ export class ApiError extends Error {
*/
export type Plane = 'full' | 'hold' | 'none'
/**
* Where the router's traffic actually ENDS UP, decided by the daemon from the
* engine config it is running (apply.Status.traffic ← generate.TrafficOf).
*
* tunnel — the default route goes into a tunnel: everything not matched by a
* more specific rule is proxied.
* split — the default leaves directly, but some rules do tunnel their traffic.
* direct — the default leaves directly and nothing is tunnelled at all.
* blocked — the default is the fail-closed backstop: unmatched traffic is
* dropped, not let out. Nothing leaks.
*
* `plane` DOES NOT ANSWER THIS and must never be read as if it did. `plane` says
* how much of the data plane is installed (nft table, policy routing, engine up);
* a router whose only rule is `default → direct` has all of it and sends the whole
* LAN out the plain WAN with its real address. That combination — plane "full",
* traffic "direct" — was live on a user's router under a green "Protected" LED.
*/
export type TrafficVerdict = 'tunnel' | 'split' | 'direct' | 'blocked'
export interface Traffic {
// '' or absent ⇒ not known (daemon that predates this field, nothing applied
// yet, or the plane is on hold). NEVER treat unknown as 'tunnel'.
verdict?: TrafficVerdict | ''
// The outbound tag the engine's default route names, in the engine's own
// vocabulary ("direct", "block", a node/group tag). Diagnostic — wording is
// driven by `verdict`, never by parsing this.
default?: string
// How many of the engine's route rules send their matched traffic into a tunnel.
// Separates "some of your traffic is protected" from "none of it is".
tunnel_rules?: number
}
/**
* One thing the last apply could not do. Deliberately fail-OPEN with a warning
* rather than refusing the whole config (the alternative was taking the network
@@ -88,6 +120,10 @@ export interface Status {
// How much of the data plane is installed. Absent on older daemons ⇒ unknown,
// in which case the UI shows nothing rather than guessing "full".
plane?: Plane
// Where the traffic actually goes under the running config. Absent on older
// daemons ⇒ unknown; see TrafficVerdict for why this is a separate question
// from `plane`.
traffic?: Traffic
// Findings from the last apply. ALWAYS an array from the daemon (never null);
// empty means the last apply was clean. Pre-sorted critical-first and capped at
// 50, where a truncated list ends with an `info` entry saying "suppressed".
@@ -354,19 +390,92 @@ export interface GroupHealth {
* any more and nothing to report here beyond the groups themselves.
*/
/**
* One hop of one chain, measured where that hop actually sits in the path.
*
* This is the reading the daemon always took and never showed. A chain is not a
* target with a single health — it is an ordered series of them, and the only
* question an operator ever asks about a broken chain is WHICH hop broke. The
* end-to-end exit reading cannot answer that: it says "the path is dead" for a
* four-hop chain and leaves the person to guess between four suspects.
*
* WIRE ORDER. `index` is 1-based and counts hops in the order the router dials
* them: hop 1 is the first physical hop, and each later hop is dialled THROUGH
* the ones before it. The hop carrying `exit: true` — always the largest index —
* is where traffic leaves for the internet. A leading `egress:` in the chain's
* configured Hops is NOT a numbered hop: the daemon lifts it into the entry
* detour of hop 1, so a chain written `egress:ewan → node:awgout → group:sub0`
* reports two hops, not three. Anything zipping this against the model's Hops
* must drop that leading egress first and give up on labelling entirely if the
* counts still disagree — a chain that splices sub-chains gets flattened here,
* and a confidently WRONG hop name is worse than no name.
*
* `tag` is the engine-side outbound (`chain-<name>-h2`). Debugging and tooltips
* only; it is never a label to put in front of a person.
*
* NODE HOP vs GROUP HOP. For `kind: "node"` the hop IS the measurement: `total`
* is 1, the counters follow its own state, and `selected` is ''. For
* `kind: "group"` the counters roll up that hop's per-hop member COPIES — the
* copies dialled through the hops in front of it, which is exactly why they can
* read alive here while the same group's standalone card reads dead. Both
* readings are true; they measure different dial paths. `selected` is the node
* NAME the hop routes through right now, and `delay_ms` / `age_seconds` belong
* to that selected member (or the freshest alive one).
*
* Invariants the daemon guarantees — never re-derive them, just read them:
* `tested === alive + dead` and `alive + dead + untested === total`.
*
* `state` is a closed set. `untested` is NEVER "dead" and never "healthy": it
* means nothing fresh enough is known, which for a used chain is seconds away
* from resolving on its own. `age_seconds: -1` means the age is unknown.
*/
export interface ChainHopHealth {
/** 1-based WIRE order. Hop 1 is dialled first; see the note above. */
index: number
/** Engine outbound tag (`chain-<name>-h2`) — tooltips/debugging, never a label. */
tag: string
/** `node` ⇒ the hop is the measurement. `group` ⇒ the counters roll up members. */
kind: 'node' | 'group'
/** This hop is where traffic leaves for the internet. Always the largest index. */
exit: boolean
/** Closed set — switch on it exhaustively. `untested` is never "dead". */
state: 'alive' | 'dead' | 'untested'
/** RTT of the selected/freshest alive member; 0 (meaningless) when not alive. */
delay_ms: number
/** Age of that measurement in seconds; -1 when unknown. */
age_seconds: number
/** Node name this GROUP hop routes through right now; '' for a node hop. */
selected: string
total: number
tested: number
alive: number
dead: number
untested: number
}
/** Per-chain reachability, the chain analogue of {@link GroupHealth}.used (plan
* §5.E): a chain no enabled routing rule routes through is outside the
* observatory's plan, so its exit is never probed and the Targets card renders it
* "unused" instead of an exit-test readout. A chain has no membership counters —
* it is a fixed path, and its end-to-end health is the exit test's job. */
* observatory's plan, so nothing probes it and the Targets card says so instead
* of rendering a health reading. A chain has no membership counters of its own —
* it is a fixed path, and its health lives on its {@link ChainHopHealth} hops. */
export interface ChainHealth {
name: string
/** An enabled routing rule (the Final target, a DNS-resolver detour, a device
* target, …) reaches this chain, so the observatory probes its exit in the
* target, …) reaches this chain, so the observatory probes its hops in the
* background. false ⇒ nothing routes through the chain: it is skipped by the
* background probing and its end-to-end health stays untested. That is an
* "unused" note about the ROUTING CONFIG, never a health problem. */
* background probing and its health stays untested. That is an "unused" note
* about the ROUTING CONFIG, never a health problem. */
used: boolean
/**
* Per-hop health in wire order (see {@link ChainHopHealth}).
*
* MAY BE ABSENT, and absent does not mean "this chain has no hops". It means
* the engine never materialised per-hop outbounds for it: the chain is unused,
* or it collapses to a single hop and the daemon points traffic straight at
* that target instead of building a copy of it. Read a missing key as "nothing
* measured per hop", never as an empty path or as a fault.
*/
hops?: ChainHopHealth[]
}
export interface GroupsHealth {
@@ -809,9 +918,17 @@ export interface Rule {
Enabled: boolean
Order: number
Src?: string[] | null
DstDomain?: string[] | null
/**
* WHERE the traffic is going — the rule's only destination matcher. Each entry
* names a {@link Ruleset}; the rule matches when ANY of them matches.
*
* There is no inline domain or address list on a rule. `dst_domain`/`dst_ip`
* were removed in schema v2, and `shaterd migrate` folds every existing one
* into a generated `rule-<name>` ruleset, so a destination list is written and
* edited in exactly one place and compiled once into a .srs that every rule
* referencing it shares.
*/
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
/**
* Narrow the rule to one transport or one sniffed application protocol. A
@@ -1218,6 +1335,17 @@ export interface RuleReach {
shadowed_by_order?: number
/** Operator-facing sentence; absent when `unreachable` is false. */
reason?: string
/**
* Whether the rule is IN FORCE right now — `Rule.Enabled` after the active WAN
* profile's overrides. This is NOT `GET /api/config`'s `Enabled`: that one is
* the desired state the page PUTs back, and on a router with profiles the two
* legitimately disagree. Draw rows from this; keep the switch on the other.
*/
effective_enabled: boolean
/** The active profile that CHANGED this rule's state; absent when none did. */
overridden_by?: string
/** Which way it went. Absent together with `overridden_by`. */
override?: 'enabled' | 'disabled'
}
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
@@ -1377,8 +1505,21 @@ export function getGroupsHealth(
}
/**
* One group's (or chain's) last test: which member the balancer picked, how fast
* it answered, and what the internet saw as the source address.
* What the OBSERVATORY measured for one group or chain — not a dial the panel
* made.
*
* This shape used to come from a fresh connection opened on demand, straight at
* the target. That was a lie on any router whose proxies are blocked when dialled
* directly and work only as a hop behind a tunnel: the card reported dead for a
* path that carries traffic all day. The daemon now has exactly one thing that
* measures — the background observatory, which probes along the REAL dial path,
* per-hop copies and all — and this endpoint reports what it found. There is no
* second measurement anywhere, and the panel never opens a connection of its own.
*
* So read the fields as a READ, not as a test run: `ok` and `delay_ms` are the
* observatory's verdict for the path traffic actually takes, and `tested_unix`
* (router clock, seconds) is when the OBSERVATORY took that measurement — which
* can be a few seconds before the refresh was asked for.
*
* `ok:true` with an EMPTY `exit_ip`/`exit_country` is a valid, successful result,
* not a partial failure: the delay was measured but the exit address could not be
@@ -1387,10 +1528,23 @@ export function getGroupsHealth(
*
* Chains ride the same endpoint. For a chain row, `group` carries the CHAIN's
* name and `selected` the node its last group hop picked ('' when the exit hop
* isn't a group). Everything else reads the same way.
* isn't a group). Per-hop detail is a different read: {@link ChainHopHealth}.
*
* `ok:false` ⇒ the test failed and `error` carries the human reason; every other
* field is meaningless. `tested_unix` is the router's clock, in seconds.
* `ok:false` ⇒ there is no usable measurement and `error` carries the human
* reason; every other field is meaningless. Four of those reasons are about the
* observatory rather than the path, and must not be rendered as "your target is
* broken":
*
* "not routed by any enabled rule, so nothing measures it — the observatory
* only probes paths the rules use"
* "the observatory has not reached this target yet — it refreshes on the
* global probe interval"
* "background probing is disabled, so there is nothing to measure this target
* with"
* "the observatory's probe through this path failed"
*
* Only the last one is a health finding. The first three say the measurement
* does not exist, which is a different thing and a different fix.
*/
export interface GroupTestResult {
group: string // group name — or a chain name for a chain row
@@ -1407,7 +1561,11 @@ export interface GroupTestResult {
* GET /api/groups/test — progress plus every result so far. `results` is ALWAYS
* an array (never null); `done`/`total` count finished vs targeted groups and
* chains while `running` is true. Idle reads `{running:false}` with the last
* run's results still attached, so a reload after a test still shows what it found.
* run's results still attached, so a reload still shows what was last read.
*
* "Running" means the observatory is working through an out-of-turn refresh pass
* over the named targets and this endpoint is collecting what it measures. It is
* not the panel dialling anything.
*/
export interface GroupTestStatus {
running: boolean
@@ -1440,10 +1598,16 @@ export interface GroupTestStart {
}
/**
* POST /api/groups/test — measure a target's delay and exit address. Pass a
* group or chain name to test one; pass nothing (or '') to test every group
* and every chain. Singleton: a second call while a run is in flight resolves
* to `{started:false, reason:'already running'}` rather than failing.
* POST /api/groups/test — ask the observatory for an out-of-turn refresh pass,
* then report what it measured. Pass a group or chain name to refresh one; pass
* nothing (or '') for every group and every chain.
*
* It does NOT dial. The observatory is the only thing in the daemon that
* measures anything, and it measures along the real dial path — so this is the
* "don't wait for the next probe interval" button, not a second opinion. The
* numbers it returns are the same numbers the cards are already showing, just
* fresher. Singleton: a second call while a pass is in flight resolves to
* `{started:false, reason:'already running'}` rather than failing.
*/
export function postGroupsTest(name = ''): Promise<GroupTestStart> {
return MOCK
+15 -1
View File
@@ -1,4 +1,5 @@
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange. */
/* Buttons — mono, uppercase. .btn is ghost; .btn.primary is solid orange;
* .btn.crit is the solid-red destructive commit. */
.btn {
display: inline-block;
padding: 7px 12px;
@@ -30,3 +31,16 @@
color: #fff;
filter: brightness(1.05);
}
/* Destructive commit. The fill is crit stepped a little toward black so white
* label text clears 4.5:1 in BOTH themes — the raw --crit is bright enough in
* dark mode to fall under it. Red here always means "this removes something". */
.btn.crit {
border-color: transparent;
background: color-mix(in srgb, var(--crit) 88%, #000);
color: #fff;
}
.btn.crit:hover {
color: #fff;
filter: brightness(1.08);
}
+15 -7
View File
@@ -1,19 +1,27 @@
import './Button.css'
import { forwardRef } from 'react'
import type { ButtonHTMLAttributes } from 'react'
export interface ButtonProps extends ButtonHTMLAttributes<HTMLButtonElement> {
/** `primary` is the solid-orange call to action; `ghost` is the default. */
variant?: 'ghost' | 'primary'
/**
* `primary` is the solid-orange call to action; `crit` is the solid-red
* destructive commit (delete, remove) — semantic crit, never the accent;
* `ghost` is the default.
*/
variant?: 'ghost' | 'primary' | 'crit'
}
export function Button({ variant = 'ghost', className, type, ...rest }: ButtonProps) {
/** Ref-forwarding so a dialog can park focus on a specific button. */
export const Button = forwardRef<HTMLButtonElement, ButtonProps>(function Button(
{ variant = 'ghost', className, type, ...rest },
ref,
) {
return (
<button
ref={ref}
type={type ?? 'button'}
className={['btn', variant === 'primary' ? 'primary' : '', className]
.filter(Boolean)
.join(' ')}
className={['btn', variant === 'ghost' ? '' : variant, className].filter(Boolean).join(' ')}
{...rest}
/>
)
}
})
+162
View File
@@ -0,0 +1,162 @@
/* <ConfirmDialog> — the safety interlock plate.
*
* This replaces the browser's native confirm dialog, which a browser can mute for
* good ("prevent this page from creating additional dialogs"): after that it
* returns false with no dialog at all, so every delete button in the panel goes
* dead and silent with no way to recover short of a page reload. We draw the
* plate ourselves, so nothing can suppress it.
*
* Faceplate language: a small rack module lifted off the panel — corner screws
* (reused from Faceplate.css), an engraved label, a groove above the actions.
* Destructive intent is carried by the crit semantic, never by the orange accent:
* accent means "this control is active", crit means "this destroys something".
*/
/* The veil is a fixed dark wash in both themes — a light scrim over a light
* panel would not read as "the panel is out of reach". Follows the tokens.css
* pattern: light base, dark via media query, data-theme overrides win both ways. */
.cfm-scrim {
--cfm-veil: rgba(33, 29, 21, 0.52);
}
@media (prefers-color-scheme: dark) {
.cfm-scrim {
--cfm-veil: rgba(0, 0, 0, 0.66);
}
}
:root[data-theme='light'] .cfm-scrim {
--cfm-veil: rgba(33, 29, 21, 0.52);
}
:root[data-theme='dark'] .cfm-scrim {
--cfm-veil: rgba(0, 0, 0, 0.66);
}
.cfm-scrim {
position: fixed;
inset: 0;
z-index: 200;
display: flex;
align-items: center;
justify-content: center;
/* Short viewports: the plate scrolls with the veil instead of being clipped. */
overflow-y: auto;
padding: calc(var(--u, 8px) * 2);
background: var(--cfm-veil);
animation: cfm-veil-in 0.14s ease-out;
}
.cfm-card {
position: relative;
width: min(32rem, 100%);
max-height: calc(100dvh - var(--u, 8px) * 4);
overflow-y: auto;
padding: calc(var(--u, 8px) * 3.25);
border: 1px solid var(--groove);
border-radius: 12px;
/* same brushed plate as <Faceplate>, one step brighter so it reads as lifted */
background:
repeating-linear-gradient(
90deg,
transparent 0 2px,
color-mix(in srgb, var(--edge) 30%, transparent) 2px 3px
),
linear-gradient(180deg, var(--raised), color-mix(in srgb, var(--raised) 82%, var(--panel)));
box-shadow:
0 1px 0 var(--edge) inset,
0 30px 60px -22px var(--shadow),
0 4px 12px var(--shadow);
animation: cfm-card-in 0.18s cubic-bezier(0.2, 0.7, 0.3, 1);
}
.cfm-card:focus {
outline: none;
}
/* `still` is set from usePrefersReducedMotion — the plate appears, it never
* travels. (The global reduced-motion rule in tokens.css also neutralises the
* duration; this keeps the intent explicit at the component.) */
.cfm-scrim.still,
.cfm-scrim.still .cfm-card {
animation: none;
}
@keyframes cfm-veil-in {
from {
opacity: 0;
}
to {
opacity: 1;
}
}
@keyframes cfm-card-in {
from {
opacity: 0;
transform: translateY(6px) scale(0.99);
}
to {
opacity: 1;
transform: none;
}
}
/* ---- header: engraved label + state LED ---- */
.cfm-hd {
display: flex;
align-items: center;
gap: 10px;
margin-bottom: calc(var(--u, 8px) * 1.5);
}
.cfm-label {
flex: 1;
font-family: var(--font-mono);
font-size: 10px;
letter-spacing: var(--track-label-wide, 0.24em);
color: var(--dim);
text-transform: uppercase;
}
/* ---- copy ---- */
.cfm-title {
margin: 0;
font-family: var(--font-mono);
font-weight: 700;
font-size: 17px;
line-height: 1.35;
color: var(--ink);
/* names can be long and unbroken — wrap rather than push the plate wide */
overflow-wrap: anywhere;
}
.cfm-body {
margin: calc(var(--u, 8px) * 1.5) 0 0;
max-width: 52ch;
font-family: var(--font-sans);
font-size: 13.5px;
line-height: 1.6;
color: var(--dim);
overflow-wrap: anywhere;
}
/* ---- action bar ---- */
.cfm-actions {
display: flex;
justify-content: flex-end;
gap: calc(var(--u, 8px));
margin-top: calc(var(--u, 8px) * 3);
padding-top: calc(var(--u, 8px) * 2);
border-top: 1px solid var(--groove);
}
@media (max-width: 420px) {
.cfm-card {
padding: calc(var(--u, 8px) * 2.5);
}
.cfm-actions {
flex-wrap: wrap;
}
.cfm-actions .btn {
flex: 1 1 auto;
text-align: center;
}
/* screws crowd a small plate — drop them rather than collide with the copy */
.cfm-card > .screw {
display: none;
}
}
+281
View File
@@ -0,0 +1,281 @@
import './ConfirmDialog.css'
import {
createContext,
useCallback,
useContext,
useEffect,
useId,
useRef,
useState,
} from 'react'
import type { ReactNode } from 'react'
import { createPortal } from 'react-dom'
import { Button } from './Button'
import { Led } from './Led'
import { usePrefersReducedMotion } from './usePrefersReducedMotion'
/**
* How the confirming button is painted.
*
* crit — the action destroys something. Semantic crit, never the accent.
* neutral — the action is a normal commit the operator should read first
* (a warning before saving); the accent's call-to-action is correct.
*/
export type ConfirmTone = 'crit' | 'neutral'
export interface ConfirmOptions {
/** Engraved eyebrow, e.g. "DELETE RULE". Names the operation, not the object. */
label?: string
/** The question. One line, ends in "?". */
title: string
/** The consequence — what changes on the router if this goes through. */
body?: ReactNode
/** Verb on the confirming button. Defaults to "Delete". */
confirmLabel?: string
/** Verb on the dismissing button. Defaults to "Cancel". */
cancelLabel?: string
/** Defaults to `crit` — the overwhelmingly common case is a delete. */
tone?: ConfirmTone
}
export interface ConfirmDialogProps extends ConfirmOptions {
open: boolean
/** Called exactly once per dialog, with the operator's answer. */
onResolve: (confirmed: boolean) => void
}
const FOCUSABLE =
'button:not([disabled]), [href], input:not([disabled]), select:not([disabled]), textarea:not([disabled]), [tabindex]:not([tabindex="-1"])'
/**
* The modal plate itself. Normally reached through `useConfirm()`; exported so a
* page that wants to own the open state can render it directly.
*
* Keyboard contract:
* - focus moves to Cancel on open, so a reflex Enter dismisses, never deletes;
* - Tab / Shift+Tab cycle inside the plate and cannot reach the page behind it;
* - Esc answers "no";
* - on close, focus returns to whatever opened the dialog.
*/
export function ConfirmDialog({
open,
onResolve,
label,
title,
body,
confirmLabel = 'Delete',
cancelLabel = 'Cancel',
tone = 'crit',
}: ConfirmDialogProps) {
const titleId = useId()
const bodyId = useId()
const cardRef = useRef<HTMLDivElement>(null)
const cancelRef = useRef<HTMLButtonElement>(null)
const openerRef = useRef<HTMLElement | null>(null)
const reduced = usePrefersReducedMotion()
// Take the page out of the tab order, park focus on Cancel, and hand focus
// back to the opener when the plate goes away.
useEffect(() => {
if (!open) return
const opener = document.activeElement
openerRef.current = opener instanceof HTMLElement ? opener : null
const prevOverflow = document.body.style.overflow
document.body.style.overflow = 'hidden'
// Cancel is the resting place: an Enter or a Space meant for the page lands
// on "no". The destructive button is one Tab away, deliberately.
;(cancelRef.current ?? cardRef.current)?.focus()
return () => {
document.body.style.overflow = prevOverflow
const back = openerRef.current
openerRef.current = null
if (back && document.contains(back)) back.focus()
}
}, [open])
// Esc answers no; Tab is caged. Capture phase so a page-level key handler
// never sees keys aimed at the dialog.
useEffect(() => {
if (!open) return
const onKey = (e: KeyboardEvent) => {
if (e.key === 'Escape') {
e.preventDefault()
e.stopPropagation()
onResolve(false)
return
}
if (e.key !== 'Tab') return
const card = cardRef.current
if (!card) return
const list = Array.from(card.querySelectorAll<HTMLElement>(FOCUSABLE))
if (list.length === 0) {
e.preventDefault()
card.focus()
return
}
const first = list[0]
const last = list[list.length - 1]
const active = document.activeElement as HTMLElement | null
if (!active || !card.contains(active)) {
e.preventDefault()
;(e.shiftKey ? last : first).focus()
} else if (e.shiftKey && active === first) {
e.preventDefault()
last.focus()
} else if (!e.shiftKey && active === last) {
e.preventDefault()
first.focus()
}
}
document.addEventListener('keydown', onKey, true)
return () => document.removeEventListener('keydown', onKey, true)
}, [open, onResolve])
if (!open) return null
return createPortal(
<div
className={['cfm-scrim', reduced ? 'still' : ''].filter(Boolean).join(' ')}
// A click on the field around the plate means "not now". Mousedown (not
// click) so a text selection dragged out of the plate can't dismiss it.
onMouseDown={(e) => {
if (e.target === e.currentTarget) onResolve(false)
}}
>
<div
className={`cfm-card tone-${tone}`}
ref={cardRef}
tabIndex={-1}
role="alertdialog"
aria-modal="true"
aria-labelledby={titleId}
aria-describedby={body != null ? bodyId : undefined}
>
<i className="screw tl" aria-hidden="true" />
<i className="screw tr" aria-hidden="true" />
<i className="screw bl" aria-hidden="true" />
<i className="screw br" aria-hidden="true" />
{/* Lamp first, then the engraved label — the way a real panel reads, and
it keeps the LED off the corner screw. */}
<div className="cfm-hd">
<Led variant={tone === 'crit' ? 'crit' : 'amber'} />
<span className="cfm-label">{label ?? (tone === 'crit' ? 'Confirm delete' : 'Confirm')}</span>
</div>
<h2 className="cfm-title" id={titleId}>
{title}
</h2>
{body != null && (
<p className="cfm-body" id={bodyId}>
{body}
</p>
)}
<div className="cfm-actions">
<Button ref={cancelRef} onClick={() => onResolve(false)}>
{cancelLabel}
</Button>
<Button variant={tone === 'crit' ? 'crit' : 'primary'} onClick={() => onResolve(true)}>
{confirmLabel}
</Button>
</div>
</div>
</div>,
document.body,
)
}
// ---- provider + hook --------------------------------------------------------
interface Request extends ConfirmOptions {
id: number
resolve: (v: boolean) => void
}
const ConfirmCtx = createContext<((o: ConfirmOptions) => Promise<boolean>) | null>(null)
/**
* Mount once at the app root. Everything below can then ask a question and await
* the answer.
*/
export function ConfirmProvider({ children }: { children: ReactNode }) {
const [req, setReq] = useState<Request | null>(null)
const pending = useRef<Request | null>(null)
const seq = useRef(0)
const confirm = useCallback(
(opts: ConfirmOptions) =>
new Promise<boolean>((resolve) => {
// A second question while one is open answers the first with "no" rather
// than leaving its promise — and its caller — hanging forever.
pending.current?.resolve(false)
seq.current += 1
const next: Request = { ...opts, id: seq.current, resolve }
pending.current = next
setReq(next)
}),
[],
)
const settle = useCallback((confirmed: boolean) => {
const open = pending.current
pending.current = null
setReq(null)
open?.resolve(confirmed)
}, [])
// Teardown must not strand a caller mid-await.
useEffect(
() => () => {
pending.current?.resolve(false)
pending.current = null
},
[],
)
// A question belongs to the page that asked it. The provider outlives the
// hash router, so a navigation would otherwise leave a stale plate floating
// over a page it has nothing to do with — answer it "no" and clear it.
useEffect(() => {
const onNav = () => {
if (pending.current) settle(false)
}
window.addEventListener('hashchange', onNav)
return () => window.removeEventListener('hashchange', onNav)
}, [settle])
return (
<ConfirmCtx.Provider value={confirm}>
{children}
{req !== null && <ConfirmDialog key={req.id} open onResolve={settle} {...req} />}
</ConfirmCtx.Provider>
)
}
/**
* Ask the operator, get a definite answer:
*
* const confirm = useConfirm()
* if (!(await confirm({ title: 'Delete rule "x"?', body: '…' }))) return
*
* The returned function is stable, so it is safe in a useCallback dep list. It
* always settles — cancel, Esc, click-outside and teardown all resolve `false`;
* only the confirming button resolves `true`.
*
* Name it `confirm` at the call site on purpose: the local binding shadows the
* global one inside that component, so an accidental bare `confirm(...)` cannot
* reach the suppressible native dialog.
*/
export function useConfirm(): (o: ConfirmOptions) => Promise<boolean> {
const ctx = useContext(ConfirmCtx)
if (!ctx) {
// Loud on purpose. A fallback that quietly resolved false would rebuild the
// exact bug this component exists to kill.
throw new Error('useConfirm() needs <ConfirmProvider> above it (mounted in main.tsx)')
}
return ctx
}
+2
View File
@@ -17,6 +17,8 @@ export { Button } from './Button'
export type { ButtonProps } from './Button'
export { Select } from './Select'
export type { SelectProps, SelectOption } from './Select'
export { ConfirmDialog, ConfirmProvider, useConfirm } from './ConfirmDialog'
export type { ConfirmDialogProps, ConfirmOptions, ConfirmTone } from './ConfirmDialog'
export { Clock } from './Clock'
export { CatSuggest } from './CatSuggest'
export { SrcPicker } from './SrcPicker'
+6 -1
View File
@@ -2,12 +2,17 @@ import { StrictMode } from 'react'
import { createRoot } from 'react-dom/client'
import './tokens.css'
import { App } from './App'
import { ConfirmProvider } from './components'
const rootEl = document.getElementById('root')
if (!rootEl) throw new Error('#root not found')
// ConfirmProvider sits ABOVE <App> so it survives App's early returns (the
// unauth / no-link plates) — useConfirm() can never find itself without a host.
createRoot(rootEl).render(
<StrictMode>
<App />
<ConfirmProvider>
<App />
</ConfirmProvider>
</StrictMode>,
)
+200 -35
View File
@@ -6,7 +6,7 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
import type { ApplyResult, ChainHealth, ChainHopHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, Profile, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning, Traffic } from './api'
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -129,9 +129,25 @@ const CONFIG: Model = {
{ Name: 'via-tunnel', Source: 'subscription', Subscription: 'primary', Strategy: 'leastping', Egress: 'awg' },
{ Name: 'fallback', Source: 'subscription', Subscription: 'backup', Strategy: 'roundrobin', Egress: '' },
],
// One multi-hop chain so `?mock` exercises the chain card's Test button and
// its result readout: enters through the awg tunnel, exits via the auto group.
Chains: [{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] }],
// Three chains, one per state the hop readout has to render.
Chains: [
// The owner's real production shape: leave through a WAN interface, cross an
// AmneziaWG node, then three subscription groups in series. The leading
// `egress:` is NOT a numbered hop — the daemon lifts it into hop 1's entry
// detour — so this reports FOUR hops, and hop 3 is dead while its neighbours
// answer. That single red notch in the middle of a live path is the entire
// reason per-hop health exists, so `?mock` must show it at a glance.
{
Name: 'ewan-wg-subs',
Hops: ['egress:wan', 'node:home-wg', 'group:auto', 'group:stealth', 'group:via-tunnel'],
},
// Used, but the observatory hasn't come round yet — every hop untested. Not
// dead and not healthy: the state the panel most easily renders as a fault.
{ Name: 'sub-fresh', Hops: ['node:home-wg', 'group:fallback'] },
// No enabled rule targets it, so the observatory skips it entirely and the
// daemon never materialises its hops: `used:false` and NO `hops` key.
{ Name: 'relay', Hops: ['egress:awg', 'group:auto'] },
],
Egresses: [
{ Name: 'wan', Type: 'interface', Interface: 'wan' },
// An AmneziaWG tunnel — the whole point of a group-level egress binding.
@@ -145,6 +161,13 @@ const CONFIG: Model = {
Rules: [
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
// These two are what make the chains USED — the observatory probes only the
// paths an enabled rule can reach, so without them every chain card would
// read "not routed" and the hop rail would never appear in `?mock`. Kept
// ABOVE the condition-less rule at Order 40, which would otherwise swallow
// everything below it and mark them "never applies".
{ Name: 'media-via-chain', Enabled: true, Order: 22, DstRuleset: ['yt-geosite'], Target: 'chain:ewan-wg-subs' },
{ Name: 'spare-via-chain', Enabled: true, Order: 24, DstRuleset: ['ad-hosts'], Target: 'chain:sub-fresh' },
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
// A SECOND condition-less rule, above the real default. It reads like a working
// rule and does nothing: a rule with no conditions becomes the router's default,
@@ -305,29 +328,75 @@ const RULESET_STATUS: RulesetStatus[] = [
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
* a rule with no conditions is the router's default, and the LAST such rule by
* Order wins — every earlier one can never apply. It reads the live CONFIG so
* edits made in `?mock` keep the badge honest. */
* edits made in `?mock` keep the badge honest.
*
* It also mirrors model.ResolveActiveProfile + ApplyProfileRuleOverrides, because
* `effective_enabled` is the whole point of the endpoint: CONFIG's `mobile-uplink`
* is active and both enables and disables rules, so `?mock` shows the same
* desired-vs-effective split the field config does. */
export async function getRulesReachability(): Promise<RulesReachability> {
await wait(60)
const rules = CONFIG.Rules ?? []
// A pin naming an existing, ENABLED profile wins outright. Otherwise auto-select:
// highest Priority among enabled profiles, ties by Name, skipping any with an
// iface condition (the WAN watcher owns those and expresses its verdict as the pin).
const profiles = CONFIG.Profiles ?? []
const pinned = String(CONFIG.Globals?.ActiveProfile ?? '').trim()
let prof: Profile | null = profiles.find((p) => p.Enabled && p.Name === pinned) ?? null
if (!prof) {
for (const p of profiles) {
if (!p.Enabled || (p.MatchIface ?? []).length > 0) continue
const pp = p.Priority ?? 0
const bp = prof?.Priority ?? 0
if (!prof || pp > bp || (pp === bp && p.Name < prof.Name)) prof = p
}
}
// Enable first, then Disable, so a name in both ends up disabled (Disable wins).
const effective = rules.map((r) => Boolean(r.Enabled))
if (prof) {
const force = (names: string[] | null | undefined, on: boolean) => {
for (const raw of names ?? []) {
const n = raw.trim()
rules.forEach((r, i) => {
if (r.Name === n) effective[i] = on
})
}
}
force(prof.EnableRules, true)
force(prof.DisableRules, false)
}
const activeProfile = prof
const out: RuleReach[] = rules.map((r, index) => ({
index,
name: String(r.Name ?? ''),
order: Number(r.Order ?? 0),
unreachable: false,
shadowed_by_index: -1,
effective_enabled: effective[index],
// Annotate only where the profile actually FLIPPED the outcome — a profile that
// disables an already-off rule has overridden nothing the operator can see.
...(activeProfile && effective[index] !== Boolean(r.Enabled)
? {
overridden_by: activeProfile.Name,
override: effective[index] ? ('enabled' as const) : ('disabled' as const),
}
: {}),
}))
const conditionless = (r: (typeof rules)[number]): boolean =>
!(r.Src ?? []).length &&
!(r.DstDomain ?? []).length &&
!(r.DstRuleset ?? []).length &&
!(r.DstIP ?? []).length &&
!String(r.DstPort ?? '').trim() &&
!String(r.Proto ?? '').trim()
const target = (r: (typeof rules)[number]): string =>
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
const defaults = rules
.map((r, index) => ({ r, index }))
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
// The EFFECTIVE flag, not the configured one: a rule the active profile
// switched off is not in force and cannot retire anything (model's
// RuleReachability runs over the effective set for the same reason).
.filter(({ r, index }) => effective[index] && conditionless(r) && target(r))
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
const winner = defaults[defaults.length - 1]
if (winner) {
@@ -407,6 +476,11 @@ export async function getRulesetCategories(source: string): Promise<RulesetCateg
// ?mock&warn=1 → a full warning set (critical + warning + info) on top
// ?mock&ks=open → healthy plane but a FAIL-OPEN kill-switch, which is what
// makes the untunnelable policy inert (F8 case 4)
// ?mock&traffic=… → with the plane FULL, where the traffic actually ends up:
// split | direct | blocked | blackout | unknown. `direct` is
// the field case the readout used to call "Protected" (one
// rule, `default → direct`); `unknown` is a daemon too old to
// report. Default: tunnel.
function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSwitch: string } {
const q = typeof location === 'undefined' ? '' : location.search
const params = new URLSearchParams(q)
@@ -418,6 +492,30 @@ function mockPlane(): { plane: 'full' | 'hold' | 'none'; engine: boolean; killSw
return { plane: 'full', engine: true, killSwitch }
}
// The daemon's verdict on where traffic goes (apply.Status.traffic). Only
// meaningful with the plane installed: with the engine down there is no running
// config to judge, and the daemon reports the unknown/zero value — so do the same
// here rather than leaving a stale "tunnel" behind a dead engine.
function mockTraffic(plane: 'full' | 'hold' | 'none'): Traffic | undefined {
if (plane !== 'full') return { verdict: '', default: '', tunnel_rules: 0 }
const params = new URLSearchParams(typeof location === 'undefined' ? '' : location.search)
switch (params.get('traffic')) {
case 'split':
return { verdict: 'split', default: 'direct', tunnel_rules: 3 }
case 'direct':
return { verdict: 'direct', default: 'direct', tunnel_rules: 0 }
case 'blocked':
return { verdict: 'blocked', default: 'block', tunnel_rules: 2 }
case 'blackout':
return { verdict: 'blocked', default: 'block', tunnel_rules: 0 }
case 'unknown':
// A daemon that predates the field sends no `traffic` at all.
return undefined
default:
return { verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }
}
}
const MOCK_WARNINGS: StatusWarning[] = [
{
severity: 'critical',
@@ -520,6 +618,7 @@ export async function getStatus(): Promise<Status> {
can_rollback: armed || hasLastGood,
engine_running: engine,
plane,
traffic: mockTraffic(plane),
warnings: mockWarnings(killSwitch),
// Process uptime. Anchored to when this tab loaded plus a fixed head start, so
// the reading ticks forward across polls exactly like the real daemon's does.
@@ -1096,14 +1195,47 @@ function healthList(): GroupHealth[] {
return (CONFIG.Groups ?? []).map((g) => summarise(g.Name, GROUP_MEMBERS.get(g.Name) ?? []))
}
/** Per-chain reachability for the Targets page's "unused" badge (plan §5.E) — the
* chain analogue of healthList's `used` field. The mock's single chain `relay` is
* NOT referenced by any rule in CONFIG.Rules (they target group:auto / block /
* direct), so it reads used=false and its card renders "unused" — exactly the case
* the badge exists to surface. A stopped engine reports no chains. */
/**
* Per-hop health, keyed by chain name — what the observatory measured at each
* position of the path, in WIRE order.
*
* `ewan-wg-subs` is the fixture that matters. Its hop 1 is the AmneziaWG node and
* answers; hops 2 and 4 are subscription groups that answer THROUGH it; hop 3 is
* a group whose members all time out at that position. Note that hop 4 rolls up
* `via-tunnel`, the same group whose standalone card reads 0 of 6 alive — alive
* as a hop, dead on its own, both true, because they measure different dial
* paths. That contradiction is the whole point of measuring per hop.
*
* `sub-fresh` is used but never yet reached: every hop untested, nothing dead.
* `relay` is absent from this map on purpose — an unused chain is never
* materialised, so the daemon sends no `hops` key at all, which is "nothing
* measured", not "no hops".
*/
const CHAIN_HOPS: Record<string, ChainHopHealth[]> = {
'ewan-wg-subs': [
{ index: 1, tag: 'chain-ewan-wg-subs-h1', kind: 'node', exit: false, state: 'alive', delay_ms: 41, age_seconds: 22, selected: '', total: 1, tested: 1, alive: 1, dead: 0, untested: 0 },
{ index: 2, tag: 'chain-ewan-wg-subs-h2', kind: 'group', exit: false, state: 'alive', delay_ms: 96, age_seconds: 18, selected: '🇳🇱 Amsterdam-01', total: 298, tested: 122, alive: 119, dead: 3, untested: 176 },
{ index: 3, tag: 'chain-ewan-wg-subs-h3', kind: 'group', exit: false, state: 'dead', delay_ms: 0, age_seconds: 15, selected: '', total: 24, tested: 24, alive: 0, dead: 24, untested: 0 },
{ index: 4, tag: 'chain-ewan-wg-subs-h4', kind: 'group', exit: true, state: 'alive', delay_ms: 148, age_seconds: 19, selected: '🇸🇬 Singapore-09', total: 6, tested: 6, alive: 6, dead: 0, untested: 0 },
],
'sub-fresh': [
{ index: 1, tag: 'chain-sub-fresh-h1', kind: 'node', exit: false, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 1, tested: 0, alive: 0, dead: 0, untested: 1 },
{ index: 2, tag: 'chain-sub-fresh-h2', kind: 'group', exit: true, state: 'untested', delay_ms: 0, age_seconds: -1, selected: '', total: 24, tested: 0, alive: 0, dead: 0, untested: 24 },
],
}
/** Per-chain reachability plus per-hop health for the Targets page. `used` is the
* chain analogue of healthList's field; `hops` is OMITTED (never null, never []),
* exactly like the daemon, for a chain the engine never materialised. A stopped
* engine reports no chains at all. */
function chainHealthList(): ChainHealth[] {
if (!mockPlane().engine) return []
return (CONFIG.Chains ?? []).map((c) => ({ name: c.Name, used: chainUsed(c.Name) }))
return (CONFIG.Chains ?? []).map((c) => {
const hops = CHAIN_HOPS[c.Name]
const h: ChainHealth = { name: c.Name, used: chainUsed(c.Name) }
if (hops) h.hops = hops.map((x) => ({ ...x }))
return h
})
}
/** A chain is "used" when some enabled routing rule (or Final, or a DNS detour)
@@ -1161,15 +1293,23 @@ class ApiErrorLike extends Error {
}
}
// Mock group/chain test. Deliberately covers every state the UI has to render,
// one per target, so a single offline run exercises all of them:
// auto → ok WITH an exit address
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
// SUCCESS, and the case the UI most easily gets wrong
// relay → the chain: same wire shape, `group` carries the CHAIN's name and
// `selected` the node its exit group picked
// fallback → a failure carrying a human reason
// Results land one per GET poll, so the running/progress state is visible too.
// Mock refresh results. The endpoint no longer dials anything: it asks the
// observatory to measure out of turn and reports what the observatory found, so
// every row here is a READ of a background measurement. Deliberately covers every
// state the UI has to render, one per target, so a single offline run exercises
// all of them:
// auto → ok WITH an exit address
// stealth → ok WITHOUT one (delay measured, address undeterminable) — a
// SUCCESS, and the case the UI most easily gets wrong
// ewan-wg-subs → the chain: same wire shape, `group` carries the CHAIN's name
// and `selected` the node its exit hop picked
// via-tunnel → the one honest health FAILURE: a probe that ran and failed
// fallback,
// relay → not routed at all, so no measurement exists to report
// sub-fresh → routed, but the observatory hasn't come round yet
// The last three are absence of measurement, not a broken target, and the copy
// has to keep them apart. Results land one per GET poll, so the running/progress
// state is visible too.
const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_unix'>> = {
auto: {
selected: 'nl-reality-2',
@@ -1187,32 +1327,57 @@ const GROUP_TEST_SHAPE: Record<string, Omit<GroupTestResult, 'group' | 'tested_u
ok: true,
error: '',
},
// The chain — Selected is the node the chain's exit group (auto) picked.
relay: {
selected: 'nl-reality-2',
delay_ms: 61,
exit_ip: '185.12.34.56',
exit_country: 'NL',
ok: true,
error: '',
// The chain, and the pairing that makes the whole feature worth building. A
// chain is one series path, so with hop 3 dead the end-to-end probe CANNOT
// succeed — this row and the hop rail have to tell one story, not two. The
// row says the path is down; the rail says which of the four hops did it,
// which is the part nobody could see before.
'ewan-wg-subs': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'the observatory’s probe through this path failed',
},
// Dead through its tunnel, exactly as its membership health says — the exit
// test and the member health tell the same story about the same group.
// The one real health failure in the fixture: the observatory's probe ran along
// this path and did not come back.
'via-tunnel': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'no member answered through egress awg (6 of 6 timed out)',
error: 'the observatory’s probe through this path failed',
},
// Not a health verdict — nothing routes here, so no measurement of it exists.
fallback: {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error: 'no reachable node in the group (all 3 members timed out)',
error:
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
},
relay: {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use',
},
// Routed, materialised, simply not reached yet. Untested is not dead.
'sub-fresh': {
selected: '',
delay_ms: 0,
exit_ip: '',
exit_country: '',
ok: false,
error:
'the observatory has not reached this target yet — it refreshes on the global probe interval',
},
}
+41 -19
View File
@@ -1,6 +1,6 @@
import './DNS.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, CatSuggest, Led, SrcPicker, Toggle } from '../components'
import { Button, CatSuggest, Led, SrcPicker, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Alert, DNSRule, Model, Resolver } from '../api'
@@ -220,6 +220,7 @@ function describeDetour(
// ---- page ------------------------------------------------------------------
export default function DNS() {
const confirm = useConfirm()
const [config, setConfig] = useState<DNSModel | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -422,15 +423,19 @@ export default function DNS() {
)
const removeBlocklist = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = blocklists[idx]
if (!window.confirm(`Delete blocklist “${target.Name}”? This removes it from the config.`))
return
const ok = await confirm({
label: 'Delete blocklist',
title: `Delete blocklist “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = blocklists.filter((_, i) => i !== idx)
void save({ ...config, Blocklists: next }, `Deleted ${target.Name}`)
},
[config, blocklists, save],
[config, blocklists, save, confirm],
)
// ---- allowlist mutations --------------------------------------------------
@@ -457,15 +462,19 @@ export default function DNS() {
)
const removeAllowlist = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = allowlists[idx]
if (!window.confirm(`Delete allowlist “${target.Name}”? This removes it from the config.`))
return
const ok = await confirm({
label: 'Delete allowlist',
title: `Delete allowlist “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = allowlists.filter((_, i) => i !== idx)
void save({ ...config, Allowlists: next }, `Deleted ${target.Name}`)
},
[config, allowlists, save],
[config, allowlists, save, confirm],
)
// ---- resolver mutations ---------------------------------------------------
@@ -492,11 +501,15 @@ export default function DNS() {
)
const removeResolver = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = resolvers[idx]
if (!window.confirm(`Delete resolver “${target.Name}”? This removes it from the config.`))
return
const ok = await confirm({
label: 'Delete resolver',
title: `Delete resolver “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = resolvers.filter((_, i) => i !== idx)
// Don't leave default/fallback pointing at a resolver that no longer exists.
const g = { ...config.Globals }
@@ -578,15 +591,19 @@ export default function DNS() {
)
const removeDNSRule = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = dnsRules[idx]
if (!window.confirm(`Delete this DNS rule? Matching queries fall back to the default resolver.`))
return
const ok = await confirm({
label: 'Delete DNS rule',
title: 'Delete this DNS rule?',
body: 'Matching queries fall back to the default resolver.',
})
if (!ok) return
const next = dnsRules.filter((_, i) => i !== idx)
void save({ ...config, DNSRules: next }, `Deleted DNS rule → ${target.Resolver}`)
},
[config, dnsRules, save],
[config, dnsRules, save, confirm],
)
// ---- alert mutations ------------------------------------------------------
@@ -610,14 +627,19 @@ export default function DNS() {
)
const removeAlert = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = alerts[idx]
if (!window.confirm(`Delete alert “${target.Name}”? This removes it from the config.`)) return
const ok = await confirm({
label: 'Delete alert',
title: `Delete alert “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = alerts.filter((_, i) => i !== idx)
void save({ ...config, Alerts: next }, `Deleted ${target.Name}`)
},
[config, alerts, save],
[config, alerts, save, confirm],
)
const setAlertVia = useCallback(
+3 -51
View File
@@ -138,57 +138,9 @@
/* inline rename: a quiet pencil affordance beside the name, and the mono input
it swaps to — in the same sink/groove tone as the domain editors. */
.dev-rename {
flex: none;
display: inline-flex;
align-items: center;
justify-content: center;
width: 22px;
height: 22px;
padding: 0;
border: 1px solid transparent;
border-radius: 5px;
background: none;
color: var(--faint);
font-size: 12px;
line-height: 1;
cursor: pointer;
transition: color 0.15s, background 0.15s, border-color 0.15s;
}
.dev-rename:hover:not(:disabled) {
color: var(--accent);
background: color-mix(in srgb, var(--accent) 12%, transparent);
}
.dev-rename:focus-visible {
color: var(--accent);
border-color: var(--accent);
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.dev-rename:disabled {
opacity: 0.5;
cursor: default;
}
.dev-name-input {
min-width: 0;
max-width: 24ch;
padding: 4px 8px;
border: 1px solid var(--accent);
border-radius: 6px;
background: var(--sink);
color: var(--ink);
font-size: 13px;
font-weight: 600;
letter-spacing: 0.01em;
box-shadow: 0 1px 2px var(--shadow) inset;
}
.dev-name-input:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
.dev-name-input:disabled {
opacity: 0.55;
}
/* The pencil button and the name input now live in App.css as .inline-rename /
.inline-rename-input — Nodes grew the same affordance and the two pages must
not drift. */
.dev-id-l2 {
display: flex;
align-items: center;
+13 -7
View File
@@ -1,6 +1,6 @@
import './Devices.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Module, Toggle } from '../components'
import { Button, Led, Module, Toggle, useConfirm } from '../components'
import type { LedVariant } from '../components'
import { apply as apiApply, getConfig, getDevices, putConfig, ApiError } from '../api'
import type { Device, DiscoveredDevice, Model } from '../api'
@@ -81,6 +81,7 @@ function networkLabel(row: DeviceRow): string {
// ---- page ------------------------------------------------------------------
export default function Devices() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -251,17 +252,22 @@ export default function Devices() {
const nameOf = (row: DeviceRow) => row.cfg?.Name || row.hostname || row.ip || 'device'
const removeControl = useCallback(
(row: DeviceRow) => {
async (row: DeviceRow) => {
if (!config) return
const devs = asArray(config.Devices)
const idx = matchDevice(devs, row.mac, row.ip)
if (idx < 0) return
const nm = devs[idx].Name || nameOf(row)
if (!window.confirm(`Stop managing “${nm}”? Its per-device rules are removed; it falls back to network defaults.`))
return
const ok = await confirm({
label: 'Stop managing device',
title: `Stop managing “${nm}”?`,
body: 'Its per-device rules are removed; it falls back to network defaults.',
confirmLabel: 'Stop managing',
})
if (!ok) return
void save({ ...config, Devices: devs.filter((_, i) => i !== idx) }, `Removed control for ${nm}`)
},
[config, save],
[config, save, confirm],
)
const loading = config === null && loadError === null && devices === null && devError === null
@@ -471,7 +477,7 @@ function DeviceCard({
{renaming ? (
<input
ref={nameInput}
className="dev-name-input mono"
className="inline-rename-input mono"
type="text"
spellCheck={false}
autoComplete="off"
@@ -497,7 +503,7 @@ function DeviceCard({
</span>
<button
type="button"
className="dev-rename"
className="inline-rename"
onClick={beginRename}
disabled={busy}
aria-label={`Rename ${name}`}
+12 -13
View File
@@ -1,6 +1,6 @@
import './Networks.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Select, Toggle } from '../components'
import { Button, Led, Select, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, putConfig, ApiError } from '../api'
import type { Inbound, Interface, Model, Status } from '../api'
import { isLanNetwork, isWanNetwork, useInterfaces } from '../srcOptions'
@@ -242,6 +242,7 @@ function computeWarnings(inbounds: Inbound[], ifaces: Interface[]): Warning[] {
// ---- page ------------------------------------------------------------------
export default function Networks({ status }: { status?: Status | null }) {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
const ifaces = useInterfaces()
@@ -392,23 +393,21 @@ export default function Networks({ status }: { status?: Status | null }) {
)
const removeInbound = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = inbounds[idx]
if (
!window.confirm(
`Delete inbound “${target.Name}”?${
intercepts(target)
? ` ${target.Network || 'Its network'} stops going through the tunnel.`
: ''
}`,
)
)
return
const ok = await confirm({
label: 'Delete inbound',
title: `Delete inbound “${target.Name}”?`,
body: intercepts(target)
? `${target.Network || 'Its network'} stops going through the tunnel.`
: undefined,
})
if (!ok) return
const next = inbounds.filter((_, i) => i !== idx)
void save({ ...config, Inbounds: next }, `Deleted ${target.Name}`)
},
[config, inbounds, save],
[config, inbounds, save, confirm],
)
return (
+50
View File
@@ -727,3 +727,53 @@ select.fp-input {
width: 9rem;
}
}
/* ---- inline node rename ----
The pencil / input pair itself is shared (.inline-rename[-input] in App.css);
only the row-local sizing and the refusal message live here. A node name is
longer than a device name (it carries a protocol and a host), so the field is
given more room than the shared 24ch default. */
.node-name-input {
max-width: 32ch;
font-family: var(--font-mono);
font-size: 12.5px;
}
/* Why a rename was refused, pinned under the row it was typed in. Semantic crit:
the name did not change, and that must not be mistaken for a saved edit. */
.row-err {
margin: 2px 0 0;
font-size: 11.5px;
line-height: 1.45;
color: var(--crit);
max-width: 68ch;
}
/* Stated once per subscription bucket: the same rule the locked control in every
row carries, so the absent rename is explained before it is looked for. */
.group-note {
margin: 0;
padding: 8px 12px;
border: 1px solid var(--groove);
border-top: 0;
background: color-mix(in srgb, var(--sink) 25%, transparent);
font-size: 11.5px;
line-height: 1.5;
color: var(--faint);
}
/* The optional name sits beside the link input on a wide row and drops onto its
own line when the row can no longer hold both. */
.add-name {
flex: 0 1 22ch;
min-width: 12ch;
}
.add-row--conf .add-name {
flex: none;
align-self: stretch;
}
@media (max-width: 640px) {
.add-row {
flex-wrap: wrap;
}
.add-name {
flex: 1 1 100%;
}
}
+486 -15
View File
@@ -2,7 +2,7 @@ import './Nodes.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import type { ReactNode } from 'react'
import type { LedVariant } from '../components'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
@@ -158,6 +158,219 @@ function uniqueName(base: string, taken: Set<string>): string {
return `${seed}-${i}`
}
// ---- node names are identity, not a caption --------------------------------
//
// A node's Name IS its sing-box outbound tag and the only thing every reference
// to it spells: a rule target `node:<name>`, a chain hop, a manual group's member
// list, a resolver detour, an alert delivery, a subscription fetch detour. Rename
// the node alone and every one of those points at nothing — and an unresolved
// target does NOT fall back to the default route, the daemon BLOCKS that traffic.
// So the rename either carries every reference with it, or it is refused.
/** Reserved outbound tags. A node called this is skipped by the generator entirely. */
const RESERVED_TAGS = ['direct', 'block']
/**
* Prefixes that `model.SplitTarget` reads as a KIND, not as part of a name. A
* name starting with one of them makes every bare reference to it ambiguous with
* a real `kind:name` reference, so it is refused rather than half-supported.
*/
const KIND_PREFIXES = ['node', 'group', 'egress', 'chain', 'direct', 'block']
/** Names are rendered into a line-oriented `uci export`; control chars are stripped there. */
function hasControlChar(s: string): boolean {
for (let i = 0; i < s.length; i++) {
const c = s.charCodeAt(i)
if (c < 0x20 || c === 0x7f) return true
}
return false
}
/**
* Why `name` cannot be a node name here, or null if it can.
*
* Every rule mirrors something the daemon actually does with the name, not a
* house style: reserved tags make generate skip the node; a duplicate makes two
* outbounds share a tag and the manager silently keeps the last one; a group of
* the same name is dropped by buildGroups ("rename the group"); an
* `egress-<name>` collision takes over a real egress outbound; and a control
* character is rewritten to a space by sanitizeUCIValue on write, so the saved
* name would not be the one you typed.
*/
function nodeNameError(
raw: string,
m: Model | null,
self: string | null,
): string | null {
const name = raw.trim()
if (!name) return 'A node needs a name.'
if (hasControlChar(name))
return 'Names can’t contain line breaks or control characters — they’re stripped when the config is written.'
if (RESERVED_TAGS.some((t) => t.toLowerCase() === name.toLowerCase()))
return `“${name}” is a reserved target name — a node called that is skipped by the engine. Pick another.`
const head = name.includes(':') ? name.slice(0, name.indexOf(':')).toLowerCase() : ''
if (head && KIND_PREFIXES.includes(head))
return `A name starting with “${head}:” reads as a ${head} reference everywhere it’s used. Pick another.`
if (!m) return null
const clash = asArray(m.Nodes).find((n) => n.Name === name && n.Name !== self)
if (clash)
return clash.FromSub
? `“${name}” is already a node from subscription “${clash.FromSub}”. Two nodes with one name share a single outbound — pick another.`
: `“${name}” is already another node. Pick another.`
if (asArray(m.Groups).some((g) => g.Name === name))
return `A group is already named “${name}”. The engine drops the group when a node takes its name — pick another.`
const egressClash = asArray(m.Egresses).find((e) => `egress-${e.Name}` === name)
if (egressClash)
return `“${name}” is the outbound tag of egress “${egressClash.Name}”. Pick another.`
return null
}
/** One place a node name is written, as a short label for the rename summary. */
interface NodeRefSite {
/** Which section — drives the "N rules, M groups" count. */
kind: 'rule' | 'group' | 'chain' | 'resolver' | 'alert' | 'subscription' | 'egress'
label: string
}
/** A target/detour string naming this node in its prefixed form (`node:<name>`). */
const isNodeRef = (v: string | undefined | null, name: string): boolean =>
(v ?? '') === `node:${name}`
/** …or in the bare form the engine also resolves (a group member, a bare hop/target). */
const isBareRef = (v: string | undefined | null, name: string): boolean => (v ?? '') === name
/**
* Every place `name` is written outside the node itself. Both spellings count:
* `resolveTarget` falls through to a bare node lookup, and a manual group's
* member list is bare by contract.
*/
function findNodeReferences(m: Model, name: string): NodeRefSite[] {
const out: NodeRefSite[] = []
for (const r of asArray(m.Rules)) {
if (isNodeRef(r.Target, name) || isBareRef(r.Target, name))
out.push({ kind: 'rule', label: `rule “${r.Name}” target` })
}
for (const g of asArray(m.Groups)) {
if (asArray(g.Nodes).some((n) => n === name))
out.push({ kind: 'group', label: `group “${g.Name}” member` })
}
for (const c of asArray(m.Chains)) {
if (asArray(c.Hops).some((h) => isNodeRef(h, name) || isBareRef(h, name)))
out.push({ kind: 'chain', label: `chain “${c.Name}” hop` })
}
for (const r of asArray(m.Resolvers)) {
if (isNodeRef(r.Detour, name)) out.push({ kind: 'resolver', label: `resolver “${r.Name}” DNS path` })
}
for (const a of asArray(m.Alerts)) {
if (isNodeRef(a.Via, name)) out.push({ kind: 'alert', label: `alert “${a.Name}” delivery` })
}
for (const s of asArray(m.Subscriptions)) {
if (isNodeRef(s.FetchDetour, name))
out.push({ kind: 'subscription', label: `subscription “${s.Name}” fetch` })
}
for (const e of asArray(m.Egresses)) {
if (isNodeRef(e.Target, name)) out.push({ kind: 'egress', label: `egress “${e.Name}” target` })
}
return out
}
/**
* What makes a rename impossible to carry rather than merely wide.
*
* A BARE reference is just a name; the engine resolves it node-first, then group.
* If something else already answers to the old name, we cannot tell which object
* a bare reference meant, and rewriting it would move a reference the operator
* never pointed at this node. That is a half-done cascade, so the rename is
* refused instead — with the collision named, so it can be fixed.
*/
function bareAmbiguity(m: Model, name: string): string | null {
const group = asArray(m.Groups).find((g) => g.Name === name)
if (!group) return null
const bare = [
...asArray(m.Rules)
.filter((r) => isBareRef(r.Target, name))
.map((r) => `rule “${r.Name}”`),
...asArray(m.Chains)
.filter((c) => asArray(c.Hops).some((h) => isBareRef(h, name)))
.map((c) => `chain “${c.Name}”`),
]
if (bare.length === 0) return null
return `A group is also named “${name}”, and ${bare.join(', ')} point${bare.length === 1 ? 's' : ''} at that bare name — there is no way to tell which of the two is meant. Rename the group first, then this node.`
}
/**
* Rewrite every reference from `from` to `to`. Returns a NEW Model with only the
* touched sections replaced; the Nodes section is the caller's business.
*
* Bare references are rewritten too — that is the whole point for a manual
* group's member list — which is safe only because `bareAmbiguity` has already
* refused the one case where a bare name could mean something else.
*/
function renameNodeReferences(m: Model, from: string, to: string): Model {
if (from === to) return m
/** Prefixed-only sites (a detour is never spelled bare). */
const pfx = (v: string | undefined) => (isNodeRef(v, from) ? `node:${to}` : v)
/** Sites that accept either spelling — each is rewritten in the spelling it already uses. */
const either = (v: string | undefined) => {
if (isNodeRef(v, from)) return `node:${to}`
if (isBareRef(v, from)) return to
return v
}
const next: Model = { ...m }
if (m.Rules) next.Rules = m.Rules.map((r) => ({ ...r, Target: either(r.Target) }))
if (m.Groups)
next.Groups = m.Groups.map((g) => ({
...g,
Nodes: g.Nodes ? g.Nodes.map((n) => (n === from ? to : n)) : g.Nodes,
}))
if (m.Chains)
next.Chains = m.Chains.map((c) => ({
...c,
Hops: c.Hops ? c.Hops.map((h) => either(h) ?? h) : c.Hops,
}))
if (m.Resolvers) next.Resolvers = m.Resolvers.map((r) => ({ ...r, Detour: pfx(r.Detour) }))
if (m.Alerts) next.Alerts = m.Alerts.map((a) => ({ ...a, Via: pfx(a.Via) }))
if (m.Subscriptions)
next.Subscriptions = m.Subscriptions.map((s) => ({ ...s, FetchDetour: pfx(s.FetchDetour) }))
if (m.Egresses) next.Egresses = m.Egresses.map((e) => ({ ...e, Target: pfx(e.Target) }))
return next
}
/** "3 rules, 1 group and 2 chains" — what the rename is about to rewrite. */
function refSummary(refs: NodeRefSite[]): string {
const plural: Record<NodeRefSite['kind'], [string, string]> = {
rule: ['rule', 'rules'],
group: ['group', 'groups'],
chain: ['chain', 'chains'],
resolver: ['resolver', 'resolvers'],
alert: ['alert', 'alerts'],
subscription: ['subscription', 'subscriptions'],
egress: ['egress', 'egresses'],
}
const order: NodeRefSite['kind'][] = [
'rule', 'group', 'chain', 'resolver', 'alert', 'subscription', 'egress',
]
const parts = order
.map((k) => [k, refs.filter((r) => r.kind === k).length] as const)
.filter(([, n]) => n > 0)
.map(([k, n]) => `${n} ${plural[k][n === 1 ? 0 : 1]}`)
if (parts.length === 1) return parts[0]
return `${parts.slice(0, -1).join(', ')} and ${parts[parts.length - 1]}`
}
/**
* The name a rename just committed to, waiting for its row to come back.
*
* A row is keyed by the node's NAME, so committing a rename unmounts the row and
* mounts a different one — carrying the focused element away with it. This baton
* survives that remount: the row that reappears under the new name claims it and
* puts the keyboard back on its own rename button, instead of dropping the user
* on <body> halfway down a list of 300 nodes.
*/
let pendingRenameFocus: string | null = null
// A subscription with more than this many nodes starts collapsed so the list
// doesn't become one endless scroll; an active search overrides it.
const LARGE_GROUP = 20
@@ -289,6 +502,7 @@ function DetourSelect({
// ---- page ------------------------------------------------------------------
export default function Nodes() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -385,6 +599,9 @@ export default function Nodes() {
const [nodeInput, setNodeInput] = useState('')
const [nodeErr, setNodeErr] = useState<string | null>(null)
const [addMode, setAddMode] = useState<'link' | 'conf'>('link')
// Optional. Empty keeps the old behaviour (a name derived from the server
// address), so "paste a link, press Add" stays a two-step path.
const [nodeName, setNodeName] = useState('')
const [importing, setImporting] = useState(false)
// ---- node search + collapsible grouping -----------------------------------
@@ -445,8 +662,12 @@ export default function Nodes() {
try {
const { uri, name } = await importWg(conf)
const taken = new Set(nodes.map((n) => n.Name))
// A typed name is used AS TYPED — uniqueName would silently turn a
// collision into "name-2", which is the confusion this field exists to
// end. It is validated instead, and a clash is refused out loud above.
const wanted = nodeName.trim()
const node: NodeCfg = {
Name: uniqueName(name || 'wireguard', taken),
Name: wanted || uniqueName(name || 'wireguard', taken),
Enabled: true,
URI: uri,
FromSub: '',
@@ -455,6 +676,7 @@ export default function Nodes() {
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
if (ok) {
setNodeInput('')
setNodeName('')
setAddMode('link')
}
} catch (e) {
@@ -463,11 +685,20 @@ export default function Nodes() {
setImporting(false)
}
},
[config, nodes, save, flash],
[config, nodes, nodeName, save, flash],
)
const addNode = useCallback(async () => {
if (!config) return
// The name is checked BEFORE the import round-trip, so a bad name costs
// nothing and the message lands in the form next to the field.
if (nodeName.trim()) {
const bad = nodeNameError(nodeName, config, null)
if (bad) {
setNodeErr(bad)
return
}
}
// Auto-detect a pasted config, whichever input it landed in.
if (nodeInput.includes(WG_MARKER)) {
await addWgConf(nodeInput)
@@ -486,10 +717,20 @@ export default function Nodes() {
const parsed = parseShareLink(uri)
const taken = new Set(nodes.map((n) => n.Name))
const base = parsed.suggested || `${parsed.proto.toLowerCase()}-${parsed.host}`.replace(/[^\w.:-]+/g, '-')
const node: NodeCfg = { Name: uniqueName(base, taken), Enabled: true, URI: uri, FromSub: '', Egress: '' }
const wanted = nodeName.trim()
const node: NodeCfg = {
Name: wanted || uniqueName(base, taken),
Enabled: true,
URI: uri,
FromSub: '',
Egress: '',
}
const ok = await save({ ...config, Nodes: [...nodes, node] }, `Added ${node.Name}`)
if (ok) setNodeInput('')
}, [config, nodeInput, nodes, save, addMode, addWgConf])
if (ok) {
setNodeInput('')
setNodeName('')
}
}, [config, nodeInput, nodeName, nodes, save, addMode, addWgConf])
const toggleNode = useCallback(
(idx: number, on: boolean) => {
@@ -501,14 +742,88 @@ export default function Nodes() {
)
const removeNode = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = nodes[idx]
if (!window.confirm(`Delete node “${target.Name}”? This removes it from the config.`)) return
const ok = await confirm({
label: 'Delete node',
title: `Delete node “${target.Name}”?`,
body: 'This removes it from the config.',
})
if (!ok) return
const next = nodes.filter((_, i) => i !== idx)
void save({ ...config, Nodes: next }, `Deleted ${target.Name}`)
},
[config, nodes, save],
[config, nodes, save, confirm],
)
/**
* Rename a manual node, carrying every reference with it.
*
* The name is this node's identity: its outbound tag, and the exact string a
* rule target, a chain hop, a manual group's member list, a resolver detour, an
* alert delivery and a subscription fetch detour all spell. So the rename is one
* atomic save of the Nodes section AND every referencing section, or it does not
* happen at all:
*
* - an invalid or colliding name is refused with the reason (`nodeNameError`);
* - a name a GROUP also answers to, with bare references pointing at it, is
* refused too — there is no way to know which object those meant, and
* guessing would move a reference the operator never pointed here;
* - anything else is shown exactly what it will rewrite, and only then saved.
*
* Errors surface through `onError` so they land in the row that was edited.
*/
const renameNode = useCallback(
async (idx: number, raw: string, onError: (msg: string) => void): Promise<boolean> => {
if (!config) return false
const target = nodes[idx]
const from = target.Name
const to = raw.trim()
if (to === from) return true
// Subscription names come back from the feed on the next update; renaming
// one would be undone without warning, so this path is manual-only.
if (target.FromSub) {
onError(`“${from}” is named by subscription “${target.FromSub}” — the feed rewrites it on the next update.`)
return false
}
const bad = nodeNameError(to, config, from)
if (bad) {
onError(bad)
return false
}
const blocked = bareAmbiguity(config, from)
if (blocked) {
onError(blocked)
return false
}
const refs = findNodeReferences(config, from)
if (refs.length > 0) {
const shown = refs.slice(0, 4).map((r) => r.label)
const more = refs.length - shown.length
const ok = await confirm({
tone: 'neutral',
label: 'Rename node',
title: `Rename “${from}” to “${to}”?`,
body: `This also updates ${refSummary(refs)} that point at it — ${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}. They are saved together, so nothing is left pointing at the old name.`,
confirmLabel: 'Rename',
})
if (!ok) return false
}
// One PUT: the node and every reference move in the same write, so no
// intermediate state exists where a reference dangles.
const carried = renameNodeReferences(config, from, to)
const next = asArray(carried.Nodes).map((n, i) => (i === idx ? { ...n, Name: to } : n))
return save(
{ ...carried, Nodes: next },
refs.length > 0
? `Renamed to ${to} — updated ${refs.length} reference${refs.length === 1 ? '' : 's'}`
: `Renamed to ${to}`,
)
},
[config, nodes, save, confirm],
)
// Pin (or clear) one node's dial egress. Same optimistic save→apply path as
@@ -563,16 +878,20 @@ export default function Nodes() {
)
const removeSub = useCallback(
(idx: number) => {
async (idx: number) => {
if (!config) return
const target = subs[idx]
const hasCache = nodes.some((n) => n.FromSub === target.Name)
const extra = hasCache ? ' Its cached nodes stay until you next apply.' : ''
if (!window.confirm(`Delete subscription “${target.Name}”?${extra}`)) return
const ok = await confirm({
label: 'Delete subscription',
title: `Delete subscription “${target.Name}”?`,
body: hasCache ? 'Its cached nodes stay until you next apply.' : undefined,
})
if (!ok) return
const next = subs.filter((_, i) => i !== idx)
void save({ ...config, Subscriptions: next }, `Deleted ${target.Name}`)
},
[config, subs, nodes, save],
[config, subs, nodes, save, confirm],
)
// Commit an options edit for one subscription. The editor hands back a fully
@@ -736,6 +1055,20 @@ export default function Nodes() {
disabled={busy || importing || !config}
/>
)}
<input
className="fp-input add-name"
type="text"
spellCheck={false}
autoComplete="off"
placeholder="Name (optional)"
aria-label="Node name — optional"
value={nodeName}
onChange={(e) => {
setNodeName(e.target.value)
if (nodeErr) setNodeErr(null)
}}
disabled={busy || importing || !config}
/>
<Button type="submit" variant="primary" disabled={busy || importing || !config}>
{importing ? 'Importing…' : saving ? 'Saving…' : addMode === 'conf' ? 'Import' : 'Add node'}
</Button>
@@ -743,7 +1076,8 @@ export default function Nodes() {
<p className="add-hint">
{addMode === 'conf'
? 'Paste a wg-quick / AmneziaWG .conf — it starts with [Interface].'
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}
: 'vless://, ss://, trojan://, hysteria2://… A pasted [Interface] config is imported automatically.'}{' '}
Leave the name empty and it’s taken from the server address; you can rename it later.
</p>
</div>
{nodeErr && (
@@ -800,6 +1134,7 @@ export default function Nodes() {
onToggle={() => toggleGroup(g)}
onToggleNode={toggleNode}
onRemoveNode={removeNode}
onRenameNode={renameNode}
onSetEgress={setNodeEgress}
/>
))}
@@ -911,6 +1246,7 @@ function NodeGroup({
onToggle,
onToggleNode,
onRemoveNode,
onRenameNode,
onSetEgress,
}: {
group: NodeGroupData
@@ -920,6 +1256,7 @@ function NodeGroup({
onToggle: () => void
onToggleNode: (idx: number, on: boolean) => void
onRemoveNode: (idx: number) => void
onRenameNode: (idx: number, name: string, onError: (msg: string) => void) => Promise<boolean>
onSetEgress: (idx: number, egress: string) => Promise<boolean>
}) {
const panelId = `node-group-${group.key || 'manual'}`
@@ -940,6 +1277,12 @@ function NodeGroup({
<span className="group-count mono">{count}</span>
</button>
</h3>
{open && group.key !== '' && (
<p className="group-note">
Names come from the subscription feed and are rewritten on every update, so nodes in this
list can’t be renamed here.
</p>
)}
{open && (
<ul id={panelId} className="rows-list group-rows">
{group.items.map(({ node, idx }) => (
@@ -950,6 +1293,7 @@ function NodeGroup({
egressNames={egressNames}
onToggle={(on) => onToggleNode(idx, on)}
onDelete={() => onRemoveNode(idx)}
onRename={(name, onError) => onRenameNode(idx, name, onError)}
onSetEgress={(egress) => onSetEgress(idx, egress)}
/>
))}
@@ -965,6 +1309,7 @@ function NodeRow({
egressNames,
onToggle,
onDelete,
onRename,
onSetEgress,
}: {
node: NodeCfg
@@ -972,6 +1317,7 @@ function NodeRow({
egressNames: string[]
onToggle: (on: boolean) => void
onDelete: () => void
onRename: (name: string, onError: (msg: string) => void) => Promise<boolean>
onSetEgress: (egress: string) => Promise<boolean>
}) {
const { proto, host, hasCreds } = useMemo(() => parseShareLink(node.URI), [node.URI])
@@ -982,6 +1328,68 @@ function NodeRow({
const [open, setOpen] = useState(false)
const panelId = `node-egress-${node.FromSub || 'manual'}-${node.Name}`
// ---- inline rename (same interaction as a device row) ---------------------
// Enter commits, Esc cancels, blur commits; a ref-guard keeps Esc-then-blur
// from committing twice. Unlike a device, the commit can be REFUSED (a name
// collision, or references that can't be carried), so the input stays open
// with the reason under it instead of closing on a change that never happened.
const [renaming, setRenaming] = useState(false)
const [draft, setDraft] = useState(node.Name)
const [renameErr, setRenameErr] = useState<string | null>(null)
const nameInput = useRef<HTMLInputElement>(null)
const renameBtn = useRef<HTMLButtonElement>(null)
const finished = useRef(false)
const beginRename = () => {
setDraft(node.Name)
setRenameErr(null)
finished.current = false
setRenaming(true)
}
const finishRename = async (commit: boolean) => {
if (finished.current) return
finished.current = true
const nm = draft.trim()
if (!commit || !nm || nm === node.Name) {
setRenaming(false)
setRenameErr(null)
return
}
// Armed BEFORE the save: the renamed row remounts the moment the config
// state lands, which is before this await resolves. Arming afterwards would
// always miss it.
pendingRenameFocus = nm
const ok = await onRename(nm, (msg) => setRenameErr(msg))
if (ok) {
setRenaming(false)
setRenameErr(null)
} else {
if (pendingRenameFocus === nm) pendingRenameFocus = null
// Refused — hold the field open on the rejected text so it can be fixed.
finished.current = false
nameInput.current?.focus()
}
}
useEffect(() => {
if (renaming) {
nameInput.current?.focus()
nameInput.current?.select()
}
}, [renaming])
// Claim the baton if this row is the one the rename produced. The row remounts
// while the PUT is still in flight, so on that first pass the button is still
// disabled and focus() would be a silent no-op — the baton is held until the
// save settles and this effect re-runs with a focusable button.
useEffect(() => {
if (pendingRenameFocus !== node.Name) return
const btn = renameBtn.current
if (!btn || btn.disabled) return
pendingRenameFocus = null
btn.focus()
}, [node.Name, busy])
return (
<li className={`row-item node-row${open ? ' node-row--open' : ''}`}>
<div className="row-head">
@@ -993,10 +1401,73 @@ function NodeRow({
/>
<div className="row-main">
<div className="row-line1">
<span className="row-name">{node.Name}</span>
{renaming ? (
<input
ref={nameInput}
className="inline-rename-input node-name-input mono"
type="text"
spellCheck={false}
autoComplete="off"
value={draft}
aria-label={`Rename node ${node.Name}`}
aria-invalid={renameErr ? true : undefined}
onChange={(e) => {
setDraft(e.target.value)
if (renameErr) setRenameErr(null)
}}
onBlur={() => void finishRename(true)}
onKeyDown={(e) => {
if (e.key === 'Enter') {
e.preventDefault()
void finishRename(true)
} else if (e.key === 'Escape') {
e.preventDefault()
void finishRename(false)
}
}}
disabled={busy}
/>
) : (
<>
<span className="row-name" title={node.Name}>
{node.Name}
</span>
{managed ? (
// Not hidden — withheld, with the reason attached. A control
// that quietly isn't there reads as a bug; this one states the
// rule, and the same sentence is on the group header above.
<button
type="button"
className="inline-rename inline-rename--locked"
disabled
aria-label={`Can’t rename ${node.Name} — its name comes from subscription “${node.FromSub}” and is rewritten on the next update`}
title={`Named by subscription “${node.FromSub}” — the feed rewrites this name on the next update. Rename it in the subscription, or add the node manually.`}
>
🔒
</button>
) : (
<button
ref={renameBtn}
type="button"
className="inline-rename"
onClick={beginRename}
disabled={busy}
aria-label={`Rename node ${node.Name}`}
title="Rename"
>
✎
</button>
)}
</>
)}
<span className="badge">{proto}</span>
{node.Stale && <span className="badge badge--warn">stale</span>}
</div>
{renameErr && (
<p className="row-err" role="alert">
{renameErr}
</p>
)}
<div className="row-line2 mono">
<span className="row-host">{host}</span>
{hasCreds && (
+17 -9
View File
@@ -341,7 +341,7 @@ export function Overview({
led={{ variant: len(config?.Rules) ? 'on' : 'amber' }}
rows={[
{ k: 'egresses', v: String(len(config?.Egresses)) },
{ k: 'default', v: defaultTarget(config), hot: true },
{ k: 'default', v: defaultTarget(status, config), hot: true },
]}
/>
@@ -563,12 +563,20 @@ const NAV_LABEL: Record<Route, string> = {
// for the apply/rollback flow, where the individual flags are the actual
// subject of the page.)
function defaultTarget(config: Model | null): string {
const rules = config?.Rules ?? []
if (rules.length === 0) return '—'
// The highest Order enabled rule is the effective catch-all.
const enabled = rules.filter((r) => r.Enabled)
if (enabled.length === 0) return 'none'
const last = enabled.reduce((a, b) => (b.Order >= a.Order ? b : a))
return last.Target || last.Egress || last.Name
/** Where everything not matched by a rule goes — the engine's route `final`.
*
* Taken from the daemon (status.traffic.default), which reads it off the config
* it is running. The guess this replaced was "the highest-Order enabled rule",
* and that is not what the default is: a rule only becomes the default by having
* NO conditions at all, whatever its Order (model.IsCatchAll), so a specific
* high-Order rule was routinely printed here as the router's default. It also
* described the config on disk rather than the one running, and could not see a
* target that failed to resolve and fell back.
*
* Falls back to the rule count only when the daemon has not reported — never to
* a guess about where traffic goes. */
function defaultTarget(status: Status | null, config: Model | null): string {
const d = status?.traffic?.default
if (d) return d
return len(config?.Rules) === 0 ? '—' : 'not reported'
}
+10 -4
View File
@@ -1,6 +1,6 @@
import './Profiles.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import { apply as apiApply, getConfig, getInterfaces, putConfig, ApiError } from '../api'
import type { Interface, Model, Profile } from '../api'
@@ -36,6 +36,7 @@ function namesOf(v: unknown): string[] {
// ---- page ------------------------------------------------------------------
export default function Profiles() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -192,9 +193,14 @@ export default function Profiles() {
)
const deleteProfile = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
if (!window.confirm(`Delete profile “${name}”? Its overrides stop applying.`)) return
const ok = await confirm({
label: 'Delete profile',
title: `Delete profile “${name}”?`,
body: 'Its overrides stop applying.',
})
if (!ok) return
const next = profiles.filter((p) => p.Name !== name)
const g =
config.Globals.ActiveProfile === name
@@ -202,7 +208,7 @@ export default function Profiles() {
: config.Globals
void save({ ...config, Profiles: next, Globals: g }, `Deleted ${name}`)
},
[config, profiles, save],
[config, profiles, save, confirm],
)
// ---- expansion (only one profile editor open at a time) -------------------
+86 -6
View File
@@ -58,6 +58,45 @@
color: var(--ink);
}
/* ---- active-profile banner ----
*
* Deliberately NOT the accent plate the save→apply bar wears above. Orange is
* "there is something for you to do" on this faceplate, and an active WAN profile
* is a standing condition, not a pending action. A quiet plate with an amber tag
* reads as "note the state" — and it is the SAME amber the overridden rows below
* carry, so the banner and its rows are visibly one story rather than two
* unrelated oddities. */
.rt-prof-banner {
display: flex;
align-items: flex-start;
gap: 10px;
margin: 0 0 calc(var(--u, 8px) * 2.5);
padding: 10px 14px;
border: 1px solid color-mix(in srgb, var(--amber) 35%, var(--groove));
border-radius: 8px;
background: color-mix(in srgb, var(--amber) 7%, transparent);
font-family: var(--font-sans);
font-size: 12px;
line-height: 1.55;
color: var(--dim);
}
.rt-prof-banner strong {
color: var(--ink);
font-weight: 600;
}
.rt-prof-tag {
flex: none;
margin-top: 1px;
padding: 2px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
font-size: 9px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--amber);
}
/* ---- empty state ---- */
.rt-empty {
padding: calc(var(--u, 8px) * 4) 0 calc(var(--u, 8px) * 3);
@@ -289,6 +328,51 @@
opacity: 0.62;
}
/* ---- a rule the active WAN profile overrides ----
*
* The row itself needs no new paint: an overridden-off rule already wears `.off`
* (it is off, whatever its switch says) and an overridden-on rule wears nothing
* (it is on). What was missing was never colour — it was the sentence naming who
* decided. So this is the per-row twin of the banner and borrows .rt-dead-note's
* type wholesale: same voice, same size, one <p> margin to reset. */
/* The same pill as .rt-badge.dead, so the two override states read as one pair,
* but in accent — a rule the profile forces ON is active, and active is orange on
* this faceplate. The pill is also what keeps it from running into the plain
* "default route · final" badge beside it, where "final on · by profile" read as
* one phrase. */
.rt-badge.prof-on {
padding: 1px 7px;
border: 1px solid var(--accent-soft);
border-radius: 999px;
background: color-mix(in srgb, var(--accent) 10%, transparent);
}
.rt-prof-note {
margin: 0;
}
.rt-prof-note strong {
color: var(--ink);
font-weight: 600;
}
/* Switch + its legend. The caption shows ONLY while a profile overrides the rule,
* and it is what keeps the control honest: the plate says what the router is
* doing, this says the switch is about the saved setting. A legend under the
* control it names is the faceplate's own idiom. */
.rt-switch {
display: inline-flex;
flex-direction: column;
align-items: center;
gap: 3px;
}
.rt-switch-note {
font-family: var(--font-mono);
font-size: 8.5px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--faint);
}
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
.rt-target {
display: inline-flex;
@@ -396,9 +480,6 @@
gap: 5px;
min-width: 0;
}
.rt-field-wide {
grid-column: span 2;
}
.rt-flabel {
font-family: var(--font-mono);
font-size: 9px;
@@ -515,6 +596,8 @@ select.rt-input {
border-color: var(--accent);
box-shadow: 0 1px 0 var(--edge) inset, 0 0 0 1px var(--accent-soft);
}
/* "no matchers" flag in the plate foot — shared by BOTH rule forms (add and
* edit), so the same non-blocking warning reads identically in either. */
.rt-edit-warn {
font-family: var(--font-mono);
font-size: 11.5px;
@@ -837,9 +920,6 @@ select.rt-input {
justify-content: flex-start;
align-self: start;
}
.rt-field-wide {
grid-column: auto;
}
.rt-rs-row {
grid-template-columns: 1fr;
row-gap: 10px;
+329 -145
View File
@@ -1,7 +1,7 @@
import './Routing.css'
import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
import type { FormEvent, ReactNode } from 'react'
import { Button, CatSuggest, SrcPicker, Toggle } from '../components'
import { Button, CatSuggest, SrcPicker, Toggle, useConfirm } from '../components'
import {
apply as apiApply,
getConfig,
@@ -22,9 +22,7 @@ import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
type RRule = Rule & {
Src?: string[] | null
DstDomain?: string[] | null
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
Proto?: string
Kill?: string
@@ -35,6 +33,29 @@ type RRule = Rule & {
// Minutes east of UTC anchoring the schedule's wall-clock times; captured
// from the editing browser on save (the router has no tzdata). 0 ⇒ UTC.
SchedUTCOffset?: number
// The daemon's unmigrated-rule tripwire (model.Rule.LegacyDst), read-only here.
// Non-empty ⇒ the config STILL carries the schema-v1 `dst_domain`/`dst_ip` that
// schema v2 removed, i.e. `shaterd migrate` never ran or could not commit. Each
// element is the raw `<option>=<value>` text so the panel can quote what was
// found. The daemon holds such a rule disabled; the panel only reports it (it
// is never rendered back to UCI, so a config write from here drains it out).
// Field name is the Go one: model.Rule has no json tags.
LegacyDst?: string[] | null
}
/**
* Whether a rule is in force, kept strictly apart from whether it is switched on.
*
* `on` is the EFFECTIVE state — what the router is actually doing — and every mark
* on the row is drawn from it. `profile`/`dir` are set only when the active WAN
* profile is the reason the two differ, so the row can name who overrode the
* saved setting instead of leaving the operator to guess why a switch that reads
* "on" routes nothing.
*/
type RuleForce = {
on: boolean
profile: string | null
dir: 'enabled' | 'disabled' | null
}
/**
@@ -97,7 +118,6 @@ function ProtoOptions({ value }: { value: string }) {
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
const byOrder = (a: RRule, b: RRule): number => a.Order - b.Order
const csv = (s: string): string[] => s.split(',').map((x) => x.trim()).filter(Boolean)
// --- ruleset helpers --------------------------------------------------------
// A `config ruleset` (api.ts Ruleset) is a named domain/ipcidr list a rule
@@ -213,18 +233,59 @@ function everyLabel(sec: number): string {
return `every ${sec}s`
}
/** A rule with no matcher of any kind is the effective catch-all (route Final). */
/** A rule with no matcher of any kind is the effective catch-all (route Final).
* Mirrors model.IsCatchAll on the daemon side — the two must agree or the
* "never applies" badge lands on a different row than the apply warning.
*
* AN UNMIGRATED RULE IS NEVER A CATCH-ALL, and that is the first thing checked
* here, exactly as on the Go side (model/reachability.go). When LegacyDst is
* non-empty the rule's destination is still written in the schema-v1 options
* the parser no longer reads, so its lack of matchers means "the destination is
* unreadable", not "matches everything" — reading it the other way is precisely
* what turned an uncommitted `shaterd migrate` into route Final for the whole
* router. The daemon holds such a rule disabled and reports false here; if this
* copy disagreed, the panel would paint the row "default route · final" while
* the daemon routes nothing through it. */
function isCatchAll(r: RRule): boolean {
if (len(r.LegacyDst) > 0) return false
return (
len(r.Src) === 0 &&
len(r.DstDomain) === 0 &&
len(r.DstRuleset) === 0 &&
len(r.DstIP) === 0 &&
!(r.DstPort && r.DstPort.trim()) &&
!(r.Proto && r.Proto.trim())
)
}
/** isCatchAll's twin for a form still being edited: the live fields of either
* rule form with no matcher left in them.
*
* Such a rule is not "matches all" in the ordinary sense — the engine emits it
* as route.Final, and among several the LAST one in rule order owns it. So what
* saving one actually does depends on what is last right now, and there are
* three cases:
* - no rules at all, or the last rule is a conditional one → the new rule
* lands last and TAKES the default, silently retargeting every otherwise
* unmatched flow (e.g. the whole LAN to `direct`, past the tunnel);
* - the last rule is already a catch-all → nextOrder() deliberately inserts
* the new one BEFORE it (and bumps the old one up), so the existing default
* keeps route.Final and the new rule is dead on arrival — the daemon
* reports it as shadowed and the row renders as such.
* Both outcomes are worth a warning, and neither form knows which it will be
* (the add form has no rule list), so the shared text says only what is certain:
* the rule has no matchers. It is a legal configuration either way, so neither
* form blocks it — they warn, from this one predicate, so the flag cannot drift
* out of sync between add and edit. */
function formHasNoMatchers(f: {
src: string[]
port: string
rulesets: string[]
proto: string
}): boolean {
return (
f.src.length === 0 && f.port.trim() === '' && f.rulesets.length === 0 && f.proto.trim() === ''
)
}
/** Effective routing target for a rule (Target wins; a bare Egress is a target too). */
function effectiveTarget(r: RRule): string {
if (r.Target && r.Target.trim()) return r.Target.trim()
@@ -263,16 +324,19 @@ interface TargetGroups {
nodes: TargetOpt[] // node:<n> (huge — rendered last)
}
// Free-text destination matchers offered by the ADD form. Domains are NOT one of
// them (the user's call): domain matching goes through named rulesets — that's
// what they exist for. 'none' = the rule matches by rulesets/source/proto alone.
// (Legacy rules that already carry DstDomain stay editable in the edit form.)
type MatchKind = 'none' | 'ip' | 'port'
// The add form's fields. WHERE traffic is going is a ruleset choice and nothing
// else — a rule has no inline domain or address list any more, so the old
// Match-kind picker (rulesets / ip / port) collapsed into a plain Port field
// beside the ruleset picker. The cost is real and accepted: routing a single
// domain is no longer done here — you leave for the Rulesets panel, create the
// list, fill it, and come back to check it. A "create a list from here" shortcut
// was proposed and rejected (DECISIONS.md D21): a second place to author a list
// is a second place for its entry semantics and duplicate-name rules to drift,
// which is the exact thing D21 removed.
interface AddForm {
name: string
src: string[]
matchKind: MatchKind
matchValue: string
port: string
rulesets: string[]
proto: string
target: string
@@ -284,8 +348,7 @@ interface AddForm {
const EMPTY_FORM: AddForm = {
name: '',
src: [],
matchKind: 'none',
matchValue: '',
port: '',
rulesets: [],
proto: '',
target: 'direct',
@@ -318,6 +381,7 @@ const browserTZName = (): string => {
}
export default function Routing() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
const [actionError, setActionError] = useState<string | null>(null)
@@ -435,7 +499,7 @@ export default function Routing() {
}, [config])
/**
* The verdict for one rule, or null when it can fire.
* The daemon's verdict for one rule, or null when we have none that describes it.
*
* Verdicts are fetched separately from the config, so between an optimistic edit
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
@@ -443,18 +507,53 @@ export default function Routing() {
* badge on a working rule: a mismatch means the verdict is not about this row,
* and no badge is the honest answer.
*/
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const verdictOf = useCallback(
(r: RRule): RuleReach | null => {
const i = modelIndex.get(r)
if (i === undefined) return null
const v = reach.get(i)
if (!v || !v.unreachable || !v.shadowed_by) return null
if (v.name !== r.Name || v.order !== r.Order) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
if (!v || v.name !== r.Name || v.order !== r.Order) return null
return v
},
[modelIndex, reach],
)
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const v = verdictOf(r)
if (!v || !v.unreachable || !v.shadowed_by) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
},
[verdictOf],
)
/**
* Whether a rule is IN FORCE, and who decided that — the two states this page
* used to conflate.
*
* `Rule.Enabled` from /api/config is the DESIRED state: what the operator saved,
* what the switch edits, what gets PUT back. The active WAN profile can override
* it in either direction, and then the desired state is no longer what the router
* is doing. Drawing the row from `Enabled` is what let a config with two rules
* `enabled '1'` show two live switches while the engine ran one chain.
*
* With no verdict — an older daemon, a stopped one, or one still describing the
* previous config — the desired state is all we know, so the row falls back to it
* and claims no profile rather than inventing one. The `typeof` guard is for the
* older daemon specifically: it answers without `effective_enabled` at all, and
* reading `undefined` as false would gray out every rule on the page.
*/
const forceOf = useCallback(
(r: RRule): RuleForce => {
const v = verdictOf(r)
if (!v || typeof v.effective_enabled !== 'boolean') {
return { on: !!r.Enabled, profile: null, dir: null }
}
return { on: v.effective_enabled, profile: v.overridden_by ?? null, dir: v.override ?? null }
},
[verdictOf],
)
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
const rulesets = useMemo<Ruleset[]>(
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
@@ -555,13 +654,17 @@ export default function Routing() {
// Delete the ruleset AND strip its name from any rule that referenced it, so no
// rule is left pointing at a matcher that no longer exists (one atomic persist).
const deleteRuleset = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const used = rulesetUsage.get(name) ?? 0
const warn = used
? `Delete ruleset "${name}"? It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
: `Delete ruleset "${name}"?`
if (!window.confirm(warn)) return
const ok = await confirm({
label: 'Delete ruleset',
title: `Delete ruleset "${name}"?`,
body: used
? `It'll be removed from ${used} rule${used === 1 ? '' : 's'} that match it.`
: undefined,
})
if (!ok) return
const nextRulesets = rulesets.filter((r) => r.Name !== name)
const nextRules = rules.map((r) => {
const cur = r.DstRuleset ?? []
@@ -572,7 +675,7 @@ export default function Routing() {
`ruleset ${name} deleted`,
)
},
[config, rules, rulesets, rulesetUsage, persist],
[config, rules, rulesets, rulesetUsage, persist, confirm],
)
const onToggle = useCallback(
@@ -612,17 +715,26 @@ export default function Routing() {
)
const onDelete = useCallback(
(name: string) => {
if (!window.confirm(`Delete rule "${name}"? Traffic it matched will fall through to the next rule.`)) return
async (name: string) => {
const ok = await confirm({
label: 'Delete rule',
title: `Delete rule "${name}"?`,
body: 'Traffic it matched will fall through to the next rule.',
})
if (!ok) return
commitRules(
rules.filter((r) => r.Name !== name),
`${name} deleted`,
)
},
[rules, commitRules],
[rules, commitRules, confirm],
)
// Insert a new rule just above the catch-all (so a specific rule can actually match).
// isCatchAll() is false for an unmigrated rule, which is the right answer here too:
// such a rule is held disabled and owns no default route, so there is nothing to
// insert ahead of — the new rule simply goes last, where a rule with no matchers
// does become the default.
const nextOrder = useCallback((): { order: number; bumpCatchAll?: { name: string; order: number } } => {
if (rules.length === 0) return { order: 10 }
const last = rules[rules.length - 1]
@@ -674,18 +786,15 @@ export default function Routing() {
return
}
setFormError(null)
const mv = form.matchValue.trim()
const rule: RRule = {
Name: name,
Enabled: true,
Order: 0,
Src: form.src,
// Domains are matched via rulesets only — the add form has no free-text
// domain matcher by design.
DstDomain: [],
// Destination = rulesets, always. Domains and addresses live in a
// `config ruleset` so one list serves every rule that needs it.
DstRuleset: form.rulesets,
DstIP: form.matchKind === 'ip' ? csv(mv) : [],
DstPort: form.matchKind === 'port' ? mv : '',
DstPort: form.port.trim(),
Proto: form.proto,
Target: form.target,
Egress: '',
@@ -756,7 +865,14 @@ export default function Routing() {
)
}
const enabledCount = rules.filter((r) => r.Enabled).length
// One force verdict per displayed rule, computed once and handed down — the row,
// the counter and the banner must all be reading the SAME answer.
const force = rules.map((r) => forceOf(r))
// EFFECTIVE, not configured. A counter that added up saved switches said "2 / 2
// active" for a config the router was running one rule of.
const enabledCount = force.filter((f) => f.on).length
const overridden = force.filter((f) => f.profile !== null)
const overrideProfile = overridden[0]?.profile ?? null
return (
<section className="page" aria-label="Routing rules">
@@ -765,12 +881,28 @@ export default function Routing() {
Rules run top to bottom on the bus — the <strong>first match wins</strong>. Traffic that
reaches the bottom follows the default route.
</p>
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules active`}>
<span className="rt-count mono" aria-label={`${enabledCount} of ${rules.length} rules in force`}>
{enabledCount}
<small> / {rules.length} active</small>
<small> / {rules.length} in force</small>
</span>
</div>
{/* Said once at the top, so the per-row badges below read as consequences of
one thing rather than as N unrelated oddities. Only shown when a profile
actually changed something: a router that uses no profiles, or one whose
profile agrees with every saved switch, gets no banner at all. */}
{overrideProfile && (
<p className="rt-prof-banner" role="status">
<span className="rt-prof-tag mono">profile</span>
<span>
<strong className="mono">{overrideProfile}</strong> is the active WAN profile and is
overriding {overridden.length === 1 ? '1 rule' : `${overridden.length} rules`} below. The
switches keep showing what you saved; the rows show what the router is running. Change
which rules a profile forces on the <strong>Profiles</strong> page.
</span>
</p>
)}
{actionError && (
<p className="page-error" role="alert">
{actionError}
@@ -789,7 +921,7 @@ export default function Routing() {
{rules.length === 0 ? (
<div className="rt-empty">
<p>No rules — all traffic follows the default route.</p>
<p className="rt-empty-sub">Add a rule below to steer a domain, address, or port.</p>
<p className="rt-empty-sub">Add a rule below to steer a destination list, source, or port.</p>
</div>
) : (
<ol className="rt-list" aria-label="Routing rules in first-match order">
@@ -818,6 +950,7 @@ export default function Routing() {
busy={saving}
editingOther={editingRule !== null}
shadow={shadowOf(r)}
force={force[i]}
onEdit={onEditRule}
onToggle={onToggle}
onMove={onMove}
@@ -868,6 +1001,7 @@ function RuleRow({
busy,
editingOther,
shadow,
force,
onEdit,
onToggle,
onMove,
@@ -880,6 +1014,9 @@ function RuleRow({
editingOther: boolean
/** Set when the daemon reports this rule can never fire; null when it can. */
shadow: { by: string; byOrder: number; reason: string } | null
/** Whether the rule is IN FORCE, and which profile decided that (see RuleForce).
* Every mark on this row comes from here; `rule.Enabled` drives only the switch. */
force: RuleForce
onEdit: (name: string) => void
onToggle: (name: string) => void
onMove: (name: string, dir: 'up' | 'down') => void
@@ -895,17 +1032,26 @@ function RuleRow({
// matched above" line): those are the claim that made two `default` rules
// indistinguishable in the first place.
const dead = shadow !== null
// The unmigrated-rule tripwire (see RRule.LegacyDst / model.IsCatchAll): this
// rule's destination is still in the removed schema-v1 options, so the daemon
// holds it disabled. It is NOT an ordinary disabled rule — nobody switched it
// off — so it gets the same "wired but not connected" amber treatment as a
// shadowed rule, plus a badge and a line saying what to run. isCatchAll()
// already refuses to call it the default, so `final` marks cannot land here.
const legacyDst = (rule.LegacyDst ?? []).filter(Boolean)
const unmigrated = legacyDst.length > 0
const inert = dead || unmigrated
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
const cls = [
'rt-rule',
rule.Enabled ? '' : 'off',
isDefault ? 'final' : '',
dead ? 'dead' : '',
]
// Dimmed by the EFFECTIVE state, never by the saved one. A rule the active
// profile switched off is not in force, and the row has to read that way even
// though its switch — which edits the saved setting — is still on.
const cls = ['rt-rule', force.on ? '' : 'off', isDefault ? 'final' : '', inert ? 'dead' : '']
.filter(Boolean)
.join(' ')
// What the switch says, spelled out, for the moment the two disagree.
const savedState = rule.Enabled ? 'on' : 'off'
return (
<li className={cls}>
@@ -919,7 +1065,7 @@ function RuleRow({
>
▲
</button>
<span className={isDefault ? 'rt-ord final' : dead ? 'rt-ord dead' : 'rt-ord'}>
<span className={isDefault ? 'rt-ord final' : inert ? 'rt-ord dead' : 'rt-ord'}>
{isDefault ? '·' : rule.Order}
</span>
<button
@@ -937,10 +1083,26 @@ function RuleRow({
<div className="rt-head">
<span className="rt-name">{rule.Name}</span>
{isDefault && <span className="rt-badge">default route · final</span>}
{unmigrated && <span className="rt-badge dead">held off · not migrated</span>}
{dead && <span className="rt-badge dead">never applies</span>}
{/* Amber for the rule the profile switched OFF (warn semantics: wired but
not connected), accent for the one it switched ON — orange is the
faceplate's active state, and a force-enabled rule is exactly that. */}
{force.dir === 'disabled' && <span className="rt-badge dead">off · by profile</span>}
{force.dir === 'enabled' && <span className="rt-badge prof-on">on · by profile</span>}
</div>
<div className="rt-match">
{dead ? (
{unmigrated ? (
// Why the rule is off and what fixes it. Same voice as the shadow note:
// state, cause, one command. The daemon says the same thing through the
// apply warnings (model.ValidateRules); this puts it on the row it is about.
<span className="rt-dead-note">
Destination still written the old way (<span className="mono">{legacyDst.join(', ')}</span>
) — this config was never migrated, so shaterd cannot read where this rule sends
traffic and holds it disabled. Run <span className="mono">shaterd migrate</span> on the
router to turn those entries into a ruleset, then enable the rule again.
</span>
) : dead ? (
// The badge says it never fires; this line says what beat it and what to
// do. Visible text, not a tooltip — the operator has to be able to find
// the other rule, and two rows can carry the same name.
@@ -955,6 +1117,22 @@ function RuleRow({
<Matchers rule={rule} />
)}
</div>
{/* Added BELOW the matchers, not instead of them: the rule's conditions are
still worth reading — the operator is deciding whether to change the
profile or the rule.
Two clauses only. The banner at the top of the page already carries the
general explanation and the way to change it, and a profile that
overrides several rules would otherwise repeat that paragraph on every
one of them. What is left is the part only this row can say: whether it
is in force, and what its own switch is showing instead. */}
{force.profile && (
<p className="rt-dead-note rt-prof-note">
{force.dir === 'disabled' ? 'Not in force' : 'In force'} — profile{' '}
<strong className="mono">{force.profile}</strong> switches this rule{' '}
{force.dir === 'disabled' ? 'off' : 'on'}. The switch still reads{' '}
<strong>{savedState}</strong>: that is the saved setting.
</p>
)}
</div>
<div className={`rt-target ${tone}`} title={`target: ${target}`}>
@@ -974,12 +1152,40 @@ function RuleRow({
>
Edit
</button>
<Toggle
pressed={rule.Enabled}
onChange={() => onToggle(rule.Name)}
label={`${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`}
disabled={frozen}
/>
{/* An unmigrated rule cannot be switched on from here, and the switch says
so rather than pretending: the daemon holds it disabled, but a config
write from the panel DROPS the unreadable legacy options (render.go
emits neither), so enabling it here would save a live rule with no
destination left at all — the catch-all this tripwire exists to
prevent. `shaterd migrate` clears LegacyDst and the switch comes back. */}
{/* The switch edits the SAVED setting and nothing else, so it keeps showing
rule.Enabled even while the active profile forces the opposite. Mirroring
the effective state here would be worse than the bug it replaces: the
operator would flip a switch that was never theirs, and the PUT would
write the profile's decision into UCI as if they had chosen it. The row
above says what the router is doing; the "saved" caption says what this
control is for. */}
<span className="rt-switch">
<Toggle
pressed={rule.Enabled}
onChange={() => onToggle(rule.Name)}
label={
unmigrated
? `Rule ${rule.Name} is held disabled until the config is migrated`
: force.profile
? `Saved setting for rule ${rule.Name} is ${savedState}; profile ${force.profile} is forcing it ${
force.dir === 'disabled' ? 'off' : 'on'
}. This switch changes the saved setting only.`
: `${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`
}
disabled={frozen || unmigrated}
/>
{force.profile && (
<span className="rt-switch-note" aria-hidden="true">
saved
</span>
)}
</span>
<button
type="button"
className="rt-del"
@@ -1012,9 +1218,7 @@ function Matchers({ rule }: { rule: RRule }): ReactNode {
)
}
listChip('src', rule.Src, 'src')
listChip('dns', rule.DstDomain, 'dom')
listChip('ruleset', rule.DstRuleset, 'rs')
listChip('ip', rule.DstIP, 'ip')
if (rule.DstPort && rule.DstPort.trim()) {
chips.push(
<span className="rt-chip" key="port">
@@ -1130,7 +1334,16 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
)
}
/** The dst_ruleset checkbox group. Renders nothing when no rulesets exist. */
/**
* The destination picker: which rulesets this rule matches (dst_ruleset).
*
* Checkboxes and nothing else. This is the ONLY way a rule names a destination,
* so it renders even when the config has no lists yet — an empty picker that says
* where lists come from is the honest answer, and hiding it would leave the rule
* form with no destination control at all. Building and filling a list is the
* Rulesets panel's job, deliberately kept out of the rule editor so a list is
* created in exactly one place.
*/
function RulesetPicker({
options,
selected,
@@ -1142,24 +1355,26 @@ function RulesetPicker({
busy: boolean
onToggle: (name: string) => void
}): ReactNode {
if (options.length === 0) return null
return (
<div className="rt-rsel">
<span className="rt-flabel">Match rulesets — dst_ruleset</span>
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
<span className="rt-flabel">Destination — dst_ruleset</span>
{options.length > 0 && (
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
)}
<p className="rt-rsel-hint">
The rule also matches any traffic in the checked list(s). Combine with a domain, address, or
port, or use a ruleset on its own.
{options.length === 0
? 'No rulesets yet. Add one under Rulesets below, then come back and check it here — a rule matches a destination through a ruleset only.'
: 'The rule matches traffic in ANY checked list. Narrow it further with a source, port or protocol.'}
</p>
</div>
)
@@ -1285,7 +1500,14 @@ function AddRule({
set('rulesets', form.rulesets.includes(n) ? form.rulesets.filter((x) => x !== n) : [...form.rulesets, n])
const toggleDay = (d: string) =>
set('schedDays', form.schedDays.includes(d) ? form.schedDays.filter((x) => x !== d) : [...form.schedDays, d])
const matchPlaceholder = form.matchKind === 'ip' ? '10.0.0.0/8, 100.64.0.0/10' : '443, 8080-8090'
// Same flag as the edit form, from the same predicate: no matcher = this rule
// asks to be route.Final. Whether it wins depends on what is last — it takes
// the default when there is no catch-all yet (or the last rule is conditional),
// and lands dead when there is one, because nextOrder() inserts ahead of it.
// Both are worth flagging, so the text states the fact and not the outcome.
// Shown, never blocking: a default route is a legal thing to write.
const noMatchers = formHasNoMatchers(form)
return (
<form className="rt-add" onSubmit={onSubmit} aria-label="Add a routing rule">
@@ -1318,32 +1540,17 @@ function AddRule({
</label>
<label className="rt-field">
<span className="rt-flabel">Match</span>
<select
<span className="rt-flabel">Port(s)</span>
<input
className="rt-input mono"
value={form.matchKind}
onChange={(e) => set('matchKind', e.target.value as MatchKind)}
>
<option value="none">rulesets only</option>
<option value="ip">ip / cidr</option>
<option value="port">port</option>
</select>
value={form.port}
onChange={(e) => set('port', e.target.value)}
placeholder="443, 8080-8090"
autoComplete="off"
spellCheck={false}
/>
</label>
{form.matchKind !== 'none' && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">{form.matchKind === 'port' ? 'Port(s)' : 'Address(es)'}</span>
<input
className="rt-input mono"
value={form.matchValue}
onChange={(e) => set('matchValue', e.target.value)}
placeholder={matchPlaceholder}
autoComplete="off"
spellCheck={false}
/>
</label>
)}
<label className="rt-field">
<span className="rt-flabel">Proto</span>
<select
@@ -1390,6 +1597,9 @@ function AddRule({
<Button type="submit" variant="primary" disabled={busy}>
{busy ? 'Saving…' : 'Add rule'}
</Button>
{noMatchers && !error && (
<span className="rt-edit-warn">no matchers — matches everything</span>
)}
{error && (
<span className="rt-add-error" role="alert">
{error}
@@ -1401,11 +1611,11 @@ function AddRule({
}
// --- edit-a-rule plate (inline, replaces the row it edits) ------------------
// Unlike AddRule, editing exposes all three destination matchers at once
// (Domain(s) / IP-CIDR(s) / Port) rather than a single Match picker — a real rule
// can carry several matcher kinds simultaneously and none may be silently dropped.
// The full original rule is spread into the result on save, so Order / Enabled /
// Kill / Egress (and anything else off-form) survive untouched.
// Same fields as AddRule, on purpose: a rule carries exactly one destination
// mechanism (rulesets) plus port/proto/source, so there is nothing an edit can
// reveal that the add form hides. The full original rule is spread into the
// result on save, so Order / Enabled / Kill / Egress (and anything else off-form)
// survive untouched.
function RuleEditForm({
initial,
names,
@@ -1425,8 +1635,6 @@ function RuleEditForm({
}) {
const [name, setName] = useState(initial.Name)
const [src, setSrc] = useState<string[]>([...(initial.Src ?? [])])
const [domain, setDomain] = useState((initial.DstDomain ?? []).join(', '))
const [ip, setIp] = useState((initial.DstIP ?? []).join(', '))
const [port, setPort] = useState(initial.DstPort ?? '')
const [proto, setProto] = useState(initial.Proto ?? '')
const [target, setTarget] = useState(effectiveTarget(initial))
@@ -1442,14 +1650,15 @@ function RuleEditForm({
const toggleDay = (d: string) =>
setSchedDays((cur) => (cur.includes(d) ? cur.filter((x) => x !== d) : [...cur, d]))
// This rule's destination may still be in the schema-v1 options (RRule.LegacyDst).
// The form cannot show or edit them — they are not fields any more — and saving
// writes the model back through render.go, which does not emit them. So a save
// here silently DISCARDS that destination. Say so before it happens; the fix is
// `shaterd migrate`, which converts them into a ruleset this form can check.
const legacyDst = (initial.LegacyDst ?? []).filter(Boolean)
// A rule with no matcher of any kind is a catch-all — legal, but worth flagging.
const noMatchers =
src.length === 0 &&
csv(domain).length === 0 &&
csv(ip).length === 0 &&
port.trim() === '' &&
rulesets.length === 0 &&
proto.trim() === ''
const noMatchers = formHasNoMatchers({ src, port, rulesets, proto })
const submit = (e: FormEvent) => {
e.preventDefault()
@@ -1471,8 +1680,6 @@ function RuleEditForm({
...initial,
Name: nm,
Src: src,
DstDomain: csv(domain),
DstIP: csv(ip),
DstPort: port.trim(),
DstRuleset: rulesets,
Proto: proto,
@@ -1491,7 +1698,15 @@ function RuleEditForm({
<form className="rt-add rt-edit-form" onSubmit={submit} aria-label={`Edit rule ${initial.Name}`}>
<div className="rt-add-hd">
<span className="rt-add-title">Edit {initial.Name}</span>
<span className="rt-add-sub">Empty a field to drop that matcher. Save, then Apply.</span>
<span className="rt-add-sub">
Uncheck a list or clear a field to drop that matcher. Save, then Apply.
</span>
{legacyDst.length > 0 && (
<span className="rt-edit-warn">
not migrated — saving discards its old destination ({legacyDst.join(', ')}); run
shaterd migrate first
</span>
)}
</div>
<div className="rt-fields">
@@ -1521,37 +1736,6 @@ function RuleEditForm({
/>
</label>
{/* Domains are matched via rulesets by design — this legacy field only
appears when the rule ALREADY carries free-text domains, so they
stay visible and clearable rather than silently preserved. */}
{(initial.DstDomain ?? []).length > 0 && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">Domain(s) — legacy</span>
<input
className="rt-input mono"
value={domain}
onChange={(e) => setDomain(e.target.value)}
placeholder="youtube.com, *.googlevideo.com"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
)}
<label className="rt-field rt-field-wide">
<span className="rt-flabel">IP / CIDR(s)</span>
<input
className="rt-input mono"
value={ip}
onChange={(e) => setIp(e.target.value)}
placeholder="10.0.0.0/8, 100.64.0.0/10"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
<label className="rt-field">
<span className="rt-flabel">Port(s)</span>
<input
+236
View File
@@ -704,6 +704,13 @@
.tg-test--bad .tg-test-msg {
color: var(--crit);
}
/* "Nothing measured this" is not a failure and must never be dressed as one: an
unlit lamp and the faintest text on the card, the same register the group
readout uses for its unmeasured state. */
.tg-test--none .tg-test-msg {
font-family: var(--font-sans);
color: var(--faint);
}
.tg-test--wait .tg-test-msg {
color: var(--amber);
}
@@ -1038,6 +1045,221 @@
}
/* ---- responsive ---- */
/* ---- chain hop rail ----
* The chain section's signature, and the one place this card spends any
* boldness: the path is drawn as a CONDUCTOR with a numbered lamp at each hop,
* and the conductor is SEVERED below the first hop that was probed and did not
* answer. A chain is a single series path, so the question is never "how many
* hops are green", it is "where does my traffic stop" — and a broken line answers
* that before a word has been read.
*
* The two marks carry two different facts and must not be conflated:
* - the LAMP is that hop's own measurement (good / warn / crit / unlit), and a
* hop below the break that answered keeps its green, because it did answer;
* - the CONDUCTOR is reachability through the path, which really does stop.
*
* Orange is untouched here. Semantics carry every colour, and everything that is
* not a lamp is groove-grey. No transitions and no animation anywhere in the
* rail, so there is nothing for reduced-motion to switch off. */
.ch-rail {
gap: 8px;
}
.ch-eyebrow {
font-family: var(--font-mono);
font-size: 9px;
letter-spacing: var(--track-label);
text-transform: uppercase;
color: var(--faint);
}
.ch-hops {
--ch-num: 1.8ch; /* the engraved hop number's gutter */
--ch-gap: 8px;
--ch-led: 10px; /* must match .led's width */
--ch-lampy: 14px; /* row top → lamp centre; the conductor's anchor */
/* x of the conductor: the number gutter, one gap, then the lamp's centre */
--ch-spine: calc(var(--ch-num) + var(--ch-gap) + var(--ch-led) / 2);
list-style: none;
margin: 0;
padding: 0;
}
.ch-hop {
position: relative;
display: grid;
grid-template-columns: var(--ch-num) var(--ch-led) minmax(0, 1fr);
column-gap: var(--ch-gap);
align-items: start;
}
/* the conductor — two halves per row, so a break lands on one link only */
.ch-hop::before,
.ch-hop::after {
content: '';
position: absolute;
left: var(--ch-spine);
width: 2px;
margin-left: -1px;
/* Brighter than a plain groove: this line IS the readout, and at groove
strength it disappeared into the panel and took the whole idea with it. */
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
}
.ch-hop::before {
top: 0;
height: calc(var(--ch-lampy) - var(--ch-led) / 2 - 3px);
}
.ch-hop::after {
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 3px);
bottom: 0;
}
/* Nothing feeds hop 1 from above, and nothing leaves the exit downward — the
path starts and ends inside this rail. */
.ch-hop--first::before {
display: none;
}
.ch-hop--exit::after {
bottom: auto;
height: 9px;
}
/* …the exit ends on a crossbar instead of trailing off: end of line. */
.ch-hop--exit .ch-socket {
position: relative;
}
.ch-hop--exit .ch-socket::after {
content: '';
position: absolute;
left: 50%;
transform: translateX(-50%);
top: calc(var(--ch-lampy) + var(--ch-led) / 2 + 12px);
width: 11px;
height: 2px;
background: color-mix(in srgb, var(--dim) 55%, var(--groove));
}
/* THE SEVER. Everything from the dead hop's outgoing link downward is drawn as a
broken conductor: unmistakably not-a-line at a glance, and unmistakably not a
colour, because a colour here would compete with the lamps that carry health. */
.ch-hop--dead::after,
.ch-hop--severed::before,
.ch-hop--severed::after {
background: repeating-linear-gradient(
to bottom,
color-mix(in srgb, var(--dim) 45%, var(--groove)) 0 3px,
transparent 3px 7px
);
}
.ch-num {
font-size: 10px;
line-height: calc(var(--ch-lampy) * 2);
text-align: right;
color: var(--faint);
}
.ch-socket {
display: flex;
align-items: center;
height: calc(var(--ch-lampy) * 2);
}
.ch-body {
min-width: 0;
/* Separates one hop from the next. The conductor runs through this space, so
too little of it and two hops read as one wrapped row. */
padding-bottom: 8px;
}
.ch-l1 {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 8px;
min-height: calc(var(--ch-lampy) * 2);
}
.ch-name {
font-size: 12px;
color: var(--ink);
overflow-wrap: anywhere;
}
.ch-hop--untested .ch-name {
color: var(--dim);
}
/* The exit marker is NEUTRAL on purpose. The config path above this rail tags its
exit green, which is free there — but in here green means "answering", and a
green badge on the last hop would read as a health claim about it. */
.ch-tag {
font-family: var(--font-mono);
font-size: 8.5px;
letter-spacing: 0.14em;
text-transform: uppercase;
color: var(--faint);
}
.ch-delay {
font-size: 11.5px;
font-weight: 700;
color: var(--ink);
}
.ch-quiet {
font-family: var(--font-sans);
font-size: 12px;
color: var(--faint);
}
.ch-age {
margin-left: auto;
font-size: 10.5px;
color: var(--faint);
white-space: nowrap;
}
/* The counters read exactly as they do on a group card — alive out of TESTED,
with the untested remainder as a quiet aside only when there is one. Same
register, same weights, deliberately not a second dialect. */
.ch-l2 {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 8px;
margin-top: 1px;
font-size: 11.5px;
letter-spacing: 0.02em;
color: var(--dim);
}
.ch-count {
font-size: 12px;
color: var(--dim);
white-space: nowrap;
}
.ch-count b {
font-size: 14px;
font-weight: 700;
color: var(--ink);
}
.ch-hop--dead .ch-count b {
color: var(--crit);
}
.ch-word {
font-size: 10.5px;
letter-spacing: 0.12em;
text-transform: uppercase;
color: var(--faint);
}
.ch-dead {
padding: 1px 6px;
border-radius: 4px;
background: color-mix(in srgb, var(--crit) 12%, transparent);
font-size: 10.5px;
color: var(--crit);
white-space: nowrap;
cursor: help;
}
.ch-rest {
font-size: 10.5px;
color: var(--faint);
}
.ch-now {
max-width: 28ch;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
font-size: 10.5px;
color: var(--dim);
}
@media (max-width: 640px) {
.tg-sec-hd {
flex-wrap: wrap;
@@ -1060,6 +1282,20 @@
.gh-now {
max-width: 100%;
}
/* On a phone the age stamp stops being pushed to a lonely right edge and just
joins the end of the hop's line; the selected node gets the full width
instead of an ellipsis it doesn't need. */
.ch-age {
margin-left: 0;
}
.ch-now {
max-width: 100%;
}
/* Every field of a hop wraps onto its own line at this width, so the gap
between hops has to grow with them or the rail reads as one block of text. */
.ch-body {
padding-bottom: 12px;
}
.gh-mems {
max-height: 260px;
}
+407 -99
View File
@@ -1,6 +1,6 @@
import './Targets.css'
import { Fragment, useCallback, useEffect, useMemo, useRef, useState } from 'react'
import { Button, Led, Toggle } from '../components'
import { Button, Led, Toggle, useConfirm } from '../components'
import type { LedVariant } from '../components'
import {
apply as apiApply,
@@ -22,6 +22,8 @@ import type {
GroupTestResult,
GroupTestStatus,
Chain,
ChainHealth,
ChainHopHealth,
Egress,
Interface,
Node,
@@ -161,18 +163,18 @@ function renameReferences(m: Model, kind: RefKind, from: string, to: string): Mo
}
/**
* The sentence a delete confirmation appends: what still points at this target,
* and what happens to it. Empty list ⇒ an explicit "nothing references it", so
* the operator can delete a stray with confidence instead of guessing.
* The body of a delete confirmation: what still points at this target, and what
* happens to it. Empty list ⇒ an explicit "nothing references it", so the
* operator can delete a stray with confidence instead of guessing.
*/
function refWarning(refs: RefSite[]): string {
if (refs.length === 0) return ' Nothing references it.'
if (refs.length === 0) return 'Nothing references it.'
const shown = refs.slice(0, 4).map((r) => r.label)
const more = refs.length - shown.length
const list = `${shown.join(', ')}${more > 0 ? `, and ${more} more` : ''}`
return refs.length === 1
? ` It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
: ` It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
? `It is referenced by ${list}, whose traffic will be blocked (an unresolved target never falls through to the default route).`
: `It is referenced by ${refs.length} places — ${list} — whose traffic will be blocked (an unresolved target never falls through to the default route).`
}
/**
@@ -436,14 +438,14 @@ const normalizeTest = (st: GroupTestStatus): GroupTestStatus => ({
})
/**
* How the header names the reach of a running exit test. The name matters more
* How the header names the reach of a running refresh pass. The name matters more
* than the number when there is only one: "auto" tells the operator which button
* they pressed; "1 target" tells them nothing they didn't already know.
*/
function scopeLabel(scope: string[], targetCount: number): string {
if (scope.length === 1) return scope[0]
if (scope.length === 0) return 'exits' // pre-scope daemon — say nothing false
return scope.length >= targetCount ? 'every exit' : `${scope.length} exits`
if (scope.length === 0) return 'targets' // pre-scope daemon — say nothing false
return scope.length >= targetCount ? 'every target' : `${scope.length} targets`
}
/** Which editor (add or edit-by-name) is open within a section. */
@@ -457,6 +459,7 @@ interface Opt {
// ---- page ------------------------------------------------------------------
export default function Targets() {
const confirm = useConfirm()
const [config, setConfig] = useState<Model | null>(null)
const [loadError, setLoadError] = useState<string | null>(null)
@@ -589,10 +592,12 @@ export default function Targets() {
[health],
)
// ---- group/chain exit test: how fast, through which node, out which address --
// The POST only kicks a run off, and a 2 s poll of the GET carries progress
// plus every result so far. One endpoint covers groups and chains alike:
// POST with a group or chain name tests that one; an empty name tests them all.
// ---- out-of-turn refresh: how fast, through which node, out which address ----
// The POST does NOT dial. It asks the observatory — the only thing in the daemon
// that measures anything, and it measures along the real dial path — to come
// round out of turn; a 2 s poll of the GET carries progress plus every reading
// so far. One endpoint covers groups and chains alike: POST with a name refreshes
// that one, an empty name refreshes them all.
const [gtest, setGtest] = useState<GroupTestStatus>(IDLE_TEST)
const [gtestErr, setGtestErr] = useState<string | null>(null)
const [polling, setPolling] = useState(false)
@@ -635,7 +640,7 @@ export default function Targets() {
void readTest().then((st) => {
if (!alive || !st || st.running) return
setPolling(false)
flash('Group test complete')
flash('Readings refreshed')
})
}, 2000)
return () => {
@@ -651,16 +656,16 @@ export default function Targets() {
if (r.started) {
setGtestErr(null)
setPolling(true)
flash(name ? `Testing ${name}…` : 'Testing every exit…')
flash(name ? `Refreshing ${name}…` : 'Refreshing every reading…')
void readTest()
} else if (r.reason === 'already running') {
setPolling(true) // pick up the run someone else started
flash('A group test is already running')
setPolling(true) // pick up the pass someone else started
flash('The prober is already refreshing')
} else {
flash(`Couldn’t start the test — ${r.reason || 'the daemon refused it'}`)
flash(`Couldn’t ask for a refresh — ${r.reason || 'the daemon refused it'}`)
}
} catch (e) {
flash(`Couldn’t start the test — ${errText(e)}`)
flash(`Couldn’t ask for a refresh — ${errText(e)}`)
}
},
[flash, readTest],
@@ -787,13 +792,18 @@ export default function Targets() {
)
const removeGroup = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'group', name)
if (!window.confirm(`Delete group “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete group',
title: `Delete group “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Groups: groups.filter((g) => g.Name !== name) }, `Deleted ${name}`)
},
[config, groups, save],
[config, groups, save, confirm],
)
// ---- chain mutations ------------------------------------------------------
@@ -820,13 +830,18 @@ export default function Targets() {
)
const removeChain = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'chain', name)
if (!window.confirm(`Delete chain “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete chain',
title: `Delete chain “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Chains: chains.filter((c) => c.Name !== name) }, `Deleted ${name}`)
},
[config, chains, save],
[config, chains, save, confirm],
)
// ---- egress mutations -----------------------------------------------------
@@ -853,13 +868,18 @@ export default function Targets() {
)
const removeEgress = useCallback(
(name: string) => {
async (name: string) => {
if (!config) return
const refs = findReferences(config, 'egress', name)
if (!window.confirm(`Delete egress “${name}”?${refWarning(refs)}`)) return
const ok = await confirm({
label: 'Delete egress',
title: `Delete egress “${name}”?`,
body: refWarning(refs),
})
if (!ok) return
void save({ ...config, Egresses: egresses.filter((e) => e.Name !== name) }, `Deleted ${name}`)
},
[config, egresses, save],
[config, egresses, save, confirm],
)
const busy = saving || applying
@@ -894,10 +914,11 @@ export default function Targets() {
<h2 className="tg-sec-title">Groups</h2>
<span className="tg-sec-count mono">{groups.length} configured</span>
{/* The observatory's background probing is invisible by design — it
keeps every used group's and chain's numbers fresh on its own. The
one manual run left is the exit test: it is scoped to the groups
and chains it names, so its progress says WHICH, and its badge
lands only on those cards. */}
keeps every used group's and chain's numbers fresh on its own, along
the path traffic actually takes. The one manual control left does
not measure anything itself: it asks that prober to come round out
of turn. It is scoped to the groups and chains it names, so its
progress says WHICH, and its badge lands only on those cards. */}
<div className="tg-sec-ctl">
{groupHealthOn && (
<>
@@ -905,11 +926,11 @@ export default function Targets() {
<span
className="tg-run tg-run--exit"
role="status"
title="An exit test sends one connection through each group or chain it covers and reports the delay and the address the internet sees."
title="The background prober is measuring the targets this refresh covers, along the path each one's traffic really takes."
>
<Led variant="amber" pulse />
<span className="tg-run-what">
exit test · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
refreshing · {scopeLabel(asArray(gtest.scope), groups.length + chains.length)}
</span>
<span className="tg-run-n mono">
{gtest.done}/{gtest.total}
@@ -919,9 +940,9 @@ export default function Targets() {
<Button
onClick={() => void runTest()}
disabled={busy || !config || (groups.length === 0 && chains.length === 0) || gtest.running}
title="Send one connection through each group and chain and report the delay and the exit address the internet sees"
title="Ask the background prober to measure every group and chain out of turn, then show what it measured. The panel opens no connection of its own."
>
{gtest.running ? 'Testing…' : 'Test every exit'}
{gtest.running ? 'Refreshing…' : 'Refresh every reading'}
</Button>
</>
)}
@@ -941,6 +962,13 @@ export default function Targets() {
dials out through a tunnel measures them through that tunnel, so the same node can be alive
in one group and dead in another.
</p>
<p className="tg-sec-note">
One thing measures, and the panel is not it. A background prober walks every path your
rules use — hop by hop, exactly as traffic goes — and every number on this page is a read
of what it found. <strong>Refresh every reading</strong> asks it to come round out of turn
instead of waiting for the next pass; it opens no connection of its own, so a target no
rule routes through has nothing to report and says so.
</p>
{groupHealthOn && healthErr && (
<p className="tg-test-err" role="alert">
@@ -951,7 +979,7 @@ export default function Targets() {
{groupHealthOn && gtestErr && (
<p className="tg-test-err" role="alert">
Couldn’t read the test results — {gtestErr}.{' '}
Couldn’t read the refreshed numbers — {gtestErr}.{' '}
<button className="linkish" onClick={() => void readTest()}>
Retry
</button>
@@ -1094,7 +1122,9 @@ export default function Targets() {
chain={c}
busy={busy}
showHealth={groupHealthOn}
used={healthByChain.get(c.Name)?.used}
// The whole chain health record, not just `.used` — the card
// renders the observatory's per-hop measurements from it.
health={healthByChain.get(c.Name)}
test={testByGroup.get(c.Name)}
// The badge is this card's business only when the run names it.
testing={gtest.running && testScope.has(c.Name)}
@@ -1216,7 +1246,7 @@ function GroupRow({
group: Group
busy: boolean
/** Group health checks are on (Settings). When false, the card drops its health
* readout, its exit-test readout and its Test button — it is config only. */
* readout, its end-to-end reading and its Refresh button — it is config only. */
showHealth: boolean
/** This group's membership health, or undefined when the engine hasn't built
* it (not applied yet, or dropped for having no usable members). */
@@ -1226,14 +1256,14 @@ function GroupRow({
healthKnown: boolean
test?: GroupTestResult
/**
* A group exit test covering THIS group is in flight.
* A refresh pass covering THIS group is in flight.
*
* Deliberately not "a test is running": the caller resolves it against the run's
* scope. There is no per-card equivalent for the health run — that one measures
* every group at once and is reported once, in the section header.
*/
testing: boolean
/** Any exit test is in flight; the daemon runs one at a time. */
/** Any refresh pass is in flight; the daemon runs one at a time. */
testBusy: boolean
onTest: () => void
onEdit: () => void
@@ -1293,7 +1323,11 @@ function GroupRow({
health={health}
healthKnown={healthKnown}
/>
<GroupTestReadout test={test} pending={testing && !test} />
<GroupTestReadout
test={test}
pending={testing && !test}
hideAbsence={health?.used === false}
/>
</>
)}
</div>
@@ -1304,7 +1338,7 @@ function GroupRow({
editLabel={`Edit group ${group.Name}`}
deleteLabel={`Delete group ${group.Name}`}
onTest={showHealth ? onTest : undefined}
testLabel={showHealth ? `Test the exit of group ${group.Name}` : undefined}
testLabel={showHealth ? `Refresh the reading for group ${group.Name}` : undefined}
testDisabled={testBusy}
/>
</li>
@@ -1363,21 +1397,7 @@ function GroupHealthReadout({
// its members would stay "untested" forever. That is a fact about the ROUTING
// CONFIG, not about the members — so instead of counters that could only ever
// read as a permanent unknown, the card says so, quietly: unused, not unwell.
if (!health.used) {
return (
<div className="gh gh--unused">
<div className="gh-line">
<span
className="gh-unused"
title="No enabled rule routes through this group, so its members are not probed. Add it to a rule to see health."
>
unused
</span>
<span className="gh-quiet">not probed — no enabled rule routes through this group</span>
</div>
</div>
)
}
if (!health.used) return <NotRoutedNote kind="group" />
const v = verdictOf(health)
const { total, tested, alive, dead, untested } = health
@@ -1540,6 +1560,56 @@ function GroupHealthReadout({
)
}
/**
* The card's answer when NOTHING ROUTES THROUGH THIS TARGET. Shared by the group
* card and the chain card, because it is the same misunderstanding on both.
*
* It has to carry two statements, and the old one-liner ("not probed — no enabled
* rule routes through this group") only carried the first. Read fast it still
* landed as a verdict: a card that normally shows health and today shows a grey
* pill reads as "the health is bad". So the two meanings are now separated, on
* purpose and in this order:
*
* 1. the ROUTING FACT — nothing routes here, so nothing measures it;
* 2. the NON-FACT — this is not a health reading at all. Absent numbers are
* absence of measurement, never failure.
*
* For a GROUP there is a third line, and it is the confusion this whole change
* exists to end: a group used only as a hop inside a chain is never routed to
* DIRECTLY, so it correctly reads unused here while carrying real traffic as a
* hop. Its health is measured at that hop, on the chain's card.
*
* Unused is neutral — groove-grey, never amber, never crit. It is a state of the
* config, and the config is not sick.
*/
function NotRoutedNote({ kind }: { kind: 'group' | 'chain' }) {
return (
<div className="gh gh--unused">
<div className="gh-line">
<span className="gh-unused">unused</span>
<span className="gh-quiet">
No enabled rule routes through this {kind}, so the observatory never probes it.
</span>
</div>
<p className="gh-say">
That is a routing fact, not a health reading. There are no numbers here because nothing
measured this {kind} — not because it failed.
</p>
{kind === 'group' ? (
<p className="gh-say">
A group used only as a hop inside a chain reads unused here on purpose: the rules point at
the chain, not at the group. Its members are measured at that hop, so its real health is on
that chain’s card, hop by hop.
</p>
) : (
<p className="gh-say">
Point a rule at this chain and the observatory starts measuring every hop within seconds.
</p>
)}
</div>
)
}
/**
* One group's member rows, fetched on demand.
*
@@ -1645,7 +1715,9 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
}
/**
* What a group test found, in the four states it actually has.
* What the OBSERVATORY measured for this target end to end, in the four states it
* actually has. Nothing here was dialled by the panel — it is a read of the
* background prober's own measurement along the real path.
*
* The one worth spelling out: `ok` with an EMPTY `exit_ip` is a SUCCESS. The
* delay was measured; only the address lookup came back empty. Rendering that as
@@ -1653,15 +1725,45 @@ function MemberRow({ member }: { member: GroupMemberHealth }) {
* traffic, so it reads as a result with the address slot marked unknown — dim,
* not red, and the LED stays green.
*/
function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending: boolean }) {
/**
* Errors that mean NO MEASUREMENT EXISTS, as opposed to "this target is broken".
*
* Three of the observatory's four failure reasons are about the observatory, not
* about the path: nothing routes here, nothing has reached it yet, or background
* probing is switched off. Painting those crit-red — which is what `ok:false`
* used to buy you — reports a fault that nobody has found, on a target that may
* be carrying traffic perfectly. Only "the observatory's probe through this path
* failed" is a health finding, and it is deliberately NOT in this list.
*
* Matched on a stable fragment rather than the whole sentence, so a daemon that
* rewords the tail still classifies. An error we don't recognise stays red: an
* unknown failure is likelier to be real than not, and that is the safe default.
*/
const NO_MEASUREMENT = [
'not routed by any enabled rule',
'has not reached this target yet',
'background probing is disabled',
]
const isAbsence = (err: string): boolean => NO_MEASUREMENT.some((frag) => err.includes(frag))
function GroupTestReadout({
test,
pending,
hideAbsence,
}: {
test?: GroupTestResult
pending: boolean
/** The card already explains why nothing measures this target (the unused
* note), so an absence error here would just say it a second time. */
hideAbsence?: boolean
}) {
if (pending) {
// "testing", never "measuring": the health run owns that word and covers every
// group at once. Two runs that read the same on a card is how one group's test
// came to look like all four were busy.
// Names who is working and on what: the prober, on this target. The badge is
// scoped to the cards the run covers, so it can say "this one" honestly.
return (
<div className="tg-test tg-test--wait" role="status">
<Led variant="amber" pulse />
<span className="tg-test-msg">testing this exit…</span>
<span className="tg-test-msg">waiting for the prober to measure this…</span>
</div>
)
}
@@ -1670,10 +1772,22 @@ function GroupTestReadout({ test, pending }: { test?: GroupTestResult; pending:
const at = test.tested_unix ? fmtClock(test.tested_unix) : ''
if (!test.ok) {
// No measurement exists. Unlit lamp, quiet text: this panel's way of saying
// "no verdict", which is precisely the state — never a red one.
if (isAbsence(test.error)) {
if (hideAbsence) return null
return (
<div className="tg-test tg-test--none" role="status">
<Led variant="off" />
<span className="tg-test-msg">{test.error}</span>
{at && <span className="tg-test-at mono">{at}</span>}
</div>
)
}
return (
<div className="tg-test tg-test--bad" role="status">
<Led variant="crit" />
<span className="tg-test-msg">{test.error || 'the test failed'}</span>
<span className="tg-test-msg">{test.error || 'the probe failed'}</span>
{at && <span className="tg-test-at mono">{at}</span>}
</div>
)
@@ -2098,7 +2212,7 @@ function ChainRow({
chain,
busy,
showHealth,
used,
health,
test,
testing,
testBusy,
@@ -2109,19 +2223,19 @@ function ChainRow({
chain: Chain
busy: boolean
/** Group health checks are on (Settings). When false, the card drops its
* exit-test readout and Test button — it is config only. */
* health readout and Refresh button — it is config only. */
showHealth: boolean
/** This chain's reachability (GroupHealth.Used's chain analogue, plan §5.E).
* undefined ⇒ the health endpoint hasn't reported this chain (not applied yet, or
* a daemon version without chains): no badge. false ⇒ no enabled rule routes
* through the chain, so the observatory never probes it and the card renders
* "unused" instead of an exit-test readout. */
used?: boolean
/** Everything the observatory knows about this chain: whether any enabled rule
* routes through it, and the per-hop measurements along it.
* undefined ⇒ the health endpoint hasn't reported this chain at all (not
* applied yet, or a daemon version without chains): the card says nothing
* rather than guessing. */
health?: ChainHealth
test?: GroupTestResult
/** An exit test covering THIS chain is in flight (the caller resolves it
/** A refresh pass covering THIS chain is in flight (the caller resolves it
* against the run's scope, exactly as for a group card). */
testing: boolean
/** Any exit test is in flight; the daemon runs one at a time. */
/** Any refresh pass is in flight; the daemon runs one at a time. */
testBusy: boolean
onTest: () => void
onEdit: () => void
@@ -2165,25 +2279,21 @@ function ChainRow({
{showHealth && (
<>
{/* A chain no enabled rule routes through is never probed (the
observatory walks only reachable paths), so instead of an exit-test
readout the card says so, quietly — the same "unused" pattern the
group card uses (GroupHealthReadout), not a new design. `used` is
undefined until the health endpoint reports this chain (or from a
daemon version without chains): no badge then. */}
{used === false && (
<div className="gh gh--unused">
<div className="gh-line">
<span
className="gh-unused"
title="No enabled rule routes through this chain, so its exit is not probed. Add it to a rule to see health."
>
unused
</span>
<span className="gh-quiet">not probed — no enabled rule routes through this chain</span>
</div>
</div>
)}
<GroupTestReadout test={test} pending={testing && !test} />
observatory walks only reachable paths), so instead of a health
readout the card says so — the same "unused" note the group card
uses, not a new design. `health` is undefined until the endpoint
reports this chain (or on a daemon without chains): say nothing
then rather than guess. */}
{health?.used === false ? (
<NotRoutedNote kind="chain" />
) : health?.used ? (
<ChainHopRail chain={chain.Name} defs={hops} hops={health.hops} />
) : null}
<GroupTestReadout
test={test}
pending={testing && !test}
hideAbsence={health?.used === false}
/>
</>
)}
</div>
@@ -2194,13 +2304,210 @@ function ChainRow({
editLabel={`Edit chain ${chain.Name}`}
deleteLabel={`Delete chain ${chain.Name}`}
onTest={showHealth ? onTest : undefined}
testLabel={showHealth ? `Test the exit of chain ${chain.Name}` : undefined}
testLabel={showHealth ? `Refresh the reading for chain ${chain.Name}` : undefined}
testDisabled={testBusy}
/>
</li>
)
}
// ---- chain hop rail --------------------------------------------------------
/**
* The API gives hops an index and no name. The page already knows the names — the
* model's own `Hops` strings ("egress:ewan", "node:awgout", "group:sub0") — so
* zip the two by POSITION.
*
* Two things make that safe rather than clever. A LEADING `egress:` is not a
* numbered hop: the daemon lifts it into hop 1's entry detour, so it is dropped
* before counting. And if the counts still disagree — a chain that splices
* sub-chains gets FLATTENED by the daemon, producing more wire hops than the
* config lists — every label is dropped. A hop labelled with its neighbour's name
* is worse than a hop with no name at all: it would send someone to fix the wrong
* target.
*/
function hopLabels(defs: string[], hops: ChainHopHealth[]): (string | undefined)[] {
const numbered = defs.length > 0 && defs[0].startsWith('egress:') ? defs.slice(1) : defs
if (numbered.length !== hops.length) return hops.map(() => undefined)
return hops.map((h) => (h.index >= 1 && h.index <= numbered.length ? numbered[h.index - 1] : undefined))
}
/**
* One hop's lamp.
*
* `dead` is crit and `untested` is an UNLIT socket — never red, because nothing
* has been measured and an unlit lamp is this panel's way of saying "no verdict".
* The fourth case is the page's existing house reading, applied here for
* consistency rather than invented: a group hop that is carrying traffic but has
* confirmed failures on its board is amber. `state` stays the daemon's word for
* "can this hop carry traffic"; the amber only qualifies HOW WELL.
*/
function hopLed(h: ChainHopHealth): LedVariant {
if (h.state === 'dead') return 'crit'
if (h.state === 'untested') return 'off'
return h.dead > 0 ? 'amber' : 'on'
}
/**
* What the observatory measured at each position of a chain — the reading the
* daemon always took and the panel never showed.
*
* THE DESIGN RISK, and the one place this card spends any boldness: the rail
* draws the CONDUCTOR as well as the lamps, and severs it below the first dead
* hop. A chain is a single series path, so the operator's real question is never
* "how many hops are green" — it is "where does my traffic stop". Four lamps in a
* column answer the first question and leave the second to arithmetic. A broken
* conductor answers the second one before you have read a single word, which is
* the whole reason this feature exists.
*
* It stays honest by keeping the two facts on two different marks. Each LAMP is
* that hop's own measurement and never changes because of a hop in front of it —
* a hop after the break that answers still shows green, because it really did
* answer. The CONDUCTOR is reachability through the path, and that genuinely does
* stop at the break. Nothing here is re-derived from the daemon's counters; the
* only thing the panel adds is what a chain structurally is.
*
* Everything around the rail is deliberately quiet: no colour but the semantic
* lamps, no motion at all, the orange accent untouched.
*/
function ChainHopRail({
chain,
defs,
hops,
}: {
chain: string
/** The chain's configured hops, straight off the model — the only source of names. */
defs: string[]
/** Absent ⇒ the engine never materialised per-hop outbounds. NOT "no hops". */
hops?: ChainHopHealth[]
}) {
const ordered = useMemo(() => [...asArray(hops)].sort((a, b) => a.index - b.index), [hops])
const labels = useMemo(() => hopLabels(defs, ordered), [defs, ordered])
// The first hop that was probed and did not answer. Everything after it is
// unreachable THROUGH THIS CHAIN, whatever its own lamp says. `untested` is
// never a break: nothing was measured, so nothing is known to be severed.
const breakAt = ordered.findIndex((h) => h.state === 'dead')
if (ordered.length === 0) {
// Say why, in one line, instead of an empty rail. The daemon collapses a
// single-target chain into a plain alias and never builds copies to measure,
// so we can tell the two absences apart from the config alone.
const numbered = defs.filter((d, i) => !(i === 0 && d.startsWith('egress:')))
return (
<div className="gh gh--absent">
<span className="gh-absent-msg">
{numbered.length <= 1
? 'This chain has a single hop, so the engine points traffic straight at that target instead of building a path to measure. Its health is on that target’s own card.'
: 'The engine hasn’t built this chain’s hops yet, so there is nothing measured per hop. They appear once it is running with this config applied.'}
</span>
</div>
)
}
return (
<div className="gh ch-rail">
<span className="ch-eyebrow">measured, hop by hop</span>
<ol className="ch-hops">
{ordered.map((h, i) => {
const label = labels[i]
const severed = breakAt >= 0 && i > breakAt
const cls = [
'ch-hop',
`ch-hop--${h.state}`,
severed ? 'ch-hop--severed' : '',
h.exit ? 'ch-hop--exit' : '',
i === 0 ? 'ch-hop--first' : '',
]
.filter(Boolean)
.join(' ')
const age = fmtAge(h.age_seconds)
return (
<li key={h.tag || h.index} className={cls}>
<span className="ch-num mono" aria-hidden="true">
{h.index}
</span>
<span className="ch-socket">
<Led variant={hopLed(h)} />
</span>
<div className="ch-body">
<div className="ch-l1">
<span className="ch-name mono" title={`engine outbound ${h.tag}`}>
{label ?? (h.kind === 'group' ? 'a group hop' : 'a node hop')}
</span>
{h.exit && <span className="ch-tag">exit</span>}
{h.state === 'alive' && h.delay_ms > 0 && (
<span className="ch-delay mono">{h.delay_ms} ms</span>
)}
{h.state === 'untested' && <span className="ch-quiet">not measured yet</span>}
{age && <span className="ch-age mono">{age}</span>}
</div>
{/* A node hop IS its own measurement (total 1), so counters would
only restate the lamp. A group hop rolls up its per-hop member
copies, and those read exactly as they do everywhere else in
this app: alive out of TESTED, with the untested remainder as a
quiet aside only when there is one. */}
{h.kind === 'group' && h.total > 0 && (
<div className="ch-l2">
{h.tested === 0 ? (
<span className="ch-rest mono">
{h.total} member{h.total === 1 ? '' : 's'}, none measured
</span>
) : (
<>
<span className="ch-count mono">
<b>{h.alive}</b> / {h.tested}
</span>
<span className="ch-word">alive</span>
{h.dead > 0 && (
<span
className="ch-dead mono"
title={`${h.dead} member${h.dead === 1 ? '' : 's'} were probed at this hop and did not answer`}
>
{h.dead} not answering
</span>
)}
{h.untested > 0 && (
<span className="ch-rest mono">
tested {h.tested} of {h.total}
</span>
)}
</>
)}
{h.selected && (
<span
className="ch-now mono"
title={`Traffic crossing hop ${h.index} of “${chain}” is on ${h.selected}`}
>
now → {h.selected}
</span>
)}
</div>
)}
</div>
</li>
)
})}
</ol>
{/* The sentence the rail's shape implies, written out — because the break is
the answer someone came here for, and a graphic alone should never be the
only place a finding exists. */}
{breakAt >= 0 && (
<p className="gh-say gh-say--bad">
Hop {ordered[breakAt].index}
{labels[breakAt] ? ` (${labels[breakAt]})` : ''} was probed and did not answer. A chain is
one path, so traffic stops there
{breakAt < ordered.length - 1
? ' — the hops after it answer on their own, but nothing reaches them through this chain.'
: '.'}
</p>
)}
</div>
)
}
function ChainEditor({
initial,
hopOptions,
@@ -2675,8 +2982,8 @@ function RowActions({
busy: boolean
editLabel: string
deleteLabel: string
// Only groups and chains can be tested, so the control is optional and absent
// everywhere else rather than a disabled stub on every row.
// Only groups and chains are probed, so the refresh control is optional and
// absent everywhere else rather than a disabled stub on every row.
onTest?: () => void
testLabel?: string
testDisabled?: boolean
@@ -2689,8 +2996,9 @@ function RowActions({
onClick={onTest}
disabled={busy || testDisabled}
aria-label={testLabel}
title="Ask the background prober to measure this target out of turn. It does not open a connection from the panel."
>
Test
Refresh
</Button>
)}
<Button className="tg-act" onClick={onEdit} disabled={busy} aria-label={editLabel}>
+150
View File
@@ -0,0 +1,150 @@
// protectionState — the one sentence the whole panel shows about "am I protected".
//
// Run with `npm test` (node's built-in test runner + native TypeScript stripping;
// no test dependency is added to the SPA, which ships inside the daemon binary).
//
// The case this file was written for is "plane full, traffic direct": the exact
// state of a live router — one enabled rule, `default → direct`, no groups, no
// rule-sets — where every part of the data plane was installed and the readout
// therefore said "Protected — traffic from your network is going through the
// tunnel", under a green LED, while the whole LAN went out the plain WAN.
//
// planeState.ts has no runtime imports (both of its imports are `import type`),
// so this runs against the real module with nothing stubbed.
import { test } from 'node:test'
import assert from 'node:assert/strict'
import { protectionState } from './planeState.ts'
import type { Status, Traffic } from './api.ts'
/** A healthy, fully-installed router; `traffic` is what each case varies. */
function status(over: Partial<Status> = {}): Status {
return {
running: true,
enabled: true,
active: true,
table: true,
hash: 'abc',
version: '1.11.0-shater',
kill_switch: 'closed',
engine_running: true,
plane: 'full',
warnings: [],
...over,
}
}
function withTraffic(traffic: Traffic | undefined): Status {
return status({ traffic })
}
// --- the field case ---------------------------------------------------------
test('plane full + default direct is NOT reported as protected', () => {
const s = protectionState(withTraffic({ verdict: 'direct', default: 'direct', tunnel_rules: 0 }))
assert.notEqual(s.headline, 'Protected')
assert.equal(s.variant, 'crit')
assert.equal(s.alarm, true)
// The claim that was false must not survive anywhere in the copy.
assert.doesNotMatch(s.detail, /going through the tunnel/)
// ...and the honest consequence must be stated, not implied.
assert.match(s.detail, /real address/)
})
// --- the other verdicts under a full plane ----------------------------------
test('plane full + default into a tunnel is protected', () => {
const s = protectionState(withTraffic({ verdict: 'tunnel', default: 'auto', tunnel_rules: 1 }))
assert.equal(s.variant, 'on')
assert.equal(s.headline, 'Protected')
assert.equal(s.alarm, false)
})
test('plane full + direct default with tunnelling rules is split, not protected', () => {
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 3 }))
assert.equal(s.variant, 'amber')
assert.notEqual(s.headline, 'Protected')
// Says how much is protected, and that the default is not.
assert.match(s.detail, /3 rules/)
assert.match(s.detail, /normal internet connection/)
// A working selective setup must not raise a banner on every other page.
assert.equal(s.alarm, false)
})
test('split names a single rule in the singular', () => {
const s = protectionState(withTraffic({ verdict: 'split', default: 'direct', tunnel_rules: 1 }))
assert.match(s.detail, /^One rule sends traffic/)
})
test('plane full + blocked default with rules leaks nothing and is never crit', () => {
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 2 }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, false)
assert.match(s.detail, /nothing is leaving unprotected/)
})
test('plane full + blocked default with no rules says the network has no way out', () => {
const s = protectionState(withTraffic({ verdict: 'blocked', default: 'block', tunnel_rules: 0 }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
assert.doesNotMatch(s.detail, /going through the tunnel/)
})
test('plane full with no verdict claims nothing either way', () => {
for (const t of [undefined, { verdict: '' as const }]) {
const s = protectionState(withTraffic(t))
assert.notEqual(s.headline, 'Protected')
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, false)
}
})
// --- the branches that were already correct ---------------------------------
test('no status yet', () => {
const s = protectionState(null)
assert.equal(s.variant, 'off')
assert.equal(s.alarm, false)
})
test('service switched off is a deliberate state, not a fault', () => {
const s = protectionState(status({ enabled: false }))
assert.equal(s.variant, 'amber')
assert.equal(s.headline, 'Turned off')
assert.equal(s.alarm, false)
})
test('hold: the kill-switch caught it — protected, offline', () => {
const s = protectionState(status({ plane: 'hold', engine_running: false, active: false }))
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
assert.match(s.headline, /blocked/)
})
test('none + fail-closed is the leak, and it is crit', () => {
const s = protectionState(status({ plane: 'none', table: false, engine_running: false }))
assert.equal(s.variant, 'crit')
assert.equal(s.alarm, true)
})
test('none + fail-open is the operator’s documented choice, stated not alarmed at', () => {
const s = protectionState(
status({ plane: 'none', table: false, engine_running: false, kill_switch: 'open' }),
)
assert.equal(s.variant, 'amber')
assert.equal(s.alarm, true)
})
test('daemon too old to send `plane` keeps its own fallback', () => {
// Nothing here may depend on `traffic`: a daemon with no `plane` has no
// `traffic` either, and this branch reads what it can observe instead.
const { plane, ...noPlane } = status()
void plane
assert.equal(protectionState(noPlane as Status).headline, 'Protected')
assert.equal(protectionState({ ...noPlane, running: false } as Status).headline, 'Service stopped')
assert.equal(
protectionState({ ...noPlane, active: false } as Status).headline,
'Starting up',
)
})
+104 -7
View File
@@ -11,7 +11,7 @@
// same router differently.
import type { LedVariant } from './components'
import type { Status } from './api'
import type { Status, Traffic } from './api'
export interface ProtectionState {
variant: LedVariant
@@ -31,6 +31,24 @@ export interface ProtectionState {
* hold — the kill-switch caught it. Protected, but offline.
* none (fail-closed) — there is no protection at all. Online, and exposed.
* Collapsing them would erase the only difference that matters.
*
* PLANE IS NOT THE WHOLE ANSWER, AND THAT USED TO BE A LIE. `plane: 'full'`
* returned "Protected — traffic from your network is going through the tunnel",
* which is a claim `plane` cannot support: it only says the nft table, the policy
* routing and the engine are all installed. Where the diverted packets go once the
* engine has them is decided by the engine's default route, and a router in the
* field ran with one rule — `default → direct`, no groups, no rule-sets. Fully
* installed plane, zero tunnel, whole LAN out the plain WAN with its real address,
* green LED, "Protected".
*
* So `full` now branches on `status.traffic`, the daemon's verdict on the config
* it is actually running (see api.ts TrafficVerdict). It is computed on the daemon
* because only the daemon knows what was GENERATED and STARTED: the panel's
* /api/config is desired state, which diverges from the running one whenever edits
* are unapplied or a rollback is pending, and re-deriving the default route from it
* would mean a second implementation of the generator's rule loop — schedules,
* shadowed catch-alls, targets that failed to resolve and fell back to direct —
* drifting against the first.
*/
export function protectionState(status: Status | null): ProtectionState {
if (!status) {
@@ -56,12 +74,7 @@ export function protectionState(status: Status | null): ProtectionState {
switch (status.plane) {
case 'full':
return {
variant: 'on',
headline: 'Protected',
detail: 'Traffic from your network is going through the tunnel.',
alarm: false,
}
return fullPlaneState(status.traffic)
case 'hold':
return {
variant: 'amber',
@@ -112,3 +125,87 @@ export function protectionState(status: Status | null): ProtectionState {
alarm: false,
}
}
/**
* The plane is fully installed — now say where the traffic it carries ends up.
*
* Only `tunnel` earns "Protected". The other verdicts each describe a real router
* someone can be sitting in front of, and they are kept apart because the thing to
* DO about them differs:
*
* split — deliberate for most people who reach it, accidental for the rest
* (a default rule that was never pointed anywhere). Amber, but no
* alarm: raising a banner on every page of a working selective setup
* is how a banner stops being read.
* direct — the engine is running and forwarding every connection out the plain
* WAN. The traffic outcome is identical to `plane: 'none'` under a
* closed kill-switch, so it gets the same weight: crit, and it
* interrupts. The wording differs because the fix does — nothing
* failed here, the routing simply says "direct".
* blocked — the fail-closed default. Nothing is leaking, so this is never crit;
* with no tunnelling rules at all it means the network has no way out
* and someone should be told why.
* unknown — an older daemon, or the seconds between this daemon starting and its
* first apply. We do not know, so we do not claim. Saying "Protected"
* here is the exact bug being removed.
*/
function fullPlaneState(traffic: Traffic | undefined): ProtectionState {
const tunnelRules = traffic?.tunnel_rules ?? 0
switch (traffic?.verdict) {
case 'tunnel':
return {
variant: 'on',
headline: 'Protected',
detail: 'Traffic from your network is going through the tunnel.',
alarm: false,
}
case 'split':
return {
variant: 'amber',
headline: 'Partly protected — the rest goes out directly',
detail: `${ruleCount(tunnelRules)} through the tunnel. Everything they don’t match leaves through your normal internet connection, with your real address.`,
alarm: false,
}
case 'direct':
return {
variant: 'crit',
headline: 'Not protected — nothing is going through the tunnel',
detail:
'The service is running, but your routing sends every connection straight out your normal internet connection, with your real address. On the Routing page, point the default rule at a group or a node.',
alarm: true,
}
case 'blocked':
return tunnelRules > 0
? {
variant: 'amber',
headline: 'Partly protected — everything else is blocked',
detail: `${ruleCount(tunnelRules)} through the tunnel. Anything they don’t match is blocked instead of being let out, so nothing is leaving unprotected.`,
alarm: false,
}
: {
variant: 'amber',
headline: 'Nothing is getting out',
detail:
'No rule sends traffic anywhere, so every connection from your network is being blocked rather than let out unprotected. Add a default rule on the Routing page.',
alarm: true,
}
default:
return {
variant: 'amber',
headline: 'Checking where traffic goes',
detail:
'The router is up and handling your traffic. It hasn’t reported yet whether that traffic is going through the tunnel.',
alarm: false,
}
}
}
/** "One rule sends traffic" / "4 rules send traffic", so the detail lines above
* can name a number the operator can go and count on the Routing page. Falls
* back to the vague form only if the daemon sent a verdict without a count. */
function ruleCount(n: number): string {
if (n <= 0) return 'Some traffic goes'
if (n === 1) return 'One rule sends traffic'
return `${n} rules send traffic`
}
+6 -1
View File
@@ -18,5 +18,10 @@
"noFallthroughCasesInSwitch": true,
"forceConsistentCasingInFileNames": true
},
"include": ["src", "vite.config.ts"]
"include": ["src", "vite.config.ts"],
// *.test.ts runs under node's built-in test runner (`npm test`), which strips
// types rather than checking them. They are excluded here because they import
// node:test / node:assert, and the SPA deliberately carries no @types/node — it
// is embedded in the daemon binary, so every devDependency is weight on a router.
"exclude": ["src/**/*.test.ts"]
}
+41 -1
View File
@@ -44,6 +44,13 @@ type URLTest struct {
group *URLTestGroup
interruptExternalConnections bool
balancer *balancer // lx: SPEC 019 — nil for least_test (default)
// lx: health board §5.C — true when options.SelfCheck == false: the group's
// OWN probing schedule (PostStart warm-up + Touch ticker) is stood down and
// the observatory is the only thing that measures its members. Stored
// INVERTED so the zero value keeps today's behaviour for every construction
// path that does not go through NewURLTest (hand-built groups in tests).
// See option.URLTestOutboundOptions.SelfCheck for the full reasoning.
selfCheckDisabled bool
}
func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLogger, tag string, options option.URLTestOutboundOptions) (adapter.Outbound, error) {
@@ -71,6 +78,9 @@ func NewURLTest(ctx context.Context, router adapter.Router, logger log.ContextLo
idleTimeout: time.Duration(options.IdleTimeout),
interruptExternalConnections: options.InterruptExistConnections,
balancer: balancer,
// nil/absent means true (self-check on) — the documented default, so a
// config written before the flag existed behaves exactly as it always has.
selfCheckDisabled: options.SelfCheck != nil && !*options.SelfCheck,
}
if len(outbound.tags) == 0 {
return nil, E.New("missing tags")
@@ -92,6 +102,9 @@ func (s *URLTest) Start() error {
return err
}
group.balancer = s.balancer // lx: SPEC 019 v2 — health-check drives the pool through it
// lx: health board §5.C — carry the stand-down flag onto the group the same
// way the balancer travels: set after construction, immutable from then on.
group.selfCheckDisabled = s.selfCheckDisabled
if s.balancer != nil {
// lx: health board §5.B — slot liveness reads through the board verdict, so a
// death recorded by any prober or a failed dial takes effect on the next pick,
@@ -319,6 +332,12 @@ type URLTestGroup struct {
lastActive common.TypedValue[time.Time]
lastSelected common.TypedValue[string] // lx: SPEC 019 — Now() in balanced modes
balancer *balancer // lx: SPEC 019 v2 — round_robin pool; nil for least_test
// lx: health board §5.C — mirrors URLTest.selfCheckDisabled (set by Start,
// immutable afterwards, zero value = probing on). Guards ONLY the group's
// own schedule: the PostStart warm-up sweep and the Touch ticker. An
// explicit CheckOutbounds/URLTest call is untouched — the flag stands down
// the schedule, not the capability.
selfCheckDisabled bool
}
func NewURLTestGroup(ctx context.Context, outboundManager adapter.OutboundManager, logger log.Logger, outbounds []adapter.Outbound, link string, interval time.Duration, tolerance uint16, idleTimeout time.Duration, interruptExternalConnections bool) (*URLTestGroup, error) {
@@ -362,14 +381,35 @@ func (g *URLTestGroup) PostStart() {
g.lastActive.Store(time.Now())
// lx: SPEC 019 v2 — seed the pool so round_robin can route from the first connection,
// before the first health-check completes (history-warm nodes first, else config order).
// The seed only READS the board, so it runs even with the self-check stood down.
g.seedPool()
go g.CheckOutbounds(false)
// lx: health board §5.C — the warm-up sweep is the first half of the group's
// own probing schedule, and it fires for EVERY group at box start, including
// groups no routing rule reaches. For those, the sweep dials every member
// directly from the router — a path nothing uses — and records the outcome
// under the members' base tags, forging the board reading the observatory
// exists to keep honest. A stood-down group therefore skips it entirely; the
// observatory (or nothing, for a truly unused group) is what measures its
// members.
if !g.selfCheckDisabled {
go g.CheckOutbounds(false)
}
}
func (g *URLTestGroup) Touch() {
if !g.started {
return
}
// lx: health board §5.C — Touch's only job is to keep the group's OWN
// probing ticker alive while traffic flows. With the self-check stood down
// there is deliberately no ticker to start or feed: the observatory owns the
// schedule, and a stray dial through an unused group (a stale rule cache, a
// manual pin) must not arm 30 minutes of direct probing under the members'
// base tags. Checked before the lock because the flag is immutable after
// Start, exactly like the started fast-path above.
if g.selfCheckDisabled {
return
}
g.access.Lock()
defer g.access.Unlock()
if g.ticker != nil {
+116
View File
@@ -0,0 +1,116 @@
package group
// lx: health board §5.C tests — SelfCheck stands the group's OWN probing
// schedule down: no PostStart warm-up sweep, no Touch ticker. The explicit
// CheckOutbounds path stays available, and the nil default keeps probing.
import (
"context"
"testing"
"time"
"github.com/sagernet/sing-box/common/urltest"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
)
// waitForHistory polls until the store holds an entry for tag or the deadline
// passes; reports whether it appeared. PostStart's sweep runs on its own
// goroutine, so both directions of the assertion need a bounded wait.
func waitForHistory(hist *urltest.HistoryStorage, tag string, deadline time.Duration) bool {
stop := time.Now().Add(deadline)
for time.Now().Before(stop) {
if hist.LoadURLTestHistory(tag) != nil {
return true
}
time.Sleep(5 * time.Millisecond)
}
return false
}
// A group with the self-check stood down writes NOTHING to the history storage
// on PostStart: the warm-up sweep — which would dial the member directly from
// the router and mark the failure under its base tag — must not fire. And
// Touch, the other half of the schedule, must not start a ticker either.
func TestSelfCheckDisabledPostStartWritesNothing(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.selfCheckDisabled = true
g.PostStart()
// The absence of a write is the assertion, so give the (non-existent) sweep
// real time to have happened before declaring victory.
if waitForHistory(hist, "a", 150*time.Millisecond) {
t.Fatal("a stood-down group's PostStart wrote to the board; the warm-up sweep must not fire")
}
g.Touch()
g.access.Lock()
ticker := g.ticker
g.access.Unlock()
if ticker != nil {
t.Fatal("Touch armed the probing ticker on a stood-down group")
}
}
// The default (SelfCheck nil, i.e. the zero-value field on a hand-built group)
// keeps today's behaviour: PostStart's warm-up sweep runs and records the
// failing member on the board.
func TestSelfCheckDefaultStillProbesOnPostStart(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.PostStart()
if !waitForHistory(hist, "a", 5*time.Second) {
t.Fatal("default group's PostStart never probed; the self-check must stay on unless stood down")
}
if v := hist.Verdict("a", 10*time.Minute); v != urltest.VerdictDead {
t.Fatalf("verdict(a) = %v, want dead from the warm-up sweep", v)
}
}
// An EXPLICIT CheckOutbounds still probes a stood-down group: the flag
// suppresses the group's own schedule, never a deliberate request (the adapter
// interface a human or an API invokes on purpose).
func TestSelfCheckDisabledExplicitCheckStillProbes(t *testing.T) {
hist := urltest.NewHistoryStorage()
a := &healthNode{tag: "a", fail: true}
manager := managerOf(a)
g := healthTestGroup(hist, manager, a)
g.selfCheckDisabled = true
g.CheckOutbounds(true)
if hist.LoadURLTestHistory("a") == nil {
t.Fatal("an explicit CheckOutbounds(true) did not probe; the flag must only stand down the schedule")
}
}
// The option → outbound plumbing: nil/absent means on, an explicit false means
// stood down, an explicit true means on. NewURLTest is the only place the
// option is read, so this is where a plumbing regression would hide.
func TestSelfCheckOptionPlumbing(t *testing.T) {
build := func(selfCheck *bool) *URLTest {
t.Helper()
opts := option.URLTestOutboundOptions{Outbounds: []string{"a"}}
opts.SelfCheck = selfCheck
ob, err := NewURLTest(context.Background(), nil, log.NewNOPFactory().Logger(), "t", opts)
if err != nil {
t.Fatalf("NewURLTest: %v", err)
}
return ob.(*URLTest)
}
if build(nil).selfCheckDisabled {
t.Fatal("nil SelfCheck must keep the self-check ON (the compatibility default)")
}
on, off := true, false
if build(&on).selfCheckDisabled {
t.Fatal("SelfCheck=true must keep the self-check on")
}
if !build(&off).selfCheckDisabled {
t.Fatal("SelfCheck=false must stand the self-check down")
}
}
+27 -13
View File
@@ -18,9 +18,10 @@
# OpenWrt package can $(INSTALL_BIN) the arch-matched artifact.
# 5. Prints a size table + a per-arch static check (ELF type / no PT_INTERP).
#
# Router build tag set = D9 (musl-static). We deliberately DROP with_purego and
# with_naive_outbound: they pull cronet-go, which forces a glibc PT_INTERP even
# with CGO_ENABLED=0, making the binary unusable on musl OpenWrt.
# Router build tag set = D9/D23 (musl-static). It is DEFINED IN, and only in,
# scripts/router-tags.sh (sourced below) — that file documents every tag and is
# machine-checked against the declared feature list by shater/buildtags's test.
# Run scripts/check-router-tags.sh after touching it.
#
# Usage:
# scripts/build-shaterd.sh [VERSION] [--fast]
@@ -82,16 +83,15 @@ fi
[ -n "$VERSION" ] || VERSION="v0.2.0-dev"
# --- config -----------------------------------------------------------------
# D9 router tag set (musl-static). Keep in sync with docs-shater/DECISIONS.md D9.
# No with_gvisor: the data plane is tproxy/redirect (netplane), generate never
# emits a tun inbound, so the userspace gvisor stack was 3.6 MB of dead weight
# (tun would fall back to the system stack anyway).
# No with_clash_api: the panel is shater's own; generate never emits a clash_api
# service ("the shater generator emits none of those" — shater/engine/engine.go).
# No with_dhcp: shater resolvers are udp/tcp/doh/dot/local/fakeip — no "dhcp://"
# DNS transport is ever generated, and the slim registry never registers it.
ROUTER_TAGS="with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command"
LDFLAGS="-X github.com/sagernet/sing-box/constant.Version=${VERSION} -checklinkname=0 -s -w -buildid="
# D9/D23 router tag set (musl-static). The set itself lives in ONE place —
# scripts/router-tags.sh — because it is also parsed by shater/buildtags's test,
# which proves it still covers every feature docs-shater/FEATURES.md declares.
# Do not re-inline it here: that split is exactly how `with_gvisor` went missing
# while `with_wireguard` stayed (D23).
# shellcheck source=router-tags.sh
. "$SCRIPT_DIR/router-tags.sh"
ROUTER_TAGS="$SHATER_ROUTER_TAGS"
LDFLAGS="-X github.com/sagernet/sing-box/constant.Version=${VERSION} ${SHATER_ROUTER_LDFLAGS} -s -w -buildid="
UPX_BIN="${UPX:-upx}"
# UPX itself treats the environment variable UPX as extra command-line options, so
@@ -111,6 +111,20 @@ echo " version : $VERSION"
echo " tags : $ROUTER_TAGS"
echo " upx : $UPX_BIN"
echo " go : $(go version)"
# --- gate: does this tag set still support what we declare? (D23) ------------
# Cheap (one tiny tag-less package, no network, ~1 s) and it travels with the
# BUILD rather than with a CI config, so an artifact produced by hand on a
# developer's machine gets the same guarantee. The heavier half — actually
# constructing every declared protocol under these tags — is
# scripts/check-router-tags.sh, which CI runs before this script.
if ! (cd "$REPO" && go test -count=1 ./shater/buildtags/ >/dev/null); then
echo >&2
echo " ABORT: the router tag set no longer covers a declared feature." >&2
echo " Details: go test ./shater/buildtags/" >&2
echo " Full check: scripts/check-router-tags.sh" >&2
exit 1
fi
echo
# --- step 1: build the SPA --------------------------------------------------
+118
View File
@@ -0,0 +1,118 @@
#!/usr/bin/env bash
#
# check-router-tags.sh — prove the SHIPPED build-tag set still supports every
# feature shater declares (D23).
#
# WHY (2026-07-25): the router tag set is a trimmed subset of upstream's, but the
# test suite builds with the FULL upstream set — so the one combination we
# actually ship was never exercised. `with_gvisor` got trimmed while
# `with_wireguard` stayed, and every shipped binary answered a WireGuard node
# with "gVisor is not included in this build". Compiling is not evidence.
#
# WHAT IT RUNS
# 1. shater/buildtags, TAG-LESS — reads scripts/router-tags.sh and fails if a
# declared feature (buildtags.Features) lost a build tag it needs. Cheap,
# hostable anywhere, catches the trim at the moment it happens.
# 2. shater/generate + shater/buildtags, WITH THE SHIPPED TAG SET on linux —
# constructs one node of every declared protocol through box.New+Start, and
# cross-checks that the tag detectors match the set the compiler was given.
# This is the half that catches "the tag is there but insufficient".
#
# The run is unprivileged (no tproxy inbound is built) and offline apart from Go
# module downloads.
#
# Usage:
# scripts/check-router-tags.sh
#
# Env:
# SHATER_GO_IMAGE docker image used to reach linux from a non-linux host
# (default golang:1.26 — keep it >= go.mod's toolchain).
# SHATER_NO_DOCKER=1 fail instead of falling back to docker.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO="$(cd "$SCRIPT_DIR/.." && pwd)"
cd "$REPO"
# shellcheck source=router-tags.sh
. "$SCRIPT_DIR/router-tags.sh"
echo "== router build-tag check =="
echo " tags : $SHATER_ROUTER_TAGS"
echo " ldflags: $SHATER_ROUTER_LDFLAGS"
echo
# --- 1. static: does the set still cover the declared features? --------------
# No tags, no OS constraint: this is the check that would have caught the outage
# on the developer's own machine.
echo "== [1/2] declared features vs. the shipped tag set (no tags needed) =="
go test -count=1 ./shater/buildtags/
echo
# --- 2. behavioural: does the shipped combination actually construct? --------
# The protocol-construction test is linux-only (box.New validates the loop-guard
# routing_mark only there). From a non-linux host, re-exec this script inside a
# golang container rather than silently skipping — a skipped guard is no guard.
if [ "$(go env GOOS)" != "linux" ] && [ "${SHATER_TAGCHECK_IN_DOCKER:-0}" != "1" ]; then
if [ "${SHATER_NO_DOCKER:-0}" = "1" ] || ! command -v docker >/dev/null 2>&1; then
echo " ERROR: step 2 needs linux (GOOS=$(go env GOOS)) and docker is unavailable/disabled." >&2
echo " Run this script on the linux CI runner or the OpenWrt VM." >&2
exit 1
fi
image="${SHATER_GO_IMAGE:-golang:1.26}"
echo "== [2/2] re-exec on linux via docker ($image) =="
host_repo="$REPO"
command -v cygpath >/dev/null 2>&1 && host_repo="$(cygpath -w "$REPO")"
# Named volumes keep the module/build cache warm between runs; MSYS2_ARG_CONV_EXCL
# stops Git Bash from rewriting the container-side paths into windows ones.
MSYS2_ARG_CONV_EXCL='*' MSYS_NO_PATHCONV=1 docker run --rm \
-v "$host_repo":/src \
-v shater-tagcheck-gomod:/go/pkg/mod \
-v shater-tagcheck-gocache:/root/.cache/go-build \
-w /src \
-e SHATER_TAGCHECK_IN_DOCKER=1 \
"$image" bash scripts/check-router-tags.sh
exit $?
fi
echo "== [2/2] every declared protocol constructs under the SHIPPED tags =="
out="$(mktemp)"
trap 'rm -f "$out"' EXIT
set +e
SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
-run 'TestShippedTagSetConstructsDeclaredProtocols|TestEveryTagGatedFeatureIsProbed' \
./shater/generate/ >"$out" 2>&1
rc_gen=$?
set -e
sed 's/^/ /' "$out"
set +e
SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
-run 'TestCompiledTagsMatchTheShippedSet' \
./shater/buildtags/ >"$out" 2>&1
rc_tags=$?
set -e
sed 's/^/ /' "$out"
if [ "$rc_gen" -ne 0 ] || [ "$rc_tags" -ne 0 ]; then
echo
echo " FAILED: the tag set we SHIP cannot do what we declare." >&2
exit 1
fi
# A guard that silently runs nothing is worse than no guard: prove the tests were
# actually compiled in and executed (build tags / file renames could exclude them).
for want in TestShippedTagSetConstructsDeclaredProtocols TestCompiledTagsMatchTheShippedSet; do
if ! SHATER_ROUTER_TAG_CHECK=1 go test -count=1 -v \
-tags "$SHATER_ROUTER_TAGS" -ldflags "$SHATER_ROUTER_LDFLAGS" \
-run "$want" ./shater/generate/ ./shater/buildtags/ 2>&1 | grep -q -- "--- PASS: $want"; then
echo " FAILED: $want did not run (build-tag/file-name drift?)" >&2
exit 1
fi
done
echo
echo "== OK: the shipped tag set covers every declared feature, and every =="
echo "== declared protocol constructs through box.New under it. =="
+78
View File
@@ -0,0 +1,78 @@
# shellcheck shell=sh
#
# router-tags.sh — THE build-tag set of the shipped `shaterd` router binary.
#
# This file is DATA, not a program: it is `.`-sourced by
# - scripts/build-shaterd.sh (the ship build)
# - scripts/check-router-tags.sh (the guard that proves the set is complete)
# and it is PARSED by shater/buildtags/buildtags_test.go, which asserts that
# every feature docs-shater/FEATURES.md declares supported has its build tags
# present here. Change the set here and nowhere else.
#
# WHY A TRIMMED SET AT ALL (D9): upstream's DEFAULT_BUILD_TAGS registers the
# whole sing-box zoo. We drop what shater/generate can never emit, because a
# router binary pays for every tag twice — flash and (UPX unpacks into anonymous
# pages) resident RAM. We do NOT drop what a declared feature needs to run.
#
# WHY THIS FILE EXISTS (the 2026-07-25 WireGuard outage): the set used to be a
# string literal inside build-shaterd.sh, with nothing connecting it to the
# feature list. `with_gvisor` was trimmed as "unreachable code" while
# `with_wireguard` stayed — so every shipped binary answered a WireGuard node
# with "gVisor is not included in this build". No test caught it: the test suite
# builds with the FULL upstream tag set, so the SHIPPED combination was never
# exercised. One file + one test now hold the two halves together (D23).
#
# ---- the set -----------------------------------------------------------------
# with_gvisor userspace netstack. REQUIRED BY with_wireguard: both
# transport/wireguard device constructors
# (newStackDevice AND newSystemStackDevice) are stubs
# returning tun.ErrGVisorNotIncluded without it, so
# box.New dies at "create WireGuard device". Not
# optional as long as we ship WireGuard/AmneziaWG.
# with_quic hysteria2 + tuic outbounds, QUIC/HTTP3 DNS
# transports, and the vless/vmess `type=quic` transport
# (shater/registry/registry_quic.go).
# with_wireguard registers the wireguard endpoint
# (shater/registry/registry_wireguard.go).
# with_awg AmneziaWG obfuscation params (jc/jmin/jmax, s1-s4,
# h1-h4, i1-i5) actually reach the device
# (transport/wireguard/device_awg.go). A driving
# product requirement — FEATURES.md marks it [MVP].
# with_utls uTLS fingerprints AND REALITY: common/tls/
# reality_client.go is itself `//go:build with_utls`.
# with_xhttp the XHTTP/SplitHTTP v2ray transport
# (shater/parse emits type=xhttp).
# badlinkname badtls fast path (common/badtls/*.go are
# `go1.25 && badlinkname`). Needs -checklinkname=0 in
# SHATER_ROUTER_LDFLAGS below or the LINK step fails.
# tfogo_checklinkname0 same deal for tfo-go's linkname use.
# with_lx_command lx daemon command server. Inert for shaterd (nothing
# under shater/ imports sing-box/daemon or libbox —
# `go list -deps ./shater/cmd/shaterd` links neither),
# kept only so the router set stays a subset of the lx
# desktop set. Costs nothing; safe to drop later.
#
# Deliberately NOT here (each is unreachable for shater, not merely unused):
# with_purego,with_naive_outbound cronet-go forces a glibc PT_INTERP even at
# CGO_ENABLED=0 -> will not run on musl (D9).
# with_clash_api the panel is shater's own web server;
# generate emits no clash_api service.
# with_dhcp resolver types are udp/tcp/doh/dot/local/
# fakeip; no dhcp:// transport is generated.
# with_tailscale,with_acme,with_ech,with_usbip,with_cloudflared,with_ocm,
# with_ccm,with_v2ray_api,with_reality_server
# nothing in shater/parse or shater/generate
# can produce them; hysteria2 `ech=` is
# refused with a warning in the parser.
#
# Adding a tag here is cheap. REMOVING one is a product decision: run
# `scripts/check-router-tags.sh` — it fails if the set no longer covers a
# declared feature, and it fails if the shipped combination cannot construct
# every protocol through box.New.
SHATER_ROUTER_TAGS="with_gvisor,with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,with_awg,with_lx_command"
# Linker flags the tag set REQUIRES (they are not optional trimming: `badlinkname`
# without -checklinkname=0 fails at link time with
# "invalid reference to crypto/tls.(*Conn).handlePostHandshakeMessage").
SHATER_ROUTER_LDFLAGS="-checklinkname=0"
+84 -12
View File
@@ -74,6 +74,19 @@ type Applier struct {
stateMu sync.RWMutex
holding bool
// traffic is WHERE THE TRAFFIC GOES under the config that is currently running:
// tunnelled, split, straight out, or blocked (see generate.TrafficOf). It is
// computed from the generated option.Options at the moment they are handed to
// the engine, so it describes what runs rather than what is on disk.
//
// Deliberately separate from `holding` and from Plane, which answer "how much of
// the data plane is installed". The panel used to read Plane == "full" as
// "protected" and said so under a green LED on a router whose only rule was
// `default -> direct`; the plane really was fully installed, and the whole LAN
// really was going out the plain WAN. Zero value = not known (no successful
// apply in this process yet), which is NOT the same as "tunnel".
traffic generate.Traffic // guarded by stateMu
// lastWarnings is the normalised warning set from the last SUCCESSFUL apply,
// published through Status so the panel can show fail-open degradations
// (a blocklist that did not load, a DoH host left reachable) instead of
@@ -192,11 +205,15 @@ func (a *Applier) configureObservatory(m *model.Model, opts option.Options) {
})
}
// TestGroups launches the engine's one-shot exit test (delay + exit address, F2)
// of the named groups/chains, returning started=false when a run is already in
// flight or the engine is absent. names empty/nil = every group and every chain.
// It takes NEITHER the apply mutex nor the flock, so kicking off a test never
// blocks behind an Apply.
// TestGroups launches the engine's one-shot group/chain test (F2): the engine
// asks its observatory for an out-of-turn pass and reports what it measured,
// plus the exit address for alive rule-routed targets — it no longer dials
// health probes of its own. Returns started=false when a run is already in
// flight or the engine is absent. names empty/nil = every group and every
// chain. probeURL is passed through for signature stability and IGNORED by the
// engine (the probe URL is a global observatory setting now). It takes NEITHER
// the apply mutex nor the flock, so kicking off a test never blocks behind an
// Apply.
func (a *Applier) TestGroups(names []string, probeURL string) (started bool) {
if a.eng == nil {
return false
@@ -514,6 +531,12 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
}
a.setHolding(false)
a.lastGood = m
// Publish where this config actually sends traffic, read off the very options
// the engine was just handed (a.eng.Apply above). The engine's hash fast-path
// may have skipped a swap, in which case these options are hash-equal to what is
// already running — either way they are the running config, which is the only
// config this verdict may describe.
a.setTraffic(generate.TrafficOf(opts))
// Publish the warnings of THIS successful apply, in one normalised set, and log
// them in one consistent format. Status carries them to the panel so a
// fail-open degradation is visible in the UI instead of only in logread.
@@ -557,6 +580,11 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
// With the kill-switch OPEN nothing is installed — fail-open is the operator's
// documented choice and this must not quietly override it.
func (a *Applier) holdLocked(m *model.Model, cause error) {
// The engine is not carrying anything, so whatever the last running config did
// with traffic is no longer true of this router. Forget it either way — a stale
// "tunnel" verdict left behind by a config that is no longer running is the same
// reassuring lie in a different place.
a.setTraffic(generate.Traffic{})
if !killSwitchClosed(m.Globals) {
a.log.Warn("engine is down and kill_switch=open: LAN traffic is NOT protected (documented fail-open): ", cause)
return
@@ -640,24 +668,52 @@ func (a *Applier) setHolding(v bool) {
a.stateMu.Unlock()
}
// Traffic returns where the traffic of the CURRENTLY RUNNING config goes. The
// zero value means no config of this process's is running (nothing applied yet,
// or the plane was torn down / put on hold), and callers must render that as
// "unknown", never as protected.
func (a *Applier) Traffic() generate.Traffic {
a.stateMu.RLock()
defer a.stateMu.RUnlock()
return a.traffic
}
func (a *Applier) setTraffic(t generate.Traffic) {
a.stateMu.Lock()
a.traffic = t
a.stateMu.Unlock()
}
func (a *Applier) setWarnings(ws []Warning) {
a.stateMu.Lock()
a.lastWarnings = ws
a.stateMu.Unlock()
}
// Warnings returns the normalised warning set from the last successful apply.
// Never nil: an empty slice means "the last apply was clean", which the panel
// must render differently from "no apply has run yet" (Active/Plane cover that).
// Warnings returns the normalised warning set from the last successful apply,
// PLUS whatever is wrong right now that no apply can describe. Never nil: an
// empty slice means "the last apply was clean", which the panel must render
// differently from "no apply has run yet" (Active/Plane cover that).
//
// The live half is currently the engine's abandoned generations
// (engineTeardownWarnings). It is computed at READ time rather than folded into
// lastWarnings on purpose: a superseded box that will not shut down is a
// condition of the process, not a property of a config. Folding it in would make
// it appear only after the NEXT successful apply and then stay published long
// after the shutdown finally completed — reporting a leak that is over, and
// staying silent about one that is not. Read-time means it shows up the instant
// it happens and clears itself the instant it resolves.
func (a *Applier) Warnings() []Warning {
var out []Warning
if a.eng != nil {
out = engineTeardownWarnings(a.eng.PendingCloses())
}
a.stateMu.RLock()
defer a.stateMu.RUnlock()
if a.lastWarnings == nil {
if out == nil && a.lastWarnings == nil {
return []Warning{}
}
out := make([]Warning, len(a.lastWarnings))
copy(out, a.lastWarnings)
return out
return append(out, a.lastWarnings...)
}
// Reconcile re-reads UCI and either tears down (disabled) or re-applies (enabled).
@@ -723,6 +779,7 @@ func (a *Applier) Teardown() error {
a.lastGood = nil
a.lastNft = ""
a.setHolding(false)
a.setTraffic(generate.Traffic{})
a.setWarnings(nil)
// The plane is gone, so the logged set no longer describes anything. Forget it,
// and the next apply re-announces its warnings in full rather than staying
@@ -1003,6 +1060,20 @@ type Status struct {
// say so rather than looking healthy.
Plane string `json:"plane"`
// Traffic is WHERE THE TRAFFIC GOES under the running config — tunnelled, split,
// straight out, or blocked (see generate.Traffic).
//
// Plane does NOT answer this, and reading it as if it did is the defect this
// field exists for. Plane == "full" only means the table, the policy routing and
// the engine are all in place; a router whose one rule is `default -> direct` has
// all three and sends every packet out the plain WAN with its real address. The
// panel showed that as "Protected — traffic is going through the tunnel".
//
// Zero value (Verdict == "") means unknown: no successful apply has run in this
// daemon process yet, or the plane is on hold / torn down. A consumer must render
// that as unknown and never as protected.
Traffic generate.Traffic `json:"traffic"`
// Warnings is the normalised warning set from the last successful apply:
// everything that was skipped, degraded or left un-applied while the apply
// still succeeded. Always non-nil so the panel can map over it unconditionally.
@@ -1082,6 +1153,7 @@ func (a *Applier) Status() Status {
Hash: a.eng.Hash(),
CanRollback: a.canRollback(),
EngineRunning: a.eng.Running(),
Traffic: a.Traffic(),
Warnings: a.Warnings(),
}
s.StartedUnix, s.UptimeSeconds = processUptime(time.Now())
+283
View File
@@ -0,0 +1,283 @@
package apply
// Regression cover for the leaked-engine-generation defect.
//
// Observed on the router: one shaterd process was carrying up to FOUR sing-box
// instances at once. sing-box stamps every log line with the elapsed seconds of
// ITS OWN instance, so the same process printed `ERROR[2015]` and `ERROR[0129]`
// in the same second — two engines half an hour apart in age, both alive, both
// dialling, both holding WireGuard devices built from the same private keys. A
// full daemon stop+start collapsed it back to one generation, which places the
// leak squarely on the config re-apply path rather than on startup.
//
// The tests below pin the two halves of the fix:
//
// TestApplySwapsLeaveExactlyOneEngineGeneration — the healthy path really
// retires the old instance (its listener is provably gone), N times in a row.
// TestStuckEngineCloseDoesNotBlockTheApply — a shutdown that never returns
// is bounded, does not stall the apply, and is REPORTED as a critical warning
// for exactly as long as it is true.
import (
"io"
"net"
"net/netip"
"strconv"
"strings"
"testing"
"time"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing/common/json/badoption"
)
// mixedOn builds a minimal but REAL engine config: a mixed inbound bound to
// 127.0.0.1:port plus a direct outbound. Two configs with different ports hash
// differently, so each Apply is a genuine swap rather than a hash-gate no-op —
// and the bound port is the observable that proves whether the old instance
// actually died.
func mixedOn(port uint16) option.Options {
listen := badoption.Addr(netip.MustParseAddr("127.0.0.1"))
return option.Options{
Log: &option.LogOptions{Level: "error"},
Inbounds: []option.Inbound{{
Type: C.TypeMixed,
Tag: "mixed-in",
Options: &option.HTTPMixedInboundOptions{
ListenOptions: option.ListenOptions{Listen: &listen, ListenPort: port},
},
}},
Outbounds: []option.Outbound{{
Type: C.TypeDirect,
Tag: "direct-out",
Options: &option.DirectOutboundOptions{},
}},
}
}
// portFree reports whether 127.0.0.1:port can be bound right now — i.e. whether
// the instance that used to listen there is really gone. Retried briefly because
// a listener is released by Close, not by the return of Close's caller.
func portFree(port uint16) bool {
deadline := time.Now().Add(3 * time.Second)
for {
ln, err := net.Listen("tcp", net.JoinHostPort("127.0.0.1", strconv.Itoa(int(port))))
if err == nil {
_ = ln.Close()
return true
}
if time.Now().After(deadline) {
return false
}
time.Sleep(20 * time.Millisecond)
}
}
// TestApplySwapsLeaveExactlyOneEngineGeneration is the core regression: after N
// sequential applies the process must be carrying ONE engine, not N.
//
// "Carrying" is checked two ways on purpose. Generations() is the engine's own
// accounting (running instance + every retirement still in flight) and would
// catch a retirement that silently never completes. The port check is
// independent of that accounting: if the superseded instance were still alive it
// would still hold its listener, and the bind would fail. A fix that only
// reset a pointer would pass the first check and fail the second.
func TestApplySwapsLeaveExactlyOneEngineGeneration(t *testing.T) {
const (
firstPort = 18801
applies = 5
)
a := New(engine.New(), nil)
t.Cleanup(func() { _ = a.eng.Close() })
for i := 0; i < applies; i++ {
port := uint16(firstPort + i)
changed, err := a.eng.Apply(mixedOn(port))
if err != nil {
t.Fatalf("apply #%d (port %d): %v", i+1, port, err)
}
if !changed {
t.Fatalf("apply #%d: every config here differs, so the swap must be real", i+1)
}
if got := a.eng.Generations(); got != 1 {
t.Fatalf("after apply #%d the process carries %d engine instances, want exactly 1 — "+
"a superseded generation is still alive (this is the four-generations-in-one-process leak)",
i+1, got)
}
if stuck := a.eng.PendingCloses(); len(stuck) != 0 {
t.Fatalf("after apply #%d: %d generation(s) abandoned, want none: %+v", i+1, len(stuck), stuck)
}
if i > 0 {
prev := uint16(firstPort + i - 1)
if !portFree(prev) {
t.Fatalf("after apply #%d the PREVIOUS generation still holds 127.0.0.1:%d — "+
"the old engine was replaced in the field but never actually stopped", i+1, prev)
}
}
// The warning set must stay clean while teardown is healthy: a critical
// warning that cries wolf on every apply is worse than none.
for _, w := range a.Warnings() {
if w.Section == "engine" {
t.Fatalf("after apply #%d a healthy swap produced an engine warning: %+v", i+1, w)
}
}
}
if err := a.eng.Close(); err != nil {
t.Fatalf("close: %v", err)
}
if got := a.eng.Generations(); got != 0 {
t.Fatalf("after Close the process carries %d engine instances, want 0", got)
}
if !portFree(firstPort + applies - 1) {
t.Fatalf("after Close the last generation still holds its listener")
}
}
// hangingCloser returns a box closer that BLOCKS the first close it is handed
// until release() is called, and performs every later close normally. That is the
// shape of the real fault: one subsystem of one generation (a WireGuard endpoint)
// refuses to come down, while the rest of the process is fine.
func hangingCloser() (closer func(io.Closer) error, release func()) {
gate := make(chan struct{})
first := make(chan struct{}, 1)
first <- struct{}{}
return func(c io.Closer) error {
select {
case <-first:
<-gate // the stuck generation: never returns until released
return c.Close() // ...and then really does close, so the port frees
default:
return c.Close()
}
}, func() {
close(gate)
}
}
// TestStuckEngineCloseDoesNotBlockTheApply pins all four requirements of the
// bounded teardown at once:
//
// 1. the apply COMPLETES — a shutdown that never returns must not hold the
// control plane (and therefore the panel) hostage;
// 2. the new engine is running afterwards — fail-closed semantics are unchanged,
// the swap succeeded;
// 3. the abandoned generation is REPORTED as a critical warning, by name, for as
// long as it is still running — it is not silently swallowed;
// 4. generations do not stack: a further apply while the leak persists leaves
// one live instance plus the one abandoned one, not three.
func TestStuckEngineCloseDoesNotBlockTheApply(t *testing.T) {
const budget = 200 * time.Millisecond
restoreBudget := engine.SetCloseBudget(budget)
defer restoreBudget()
a := New(engine.New(), nil)
// Generation 1 comes up with the REAL closer still installed.
if _, err := a.eng.Apply(mixedOn(18811)); err != nil {
t.Fatalf("apply #1: %v", err)
}
closer, release := hangingCloser()
restoreCloser := engine.SetBoxCloser(closer)
released := false
defer func() {
if !released {
release()
}
restoreCloser()
_ = a.eng.Close()
}()
// (1) Generation 2: the retirement of generation 1 will never return.
start := time.Now()
changed, err := a.eng.Apply(mixedOn(18812))
elapsed := time.Since(start)
if err != nil {
t.Fatalf("apply #2 must SUCCEED despite the stuck teardown: %v", err)
}
if !changed {
t.Fatalf("apply #2: expected a real swap")
}
// Generously bounded: the budget plus box.New+Start. The point is that it
// returned at all — before the fix this waited on Box.Close forever.
if elapsed > budget+20*time.Second {
t.Fatalf("apply #2 took %s: a stuck teardown must not stall the apply", elapsed)
}
// (2) fail-closed semantics unchanged: the new engine really is up.
if !a.eng.Running() {
t.Fatalf("apply #2: the new engine must be running")
}
// (3) the leak is visible, named, and critical.
stuck := a.eng.PendingCloses()
if len(stuck) != 1 {
t.Fatalf("PendingCloses() = %+v, want exactly the one abandoned generation", stuck)
}
if stuck[0].Generation != 1 {
t.Errorf("abandoned generation = %d, want 1", stuck[0].Generation)
}
ws := a.Status().Warnings // the exact set `shaterd status` and the panel read
var found *Warning
for i := range ws {
if ws[i].Section == "engine" {
found = &ws[i]
break
}
}
if found == nil {
t.Fatalf("a superseded engine that will not shut down produced NO warning; "+
"Status would show a healthy router: %+v", ws)
}
if found.Severity != SeverityCritical {
t.Errorf("stuck-teardown warning severity = %q, want %q", found.Severity, SeverityCritical)
}
if !strings.Contains(found.Name, "generation 1") {
t.Errorf("stuck-teardown warning must name the generation, got Name=%q", found.Name)
}
if !strings.Contains(found.Message, "STILL RUNNING") {
t.Errorf("stuck-teardown warning must say the instance is still running, got %q", found.Message)
}
if got := a.eng.Generations(); got != 2 {
t.Fatalf("Generations() = %d, want 2 (one live + one abandoned)", got)
}
// (4) another apply while the leak persists must not add a THIRD generation:
// with a generation abandoned the swap goes close-old-then-start-new, so the
// process still holds one live instance plus the one that will not die.
if _, err := a.eng.Apply(mixedOn(18813)); err != nil {
t.Fatalf("apply #3: %v", err)
}
if got := a.eng.Generations(); got != 2 {
t.Fatalf("Generations() = %d after a third apply, want 2 — generations are stacking, "+
"which is exactly the four-live-engines fault", got)
}
if stuck := a.eng.PendingCloses(); len(stuck) != 1 || stuck[0].Generation != 1 {
t.Fatalf("PendingCloses() = %+v, want only the original abandoned generation 1", stuck)
}
// (5) and it CLEARS: when the shutdown finally completes the warning goes away
// on its own. A leak report that outlives the leak trains the operator to
// ignore the panel.
release()
released = true
deadline := time.Now().Add(5 * time.Second)
for len(a.eng.PendingCloses()) > 0 && time.Now().Before(deadline) {
time.Sleep(10 * time.Millisecond)
}
if got := a.eng.PendingCloses(); len(got) != 0 {
t.Fatalf("the finished shutdown is still reported as abandoned: %+v", got)
}
for _, w := range a.Warnings() {
if w.Section == "engine" {
t.Fatalf("the engine warning outlived the leak it describes: %+v", w)
}
}
if got := a.eng.Generations(); got != 1 {
t.Fatalf("Generations() = %d after the stuck shutdown completed, want 1", got)
}
}
+47
View File
@@ -0,0 +1,47 @@
package apply
import (
"errors"
"testing"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/generate"
"github.com/sagernet/sing-box/shater/model"
)
// TestStatusReportsTraffic pins the wiring the panel's headline depends on.
//
// Plane says how much of the data plane is installed; it does NOT say where the
// traffic goes, and reading it as if it did put "Protected — traffic is going
// through the tunnel" on a router whose only rule was `default -> direct`. The
// verdict that answers the real question travels in Status.Traffic, so it must
// (a) start unknown, (b) surface what the last successful apply published, and
// (c) go back to unknown the moment the engine stops carrying that config.
func TestStatusReportsTraffic(t *testing.T) {
a := New(engine.New(), nil)
// A fresh applier has applied nothing, so it knows nothing. The zero value must
// NOT read as any verdict — least of all "tunnel".
if got := a.Status().Traffic; got.Verdict != "" {
t.Fatalf("fresh applier: Traffic.Verdict = %q, want \"\" (unknown)", got.Verdict)
}
a.setTraffic(generate.Traffic{Verdict: generate.VerdictTunnel, Default: "auto", TunnelRules: 2})
got := a.Status().Traffic
if got.Verdict != generate.VerdictTunnel || got.Default != "auto" || got.TunnelRules != 2 {
t.Fatalf("Status().Traffic = %+v, want the verdict the last apply published", got)
}
// The engine is down and the config it was running is no longer in force. A
// verdict left over from it would be the same reassuring lie, one layer down.
// (kill_switch=open takes holdLocked's early return, which is precisely the path
// that must still forget the verdict.)
g := model.DefaultGlobals()
g.KillSwitch = "open"
a.mu.Lock()
a.holdLocked(&model.Model{Globals: g}, errors.New("engine start failed"))
a.mu.Unlock()
if got := a.Status().Traffic; got.Verdict != "" {
t.Fatalf("after hold: Traffic.Verdict = %q, want \"\" (unknown) — the engine is not carrying that config", got.Verdict)
}
}
+172 -6
View File
@@ -20,6 +20,7 @@ import (
"net/netip"
"strings"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
@@ -30,6 +31,110 @@ import (
// engine.RuleSetIPCIDRs).
type ruleSetCIDRs func(tag string) ([]netip.Prefix, bool)
// SCHEMA v2 MOVED A RULE'S DESTINATION OUT OF THE RULE, AND THIS FILE HAD TO FOLLOW.
//
// Until D21 a routing rule spelled its destination inline: `dst_domain` became
// DefaultRule.Domain*, `dst_ip` became DefaultRule.IPCIDR. Both were visible right
// on the rule, so matchesUntunnelable could read them off it and the ADDRESS half
// could only ever be resolved through the engine for the rare geosite/geoip case.
//
// Since v2 a rule names a `config ruleset` and generate materialises it into
// Route.RuleSet, so BOTH halves arrive by reference. That broke two things at once,
// in opposite directions:
//
// - DOMAINS WENT INVISIBLE. A v1 rule `dst_domain=[bank.ru] dst_ip=[10/8]` ANDed
// its matchers, so it could not match an ICMP packet (no domain in one) and this
// file skipped it. After migration the same rule reads `rule_set: [rule-x,
// rule-x-ip]`, whose two sets are ORed — so the address set alone would make the
// rule look applicable, and the plan would start emitting an ALLOW (target
// direct) or a DENY (target block) for 10/8 that the operator never asked for.
// The ALLOW is the dangerous half: it sends previously-tunnelled ICMP out with
// the client's real address, i.e. an upgrade silently opening a leak. So a rule
// whose rule-sets are KNOWN to carry domain matchers is treated exactly as its
// v1 form was: inapplicable to untunnelable traffic. (The OR is deliberate and
// documented in D21 for the ENGINE's TCP/UDP path; it does not follow that a
// leak-guard should widen itself during an upgrade without a word.)
// - ADDRESSES WENT ENGINE-ONLY. `rule-x-ip` is an INLINE rule-set: its prefixes
// are sitting right there in the generated options. Asking the engine for them
// — and, when the engine is not up, truncating the whole walk and denying
// everything — was needless. Inline sets are now read from the options, so an
// engine that has not started yet no longer costs the operator their ping.
//
// Both come from the same index, built once per plan: what the CONFIG ITSELF can
// say about each rule-set. Remote/local sets stay the engine's business.
// ruleSetFacts is what the generated options alone reveal about one rule-set.
//
// opaque means "not knowable from here" — a local/remote set (only the engine has
// its contents) or an inline set with a rule shape this code does not model. An
// opaque set is never used to conclude anything; it falls through to the engine
// lookup exactly as before.
type ruleSetFacts struct {
opaque bool
hasDomain bool // carries a matcher only a NAMED destination can satisfy
cidrs []netip.Prefix // addresses, when they are knowable here
}
// indexRuleSets summarises every rule-set in the generated options by tag.
func indexRuleSets(sets []option.RuleSet) map[string]ruleSetFacts {
out := make(map[string]ruleSetFacts, len(sets))
for _, rs := range sets {
var f ruleSetFacts
// option.RuleSet leaves Type empty for inline in some marshalled forms; the
// generator always sets it, and both spellings mean the same thing here.
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
f.opaque = true
out[rs.Tag] = f
continue
}
for _, hr := range rs.InlineOptions.Rules {
if hr.Type != C.RuleTypeDefault && hr.Type != "" {
// A logical headless rule, or a shape sing-box adds later. Never guessed
// at — the same rule the walk itself follows for a logical route rule.
f = ruleSetFacts{opaque: true}
break
}
d := hr.DefaultOptions
if headlessNeedsUnavailableMatcher(d) {
f.hasDomain = true
}
f.cidrs = append(f.cidrs, parsePrefixes(d.IPCIDR)...)
}
out[rs.Tag] = f
}
return out
}
// headlessNeedsUnavailableMatcher is matchesUntunnelable's predicate for the
// HEADLESS rule shape a rule-set carries. It reports whether the rule demands
// something a packet with no domain, no ports, no sniffed protocol and no process
// identity cannot supply — in which case that rule-set cannot claim such a packet.
//
// The generator only ever emits domain matchers or ip_cidr into an inline routing
// rule-set, so in practice this is the domain test; the rest is there so a future
// matcher is refused rather than silently ignored.
func headlessNeedsUnavailableMatcher(d option.DefaultHeadlessRule) bool {
switch {
case len(d.Domain) > 0, len(d.DomainSuffix) > 0, len(d.DomainKeyword) > 0,
len(d.DomainRegex) > 0, len(d.AdGuardDomain) > 0:
return true // no domain in an ICMP/ESP/GRE packet
case len(d.Port) > 0, len(d.PortRange) > 0, len(d.SourcePort) > 0, len(d.SourcePortRange) > 0:
return true
case len(d.Network) > 0, len(d.QueryType) > 0:
return true
case len(d.ProcessName) > 0, len(d.ProcessPath) > 0, len(d.ProcessPathRegex) > 0,
len(d.PackageName) > 0, len(d.PackageNameRegex) > 0:
return true
case len(d.WIFISSID) > 0, len(d.WIFIBSSID) > 0:
return true
case d.Invert:
// Same reasoning as the route-rule case: inverting flips every conservative
// assumption, so the set may only ever withhold, never grant.
return true
}
return false
}
// buildUntunnelablePlan reduces the resolved routing rules to the first-match
// walk the nft forward chain can evaluate for a packet that has no domain, no
// ports and no sniffable protocol.
@@ -57,6 +162,9 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
// No routing at all: nothing is provably direct.
return plan
}
// Since schema v2 a rule's destination lives in a rule-set, so the rule-sets
// have to be read alongside the rules. See the block above ruleSetFacts.
sets := indexRuleSets(opts.Route.RuleSet)
for _, r := range opts.Route.Rules {
if r.Type == "logical" {
@@ -78,7 +186,10 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
// and the destination logic never allowed anything. Its `protocol: dns`
// matcher makes it inapplicable to a packet with no stream to sniff, which
// is the correct and obvious reading once the checks are in this order.
if !matchesUntunnelable(d) {
if !matchesUntunnelable(d, sets) {
if note, ok := namedDestinationNote(d, sets); ok {
plan.Warnings = append(plan.Warnings, note)
}
continue // needs a domain/port/protocol: cannot apply to this traffic
}
@@ -96,21 +207,28 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
mt := netplane.UntunnelableMatch{Allow: verdict}
// Destination predicate: explicit ip_cidr plus every rule-set's addresses.
// An INLINE rule-set is answered straight out of the config — it is fully
// present there — so a rule whose lists are all inline no longer depends on
// the engine being up. Only local/remote sets still have to be asked for.
dst := parsePrefixes(d.IPCIDR)
resolvable := true
var unresolved string
for _, tag := range d.RuleSet {
if f, known := sets[tag]; known && !f.opaque {
dst = append(dst, f.cidrs...)
continue
}
prefixes, ok := lookup(tag)
if !ok {
resolvable = false
unresolved = tag
break
}
dst = append(dst, prefixes...)
}
if !resolvable {
if unresolved != "" {
plan.Warnings = append(plan.Warnings, fmt.Sprintf(
"list %q is not loaded yet, so ping/IPTV/VPN passthrough stays blocked for the "+
"addresses it covers (it resolves itself once the list downloads)",
strings.Join(d.RuleSet, ", ")))
unresolved))
return plan
}
hasDstMatcher := len(d.IPCIDR) > 0 || len(d.RuleSet) > 0
@@ -198,7 +316,20 @@ func classifyAction(a option.RuleAction) (isDirect bool, class int) {
// Every matcher listed here is one the packet cannot satisfy, so a rule carrying
// any of them is inapplicable rather than universally matching. Matchers ANDed
// within a rule mean one unsatisfiable matcher makes the whole rule unsatisfiable.
func matchesUntunnelable(d option.DefaultRule) bool {
//
// sets carries the same test for the matchers that are no longer written on the
// rule: since schema v2 a destination is a rule-set reference, so a rule's domains
// live in Route.RuleSet rather than in d.Domain*. A rule referencing a set that is
// KNOWN to match by name is inapplicable here for exactly the reason a d.Domain
// rule is — the packet has no name to match. Sets whose contents this code cannot
// see (local/remote) say nothing either way; the address walk below already refuses
// to conclude anything from them.
func matchesUntunnelable(d option.DefaultRule, sets map[string]ruleSetFacts) bool {
for _, tag := range d.RuleSet {
if sets[tag].hasDomain {
return false
}
}
switch {
case len(d.Domain) > 0, len(d.DomainSuffix) > 0, len(d.DomainKeyword) > 0,
len(d.DomainRegex) > 0, len(d.Geosite) > 0:
@@ -229,6 +360,41 @@ func matchesUntunnelable(d option.DefaultRule) bool {
return true
}
// namedDestinationNote explains the one skip an operator can actually be
// surprised by: a rule that names BOTH a domain list and an address list.
//
// A domain-only rule contributes nothing here and always has, so saying so would be
// noise. A mixed rule is different — its addresses look like something this plan
// could act on, and it is precisely the shape `shaterd migrate` produces from a v1
// rule that carried dst_domain and dst_ip together. The rule is skipped so the
// upgrade cannot quietly change what ping does (see the block above ruleSetFacts);
// that decision is worth one line in the operator's warning list rather than none.
//
// ok=false when there is nothing surprising to report.
func namedDestinationNote(d option.DefaultRule, sets map[string]ruleSetFacts) (string, bool) {
var named []string
addressed := len(d.IPCIDR) > 0
for _, tag := range d.RuleSet {
f, known := sets[tag]
if known && f.hasDomain {
named = append(named, tag)
continue
}
// An unknown/opaque set may well hold addresses; so may a knowable one.
if !known || f.opaque || len(f.cidrs) > 0 {
addressed = true
}
}
if len(named) == 0 || !addressed {
return "", false
}
return fmt.Sprintf(
"a routing rule matches by name (%s) as well as by address, and ping/IPTV/VPN-passthrough traffic "+
"carries no name — so that rule is left out of the ping/IPTV decision entirely and later rules "+
"decide those addresses. Split it into a name rule and an address rule if you want the addresses "+
"decided here", strings.Join(named, ", ")), true
}
// parsePrefixes converts CIDR/bare-address strings to prefixes, dropping anything
// unparseable (it cannot be matched on, so it must not silently widen a set).
func parsePrefixes(in []string) []netip.Prefix {
+156 -9
View File
@@ -22,6 +22,16 @@ func tunnelModel() *model.Model {
return m
}
// pinnedIPSet declares the inline `type=ipcidr` rule-set a rule pins its
// destination addresses with. Since schema v2 a routing rule has no dst_ip of its
// own: addresses are a `config ruleset`, so the generated rule carries a rule_set
// reference and the plan resolves the actual prefixes through the lookup below —
// i.e. through the RUNNING engine in production (engine.RuleSetIPCIDRs), not out
// of the config text. planFor's table stands in for that.
func pinnedIPSet(name string, cidrs ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: cidrs}
}
// planFor generates m and builds the untunnelable plan, resolving rule-set tags
// from the supplied table. A tag absent from the table reports "not loaded".
func planFor(t *testing.T, m *model.Model, sets map[string][]string) *netplane.UntunnelablePlan {
@@ -58,11 +68,12 @@ func renderPlan(t *testing.T, m *model.Model, plan *netplane.UntunnelablePlan) s
// only 8.8.8.8 may be un-pingable and the rest of the internet must answer.
func TestOnlyPinnedAddressIsTunnelled(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}})
if !plan.DefaultAllow {
t.Fatalf("a catch-all `direct` rule must make the default ALLOW; plan=%+v", plan)
@@ -221,14 +232,148 @@ func TestDomainRuleDoesNotAffectUntunnelable(t *testing.T) {
}
}
// TestBlockedRuleDenies: an explicitly blocked destination stays dropped.
func TestBlockedRuleDenies(t *testing.T) {
// --- schema v2: the destination moved into rule-sets, and so did the analysis ---
// migratedRule reproduces exactly what `shaterd migrate` makes of a v1 rule that
// carried BOTH lists: `dst_domain=[bank.ru] dst_ip=[203.0.113.0/24]` becomes two
// inline rule-sets, `rule-x` and `rule-x-ip`, and the rule references both.
func migratedRule(target string) *model.Model {
m := tunnelModel()
m.Rulesets = []model.Ruleset{
{Name: "rule-x", Type: "domain", Source: "inline", Entries: []string{"full:bank.ru"}},
pinnedIPSet("rule-x-ip", "203.0.113.0/24"),
}
m.Rules = []model.Rule{
{Name: "bad", Enabled: true, Order: 10, DstIP: []string{"203.0.113.0/24"}, Target: "block"},
{Name: "x", Enabled: true, Order: 10,
DstRuleset: []string{"rule-x", "rule-x-ip"}, Target: target},
{Name: "dflt", Enabled: true, Order: 99, Target: "group:auto"},
}
return m
}
// TestMigratedDomainAndIPRuleStaysOutOfTheUntunnelablePlan is the upgrade
// regression. In v1 the rule ANDed its domain and its addresses, so it could never
// claim a packet that carries no domain and this plan skipped it. After the
// migration the same rule reads `rule_set: [rule-x, rule-x-ip]`, and rule_set
// entries are ORed — so the address set alone made the rule look applicable and the
// plan started emitting a verdict for 203.0.113.0/24 that nobody asked for.
//
// With target=direct that verdict is an ALLOW, i.e. an upgrade quietly sending
// previously-tunnelled ICMP out with the client's real source address. That is the
// half that makes this a leak and not just a surprise, so it is checked first.
func TestMigratedDomainAndIPRuleStaysOutOfTheUntunnelablePlan(t *testing.T) {
// Both engine states, because they fail differently: with the engine UP the old
// code emitted a live verdict for the addresses, and with it DOWN the same
// mistake hid behind the unresolvable-list truncation. Neither may happen.
for _, target := range []string{"direct", "block", "group:auto"} {
for _, engine := range []string{"up", "down"} {
t.Run(target+"/engine-"+engine, func(t *testing.T) {
m := migratedRule(target)
var loaded map[string][]string
if engine == "up" {
loaded = map[string][]string{"rs-rule-x": {}, "rs-rule-x-ip": {"203.0.113.0/24"}}
}
plan := planFor(t, m, loaded)
if len(plan.Matches) != 0 {
t.Fatalf("a rule that matches by NAME as well as by address must contribute no "+
"address step (it cannot match a packet that carries no name): %+v", plan.Matches)
}
if plan.DefaultAllow {
t.Errorf("the catch-all routes into the tunnel, so the default must still deny")
}
fwd := renderPlan(t, m, plan)
if strings.Contains(fwd, "203.0.113.0/24") {
t.Errorf("the migrated address list must not reach the data plane through the "+
"untunnelable policy:\n%s", fwd)
}
})
}
}
}
// TestMigratedDomainAndIPRuleSaysWhyItWasSkipped: skipping is the safe answer, but
// a silently different ping after an upgrade is the complaint this whole change
// exists to answer. The mixed rule — the exact shape migration produces — is named.
func TestMigratedDomainAndIPRuleSaysWhyItWasSkipped(t *testing.T) {
plan := planFor(t, migratedRule("direct"), nil)
var seen bool
for _, w := range plan.Warnings {
if strings.Contains(w, "matches by name") && strings.Contains(w, "rs-rule-x") {
seen = true
}
}
if !seen {
t.Fatalf("the skip must be explained and the list named; warnings=%v", plan.Warnings)
}
}
// TestDomainOnlyRuleSetIsQuiet: a rule whose only destination is a NAME list has
// always contributed nothing here and is not a surprise, so it must not produce a
// note. Only the mixed shape is worth a line.
func TestDomainOnlyRuleSetIsQuiet(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{
{Name: "names", Type: "domain", Source: "inline", Entries: []string{"full:bank.ru"}},
}
m.Rules = []model.Rule{
{Name: "names", Enabled: true, Order: 10, DstRuleset: []string{"names"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
if !plan.DefaultAllow {
t.Errorf("a name-only rule must not withhold the catch-all allow; plan=%+v", plan)
}
for _, w := range plan.Warnings {
if strings.Contains(w, "matches by name") {
t.Errorf("a name-only rule is not a surprise and needs no note: %q", w)
}
}
}
// TestInlineRuleSetNeedsNoEngine: an inline rule-set's addresses are sitting in the
// generated config. Asking the engine for them — and, while it is down, truncating
// the whole walk so ping/IPTV/VPN passthrough dies everywhere — was a conservative
// answer to a question that did not have to be asked. The verdict must now be the
// same whether or not the engine is up.
func TestInlineRuleSetNeedsNoEngine(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
down := planFor(t, m, nil) // engine not up
up := planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}}) // engine up
for name, plan := range map[string]*netplane.UntunnelablePlan{"engine-down": down, "engine-up": up} {
if !plan.DefaultAllow {
t.Fatalf("%s: the catch-all direct must allow the rest of the internet; plan=%+v", name, plan)
}
if len(plan.Matches) != 1 || plan.Matches[0].Allow {
t.Fatalf("%s: the tunnelled address must produce one DENY step, got %+v", name, plan.Matches)
}
if got := plan.Matches[0].Dst4; len(got) != 1 || got[0] != "8.8.8.8/32" {
t.Fatalf("%s: step destinations = %v, want [8.8.8.8/32]", name, got)
}
}
for _, w := range down.Warnings {
if strings.Contains(w, "not loaded yet") {
t.Errorf("an inline list is never 'not loaded': %q", w)
}
}
}
// TestBlockedRuleDenies: an explicitly blocked destination stays dropped.
func TestBlockedRuleDenies(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("bad", "203.0.113.0/24")}
m.Rules = []model.Rule{
{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "block"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, map[string][]string{"rs-bad": {"203.0.113.0/24"}})
if len(plan.Matches) != 1 || plan.Matches[0].Allow {
t.Fatalf("a blocked destination must deny untunnelable traffic too: %+v", plan.Matches)
}
@@ -361,11 +506,12 @@ func TestPlanNeverAcceptsTCPOrUDP(t *testing.T) {
m := tunnelModel()
m.Globals.IPv6 = true
m.Globals.Untunnelable = policy
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
fwd := renderPlan(t, m, planFor(t, m, nil))
fwd := renderPlan(t, m, planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}}))
// Only the lines the untunnelable policy emits are in scope: the tproxy
// diverts in prerouting legitimately match TCP/UDP, which is their job.
@@ -424,12 +570,13 @@ func TestLocalPlaneSurvivesEveryPlan(t *testing.T) {
// only grant those clients, not everyone.
func TestSourceScopedRuleNarrowsTheAllow(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("lab", "198.51.100.0/24")}
m.Rules = []model.Rule{
{Name: "lab", Enabled: true, Order: 10, Src: []string{"192.168.9.0/24"},
DstIP: []string{"198.51.100.0/24"}, Target: "direct"},
DstRuleset: []string{"lab"}, Target: "direct"},
{Name: "dflt", Enabled: true, Order: 99, Target: "group:auto"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-lab": {"198.51.100.0/24"}})
if len(plan.Matches) != 1 {
t.Fatalf("expected one step, got %+v", plan.Matches)
}
+46
View File
@@ -24,7 +24,9 @@ import (
"sort"
"strconv"
"strings"
"time"
"github.com/sagernet/sing-box/shater/engine"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
)
@@ -269,6 +271,50 @@ func untunnelablePolicyWarnings(g model.Globals, planNotes []string) []Warning {
}
}
// engineTeardownWarnings turns the engine's ABANDONED generations — superseded
// sing-box instances whose shutdown overran the hard close budget and are still
// running inside this process — into operator-facing warnings.
//
// Critical, without hesitation. A leaked generation is not untidiness:
//
// - it still holds its WireGuard devices, and two devices built from the same
// private key evict each other at the peer (one session per public key), so
// the leak reproduces BETWEEN generations exactly the fault
// generate/wgdedup.go removes WITHIN a config — the tunnel flaps and neither
// end can say why;
// - it still holds its outbound connections and keeps probing nodes, so the
// log fills with errors attributed to a config that is no longer applied;
// - on a 512 MiB router each one costs real memory that is never returned.
//
// The generation number is carried in Name so two consecutive status reads can
// tell "the same stuck generation" from "another one just leaked", and the
// elapsed time is in the message because a shutdown at 8s and one at 40 minutes
// are different problems.
func engineTeardownWarnings(stuck []engine.StuckClose) []Warning {
if len(stuck) == 0 {
return nil
}
out := make([]Warning, 0, len(stuck))
for _, s := range stuck {
config := "unknown config"
if len(s.Hash) >= 12 {
config = "config " + s.Hash[:12]
}
out = append(out, Warning{
Severity: SeverityCritical,
Section: "engine",
Name: fmt.Sprintf("generation %d", s.Generation),
Message: fmt.Sprintf(
"a superseded engine instance (%s) has been shutting down for %s and is STILL RUNNING: "+
"it keeps its outbound connections and its WireGuard devices, so it can evict the "+
"live tunnel at the peer and it keeps writing to the log. The current configuration "+
"is applied and running; restart shaterd if this does not clear.",
config, s.Elapsed.Round(time.Second)),
})
}
return out
}
func severityRank(s string) int {
switch s {
case SeverityCritical:
+206
View File
@@ -0,0 +1,206 @@
// Package buildtags is the contract between what shater DECLARES it supports
// and the build tags the shipped router binary is actually compiled with.
//
// # Why this package exists
//
// The router binary is built with a deliberately trimmed tag set (D9/D23,
// scripts/router-tags.sh) — upstream's full set registers a zoo shater/generate
// can never emit, and a router pays for every tag in flash and in RAM. Trimming
// is right; trimming BLIND is not. On 2026-07-25 a production router answered a
// configured WireGuard node with
//
// create instance: initialize endpoint[0]: create WireGuard device:
// gVisor is not included in this build, rebuild with -tags with_gvisor
//
// because `with_gvisor` had been trimmed as "unreachable code" (true for the tun
// inbound we never emit — false for the WireGuard endpoint we ship and declare
// [MVP]) while `with_wireguard` stayed. Nothing caught it: the test suite builds
// with the FULL upstream tag set, so the SHIPPED tag combination was, at that
// point, the one configuration nothing in the repo ever exercised.
//
// # What holds it together now
//
// 1. Features below names each declared feature and the build tags it needs to
// RUN (not merely to compile). shater/buildtags's own test parses
// scripts/router-tags.sh and fails if the shipped set does not cover them —
// it needs no tags, no Linux and no network, so it runs in every plain
// `go test ./...`.
// 2. shater/generate's TestShippedTagSetConstructsDeclaredProtocols drives one
// node of every declared protocol through box.New under whatever tags the
// test binary was built with, skipping only what is genuinely not compiled
// in. scripts/check-router-tags.sh runs it with the SHIPPED set, so the
// combination we ship is proven to construct, not merely to link.
//
// (1) catches a trimmed dependency the moment it is trimmed; (2) catches the
// class of failure (1) cannot model — a tag that is present but insufficient.
//
// Adding a protocol to shater/parse + shater/generate means adding a row here.
package buildtags
import "sort"
// Feature is one capability the product declares, together with the build tags
// the binary must carry for it to work at runtime.
type Feature struct {
// Name is the feature as a user would name it.
Name string
// Declared points at where we promise it (docs-shater/FEATURES.md section,
// or the generator/registry that emits it).
Declared string
// Tags are ALL build tags required for the feature to work — including
// transitive ones (with_awg alone is useless without with_wireguard, which
// is useless without with_gvisor). Listing them transitively is deliberate:
// the check must not depend on a dependency graph nobody maintains.
Tags []string
// Why explains what breaks without those tags, with the code anchor. It is
// printed by the failing test, so a future trimmer reads the reason instead
// of rediscovering it on a router.
Why string
}
// Features is the authoritative list. Only tag-GATED capabilities belong here:
// tproxy, routing rules, rule-sets, the DNS filter, nft/policy routing and the
// panel are compiled unconditionally and cannot be lost to a tag trim.
var Features = []Feature{
{
Name: "WireGuard nodes (wg:// / wireguard:// links, wg-quick .conf import)",
Declared: "FEATURES.md §Proxy engine — “VLESS, VMess, Trojan, Shadowsocks, WireGuard” [MVP]",
Tags: []string{"with_wireguard", "with_gvisor"},
Why: "with_wireguard registers the endpoint (shater/registry/registry_wireguard.go); " +
"with_gvisor supplies the userspace netstack EVERY WireGuard device needs — without it " +
"transport/wireguard/device_stack_stub.go returns tun.ErrGVisorNotIncluded from BOTH " +
"newStackDevice and newSystemStackDevice, so box.New fails with " +
"\"create WireGuard device: gVisor is not included in this build\" and the node is dead. " +
"system_interface=true is not an escape hatch: it hits the same stub.",
},
{
Name: "AmneziaWG obfuscation (awg:// links; jc/jmin/jmax, s1-s4, h1-h4, i1-i5)",
Declared: "FEATURES.md §Proxy engine — “AmneziaWG 2.0 … a driving requirement” [MVP]",
Tags: []string{"with_awg", "with_wireguard", "with_gvisor"},
Why: "with_awg makes the AWG params reach the device (transport/wireguard/device_awg.go); " +
"without it they parse and are silently ignored (option/wireguard.go). It rides on the " +
"WireGuard endpoint, so it needs that feature's tags too.",
},
{
Name: "Hysteria2 nodes (hysteria2:// / hy2://)",
Declared: "FEATURES.md §Proxy engine [T1]; shater/registry registerQUICOutbounds",
Tags: []string{"with_quic"},
Why: "hysteria2.RegisterOutbound is compiled only under with_quic (shater/registry/registry_quic.go); without it box.New rejects the outbound as an unknown type.",
},
{
Name: "TUIC nodes (tuic://)",
Declared: "FEATURES.md §Proxy engine [T1]; shater/registry registerQUICOutbounds",
Tags: []string{"with_quic"},
Why: "tuic.RegisterOutbound is compiled only under with_quic (shater/registry/registry_quic.go).",
},
{
Name: "VLESS/VMess over the QUIC v2ray transport (type=quic)",
Declared: "FEATURES.md §Proxy engine — “Transports: TCP/WS/gRPC/HTTPUpgrade/H2/QUIC” [MVP]",
Tags: []string{"with_quic"},
Why: "transport/v2rayquic registers its constructor from an init() blank-imported only under with_quic; without it NewQUICClient returns os.ErrInvalid at dial time.",
},
{
Name: "QUIC / HTTP3 DNS transports (quic://, h3://)",
Declared: "shater/registry registerQUICTransports",
Tags: []string{"with_quic"},
Why: "dns/transport/quic is registered only under with_quic (shater/registry/registry_quic.go).",
},
{
Name: "REALITY (vless security=reality, pbk/sid)",
Declared: "FEATURES.md §Proxy engine — “Reality/XTLS” [MVP]; shater/parse security=reality",
Tags: []string{"with_utls"},
Why: "the REALITY client lives in common/tls/reality_client.go, which is itself `//go:build with_utls`; without the tag a reality config is rejected by the TLS layer.",
},
{
Name: "uTLS ClientHello fingerprints (fp=chrome/firefox/safari/…)",
Declared: "shater/generate/outbound.go TLS mapping (UTLS options)",
Tags: []string{"with_utls"},
Why: "common/tls/utls_client.go is `//go:build with_utls`; the stub (utls_stub.go) refuses a config that sets a fingerprint.",
},
{
Name: "XHTTP / SplitHTTP transport (type=xhttp, type=splithttp)",
Declared: "FEATURES.md §Proxy engine [T1]; shater/parse/sharelink.go case \"xhttp\"",
Tags: []string{"with_xhttp"},
Why: "transport/v2rayxhttp registers the \"xhttp\" transport from an init() blank-imported only under with_xhttp (shater/registry/registry_xhttp.go); without it the transport type is unknown at box.New.",
},
{
Name: "badtls fast path (zero-copy TLS read-wait / ktls, used by every TLS outbound)",
Declared: "common/badtls — linked unconditionally by the TLS client",
Tags: []string{"badlinkname", "tfogo_checklinkname0"},
Why: "common/badtls/*.go are `go1.25 && badlinkname`; without the tag the package degrades to read_wait_stub.go. " +
"These two tags additionally REQUIRE -checklinkname=0 in the linker flags — the build fails at link time otherwise " +
"(\"invalid reference to crypto/tls.(*Conn).handlePostHandshakeMessage\"), which is why " +
"scripts/router-tags.sh carries SHATER_ROUTER_LDFLAGS next to the tag set.",
},
}
// RequiredTags is the union of every declared feature's tags, sorted.
func RequiredTags() []string {
seen := map[string]bool{}
for _, f := range Features {
for _, t := range f.Tags {
seen[t] = true
}
}
return sortedKeys(seen)
}
// Compiled reports the shater-relevant build tags THIS binary was compiled with,
// sorted. It is populated by the one-line tag_*.go twins in this package; a tag
// with no file here is simply not tracked (and must not appear in Features).
func Compiled() []string { return sortedKeys(compiled) }
// Has reports whether this binary was compiled with tag.
func Has(tag string) bool { return compiled[tag] }
// MissingTags returns the tags f needs that this binary lacks, sorted. Empty
// means the feature is fully compiled in.
func MissingTags(f Feature) []string {
missing := map[string]bool{}
for _, t := range f.Tags {
if !compiled[t] {
missing[t] = true
}
}
return sortedKeys(missing)
}
// Tracked reports whether tag has a detector file (tag_*.go) in this package.
// Features must only reference tracked tags — an untracked tag would silently
// read as "not compiled" and turn a real check into a skip. TestFeatureTagsAreTracked
// enforces that, and scripts/check-router-tags.sh additionally proves the
// detectors match the tag set the compiler was actually handed.
func Tracked(tag string) bool { return tracked[tag] }
// TrackedTags is every tag this package can observe, i.e. exactly the tags with
// a tag_*.go detector. Keep the two in sync — the check script fails loudly if
// they drift.
func TrackedTags() []string { return sortedKeys(tracked) }
var tracked = map[string]bool{
"with_gvisor": true,
"with_quic": true,
"with_wireguard": true,
"with_awg": true,
"with_utls": true,
"with_xhttp": true,
"with_lx_command": true,
"badlinkname": true,
"tfogo_checklinkname0": true,
}
// compiled is filled by the tag_*.go detectors' init(). A tag with no detector
// file compiled in is absent from the map, which reads as "not compiled".
var compiled = map[string]bool{}
// mark records that tag is compiled into this binary.
func mark(tag string) { compiled[tag] = true }
func sortedKeys(m map[string]bool) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
sort.Strings(out)
return out
}
+175
View File
@@ -0,0 +1,175 @@
package buildtags
import (
"os"
"path/filepath"
"regexp"
"sort"
"strings"
"testing"
)
// repoFile reads a file relative to the repo root (this package sits at
// <repo>/shater/buildtags).
func repoFile(t *testing.T, rel string) string {
t.Helper()
b, err := os.ReadFile(filepath.Join("..", "..", filepath.FromSlash(rel)))
if err != nil {
t.Fatalf("read %s: %v", rel, err)
}
return string(b)
}
// shVar pulls VAR="…" out of a POSIX sh fragment.
func shVar(t *testing.T, script, name string) string {
t.Helper()
re := regexp.MustCompile(`(?m)^` + regexp.QuoteMeta(name) + `="([^"]*)"`)
m := re.FindStringSubmatch(script)
if m == nil {
t.Fatalf("scripts/router-tags.sh: %s=\"…\" not found (single line, double quotes)", name)
}
return m[1]
}
// routerTagSet returns the shipped tag set as a set, read from the ONE file that
// defines it.
func routerTagSet(t *testing.T) map[string]bool {
t.Helper()
set := map[string]bool{}
for _, tag := range strings.Split(shVar(t, repoFile(t, "scripts/router-tags.sh"), "SHATER_ROUTER_TAGS"), ",") {
if tag = strings.TrimSpace(tag); tag != "" {
set[tag] = true
}
}
if len(set) == 0 {
t.Fatal("SHATER_ROUTER_TAGS is empty")
}
return set
}
// TestRouterTagSetCoversDeclaredFeatures is THE guard the 2026-07-25 WireGuard
// outage was missing (D23): it reads the tag set the router binary is actually
// built with and fails if a feature we DECLARE supported has lost the build tag
// it needs to run.
//
// It deliberately needs no build tags, no Linux, no privileges and no network,
// so it runs in every plain `go test ./...` — including on the Windows dev host,
// where nothing else can exercise the shipped configuration. The behavioural
// half (does the shipped combination actually CONSTRUCT?) is
// shater/generate.TestShippedTagSetConstructsDeclaredProtocols, run with this
// same set by scripts/check-router-tags.sh.
func TestRouterTagSetCoversDeclaredFeatures(t *testing.T) {
shipped := routerTagSet(t)
for _, f := range Features {
var missing []string
for _, tag := range f.Tags {
if !shipped[tag] {
missing = append(missing, tag)
}
}
if len(missing) > 0 {
t.Errorf("the shipped router binary would NOT support a feature we declare.\n"+
" feature : %s\n"+
" declared: %s\n"+
" missing : %s (not in SHATER_ROUTER_TAGS, scripts/router-tags.sh)\n"+
" why : %s\n"+
"Either add the tag back, or stop declaring the feature — those are the only two honest options.",
f.Name, f.Declared, strings.Join(missing, ", "), f.Why)
}
}
}
// TestFeatureTagsAreTracked keeps Features honest: every tag it names must have
// a tag_*.go detector, or Compiled()/MissingTags() would report it absent even
// when it is compiled in — and the behavioural test would silently SKIP the
// feature instead of checking it. A false green is worse than a red.
func TestFeatureTagsAreTracked(t *testing.T) {
for _, f := range Features {
for _, tag := range f.Tags {
if !Tracked(tag) {
t.Errorf("feature %q requires tag %q, which has no detector: add shater/buildtags/tag_%s.go and the entry in the tracked map", f.Name, tag, tag)
}
}
}
}
// TestTrackedTagsHaveDetectorFiles pairs the tracked map with the files on disk,
// so a renamed/deleted detector cannot quietly make a tag read as absent.
func TestTrackedTagsHaveDetectorFiles(t *testing.T) {
for _, tag := range TrackedTags() {
name := "tag_" + tag + ".go"
body, err := os.ReadFile(name)
if err != nil {
t.Errorf("tracked tag %q has no detector file %s: %v", tag, name, err)
continue
}
if !strings.Contains(string(body), "//go:build "+tag) || !strings.Contains(string(body), `mark("`+tag+`")`) {
t.Errorf("%s must be `//go:build %s` and call mark(%q)", name, tag, tag)
}
}
files, err := filepath.Glob("tag_*.go")
if err != nil {
t.Fatal(err)
}
for _, f := range files {
tag := strings.TrimSuffix(strings.TrimPrefix(f, "tag_"), ".go")
if !Tracked(tag) {
t.Errorf("detector %s exists but %q is not in the tracked map", f, tag)
}
}
}
// TestBuildScriptUsesTheSharedTagSet stops the split that caused the outage from
// coming back: the ship build must SOURCE scripts/router-tags.sh, not carry its
// own copy of the tag list. A second copy is a second truth, and the second one
// is the one nobody checks.
func TestBuildScriptUsesTheSharedTagSet(t *testing.T) {
build := repoFile(t, "scripts/build-shaterd.sh")
if !strings.Contains(build, "router-tags.sh") {
t.Fatal("scripts/build-shaterd.sh must source scripts/router-tags.sh")
}
if regexp.MustCompile(`(?m)^\s*ROUTER_TAGS="with_`).MatchString(build) {
t.Fatal("scripts/build-shaterd.sh re-inlines a literal tag list; the set must come from scripts/router-tags.sh only")
}
}
// TestRouterLdflagsSatisfyTagRequirements: `badlinkname` is not self-contained —
// the LINK step fails without -checklinkname=0. The flag therefore belongs to
// the tag set, and lives beside it; assert the pair never separates.
func TestRouterLdflagsSatisfyTagRequirements(t *testing.T) {
script := repoFile(t, "scripts/router-tags.sh")
ldflags := shVar(t, script, "SHATER_ROUTER_LDFLAGS")
if routerTagSet(t)["badlinkname"] && !strings.Contains(ldflags, "-checklinkname=0") {
t.Fatalf("SHATER_ROUTER_TAGS carries badlinkname but SHATER_ROUTER_LDFLAGS (%q) lacks -checklinkname=0: the build will fail at link time", ldflags)
}
if !strings.Contains(repoFile(t, "scripts/build-shaterd.sh"), "SHATER_ROUTER_LDFLAGS") {
t.Fatal("scripts/build-shaterd.sh must use $SHATER_ROUTER_LDFLAGS, not a hand-copied -checklinkname=0")
}
}
// TestCompiledTagsMatchTheShippedSet proves the DETECTORS are telling the truth:
// when the test binary is compiled with exactly the shipped tag set, Compiled()
// must equal that set (restricted to tracked tags). Without this, a typo'd or
// deleted detector would make the behavioural test skip a protocol and pass.
//
// It only runs under scripts/check-router-tags.sh (which compiles with that very
// set and exports SHATER_ROUTER_TAG_CHECK=1); a plain `go test ./...` compiles
// with no tags at all, where the comparison is meaningless.
func TestCompiledTagsMatchTheShippedSet(t *testing.T) {
if os.Getenv("SHATER_ROUTER_TAG_CHECK") != "1" {
t.Skip("not a router-tag-set run; use scripts/check-router-tags.sh")
}
var want []string
for tag := range routerTagSet(t) {
if Tracked(tag) {
want = append(want, tag)
}
}
sort.Strings(want)
got := Compiled()
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("compiled tags do not match the shipped set\n compiled: %v\n shipped : %v\n"+
"Either the build ran with the wrong -tags, or a tag_*.go detector is broken.", got, want)
}
}
+7
View File
@@ -0,0 +1,7 @@
//go:build badlinkname
package buildtags
// Detector for the badlinkname build tag — see buildtags.go. There is no !badlinkname twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("badlinkname") }
@@ -0,0 +1,7 @@
//go:build tfogo_checklinkname0
package buildtags
// Detector for the tfogo_checklinkname0 build tag — see buildtags.go. There is no !tfogo_checklinkname0 twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("tfogo_checklinkname0") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_awg
package buildtags
// Detector for the with_awg build tag — see buildtags.go. There is no !with_awg twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_awg") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_gvisor
package buildtags
// Detector for the with_gvisor build tag — see buildtags.go. There is no !with_gvisor twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_gvisor") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_lx_command
package buildtags
// Detector for the with_lx_command build tag — see buildtags.go. There is no !with_lx_command twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_lx_command") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_quic
package buildtags
// Detector for the with_quic build tag — see buildtags.go. There is no !with_quic twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_quic") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_utls
package buildtags
// Detector for the with_utls build tag — see buildtags.go. There is no !with_utls twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_utls") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_wireguard
package buildtags
// Detector for the with_wireguard build tag — see buildtags.go. There is no !with_wireguard twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_wireguard") }
+7
View File
@@ -0,0 +1,7 @@
//go:build with_xhttp
package buildtags
// Detector for the with_xhttp build tag — see buildtags.go. There is no !with_xhttp twin:
// an absent detector means "not compiled in", which is exactly the truth.
func init() { mark("with_xhttp") }
+7
View File
@@ -68,6 +68,12 @@ type nodeView struct {
// ParseError is the reason the engine will SKIP this node, verbatim from
// parse.ParseShareLink. Empty on every usable node.
ParseError string `json:"parse_error,omitempty"`
// ParseWarnings lists what the share link asked for that the engine cannot
// do (hysteria2 port hopping, tuic congestion control, …). The node IS
// usable — that is the difference from parse_error — but it does not behave
// exactly as its link describes, and that gap belongs on screen rather than
// in a code comment.
ParseWarnings []string `json:"parse_warnings,omitempty"`
}
// readModel is the model source. A package var so the tests can exercise the
@@ -111,6 +117,7 @@ func nodeViews(m *model.Model) []nodeView {
}
v.Server = p.Server
v.Port = p.Port
v.ParseWarnings = p.Warnings
} else if n.URI != "" {
v.ParseError = err.Error()
}
+126 -31
View File
@@ -61,6 +61,25 @@ type Engine struct {
current option.Options // options the running Box was built from
hash string // stable hash of current (hex sha256 of canonical JSON)
// instanceCancel cancels the context THIS box was built on — its own
// cancellable child of e.ctx, one per box (see newBox). box.New does not
// derive a cancellable context of its own, so without this every goroutine
// inside a retired box that waits on ctx.Done() waits forever. Upstream's
// runner does exactly this and calls cancel before Close
// (cmd/sing-box/cmd_run.go). nil when nothing is running.
instanceCancel context.CancelFunc
// instanceGen is the 1-based sequence number of the running box, bumped on
// every adopted swap. It is what names a generation in the log and in
// PendingCloses when a shutdown overruns its budget.
instanceGen uint64
// pending holds the retirements in flight (see teardown.go). Its OWN leaf
// lock, deliberately not mu: PendingCloses is read by the status path, and
// the moment that read matters most is while an apply is holding mu waiting
// out a shutdown that will not finish.
pendingMu sync.Mutex
pending []*pendingClose
// defaultLogWriter, when non-nil, is handed to EVERY box.New this engine
// performs (box.Options.DefaultLogWriter): the daemon points it at its
// long-lived logsink once, and each Apply-swapped box then logs into that
@@ -174,6 +193,17 @@ func (e *Engine) Apply(opts option.Options) (changed bool, err error) {
}
func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// Stand down the self-check of urltest groups no rule reaches (see
// selfcheck.go for the whole argument). This MUTATES opts in place — the
// option structs are pointers behind an `any` — and it must run BEFORE the
// hash below, so the hash describes the config that is really built: a rule
// change that flips a group used<->unused is then a real change that
// triggers a swap, and an unchanged config hashes identically on every
// reconcile because the stand-down is deterministic.
if stood := standDownUnusedSelfCheck(opts); stood > 0 && e.log != nil {
e.log.Info("apply: stood down self-check on ", stood, " unused urltest group(s); the observatory is their only prober")
}
newHash, err := e.hashOptions(opts)
if err != nil {
return false, E.Cause(err, "hash options")
@@ -184,12 +214,18 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
return false, nil
}
// (2) build + validate. box.New constructs and validates every adapter.
nb, err := box.New(box.Options{
Context: e.ctx,
Options: opts,
DefaultLogWriter: e.defaultLogWriter,
})
// (1b) Do not stack generations. A superseded box whose shutdown overran its
// budget is still holding its outbound connections and its WireGuard devices;
// building another one on top of it is how one process ends up carrying four
// live generations. Give any abandoned teardown one more budget to finish
// BEFORE we create anything (see teardown.go awaitAbandonedLocked). Placed
// after the hash gate on purpose: a no-op reconcile — cron, every minute —
// must stay free.
stuck := e.awaitAbandonedLocked()
// (2) build + validate on its OWN cancellable context. box.New constructs and
// validates every adapter.
nb, nbCancel, err := e.newBox(opts)
if err != nil {
// Validation failed: keep the running instance, do not swap.
return false, E.Cause(err, "create instance")
@@ -211,8 +247,15 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// close-old-then-start-new DIRECTLY — skipping the stall. (box.New above
// already validated opts, so we never tear the old box down for a config
// that would fail to build.)
if e.instance != nil && sharesCacheFileLock(e.current, opts) {
return e.closeOldThenStart(nb, opts, newHash)
//
// A generation that is STILL abandoned after the wait above forces the same
// path for a different reason: start-new-first would put a second LIVE box
// alongside a third that refuses to die, all three contending for the same
// tproxy port, the same cache file and — the expensive one — the same
// WireGuard private keys. Closing the current box first keeps the process to
// at most one live instance plus the abandoned one.
if e.instance != nil && (stuck > 0 || sharesCacheFileLock(e.current, opts)) {
return e.closeOldThenStart(nb, nbCancel, opts, newHash)
}
// (3b) start-new-first (zero-downtime when there is no resource conflict).
@@ -223,25 +266,59 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// close-old-then-start-new then. Any OTHER Start failure keeps today's
// behavior: discard the new box, keep the old one running.
if e.instance == nil || !isSwapConflict(err) {
nbCancel()
_ = nb.Close()
return false, E.Cause(err, "start instance")
}
return e.closeOldThenStart(nb, opts, newHash)
return e.closeOldThenStart(nb, nbCancel, opts, newHash)
}
// Swap succeeded. The previously running config becomes last-good.
old := e.instance
if old != nil {
if e.instance != nil {
e.lastGood = e.current
e.hasLastGood = true
_ = old.Close()
// Retire the old generation under the hard budget. Its error — including
// ErrCloseTimeout — is deliberately NOT returned: the new box is started
// and carrying traffic, so this apply SUCCEEDED, and failing it here would
// abort the caller's netplane stage and leave a stale ruleset loaded over a
// perfectly healthy engine. An abandoned generation is surfaced through
// PendingCloses() (critical warning in `shaterd status` and the panel) and
// an ERROR line naming it — visible, but not mistaken for a failed apply.
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
}
e.instance = nb
e.current = opts
e.hash = newHash
e.adoptLocked(nb, nbCancel, opts, newHash)
return true, nil
}
// newBox builds a box on its OWN cancellable child of the engine context and
// returns the cancel alongside it. Every goroutine the box starts inherits that
// context, so cancelling it is what unwinds the ones Close does not reach; see
// teardown.go for why the shared, never-cancelled context was the defect.
func (e *Engine) newBox(opts option.Options) (*box.Box, context.CancelFunc, error) {
ctx, cancel := context.WithCancel(e.ctx)
b, err := box.New(box.Options{
Context: ctx,
Options: opts,
DefaultLogWriter: e.defaultLogWriter,
})
if err != nil {
cancel()
return nil, nil, err
}
return b, cancel, nil
}
// adoptLocked installs a started box as THE running instance and gives it the
// next generation number. Caller holds e.mu and has already retired whatever was
// running before.
func (e *Engine) adoptLocked(b *box.Box, cancel context.CancelFunc, opts option.Options, hash string) {
e.instanceGen++
e.instance = b
e.instanceCancel = cancel
e.current = opts
e.hash = hash
}
// closeOldThenStart is the close-old-then-start-new swap. It is taken both
// proactively (the incoming config shares the running box's cache_file lock, so
// start-new-first cannot work) and reactively (start-new-first hit a swap
@@ -256,27 +333,36 @@ func (e *Engine) applyLocked(opts option.Options) (bool, error) {
// already validated by the caller's box.New, so the rebuild below is expected to
// succeed; the restore path guards the rare case it does not. The caller holds
// e.mu.
func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHash string) (bool, error) {
func (e *Engine) closeOldThenStart(discard *box.Box, discardCancel context.CancelFunc, opts option.Options, newHash string) (bool, error) {
// (a) Drop the pre-built box (a Start-failed box cannot be restarted, and the
// proactively-built one must not hold the cache_file lock while we rebuild).
// It was never adopted, so it is not a generation — close it inline, but
// cancel its context first exactly like a retired one.
if discardCancel != nil {
discardCancel()
}
if discard != nil {
_ = discard.Close()
}
// (b) Free the port + cache_file lock by closing the old instance. Remember
// its config so we can restore it if the fresh box cannot come up.
// (b) Free the port + cache_file lock by closing the old instance, under the
// hard budget. Remember its config so we can restore it if the fresh box
// cannot come up. An overrun here is reported by retireLocked and tracked in
// PendingCloses; we still proceed, because the resources it was supposed to
// free are exactly what the operator is waiting on.
prevOpts := e.current
prevHash := e.hash
prevLastGood := e.lastGood
prevHasLastGood := e.hasLastGood
_ = e.instance.Close()
e.instance = nil
_ = e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, prevHash)
e.instance, e.instanceCancel = nil, nil
// (c) Build a FRESH box for opts (the discarded one cannot be reused).
nb2, err := box.New(box.Options{Context: e.ctx, Options: opts, DefaultLogWriter: e.defaultLogWriter})
nb2, cancel2, err := e.newBox(opts)
if err == nil {
err = nb2.Start()
if err != nil {
cancel2()
_ = nb2.Close()
}
}
@@ -284,28 +370,25 @@ func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHas
// (d) Success: the old config we just closed becomes last-good.
e.lastGood = prevOpts
e.hasLastGood = true
e.instance = nb2
e.current = opts
e.hash = newHash
e.adoptLocked(nb2, cancel2, opts, newHash)
return true, nil
}
// (e) The fresh box could not come up and the old one is already closed —
// interception is currently down. Try to RESTORE the previous config so we
// do not leave the tunnel dead.
rb, rerr := box.New(box.Options{Context: e.ctx, Options: prevOpts, DefaultLogWriter: e.defaultLogWriter})
rb, rcancel, rerr := e.newBox(prevOpts)
if rerr == nil {
rerr = rb.Start()
if rerr != nil {
rcancel()
_ = rb.Close()
}
}
if rerr == nil {
// Old config restored: keep current/hash/last-good exactly as they were
// (do NOT advance them). Report that opts was not applied.
e.instance = rb
e.current = prevOpts
e.hash = prevHash
e.adoptLocked(rb, rcancel, prevOpts, prevHash)
e.lastGood = prevLastGood
e.hasLastGood = prevHasLastGood
return false, E.Cause(err, "start instance (config not applied; previous config restored)")
@@ -314,7 +397,7 @@ func (e *Engine) closeOldThenStart(discard *box.Box, opts option.Options, newHas
// Restore ALSO failed: the engine is now STOPPED. e.instance stays nil (never
// pointing at a closed box). The fail-closed nft kill-switch keeps the LAN
// safe (no unproxied leak) even though interception is down.
e.instance = nil
e.instance, e.instanceCancel = nil, nil
e.hash = ""
return false, E.Cause(E.Errors(err, rerr), "start instance failed and could not restore previous config; engine stopped")
}
@@ -410,14 +493,26 @@ func (e *Engine) HasLastGood() bool {
}
// Close stops the running instance, if any. It is idempotent.
//
// Unlike the swap path, this one DOES return ErrCloseTimeout: here the shutdown
// is the whole operation, so "it did not stop" is the result, not a footnote. The
// caller (apply.Teardown) records it as the teardown's error while still
// completing the netplane teardown — the data plane must come down even when a
// box will not.
func (e *Engine) Close() error {
e.mu.Lock()
defer e.mu.Unlock()
if e.instance == nil {
// Nothing running, but a previously abandoned generation may still be:
// give it a last budget so a teardown followed by a restart does not carry
// the leak across.
if e.awaitAbandonedLocked() > 0 {
return ErrCloseTimeout
}
return nil
}
err := e.instance.Close()
e.instance = nil
err := e.retireLocked(e.instance, e.instanceCancel, e.instanceGen, e.hash)
e.instance, e.instanceCancel = nil, nil
e.hash = ""
return err
}
+203 -14
View File
@@ -1,6 +1,7 @@
package engine
import (
"sort"
"strings"
"github.com/sagernet/sing-box/adapter"
@@ -320,11 +321,15 @@ func parseGroupCopyTag(group, tag string) (member string, ok bool) {
return member, true
}
// ChainHealth is one configured chain's reachability, mirroring GroupHealth.Used
// for the chain card (plan §5.E): a chain no enabled rule routes through is never
// probed — the observatory walks only reachable paths — and the panel renders it
// "unused" rather than as a health problem. A chain has no membership counters: it
// is a fixed path, and its end-to-end health is the exit test's job, not a roll-up.
// ChainHealth is one configured chain's reachability plus its PER-HOP health,
// mirroring GroupHealth.Used for the chain card (plan §5.E): a chain no enabled
// rule routes through is never probed — the observatory walks only reachable
// paths — and the panel renders it "unused" rather than as a health problem.
// A chain has no membership counters of its own: it is a fixed path, and its
// end-to-end health is the exit probe's job. What it DOES have is hops, and
// since the observatory now probes every hop wrapper (probeplan.go walkDetour),
// each hop's health is on the board and is projected here so an operator can
// see WHICH hop died instead of only that the chain did.
type ChainHealth struct {
// Name is the chain's model name (config chain "chain:<name>"), what the Targets
// page lists and what a rule targets.
@@ -336,34 +341,218 @@ type ChainHealth struct {
// when the observatory is disabled or not yet configured: no badge is better
// than a wrong one (the same rule as GroupHealth.Used).
Used bool `json:"used"`
// Hops is the per-hop health readout, L1..Ln in wire order (see ChainHopHealth).
// It is EMPTY for a chain the running box never materialised: a chain no rule
// references is resolved lazily and never built, and a 1-hop chain without an
// egress entry resolves straight to its target with no wrapper — in both cases
// there are no "chain-<name>-h…" outbounds to project. An absent "hops" key
// therefore means "nothing materialised to report on", NEVER "this chain has
// no hops" — the model, not this projection, knows how many hops were
// configured.
Hops []ChainHopHealth `json:"hops,omitempty"`
}
// ChainHealth reports the reachability (used/unused) of every named chain, one row
// per name, in the order given. It is the chain analogue of GroupHealth.Used: a
// chain the observatory's used-set does not cover is reported Used=false so the
// panel can mark it "unused" instead of running an exit test against a path nothing
// routes through.
// ChainHopHealth is one hop of one materialised chain, as the health board saw
// it — a projection, like everything in this file: nothing here dials.
//
// A NODE hop is a single measurement: the observatory dials the hop wrapper
// "chain-<name>-h<i>", which pulls exactly the path prefix up to and including
// this hop, so Total=1, the counters follow the hop's own state, DelayMs and
// AgeSeconds are its own observation, and Selected is "" (a fixed hop selects
// nothing).
//
// A GROUP hop rolls up its member copies "chain-<name>-h<i>-<member>", each of
// which the observatory probes through its own prefix of the chain. The
// counters obey the same invariants as GroupHealth — Tested == Alive+Dead and
// Alive+Dead+Untested == Total — so the panel needs no arithmetic of its own.
// State summarises them: "alive" when at least one member is alive (the hop can
// carry traffic), "dead" when at least one was tested and none is alive (a
// positive finding of a dead hop), "untested" when nothing was tested. Selected
// is the node NAME the wrapper currently picks; DelayMs/AgeSeconds are the
// SELECTED member's observation, or the freshest ALIVE member's when the
// selection has no measurement of its own — the number shown must always be a
// measurement somebody took, never an average nobody did.
type ChainHopHealth struct {
Index int `json:"index"` // 1-based position on the wire, L1..Ln
Tag string `json:"tag"` // "chain-<name>-h<i>" — the wrapper actually dialled
Kind string `json:"kind"` // "node" | "group"
Exit bool `json:"exit"` // the LAST hop: where traffic leaves to the internet
State string `json:"state"` // "alive" | "dead" | "untested"
DelayMs int `json:"delay_ms"`
AgeSeconds int64 `json:"age_seconds"` // -1 when unknown
Selected string `json:"selected"` // group hop: the node NAME it currently selects; "" otherwise
Total int `json:"total"`
Tested int `json:"tested"`
Alive int `json:"alive"`
Dead int `json:"dead"`
Untested int `json:"untested"`
}
// ChainHealth reports the reachability (used/unused) of every named chain plus
// the per-hop health of each one the running box materialised, one row per
// name, in the order given. The Used half is the chain analogue of
// GroupHealth.Used: a chain the observatory's used-set does not cover is
// reported Used=false so the panel can mark it "unused" instead of a health
// readout against a path nothing routes through.
//
// names come from the desired-state model, NOT the running box: a chain no rule
// references is never materialised (generate/chain.go resolveChain is lazy), so it
// is invisible to a box-only enumeration — yet the panel lists it from the config
// and must be able to badge it. The engine supplies the only fact a box read can
// add here, the observatory's published used-set. Pure apart from that read; nil
// names or a stopped engine (nil used-set) yield an empty/used-everything result.
// and must be able to badge it. The engine supplies what only it can: the
// observatory's published used-set, and the running box's outbound/endpoint
// pool the hop projection reads. Still no dialling anywhere; nil names or a
// stopped engine (nil used-set, empty pool) yield an empty/used-everything
// result with no hops.
func (e *Engine) ChainHealth(names []string) []ChainHealth {
out := make([]ChainHealth, 0, len(names))
used := e.observatoryUsed()
usedChains := usedChainNames(used)
pool := e.runningPool()
view := e.HealthView()
for _, name := range names {
name = strings.TrimSpace(name)
if name == "" {
continue
}
out = append(out, ChainHealth{Name: name, Used: used == nil || usedChains[name]})
out = append(out, ChainHealth{
Name: name,
Used: used == nil || usedChains[name],
Hops: chainHopHealthOf(pool, name, view),
})
}
return out
}
// runningPool snapshots the running box's outbounds AND endpoints into one
// list. The endpoints matter: a chain hop rebuilt from a wireguard/AmneziaWG
// node is an ENDPOINT copy, invisible in Outbounds() — the same trap
// groupTargets documents — and a hop projection that missed it would silently
// drop the very hop this feature exists to localise (the production chain's
// first hop IS an AWG endpoint). Empty (never nil-unsafe) on a stopped engine.
func (e *Engine) runningPool() []adapter.Outbound {
var pool []adapter.Outbound
inst := e.Instance()
if inst == nil {
return pool
}
if om := inst.Outbound(); om != nil {
pool = append(pool, om.Outbounds()...)
}
if em := inst.Endpoint(); em != nil {
for _, ep := range em.Endpoints() {
pool = append(pool, ep)
}
}
return pool
}
// chainHopHealthOf projects one chain's hop wrappers out of an outbound pool
// against one health view. Pure apart from the view reads — the same
// unit-testing contract as groupHealthOf: hand it a fake pool and a hand-built
// store and every branch is reachable without a box.
//
// A pool entry belongs to chain <name> when its tag is exactly
// "chain-<name>-h<digits>" — the hop WRAPPER the observatory dials. Member
// copies ("chain-<name>-h<i>-<member>") are not hops themselves; they are
// reached through the wrapper's own member list (adapter.OutboundGroup.All), so
// the roll-up sees exactly what the wrapper can select, in its order.
func chainHopHealthOf(pool []adapter.Outbound, name string, view HealthView) []ChainHopHealth {
prefix := "chain-" + name + "-h"
var hops []ChainHopHealth
for _, ob := range pool {
tag := ob.Tag()
rest, ok := strings.CutPrefix(tag, prefix)
if !ok {
continue
}
idx, ok := parseAllDigits(rest)
if !ok {
continue // a member copy, or another chain sharing the prefix
}
if g, isGroup := ob.(adapter.OutboundGroup); isGroup {
hops = append(hops, chainGroupHop(g, name, idx, view))
} else {
hops = append(hops, chainNodeHop(tag, idx, view))
}
}
sort.Slice(hops, func(i, j int) bool { return hops[i].Index < hops[j].Index })
if len(hops) > 0 {
// The largest index is the exit — the wrapper whose probe leaves to the
// internet. Marked after sorting so the flag cannot depend on pool order.
hops[len(hops)-1].Exit = true
}
return hops
}
// chainNodeHop is the one-measurement hop: the wrapper itself was dialled by
// the observatory, so its own board state IS the hop's health and the counters
// degenerate to whichever bucket that state fills.
func chainNodeHop(tag string, idx int, view HealthView) ChainHopHealth {
hop := ChainHopHealth{Index: idx, Tag: tag, Kind: "node", Total: 1}
hop.State, hop.DelayMs, hop.AgeSeconds = view.State(tag)
switch hop.State {
case HealthAlive:
hop.Alive = 1
case HealthDead:
hop.Dead = 1
default:
hop.Untested = 1
}
hop.Tested = hop.Alive + hop.Dead
return hop
}
// chainGroupHop rolls a group hop up over its member copies. The counters carry
// the GroupHealth invariants; the summary State answers the only question a hop
// row asks — "can this hop carry the chain": alive while anything answers,
// dead only on a positive all-tested-dead finding, untested when nothing is
// known (never dead-by-absence, the same honesty rule as everywhere else).
func chainGroupHop(g adapter.OutboundGroup, chain string, idx int, view HealthView) ChainHopHealth {
hop := ChainHopHealth{Index: idx, Tag: g.Tag(), Kind: "group", AgeSeconds: -1}
selected := g.Now()
if selected != "" {
hop.Selected = chainMemberName(selected, chain)
}
// The number a hop row shows must be a real observation: the selected
// member's when it has one, else the freshest alive member's.
freshDelay, freshAge := 0, int64(-1)
selDelay, selAge := 0, int64(-1)
for _, tag := range g.All() {
state, delayMs, age := view.State(tag)
switch state {
case HealthAlive:
hop.Alive++
if age >= 0 && (freshAge < 0 || age < freshAge) {
freshDelay, freshAge = delayMs, age
}
case HealthDead:
hop.Dead++
default:
hop.Untested++
}
hop.Total++
if tag == selected && age >= 0 {
selDelay, selAge = delayMs, age
}
}
hop.Tested = hop.Alive + hop.Dead
switch {
case hop.Alive > 0:
hop.State = HealthAlive
case hop.Tested > 0:
hop.State = HealthDead
default:
hop.State = HealthUntested
}
if selAge >= 0 {
hop.DelayMs, hop.AgeSeconds = selDelay, selAge
} else {
hop.DelayMs, hop.AgeSeconds = freshDelay, freshAge
}
return hop
}
// usedChainNames recovers the set of chain NAMES the observatory's used-set covers.
//
// The used-set is keyed by outbound TAG (the generator's schema, materialised in
+138 -1
View File
@@ -309,8 +309,145 @@ func TestChainHealthNilUsedSet(t *testing.T) {
t.Fatalf("ChainHealth = %+v, want %+v (nil used-set ⇒ all used, blanks dropped)", got, want)
}
for i, w := range want {
if got[i] != w {
if got[i].Name != w.Name || got[i].Used != w.Used {
t.Errorf("ChainHealth[%d] = %+v, want %+v", i, got[i], w)
}
// A stopped engine materialised nothing: hops must be empty, and the
// contract says empty means "nothing materialised", not "no hops".
if len(got[i].Hops) != 0 {
t.Errorf("ChainHealth[%d].Hops = %+v, want empty on a stopped engine", i, got[i].Hops)
}
}
}
// TestChainHopHealthProjection drives the per-hop readout over a 3-hop chain —
// node hop L1, group hop L2, group hop L3 (the exit) — with exactly ONE hop's
// members dead. The dead state must land on THAT hop's index and nowhere else,
// the counters must obey the GroupHealth invariants on every hop, and the exit
// flag must sit on the largest index. This is the "which hop died" question the
// whole hop surface exists to answer.
func TestChainHopHealthProjection(t *testing.T) {
hist := urltest.NewHistoryStorage()
now := time.Now()
// L1 (node hop wrapper): alive — the prefix up to hop 1 works.
hist.StoreURLTestHistory("chain-c-h1", &adapter.URLTestHistory{LastOK: now.Add(-5 * time.Second), Delay: 40})
// L2 (group hop): BOTH member copies dead — this is the hop that died.
hist.StoreURLTestHistory("chain-c-h2-x", &adapter.URLTestHistory{LastFail: now.Add(-3 * time.Second)})
hist.StoreURLTestHistory("chain-c-h2-y", &adapter.URLTestHistory{LastFail: now.Add(-2 * time.Second)})
// L3 (group hop, exit): one member alive, one never measured.
hist.StoreURLTestHistory("chain-c-h3-a", &adapter.URLTestHistory{LastOK: now.Add(-7 * time.Second), Delay: 200})
// chain-c-h3-b: nothing at all.
pool := []adapter.Outbound{
&depOutbound{failingOutbound{tag: "chain-c-h1"}, nil},
&depGroup{fakeGroup{
tag: "chain-c-h2", kind: C.TypeSelector,
all: []string{"chain-c-h2-x", "chain-c-h2-y"},
now: "chain-c-h2-x",
}, []string{"chain-c-h2-x", "chain-c-h2-y"}},
&depGroup{fakeGroup{
tag: "chain-c-h3", kind: C.TypeSelector,
all: []string{"chain-c-h3-a", "chain-c-h3-b"},
now: "chain-c-h3-a",
}, []string{"chain-c-h3-a", "chain-c-h3-b"}},
// Noise the projection must ignore: a member copy is not a hop, another
// chain's wrapper is not this chain's.
&depOutbound{failingOutbound{tag: "chain-c-h2-x"}, nil},
&depOutbound{failingOutbound{tag: "chain-other-h1"}, nil},
}
hops := chainHopHealthOf(pool, "c", newHealthView(hist, healthTTLFloor, now))
if len(hops) != 3 {
t.Fatalf("got %d hops, want 3: %+v", len(hops), hops)
}
// Ordered by index, exit on the largest.
for i, wantIdx := range []int{1, 2, 3} {
if hops[i].Index != wantIdx {
t.Fatalf("hops out of order: %+v", hops)
}
if got, want := hops[i].Exit, wantIdx == 3; got != want {
t.Errorf("hop %d Exit = %v, want %v", wantIdx, got, want)
}
}
h1, h2, h3 := hops[0], hops[1], hops[2]
// L1: one measurement, its own numbers, no selection.
if h1.Kind != "node" || h1.State != HealthAlive || h1.Total != 1 || h1.Alive != 1 ||
h1.DelayMs != 40 || h1.AgeSeconds != 5 || h1.Selected != "" {
t.Errorf("h1 = %+v, want an alive node hop with its own 40ms/5s and no selection", h1)
}
// L2: the dead hop. Both members tested, none alive => a POSITIVE dead
// finding on exactly this index. The selected member (x) is dead, and a
// dead observation carries delay 0 with the failure's age.
if h2.Kind != "group" || h2.State != HealthDead {
t.Fatalf("h2 = %+v, want the DEAD group hop — this is the answer to 'which hop died'", h2)
}
if h2.Total != 2 || h2.Tested != 2 || h2.Alive != 0 || h2.Dead != 2 || h2.Untested != 0 {
t.Errorf("h2 counters = %+v, want 2 tested / 2 dead", h2)
}
if h2.Selected != "x" {
t.Errorf("h2 Selected = %q, want the node NAME x", h2.Selected)
}
if h2.DelayMs != 0 || h2.AgeSeconds != 3 {
t.Errorf("h2 delay/age = %d/%d, want 0/3 (the selected member's failure observation)", h2.DelayMs, h2.AgeSeconds)
}
// L3: alive (one member answers), one member honestly untested — never
// folded into dead. The selected member is the alive one, so its numbers show.
if h3.Kind != "group" || h3.State != HealthAlive {
t.Fatalf("h3 = %+v, want an alive exit hop", h3)
}
if h3.Total != 2 || h3.Tested != 1 || h3.Alive != 1 || h3.Dead != 0 || h3.Untested != 1 {
t.Errorf("h3 counters = %+v, want 1 alive / 1 untested", h3)
}
if h3.Selected != "a" || h3.DelayMs != 200 || h3.AgeSeconds != 7 {
t.Errorf("h3 = %+v, want selection a with its 200ms/7s", h3)
}
// The invariants, on every hop, exactly as GroupHealth promises.
for _, h := range hops {
if h.Tested != h.Alive+h.Dead {
t.Errorf("hop %d: Tested = %d, want alive+dead = %d", h.Index, h.Tested, h.Alive+h.Dead)
}
if h.Alive+h.Dead+h.Untested != h.Total {
t.Errorf("hop %d: alive+dead+untested = %d, want total = %d",
h.Index, h.Alive+h.Dead+h.Untested, h.Total)
}
}
}
// A group hop whose SELECTION has no measurement shows the freshest ALIVE
// member's numbers instead — the shown number must always be an observation
// somebody took.
func TestChainHopSelectionWithoutMeasurementFallsBack(t *testing.T) {
hist := urltest.NewHistoryStorage()
now := time.Now()
hist.StoreURLTestHistory("chain-c-h1-a", &adapter.URLTestHistory{LastOK: now.Add(-30 * time.Second), Delay: 90})
hist.StoreURLTestHistory("chain-c-h1-b", &adapter.URLTestHistory{LastOK: now.Add(-4 * time.Second), Delay: 150})
// The selection points at c, which has no observation at all.
pool := []adapter.Outbound{
&depGroup{fakeGroup{
tag: "chain-c-h1", kind: C.TypeSelector,
all: []string{"chain-c-h1-a", "chain-c-h1-b", "chain-c-h1-c"},
now: "chain-c-h1-c",
}, nil},
}
hops := chainHopHealthOf(pool, "c", newHealthView(hist, healthTTLFloor, now))
if len(hops) != 1 {
t.Fatalf("got %d hops, want 1", len(hops))
}
h := hops[0]
if h.Selected != "c" {
t.Errorf("Selected = %q, want c (the selection is reported even unmeasured)", h.Selected)
}
if h.DelayMs != 150 || h.AgeSeconds != 4 {
t.Errorf("delay/age = %d/%d, want the freshest ALIVE member's 150/4", h.DelayMs, h.AgeSeconds)
}
if h.State != HealthAlive || h.Alive != 2 || h.Untested != 1 {
t.Errorf("hop = %+v, want alive with 2 alive / 1 untested", h)
}
}
+228 -59
View File
@@ -12,33 +12,46 @@ import (
"time"
"github.com/sagernet/sing-box/adapter"
"github.com/sagernet/sing-box/common/urltest"
C "github.com/sagernet/sing-box/constant"
)
// Exit test — "what am I actually going out through, and how fast" (F2, plan §5.E).
// Group/chain test — "ask the observatory to refresh, then report what it
// measured, plus the exit address" (F2, plan §5.E; reworked under the one-probe
// rule).
//
// # Why this is not just another node probe
// # This file no longer measures latency. On purpose.
//
// The observatory (observatory.go) answers "which nodes are alive". It cannot
// answer the question an operator actually asks after switching a group: *through
// which address am I leaving the country right now*. A group is an indirection —
// a selector or urltest over many members — so the delay of the group and the
// public address it exits from are properties of the CURRENT selection, not of
// any node the panel can point at. A CHAIN is the same question over a longer
// path: its exit wrapper "chain-<name>-hN" tunnels through every hop, so dialling
// it measures the whole L1..Ln path end to end.
// It used to: one urltest.URLTest per target, dialled right here. That was the
// defect, not a feature. The observatory (observatory.go) probes every path the
// routing rules actually use — per-chain member copies, egress-bound group
// copies, chain exits — and it is the ONLY thing allowed to dial for health.
// A second dial path from this file measured the WRONG thing: pressing "Test"
// dialled the base groups directly from the router over the default WAN, a path
// no rule routes through, and (worse) URLTest's DialContext touched the group,
// arming its own 30-minute probing ticker, which kept writing those false
// direct measurements under the members' base tags. For a node that is blocked
// on the direct WAN and alive only behind a tunnel, that reading is not stale —
// it is FALSE, and it poisoned selection and the panel alike.
//
// So one test per target (group or chain) measures two things:
// So a manual test is now a READ with a refresh request in front of it:
//
// delay_ms — reusing urltest.URLTest, the SAME primitive the observatory uses
// for a node. There is deliberately no second latency mechanism
// here: the target is just an outbound, and probing it exercises
// exactly the path traffic will take.
// exit_ip — a real HTTP request THROUGH the target to a service that echoes
// the client address back. Nothing else can produce this number: the
// router cannot know its own public address, and the proxy protocol
// does not report it.
// delay_ms / ok — the observatory's OWN measurement of the target's dial
// path. TestGroups rewinds the observatory (RefreshObservatory)
// so the numbers are fresh — taken AFTER the button press —
// then each target waits for an observation newer than the
// run's start instant and reports it verbatim. tested_unix is
// the instant that observation was taken, not "now".
// exit_ip — a real HTTP request THROUGH the target to a service that
// echoes the client address back. This is the ONE connection
// this file still opens, and it stays because no probe can
// answer it: the router cannot know its own public address,
// and the proxy protocol does not report it. It is NOT a
// second health probe — it travels the target's own routed
// path (the same outbound object the rules dial), and it is
// only ever issued for a target that (a) the observatory's
// used-set covers and (b) just resolved ALIVE. A target
// outside the rules, or one whose path is down, gets an empty
// exit_ip — never a connection.
//
// # The direct-egress trap
//
@@ -51,16 +64,15 @@ import (
// exit_ip instead of somebody else's number.
const (
// groupTestDelayTimeout bounds the latency probe of one group.
groupTestDelayTimeout = 5 * time.Second
// groupTestExitTimeout bounds the whole exit-address lookup for one group
// (connect through the tunnel + TLS + response). Short on purpose: this runs
// while a human waits on a panel button, and a slow answer is worth less than a
// prompt "could not determine".
groupTestExitTimeout = 6 * time.Second
// groupTestConcurrency bounds how many groups are tested in parallel. Small:
// every one of these opens a real tunnelled connection, and a router with a
// handful of groups is the normal case.
// groupTestConcurrency bounds how many EXIT-ADDRESS lookups run in parallel.
// Small: each one opens a real tunnelled connection, and a router with a
// handful of groups is the normal case. The board polling itself is not
// bounded by this — it is sleep-and-read, no I/O.
groupTestConcurrency = 4
// exitBodyLimit caps what is read from the exit-address service. These responses
// are a few hundred bytes; anything larger is a hijacked/captive-portal answer
@@ -68,19 +80,58 @@ const (
exitBodyLimit = 4 << 10
)
// groupTestWaitDeadline / groupTestPollEvery pace the wait for the observatory:
// each covered target polls the health board once a second until its measured
// tag carries an observation newer than the run's start, giving up after the
// deadline. 120s covers a forced pass of a large plan (batches chain
// back-to-back during a forced pass, see observatoryTickOnce) with room for the
// probe timeouts of a mostly-dead population. Package variables, not constants,
// for exactly one reason: the timeout tests must not take two minutes.
var (
groupTestWaitDeadline = 120 * time.Second
groupTestPollEvery = time.Second
)
// The fixed result texts for the targets that are never (or not yet) measured.
// They are contract, not decoration — the panel shows them verbatim.
const (
// groupTestErrNotRouted: the target is outside the observatory's used-set,
// so no rule routes through it and nothing measures it. Dialling it anyway
// would recreate the false direct measurement this rework removed.
groupTestErrNotRouted = "not routed by any enabled rule, so nothing measures it — the observatory only probes paths the rules use"
// groupTestErrNotReached: the deadline passed without a fresh observation.
groupTestErrNotReached = "the observatory has not reached this target yet — it refreshes on the global probe interval"
// groupTestErrProbingOff: the observatory is disabled (the GroupHealth
// master switch), so a refresh request has nothing to wake and waiting for
// the deadline would just delay the same answer by two minutes.
groupTestErrProbingOff = "background probing is disabled, so there is nothing to measure this target with"
// groupTestErrPathDead: the fresh observation exists and it is a FAILURE —
// the observatory probed the target's path after the button press and the
// path did not answer. An honest negative, not a missing measurement.
groupTestErrPathDead = "the observatory's probe through this path failed"
)
// GroupTestResult is one target's test outcome — a group's or a chain's (Group
// then carries the chain's model name). The JSON tags are the panel contract —
// see the shater API docs for /api/groups/test.
// see the shater API docs for /api/groups/test. The field names and types are
// FROZEN; what changed in the rework is where the numbers come from.
//
// DelayMs and OK are a READ of the observatory's measurement of the target's
// dial path — not a fresh dial performed by this file. OK is true exactly when
// the health board's state for the measured tag is alive; DelayMs is that
// observation's RTT; TestedUnix is the instant the OBSERVATION was taken (now
// minus its age), so a result honestly says how old its number is instead of
// stamping the poll time over it.
//
// Selected is the group's current pick (OutboundGroup.Now()); for a chain it is
// the node NAME the chain's last group hop currently selects, "" when the chain
// has no group hop (a fixed path selects nothing).
//
// OK reports whether the LATENCY measurement succeeded, which is the test's primary
// question. A failed exit-address lookup deliberately does NOT clear it: knowing the
// target is up and fast is useful on its own, and a probe service being unreachable
// says nothing about the tunnel. In that case OK stays true and ExitIP is empty —
// "not determined", never a guess and never somebody else's address.
// A failed exit-address lookup deliberately does NOT clear OK: knowing the
// target is up and fast is useful on its own, and a probe service being
// unreachable says nothing about the tunnel. In that case OK stays true and
// ExitIP is empty — "not determined", never a guess and never somebody else's
// address.
type GroupTestResult struct {
Group string `json:"group"`
Selected string `json:"selected"`
@@ -259,19 +310,30 @@ func parseAllDigits(s string) (int, bool) {
// and chains (empty/nil = every group and every chain in the running box),
// returning started=false when a run is already in flight.
//
// It is a SINGLETON: a second request while a run is in flight is refused rather
// than queued or run in parallel, because these runs open real tunnelled
// connections and a panel that double-fires a button must not multiply the load
// on the uplink. The observatory checks the same guard and skips its tick while
// a run is in flight (observatory.go), so a manual test never competes with
// background probing for the uplink.
// It does NOT probe. It records the run's start instant, asks the observatory
// for an out-of-turn full pass (RefreshObservatory), and then each target waits
// for the health board to carry an observation NEWER than that instant on the
// tag the observatory actually measures for it — see testOneTarget. The only
// connection a run may still open is the exit-address lookup, and only for a
// target that resolved alive on a rule-routed path.
//
// probeURL is the latency-probe URL; "" falls back to urltest's gstatic default.
// It is a SINGLETON: a second request while a run is in flight is refused
// rather than queued, because two overlapping runs would each rewind the
// observatory's cursor and neither pass would ever complete — and the panel's
// progress contract assumes one run's counters at a time anyway.
//
// Apply-swap safety: the target outbounds are snapshotted up front, so a config swap
// mid-run cannot change what is being tested. A target torn down mid-run simply fails
// its probe and is reported not-ok.
// probeURL is accepted and IGNORED. The probe URL is a global observatory
// setting now (ObservatoryConfig.ProbeURL, set at apply time); a per-run URL
// would mean this run measures something different from what the board holds,
// which is exactly the two-instruments split the rework removed. The parameter
// stays so the callers (shater/apply, shater/panel) keep compiling and the
// control-plane API shape does not churn.
//
// Apply-swap safety: the target outbounds are snapshotted up front, so a config
// swap mid-run cannot change what is being tested. A target torn down mid-run
// simply never receives a fresh observation and resolves on the deadline.
func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
_ = probeURL // ignored — see the doc comment above
if !e.groupTestRunning.CompareAndSwap(false, true) {
return false
}
@@ -285,6 +347,13 @@ func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
// GroupTestStatus for why the scope is a set of names.
e.setGroupTestScope(scopeOf(targets, missing))
// The freshness watermark: only an observation taken AFTER this instant may
// answer this run. Recorded BEFORE the refresh request so a probe that lands
// between the two can never be missed, only double-counted as fresh — the
// harmless direction.
t0 := time.Now()
e.RefreshObservatory()
go func() {
defer e.groupTestRunning.Store(false)
@@ -303,15 +372,22 @@ func (e *Engine) TestGroups(names []string, probeURL string) (started bool) {
e.setGroupTestResults(results)
e.groupTestDone.Add(int64(len(missing)))
// The used-set and the enabled bit are snapshotted once for the whole
// run: they only change on an apply, and a run that straddles an apply
// is already best-effort (see the swap-safety note above).
used := e.observatoryUsed()
obsEnabled, _, _ := e.ObservatoryStatus()
// One goroutine per target: they spend their life sleeping on the board
// poll, so there is nothing to bound — the semaphore below bounds the
// exit-address lookups, the only real connections left in a run.
sem := make(chan struct{}, groupTestConcurrency)
var wg sync.WaitGroup
for i, tgt := range targets {
wg.Add(1)
sem <- struct{}{}
go func(i int, tgt groupTestTarget) {
defer wg.Done()
defer func() { <-sem }()
res := e.testOneTarget(tgt, probeURL)
res := e.testOneTarget(tgt, used, obsEnabled, t0, sem)
e.storeGroupTestResult(i, res)
e.groupTestDone.Add(1)
}(i, tgt)
@@ -393,29 +469,122 @@ func (e *Engine) groupTargets(names []string) (targets []groupTestTarget, missin
return targets, missing
}
// testOneTarget measures one target: what it currently selects, the latency
// through it, and the public address it exits from. For a chain the dialled
// outbound is the exit wrapper, so the delay and the exit address are end-to-end
// properties of the whole L1..Ln path.
func (e *Engine) testOneTarget(t groupTestTarget, probeURL string) GroupTestResult {
// testOneTarget resolves one target WITHOUT probing it: it decides whether the
// observatory measures this target at all, and if so waits for a fresh
// observation and reports it. The three ways out, in order:
//
// 1. the used-set does not cover the target — no enabled rule routes through
// it, so nothing measures it and nothing SHOULD: resolved immediately with
// groupTestErrNotRouted, never dialled, empty exit address. A nil used-set
// means "unknown" (observatory not yet configured) and is treated as
// covered — no refusal is better than a wrong one;
// 2. the observatory is disabled — the refresh request went nowhere, so the
// wait below could only ever end on its deadline: resolved immediately
// with groupTestErrProbingOff instead of stalling the panel for two
// minutes to say the same thing;
// 3. covered and enabled — poll the health board once a second until the
// MEASURED TAG carries an observation newer than since, then report that
// observation verbatim (readFreshObservation). On the deadline:
// groupTestErrNotReached.
//
// The exit-address lookup runs ONLY on the alive path of (3) — a rule-routed
// target whose path just answered a probe — and through sem, so a run never
// opens more than groupTestConcurrency tunnelled connections at once.
func (e *Engine) testOneTarget(t groupTestTarget, used map[string]bool, obsEnabled bool, since time.Time, sem chan struct{}) GroupTestResult {
res := GroupTestResult{Group: t.name, TestedUnix: time.Now().Unix()}
if t.sel != nil {
res.Selected = t.sel()
}
ctx, cancel := context.WithTimeout(context.Background(), groupTestDelayTimeout)
delay, err := urltest.URLTest(ctx, probeURL, t.ob)
cancel()
if err != nil {
res.Error = err.Error()
if used != nil && !used[t.ob.Tag()] {
res.Error = groupTestErrNotRouted
return res
}
if !obsEnabled {
res.Error = groupTestErrProbingOff
return res
}
res.OK = true
res.DelayMs = int(delay)
// The exit address is best-effort by design: see GroupTestResult.OK.
res.ExitIP, res.ExitCountry = e.exitAddress(t.ob)
return res
deadline := time.Now().Add(groupTestWaitDeadline)
for {
if e.readFreshObservation(&res, t, since) {
if res.OK {
// The exit address is best-effort by design: see GroupTestResult.
sem <- struct{}{}
res.ExitIP, res.ExitCountry = e.exitAddress(t.ob)
<-sem
}
return res
}
if !time.Now().Before(deadline) {
res.Error = groupTestErrNotReached
return res
}
time.Sleep(groupTestPollEvery)
}
}
// measuredTagOf is the tag the observatory actually probes for this target's
// dial path — the tag whose board entry answers "how is this target doing":
//
// - a GROUP target dials whatever it currently selects, so the group's health
// IS its selection's health: OutboundGroup.Now(). For an egress-bound group
// that is a per-group copy tag, which is exactly what the plan probes;
// - a CHAIN whose exit wrapper is a PLAIN outbound is probed end-to-end under
// that wrapper tag (probeplan.go: a chain exit is its own measurement);
// - a CHAIN whose exit wrapper is a GROUP (the last hop is a group) has its
// member copies probed instead of the wrapper, so the wrapper's Now() — the
// member copy the chain currently dials through — is the measured tag. The
// wrapper here IS the last group hop, so this is chainLastGroupHop's Now()
// without a second lookup;
// - fallback: the target's own tag, for a group that has not selected yet
// (cold selector mid-swap). Its board entry is almost certainly empty, and
// the caller then honestly reports "not reached" rather than inventing one.
func measuredTagOf(t groupTestTarget) string {
if g, ok := t.ob.(adapter.OutboundGroup); ok {
if now := g.Now(); now != "" {
return now
}
}
return t.ob.Tag()
}
// readFreshObservation reads the board once: if the target's measured tag holds
// an observation taken at or after since, it is written into res (state, delay,
// the observation's own timestamp, and the current selection so Selected and
// the measurement describe the same pick) and true is returned. Otherwise res
// is left for the next poll.
//
// Precision note: the board reports ages in whole seconds, so "at or after
// since" is accurate to one second — an observation taken up to a second
// BEFORE the refresh can slip through as fresh. That is the acceptable
// direction: it is still a real measurement of the same path, at most a second
// older than requested; the strict direction (discarding genuinely fresh
// observations) would make every run one probe interval slower for nothing.
func (e *Engine) readFreshObservation(res *GroupTestResult, t groupTestTarget, since time.Time) bool {
// Re-read the selection at every poll: a forced observatory pass is exactly
// the kind of event that makes a urltest group switch members, and the
// measurement below is taken against the CURRENT pick.
if t.sel != nil {
res.Selected = t.sel()
}
tag := measuredTagOf(t)
now := time.Now()
state, delayMs, age := e.HealthView().State(tag)
if age < 0 {
return false // no observation at all (untested)
}
observedAt := now.Add(-time.Duration(age) * time.Second)
if observedAt.Before(since) {
return false // an old reading; the refresh has not reached this tag yet
}
res.OK = state == HealthAlive
res.DelayMs = delayMs
res.TestedUnix = observedAt.Unix()
if !res.OK {
res.Error = groupTestErrPathDead
}
return true
}
// exitProbe is one exit-address service: a URL and the parser for its body.
+142 -15
View File
@@ -3,6 +3,7 @@ package engine
import (
"context"
"crypto/tls"
"fmt"
"net"
"net/http"
"net/http/httptest"
@@ -143,10 +144,10 @@ func TestChainMemberName(t *testing.T) {
}
}
// dialableOutbound routes every dial to a fixed local address — a stand-in for a
// chain exit whose whole path is up. urltest.URLTest dials the probe URL's host
// through the outbound, so pointing every dial at a local HTTP server makes the
// latency probe succeed without any network.
// dialableOutbound routes every dial to a fixed local address — the sink the
// exit-address lookup lands in during tests. Since the rework nothing else in
// this file dials at all; the local server refuses to be an exit-address
// service (wrong status, no TLS), so the lookup honestly comes back empty.
type dialableOutbound struct {
failingOutbound
addr string
@@ -158,20 +159,42 @@ func (d *dialableOutbound) DialContext(ctx context.Context, network string, _ M.
return (&net.Dialer{}).DialContext(ctx, network, d.addr)
}
// TestChainExitTestMeasuresEndToEnd is the §6-S4 acceptance path for chains: the
// exit test dials the chain's EXIT TAG, returns a measured delay, carries the
// chain's model name (not the wrapper tag) as the result's Group, and reports the
// last group hop's pick as Selected. The exit address is measured through the
// same outbound (unreachable from a test => empty, never a guess).
func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
probe := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
// countingOutbound counts every DialContext/ListenPacket. It is the tripwire of
// the rework: a target the observatory does not cover must NEVER be dialled,
// and the counter is the proof.
type countingOutbound struct {
failingOutbound
dials int
}
func (c *countingOutbound) DialContext(ctx context.Context, network string, dst M.Socksaddr) (net.Conn, error) {
c.dials++
return nil, fmt.Errorf("dial refused (test)")
}
func (c *countingOutbound) ListenPacket(ctx context.Context, dst M.Socksaddr) (net.PacketConn, error) {
c.dials++
return nil, fmt.Errorf("dial refused (test)")
}
// groupTestSem is a fresh exit-lookup semaphore for direct testOneTarget calls.
func groupTestSem() chan struct{} { return make(chan struct{}, groupTestConcurrency) }
// TestChainTestReportsObservatoryMeasurement is the §6-S4 acceptance path for
// chains under the one-probe rule: the manual test does NOT dial the exit — it
// reads the observatory's board entry for the exit tag, reports its delay and
// ITS timestamp, carries the chain's model name as Group and the last group
// hop's pick as Selected. The exit-address lookup still travels the exit
// outbound (unreachable from a test => empty, never a guess).
func TestChainTestReportsObservatoryMeasurement(t *testing.T) {
sink := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusNoContent)
}))
defer probe.Close()
defer sink.Close()
exit := &dialableOutbound{
failingOutbound: failingOutbound{tag: "chain-x-h2"},
addr: probe.Listener.Addr().String(),
addr: sink.Listener.Addr().String(),
deps: []string{"chain-x-h1"},
}
pool := []adapter.Outbound{
@@ -187,14 +210,31 @@ func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
if len(targets) != 1 || targets[0].ob.Tag() != "chain-x-h2" {
t.Fatalf("targets = %+v, want the chain dialled via its exit tag", targets)
}
if got := measuredTagOf(targets[0]); got != "chain-x-h2" {
t.Fatalf("measuredTagOf = %q, want the plain exit wrapper itself", got)
}
e := New()
res := e.testOneTarget(targets[0], probe.URL)
// The observatory measured the exit 10 seconds ago, AFTER the run's start
// instant below: that observation — delay and timestamp both — is the answer.
observedAt := time.Now().Add(-10 * time.Second)
e.URLTestHistory().StoreURLTestHistory("chain-x-h2", &adapter.URLTestHistory{LastOK: observedAt, Delay: 77})
since := time.Now().Add(-time.Minute)
res := e.testOneTarget(targets[0], map[string]bool{"chain-x-h2": true}, true, since, groupTestSem())
if res.Group != "x" {
t.Errorf("Group = %q, want the chain's model name x", res.Group)
}
if !res.OK || res.Error != "" {
t.Fatalf("result = %+v, want a successful measurement through the exit tag (a local roundtrip may legitimately read 0ms)", res)
t.Fatalf("result = %+v, want the board's alive observation reported", res)
}
if res.DelayMs != 77 {
t.Errorf("DelayMs = %d, want the observatory's 77 — this file measures nothing itself", res.DelayMs)
}
// tested_unix is the OBSERVATION's instant (whole-second precision), never
// the poll's.
if got, want := res.TestedUnix, observedAt.Unix(); got < want-1 || got > want+1 {
t.Errorf("TestedUnix = %d, want the observation's instant ~%d", got, want)
}
if res.Selected != "relay" {
t.Errorf("Selected = %q, want the last group hop's pick relay", res.Selected)
@@ -204,6 +244,93 @@ func TestChainExitTestMeasuresEndToEnd(t *testing.T) {
}
}
// TestTestOneTargetUnroutedNeverDialled is the rework's core promise: a target
// outside the observatory's used-set is resolved immediately with the
// documented explanation and ZERO dials — no latency probe, no exit-address
// lookup, nothing. Dialling it would manufacture exactly the false direct
// measurement the rework removed.
func TestTestOneTargetUnroutedNeverDialled(t *testing.T) {
e := New()
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "idle"}}
tgt := groupTestTarget{name: "idle", ob: ob}
res := e.testOneTarget(tgt, map[string]bool{"something-else": true}, true, time.Now(), groupTestSem())
if ob.dials != 0 {
t.Fatalf("an unrouted target was dialled %d time(s); it must never be", ob.dials)
}
if res.OK {
t.Fatal("an unrouted target reported ok=true")
}
if res.Error != groupTestErrNotRouted {
t.Fatalf("Error = %q, want the not-routed explanation", res.Error)
}
if res.ExitIP != "" || res.ExitCountry != "" {
t.Fatalf("an unrouted target must carry no exit address: %+v", res)
}
}
// A disabled observatory resolves covered targets immediately too: the refresh
// request went nowhere, so waiting out the deadline would only delay the same
// honest answer — and still, nothing is dialled.
func TestTestOneTargetProbingOffNeverDialled(t *testing.T) {
e := New()
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
tgt := groupTestTarget{name: "auto", ob: ob}
res := e.testOneTarget(tgt, nil, false, time.Now(), groupTestSem())
if ob.dials != 0 {
t.Fatalf("target dialled %d time(s) with probing off; want 0", ob.dials)
}
if res.OK || res.Error != groupTestErrProbingOff {
t.Fatalf("result = %+v, want ok=false with the probing-off explanation", res)
}
}
// The deadline path: a covered target whose measured tag never receives a fresh
// observation resolves with the not-reached explanation — and, again, without a
// single dial of its own.
func TestTestOneTargetDeadlineWithoutObservation(t *testing.T) {
// Shrink the wait so the test answers in milliseconds; restore afterwards.
oldDeadline, oldPoll := groupTestWaitDeadline, groupTestPollEvery
groupTestWaitDeadline, groupTestPollEvery = 30*time.Millisecond, 5*time.Millisecond
defer func() { groupTestWaitDeadline, groupTestPollEvery = oldDeadline, oldPoll }()
e := New()
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
tgt := groupTestTarget{name: "auto", ob: ob}
res := e.testOneTarget(tgt, map[string]bool{"auto": true}, true, time.Now(), groupTestSem())
if ob.dials != 0 {
t.Fatalf("target dialled %d time(s) while waiting on the board; want 0", ob.dials)
}
if res.OK || res.Error != groupTestErrNotReached {
t.Fatalf("result = %+v, want ok=false with the not-reached explanation", res)
}
}
// A fresh DEAD observation is an answer, not a timeout: ok=false with the
// path-dead explanation, delay 0, and no exit-address connection for a path
// that just failed its probe.
func TestTestOneTargetReportsFreshDeath(t *testing.T) {
e := New()
ob := &countingOutbound{failingOutbound: failingOutbound{tag: "auto"}}
tgt := groupTestTarget{name: "auto", ob: ob}
since := time.Now().Add(-time.Minute)
e.URLTestHistory().MarkFailed("auto")
res := e.testOneTarget(tgt, map[string]bool{"auto": true}, true, since, groupTestSem())
if ob.dials != 0 {
t.Fatalf("a dead target was dialled %d time(s); the exit lookup is for alive paths only", ob.dials)
}
if res.OK || res.Error != groupTestErrPathDead {
t.Fatalf("result = %+v, want ok=false with the path-dead explanation", res)
}
if res.DelayMs != 0 || res.ExitIP != "" {
t.Fatalf("a dead path must carry no delay and no exit address: %+v", res)
}
}
// TestGroupTestSingleton is the "no parallel runs" invariant: while a run holds the
// guard, a second TestGroups must be refused rather than starting a concurrent run
// (each of these opens real tunnelled connections).
+74 -10
View File
@@ -100,6 +100,14 @@ type observatoryState struct {
cycles uint64
// busy guards against a slow tick overlapping the next one.
busy bool
// force is the ONE-SHOT "check now" flag raised by RefreshObservatory: while
// it is up the freshness gate is bypassed, so a manual refresh really
// RE-MEASURES everything on the plan instead of skipping whatever happens to
// be fresh, and each finished batch immediately nudges the next one so the
// pass runs back-to-back rather than on the 10s tick. It is cleared the
// moment the forced pass schedules its last batch (cursor reaches the end of
// the plan) — one full walk, then back to the polite steady state.
force bool
}
// ConfigureObservatory installs (or re-tunes, or stops) the observatory. It is
@@ -120,6 +128,7 @@ func (e *Engine) ConfigureObservatory(cfg ObservatoryConfig) {
if !cfg.Enabled {
e.obs.plan, e.obs.used, e.obs.cursor = nil, nil, 0
e.obs.force = false
if e.obs.stop != nil {
close(e.obs.stop)
e.obs.stop, e.obs.done, e.obs.nudge = nil, nil, nil
@@ -163,6 +172,7 @@ func (e *Engine) StopObservatory() {
e.obs.stop, e.obs.done, e.obs.nudge = nil, nil, nil
e.obs.enabled = false
e.obs.plan, e.obs.used, e.obs.cursor = nil, nil, 0
e.obs.force = false
if stop != nil {
close(stop)
}
@@ -172,6 +182,38 @@ func (e *Engine) StopObservatory() {
}
}
// RefreshObservatory asks for an OUT-OF-TURN full pass of the probe plan: the
// cursor is rewound to the top, the one-shot force flag is raised so the pass
// bypasses the freshness gate (a "check now" that skipped everything fresh
// would measure nothing and answer with yesterday's numbers), and the loop is
// nudged so the first batch starts immediately instead of on the next 10s tick.
//
// This is the ONLY way the panel's "Test" button touches the network now: the
// manual group test (grouptest.go) no longer dials anything itself — it calls
// this, then watches the health board for observations newer than its start
// instant. One prober, one set of dial paths, one truth.
//
// No-op when the observatory is disabled or has no plan: there is nothing to
// walk, and raising force with no loop to consume it would only arm a stale
// flag for a future enable. The caller (TestGroups) reports that situation
// through its own results; this method stays silent about it on purpose —
// it is a request, not a query.
func (e *Engine) RefreshObservatory() {
e.obsMu.Lock()
defer e.obsMu.Unlock()
if !e.obs.enabled || len(e.obs.plan) == 0 {
return
}
e.obs.cursor = 0
e.obs.force = true
if e.obs.nudge != nil {
select {
case e.obs.nudge <- struct{}{}:
default:
}
}
}
// ObservatoryStatus reports whether the observatory is on, how many measurements
// the current plan holds, and how many full passes have completed. Cheap; used by
// tests and available for diagnostics.
@@ -240,17 +282,18 @@ func (e *Engine) observatoryLoop(stop <-chan struct{}, done chan<- struct{}, nud
// observatoryTickOnce runs one batch. It is deliberately conservative about when
// it does nothing:
//
// - a manual exit test (grouptest.go) is in flight => SKIP. That run is the
// human's explicit request; the observatory defers rather than competing for
// the uplink. It checks the guard but never CLAIMS it, so a pressed Test
// button can never be answered "already running" because of background work;
// - the previous tick has not finished => SKIP, so slow probes never stack;
// - the engine is stopped => jobs no longer resolve and are skipped (probeJob).
//
// There is deliberately NO deferral to a manual group test any more. The guard
// that used to sit here ("skip the tick while e.groupTestRunning is up") made
// sense when the manual run dialled the uplink itself and the two would have
// competed for it; that run no longer probes ANYTHING — it asks for a forced
// pass via RefreshObservatory and then reads the board. Keeping the guard would
// therefore be worse than pointless: a manual run now DEPENDS on the
// observatory ticking, and skipping ticks for its whole duration would deadlock
// the very refresh it is waiting on (up to its full 120s deadline).
func (e *Engine) observatoryTickOnce() {
if e.groupTestRunning.Load() {
return
}
e.obsMu.Lock()
if e.obs.busy || !e.obs.enabled || len(e.obs.plan) == 0 {
e.obsMu.Unlock()
@@ -258,7 +301,8 @@ func (e *Engine) observatoryTickOnce() {
}
if e.obs.cursor >= len(e.obs.plan) {
// Cycle complete: count it and start the next pass from the top. This is
// the ONE place the cursor is deliberately rewound.
// the ONE place the cursor is rewound on the observatory's own schedule
// (RefreshObservatory rewinds it too, but that is the caller's request).
e.obs.cycles++
e.obs.cursor = 0
}
@@ -270,12 +314,32 @@ func (e *Engine) observatoryTickOnce() {
batch := append([]ProbeJob(nil), e.obs.plan[start:end]...)
e.obs.cursor = end
interval := e.obs.interval
// A forced pass (RefreshObservatory) bypasses the freshness gate for its one
// walk of the plan. The flag is captured for THIS batch and cleared the
// moment the walk's last batch is scheduled: the remaining jobs of this very
// batch still run unfiltered off the captured copy, and the next pass is an
// ordinary polite one again.
force := e.obs.force
if force && end >= len(e.obs.plan) {
e.obs.force = false
}
e.obs.busy = true
e.obsMu.Unlock()
defer func() {
e.obsMu.Lock()
e.obs.busy = false
// Mid-forced-pass: chain straight into the next batch instead of waiting
// out the 10s tick. A human pressed "check now"; walking a subscription-
// sized plan at one batch per tick would stretch the answer past the
// manual run's deadline for no reason — the batch size still bounds how
// much is in flight at once, which is the limit that actually matters.
if e.obs.force && e.obs.nudge != nil {
select {
case e.obs.nudge <- struct{}{}:
default:
}
}
e.obsMu.Unlock()
}()
@@ -287,7 +351,7 @@ func (e *Engine) observatoryTickOnce() {
sem := make(chan struct{}, observatoryConcurrency)
var wg sync.WaitGroup
for _, j := range batch {
if !observatoryShouldProbe(hist, j, interval, now) {
if !force && !observatoryShouldProbe(hist, j, interval, now) {
continue
}
wg.Add(1)
+69 -8
View File
@@ -100,28 +100,89 @@ func TestObservatoryTickStoppedEngine(t *testing.T) {
}
}
// The exit-test gate: while a manual group/chain test is in flight the tick is a
// no-op — the observatory checks the guard but never claims it, so a pressed
// Test button is never refused because of background work.
func TestObservatoryDefersToExitTest(t *testing.T) {
// The INVERSE of the old exit-test gate, pinned so it cannot come back: the
// tick must RUN while a manual group test is in flight. The manual run no
// longer probes anything — it asks for a forced pass and then waits on the
// board — so a tick that deferred to it would deadlock the very refresh the
// run is polling for, until its 120s deadline reported "not reached" about
// paths nobody attempted.
func TestObservatoryTicksDuringManualRun(t *testing.T) {
e := New()
defer e.StopObservatory()
e.ConfigureObservatory(ObservatoryConfig{Enabled: true, Options: obsFixture(), ProbeInterval: time.Minute})
e.groupTestRunning.Store(true)
e.observatoryTickOnce()
if _, cursor, _ := obsState(e); cursor != 0 {
t.Fatal("tick ran while an exit test was in flight")
if _, cursor, _ := obsState(e); cursor == 0 {
t.Fatal("tick deferred to an in-flight manual run; the run depends on the tick now")
}
e.groupTestRunning.Store(false)
// And the guard is free: a manual run can start immediately.
// The observatory never claims the manual-run guard either way round: a
// manual run can still start immediately.
if !e.TestGroups(nil, "") {
t.Fatal("a manual exit test was refused; the observatory must never hold groupTestRunning")
t.Fatal("a manual test was refused; the observatory must never hold groupTestRunning")
}
waitGroupTestIdle(t, e)
}
// RefreshObservatory is the manual run's whole probing story: it rewinds the
// cursor, raises the one-shot force flag (bypassing the freshness gate for one
// full walk), and the flag clears itself when the forced pass schedules its
// last batch. Disabled or plan-less observatories ignore it entirely.
func TestRefreshObservatoryForcesOnePass(t *testing.T) {
e := New()
defer e.StopObservatory()
// Disabled: a refresh request is a documented no-op.
e.RefreshObservatory()
e.obsMu.Lock()
force := e.obs.force
e.obsMu.Unlock()
if force {
t.Fatal("RefreshObservatory raised force on a disabled observatory")
}
e.ConfigureObservatory(ObservatoryConfig{Enabled: true, Options: obsFixture(), ProbeInterval: time.Minute})
// Fresh observations on every planned tag: the polite gate would skip them
// all, which is exactly what a forced pass must NOT do.
hist := e.URLTestHistory()
now := time.Now()
hist.StoreURLTestHistory("n1", &adapter.URLTestHistory{LastOK: now, Delay: 5})
hist.StoreURLTestHistory("n2", &adapter.URLTestHistory{LastOK: now, Delay: 5})
// Pretend a walk was mid-plan; the refresh must rewind it.
e.obsMu.Lock()
e.obs.cursor = 1
e.obsMu.Unlock()
e.RefreshObservatory()
e.obsMu.Lock()
cursor, force := e.obs.cursor, e.obs.force
e.obsMu.Unlock()
if cursor != 0 || !force {
t.Fatalf("after refresh: cursor=%d force=%v, want 0/true", cursor, force)
}
// One tick covers the whole 2-job plan (batch is 24), so the forced pass
// completes and the flag clears itself — one full walk, not a permanent mode.
e.observatoryTickOnce()
e.obsMu.Lock()
cursor, force = e.obs.cursor, e.obs.force
e.obsMu.Unlock()
if force {
t.Fatal("force flag survived the forced pass; it must be one-shot")
}
if cursor != 2 {
t.Fatalf("cursor after forced tick = %d, want 2 (the pass really walked the plan)", cursor)
}
// The stopped engine resolves no outbounds, so the probes were skipped —
// but the fresh history proves the GATE was bypassed only if the jobs were
// attempted; attempt-tracking lives in probeJob, which needs a box. What is
// pinned here is the flag lifecycle and the rewind, the two halves
// TestGroups depends on.
}
// The freshness gate: a job whose every covered tag has an observation younger
// than the global probe interval is skipped — that set is exactly what an ACTIVE
// group is measuring itself, and the observatory must not duplicate or suppress
+39 -20
View File
@@ -27,8 +27,12 @@ import (
// - a CHAIN entry tag "chain-<n>-hN" (the exit wrapper, generate/chain.go) is
// probed ITSELF: dialling the copy pulls the whole L1→…→Ln path through its
// Detour links, so one probe is the chain's end-to-end health. Intermediate
// NODE hops are not probed separately — their death is visible in the exit
// probe and no selection depends on them;
// NODE hop wrappers ("chain-<n>-h<i>" that are plain outbounds) are probed
// TOO — not because selection needs them (it does not), but because an
// operator does: dialling "chain-<n>-h1" measures exactly the prefix of the
// path up to and including hop 1, so when the exit probe dies the per-hop
// probes say WHICH hop died instead of only that the chain did (the
// ChainHealth hops surface, grouphealth.go);
// - a GROUP hop inside a chain ("chain-<n>-h<i>", a wrapper selector reached by
// walking the exit's Detour links) additionally has its member copies
// "chain-<n>-h<i>-<member>" probed: they are the wrapper's selection
@@ -291,8 +295,20 @@ func (w *planWalk) visitMember(group, member string) {
// walkDetour follows a probed tag's Detour links toward the router. A group met
// on the way is a chain's GROUP HOP wrapper: its member copies are selection
// candidates and are probed, each through its own prefix of the path. A plain
// outbound met on the way is an intermediate node hop — marked used, never probed
// on its own — and the walk continues through it.
// outbound met on the way falls in two halves:
//
// - a chain NODE HOP wrapper ("chain-<n>-h<i>", parseChainExitTag) is probed
// as a measurement of its own. Dialling it pulls exactly the prefix of the
// chain up to and including hop i through the Detour links, so its verdict
// LOCALISES a failure: the exit probe says the chain died, the hop probes
// say where. It used to be skipped here on the argument that its death
// shows up in the exit probe anyway — true for end-to-end health, useless
// for an operator staring at a dead 4-hop chain. The measurement is stored
// under the wrapper tag alone: a chain prefix is nobody else's dial path,
// so no alias may borrow its verdict;
// - anything else on the detour path — an "egress-…" interface outbound, an
// ordinary node's egress detour — keeps today's behaviour: marked used,
// never probed on its own, and the walk continues through it.
func (w *planWalk) walkDetour(tag string) {
if tag == "" || w.walked[tag] {
return
@@ -303,25 +319,28 @@ func (w *planWalk) walkDetour(tag string) {
return
}
w.used[tag] = true
if ent.group {
next := ""
for _, m := range ent.members {
ment, ok := w.byTag[m]
if !ok {
continue
}
w.used[m] = true
if !ment.group && probeablePlanTag(ment.typ, m) {
w.addTarget(dialKeyTag(m), m)
}
if next == "" {
next = ment.detour // every member detours into the same previous hop
}
if !ent.group {
if _, isHopWrapper := parseChainExitTag(tag); isHopWrapper && probeablePlanTag(ent.typ, tag) {
w.addTarget(dialKeyTag(tag), tag)
}
w.walkDetour(next)
w.walkDetour(ent.detour)
return
}
w.walkDetour(ent.detour)
next := ""
for _, m := range ent.members {
ment, ok := w.byTag[m]
if !ok {
continue
}
w.used[m] = true
if !ment.group && probeablePlanTag(ment.typ, m) {
w.addTarget(dialKeyTag(m), m)
}
if next == "" {
next = ment.detour // every member detours into the same previous hop
}
}
w.walkDetour(next)
}
// dialKeyTag / dialKeyCopy build the dedup key for one measurement target.
+58
View File
@@ -218,6 +218,64 @@ func TestObservatoryPlanChain(t *testing.T) {
}
}
// A chain with a plain NODE hop wrapper on the way to the exit: the wrapper IS
// a measurement of its own now (walkDetour) — dialling "chain-<c>-h1" measures
// exactly the path prefix up to hop 1, which is what localises a dead hop. And
// the other half of the honesty contract: a base node that appears ONLY as a
// chain group-hop member is never measured under its BASE tag — the member
// copy's observation stays on the copy, because the copy's dial path (through
// the chain prefix) is not the base outbound's dial path, and no job may write
// a direct-looking verdict onto a tag it did not dial.
func TestObservatoryPlanChainHopWrappersProbed(t *testing.T) {
opts := option.Options{
Outbounds: []option.Outbound{
// L1: a plain node hop wrapper.
fixNode("chain-c2-h1", ""),
// L2: a group hop over one member copy of base node N.
fixNode("chain-c2-h2-N", "chain-c2-h1"),
fixGroup(C.TypeSelector, "chain-c2-h2", "chain-c2-h2-N"),
// L3: the exit, a plain node copy.
fixNode("chain-c2-h3", "chain-c2-h2"),
// The base node behind the L2 member copy: emitted, referenced by
// NOTHING but the copy's name.
fixNode("N", ""),
},
Route: fixRoute("", "chain-c2-h3"),
}
byDial, used := planOf(t, opts)
// The node hop wrapper is a job of its own, stored only under itself.
j := assertPlanHas(t, byDial, "chain-c2-h1")
if len(j.Store) != 1 || j.Store[0] != "chain-c2-h1" {
t.Errorf("node hop wrapper store = %v, want only itself (a chain prefix is nobody else's path)", j.Store)
}
// The exit and the group-hop member copy are jobs, as before.
assertPlanHas(t, byDial, "chain-c2-h3")
assertPlanHas(t, byDial, "chain-c2-h2-N")
// NO job may record anything under the base tag N: not as a dial, not as a
// store alias. A direct measurement can never land on it from this config.
if _, ok := byDial["N"]; ok {
t.Error("plan dials the base node N, which no rule reaches")
}
for dial, job := range byDial {
for _, tag := range job.Store {
if tag == "N" {
t.Errorf("job %q stores under the base tag N; a chain member copy must never alias its base", dial)
}
}
}
if used["N"] {
t.Error("used-set claims the base node N; only the chain copies are on the path")
}
// The wrappers themselves are used (they are the path).
for _, u := range []string{"chain-c2-h1", "chain-c2-h2", "chain-c2-h3", "chain-c2-h2-N"} {
if !used[u] {
t.Errorf("used-set is missing %q", u)
}
}
}
// DNS server detours and route.Final are roots too — a group referenced only as a
// resolver's detour is still a used, probed path.
func TestObservatoryPlanDNSDetourRoot(t *testing.T) {
+70
View File
@@ -0,0 +1,70 @@
package engine
import (
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
)
// Standing down the self-check of unused urltest groups (health board §5.C).
//
// # Why the engine does this, and not the generator
//
// A urltest group probes its own members: a warm-up sweep at PostStart and a
// ticker for as long as traffic touches it (protocol/group/urltest.go). For a
// group the routing rules actually use, that probing travels the same path the
// traffic does and is welcome. For a group NO rule reaches, it dials the
// members' base outbounds directly from the router — a path nothing uses — and
// stores the results under the base tags, where every health consumer reads
// them as "the node's health". A node that only works behind a tunnel then
// reads dead on the shared board because an idle group measured it over the
// blocked direct WAN. The observatory's plan (probeplan.go) is the single
// statement of what gets probed and along which paths; this file makes the
// config that reaches box.New SAY so, by flipping SelfCheck off on every
// urltest group outside the plan's used-set.
//
// The generator cannot do this: whether a group is used is a property of the
// EMITTED rules and DNS detours as a whole, and BuildObservatoryPlan is the one
// place that reachability is computed. Recomputing it in the generator would be
// a second copy of the same walk, and second copies drift.
// standDownUnusedSelfCheck flips SelfCheck to false on every urltest group in
// opts that the observatory's used-set does not cover, and returns how many it
// stood down (for the apply log and the tests). Selector groups are left alone
// — they have no self-check to stand down — and a group any rule, route.Final
// or DNS detour reaches keeps its default (nil == on).
//
// It MUTATES the option structs the caller handed in: opts.Outbounds carries
// the generator's option values as pointers behind an `any` (the generator
// always emits *option.URLTestOutboundOptions), so writing through them changes
// the caller's config. That is deliberate, not an accident to guard against:
// the plan is the single source of what gets probed, and the config that
// reaches box.New must already say so — a copy-on-write here would build a box
// whose groups probe paths the hash and the plan claim nobody probes. It runs
// BEFORE hashOptions in applyLocked for the same reason: the hash must describe
// what is really built, so a rule change that flips a group used<->unused is a
// real config change and triggers a swap.
//
// A urltest outbound whose Options is not the expected pointer type (a value,
// or something foreign) is skipped rather than guessed at: it did not come from
// our generator, and silently rebuilding somebody else's option struct is worse
// than leaving one group's self-check up.
func standDownUnusedSelfCheck(opts option.Options) int {
_, used := BuildObservatoryPlan(opts, "")
stood := 0
for i := range opts.Outbounds {
ob := &opts.Outbounds[i]
if ob.Type != C.TypeURLTest || used[ob.Tag] {
continue
}
utOpts, ok := ob.Options.(*option.URLTestOutboundOptions)
if !ok {
continue
}
// One fresh pointer per group, so no two option structs alias a shared
// bool that a later caller could flip for both at once.
off := false
utOpts.SelfCheck = &off
stood++
}
return stood
}
+204
View File
@@ -0,0 +1,204 @@
package engine
import (
"testing"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
)
// standDownUnusedSelfCheck is the enforcement half of the one-probe rule: a
// urltest group no rule reaches gets SelfCheck=false written into its option
// struct (in place — the struct the box will be built from), a rule-reachable
// one keeps its nil default, and selectors are never touched at all (they have
// no self-check to stand down). The count it returns is what applyLocked logs.
func TestStandDownUnusedSelfCheck(t *testing.T) {
opts := option.Options{
Outbounds: []option.Outbound{
fixNode("n1", ""),
fixNode("idle1", ""),
fixNode("pin1", ""),
// Reached by a rule: must keep probing itself.
fixGroup(C.TypeURLTest, "auto", "n1"),
// Reached by nothing: its self-check dials a path nobody uses and
// must be stood down.
fixGroup(C.TypeURLTest, "idle", "idle1"),
// A selector reached by nothing: left alone — no self-check exists.
fixGroup(C.TypeSelector, "pins", "pin1"),
},
Route: fixRoute("auto"),
}
stood := standDownUnusedSelfCheck(opts)
if stood != 1 {
t.Fatalf("stood down %d group(s), want exactly 1 (idle)", stood)
}
find := func(tag string) option.Outbound {
t.Helper()
for _, ob := range opts.Outbounds {
if ob.Tag == tag {
return ob
}
}
t.Fatalf("outbound %q missing from opts", tag)
return option.Outbound{}
}
// The unused urltest group: SelfCheck is now an explicit false in the very
// struct the caller handed in (the mutation is the point — the config that
// reaches box.New must say what the plan says).
idle := find("idle").Options.(*option.URLTestOutboundOptions)
if idle.SelfCheck == nil || *idle.SelfCheck {
t.Fatalf("idle group SelfCheck = %v, want an explicit false", idle.SelfCheck)
}
// The used urltest group keeps the nil default (self-check on): its probing
// travels the path traffic actually takes and is welcome.
auto := find("auto").Options.(*option.URLTestOutboundOptions)
if auto.SelfCheck != nil {
t.Fatalf("used group SelfCheck = %v, want nil (untouched default)", *auto.SelfCheck)
}
// The selector's option struct has no SelfCheck field at all; what is
// pinned here is that stand-down neither counted it nor mangled its type.
if _, ok := find("pins").Options.(*option.SelectorOutboundOptions); !ok {
t.Fatal("selector options were rebuilt; stand-down must leave selectors alone")
}
// Idempotence: a second pass finds the same one group (already-false is
// still counted — the function reports plan membership, not novelty) and
// changes nothing further. This is what keeps the apply hash stable across
// the every-minute reconcile.
if again := standDownUnusedSelfCheck(opts); again != 1 {
t.Fatalf("second pass stood down %d, want 1 (deterministic on identical opts)", again)
}
if idle.SelfCheck == nil || *idle.SelfCheck {
t.Fatal("second pass flipped the idle group's SelfCheck back")
}
}
// TestStandDownNeverTouchesChainHopWrappers pins the property the production
// config lives or dies by, DIRECTLY rather than transitively through the plan
// tests: a CHAIN HOP WRAPPER group ("chain-<name>-h<i>", type urltest) must
// NEVER be stood down.
//
// The failure mode this guards against is subtle and silent. The owner's live
// rule routes through egress:ewan -> node:awgout -> group:sub0 -> group:sub1
// -> group:sub2 — a chain whose group hops are rebuilt as urltest wrappers
// over per-chain member copies. Those wrappers are the ONLY probing that
// travels the real path: their own schedule is the on-path measurement, and
// the selection inside the chain (which member each hop dials) depends on the
// verdicts it writes. Reachability of a wrapper is established by walkDetour
// (probeplan.go) walking the exit's Detour links — a DIFFERENT code path from
// the root/member expansion the other tests exercise. If a future change to
// that walk dropped the wrappers from the used-set, standDownUnusedSelfCheck
// would obediently flip their SelfCheck off, the chain would stop measuring
// itself, and nothing else would ever measure it: we would have replaced the
// false reading this rework removed with NO reading at all. Both existing
// tests could stay green while that happened — the plan test asserts the
// used-set, not the stand-down; the stand-down test above never mentions
// chains. This one closes that gap by asserting the stand-down's OUTPUT on
// the production topology.
//
// The fixture mirrors what the generator actually emits for that config, not
// a toy: an exit wrapper chain-p-h4 (urltest over member copies) whose copies
// detour into chain-p-h3 (urltest), whose copies detour into chain-p-h2
// (urltest), whose copies detour into chain-p-h1 (the plain AmneziaWG hop
// copy), which detours into egress-ewan — the detour lives on the member
// COPIES, not on the wrappers, exactly as generate/chain.go builds it. The
// base groups sub0/sub1/sub2 are ALSO emitted, over the base node tags,
// reached by nothing — which is precisely what buildGroups does today and
// precisely the false direct prober the stand-down exists to silence.
func TestStandDownNeverTouchesChainHopWrappers(t *testing.T) {
opts := option.Options{
Outbounds: []option.Outbound{
// The egress the whole chain leaves through.
fixUtility(C.TypeDirect, "egress-ewan"),
// L1: the plain AWG hop copy (in production an endpoint; a plain
// outbound here exercises the same walk — walkDetour treats both as
// non-group entries).
fixNode("chain-p-h1", "egress-ewan"),
// L2..L4: urltest group hop wrappers over per-chain member copies;
// each COPY detours into the previous hop.
fixNode("chain-p-h2-a", "chain-p-h1"),
fixNode("chain-p-h2-b", "chain-p-h1"),
fixGroup(C.TypeURLTest, "chain-p-h2", "chain-p-h2-a", "chain-p-h2-b"),
fixNode("chain-p-h3-a", "chain-p-h2"),
fixNode("chain-p-h3-b", "chain-p-h2"),
fixGroup(C.TypeURLTest, "chain-p-h3", "chain-p-h3-a", "chain-p-h3-b"),
fixNode("chain-p-h4-a", "chain-p-h3"),
fixNode("chain-p-h4-b", "chain-p-h3"),
fixGroup(C.TypeURLTest, "chain-p-h4", "chain-p-h4-a", "chain-p-h4-b"),
// The base nodes and the base groups over them: emitted alongside
// the chain, reached by no rule — the generator keeps emitting them
// today, and they are exactly the direct-WAN probers that poisoned
// the board.
fixNode("a", ""),
fixNode("b", ""),
fixGroup(C.TypeURLTest, "sub0", "a", "b"),
fixGroup(C.TypeURLTest, "sub1", "a", "b"),
fixGroup(C.TypeURLTest, "sub2", "a", "b"),
},
// One rule, targeting the chain's exit wrapper — the production shape.
Route: fixRoute("", "chain-p-h4"),
}
stood := standDownUnusedSelfCheck(opts)
if stood != 3 {
t.Fatalf("stood down %d group(s), want exactly the 3 base groups sub0/sub1/sub2", stood)
}
utOptions := func(tag string) *option.URLTestOutboundOptions {
t.Helper()
for _, ob := range opts.Outbounds {
if ob.Tag == tag {
o, ok := ob.Options.(*option.URLTestOutboundOptions)
if !ok {
t.Fatalf("outbound %q is not a urltest group", tag)
}
return o
}
}
t.Fatalf("outbound %q missing from opts", tag)
return nil
}
// Every hop wrapper keeps its self-check: nil, the untouched default. An
// explicit false on ANY of these is the chain going blind.
for _, wrapper := range []string{"chain-p-h2", "chain-p-h3", "chain-p-h4"} {
if sc := utOptions(wrapper).SelfCheck; sc != nil {
t.Errorf("hop wrapper %q SelfCheck = %v, want nil — its own schedule IS the on-path measurement", wrapper, *sc)
}
}
// Every base group is stood down: nothing routes through them, and their
// probing would dial the base nodes over the direct WAN.
for _, base := range []string{"sub0", "sub1", "sub2"} {
if sc := utOptions(base).SelfCheck; sc == nil || *sc {
t.Errorf("base group %q SelfCheck = %v, want an explicit false", base, sc)
}
}
// And the reason the stand-down of sub0/sub1/sub2 is SAFE in this config:
// with the base groups silenced, the base node tags a/b have no measurement
// path at all — no job dials them and no job stores under them (the chain's
// member copies record only under themselves). Their board entries simply
// stop being written, honestly untested, instead of carrying a false
// direct-WAN verdict.
jobs, used := BuildObservatoryPlan(opts, "")
for _, baseTag := range []string{"a", "b"} {
if used[baseTag] {
t.Errorf("used-set claims base node %q; only the chain copies are on the path", baseTag)
}
for _, j := range jobs {
if j.Dial == baseTag {
t.Errorf("plan dials base node %q, which no rule reaches", baseTag)
}
for _, s := range j.Store {
if s == baseTag {
t.Errorf("job %q stores under base node %q; a chain copy must never alias its base", j.Dial, baseTag)
}
}
}
}
}
+344
View File
@@ -0,0 +1,344 @@
package engine
// Bounded, deterministic teardown of a SUPERSEDED sing-box instance.
//
// # The defect this file exists for
//
// sing-box's Box has no live reload, so every applied config change builds a
// fresh Box and retires the old one (see engine.go). Retiring it used to be one
// line — `_ = old.Close()` — and that line had three separate problems, all of
// which had to be true at once for the observed failure:
//
// 1. NO CANCELLATION. Upstream's own runner gives every Box its OWN cancellable
// context and calls cancel() BEFORE Close (cmd/sing-box/cmd_run.go:136 and
// :188-190). This engine handed EVERY box.New the one shared, never-cancelled
// context built in New (engine.go), and box.New does not derive a cancellable
// child of what it is given (box.go: `ctx := options.Context` and nothing
// else). So every goroutine inside a retired box that would have stopped on
// context cancellation simply never stopped, and only what each adapter's
// Close() explicitly tears down actually went away.
//
// 2. NO TIME BUDGET. Box.Close walks its subsystems SEQUENTIALLY and waits for
// each one forever; the only thing watching the clock is a taskmonitor that
// prints "close endpoint/wireguard[...] take too much time to finish!" after
// C.StopTimeout and then keeps waiting anyway (common/taskmonitor/monitor.go).
// A wireguard endpoint that will not come down therefore stalls the rest of
// the close list, and every subsystem AFTER it in the walk is never reached.
//
// 3. NO EVIDENCE. The error was discarded (`_ =`) and the pointer to the old box
// was overwritten in the same breath, so after the swap nothing in the process
// could tell — or even ask — whether the previous generation had actually
// died. On the router this accumulated: up to four generations logging side by
// side inside one shaterd process, each with its own outbound connections and,
// worse, its own WireGuard devices. Two devices sharing a private key evict
// each other at the peer, which is exactly the fault generate/wgdedup.go
// removes WITHIN a config — reintroduced here BETWEEN generations.
//
// # The contract now
//
// - Every box gets its own cancellable context, cancelled before Close.
// - Close runs against a HARD budget (closeBudget). When it expires the apply
// moves on: the new box is already started and serving, and blocking the
// control plane on a shutdown that is not going to finish would only add an
// unresponsive panel to the problem.
// - An overrun is LOUD, not swallowed: an ERROR line naming the generation, and
// a critical entry in the operator-facing warning set for as long as the
// abandoned generation is still running (see PendingCloses).
// - Generations do not stack: before an apply builds another instance it gives
// any abandoned teardown one more budget to finish, and if one is still alive
// the swap is forced onto the close-old-then-start-new path so the process
// never holds two LIVE boxes on top of an abandoned one.
import (
"errors"
"io"
"sync"
"time"
box "github.com/sagernet/sing-box"
)
// ErrCloseTimeout reports that a superseded instance did not finish shutting
// down inside the close budget and has been ABANDONED (it may still be running).
//
// It is deliberately NOT propagated out of Apply. The swap it accompanies
// succeeded — the new box is built, started and carrying traffic — and returning
// a failure there would make the caller abort the netplane stage and, with no
// engine fault to point at, leave a stale ruleset loaded over a healthy engine.
// The fact is surfaced through PendingCloses() instead, which reaches
// `shaterd status` and the panel and clears itself the moment the shutdown
// finally completes. Close()/Teardown DO return it: there the failure to stop is
// the whole point of the operation.
var ErrCloseTimeout = errors.New("engine instance did not shut down within the close budget")
// defaultCloseBudget is the HARD wall-clock budget for retiring one superseded
// box.
//
// Five seconds, for three reasons that all point at the same number:
//
// - it is exactly sing-box's own C.StopTimeout — the threshold at which the
// engine itself declares a single lifecycle stop to be taking too long. A
// close that blows past the budget upstream considers excessive is by
// definition not a slow close, it is a stuck one;
// - it stays under C.FatalStopTimeout (10s), which is when upstream's CLI gives
// up and calls the process unclosable. We want to have already reacted by
// then;
// - it is paid AFTER the replacement box is started and serving, and only on an
// apply that really changed something (the hash gate makes a no-op reconcile
// free), so the worst case is five seconds added to one apply — not to the
// every-minute cron reconcile, and never to a status poll.
const defaultCloseBudget = 5 * time.Second
// closeBudget/closeBoxFn are package-wide seams. Production never touches them;
// tests in this package and in shater/apply install a short budget and a
// deliberately hung closer to exercise the abandoned-generation path without
// standing up a box that really refuses to die.
var (
seamMu sync.RWMutex
closeBudget = defaultCloseBudget
closeBoxFn = func(c io.Closer) error { return c.Close() }
)
// SetCloseBudget overrides the close budget and returns a function restoring the
// previous value. TEST SEAM — production uses defaultCloseBudget.
func SetCloseBudget(d time.Duration) (restore func()) {
seamMu.Lock()
prev := closeBudget
closeBudget = d
seamMu.Unlock()
return func() {
seamMu.Lock()
closeBudget = prev
seamMu.Unlock()
}
}
// SetBoxCloser overrides HOW a superseded instance is closed and returns a
// function restoring the previous closer. TEST SEAM — production calls
// Box.Close. A closer that never returns is how the abandoned-generation path is
// tested.
func SetBoxCloser(fn func(io.Closer) error) (restore func()) {
seamMu.Lock()
prev := closeBoxFn
closeBoxFn = fn
seamMu.Unlock()
return func() {
seamMu.Lock()
closeBoxFn = prev
seamMu.Unlock()
}
}
func currentCloseBudget() time.Duration {
seamMu.RLock()
defer seamMu.RUnlock()
return closeBudget
}
func currentBoxCloser() func(io.Closer) error {
seamMu.RLock()
defer seamMu.RUnlock()
return closeBoxFn
}
// pendingClose tracks ONE retirement in flight. It lives in Engine.pending from
// the moment the retirement starts until Close returns, and `abandoned` records
// whether the budget expired while it was still running — i.e. whether this
// generation is a leak the operator must be told about, or merely a close that
// is in progress under a lock the caller is holding anyway.
type pendingClose struct {
gen uint64
hash string // config hash the retired box was built from
since time.Time
done chan struct{} // closed when the underlying Close finally returns
abandoned bool // guarded by Engine.pendingMu
finished bool // guarded by Engine.pendingMu
}
// StuckClose is the read-side view of a superseded generation that overran the
// close budget and is STILL running inside this process.
type StuckClose struct {
// Generation is the 1-based sequence number of the box that will not die.
// It matches nothing in the sing-box log by itself, but it lets two status
// reads tell "the same old generation" from "another one just leaked".
Generation uint64
// Hash is the config hash the leaked box was built from — the same value
// `shaterd status` reported as `hash` while that config was the running one.
// It is what connects "an engine is stuck" to WHICH configuration is stuck.
Hash string
// Since is when its shutdown was started.
Since time.Time
// Elapsed is how long it has been shutting down, as of the read.
Elapsed time.Duration
}
// retireLocked tears down a superseded box under the close budget. The caller
// holds e.mu and must clear e.instance itself.
//
// Order matters and mirrors upstream: cancel the box's context FIRST so every
// goroutine keyed on it unwinds, and only then call Close, which is what actually
// releases the listeners, the WireGuard devices and the cache-file lock.
//
// Returns nil on a clean shutdown, ErrCloseTimeout when the budget expired (the
// close keeps running in the background and is reaped when it finishes), or the
// close's own error.
func (e *Engine) retireLocked(b *box.Box, cancel func(), gen uint64, hash string) error {
if cancel != nil {
cancel()
}
if b == nil {
return nil
}
pc := &pendingClose{gen: gen, hash: hash, since: time.Now(), done: make(chan struct{})}
e.pendingMu.Lock()
e.pending = append(e.pending, pc)
e.pendingMu.Unlock()
closeFn := currentBoxCloser()
errc := make(chan error, 1)
go func() {
err := closeFn(b)
e.pendingMu.Lock()
pc.finished = true
wasAbandoned := pc.abandoned
e.removePendingLocked(pc)
e.pendingMu.Unlock()
if wasAbandoned && e.log != nil {
// The counterpart of the ERROR below: an operator who saw the leak
// reported must be able to see it resolve without restarting anything.
e.log.Warn("engine generation ", gen, " finally finished shutting down after ",
time.Since(pc.since).Round(time.Millisecond), "; it is no longer running")
}
close(pc.done)
errc <- err
}()
budget := currentCloseBudget()
timer := time.NewTimer(budget)
defer timer.Stop()
select {
case err := <-errc:
return err
case <-timer.C:
}
// The budget expired. Claim the generation as abandoned — unless the close
// happened to land in the same instant, in which case there is nothing to
// report and we take its real result.
e.pendingMu.Lock()
abandoned := !pc.finished
if abandoned {
pc.abandoned = true
}
e.pendingMu.Unlock()
if !abandoned {
return <-errc
}
if e.log != nil {
e.log.Error("engine generation ", gen, " (config ", shortHash(hash), ") did NOT shut down within ", budget,
" and has been ABANDONED: it may still hold its outbound connections and its ",
"WireGuard devices (two devices with the same private key evict each other at ",
"the peer), and it keeps writing to this log. The new configuration is running; ",
"restart shaterd if this generation never clears.")
}
return ErrCloseTimeout
}
// shortHash renders a config hash the way the log and the panel both need it:
// enough to identify the configuration, short enough to read in a syslog line.
func shortHash(h string) string {
if len(h) < 12 {
return "unknown"
}
return h[:12]
}
// removePendingLocked drops pc from the pending list. Caller holds pendingMu.
func (e *Engine) removePendingLocked(pc *pendingClose) {
for i, p := range e.pending {
if p == pc {
e.pending = append(e.pending[:i], e.pending[i+1:]...)
return
}
}
}
// pendingSnapshot copies the in-flight retirements. Leaf lock only.
func (e *Engine) pendingSnapshot() []*pendingClose {
e.pendingMu.Lock()
defer e.pendingMu.Unlock()
return append([]*pendingClose(nil), e.pending...)
}
// PendingCloses lists the superseded generations that overran their close budget
// and are still running. Empty is the healthy answer.
//
// It takes ONLY the pending leaf lock — never e.mu — so the panel and
// `shaterd status` can report a leaking teardown even while the apply that
// produced it is still holding the engine mutex. That is deliberate: the one
// moment this information matters most is while an apply is stalled behind a
// shutdown that will not finish.
func (e *Engine) PendingCloses() []StuckClose {
now := time.Now()
e.pendingMu.Lock()
defer e.pendingMu.Unlock()
out := make([]StuckClose, 0, len(e.pending))
for _, p := range e.pending {
if !p.abandoned {
continue
}
out = append(out, StuckClose{
Generation: p.gen,
Hash: p.hash,
Since: p.since,
Elapsed: now.Sub(p.since),
})
}
return out
}
// Generations reports how many sing-box instances this process is carrying: the
// running one (0 or 1) plus every superseded generation whose shutdown has not
// finished. ONE is the healthy answer for a started engine; anything above it is
// the leak this file exists to make impossible to hide.
func (e *Engine) Generations() int {
e.mu.Lock()
live := 0
if e.instance != nil {
live = 1
}
e.mu.Unlock()
e.pendingMu.Lock()
defer e.pendingMu.Unlock()
return live + len(e.pending)
}
// awaitAbandonedLocked gives every already-abandoned generation ONE more close
// budget to finish, and returns how many are still running afterwards. The caller
// holds e.mu and is about to build another instance.
//
// This is what keeps generations from stacking. Without it, a box that will not
// come down means the NEXT apply quietly adds a third instance to the process,
// and the one after that a fourth — which is precisely the accumulation observed
// on the router. The wait is bounded by one budget for the whole set (not per
// entry), so a permanently stuck generation costs a single extra budget on an
// apply that actually changes the config, and nothing at all on the every-minute
// no-op reconcile, which never gets past the hash gate.
func (e *Engine) awaitAbandonedLocked() int {
pending := e.pendingSnapshot()
if len(pending) == 0 {
return 0
}
deadline := time.NewTimer(currentCloseBudget())
defer deadline.Stop()
for _, p := range pending {
select {
case <-p.done:
case <-deadline.C:
return len(e.PendingCloses())
}
}
return len(e.PendingCloses())
}
+247
View File
@@ -0,0 +1,247 @@
// lx: pins the invariant that a chain copy of an AmneziaWG node carries the
// SAME obfuscation parameters as the base node, field for field.
//
// Context: on the router a node placed behind an egress hop
// (egress:ewan -> node:awgout, tag chain-<name>-h1) handshook forever and never
// passed a transport packet, while the same node standalone was healthy. One
// candidate explanation was that rebuildNode loses AWG params on the copy — it
// does not (both the base and the copy go through the same
// builder.wireguardEndpoint / amneziaOptions, the copy differing only in Tag and
// DialerOptions.Detour), and this test holds that line. The real cause was the
// bind swap the detour triggers: a detour makes Endpoint.Start pick ClientBind
// instead of conn.StdNetBind, and ClientBind unconditionally overwrote bytes 1-3
// of every datagram, shredding the h4 magic header of transport packets. See
// transport/wireguard/client_bind.go (hasReserved) and its regression tests.
//
// Honest note: this test passes both before and after that fix — it is a pin on
// a path that was never broken, not the reproducer for the bug.
//
// Unlike generate_test.go this file is NOT Linux-gated: it stops at
// GenerateWithWarnings and never calls engine.Apply / box.New, so it needs no
// routing_mark validation and runs on every platform.
package generate
import (
"encoding/base64"
"fmt"
"net/url"
"reflect"
"testing"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// awgKey returns a valid 32-byte base64 WireGuard key seeded by fill. Local to
// this file so it does not depend on the Linux-only suite's helpers.
func awgKey(fill byte) string {
b := make([]byte, 32)
for i := range b {
b[i] = fill + byte(i)
}
return base64.StdEncoding.EncodeToString(b)
}
// TestChainCopyPreservesAmneziaWGOptions builds a chain whose terminal hop is an
// AmneziaWG node and asserts the chain copy's AmneziaWGOptions equals the base
// endpoint's, comparing every field of the struct (so a field added later
// without being mapped in amneziaOptions is caught here too).
func TestChainCopyPreservesAmneziaWGOptions(t *testing.T) {
priv := awgKey(1)
pub := awgKey(9)
// Every knob the option struct carries that the share-link can express:
// jc/jmin/jmax, s1-s3, ranged h1-h4 (AWG 2.0), i1-i5. s4 is deliberately left
// unset (0) — that is the real-world shape in which the transport magic lands
// in bytes 0-3 and the ClientBind bug bit.
uri := fmt.Sprintf(
"awg://%s@203.0.113.10:51820?publickey=%s&address=10.13.13.2/32&allowedips=0.0.0.0/0"+
"&jc=4&jmin=40&jmax=70&s1=86&s2=57&s3=13"+
"&h1=1618116899-1618116949&h2=1795397486-1795397536&h3=3333333333&h4=1618116899-1618116949"+
"&i1=%s&i2=%s&i3=%s&i4=%s&i5=%s#awgout",
url.QueryEscape(priv), url.QueryEscape(pub),
url.QueryEscape("<b 0xf0>"), url.QueryEscape("<c>"), url.QueryEscape("<t>"),
url.QueryEscape("<r 10>"), url.QueryEscape("<b 0xab>"),
)
node := model.Node{Name: "awgout", Enabled: true, URI: uri}
// Two SEPARATE configs, not one. wgdedup refuses to materialise a WireGuard
// node twice from one private key (two devices sharing a key evict each other's
// session), so a single config can hold either the base endpoint or the chain
// copy — never both. That is also why on the router this node exists only as
// "chain-<name>-h1". The invariant under test is therefore cross-config: the
// same node, referenced directly vs referenced through a chain, must yield the
// same AmneziaWG parameters.
direct := &model.Model{
Globals: model.DefaultGlobals(),
Nodes: []model.Node{node},
Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 100, Target: "node:awgout"},
},
}
chained := &model.Model{
Globals: model.DefaultGlobals(),
Nodes: []model.Node{node},
Egresses: []model.Egress{
{Name: "ewan", Type: "interface", Interface: "wan"},
},
Chains: []model.Chain{
{Name: "viaewan", Hops: []string{"egress:ewan", "node:awgout"}},
},
Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 100, Target: "chain:viaewan"},
},
}
directOpts, warns, err := GenerateWithWarnings(direct)
if err != nil {
t.Fatalf("GenerateWithWarnings(direct): %v (warnings: %v)", err, warns)
}
chainedOpts, chainWarns, err := GenerateWithWarnings(chained)
if err != nil {
t.Fatalf("GenerateWithWarnings(chained): %v (warnings: %v)", err, chainWarns)
}
base, ok := awgOptionsForTag(directOpts, "awgout")
if !ok {
t.Fatalf("base endpoint %q not found; tags = %v (warnings: %v)",
"awgout", endpointTags(directOpts), warns)
}
if !base.IsSet() {
t.Fatalf("base endpoint has no AmneziaWG params: %+v", base)
}
// Absolute expectation, not just parity. A "copy == base" assertion alone is
// satisfied when BOTH lose a field (they share amneziaOptions), so pin the
// literal values the share-link carries. Adding a field to
// option.AmneziaWGOptions without mapping it in amneziaOptions fails the
// exhaustiveness check below.
want := option.AmneziaWGOptions{
Jc: 4, Jmin: 40, Jmax: 70,
S1: 86, S2: 57, S3: 13, S4: 0,
H1: "1618116899-1618116949",
H2: "1795397486-1795397536",
H3: "3333333333",
H4: "1618116899-1618116949",
I1: "<b 0xf0>", I2: "<c>", I3: "<t>", I4: "<r 10>", I5: "<b 0xab>",
}
if base != want {
t.Fatalf("base AmneziaWG params lost/garbled in parsing:\n got %+v\nwant %+v", base, want)
}
// Exhaustiveness: every field of the struct must be exercised above, so a
// newly added knob cannot slip through unmapped and unnoticed.
assertAWGFieldsCovered(t, want)
// The chain hop copy: buildHopWrapper tags it chain-<chain>-h<idx>. Find it as
// "the wireguard endpoint in the chained config" rather than hardcoding the
// index, so a change in hop numbering does not silently no-op this test.
var (
copyTag string
copyOptions option.AmneziaWGOptions
found bool
)
for _, endpoint := range chainedOpts.Endpoints {
wg, isWG := endpoint.Options.(*option.WireGuardEndpointOptions)
if !isWG {
continue
}
copyTag, copyOptions, found = endpoint.Tag, wg.AmneziaWGOptions, true
break
}
if !found {
t.Fatalf("no chain copy endpoint emitted; endpoint tags = %v (warnings: %v)",
endpointTags(chainedOpts), chainWarns)
}
if copyTag == "awgout" {
t.Fatalf("expected a per-hop chain copy tag, got the base tag %q", copyTag)
}
// Field-by-field, via reflection: any field of AmneziaWGOptions the copy fails
// to carry is named explicitly rather than hidden behind one struct diff.
baseValue := reflect.ValueOf(base)
copyValue := reflect.ValueOf(copyOptions)
for i := 0; i < baseValue.NumField(); i++ {
field := baseValue.Type().Field(i)
want := baseValue.Field(i).Interface()
got := copyValue.Field(i).Interface()
if !reflect.DeepEqual(want, got) {
t.Errorf("chain copy %q lost AmneziaWG field %s: got %#v, want %#v",
copyTag, field.Name, got, want)
}
}
if !t.Failed() && base != copyOptions {
t.Fatalf("chain copy %q AmneziaWGOptions differ from base: %+v vs %+v",
copyTag, copyOptions, base)
}
// The copy must additionally differ from the base in exactly the way the chain
// intends: it detours through the egress hop, the base does not.
baseDialer, okBase := dialerForTag(directOpts, "awgout")
copyDialer, okCopy := dialerForTag(chainedOpts, copyTag)
if !okBase || !okCopy {
t.Fatalf("dialer options missing: base=%v copy=%v", okBase, okCopy)
}
if baseDialer.Detour != "" {
t.Errorf("base endpoint must dial directly, got Detour=%q", baseDialer.Detour)
}
if copyDialer.Detour == "" {
t.Errorf("chain copy %q must carry the egress hop as Detour, got empty", copyTag)
}
}
// assertAWGFieldsCovered fails if any field of option.AmneziaWGOptions is left
// at its zero value in the expectation, other than the ones deliberately unset:
//
// S4 — kept 0 on purpose; that is the shape in which the transport magic
// lands in bytes 0-3, i.e. the configuration that broke on the router.
// Id/Ip/Ib — WireSock masquerade sugar, mutually exclusive with an explicit I1
// (device_awg.go rejects the combination), and I1 is set here.
//
// The point is that adding a knob to the struct without extending this test
// turns into a failure here rather than silent non-coverage.
func assertAWGFieldsCovered(t *testing.T, want option.AmneziaWGOptions) {
t.Helper()
deliberatelyUnset := map[string]bool{"S4": true, "Id": true, "Ip": true, "Ib": true}
value := reflect.ValueOf(want)
for i := 0; i < value.NumField(); i++ {
name := value.Type().Field(i).Name
if deliberatelyUnset[name] {
continue
}
if value.Field(i).IsZero() {
t.Errorf("AmneziaWGOptions.%s is not exercised by this test "+
"(add it to the share-link and to `want`, or to deliberatelyUnset)", name)
}
}
}
// awgOptionsForTag returns the AmneziaWGOptions of the wireguard endpoint tagged
// tag.
func awgOptionsForTag(opts option.Options, tag string) (option.AmneziaWGOptions, bool) {
for _, endpoint := range opts.Endpoints {
if endpoint.Tag != tag {
continue
}
wg, ok := endpoint.Options.(*option.WireGuardEndpointOptions)
if !ok {
return option.AmneziaWGOptions{}, false
}
return wg.AmneziaWGOptions, true
}
return option.AmneziaWGOptions{}, false
}
// dialerForTag returns the DialerOptions of the wireguard endpoint tagged tag.
func dialerForTag(opts option.Options, tag string) (option.DialerOptions, bool) {
for _, endpoint := range opts.Endpoints {
if endpoint.Tag != tag {
continue
}
wg, ok := endpoint.Options.(*option.WireGuardEndpointOptions)
if !ok {
return option.DialerOptions{}, false
}
return wg.DialerOptions, true
}
return option.DialerOptions{}, false
}
+37 -21
View File
@@ -610,44 +610,60 @@ func splitDomainMarker(e string) (marker, value string, ok bool) {
// reader can tell a deliberate difference from an oversight (R9.3). Verified
// against the code on both sides, not against upstream docs — this is a fork.
//
// marker domain lists (this file) routing rules (ruleMatchers, route.go)
// There are now only TWO contexts, not three. A routing rule no longer classifies
// domains at all: `dst_domain` was removed in schema v2 and a rule names its
// destination through a `config ruleset`, so every domain entry in the system —
// filter list, device list, inline rule-set — arrives at classifyDomainEntries
// below. The one caller that adds something on top is inlineRulesetRule
// (ruleset.go), which peels `regexp:` off first.
//
// marker DNS filter / device lists inline rule-sets (inlineRulesetRule)
// ---------- ---------------------------- --------------------------------------
// full: yes — exact domain yes — exact domain
// suffix: yes — domain + subdomains yes — domain + subdomains
// .example SAME AS suffix: (dot stripped) SAME AS suffix: (dot stripped)
// keyword: yes — substring yes — substring
// (bare) SUFFIX in filter/device/ EXACT domain
// inline-ruleset lists
// (bare) SUFFIX (domain + subdomains) SUFFIX (domain + subdomains)
// regexp: NO — warns yes — DomainRegex (pattern validated)
// geosite: NO — warns NO — warns, rule matcher omitted
// geosite: NO — warns NO — warns, entry dropped
//
// Two corrections this table used to get wrong, both of the "claimed a behaviour
// that does not exist" kind:
// A THIRD context exists and has NO row here on purpose: the body of a
// `source=url` plain-text list. That is a hosts/one-domain-per-line FILE parsed by
// parseDomainList (ruleset.go), not a list of typed entries, and it understands no
// marker at all — every line is a domain plus its subdomains. An entry written in
// the vocabulary above is dropped there (a domain cannot contain ":") and reported
// per list by warnListEntryVocabulary. The long block above parseDomainList states
// why unifying the two was rejected; do not read this table as covering it.
//
// - `.example.com` was documented as "subdomains only" on both sides. It is not:
// The bare-entry row is now the SAME on both sides, and that uniformity is the
// point of schema v2 — a routing rule used to read a bare entry as an EXACT host
// while every list read it as a suffix, a difference nothing in the UI showed.
// model.migrate1to2 is what preserves the old meaning of existing configs: it
// rewrites a rule's bare `dst_domain` entry as `full:` when it moves it into the
// generated rule-set.
//
// Corrections this table used to get wrong, all of the "claimed a behaviour that
// does not exist" kind:
//
// - `.example.com` was documented as "subdomains only". It is not:
// classifyDomainEntries strips the dot, so it is a synonym of `suffix:` and
// matches the apex too. See the leading-dot branch above for why the synonym
// is kept rather than the distinction implemented.
// - `geosite:` was documented as a working routing matcher. It is not: the
// route-rule geosite field was REMOVED from this engine, so route.go warns and
// omits the matcher (a rule left with no other matcher is skipped entirely).
// Use a `config ruleset` with source=geosite.
// route-rule geosite field was REMOVED from this engine. Use a `config
// ruleset` with source=geosite and category chips.
//
// The two absences in the domain-list column are DELIBERATE, not gaps:
// The two absences in the FILTER column are DELIBERATE, not gaps:
//
// - regexp: an invalid regular expression is only detected when the rule is
// built, where it aborts box.New and takes the whole config down — the exact
// fail-open violation this audit spent its time removing. Supporting it here
// would require compiling and validating every pattern at generate time.
// Warning is the honest answer until that is done.
// - geosite: filter lists and rulesets already express geosite properly, via
// - regexp: an invalid regular expression aborts box.New and takes the whole
// config down, so it may only be accepted where every pattern is compiled and
// validated at generate time. Inline rule-sets do exactly that
// (peelDomainRegexes, ruleset.go), which is why the right-hand column says
// yes; the DNS-filter path has no such validation and warns instead.
// - geosite: filter lists and rule-sets already express geosite properly, via
// `source=geosite` plus category chips, which fetches the official compiled
// .srs. A `geosite:` entry inside an inline list would be a second, weaker
// path to the same feature.
//
// The BARE-entry difference is also deliberate and long-standing: a blocklist
// entry is meant to cover subdomains, while a routing rule's bare entry is an
// exact host. Both are documented at their call sites.
// toASCIIDomain punycodes a unicode domain entry so it can match the punycoded
// names that actually arrive in DNS queries. Lenient by design: an entry idna
+9 -1
View File
@@ -185,7 +185,15 @@ func TestDNSFilterRemoteBlocklistHTTPClient(t *testing.T) {
{Name: "cf", Type: "doh", Address: "https://1.1.1.1/dns-query", Detour: "direct"},
},
Blocklists: []model.Blocklist{
{Name: "remote-ads", Enabled: true, Source: "url", URL: srv.URL, Response: "nxdomain", UpdateInterval: "24h"},
// The ".srs" suffix is LOAD-BEARING, not decoration: ruleSetURLIsEngineNative
// (ruleset.go) decides remote-vs-compiled-local by URL EXTENSION alone, and
// this test is about the REMOTE path — the engine fetching the compiled set
// itself through the direct outbound. httptest.NewServer's bare
// "http://127.0.0.1:<port>" has no extension, so it fell into the TEXT-list
// path instead: the list was downloaded by generate's own listFetcher, parsed
// as a hosts file and compiled into a LOCAL rule-set, which every assertion
// below then contradicted. Do not trim the suffix.
{Name: "remote-ads", Enabled: true, Source: "url", URL: srv.URL + "/blocklist.srs", Response: "nxdomain", UpdateInterval: "24h"},
},
}
+2 -1
View File
@@ -289,8 +289,9 @@ func TestFailoverWarnsOnceAcrossChainCopy(t *testing.T) {
m := twoNodeGroupModel("failover")
m.Nodes = append(m.Nodes, model.Node{Name: "hop", Enabled: true, URI: ss("203.0.113.9")})
m.Chains = []model.Chain{{Name: "ch", Hops: []string{"node:hop", "group:g"}}}
m.Rulesets = []model.Ruleset{inlineDomainSet("ex", "example.com")}
m.Rules = []model.Rule{
{Name: "r", Enabled: true, DstDomain: []string{"example.com"}, Target: "chain:ch"},
{Name: "r", Enabled: true, DstRuleset: []string{"ex"}, Target: "chain:ch"},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
+27 -9
View File
@@ -310,6 +310,19 @@ func GenerateWithWarningsAt(m *model.Model, now time.Time) (option.Options, []st
Route: route,
DNS: dns,
}
// LAST, on the finished config: fold away duplicate WireGuard devices.
//
// A WG/AWG node may be materialised several times over — the always-emitted
// base endpoint (outbound.go), a per-chain hop copy and a chain group hop's
// per-member copy (chain.go) — and unlike a TCP proxy copy, each of those is a
// real device holding the SAME private key. A WireGuard peer keeps one session
// per key, so the copies evict each other and none of them passes traffic:
// every chain containing a WG node was dead for exactly this reason. Running
// on the assembled opts (rather than at each producer) is what makes the
// guarantee hold for all three paths and any future fourth. See wgdedup.go.
b.dedupWireGuardEndpoints(&opts)
return opts, b.warnings, nil
}
@@ -506,16 +519,21 @@ func validPortRange(s string) bool {
// the classifier actually drops.
//
// `suffix:` belongs here. The list used to omit it on the theory that suffix:/
// regexp: are route-rule-only and are peeled off by ruleMatchers before the shared
// classifier sees them. That is true of route.go — which handles a bare `suffix:`
// in its own switch and warns there, so this list cannot double-report — but NOT
// of the classifier itself: classifyDomainEntries has a `case "suffix"`, and every
// other caller (devices.go, dnsfilter.go) reaches it directly. A lone `suffix:`
// arriving that way was dropped by add()'s empty-value guard and reported by
// nobody, which is the one outcome this pair of functions exists to prevent.
// regexp: were route-rule-only and were peeled off before the shared classifier
// saw them. classifyDomainEntries has a `case "suffix"`, and every caller
// (devices.go, dnsfilter.go, inlineRulesetRule) reaches it directly, so a lone
// `suffix:` was dropped by add()'s empty-value guard and reported by nobody —
// the one outcome this pair of functions exists to prevent. (The route-rule half
// of that old reasoning is gone entirely: `dst_domain` was removed in schema v2,
// so route.go classifies no domains at all and there is no second warner to
// double-report with.)
//
// `regexp:` is correctly absent: the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix and is reported by unrecognisedDomainPrefix.
// `regexp:` is correctly absent, for two different reasons depending on the
// caller. In a filter/device list the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix reported by unrecognisedDomainPrefix. In an
// inline rule-set peelDomainRegexes (ruleset.go) strips every `regexp:` entry
// BEFORE this predicate runs and reports a valueless one itself, so it can never
// reach here either way.
var domainMarkers = []string{"keyword:", "full:", "suffix:", "."}
// isDomainMarkerOnly reports whether an entry is a bare classification marker
+31 -5
View File
@@ -227,6 +227,12 @@ func TestAllReachableProtocols(t *testing.T) {
{Name: "vmess-ws", Enabled: true, URI: vmessLink},
{Name: "trojan1", Enabled: true, URI: "trojan://password@example.com:443?sni=example.com#trojan1"},
{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"},
// QUIC protocols travel the same road: share-link -> parse.Proxy ->
// outbound -> box.New. Before parse grew these two branches the links
// were dropped at the parser and the outbounds below never existed.
{Name: "hy2-1", Enabled: true, URI: "hysteria2://secret@example.com:8443/?sni=example.com&alpn=h3#hy2-1"},
{Name: "hy2-alias", Enabled: true, URI: "hy2://secret@example.org#hy2-alias"},
{Name: "tuic1", Enabled: true, URI: "tuic://22222222-2222-2222-2222-222222222222:secret@example.com:443/?sni=example.com&alpn=h3&congestion_control=cubic&udp_relay_mode=native#tuic1"},
},
}
opts, warns, changed := applyAndClose(t, m)
@@ -236,16 +242,34 @@ func TestAllReachableProtocols(t *testing.T) {
if len(warns) != 0 {
t.Fatalf("unexpected warnings: %v", warns)
}
for _, tag := range []string{"vless-ws", "vless-grpc", "vmess-ws", "trojan1", "ss1"} {
for _, tag := range []string{"vless-ws", "vless-grpc", "vmess-ws", "trojan1", "ss1", "hy2-1", "hy2-alias", "tuic1"} {
if findOutbound(opts, tag) == nil {
t.Fatalf("outbound %q not emitted", tag)
}
}
// The hy2 alias link carries no port and no sni: parse must have supplied the
// scheme default (443) and the server name, or the node would dial nowhere.
if ob := findOutbound(opts, "hy2-alias"); ob != nil {
o, ok := ob.Options.(*option.Hysteria2OutboundOptions)
if !ok {
t.Fatalf("hy2-alias options type = %T", ob.Options)
}
if o.ServerPort != 443 {
t.Fatalf("hy2-alias port = %d, want the scheme default 443", o.ServerPort)
}
if o.TLS == nil || o.TLS.ServerName != "example.org" {
t.Fatalf("hy2-alias tls = %+v", o.TLS)
}
if o.TLS.UTLS != nil {
t.Fatalf("uTLS must never reach a QUIC outbound: %+v", o.TLS.UTLS)
}
}
}
// --- White-box: hysteria2 / tuic / shadowtls option mapping validates. -------
// These protocols are not yet produced by parse.ParseShareLink, so we drive the
// mapping directly with synthetic parse.Proxy values and validate via box.New.
// shadowtls is not produced by parse.ParseShareLink (hysteria2 and tuic now
// are, see TestAllReachableProtocols), so the mapping is driven directly with
// synthetic parse.Proxy values and validated via box.New.
func TestQUICAndShadowTLSMappingValidates(t *testing.T) {
b := newBuilder(&model.Model{Globals: model.DefaultGlobals()})
@@ -323,8 +347,9 @@ func TestEgressDPISpoofValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12366, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "spf", Type: "direct", DPI: "spoof"}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "spoof-rule", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:spf"},
{Name: "spoof-rule", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:spf"},
},
}
opts, warns, changed := applyAndClose(t, m)
@@ -346,8 +371,9 @@ func TestByedpiEgressValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12367, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: 1080}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:bd"},
{Name: "desync", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:bd"},
},
}
opts, warns, changed := applyAndClose(t, m)
@@ -162,14 +162,23 @@ func TestObservatoryPlanFromGeneratedConfig(t *testing.T) {
t.Errorf("chain exit %q store = %v, want only itself (a chain prefix is nobody else's path)", chainExit, store)
}
}
// The chain's intermediate node hop n1 is used (the exit detours through it) but
// NOT probed separately — its death is visible in the exit probe, and no
// selection depends on it. n1 IS probed via the plain "auto" group, so the
// assertion is "no SEPARATE chain-hop-h1 job", not "n1 never dialled".
if jobDials(jobs, "chain-hop-h1") {
t.Errorf("plan dials intermediate chain hop chain-hop-h1 — only the exit is probed end-to-end")
// The chain's intermediate NODE hop wrapper IS probed now, as a measurement
// of its own: dialling chain-hop-h1 measures exactly the prefix of the path
// up to and including hop 1, which is what lets an operator see WHICH hop
// died instead of only that the chain did (engine/probeplan.go walkDetour,
// surfaced through ChainHealth.Hops). Its store is only itself — a chain
// prefix is nobody else's dial path, so the base node n1 must never inherit
// a verdict measured through the chain.
chainHop := "chain-hop-h1"
if !jobDials(jobs, chainHop) {
t.Errorf("plan does not probe intermediate chain hop %q; jobs=%v", chainHop, dialsOfJobs(jobs))
} else {
store := jobStore(jobs, chainHop)
if len(store) != 1 || store[0] != chainHop {
t.Errorf("chain hop %q store = %v, want only itself (no base alias for a chain prefix)", chainHop, store)
}
}
if !used[chainExit] || !used["chain-hop-h1"] {
if !used[chainExit] || !used[chainHop] {
t.Errorf("used-set is missing chain hop tags; used=%v", used)
}
// The used groups themselves are in the used-set but never dialled (a group is
+3 -2
View File
@@ -21,9 +21,10 @@ func TestProfileAppliesCleanly(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12370, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{inlineDomainSet("ads", "ads.example")},
Rules: []model.Rule{
{Name: "lan-proxy", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "node:ss1"},
{Name: "adblock", Enabled: true, Order: 20, DstDomain: []string{"ads.example"}, Target: "block"},
{Name: "adblock", Enabled: true, Order: 20, DstRuleset: []string{"ads"}, Target: "block"},
},
Profiles: []model.Profile{
{Name: "home", Enabled: true, Priority: 1, DisableRules: []string{"adblock"}},
@@ -35,7 +36,7 @@ func TestProfileAppliesCleanly(t *testing.T) {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
// Profile-disabled 'adblock' rule must be absent.
if opts.Route == nil || hasDomainRule(opts.Route, "ads.example") {
if opts.Route == nil || hasRulesetRule(opts.Route, "ads") {
t.Fatalf("profile-disabled rule 'adblock' should not be emitted")
}
// The lan-proxy rule still routes to ss1.
+41 -23
View File
@@ -49,9 +49,13 @@ func TestNoProfilesUnchanged(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(), // kill-switch closed => Final "block"
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("a", "a.example"),
inlineDomainSet("b", "b.example"),
},
Rules: []model.Rule{
{Name: "a", Enabled: true, Order: 10, DstDomain: []string{"a.example"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstDomain: []string{"b.example"}, Target: "node:ss1"},
{Name: "a", Enabled: true, Order: 10, DstRuleset: []string{"a"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstRuleset: []string{"b"}, Target: "node:ss1"},
},
}
rt, b := buildRouteAt(m, wed12UTC)
@@ -70,7 +74,7 @@ func TestNoProfilesUnchanged(t *testing.T) {
if len(rt.Rules) != 4 {
t.Fatalf("route rule count = %d, want 4 (sniff+hijack+2 user)", len(rt.Rules))
}
if !hasDomainRule(rt, "a.example") || !hasDomainRule(rt, "b.example") {
if !hasRulesetRule(rt, "a") || !hasRulesetRule(rt, "b") {
t.Fatalf("both user rules should survive unchanged")
}
}
@@ -83,9 +87,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("blocked", "blocked.example"),
inlineDomainSet("keep", "keep.example"),
},
Rules: []model.Rule{
{Name: "blockme", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstDomain: []string{"keep.example"}, Target: "direct"},
{Name: "blockme", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstRuleset: []string{"keep"}, Target: "direct"},
},
Profiles: []model.Profile{
// Manual pin: honored regardless of conditions. Disables "blockme".
@@ -94,10 +102,10 @@ func TestActiveProfileDisablesRule(t *testing.T) {
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "blocked.example") {
if hasRulesetRule(rt, "blocked") {
t.Fatalf("profile disabled rule 'blockme' but its matcher is still emitted")
}
if !hasDomainRule(rt, "keep.example") {
if !hasRulesetRule(rt, "keep") {
t.Fatalf("rule 'keep' should be untouched")
}
if rt.Final != tagBlock {
@@ -115,9 +123,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
func TestAutoSelectHighestPriority(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
inlineDomainSet("lo", "lo.example"),
inlineDomainSet("hi", "hi.example"),
},
Rules: []model.Rule{
{Name: "rLo", Enabled: true, Order: 10, DstDomain: []string{"lo.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rLo", Enabled: true, Order: 10, DstRuleset: []string{"lo"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "lo", Enabled: true, Priority: 5, DisableRules: []string{"rLo"}},
@@ -125,10 +137,10 @@ func TestAutoSelectHighestPriority(t *testing.T) {
},
}
rt, _ := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("the highest-priority profile must win and disable rHi")
}
if !hasDomainRule(rt, "lo.example") {
if !hasRulesetRule(rt, "lo") {
t.Fatalf("only the winning profile applies (rLo should survive)")
}
}
@@ -142,9 +154,13 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
g.ActiveProfile = active
return &model.Model{
Globals: g,
Rulesets: []model.Ruleset{
inlineDomainSet("hi", "hi.example"),
inlineDomainSet("if", "if.example"),
},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstDomain: []string{"if.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstRuleset: []string{"if"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
@@ -156,19 +172,19 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
// Auto-select (no pin): iface profile skipped (the watcher owns it); 'plain' wins.
rt, _ := buildRouteAt(mk(""), wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("auto-select should apply 'plain' and disable rHi")
}
if !hasDomainRule(rt, "if.example") {
if !hasRulesetRule(rt, "if") {
t.Fatalf("iface profile 'wwan' must be skipped by auto-select (rIf should survive)")
}
// Named explicitly: iface profile honored regardless of its iface condition.
rt, _ = buildRouteAt(mk("wwan"), wed12UTC)
if hasDomainRule(rt, "if.example") {
if hasRulesetRule(rt, "if") {
t.Fatalf("explicitly named iface profile must be honored and disable rIf")
}
if !hasDomainRule(rt, "hi.example") {
if !hasRulesetRule(rt, "hi") {
t.Fatalf("only the named profile applies; rHi should survive")
}
}
@@ -179,14 +195,15 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
g := model.DefaultGlobals()
g.ActiveProfile = "ghost"
m := &model.Model{
Globals: g,
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstDomain: []string{"x.example"}, Target: "direct"}},
Globals: g,
Rulesets: []model.Ruleset{inlineDomainSet("x", "x.example")},
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstRuleset: []string{"x"}, Target: "direct"}},
Profiles: []model.Profile{
{Name: "auto", Enabled: true, Priority: 1, DisableRules: []string{"rX"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "x.example") {
if hasRulesetRule(rt, "x") {
t.Fatalf("fallback auto-select should apply 'auto' and disable rX")
}
if !hasWarning(b, "falling back to auto-select") {
@@ -208,16 +225,17 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
// suppress it or to warn about it.
func TestPlainProfileIsSelectableAfterProbeRemoval(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{inlineDomainSet("hi", "hi.example")},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("a plain profile must apply and disable rHi")
}
// No leftover diagnostic: warning about an option the parser no longer reads
+8 -110
View File
@@ -1,8 +1,6 @@
package generate
import (
"fmt"
"regexp"
"sort"
"strings"
@@ -321,114 +319,14 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
matched = true
}
// Destination domains. The route-rule-specific prefixes (geosite:/regexp:/
// suffix:) are peeled off here; everything else goes through the shared
// classifier in dnsfilter.go with bareIsSuffix=FALSE — a bare entry in a
// ROUTING rule is an exact domain, unlike the DNS/filter lists where it means
// "and all subdomains". classifyDomainEntries is also what drops marker-only
// entries ("." / "full:" / "keyword:"), which must never reach the engine:
// an empty domain/domain_suffix aborts box.New, and an empty domain_keyword
// is strings.Contains(host, "") — a silent match on EVERY host.
var explicitSuffix []string
var plain []string
for _, d := range r.DstDomain {
d = strings.TrimSpace(d)
if d == "" {
continue
}
switch {
case strings.HasPrefix(d, "geosite:"):
// A `geosite:` matcher is NOT emittable: the route-rule geosite field was
// removed in this engine and route/rule.NewDefaultRule hard-errors on it,
// which aborts box.New for the WHOLE config. Mirroring the geoip handling
// below, it is treated as inert: warned and omitted, so a legacy/UCI rule
// carrying one degrades to "this rule does nothing" instead of taking the
// entire tunnel down. Use a `config ruleset` with source=geosite instead.
b.warnf("rule %q: geosite matcher %q is removed from this engine — use a ruleset with source=geosite instead, omitted (inert)", r.Name, d)
case strings.HasPrefix(d, "regexp:"):
// route/rule.NewDomainRegexItem hard-errors on an uncompilable pattern and
// takes box.New with it; validate here and drop the bad one with a warning.
// A BARE "regexp:" compiles fine but matches every host — same silent
// match-all hazard as an empty keyword, so it is dropped too.
re := strings.TrimSpace(strings.TrimPrefix(d, "regexp:"))
if re == "" {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", r.Name, d)
break
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("rule %q: domain regexp %q is invalid (%v), omitted", r.Name, re, err)
break
}
raw.DomainRegex = append(raw.DomainRegex, re)
matched = true
case strings.HasPrefix(d, "suffix:"):
// Label-aware suffix (apex + subdomains): sing-box domain_suffix stored
// in bare form matches both "example.com" and "*.example.com" (see
// sing/common/domain matcher). A leading-dot entry, by contrast, matches
// subdomains ONLY, so the explicit `suffix:` form is how presets/rules
// ask for apex-inclusive suffix matching.
if s := strings.TrimSpace(strings.TrimPrefix(d, "suffix:")); s != "" {
explicitSuffix = append(explicitSuffix, s)
} else {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New)", r.Name, d)
}
default:
plain = append(plain, d)
}
}
domain, suffix, keyword := classifyDomainEntries(plain, false)
for _, d := range plain {
if isDomainMarkerOnly(d) {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword matches EVERY host)", r.Name, d)
}
}
// An unrecognised `word:` prefix is DROPPED by classifyDomainEntries (a domain
// cannot contain ":", so keeping it would load a provably unmatchable literal).
// It has to be reported here or the rule silently loses that destination — the
// v0.1/xray spelling `domain:example.com` is exactly what someone migrating
// writes, and it used to disappear without a trace. geosite:/regexp:/suffix:
// were already peeled off above, so `plain` carries only the shared markers.
b.warnUnrecognisedPrefixes(fmt.Sprintf("rule %q", r.Name), plain)
suffix = append(explicitSuffix, suffix...)
if len(domain) > 0 {
raw.Domain = badoption.Listable[string](domain)
matched = true
}
if len(suffix) > 0 {
raw.DomainSuffix = badoption.Listable[string](suffix)
matched = true
}
if len(keyword) > 0 {
raw.DomainKeyword = badoption.Listable[string](keyword)
matched = true
}
// Destination IPs. A `geoip:<code>` entry is NOT a CIDR: route-rule geoip was
// removed in this engine (box.New hard-errors on it), so — mirroring how the
// ruleset geoip source is handled — it is treated as inert: warned and omitted
// (never emitted as an ip_cidr, which would also abort box.New). This keeps a
// geoip-driven rule/preset (e.g. the ru-bypass pack) fail-open instead of fatal.
var ipcidr []string
for _, ip := range r.DstIP {
ip = strings.TrimSpace(ip)
if ip == "" {
continue
}
if strings.HasPrefix(strings.ToLower(ip), "geoip:") {
b.warnf("rule %q: geoip matcher %q is removed from this engine — use a ruleset with source=geoip instead, omitted (inert)", r.Name, ip)
continue
}
ipcidr = append(ipcidr, ip)
}
// Same fail-open validation as the source list: an unparseable ip_cidr aborts
// box.New for the whole config, so drop it with a warning instead.
ipcidr, badIP := validPrefixes(ipcidr)
for _, s := range badIP {
b.warnf("rule %q: destination %q is not a valid IP/CIDR, omitted", r.Name, s)
}
if len(ipcidr) > 0 {
raw.IPCIDR = badoption.Listable[string](ipcidr)
matched = true
}
// Destination: a rule's ONLY destination matcher is DstRuleset (schema v2).
// The inline `dst_domain` / `dst_ip` lists that used to be classified here are
// gone. A destination list is a `config ruleset` — compiled once into a .srs and
// shared by every rule that references it — so the prefix vocabulary
// (full:/suffix:/keyword:/regexp:, a leading dot) and the geosite/geoip sources
// live in exactly one place (generate/ruleset.go inlineRulesetRule). Existing
// configs were rewritten by model.migrate1to2, which preserves each entry's
// meaning 1:1.
// DstRuleset -> reference the rs-<name> rule-sets materialised by
// buildRoutingRuleSets. A dst_ruleset naming an UNDEFINED ruleset is warned +
+127 -72
View File
@@ -8,6 +8,7 @@ import (
"strings"
"testing"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
@@ -22,26 +23,45 @@ func killModel(kill string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "social", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "group:grp", Kill: kill},
{Name: "social", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "group:grp", Kill: kill},
{Name: "catch-tcp", Enabled: true, Order: 20, Proto: "tcp", Target: "direct"},
},
}
}
// domainRuleTarget returns the outbound the rule matching domain routes to.
func domainRuleTarget(rt *option.RouteOptions, domain string) (string, bool) {
for _, r := range generalRules(rt) {
for _, d := range r.DefaultOptions.RawDefaultRule.Domain {
if d == domain {
return r.DefaultOptions.RuleAction.RouteOptions.Outbound, true
}
}
// rulesetRuleTarget returns the outbound the rule referencing the rule-set named
// name routes to. It is the post-schema-v2 replacement for looking a rule up by
// its dst-domain matcher: a rule's destination is a rule_set reference now, so
// the matcher that identifies it is the rs-<name> tag.
func rulesetRuleTarget(rt *option.RouteOptions, name string) (string, bool) {
dr := findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name)
if dr == nil {
return "", false
}
return "", false
return dr.RuleAction.RouteOptions.Outbound, true
}
// soleInlineRule returns the single default headless rule an inline rule-set is
// expected to carry, failing the ASSERTION (rather than panicking on an index) when
// the set turns out to be remote/local or to hold a different rule shape.
func soleInlineRule(t *testing.T, rs option.RuleSet) option.DefaultHeadlessRule {
t.Helper()
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
t.Fatalf("rule-set %q is type %q, not inline — it has no rules in the config to inspect", rs.Tag, rs.Type)
}
if len(rs.InlineOptions.Rules) != 1 {
t.Fatalf("rule-set %q must carry exactly one headless rule, got %d", rs.Tag, len(rs.InlineOptions.Rules))
}
hr := rs.InlineOptions.Rules[0]
if hr.Type != C.RuleTypeDefault && hr.Type != "" {
t.Fatalf("rule-set %q carries a %q headless rule, not a default one", rs.Tag, hr.Type)
}
return hr.DefaultOptions
}
// TestRuleKillDefaultBlocks: kill="" / "default" is fail-closed — the rule is
@@ -54,7 +74,7 @@ func TestRuleKillDefaultBlocks(t *testing.T) {
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("%q: rule must be emitted routing to block, but it was dropped", kill)
}
@@ -76,7 +96,7 @@ func TestRuleKillClosedBlocksHere(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=closed must still emit the rule (warnings %v)", warns)
}
@@ -103,7 +123,7 @@ func TestRuleKillOpenGoesDirect(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=open must still emit the rule (warnings %v)", warns)
}
@@ -122,7 +142,7 @@ func TestRuleKillUnknownBlocks(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != tagBlock {
t.Fatalf("unknown kill policy must block, got %q (ok=%v)", got, ok)
}
@@ -145,7 +165,7 @@ func TestRuleKillPreservesFailClosedInvariant(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("%q: Final = %q, want block", kill, opts.Route.Final)
}
if tgt, ok := domainRuleTarget(opts.Route, "social.example"); ok && tgt == tagDirect {
if tgt, ok := rulesetRuleTarget(opts.Route, "social"); ok && tgt == tagDirect {
t.Fatalf("%q: dead group leaked direct", kill)
}
}
@@ -158,17 +178,18 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "ok", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "node:n1", Kill: kill},
{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "node:n1", Kill: kill},
},
}
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != "n1" {
t.Fatalf("%q: healthy target must win, got %q (ok=%v)", kill, got, ok)
}
@@ -176,29 +197,54 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
}
// --- R4: marker-only domain entries ------------------------------------------
//
// R4 moved with the destination list itself: a rule's domains are an inline
// `config ruleset` now (schema v2), so the marker-only guard has to hold in
// inlineRulesetRule rather than in ruleMatchers. The hazard is unchanged.
// TestRuleMarkerOnlyDomainEntriesDropped: an entry that is nothing but its marker
// must never reach the engine. An empty domain/domain_suffix makes NewDomainItem
// return "empty item is not allowed" and aborts box.New (whole LAN down from one
// stray "."); an empty domain_keyword is SILENT and matches every host.
func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
// TestRulesetMarkerOnlyDomainEntriesDropped: an entry that is nothing but its
// marker must never reach the engine. An empty domain/domain_suffix makes
// NewDomainItem return "empty item is not allowed" and aborts box.New (whole LAN
// down from one stray "."); an empty domain_keyword is SILENT and matches every
// host. As the sole entry it also leaves the rule-set with nothing usable, so the
// list is skipped and the rule referencing it is not emitted — never promoted to
// "matches everything".
func TestRulesetMarkerOnlyDomainEntriesDropped(t *testing.T) {
for _, entry := range []string{".", "full:", "keyword:", "suffix:", "regexp:", " keyword: "} {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{entry}, Target: "node:n1",
})
for _, r := range rt.Rules {
raw := r.DefaultOptions.RawDefaultRule
for _, list := range [][]string{raw.Domain, raw.DomainSuffix, raw.DomainKeyword, raw.DomainRegex} {
for _, v := range list {
if strings.TrimSpace(v) == "" {
t.Fatalf("%q: emitted an EMPTY matcher token", entry)
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m", entry)},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
// Only an INLINE rule-set has its rules in the options at all; a remote or
// local one (a geosite chip, a compiled url list) carries a URL or a path and
// an empty InlineOptions. Indexing Rules[0] unconditionally turned any such
// fixture into an index-out-of-range PANIC instead of a failed assertion, so
// the shape is checked rather than assumed.
for _, rs := range rt.RuleSet {
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
continue
}
for i, hrule := range rs.InlineOptions.Rules {
if hrule.Type != C.RuleTypeDefault && hrule.Type != "" {
continue // a logical headless rule has no matcher lists of its own
}
hr := hrule.DefaultOptions
for name, list := range map[string][]string{
"domain": hr.Domain, "domain_suffix": hr.DomainSuffix,
"domain_keyword": hr.DomainKeyword, "domain_regex": hr.DomainRegex,
} {
for _, v := range list {
if strings.TrimSpace(v) == "" {
t.Fatalf("%q: rule-set %q rule %d emitted an EMPTY %s token (an empty domain aborts box.New; an empty keyword matches every host)",
entry, rs.Tag, i, name)
}
}
}
}
}
// Sole entry => nothing left to match on => the rule must be skipped, never
// silently promoted to "matches everything".
if _, ok := ruleSetByTag(rt, "rs-m"); ok {
t.Fatalf("%q: a marker-only list must not materialise a rule-set", entry)
}
if n := len(generalRules(rt)); n != 0 {
t.Fatalf("%q: marker-only sole entry must skip the rule, got %d rules", entry, n)
}
@@ -208,46 +254,55 @@ func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
}
}
// TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must survive
// the marker-only ones.
func TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{".", "keyword:", "exact.example", "keyword:ads", ".sub.example", "suffix:apex.example"},
Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d (warnings %v)", len(gen), warns)
// TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must
// survive the marker-only ones.
func TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m",
".", "keyword:", "full:exact.example", "keyword:ads", ".sub.example", "suffix:apex.example")},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-m")
if !ok {
t.Fatalf("the usable entries must keep the list alive (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example] (bare entry in a ROUTING rule is exact)", raw.Domain)
hr := soleInlineRule(t, rs)
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example]", hr.Domain)
}
if len(raw.DomainKeyword) != 1 || raw.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", raw.DomainKeyword)
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", hr.DomainKeyword)
}
if len(raw.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", raw.DomainSuffix)
if len(hr.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", hr.DomainSuffix)
}
if findRouteRuleWithRuleSet(rt, "rs-m") == nil {
t.Fatalf("the rule must be emitted referencing rs-m; rules=%+v", rt.Rules)
}
}
// TestRuleBareEntryIsExactDomain pins the routing convention (bare == exact),
// which deliberately differs from the DNS/filter lists (bare == suffix).
func TestRuleBareEntryIsExactDomain(t *testing.T) {
rt, _ := genRules(t, model.Rule{
Name: "b", Enabled: true, Order: 10,
DstDomain: []string{"example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d", len(gen))
// TestRulesetBareEntryIsDomainSuffix pins the convention a destination list now
// follows — and it is the OPPOSITE of the one the old dst_domain used.
//
// A bare entry in a routing rule meant one EXACT domain; a bare entry in a
// rule-set means the domain AND its subdomains (the DNS/filter/device convention,
// classifyDomainEntries with bareIsSuffix=true). `full:` is how an exact match is
// written now. model.migrate1to2 rewrites old dst_domain entries accordingly, so
// this asymmetry is the thing that migration has to get right.
func TestRulesetBareEntryIsDomainSuffix(t *testing.T) {
rt, _ := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("b", "example.com", "full:exact.example")},
model.Rule{Name: "b", Enabled: true, Order: 10, DstRuleset: []string{"b"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-b")
if !ok {
t.Fatalf("rs-b not emitted; route=%+v", rt)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("bare entry must be an exact Domain, got domain=%v suffix=%v", raw.Domain, raw.DomainSuffix)
hr := soleInlineRule(t, rs)
if len(hr.DomainSuffix) != 1 || hr.DomainSuffix[0] != "example.com" {
t.Fatalf("bare entry must become a domain_suffix, got suffix=%v domain=%v", hr.DomainSuffix, hr.Domain)
}
if len(raw.DomainSuffix) != 0 {
t.Fatalf("bare entry must NOT become a suffix, got %v", raw.DomainSuffix)
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("full: must be the exact form, got domain=%v", hr.Domain)
}
}
+35 -14
View File
@@ -18,10 +18,13 @@ import (
)
// TestMalformedMatchersStillApply drives ONE model carrying every previously
// fatal matcher — a geosite: domain, a bad ip_cidr, a bad source cidr, an
// fatal matcher — a geosite: entry, a bad ip_cidr, a bad source cidr, an
// uncompilable regexp and a malformed port range — plus a healthy rule, through
// engine.Apply. It must come up, the healthy rule must survive, and the
// kill-switch backstop must stay closed.
// kill-switch backstop must stay closed. The destination lists are inline
// `config ruleset`s (schema v2), which is where the domain/IP guards live now;
// a rule-set left with no usable entry is skipped, and so is the rule whose only
// matcher it was.
func TestMalformedMatchersStillApply(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
@@ -29,13 +32,19 @@ func TestMalformedMatchersStillApply(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12395, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("geosite", "geosite:youtube"),
inlineIPSet("badip", "999.1.1.1/24"),
inlineDomainSet("badre", "regexp:*broken("),
inlineDomainSet("ok", "ok.example"),
},
Rules: []model.Rule{
{Name: "geosite", Enabled: true, Order: 1, DstDomain: []string{"geosite:youtube"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstIP: []string{"999.1.1.1/24"}, Target: "node:n1"},
{Name: "geosite", Enabled: true, Order: 1, DstRuleset: []string{"geosite"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstRuleset: []string{"badip"}, Target: "node:n1"},
{Name: "badsrc", Enabled: true, Order: 3, Src: []string{"192.168.0.0/99"}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstDomain: []string{"regexp:*broken("}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstRuleset: []string{"badre"}, Target: "node:n1"},
{Name: "badport", Enabled: true, Order: 5, DstPort: "a-b", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstDomain: []string{"ok.example"}, DstPort: "443", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstRuleset: []string{"ok"}, DstPort: "443", Target: "node:n1"},
},
}
@@ -46,14 +55,21 @@ func TestMalformedMatchersStillApply(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("Final = %q, want block", opts.Route.Final)
}
if !hasDomainRule(opts.Route, "ok.example") {
if !hasRulesetRule(opts.Route, "ok") {
t.Fatalf("the healthy rule must survive alongside the malformed ones")
}
if !hasRouteToOutbound(opts, "n1") {
t.Fatalf("expected a route rule to n1")
}
// Every malformed matcher reported itself rather than failing silently.
for _, want := range []string{"geosite matcher", "is not a valid IP/CIDR", "domain regexp", "is not a valid port/range"} {
for _, want := range []string{
"unrecognised prefix", // geosite: in a domain list
"no usable entries", // ...leaving that list empty
"bad ip_cidr entry", // 999.1.1.1/24
"is not a valid IP/CIDR", // the source cidr
"domain regexp", // regexp:*broken(
"is not a valid port/range", // a-b
} {
if !routeWarnsHave(warns, want) {
t.Fatalf("missing diagnostic %q in %v", want, warns)
}
@@ -155,10 +171,15 @@ func TestRuleKillPoliciesApply(t *testing.T) {
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12398, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "dead", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#dead"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"dead"}}},
Rulesets: []model.Ruleset{
inlineDomainSet("k-closed", "closed.example"),
inlineDomainSet("k-open", "open.example"),
inlineDomainSet("k-default", "default.example"),
},
Rules: []model.Rule{
{Name: "k-closed", Enabled: true, Order: 1, DstDomain: []string{"closed.example"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstDomain: []string{"open.example"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstDomain: []string{"default.example"}, Target: "group:grp"},
{Name: "k-closed", Enabled: true, Order: 1, DstRuleset: []string{"k-closed"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstRuleset: []string{"k-open"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstRuleset: []string{"k-default"}, Target: "group:grp"},
},
}
@@ -166,13 +187,13 @@ func TestRuleKillPoliciesApply(t *testing.T) {
if !changed {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
if got, ok := domainRuleTarget(opts.Route, "closed.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-closed"); !ok || got != tagBlock {
t.Fatalf("kill=closed must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "open.example"); !ok || got != tagDirect {
if got, ok := rulesetRuleTarget(opts.Route, "k-open"); !ok || got != tagDirect {
t.Fatalf("kill=open must emit a rule routed to direct, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "default.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-default"); !ok || got != tagBlock {
t.Fatalf("kill=default must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if opts.Route.Final != tagBlock {
+209 -85
View File
@@ -45,6 +45,40 @@ func genRules(t *testing.T, rules ...model.Rule) (*option.RouteOptions, []string
return opts.Route, warns
}
// ruleModelWithSets is ruleModel plus the `config ruleset` definitions the rules'
// DstRuleset entries point at. A rule's ONLY destination matcher is a rule-set
// (schema v2), so every case below that just needs "some destination the engine
// can match on" declares one here rather than writing an inline domain/IP list.
func ruleModelWithSets(sets []model.Ruleset, rules ...model.Rule) *model.Model {
m := ruleModel(rules...)
m.Rulesets = sets
return m
}
// genRulesWithSets is genRules for a model that also carries rule-sets.
func genRulesWithSets(t *testing.T, sets []model.Ruleset, rules ...model.Rule) (*option.RouteOptions, []string) {
t.Helper()
opts, warns, err := GenerateWithWarnings(ruleModelWithSets(sets, rules...))
if err != nil {
t.Fatalf("Generate: %v", err)
}
return opts.Route, warns
}
// inlineDomainSet builds an inline DOMAIN `config ruleset`.
//
// Mind the convention the move to rule-sets brought with it: a BARE entry here is
// a DomainSuffix (the apex AND its subdomains), whereas the routing rule's old
// dst_domain read a bare entry as an EXACT domain. `full:` is the exact form.
func inlineDomainSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "domain", Source: "inline", Entries: entries}
}
// inlineIPSet builds an inline IP-RANGE `config ruleset`.
func inlineIPSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: entries}
}
func routeWarnsHave(warns []string, substr string) bool {
for _, w := range warns {
if strings.Contains(w, substr) {
@@ -68,69 +102,117 @@ func generalRules(rt *option.RouteOptions) []option.Rule {
// --- geosite: the landmine ---------------------------------------------------
// TestGeositeMatcherIsInertNotFatal: route-rule `geosite` was REMOVED from this
// engine — route/rule.NewDefaultRule returns "geosite database is deprecated ...
// removed in sing-box 1.12.0" for a non-empty Geosite list, and that error aborts
// box.New for the whole config. A `geosite:` entry must therefore be warned and
// omitted (exactly like the geoip matcher below), never emitted.
func TestGeositeMatcherIsInertNotFatal(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "geo", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube"}, Target: "node:n1",
})
// TestGeositeEntryInRulesetIsInertNotFatal: route-rule `geosite` was REMOVED from
// this engine — route/rule.NewDefaultRule returns "geosite database is deprecated
// ... removed in sing-box 1.12.0" for a non-empty Geosite list, and that error
// aborts box.New for the whole config. Destinations now live in a `config
// ruleset`, so a `geosite:` entry lands in an inline domain list, where it is an
// unrecognised `word:` prefix: dropped by the classifier, reported, and — being
// the list's only entry — leaving the rule-set with nothing to match, so it is
// skipped and the rule that referenced it is not emitted either. Nothing about
// that path may ever put a value in RawDefaultRule.Geosite. (`source=geosite` on
// the ruleset itself is the working way to route a category; see
// TestRoutingRuleSetGeositeCategory.)
func TestGeositeEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("geo", "geosite:youtube")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.Geosite) > 0 {
t.Fatalf("geosite must never be emitted (aborts box.New), got %v", r.DefaultOptions.RawDefaultRule.Geosite)
}
}
if !routeWarnsHave(warns, "geosite matcher") {
t.Fatalf("expected an inert-geosite warning, got %v", warns)
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a rule-set with no usable entry must not be emitted (an empty list would match everything)")
}
if findRouteRuleWithRuleSet(rt, "rs-geo") != nil {
t.Fatalf("no rule may reference the skipped rule-set")
}
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("expected an unrecognised-prefix warning for geosite:, got %v", warns)
}
if !routeWarnsHave(warns, "no usable entries") {
t.Fatalf("expected a no-usable-entries warning, got %v", warns)
}
}
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule carrying BOTH a geosite entry
// and a real domain keeps the real matcher and still routes — only the geosite
// part is dropped.
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule-set carrying BOTH a geosite
// entry and a real domain keeps the real matcher, and the rule referencing it
// still routes — only the geosite entry is dropped.
func TestGeositeMixedWithRealDomainKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube", "example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d (warnings %v)", len(gen), warns)
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mixed", "geosite:youtube", "full:example.com")},
model.Rule{Name: "mixed", Enabled: true, Order: 10, DstRuleset: []string{"mixed"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-mixed")
if !ok {
t.Fatalf("rs-mixed must survive the geosite entry (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Geosite) != 0 {
t.Fatalf("geosite leaked: %v", raw.Geosite)
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", hr)
}
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", raw.Domain)
if len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.DomainRegex) != 0 {
t.Fatalf("geosite: must be dropped, not reinterpreted: %+v", hr)
}
if got := gen[0].DefaultOptions.RuleAction.RouteOptions.Outbound; got != "n1" {
dr := findRouteRuleWithRuleSet(rt, "rs-mixed")
if dr == nil {
t.Fatalf("the rule must be emitted referencing rs-mixed; rules=%+v", rt.Rules)
}
if got := dr.RuleAction.RouteOptions.Outbound; got != "n1" {
t.Fatalf("target = %q, want n1", got)
}
}
// TestGeoipEntryInRulesetIsInertNotFatal is the IP-side twin: model.migrate1to2
// moves an old `dst_ip geoip:ru` verbatim into an inline type=ipcidr rule-set
// (deliberately — see TestMigrate1to2KeepsGeoMarkersInert: it must not silently
// become a working geoip source, because the operator never asked to download
// anything). Here it is an unparseable prefix, so it must be dropped LOUDLY
// rather than reach NewIPCIDRItem, which errors and aborts box.New.
func TestGeoipEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("geo", "geoip:ru")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.GeoIP) > 0 {
t.Fatalf("geoip must never be emitted, got %v", r.DefaultOptions.RawDefaultRule.GeoIP)
}
}
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a list whose only entry is unparseable must not materialise")
}
if !routeWarnsHave(warns, `bad ip_cidr entry "geoip:ru"`) {
t.Fatalf("expected a bad-ip_cidr warning naming the entry, got %v", warns)
}
}
// --- malformed matchers that used to abort box.New ---------------------------
// TestBadDstCIDRWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. It must be dropped.
func TestBadDstCIDRWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstIP: []string{"999.1.1.1/24", "198.51.100.0/24"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadIPCIDREntryWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. Destination addresses are
// an inline `type=ipcidr` rule-set now, so the guard lives in inlineRulesetRule:
// the typo is dropped, the valid entry survives and the rule still routes.
func TestRulesetBadIPCIDREntryWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("bad", "999.1.1.1/24", "198.51.100.0/24")},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one bad entry must not take the whole list down (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.IPCIDR) != 1 || raw.IPCIDR[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", raw.IPCIDR)
got := rs.InlineOptions.Rules[0].DefaultOptions.IPCIDR
if len(got) != 1 || got[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", got)
}
if !routeWarnsHave(warns, `destination "999.1.1.1/24" is not a valid IP/CIDR`) {
t.Fatalf("expected a bad-destination warning, got %v", warns)
if !routeWarnsHave(warns, `bad ip_cidr entry "999.1.1.1/24"`) {
t.Fatalf("expected a bad-ip_cidr warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
@@ -152,24 +234,30 @@ func TestBadSrcCIDRWarnsAndSkips(t *testing.T) {
}
}
// TestBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New.
func TestBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstDomain: []string{"regexp:*broken(", `regexp:^ok\.example$`}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New. The pattern vocabulary
// moved into the inline rule-set with the rest of the destination list, so the
// validation moved with it (peelDomainRegexes): the broken pattern is dropped and
// the compilable one survives.
func TestRulesetBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("bad", "regexp:*broken(", `regexp:^ok\.example$`)},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one broken pattern must not take the whole list down (warnings %v)", warns)
}
got := gen[0].DefaultOptions.RawDefaultRule.DomainRegex
got := rs.InlineOptions.Rules[0].DefaultOptions.DomainRegex
if len(got) != 1 || got[0] != `^ok\.example$` {
t.Fatalf("domain_regex = %v, want only the compilable one", got)
}
if !routeWarnsHave(warns, "domain regexp") {
t.Fatalf("expected a bad-regexp warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
// TestBadPortRangeWarnsAndSkips: a malformed range reaches
@@ -911,10 +999,12 @@ func TestRuleProtoKnownValuesAreSilent(t *testing.T) {
// that would break existing configs either open or closed), but the widening is
// now reported with its consequence.
func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstDomain: []string{"youtube.com"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "applies to EVERY client") {
t.Fatalf("expected a rule-widening warning, got %v", warns)
}
@@ -930,11 +1020,13 @@ func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
// TestMixedSourceRuleReportsTheDroppedHalf: with one usable IP source alongside
// an unmatchable one, the rule narrows to the IP source only.
func TestMixedSourceRuleReportsTheDroppedHalf(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstDomain: []string{"youtube.com"}, Target: "block",
})
_, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "could not be used") {
t.Fatalf("expected a dropped-source warning, got %v", warns)
}
@@ -1041,41 +1133,73 @@ func TestRuleTargetKindsResolve(t *testing.T) {
}
}
// TestRuleDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped by
// the shared classifier (a domain cannot contain ":"), so the rule silently lost
// that destination. `domain:example.com` is the v0.1/xray spelling a migrating
// user writes, and it must not disappear without a trace.
func TestRuleDomainUnrecognisedPrefixWarns(t *testing.T) {
// TestRulesetDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped
// by the shared classifier (a domain cannot contain ":"), so the list silently
// lost that destination. `domain:example.com` is the v0.1/xray spelling a
// migrating user writes, and it must not disappear without a trace — the more so
// now that a destination list is ALWAYS a rule-set, i.e. the one place a typo can
// hide.
func TestRulesetDomainUnrecognisedPrefixWarns(t *testing.T) {
for _, entry := range []string{"domain:example.com", "regex:example.com", "ext:foo.dat:cn"} {
rt, warns := genRules(t, model.Rule{
Name: "mig", Enabled: true, Order: 10,
DstDomain: []string{entry, "keep.example"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mig", entry, "keep.example")},
model.Rule{Name: "mig", Enabled: true, Order: 10, DstRuleset: []string{"mig"}, Target: "block"},
)
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("%q: expected an unrecognised-prefix warning, got %v", entry, warns)
}
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("%q: want 1 rule, got %d", entry, len(gen))
rs, ok := ruleSetByTag(rt, "rs-mig")
if !ok {
t.Fatalf("%q: the usable entry must keep the rule-set alive; warnings %v", entry, warns)
}
for _, d := range gen[0].DefaultOptions.RawDefaultRule.Domain {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
hr := rs.InlineOptions.Rules[0].DefaultOptions
for _, list := range [][]string{hr.Domain, hr.DomainSuffix, hr.DomainKeyword, hr.DomainRegex} {
for _, d := range list {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
}
}
}
if findRouteRuleWithRuleSet(rt, "rs-mig") == nil {
t.Fatalf("%q: the rule must still be emitted; rules=%+v", entry, rt.Rules)
}
}
}
// TestRuleDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary routing rules really support.
func TestRuleDomainKnownPrefixesAreSilent(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "ok", Enabled: true, Order: 10, Target: "block",
DstDomain: []string{"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example"},
})
// TestRulesetDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary an inline domain rule-set really supports.
//
// `regexp:` is the load-bearing case: the shared classifier has no branch for it,
// so it WOULD be reported as an unknown prefix — inlineRulesetRule peels the
// regexes off first (peelDomainRegexes) precisely so it is not. A regression there
// would both warn about a working matcher and drop it.
func TestRulesetDomainKnownPrefixesAreSilent(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("ok",
"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example")},
model.Rule{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"ok"}, Target: "block"},
)
if routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("the supported prefixes must not warn: %v", warns)
}
rs, ok := ruleSetByTag(rt, "rs-ok")
if !ok {
t.Fatalf("rs-ok not emitted; warnings %v", warns)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "a.example" {
t.Fatalf("full: must be an exact Domain, got %v", hr.Domain)
}
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^d\.` {
t.Fatalf("regexp: must survive as a domain_regex matcher, got %v", hr.DomainRegex)
}
// suffix:, the leading dot and the BARE entry all collapse to domain_suffix.
if len(hr.DomainSuffix) != 3 {
t.Fatalf("domain_suffix = %v, want b/e/f.example (bare entry is a suffix in a rule-set)", hr.DomainSuffix)
}
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "c" {
t.Fatalf("domain_keyword = %v, want [c]", hr.DomainKeyword)
}
}
// TestDeviceDomainUnrecognisedPrefixWarns is the same guarantee in the place it
+231 -7
View File
@@ -25,6 +25,7 @@ import (
"net/netip"
"os"
"path/filepath"
"regexp"
"runtime/debug"
"strings"
"sync"
@@ -500,6 +501,47 @@ func ruleSetURLIsEngineNative(rawURL string) (format string, native bool) {
}
}
// TWO WAYS TO WRITE A DESTINATION LIST, AND WHY THEY DO NOT SHARE A VOCABULARY.
//
// D21 promises "one destination mechanism, one vocabulary". The vocabulary half of
// that promise is about the ENTRIES AN OPERATOR TYPES, and those live in exactly
// one place: an inline `config ruleset` (inlineRulesetRule -> peelDomainRegexes +
// classifyDomainEntries), where `full:` / `suffix:` / `keyword:` / `regexp:` / a
// leading dot / a bare name all mean what docs-shater/DECISIONS.md D21 says.
//
// A `source=url` list that is not engine-native is NOT another spelling of that.
// It is a FILE FORMAT — the hosts / plain-domain / AdBlock-ish text third parties
// publish — and parseDomainList is a parser for that format, not for our entry
// vocabulary. `source=file` is a third thing again: a compiled .srs or a rule-set
// .json handed straight to the engine, which never sees shater's entry syntax at
// all. (An earlier review read this as "the same list written two ways behaves
// differently"; it is closer to "a typed list and a downloaded file are different
// artifacts". The diagnostics below exist so an operator never has to guess which
// one they are looking at.)
//
// Unifying them was considered and REJECTED, on three grounds:
//
// - The formats collide. A real AdGuard/OISD list is full of colon-bearing lines
// that are not our markers at all (`example.com##.banner:has(...)`, `$domain=`
// options, absolute URL rules). Feeding those through the marker classifier
// would either mis-import them or, if we reported every `word:`-shaped token as
// an unknown prefix, drown the operator in hundreds of warnings per list — a
// louder dishonesty than the quiet one it replaces.
// - The shapes collide. A hosts line carries SEVERAL names ("127.0.0.1 a.com
// b.com"), so this parser works per TOKEN; the entry vocabulary works per LINE
// and allows a space after the marker ("keyword: ads"). There is no split rule
// that serves both.
// - `regexp:` from a URL is regex supplied by a third party, compiled into the
// router's matcher and evaluated per query on a 512 MB box. The inline path can
// accept it because the operator typed it; a downloaded list is not that.
//
// So the difference STAYS, and is paid for in diagnostics instead: a text list that
// carries our marker vocabulary is reported per list (see listEntryMarker and
// warnListEntryVocabulary), naming the entries and where they DO work. The check is
// narrow on purpose — only the four markers D21 defines, never the general `word:`
// shape — so it fires on an operator's mistake and stays silent on ordinary filter
// syntax.
//
// parseDomainList extracts domains from the formats public blocklists ship in:
//
// - HOSTS "0.0.0.0 ads.example.com", "127.0.0.1 a.com b.com"
@@ -517,8 +559,12 @@ func ruleSetURLIsEngineNative(rawURL string) (format string, native bool) {
// de-duplication map is kept — domain.NewMatcher already de-duplicates internally
// while building the succinct set, so a second map would just double the largest
// allocation in the pipeline. See listMaxDomains for the measured budget.
func parseDomainList(content []byte) []string {
//
// The second return reports the inline-vocabulary entries seen on the way past, so
// the caller can say so instead of dropping them without a word.
func parseDomainList(content []byte) ([]string, listMarkerNote) {
out := make([]string, 0, 4096)
var note listMarkerNote
scanner := bufio.NewScanner(bytes.NewReader(content))
// Public lists are one domain per line; 64 KiB is far beyond any real line, and
// an over-long line is skipped rather than aborting the parse.
@@ -544,17 +590,123 @@ func parseDomainList(content []byte) []string {
// AdBlock-ish "||domain^" -> domain.
f = strings.TrimPrefix(f, "||")
f = strings.TrimSuffix(f, "^")
if _, isMarker := listEntryMarker(f); isMarker {
// normaliseListDomain would drop this silently (a domain cannot contain
// ":"). Record it so the caller can name it; it is the one class of junk
// in a text list that is provably an operator mistake rather than filter
// syntax we simply do not import.
note.record(f)
continue
}
d, ok := normaliseListDomain(f)
if !ok {
continue
}
out = append(out, d)
if len(out) >= listMaxDomains {
return out
return out, note
}
}
}
return out
return out, note
}
// listEntryVocabulary is EXACTLY the marker set an INLINE rule-set entry may use
// (D21). It is deliberately NOT the general `word:` shape unrecognisedDomainPrefix
// tests for: a published filter list legitimately contains hundreds of colon-
// bearing tokens, and reporting those would make the diagnostic useless. These
// four, by contrast, appear in a downloaded text list only when a human wrote them
// there expecting shater to honour them.
var listEntryVocabulary = []string{"full:", "suffix:", "keyword:", "regexp:"}
// listEntryMarker reports whether a text-list token is written in the inline entry
// vocabulary, and which marker it used.
func listEntryMarker(token string) (string, bool) {
lower := strings.ToLower(strings.TrimSpace(token))
for _, m := range listEntryVocabulary {
if strings.HasPrefix(lower, m) {
return m, true
}
}
return "", false
}
// listMarkerSamples bounds how many offending entries a warning quotes. A list is
// remote content: it must not be able to write an unbounded amount into our log.
const listMarkerSamples = 5
// listMarkerNote records the inline-vocabulary entries one plain-text list carried.
type listMarkerNote struct {
Count int
Samples []string
}
func (n *listMarkerNote) record(entry string) {
n.Count++
if len(n.Samples) < listMarkerSamples {
n.Samples = append(n.Samples, strings.TrimSpace(entry))
}
}
func (n listMarkerNote) empty() bool { return n.Count == 0 }
// listMarkerMemo remembers, per list URL, what the last COMPILATION of that list
// found. Without it the diagnostic would exist for exactly one reconcile — the one
// that happened to refresh the artifact — and then vanish for a whole
// update_interval, which is precisely the "reported to nobody" failure it is meant
// to fix. Same shape (and same reasoning) as ruleSetProbeCache above; a refresh
// that finds nothing clears the entry, so fixing the list silences it.
var (
listMarkerMu sync.Mutex
listMarkerMemo = map[string]listMarkerNote{}
)
func rememberListMarkers(url string, note listMarkerNote) {
listMarkerMu.Lock()
if note.empty() {
delete(listMarkerMemo, url)
} else {
listMarkerMemo[url] = note
}
listMarkerMu.Unlock()
}
func recallListMarkers(url string) listMarkerNote {
listMarkerMu.Lock()
defer listMarkerMu.Unlock()
return listMarkerMemo[url]
}
// resetListMarkerMemo clears the memo (tests).
func resetListMarkerMemo() {
listMarkerMu.Lock()
listMarkerMemo = map[string]listMarkerNote{}
listMarkerMu.Unlock()
}
// warnListEntryVocabulary tells the operator that entries written in the INLINE
// entry vocabulary were found in a downloaded TEXT list, where they mean nothing.
// See the parseDomainList block above for why the two vocabularies are separate
// and why saying so is the whole of the fix.
func (b *builder) warnListEntryVocabulary(diag, url string) {
note := recallListMarkers(url)
if note.empty() {
return
}
b.warnf("%s: %q is a plain-text list (a hosts file or one domain per line), but %d of its entries are written in the "+
"inline rule-set vocabulary (e.g. %s) — a plain-text list has NO markers, so every line is read as a domain plus its "+
"subdomains and anything containing \":\" is dropped, because a domain name cannot contain one. Those entries match NOTHING. "+
"Put them in a rule-set with source=inline, which is the one place full:/suffix:/keyword:/regexp: are honoured.",
diag, url, note.Count, quoteList(note.Samples))
}
// quoteList renders sample entries for a diagnostic.
func quoteList(in []string) string {
out := make([]string, 0, len(in))
for _, s := range in {
out = append(out, fmt.Sprintf("%q", s))
}
return strings.Join(out, ", ")
}
// hostsBoilerplate are the names every hosts file carries for its own bookkeeping.
@@ -675,6 +827,11 @@ func (b *builder) compiledListRuleSet(tag, url, updateInterval, diag string) (op
}
}
// Reported on EVERY generate, not only on the one that refreshed the artifact:
// an entry that matches nothing is exactly as wrong the day after it was
// compiled as the moment it was.
b.warnListEntryVocabulary(diag, url)
return option.RuleSet{
Type: C.RuleSetTypeLocal,
Tag: tag,
@@ -716,8 +873,11 @@ func (b *builder) refreshCompiledList(path, url, diag string) error {
if err != nil {
return err
}
domains := parseDomainList(body)
domains, markers := parseDomainList(body)
body = nil // release the source text before the matcher allocates
// Remember (or clear) what this compilation saw, so the diagnostic survives the
// reconciles that do no I/O at all. See warnListEntryVocabulary.
rememberListMarkers(url, markers)
if len(domains) == 0 {
return fmt.Errorf("no usable domains found at %s (fetched %s, but nothing in it parsed as a domain)", url, "the file")
@@ -1163,7 +1323,8 @@ func (b *builder) buildRoutingRuleSetRaw(rs model.Ruleset) ([]option.RuleSet, []
// its Entries, keyed by Type: an ipcidr ruleset fills ip_cidr; a domain ruleset
// (the default) is classified with inlineDomainRule — bare entry => DomainSuffix
// (so subdomains match), full: => Domain, keyword: => DomainKeyword, . => suffix
// — the same classification the DNS filter uses. ok=false when nothing usable.
// — the same classification the DNS filter uses, plus `regexp:` (see
// peelDomainRegexes). ok=false when nothing usable.
func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRule, bool) {
b.warnUnknownRuleSetType(fmt.Sprintf("ruleset %q", rs.Name), rs.Type)
if ruleSetTypeIsIPCIDR(rs.Type) {
@@ -1190,10 +1351,73 @@ func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRul
return option.DefaultHeadlessRule{IPCIDR: badoption.Listable[string](cidrs)}, true
}
// "domain" (and empty, and anything unrecognised => domain, warned above).
// `regexp:` is peeled off first: it is a routing-rule matcher the DNS-filter
// classifier does not know, and it must not be reported as an unknown prefix.
diag := fmt.Sprintf("ruleset %q", rs.Name)
rest, regexes := b.peelDomainRegexes(diag, rs.Entries)
// R4, on the path every destination list now takes. classifyDomainEntries
// DROPS an entry that is nothing but its marker (".", "full:", "keyword:"),
// silently — and the silence is the dangerous half: an empty domain token
// aborts box.New for the whole config, and an empty keyword is
// strings.Contains(host, "") i.e. EVERY host. The drop is right; not saying so
// is not. (devices.go reports the same class for a device's own lists; the bare
// `regexp:` form is reported by peelDomainRegexes above, which is why it is
// peeled off before this loop and cannot be double-reported.)
for _, e := range rest {
if isDomainMarkerOnly(e) {
b.warnf("%s: entry %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword would match EVERY host)", diag, strings.TrimSpace(e))
}
}
// Only the DOMAIN branch reports unknown `word:` prefixes — the ipcidr branch
// above is full of legitimate colons (IPv6) and must never be checked (R9.2).
b.warnUnrecognisedPrefixes(fmt.Sprintf("ruleset %q", rs.Name), rs.Entries)
return inlineDomainRule(rs.Entries)
b.warnUnrecognisedPrefixes(diag, rest)
hr, ok := inlineDomainRule(rest)
if len(regexes) > 0 {
hr.DomainRegex = badoption.Listable[string](regexes)
ok = true
}
return hr, ok
}
// peelDomainRegexes splits `regexp:<pattern>` entries out of a domain rule-set's
// entry list, returning the remaining entries and the validated patterns.
//
// It exists because a destination list is now ALWAYS a rule-set (schema v2), so
// every matcher a `dst_domain` used to express has to be expressible here —
// including the regex form, which the shared DNS-filter classifier
// (classifyDomainEntries) deliberately does not know about. Validation mirrors
// what the routing rule did before the move, and for the same reason:
// route/rule.NewDomainRegexItem returns an error for an uncompilable pattern and
// that aborts box.New for the WHOLE config, so a bad pattern must degrade to a
// warning. A BARE `regexp:` compiles fine but matches every host — the same
// silent match-all hazard as an empty keyword — so it is dropped too.
//
// SCOPE: this is the INLINE entry path only, and deliberately so. A `source=url`
// text list is a hosts/plain-domain FILE, parsed by parseDomainList, which has no
// marker vocabulary at all — see the block above parseDomainList for why the two
// are not unified and how an entry written in the wrong one is reported.
func (b *builder) peelDomainRegexes(diag string, entries []string) (rest, regexes []string) {
for _, e := range entries {
e = strings.TrimSpace(e)
if e == "" {
continue
}
if !strings.HasPrefix(strings.ToLower(e), "regexp:") {
rest = append(rest, e)
continue
}
re := strings.TrimSpace(e[len("regexp:"):])
if re == "" {
b.warnf("%s: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", diag, e)
continue
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("%s: domain regexp %q is invalid (%v), omitted", diag, re, err)
continue
}
regexes = append(regexes, re)
}
return rest, regexes
}
// Ruleset.Type — the two shapes a rule-set can match, and the accepted spellings.
+55 -3
View File
@@ -33,6 +33,14 @@ func findRouteRuleWithRuleSet(rt *option.RouteOptions, tag string) *option.Defau
return nil
}
// hasRulesetRule reports whether any emitted route rule references the rule-set
// named name (tag rs-<name>). Since schema v2 a rule's destination is ALWAYS a
// rule-set reference, so this is how a test says "that rule was emitted" — the
// former "does any rule carry this dst domain" question has no answer any more.
func hasRulesetRule(rt *option.RouteOptions, name string) bool {
return findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name) != nil
}
func ruleSetByTag(rt *option.RouteOptions, tag string) (option.RuleSet, bool) {
if rt == nil {
return option.RuleSet{}, false
@@ -117,6 +125,50 @@ func TestRoutingRuleSetInlineDomain(t *testing.T) {
}
}
// TestRoutingRuleSetInlineDomainRegex: `regexp:` is a matcher the shared domain
// classifier does NOT know — it belongs to the routing plane, and it used to be
// peeled off inside ruleMatchers, which no longer sees any domains at all. It
// therefore had to move into the inline rule-set with the rest of the destination
// vocabulary (peelDomainRegexes), or every migrated `regexp:` entry would have
// been reported as an unknown prefix and silently dropped: a routing rule that
// looks configured and matches nothing.
func TestRoutingRuleSetInlineDomainRegex(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
{Name: "ads", Type: "domain", Source: "inline", Entries: []string{`regexp:^ads\.`}},
},
Rules: []model.Rule{
{Name: "block-ads", Enabled: true, Order: 10, DstRuleset: []string{"ads"}, Target: "block"},
},
}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if len(warns) != 0 {
t.Fatalf("a valid regexp: entry must not warn: %v", warns)
}
rs, ok := ruleSetByTag(opts.Route, "rs-ads")
if !ok {
t.Fatalf("a regexp-only rule-set must still materialise; route=%+v", opts.Route)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^ads\.` {
t.Fatalf("domain_regex = %+v, want [^ads\\.]", hr.DomainRegex)
}
if len(hr.Domain)+len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.IPCIDR) != 0 {
t.Fatalf("the regexp entry must not leak into another matcher: %+v", hr)
}
dr := findRouteRuleWithRuleSet(opts.Route, "rs-ads")
if dr == nil {
t.Fatalf("no route rule references rs-ads; rules=%+v", opts.Route.Rules)
}
if dr.RuleAction.RouteOptions.Outbound != tagBlock {
t.Fatalf("route rule must route to %q, got %+v", tagBlock, dr.RuleAction)
}
}
// TestRoutingRuleSetInlineIPCIDR: an ipcidr ruleset fills ip_cidr (not domain*),
// and the route rule references it.
func TestRoutingRuleSetInlineIPCIDR(t *testing.T) {
@@ -1501,7 +1553,7 @@ func TestURLBlocklistNotRefetchedWhileFresh(t *testing.T) {
// wildcards, IPs and bare labels can never match a domain query, so importing them
// would be a silent dud (the R9 lesson applied to fetched content).
func TestParseDomainListRejectsJunk(t *testing.T) {
got := parseDomainList([]byte(strings.Join([]string{
got, _ := parseDomainList([]byte(strings.Join([]string{
"good.example.com",
"*.wildcard.example", // wildcard syntax
"/regex/", // regex rule
@@ -1532,12 +1584,12 @@ func TestParseDomainListRejectsJunk(t *testing.T) {
// TestParseDomainListHostsEdgeCases covers the messy real-world shapes.
func TestParseDomainListHostsEdgeCases(t *testing.T) {
got := parseDomainList([]byte(stevenBlackSample))
got, _ := parseDomainList([]byte(stevenBlackSample))
if len(got) != 5 {
t.Fatalf("expected 5 domains from the sample, got %d (%v)", len(got), got)
}
// Unicode is punycoded, matching what actually arrives in a DNS query.
uni := parseDomainList([]byte("0.0.0.0 реклама.рф\n"))
uni, _ := parseDomainList([]byte("0.0.0.0 реклама.рф\n"))
if len(uni) != 1 || !strings.HasPrefix(uni[0], "xn--") {
t.Fatalf("a unicode entry must be punycoded, got %v", uni)
}

Some files were not shown because too many files have changed in this diff Show More