Compare commits

...
16 Commits
Author SHA1 Message Date
omarandClaude Opus 5 eccfc6136c fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review of a8ef887c5 + 244b7c419 ("a rule's destination is a rule-set,
and nothing else") found that the change rested on a comment that was not
true. ParseUCIExport dropped dst_domain/dst_ip on the strength of "the
migration is re-run on every load"; model.Migrate() actually runs only from
`shaterd migrate`, i.e. the service init and uci-defaults. The daemon's run
path, the SIGHUP reconcile and the panel's config write never migrate.

So an uncommitted migration (a full /overlay is the documented way that
happens) turned `list dst_domain 'bank.ru'` + `target direct` into a rule
with NO matchers, which IS the spelling of a catch-all: generate points
route.Final at it and the LAST such rule wins. One failed `uci commit` sent
every packet on the router out the plain WAN, silently.

Rule.LegacyDst is the tripwire. It is non-empty exactly when the config
still carries the removed options, and three locks hang off it:
  - ParseUCIExport holds such a rule DISABLED. Chosen over "make IsCatchAll
    false" alone, which only covers matcher-less rules: `dst_domain` plus a
    `src` was never a catch-all, and routing it without its destination
    would still have sent a whole subnet direct.
  - IsCatchAll returns false for it, so it can never own route.Final even
    if something hands its Enabled bit back.
  - ApplyProfileRuleOverrides refuses to enable it (a profile with
    `list enable_rule` would otherwise have defeated the parser).
ValidateRules reports it through the existing warning channel, before the
Enabled gate, so the one message explaining the outage is not suppressed by
the fact that caused it. The init script logs a failed migration to syslog
instead of discarding its exit code and stderr.

The write path had none of this. PUT /api/config decodes a Model straight
from the request body and render.go wrote `enabled` from it, so a panel
save erased the operator's lists (as did the subscription cron, which
re-renders the whole package), and a crafted body with Enabled:true and no
LegacyDst put a live matcher-less rule on disk -- the same whole-router
leak, re-entered from the other side. WriteUCI now reads DISK state and
refuses a rule-changing write over an unmigrated config (409, not 500);
non-rule writers pass and legacyDstOpts carries the options across so cron
preserves them; withDiskLegacyDst takes the field from disk so a fabricated
one can never reach the renderer.

Migration hardening: an entry list that migrates to nothing no longer has
its legacy option deleted (that made "matches nothing" silently become
"matches everything"); a hand-written rule-set whose name collides is no
longer allowed to swallow the entries; delete failures propagate instead of
bumping schema_version past them forever; every error path reverts the
staged uci delta so another process's commit cannot flush a half-migration.

untunnelable.go follows the destination out of the rule: a rule whose
rule-sets are known to match by name is still skipped by the ping/IPTV/VPN
plan, as its v1 form was. D21 documents the AND->OR widening for the
engine's TCP/UDP path; it does not follow that a leak-guard should widen
itself during an upgrade, and with target=direct that meant previously
tunnelled ICMP leaving with the client's real address. Inline rule-sets are
now read from the options, so an engine that has not started yet no longer
costs the operator their ping.

Rule-set vocabulary: `full:`/`suffix:`/`keyword:`/`regexp:` in a text list
fetched by URL were dropped with no diagnostic at all (normaliseListDomain
rejects any token with a colon) -- not "reported as an unknown prefix".
Unifying was rejected: published filter lists are full of colon-bearing
syntax, and a third-party `regexp:` is compiled into the router's matcher
and run per query. The difference stands and is paid for in diagnostics,
per list, on every generate. D21 gains the source/vocabulary table.

Panel: the add form warns about a matcher-less rule exactly as the edit
form does, from one shared predicate; its isCatchAll matches the daemon's
new one; an unmigrated rule reads as held-off rather than merely switched
off. The comment promising a "New list" button that D21 rejected is gone.

go build ./..., go vet ./shater/..., go test ./shater/... (13 packages) and
panel `npm run build` are green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:04:21 +03:00
omarandClaude Opus 5 244b7c4199 feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.

The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.

The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.

isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.

Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 8fd5c52488 fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.

Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.

With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.

Fix (upstream file, lx:tproxy_writeback_connect):
  * CONNECT the write-back socket to the one peer it ever talks to. The kernel's
    compute_score() rejects a connected socket for any other peer, and a
    connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
    reuseport selection outright — so it can no longer be handed a datagram it
    will not read. Nothing about the reply changes: same spoofed source, same
    single peer, Write instead of WriteTo. The unconnected path is kept verbatim
    for a destination that cannot be bound (domain socksaddr).
  * A failed cached write now CLOSES the socket instead of only dropping the
    reference (upstream left the fd to the GC finalizer).
  * TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
    but the cache evicts lazily, so after the inbound is gone nothing wakes the
    live sessions and each strands its write-back socket. Invisible upstream
    (one close at shutdown); on this fork the engine is rebuilt on every apply,
    so it was one stranded generation per apply.

Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.

The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.

Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:41:02 +03:00
omarandClaude Opus 5 32e8f8ff0b fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.

Cause: `restart` is not synchronised end to end.

  * procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
    `stop; start`, and the `service delete` ubus call returns as soon as
    SIGTERM has been SENT. `start_service` therefore re-adds the instance
    (and runs `shaterd migrate`) while the outgoing `shaterd run` is still
    executing its honest teardown.
  * The successor's only defence was `daemonAlive()` -> exit(1), leaning on
    procd's `respawn 3600 5 0` to try again five seconds later. That is a
    blind retry, not synchronisation: it neither knows nor waits for the
    teardown, and it turns every restart into a logged crash plus a
    five-second hole with no data plane.
  * `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
    engine.Close of a several-hundred-outbound box flushes cache.db to
    flash before the netplane teardown even starts — aborting the teardown
    at an arbitrary point and leaving the plane HALF removed.
  * Nothing in the tree ever touched conntrack, so flows that crossed one
    of those windows kept entries formed against a plane that no longer
    exists. For UDP there is no handshake to resynchronise on and every
    retry merely refreshes the entry, so the flow stays wedged for as long
    as the client keeps asking — a flow-scoped, permanent failure that no
    ruleset rebuild can reach.
  * RoutingPresent() reported "plane intact" from the ip RULE alone, while
    ApplyRouting installs a rule AND a `local default dev lo` route removed
    by two independent commands. A teardown interrupted between them was
    therefore invisible, applyLocked's fast-path skipped ApplyRouting
    forever, and no reconcile could repair it.

Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):

  * init: `start_service` waits for a live predecessor pidfile to clear
    before opening the instance, so restart == stop + pause + start. Zero
    cost at boot. term_timeout 10 -> 30 so an honest teardown is never
    killed halfway.
  * daemon: the single-owner guard WAITS for the predecessor (bounded,
    60s) instead of exiting 1; it still refuses if the budget expires.
  * netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
    a blanket flush would drop the admin's own SSH/LuCI sessions) called
    on every plane transition: after a ruleset loads, after the table is
    removed, and once more in applyLocked when the whole plane (table +
    policy routing + sysctls) is assembled.
  * netplane: RoutingPresent() now verifies both halves it installs.

Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:51:19 +03:00
omarandClaude Opus 5 c562579ef3 docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all
unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute
does not emit a condition-less rule as a match-all route rule: it sets
route.Final and continues, so the LAST condition-less rule by order wins, and it
can never shadow a rule that has conditions (those are emitted ahead of Final
regardless of order).

For the config on the router this inverts the conclusion: traffic IS going
through the proxy (order=100 -> group:auto is the live default) and the dead knob
is the order=20 `direct` one. Severity downgraded from high to medium
accordingly — a dead setting, not a leak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:41:58 +03:00
omarandClaude Opus 5 6f89acbae7 feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final",
so a config with two of them showed two identical claims and no hint that
only the last one is the default the router uses.

A superseded rule now loses those marks — it keeps its real Order in the rail
instead of the "·" that means final — and gains a "never applies" badge plus
a line naming the rule that beat it and what to do about it: give this one a
condition, or delete one of the two. Warn semantics throughout (--amber,
dashed frame, dimmed target chip): orange is the ACTIVE state on this
faceplate, and a rule the router ignores is the opposite of active.

Verdicts come from GET /api/rules/reachability and are keyed by the rule's
index in Rules, never by name — the config that prompted this had two rules
both called `default`. They are re-fetched after every save, and a verdict
whose echoed name/order no longer matches the row is dropped rather than
shown, so the window between an optimistic edit and the refetch cannot badge
a working rule.

Rule rows were also keyed by name in React, which silently collapses two rows
that share one; the key now carries the model index.

The mock fixture gains a second condition-less rule so `?mock` renders the
state, and mock.getRulesReachability derives its verdicts from the live
fixture config rather than hard-coding them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 88a82c7297 feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:

  * two condition-less rules retire each other, and the LAST one by Order
    wins, so an earlier "default -> direct" is dead while looking live;
  * a condition-less rule can NEVER retire a rule that HAS conditions —
    those are emitted ahead of Final whatever their Order.

A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.

model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.

Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.

Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.

Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.

GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.

Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 a8f2b0f068 ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.

ci/version.sh is now the single source of truth. It derives the version
from `git describe`:

    tag `vX.Y.Z`   -> PKG_VERSION=X.Y.Z  PKG_RELEASE=1
    off-tag build  -> nearest tag + PKG_RELEASE=<commits since it> + 1
    no tag/no git  -> 0.0.0-r1 (below everything ever published)

Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.

The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.

The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.

byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.

Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:32:38 +03:00
omarandClaude Opus 5 0b32a6d58b fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:

  daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
  [ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...

`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.

Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:

  * control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
    DisableColors at its false zero value. It now comes from
    controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
    lives in an untagged file (same split as profilewatch.go) so it is
    unit-testable off the linux target.
  * engine (shater/generate): the generated option.LogOptions never set
    DisableColor, so box.New built a colouring formatter over the shared
    sink. logOptions() now sets it from the same TTY gate (seam:
    logColorAllowed).

logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.

The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.

Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:31:01 +03:00
omarandClaude Opus 5 893fdc500c fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as 9dc954029 / aec82d444.

`nodes` now reads model.ReadUCI() — `uci export shater` merged with the
per-subscription JSON caches — which is literally the call GET
/api/config serves and generate builds the engine from, so the verb
cannot drift from the panel or from the running engine: there is no
second assembly here to drift. Both ends use the same nodesJSON():
the daemon answers over the control socket (like `stats`), and the CLI
falls back to reading the same on-disk state when no daemon is running
(like `status`). A read failure goes to stderr with a non-zero exit
instead of printing `[]`, so an empty list on stdout now means one thing.

Output is a purpose-built view rather than raw model.Node: the share-link
URI is a credential and CLI output ends up in tickets and cron mail, so
the view reports what the link decodes to (protocol/server/port) plus the
model's own facts (enabled/sub/egress/stale/fingerprint). Nodes whose URI
does not parse are still listed, with the reason in `parse_error` — the
engine skips exactly those, and hiding them would be the same lie smaller.

cmdReadStub keeps `stats`, where the default IS the truth (nothing was
counted without an engine), and now says so in its doc comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:30:27 +03:00
omarandClaude Opus 5 02c266188f docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the
previous install, install from the signed apk feed, verify the default state,
restore a working config with 315 subscription nodes, then exercise the data
plane, panel API, config lifecycle, resilience and DNS.

74 PASS. Findings (detailed separately): two catch-all `default` rules where the
first sends all unspecific traffic direct and makes the second unreachable;
`shaterd nodes` is a stub returning [] while usage promises the node list; DNS to
the router LAN address dies after `service shater restart` (stop+pause+start is
fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI
colour codes reach syslog.

Also records the four-iteration CI hunt that ended in the green apk lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:09:52 +03:00
omarandClaude Opus 5 024e9308c9 fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.

The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:

    config PACKAGE_kmod-mlx5-core
            tristate
            default m

No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.

The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.

Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:27:37 +03:00
omarandClaude Opus 5 eab1db2c3f fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.

They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.

Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
  - an explicit value in .config beats a `default` (same rule that kept our
    `# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
  - `select` is OR-ed in after the user value, so shater-core's
    `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
    transitive kmods) back on their own.

Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.

Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:16:05 +03:00
omarandClaude Opus 5 22d7161c08 fix(ci/apk): disable the SDK's ALL/ALL_KMODS mass-select
release / aarch64_cortex-a53 (push) Successful in 3m25s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m3s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.

The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:

    config ALL_NONSHARED ... default ALL
    config ALL_KMODS     ... default ALL
    config ALL           ... default y

In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.

Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.

Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.

Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:57:19 +03:00
omarandClaude Opus 5 24c5a1615d fix(ci/apk): build .config from scratch — stop packing all 3593 kmods
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m18s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m2s
release / release (push) Successful in 7s
release / release apk (push) Successful in 5s
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was
running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn —
none of which we ship) before it ever got near our four.

Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the
.config that ships inside the ImmortalWrt SDK tarball. That file is the
buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus
CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so
`make defconfig` re-selected every kernel module of the target as =m and
package/kernel/linux/compile — pulled in via shater-core's nft kmod
deps — packed the lot.

Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds":
start the .config EMPTY so kconfig can only pull in what our packages
actually select. Carried over from the SDK's .config, nothing more:
the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a
wrong guess here means silently cross-compiling for another arch),
CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and
CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one
makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel).

Also adds the diagnostics this lane never had, since a failed run leaves
a 27 MB log: the carried-over identity lines, the post-defconfig kmod
count and target readout, a hard check that all four of our packages
survived defconfig, an abort if the kmod count is back in the hundreds,
and du/df after compile.

opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR,
CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:31:28 +03:00
77 changed files with 7293 additions and 663 deletions
+60 -18
View File
@@ -32,6 +32,23 @@
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
# API via curl (ci/gitea-release.sh) — no external action needed.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
# used to be, and nobody bumped them: v0.2.2…v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with different binaries inside, so `apk update` never saw
# a new version and routers could not be updated at all. Now `ci/version.sh`
# derives them from the git tag ONCE per job (the "Compute version" step,
# exported via $GITHUB_ENV):
# tag `vX.Y.Z` -> X.Y.Z-r1
# anything else -> <nearest tag>-r<commits since it + 1>
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
# built .ipk/.apk really carry that version, so the failure can never be
# silent again. This is also why both build jobs check out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
#
# APK LANE (25.12+, ADDITIVE — T2)
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
@@ -118,8 +135,14 @@ jobs:
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
steps:
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the engine via a go.mod
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
@@ -129,6 +152,13 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
- name: Set up Go
uses: actions/setup-go@v5
@@ -197,25 +227,23 @@ jobs:
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
# the SDK package build (the openwrt/shaterd package installs the staged
# artifact). VERSION is stamped into constant.Version. On an exact
# artifact). $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
- name: Build signed feed (SDK)
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
@@ -253,8 +281,12 @@ jobs:
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
steps:
# fetch-depth: 0 — see the opkg lane: the package version comes from
# `git describe`, which needs tags + ancestry.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the AmneziaWG-patched wireguard-go via a
# go.mod `replace => ./submodules/wireguard-go`, so that submodule must be
@@ -263,6 +295,11 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# Same single version computation as the opkg lane — both lanes MUST agree
# on the version, they package the identical tree.
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
- name: Set up Go
uses: actions/setup-go@v5
with:
@@ -346,15 +383,10 @@ jobs:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages as .apk through the ImmortalWrt 25.12 SDK and
# sign the per-arch packages.adb with the EC key (secret KEY_APK).
@@ -462,7 +494,7 @@ jobs:
luci-app-shater (arch=all).
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
── Add as an opkg feed (recommended — then `opkg upgrade` just works) ──
── Add as an opkg feed (recommended — then updating is one command) ──
This release is itself a SIGNED package feed; opkg filters by
architecture, so the same lines work on every device:
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
@@ -472,6 +504,10 @@ jobs:
The public-key install is one-time; after it, `opkg update/upgrade`
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
── Update (name our packages — never a bare `opkg upgrade`) ──
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
── Or install the loose .ipk directly / from the tarball feed ──
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
@@ -531,13 +567,19 @@ jobs:
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
── Add as an apk repository (auto-updates via \`apk upgrade\`) ──
── Add as an apk repository ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
Update: apk update && apk upgrade shaterd shater-core luci-app-shater byedpi
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
A bare \`apk upgrade\` reconciles EVERY installed package against every
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
those packages are upgraded along with needed dependencies\").
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
echo "[release-apk] publishing $TAG from $d"
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
+27 -3
View File
@@ -154,8 +154,14 @@ opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # опционально: ByeDPI desync-egress
```
Обновление: `opkg update && opkg upgrade shaterd shater-core luci-app-shater byedpi`
(обновляйте только эти четыре пакета, не системные).
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
он тянет обновления и на системные пакеты, это классический способ окирпичить
роутер):
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
```
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
@@ -176,7 +182,25 @@ apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # опционально: ByeDPI desync-egress
```
Обновление: `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`.
Обновление — **перечисляйте пакеты явно, голый `apk upgrade` не запускайте**: без
аргументов apk пересобирает состояние ВСЕХ установленных пакетов по ВСЕМ
подключённым репозиториям и может задеть (в т.ч. откатить) посторонние системные
пакеты.
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
# эквивалент, дополнительно закрепляющий пакеты в world:
# apk add -u shaterd shater-core luci-app-shater byedpi
```
Документация apk-tools 3 про `apk upgrade`: *«If list of packages is provided,
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
+12
View File
@@ -55,6 +55,16 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
# standalone. Passed into the container below and re-exported to the
# unprivileged build user in ci/sdk-build-apk.sh.
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) runner-side caches --------------------------------------------------
# All under $REPO/.cache so (a) actions/cache in the workflow can persist them
# between runs and (b) the nested container sees them via --volumes-from.
@@ -94,6 +104,8 @@ docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e SDK_URL="$SDK_URL" \
-e SDK_TAR="$SDK_TAR" -e DL_DIR="$CACHE/dl" -e APT_CACHE="$CACHE/apt" \
-e FEEDS_CACHE="$FEEDS_CACHE" -e KEY_APK="${KEY_APK:-}" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
debian:bookworm bash "$REPO/ci/sdk-build-apk.sh"
# --- 2) sanity: the per-arch apk repo dir must be complete -------------------
+14
View File
@@ -47,6 +47,18 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# The workflow normally puts these in the job env (ci/version.sh --env >>
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
# 0.2.0-r3 and no router could ever see an update).
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
# Workspace dir restored/saved by actions/cache in the workflow and shared into
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
@@ -81,6 +93,8 @@ docker pull "openwrt/sdk:$SDK_TAG"
docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
-e FEEDS_CACHE="$FEEDS_CACHE" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
"openwrt/sdk:$SDK_TAG" \
sh "$REPO/ci/sdk-build.sh"
+244 -1
View File
@@ -28,6 +28,11 @@ SDK_URL="${SDK_URL:?SDK_URL env required}"
echo "[apk-sdk] arch=$ARCH repo=$REPO out=$OUT"
echo "[apk-sdk] sdk=$SDK_URL"
# Package version derived from the git tag by ci/version.sh (bug B4). Forwarded
# to the unprivileged build user on the `su` line at the bottom of this file;
# openwrt/{shaterd,shater-core,luci-app-shater}/Makefile pick it up from the
# environment. byedpi keeps upstream ByeDPI's own version (see its Makefile).
echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -133,6 +138,101 @@ fi
echo "[apk-sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
# --- strip the SDK's generated per-package `default m` blocks ----------------
# Run 60 settled the question that runs 58 and 59 left open. Writing an explicit
# `# CONFIG_PACKAGE_kmod-x is not set` for all 1126 of them and re-running
# defconfig deselected exactly nothing: the count came back 1078, unchanged.
# Meanwhile the very same explicit form DID stick for CONFIG_ALL/ALL_KMODS/
# ALL_NONSHARED. The difference is prompts. kconfig only honours a user value for
# a symbol that has one (sym_calc_value ignores S_DEF_USER for a promptless
# symbol and falls back to its `default`), and the ALL* symbols carry prompts in
# the SDK's own Config.in while these generated blocks are bare:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# So no value we write into .config can ever turn them off — the fix has to
# remove the `default m` itself. That is what this does: drop every generated
# `config PACKAGE_*` block from the SDK's Config-build.in before the first
# defconfig. Nothing is lost by it — these blocks only replay which packages the
# BUILDBOT happened to build; the packages themselves are still declared, with
# prompts, by the package tree (tmp/.config-package.in), which is what makes our
# four selectable and what `select` acts on. KERNEL_*/LIBC/TOOLCHAIN blocks are
# left untouched, so the SDK still reproduces its own toolchain settings.
CB=$(find . -maxdepth 2 -name 'Config-build.in' -print -quit 2>/dev/null || true)
if [ -n "$CB" ] && command -v perl >/dev/null 2>&1; then
pkg_before=$(grep -c '^config PACKAGE_' "$CB" || true)
# Paragraph-wise delete: a block is `config PACKAGE_x`, its indented body, and
# the blank line that ends it. Anchored per-line (/m) so nothing else matches.
perl -0777 -pi -e 's/^config PACKAGE_\S+\n(?:[ \t]+\S[^\n]*\n)+\n//gm' "$CB"
pkg_after=$(grep -c '^config PACKAGE_' "$CB" || true)
echo "[apk-sdk] $CB: stripped $((pkg_before - pkg_after)) generated PACKAGE default blocks ($pkg_before -> $pkg_after)"
else
echo "[apk-sdk] WARNING: no Config-build.in found (or no perl) — per-package"
echo "[apk-sdk] 'default m' blocks stay; the kmod tripwire will catch it"
fi
# --- .config: turn OFF the SDK's mass-select defaults ------------------------
# Symptom (v0.2.2, and still v0.2.3 run 58): the SDK ran `apk mkpkg` on ~1100
# kmod-* packages — mlx5, amdgpu, ata, isdn, none of which we ship — and died
# with `Disk quota exceeded` on the runner's 64 GB ZFS quota. Our kmod deps pull
# in `package/kernel/linux/compile`, which packs every module marked =m.
#
# Why they are =m has nothing to do with anything we write here. An OpenWrt SDK
# carries its OWN top-level Config.in (target/sdk/files/Config.in), and it reads:
#
# config ALL_NONSHARED
# bool "Select all target specific packages by default"
# default ALL
# config ALL_KMODS
# bool "Select all kernel module packages by default"
# default ALL
# config ALL
# bool "Select all userspace packages by default"
# default y <-- y, not n, and ONLY inside the SDK
#
# In the main tree those three default to n; the SDK flips ALL to y so that
# `make world` in a bare SDK builds something useful. So `make defconfig` on ANY
# .config — empty or not — selects the entire kernel. This is stock OpenWrt, not
# an ImmortalWrt quirk: openwrt/openwrt's target/sdk/files/Config.in is identical.
# (It also means the reference we copied, Slava-Shchipunov/awg-openwrt, builds
# every kmod too — it just never hits a disk quota on GitHub's runners.)
#
# Fix: state all three explicitly. They carry prompts in the SDK's Config.in, so
# they are user-settable and an explicit value beats the `default`. Note the FORM:
# kconfig writes a false bool as `# CONFIG_X is not set` and `CONFIG_X=n` is not
# reliably honoured, so `is not set` is the only form used here. All three are set
# rather than just the root `ALL`, so this keeps working whichever symbol a future
# SDK makes the root of the chain.
# Stash anything the SDK shipped (see below — today there is nothing) and start
# from a known-empty file, so what we build here is exactly what we intended.
if [ -s .config ]; then mv -f .config .config.sdk; fi
: > .config
for s in ALL ALL_KMODS ALL_NONSHARED; do
echo "# CONFIG_$s is not set" >> .config
done
# About that stash: an SDK tarball ships NO top-level .config (run 58 logged
# `grep: .config: No such file or directory` — the only `.config` inside the
# tarball is the prebuilt KERNEL's, under the linux dir). This is also why the
# first version of this fix was aimed at the wrong thing: there was never a
# buildbot .config here to append to. Nothing needs carrying over from it either,
# because
# target/sdk/Makefile bakes the buildbot's non-package settings — every
# CONFIG_KERNEL_* included — into the SDK's generated Config-build.in as kconfig
# `default`s (target/sdk/convert-config.pl). defconfig therefore reproduces the
# exact toolchain/kernel settings the SDK was built with, on its own; an earlier
# attempt to copy those lines by hand was redundant and is gone.
# Should a future SDK start shipping a .config, this keeps the two things that
# would then be worth honouring — the target identity and the package format —
# and still lets the lines above override the mass-select.
if [ -s .config.sdk ]; then
echo "[apk-sdk] SDK shipped a .config — carrying over target identity + format:"
grep -E '^CONFIG_TARGET_[a-z0-9_]+=y$|^CONFIG_TARGET_(BOARD|SUBTARGET|ARCH_PACKAGES)=|^CONFIG_USE_APK=' \
.config.sdk | tee -a .config | sed 's/^/[apk-sdk] /' || true
fi
for p in shaterd shater-core byedpi luci-app-shater; do
echo "CONFIG_PACKAGE_$p=m" >> .config
done
@@ -151,11 +251,135 @@ fi
echo "[apk-sdk] defconfig"
make defconfig >/dev/null
# --- second pass: deselect the kernel, keep only what our packages select -----
# Turning ALL/ALL_KMODS/ALL_NONSHARED off (above) provably worked — run 59 shows
# all three as `is not set` after defconfig — and changed the kmod count by
# exactly zero, 1078 both times. The kmods are not selected through ALL_KMODS at
# all. They are selected one by one, and here is where from:
#
# target/sdk/Makefile:
# ./convert-config.pl $(TOPDIR)/.config > $(SDK_BUILD_DIR)/Config-build.in
#
# The SDK's Config-build.in is GENERATED from the buildbot's .config — a config
# in which ALL_KMODS=y had already expanded into a `CONFIG_PACKAGE_kmod-*=m` line
# per module. convert-config.pl turns every `CONFIG_X=<val>` line into a kconfig
# symbol carrying an unconditional `default <val>`; its `next if
# /^(# )?CONFIG_PACKAGE/` filter sits in the `else` branch, which a line with an
# `=` in it never reaches. So the SDK ships, verbatim, 1078 blocks of:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# Nothing there consults ALL_KMODS, which is why switching it off was inert.
#
# Fix: give those symbols an explicit user value. We cannot do it before the
# first defconfig — the list of names only exists once kconfig has expanded the
# tree — so this is a second pass: rewrite every selected kmod to `is not set`
# and re-run defconfig. Two kconfig rules make the result exactly what we want,
# and both are already demonstrated in our own logs:
# * an explicit value in .config beats a `default` (this is precisely why the
# `# CONFIG_ALL* is not set` lines survived defconfig in run 59), so the
# ~1078 kmods we do not need stay off;
# * `select` is a reverse dependency, OR-ed into the symbol's value AFTER the
# user value in sym_calc_value(), so it cannot be overridden by an explicit
# `n`. shater-core's `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` becomes
# `select PACKAGE_kmod-nft-tproxy` (scripts/package-metadata.pl: a `+` flag
# sets `$m = "select"`, and it re-emits the dependency's own depends too, so
# transitive kmods follow). Those come back on their own.
# Net effect: we build the handful of kmods our packages actually pull in.
#
# Rejected alternatives:
# * limiting what `package/kernel/linux/compile` packs — that target has no
# such knob; it iterates the selected set, so the selection IS the knob;
# * `package/kernel/linux/clean` + a targeted build — the kernel package would
# simply be rebuilt in full as a dependency of shater-core, same cost;
# * copying OpenWrt's own feed CI (openwrt/gh-action-sdk) — it does nothing
# about this; it just runs `make defconfig` and builds. Its one disk-related
# setting, CONFIG_AUTOREMOVE=y, is already the SDK's default;
# * editing the SDK's generated Config-build.in to strip the offending blocks —
# it would work, but it means parsing a generated kconfig file by hand and a
# format change would corrupt it silently. The two-pass approach uses only
# kconfig's documented semantics and leaves the evidence in .config.
kmods_all=$(grep -c '^CONFIG_PACKAGE_kmod-[^=]*=[my]$' .config || true)
if [ "$kmods_all" -gt 0 ]; then
echo "[apk-sdk] deselecting $kmods_all kmod packages, then defconfig again"
sed -i -E 's/^CONFIG_(PACKAGE_kmod-[^=]*)=[my]$/# CONFIG_\1 is not set/' .config
make defconfig >/dev/null
fi
# --- post-defconfig sanity + disk-cost readout -------------------------------
# A failed run leaves a ~27 MB log; digging the cause out of it is miserable, so
# print the handful of numbers that decide whether this run survives the
# runner's disk quota BEFORE anything is compiled.
kmods=$(grep -c '^CONFIG_PACKAGE_kmod.*=m' .config || true)
echo "[apk-sdk] target: board=$(sed -n 's/^CONFIG_TARGET_BOARD=//p' .config)" \
"subtarget=$(sed -n 's/^CONFIG_TARGET_SUBTARGET=//p' .config)" \
"arch_packages=$(sed -n 's/^CONFIG_TARGET_ARCH_PACKAGES=//p' .config)"
echo "[apk-sdk] kmod packages selected (=m): $kmods"
# Proof the mass-select stayed off: these three must come back out of defconfig
# as `is not set`. If any reads `=y`, the SDK's `default ALL`/`default y` won and
# the kmod count above will be in the four digits.
echo "[apk-sdk] mass-select symbols after defconfig:"
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/[apk-sdk] /' || true
# After the second pass the only kmods left are the ones shater-core's
# `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` turns into kconfig `select`s, plus
# whatever those select in turn — a handful. Worth printing verbatim while the
# list is short. A count of 0 is NOT fatal: those kmods ship in the router's own
# base feed, so apk resolves them there; but it would mean the selects did not
# fire, and that is something we want to see in the log rather than guess at.
if [ "$kmods" -le 30 ]; then
grep '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/[apk-sdk] /' || true
fi
# The two cache knobs are written before the first defconfig and have to survive
# both of them — losing DOWNLOAD_FOLDER silently costs us the dl/ cache, and
# losing LOCALMIRROR brings back the sourceware.org stalls. Cheap to just look.
echo "[apk-sdk] cache settings after defconfig:"
grep -E '^CONFIG_(LOCALMIRROR|DOWNLOAD_FOLDER)=' .config | sed 's/^/[apk-sdk] /' || true
echo "[apk-sdk] our packages after defconfig:"
grep -E '^CONFIG_PACKAGE_(shaterd|shater-core|byedpi|luci-app-shater)=' .config \
| sed 's/^/[apk-sdk] /' || true
# Each of our 4 must have SURVIVED defconfig. If kconfig dropped one, it is
# because a symbol it `select`s (a DEPENDS entry) does not exist in the installed
# feeds — with the old append-everything .config that was masked by the SDK
# pre-selecting half the distro. `make package/<p>/compile` would then die with a
# cryptic "No rule to make target", far from the real cause.
for p in shaterd shater-core byedpi luci-app-shater; do
grep -q "^CONFIG_PACKAGE_$p=m" .config || {
echo "[apk-sdk] ERROR: $p is NOT selected after defconfig."
echo " kconfig dropped it -> one of its DEPENDS is missing from the"
echo " installed feeds (check the 'feeds install' step above)."; exit 10; }
done
# Only our two nft kmods (+ whatever they themselves depend on) have any business
# being selected here — a dozen at the very most. A count in the hundreds means an
# ALL_KMODS-style mass-select crept back in, and the run would spend ~40 min
# packing the kernel before dying on `Disk quota exceeded`. Fail now instead.
[ "$kmods" -le 200 ] || {
echo "[apk-sdk] ERROR: $kmods kmod packages selected — that is the whole kernel."
echo " Aborting before this fills the runner's disk. Two causes are"
echo " possible, and the lines below tell them apart:"
echo " (a) the mass-select is back on -> a CONFIG_ALL* line reads =y;"
echo " (b) the second pass did not take -> ALL* are 'is not set' but the"
echo " kmods returned anyway, i.e. the per-kmod 'default m' from the"
echo " SDK's generated Config-build.in outlived our explicit 'n'."
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/ /' || true
echo " first few kmods still selected:"
grep -m5 '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/ /' || true
exit 11; }
for p in shaterd shater-core byedpi luci-app-shater; do
echo "[apk-sdk] === build $p ==="
make "package/$p/compile" V=s -j"$(nproc)"
done
# What the build actually cost on disk. The runner's 64 GB ZFS quota is the
# binding constraint on this lane, so record it while the tree still exists.
echo "[apk-sdk] disk usage after compile:"
du -sh build_dir staging_dir bin 2>/dev/null || true
df -h /home/build || true
# A 25.12 apk-SDK must emit .apk — finding only .ipk means a wrong SDK was fed in.
anyapk=$(find bin -type f -name '*.apk' | wc -l)
[ "$anyapk" -gt 0 ] || {
@@ -173,6 +397,25 @@ done
[ "$found" -ge 4 ] || { echo "[apk-sdk] ERROR: expected >=4 of OUR .apk, collected $found"; echo "[apk-sdk] (all .apk under bin/:)"; find bin -type f -name '*.apk' | head -20; exit 6; }
echo "[apk-sdk] collected $found of our .apk"
# --- assert the tag-derived version actually reached the packages -------------
# B4's failure mode is a wrong-but-plausible version shipping silently, so the
# env -> make hand-off is verified, not trusted: each of our three tag-versioned
# packages must be named `<name>-<ver>-r<rel>.apk`. byedpi is excluded on purpose
# (it carries upstream ByeDPI's own version). This runs BEFORE `apk mkndx`, so a
# stale version can never even reach the index.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
[ -f "$OUT/${p}-${want}.apk" ] || {
echo "[apk-sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[apk-sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[apk-sdk] version check OK — our 3 packages are $want"
fi
# --- index + sign: exactly how the OpenWrt 25.12 buildsystem does it ---------
# apk mkndx --root T --keys-dir T [--sign key] --allow-untrusted \
# --output packages.adb *.apk
@@ -204,7 +447,7 @@ INNER
chmod 0644 /home/build/inner.sh
su build -s /bin/bash -c \
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' bash /home/build/inner.sh"
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' SHATER_PKG_VERSION='${SHATER_PKG_VERSION:-}' SHATER_PKG_RELEASE='${SHATER_PKG_RELEASE:-}' bash /home/build/inner.sh"
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[apk-sdk] OK arch=$ARCH — apk feed dir:"
+27
View File
@@ -27,6 +27,13 @@ OUT="${OUT:?OUT env required}"
mkdir -p "$OUT"
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
# Package version, derived from the git tag by ci/version.sh and handed in by
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
# imports every environment variable as a variable, and it propagates through
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
# byedpi deliberately keeps its own upstream version (see its Makefile).
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -111,6 +118,26 @@ for p in shaterd shater-core byedpi luci-app-shater; do
done
done
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
# --- assert the tag-derived version actually reached the packages -------------
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
# env -> make hand-off has several layers (docker -e, make's env import, the
# metadata dump), so verify the result instead of trusting it: every one of our
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
echo "[sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[sdk] version check OK — our 3 packages are $want"
fi
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
ls -l "$OUT"
Executable
+133
View File
@@ -0,0 +1,133 @@
#!/bin/sh
# ci/version.sh — the SINGLE source of truth for "what version is this build?".
#
# WHY THIS EXISTS (bug B4)
# -----------------------
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
# feed's version string differs from the installed one, `apk update` saw nothing
# new and the routers could not be updated through the normal path at all.
#
# So the version is now DERIVED, in CI, from the git tag, and the package
# Makefiles only carry a fallback for manual/offline builds.
#
# THE SCHEME
# ----------
# tag push `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
# any other build -> PKG_VERSION=X.Y.Z of the NEAREST reachable tag,
# (workflow_dispatch, PKG_RELEASE=<commits since that tag> + 1
# rolling `latest`)
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
#
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
# part first (numerically, component by component), the `r<rel>` only as a
# tie-break. Verified against the real tools, not from memory:
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
# identical results (opkg implements the Debian algorithm).
# That is exactly the ordering this scheme needs:
# * a release always outranks every rolling build that preceded it
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
# * rolling builds between two releases grow monotonically (r2 < r10 < r11),
# so a rolling build can never look newer than the next release, and the
# `latest` feed still moves forward on every dispatch.
#
# +1 on the commit count (rather than the raw count) only avoids `-r0` and makes
# a dispatch build of the tagged commit itself identical to the release build of
# that same commit — which is the truth: same tree, same binary.
#
# `byedpi` is deliberately NOT versioned from our tag — see openwrt/byedpi/Makefile.
#
# USAGE
# ci/version.sh # or --env: eval-able / $GITHUB_ENV-able lines
# ci/version.sh --pkg-version # X.Y.Z
# ci/version.sh --pkg-release # R
# ci/version.sh --binary # vX.Y.Z-rR[-g<sha>] for constant.Version
#
# Env:
# SHATER_REF / GITHUB_REF when it is `refs/tags/<tag>` that tag wins and no
# git history is needed (the tag-push path is exact
# even on a shallow checkout).
set -eu
REPO="$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)"
TAG=""
EXACT=0
N=0
SHA=""
# --- 1) an explicit tag ref is authoritative (and needs no git) --------------
REF="${SHATER_REF:-${GITHUB_REF:-}}"
case "$REF" in
refs/tags/*) TAG="${REF#refs/tags/}"; EXACT=1 ;;
esac
# --- 2) otherwise ask git for the nearest reachable release tag --------------
# `--match 'v[0-9]*'` keeps non-release tags (latest, sdk-cache, apk-latest-*,
# musl-toolchain-cache) out. This repo is a sing-box FORK and therefore also
# carries upstream's v1.x tags — `git describe` picks the CLOSEST tag by commit
# distance, so our own v0.2.x (a handful of commits back) always wins over
# upstream's v1.x (thousands of commits back). The tag it picked is logged
# below, so a surprise is visible in the CI log rather than silently shipped.
if [ "$EXACT" -eq 0 ]; then
if D="$(git -C "$REPO" describe --tags --long --match 'v[0-9]*' 2>/dev/null)"; then
# `v0.2.6-1-g02c266188` -> TAG=v0.2.6 N=1 SHA=g02c266188.
# `%` strips the SHORTEST matching suffix, so a tag that itself contains a
# dash (`v0.2.0-healthplan`) survives intact.
TAG="${D%-*-g*}"
REST="${D#"$TAG"-}"
N="${REST%%-*}"
SHA="${REST#*-}"
if [ "$N" -eq 0 ]; then EXACT=1; fi
fi
fi
# --- 3) tag -> numeric PKG_VERSION ------------------------------------------
# Keep the leading dotted-numeric run only: `v0.2.0-healthplan` -> `0.2.0`.
VER=""
if [ -n "$TAG" ]; then
VER="$(printf '%s' "${TAG#v}" | sed -n 's/^\([0-9][0-9.]*\).*/\1/p' | sed 's/\.*$//')"
fi
if [ -z "$VER" ]; then
# No release tag anywhere (shallow clone with no tags, a tarball export, a
# fresh fork). 0.0.0 is BELOW every version we have ever published, so such a
# build can never masquerade as an upgrade on a real router; the commit count
# still makes successive dev builds distinguishable.
VER="0.0.0"
EXACT=0
N="$(git -C "$REPO" rev-list --count HEAD 2>/dev/null || echo 0)"
SHA="$(git -C "$REPO" rev-parse --short HEAD 2>/dev/null || echo '')"
[ -z "$SHA" ] || SHA="g$SHA"
echo "[version] WARNING: no reachable vX.Y.Z tag (and/or no git) -> $VER" >&2
fi
# --- 4) PKG_RELEASE + the string stamped into the binary --------------------
if [ "$EXACT" -eq 1 ]; then
REL=1
FULL="v${VER}-r${REL}"
else
REL=$((N + 1))
FULL="v${VER}-r${REL}${SHA:+-$SHA}"
fi
echo "[version] tag='${TAG:-none}' commits_since=$N exact=$EXACT -> ${VER}-r${REL} (binary: $FULL)" >&2
case "${1:---env}" in
--env|"")
printf 'SHATER_PKG_VERSION=%s\n' "$VER"
printf 'SHATER_PKG_RELEASE=%s\n' "$REL"
printf 'SHATER_VERSION=%s\n' "$FULL"
;;
--pkg-version) printf '%s\n' "$VER" ;;
--pkg-release) printf '%s\n' "$REL" ;;
--binary|--version) printf '%s\n' "$FULL" ;;
*)
echo "usage: $0 [--env|--pkg-version|--pkg-release|--binary]" >&2
exit 2 ;;
esac
+105
View File
@@ -424,3 +424,108 @@ on the next render.
Consequence: all delay numbers are comparable (least ping ranks apples against
apples), and group settings lose two footgun fields while Settings keeps the
two that actually govern every check.
## D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). `config rule` carried THREE ways to say
where traffic is going: `dst_domain` (an inline domain list), `dst_ip` (an inline
CIDR list) and `dst_ruleset` (a reference to a `config ruleset`). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a `.srs`, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
**Decision: `dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is
the only destination matcher.** `Src`, `dst_port` and `proto` are untouched —
they are not lists of destinations and have no rule-set form.
- **Rejected: keep the inline lists as a shorthand.** "One obvious way" is the
whole point; a shorthand that quietly means something different from the long
form (see the bare-entry trap) is worse than no shorthand.
- **Rejected: promote inline lists to rule-sets lazily at generate time.** The
config on disk would then not say what the router does, and the panel would
have to render a list the user cannot find or edit.
### The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (`dst_domain`) | in a rule-set (`entry`) | migrated to |
|--------------------|--------------------------|-------------------------|--------------|
| `example.com` | **exact host** | **host + subdomains** | `full:example.com` |
| `full:example.com` | exact host | exact host | unchanged |
| `suffix:example.com` / `.example.com` | host + subdomains | host + subdomains | unchanged |
| `keyword:ads` | substring | substring | unchanged |
| `regexp:^ads\.` | pattern | pattern *(added here)* | unchanged |
| `geosite:x` / `geoip:x` | inert (engine field removed) | inert (unknown prefix) | unchanged |
`shaterd migrate` (schema v1→v2, `shater/model/migrate.go`) creates one inline
`config ruleset` per rule that still carries a legacy list — `rule-<rule name>`
for domains, `rule-<rule name>-ip` for addresses — moves the entries across with
the conversion above, appends the new name to `dst_ruleset`, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
`rule-<name>-2`). `regexp:` support was added to inline rule-sets in the same
change precisely so the move can be lossless.
`geosite:`/`geoip:` entries are copied VERBATIM rather than promoted to a
`source=geosite` rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
**One deliberate semantic change, called out:** a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because `rule_set: [a, b]` matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
That AND→OR change is about the ENGINE's TCP/UDP path, and it deliberately does
**not** extend to the untunnelable-protocol plane (`shater/apply/untunnelable.go`,
the ping / IPTV / VPN-passthrough policy in nftables). There, a v1
`dst_domain + dst_ip` rule could never claim a packet that carries no domain, so
the plan skipped it; reading the migrated form as OR would have made the address
half suddenly decisive, and with `target=direct` that means an upgrade quietly
sending previously-tunnelled ICMP out with the client's real source address. A
rule whose rule-sets are known to match by NAME is therefore still skipped by that
plan, and the skip is reported ("a routing rule matches by name … as well as by
address"). Split the rule in two if you want the addresses decided there.
### The vocabulary is about ENTRIES YOU TYPE, not about every list body
The table above is the vocabulary of an **inline** rule-set's `entry` values (and
of the DNS-filter/device lists, which share the classifier). The other two rule-set
sources are not other spellings of it:
| source | what it is | vocabulary |
|------------------------------|----------------------------------|------------|
| `inline` | entries you type | the table above |
| `url` → `.srs` / `.json` | a compiled rule-set, engine-owned | the engine's, not ours |
| `url` → anything else | a hosts / one-domain-per-line / AdBlock TEXT FILE | **none** — every line is a domain plus its subdomains |
| `file` | a local `.srs` / `.json` | the engine's, not ours |
**Rejected: run text lists through the entry classifier too.** A published
AdGuard/OISD list is full of colon-bearing tokens that are ordinary filter syntax
(`##…:has(…)`, `$domain=`, absolute URLs); classifying them would either mis-import
them or bury the operator under hundreds of "unrecognised prefix" warnings per
list. The formats also disagree structurally — a hosts line carries several names,
so the text parser works per token, while an entry is a whole line. And `regexp:`
arriving from a third-party URL is a pattern compiled into the router's matcher and
evaluated per query, which is a very different proposition from one the operator
typed.
So the difference stands and is paid for in diagnostics instead: a text list
containing `full:` / `suffix:` / `keyword:` / `regexp:` is reported per list, on
every generate, naming the entries and pointing at `source=inline` where they work
(`warnListEntryVocabulary`, `shater/generate/ruleset.go`). The check tests only
those four markers, never the general `word:` shape, so it fires on a human's
mistake and stays quiet on published filter syntax.
Consequence: one destination mechanism, one vocabulary, one place a list is
edited; every list is compiled once and reused. The panel's rule editor drops its
Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of
the rulesets that already exist, and nothing more. Creating and filling a list
stays in the Rulesets panel — **rejected: a "create a list from here" shortcut in
the rule editor**, because a second place to author a list is a second place for
its semantics and its duplicate-name rules to drift, and the whole point of this
decision was to stop having two.
+6 -2
View File
@@ -13,8 +13,12 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination (domain/suffix/keyword/geosite), reusable domain/IP lists, port,
proto → target (outbound/selector/chain/direct/block) + egress.
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
(inline domains/CIDRs, a local or remote file, or a geosite/geoip category) that is
compiled once into a `.srs` and shared by every rule that references it. Domain
entries take `full:` (exact), `suffix:` / a leading dot (host + subdomains),
`keyword:` (substring) and `regexp:`; a bare entry means host + subdomains.
- **[MVP]** Node groups with balancer/observatory (least-ping/failover/round-robin).
- **[T1]** Multi-hop chains (L1→Ln); per-rule egress selection; egress via any
interface/tunnel (e.g. an AmneziaWG tunnel).
+81 -15
View File
@@ -26,7 +26,9 @@ What it does:
Arg / env:
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
`$SHATER_VERSION` → `git describe --tags` → `v0.2.0-dev`.
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
the **same** computation the package version comes from (§2.1), so the string
the panel shows always matches what `apk info shaterd` / `opkg status` report.
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
@@ -84,15 +86,52 @@ it. Because the binary is UPX-packed, the package disables the SDK's default str
feed installed and run `make package/shaterd/compile` (and the others) per target.
See `openwrt-package-build-ci` for SDK/feed mechanics.
### 2.1 Package versions come from the git tag
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
Both package managers offer an upgrade only when the feed's version string
differs from the installed one, so `apk update` saw nothing new and the routers
could not be updated through the normal path at all.
`ci/version.sh` now derives them from `git describe`, once per CI job:
| Build | `PKG_VERSION` | `PKG_RELEASE` | `constant.Version` |
|---|---|---|---|
| tag push `v0.2.7` | `0.2.7` | `1` | `v0.2.7-r1` |
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
Ordering is what makes this safe, and both managers agree on it (checked with
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
outranks every rolling build before it, rolling builds between two releases grow
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
upgrade.
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
environment; the Makefiles read it with a literal fallback for manual/offline
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
so a lost variable fails the build instead of shipping a stale version.
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
installed. Stamping our tag on it would also be a downgrade: every comparator
reads `0.2.7 < 0.17.3` (component-wise, `2 < 17`). Bump its `PKG_RELEASE` by hand
when our packaging of it changes.
## 3. Install on a router
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
```sh
opkg install shaterd_0.2.0-1_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_0.2.0-1_all.ipk
opkg install luci-app-shater_0.2.0-1_all.ipk
opkg install byedpi_0.17.3-1_<arch>.ipk # optional: ByeDPI egress
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_<ver>_all.ipk
opkg install luci-app-shater_<ver>_all.ipk
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
```
Installing from a signed feed instead:
@@ -158,16 +197,25 @@ every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
### Updating
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
tries to upgrade *every* installed package from *every* configured feed, which on
OpenWrt means base/system packages on the overlay and is a well-known way to
brick a router.
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
```
Updates are only offered when the feed's `Version` differs from the installed one,
so **bump `PKG_RELEASE`** (or `PKG_VERSION`) in the package Makefile on every
shipped change — otherwise `opkg upgrade` sees the same version and does nothing.
Do **not** `opkg upgrade` base/system packages from this feed; upgrade only the
four shater packages above.
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
when the feed's `Version` differs from the installed one — that is exactly what
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
the version from the git tag on every build (§2.1), so there is nothing to bump
by hand any more; check with:
```sh
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
```
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
@@ -214,15 +262,33 @@ apk add byedpi # optional: ByeDPI desync egress
### Updating
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
installed package against *every* configured repository at once; on a router
whose distfeeds point at a moving snapshot that can pull in — or roll back —
unrelated system packages. Always name ours:
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
apk upgrade shaterd shater-core luci-app-shater byedpi
```
Same rule as opkg: an upgrade is only offered when the feed version differs, so
bump `PKG_RELEASE`/`PKG_VERSION` on every shipped change (apk shows it as
`0.2.0-r1`). Pin a version instead of tracking rolling by pointing the repo line
at `.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
apk-tools 3 documents exactly this behaviour for `apk upgrade`: *"When no
packages are specified, all packages are upgraded if possible. If list of
packages is provided, only those packages are upgraded along with needed
dependencies."* The equivalent form, which additionally re-pins the packages in
`world`, is:
```sh
apk add -u shaterd shater-core luci-app-shater byedpi # -u = --upgrade
```
Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
rolling by pointing the repo line at
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
### BananaWRT `25.12-mtk-vendor` compatibility
+7 -3
View File
@@ -195,7 +195,7 @@ type Chain struct { Name string; Hops []string } // "group:<n>" | "node:<n>", L1
type Egress struct { Name,Type,Interface,Target string } // interface|proxy|direct|block
type Rule struct {
Name string; Enabled bool; Order int
Src []string; DstDomain,DstRuleset,DstIP []string; DstPort,Proto string
Src []string; DstRuleset []string; DstPort,Proto string // dst = ruleset only (v0.2 schema v2)
Target string // chain:|group:|node:|direct|block
Egress,Kill string
SchedEnabled bool; SchedDays []string; SchedStart,SchedEnd string; SchedUTCOffset int
@@ -257,8 +257,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
- `config chain`: name, list hop. `config egress`: name, type, interface, target.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url), url, path, format, update_interval, list entry.
- `config rule`: name, enabled, order, list src/dst_domain/dst_ruleset/dst_ip, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end/tz.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url|geosite|geoip), url, path, format, update_interval, list category, list entry.
- `config rule`: name, enabled, order, list src, list dst_ruleset, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end, sched_utc_offset.
v0.1 carried `dst_domain`/`dst_ip` on the rule itself; **schema v2 removed both** — a
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
@@ -0,0 +1,170 @@
# Живое тестирование shater v0.2.6 на mini_router
**Дата:** 2026-07-25
**Устройство:** Bananapi BPi-R3 Mini · ImmortalWrt **25.12-linkup** · `aarch64_cortex-a53`
**Установка:** из подписанного apk-фида `apk-v0.2.6-aarch64_cortex-a53`
**Пакеты:** `shaterd 0.2.0-r3`, `shater-core 0.2.0-r3`, `luci-app-shater 0.2.0-r2`, `byedpi 0.17.3-r1`
**Сборка:** CI run 61, коммит `024e9308c` (вершина `main`)
Сценарий: полное удаление предыдущей установки → чистая установка из фида →
проверка дефолтного состояния → восстановление рабочего конфига с подписками
(315 узлов) → функциональная проверка.
**Итог: 79 проверок, 74 PASS, 5 находок** (детали и разбор — в
`shater-bugs-2026-07-25.md` на рабочем столе).
---
## 1. Релиз и фид
| # | Проверка | Результат |
|---|---|---|
| T1 | Публикация `apk-v0.2.6-<arch>` для обеих архитектур | PASS |
| T2 | Ассеты: 4 `.apk` + `packages.adb` + `shater-apk.pem` | PASS |
| T3 | `apk update` принимает индекс (проверка EC-подписи) | PASS |
| T4 | Пакеты видны в нужных версиях (r3/r3/r2) | PASS |
| T5 | Диагностика сборки: `kmod packages selected (=m): 0` (было 1078) | PASS |
| T6 | Собраны ровно наши 4 пакета | PASS |
| T7 | opkg-лейн v0.2.6 (24.10) тоже зелёный | PASS |
## 2. Установка
| # | Проверка | Результат |
|---|---|---|
| T8 | `apk add luci-app-shater byedpi` — 4 пакета | PASS |
| T9 | Зависимости `kmod-nft-tproxy`/`kmod-nft-socket` из базового фида | PASS |
| T10 | Целостность: `apk manifest` = sha256 файла на диске | PASS |
| T11 | Установлен именно бинарь v0.2.6 (5 491 616 Б vs 5 488 336 Б в r2) | PASS |
| T12 | init-скрипты `shater`, `shater-cron` | PASS |
| T13 | `sysctl.d/99-shater.conf`, `hotplug.d/iface/99-shater` | PASS |
| T14 | boot-линки `S99shater`, `K10shater`, `S96shater-cron` | PASS |
## 3. Дефолтное состояние (чистая установка)
| # | Проверка | Результат |
|---|---|---|
| T15 | Дефолтный конфиг создан uci-defaults (27 строк) | PASS |
| T16 | `enabled='0'` — плоскость не ставится без согласия | PASS |
| T17 | Заготовлен tproxy-inbound на LAN, пресеты выключены | PASS |
| T18 | Демон стартует, `plane=none`, `table=false` | PASS |
| T19 | Права конфига `-rw-------` (0600) | PASS |
## 4. Восстановление рабочего конфига
| # | Проверка | Результат |
|---|---|---|
| T20 | Восстановление из бэкапа (3213 UCI-строк) | PASS |
| T21 | Кэш подписок цел: 315 узлов в 4 файлах | PASS |
| T22 | `shaterd migrate` → `ok`, схема v1 | PASS |
| T23 | Старт с реальным конфигом: `active`, `engine_running`, `plane=full` | PASS |
## 5. Data plane
| # | Проверка | Результат |
|---|---|---|
| T24 | Таблица `inet shater` создана (9 цепочек/сетов) | PASS |
| T25 | 16 tproxy-правил | PASS |
| T26 | `ip rule from all fwmark 0x2000 lookup shater` | PASS |
| T27 | `accept_local=1` на `br-lan` | PASS |
| T28 | DNS-divert: `dport 53 → tproxy :12345` для LAN-интерфейсов | PASS |
| T29 | DoT заблокирован: `dport 853 reject` | PASS |
| T30 | `block_doh=1`, правила присутствуют | PASS |
| T31 | **Kill-switch fail-closed**: цепочка `forward` завершается `drop` для LAN (v4+v6) | PASS |
| T32 | fw4 и dnsmasq не тронуты (свои таблицы целы) | PASS |
## 6. Панель и API
| # | Проверка | Результат |
|---|---|---|
| T33 | SPA отдаётся на `:8088` | PASS |
| T34 | `shaterd mint-token` выдаёт одноразовый токен | PASS |
| T35 | `/api/status` без сессии → **401** | PASS |
| T36 | `/api/session` (POST, JSON) → 200 + cookie `HttpOnly; SameSite=Strict; Max-Age=28800` | PASS |
| T37 | `/api/status` по cookie отдаёт данные, совпадающие с CLI | PASS |
| T38 | `/api/config` — 340 записей узлов | PASS |
| T39 | `/api/groups/health` — 103 протестировано, 13 живых, выбран `FR-vless-8` | PASS |
| T40 | `/api/devices` — устройства с IPv4/IPv6/MAC | PASS |
| T41 | `/api/interfaces` — `ewan/eth1 10.0.0.125/24 zone=wan` | PASS |
| T42 | `/api/ruleset/status` — remote-ruleset обновлён сегодня | PASS |
| T43 | `/api/stats` — memory backend, счётчики и top-domains | PASS |
| T44 | `/api/stats/log` — query-log с доменом, qtype, rcode, сервером | PASS |
| T45 | `/api/log?range=100` — пусто (следствие `log_file='0'`, не дефект) | OK |
## 7. Жизненный цикл конфигурации
| # | Проверка | Результат |
|---|---|---|
| T46 | `shaterd apply` → `{"changed":false}`, `can_rollback=true` | PASS |
| T47 | `shaterd confirm` снимает авто-откат (`can_rollback=false`) | PASS |
| T48 | `shaterd rollback` после confirm корректно сообщает об отсутствии last-good | PASS |
| T49 | `shaterd reconcile` (SIGHUP) не роняет движок | PASS |
| T50 | `shaterd sub update all-qomar` — реально обновил 143 узла | PASS |
| T51 | `shaterd blocklist update` → reconcile signalled | PASS |
| T52 | `shaterd schedule due` → reconcile signalled | PASS |
## 8. Устойчивость
| # | Проверка | Результат |
|---|---|---|
| T53 | `kill -9` демона → procd поднимает новый PID | PASS |
| T54 | После respawn: `engine_running=true`, `plane=full` | PASS |
| T55 | `stop` снимает таблицу `inet shater` полностью | PASS |
| T56 | `stop` → пауза → `start`: плоскость восстанавливается | PASS |
| T57 | Сеть при остановленном shater не деградирует | PASS |
| T58 | Память: 253 МБ занято из 2 ГБ при работающем движке | PASS |
## 9. DNS
| # | Проверка | Результат |
|---|---|---|
| T59 | Резолв через `127.0.0.1` | PASS |
| T60 | LAN-клиенты резолвят через движок (query-log растёт) | PASS |
| T61 | `.lan`-домены остаются за dnsmasq | PASS |
| T62 | dnsmasq жив и слушает на всех адресах | PASS |
| T63 | **Резолв через LAN-адрес `10.67.0.1` после `restart`** | **FAIL — B3** |
| T64 | Тот же резолв после `stop` → пауза → `start` | PASS |
## 10. Конфигурация и логи
| # | Проверка | Результат |
|---|---|---|
| T65 | 5 правил маршрутизации, 2 профиля, активен `ethernet-uplink` | PASS |
| T66 | **Два правила `default`, оба catch-all — нижнее живое, верхнее мертво** | **FAIL — B1** |
| T67 | **`shaterd nodes` всегда возвращает `[]`** | **FAIL — B2** |
| T68 | Логи уходят в syslog (`log_syslog=1`, 22 записи) | PASS |
| T69 | **ANSI-escape коды в syslog** | **FAIL — B5** |
| T70 | `loglevel=warning` соблюдается | PASS |
| T71–T79 | Прочие проверки состояния (статус-поля, права, uptime, счётчики, целостность таблиц) | PASS |
---
## Находки
| ID | Суть | Важность |
|---|---|---|
| **B1** | Два catch-all правила `default`; одно из них не работает никогда. **Поправка к первоначальному диагнозу:** правило без условий задаёт `route.Final`, а не выпускается как match-all, поэтому выигрывает ПОСЛЕДНЕЕ (`order=100 → group:auto`) — трафик идёт через прокси, а мёртвая настройка это `order=20 → direct` | средняя |
| **B2** | `shaterd nodes` — заглушка, всегда `[]`, хотя usage обещает список узлов (в кэше 315, в `/api/config` 340) | средняя |
| **B3** | После `service shater restart` резолв к LAN-адресу роутера не работает и не восстанавливается; `stop`+пауза+`start` — работает (гонка) | средняя |
| **B4** | `PKG_RELEASE` не менялся с v0.2.1 → v0.2.2…v0.2.6 выходят как `r3` при разном содержимом; `apk upgrade` не увидит обновления | средняя |
| **B5** | ANSI-раскраска попадает в syslog | низкая |
Разбор с воспроизведением — в `shater-bugs-2026-07-25.md`.
## История CI по этому релизу
Путь до зелёной сборки apk-лейна занял четыре итерации, каждая вскрывала
следующий слой одной причины:
| Тег | Что чинили | Итог |
|---|---|---|
| v0.2.2 | — (первый прогон с фиксами аудита) | `Disk quota exceeded`, 3593 `apk mkpkg kmod-*` |
| v0.2.3 | `.config` строится с нуля, а не дописывается | 1078 kmod — SDK вообще не везёт `.config` |
| v0.2.4 | Выключены `ALL`/`ALL_KMODS`/`ALL_NONSHARED` | 1078 kmod — они выбираются не через `ALL_KMODS` |
| v0.2.5 | Второй проход: явное `is not set` для каждого kmod | 1078 kmod — kconfig игнорирует user-значение у беспромптовых символов |
| **v0.2.6** | Удаление сгенерированных блоков `config PACKAGE_*` (`default m`) из `Config-build.in` | **0 kmod, сборка зелёная** |
Корень: `target/sdk/Makefile` генерирует `Config-build.in` прогоном
`convert-config.pl` по конфигу бильдбота, где `ALL_KMODS=y` уже развернулся в
`CONFIG_PACKAGE_kmod-*=m` на каждый модуль. Фильтр `next if /^(# )?CONFIG_PACKAGE/`
в скрипте стоит в ветке `else`, куда строка со знаком `=` не попадает, поэтому
каждый kmod приезжает в SDK как безусловный `default m`.
+11
View File
@@ -15,6 +15,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=byedpi
# DELIBERATELY NOT auto-versioned from our git tag (unlike shaterd/shater-core/
# luci-app-shater, which take SHATER_PKG_VERSION/SHATER_PKG_RELEASE from
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
PKG_VERSION:=0.17.3
PKG_RELEASE:=1
+7 -2
View File
@@ -24,8 +24,13 @@ LUCI_TITLE:=LuCI thin launcher for Shater (mini dashboard + panel handoff)
LUCI_DEPENDS:=+shater-core +rpcd
LUCI_PKGARCH:=all
PKG_VERSION:=0.2.0
PKG_RELEASE:=2
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4). The literals are the manual/offline fallback only.
# These MUST stay above the luci.mk include: luci.mk only defaults PKG_VERSION/
# PKG_RELEASE when they are still unset, and the i18n subpackages inherit them.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+7 -2
View File
@@ -13,8 +13,13 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shater-core
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4: v0.2.2…v0.2.6 all shipped as 0.2.0-r3). The literals
# are the manual/offline fallback only.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-2.0-or-later
+29 -2
View File
@@ -62,7 +62,29 @@ config inbound
# list node 'my-node'
#
# A routing rule. target: chain:<n>|group:<n>|node:<n>|egress:<n>|direct|block.
# Match on src / dst_domain / dst_ruleset / dst_ip / dst_port / proto.
# Match on src / dst_ruleset / dst_port / proto. A rule with NO matcher at all is
# the default route for everything that reached it.
#
# WHERE the traffic is going is named ONLY by dst_ruleset — one or more
# `config ruleset` names; the rule matches when ANY of them matches. There is no
# inline domain or address list on a rule (`dst_domain`/`dst_ip` were removed in
# schema v2): a destination list is written once as a ruleset, compiled into a
# .srs and shared by every rule that references it. `shaterd migrate` converts
# older configs automatically, creating a `rule-<name>` ruleset per rule.
#config ruleset
# option name 'blocked-video'
# option type 'domain'
# option source 'inline'
# list entry 'youtube.com'
# list entry 'suffix:googlevideo.com'
#
#config rule
# option name 'video-via-main'
# option enabled '1'
# option order '50'
# list dst_ruleset 'blocked-video'
# option target 'group:main'
#
#config rule
# option name 'all-via-main'
# option enabled '1'
@@ -83,11 +105,16 @@ config inbound
# option type 'direct'
# option dpi 'fragment'
#
#config ruleset
# option name 'youtube'
# option source 'geosite'
# list category 'youtube'
#
#config rule
# option name 'youtube-fragment'
# option enabled '1'
# option order '50'
# list dst_domain 'geosite:youtube'
# list dst_ruleset 'youtube'
# option target 'egress:frag'
#
# A DNS resolver (type: doh|dot|plain|local|fakeip). `detour` routes its queries
+94 -6
View File
@@ -47,6 +47,13 @@ PROG=/usr/bin/shaterd
# hotplug/shater-cron touch the data plane. tmpfs => cleared by reboot, so
# nothing reconciles before this init has run at boot.
ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# --- helpers ---------------------------------------------------------------
@@ -66,6 +73,54 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
# data plane has not been dismantled yet".
shater_daemon_pid() {
local pid
pid=$(cat "$PIDFILE" 2>/dev/null) || return 1
[ -n "$pid" ] || return 1
kill -0 "$pid" 2>/dev/null || return 1
echo "$pid"
}
# Block until no predecessor daemon is left, bounded by STOP_WAIT_SECS.
#
# WHY THIS EXISTS. procd's `stop` is ASYNCHRONOUS: rc.common's `restart` is
# literally `stop; start`, and the `service delete` ubus call returns the moment
# procd has SENT SIGTERM — not when the instance is gone. `start` therefore
# re-adds the instance while the outgoing `shaterd run` is still executing its
# honest teardown (engine close, then `nft delete table`, `ip rule`/`ip route`
# removal and the per-iface sysctl restore). The result is that `restart` is NOT
# equivalent to `stop` + pause + `start`: the new plane is stood up on top of
# kernel state the old one has not finished removing, which is what B3 (DNS to
# the router's own LAN address dead after a restart, and never recovering) came
# out of. Waiting here restores the equivalence, and costs literally nothing when
# there is no predecessor — the check runs before the first sleep.
#
# Returning non-zero does NOT abort the start: the daemon carries its own
# single-owner guard and will refuse (or wait) on its side. Better to hand the
# decision to the process that can actually see the plane than to leave the box
# with no service at all.
shater_wait_stopped() {
local i=0 pid
pid=$(shater_daemon_pid) || return 0
_slog -p daemon.info \
"restart: waiting for the previous shaterd (pid $pid) to finish tearing the data plane down"
while [ "$i" -lt "$STOP_WAIT_SECS" ]; do
sleep 1
i=$((i + 1))
shater_daemon_pid >/dev/null || {
_slog -p daemon.info "restart: previous shaterd exited after ${i}s; starting a fresh one"
return 0
}
done
_slog -p daemon.warn \
"restart: previous shaterd (pid $pid) still alive after ${STOP_WAIT_SECS}s — starting anyway"
return 1
}
# --- procd lifecycle -------------------------------------------------------
start_service() {
@@ -87,9 +142,31 @@ start_service() {
return 0
fi
# Do not stand a new data plane up on top of one that is still being taken
# down. On `restart` procd has only just SIGTERMed the previous instance and
# returned; this is the handshake that makes `restart` == `stop` + pause +
# `start`. It also keeps `migrate` below from rewriting UCI underneath a
# daemon that is still reading it. No-op (and no delay) when nothing is
# running, which is the boot case.
shater_wait_stopped
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
"$PROG" migrate >/dev/null 2>&1
#
# THE FAILURE IS LOGGED, NOT SWALLOWED. This is the only place the schema
# migration runs at boot (`shaterd run`, the SIGHUP reconcile and the panel's
# config write all read UCI directly), so if it fails here it does not get
# retried until the next start. And it CAN fail for a mundane reason — a full
# /overlay makes `uci commit` fail — after which the config still carries the
# schema-v1 `dst_domain`/`dst_ip` options. The daemon holds every rule that
# still has them DISABLED and reports it, so nothing is silently misrouted, but
# rules the operator wrote are then not in force and the reason has to be
# visible somewhere. Hence: log the binary's own stderr, and start anyway —
# refusing to start would take the admin panel down with it, and the panel is
# the only way to fix the box.
local migrate_out
migrate_out=$("$PROG" migrate 2>&1) || _slog -p daemon.err \
"UCI schema migration FAILED: ${migrate_out:-no output from $PROG migrate}. Starting anyway; routing rules that still carry the removed dst_domain/dst_ip options stay DISABLED until this succeeds. Free space on /overlay and re-run '$PROG migrate', or restart the service."
procd_open_instance shater
# shaterd runs in the FOREGROUND under procd (must never daemonize). `run` is
@@ -111,7 +188,16 @@ start_service() {
procd_set_param stderr 1
# Give the daemon room to run its honest teardown (engine.Close + netplane
# restore) before procd SIGKILLs it.
procd_set_param term_timeout 10
#
# 30s, not 10s: an engine holding a few hundred outbounds closes its
# urltest/observatory goroutines and flushes experimental.cache_file to FLASH
# before the netplane teardown even starts, and on eMMC/NAND that alone can
# outlast 10s. A SIGKILL there aborts the teardown at an arbitrary point and
# leaves the plane HALF removed — the nft table gone but the policy routing
# still installed, or vice versa — which is precisely the class of leftover
# state the successor's idempotent fast-path cannot see and never repairs.
# Shutdown is bounded by procd either way; we are only choosing where.
procd_set_param term_timeout 30
procd_close_instance
# Mark the stack live for hotplug/cron — but ONLY when interception is
@@ -141,10 +227,12 @@ stop_service() {
reload_service() {
# Fired by the `shater` config.change reload-trigger (LuCI Save & Apply /
# reload_config). Simplest correct behaviour: stop + start. `stop` clears the
# flag and SIGTERMs the daemon (honest teardown); `start` re-guards on
# enabled and, if still enabled, launches a fresh `shaterd run` that reads
# the new UCI and applies it. When the stack is disabled, `start` is a no-op,
# so a disable+apply cleanly tears everything down.
# flag and SIGTERMs the daemon (honest teardown); `start` WAITS for that
# teardown to actually finish (shater_wait_stopped) and then launches a fresh
# `shaterd run` that reads the new UCI and applies it. When the stack is
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
stop
start
}
+11 -2
View File
@@ -34,8 +34,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shaterd
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
# that version, so a lost env can never silently ship a stale one again.
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
# are not "the release version"; releases are named by the tag.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+54 -2
View File
@@ -809,9 +809,17 @@ export interface Rule {
Enabled: boolean
Order: number
Src?: string[] | null
DstDomain?: string[] | null
/**
* WHERE the traffic is going — the rule's only destination matcher. Each entry
* names a {@link Ruleset}; the rule matches when ANY of them matches.
*
* There is no inline domain or address list on a rule. `dst_domain`/`dst_ip`
* were removed in schema v2, and `shaterd migrate` folds every existing one
* into a generated `rule-<name>` ruleset, so a destination list is written and
* edited in exactly one place and compiled once into a .srs that every rule
* referencing it shares.
*/
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
/**
* Narrow the rule to one transport or one sniffed application protocol. A
@@ -1187,6 +1195,50 @@ export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEn
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
* One routing rule's reachability verdict — the rule analogue of
* {@link ChainHealth}.used: a quiet note about the ROUTING CONFIG, never a health
* signal.
*
* `unreachable` means the rule can NEVER take effect, whatever the traffic. Today
* the daemon reports exactly one certain case, and it is a subtle one: a rule with
* no conditions at all is not matched in sequence — it becomes the router's
* default. Two such rules therefore retire each other, and the LAST one by Order
* wins, so an earlier "default → direct" is dead even though it sorts first. A
* condition-less rule never retires a rule that HAS conditions: those are matched
* ahead of the default whatever their Order.
*
* `index` is the rule's position in GET /api/config's `Rules`, which is how a
* verdict is matched to a row — rule names are not unique, and the config that
* prompted this had two rules both called `default`. `name`/`order` are echoed so
* a page holding a verdict fetched before an edit can check it still describes the
* row it is about to badge, and drop it silently otherwise.
*/
export interface RuleReach {
index: number
name: string
order: number
unreachable: boolean
/** The rule that supersedes this one; absent when `unreachable` is false. */
shadowed_by?: string
/** Its index in `Rules`, or -1 when there is none. */
shadowed_by_index: number
shadowed_by_order?: number
/** Operator-facing sentence; absent when `unreachable` is false. */
reason?: string
}
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
* the same order as GET /api/config's `Rules`. */
export interface RulesReachability {
rules: RuleReach[]
}
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
+49 -1
View File
@@ -6,7 +6,7 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -146,6 +146,12 @@ const CONFIG: Model = {
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
// A SECOND condition-less rule, above the real default. It reads like a working
// rule and does nothing: a rule with no conditions becomes the router's default,
// and the last such rule by Order wins — so this one never applies. It is in the
// fixture on purpose, to exercise the "never applies" badge; the field config
// that prompted it had two rules BOTH named `default` (orders 20 and 100).
{ Name: 'default-bypass', Enabled: true, Order: 40, Target: 'direct' },
{ Name: 'default-tunnel', Enabled: true, Order: 900, Target: 'group:auto' },
],
// Named match-lists a rule points DstRuleset at. url + geosite + geoip are remote
@@ -296,6 +302,48 @@ const RULESET_STATUS: RulesetStatus[] = [
{ tag: 'rs-ru-geoip-ru', name: 'ru-geoip', category: 'ru', kind: 'ruleset', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
]
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
* a rule with no conditions is the router's default, and the LAST such rule by
* Order wins — every earlier one can never apply. It reads the live CONFIG so
* edits made in `?mock` keep the badge honest. */
export async function getRulesReachability(): Promise<RulesReachability> {
await wait(60)
const rules = CONFIG.Rules ?? []
const out: RuleReach[] = rules.map((r, index) => ({
index,
name: String(r.Name ?? ''),
order: Number(r.Order ?? 0),
unreachable: false,
shadowed_by_index: -1,
}))
const conditionless = (r: (typeof rules)[number]): boolean =>
!(r.Src ?? []).length &&
!(r.DstRuleset ?? []).length &&
!String(r.DstPort ?? '').trim() &&
!String(r.Proto ?? '').trim()
const target = (r: (typeof rules)[number]): string =>
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
const defaults = rules
.map((r, index) => ({ r, index }))
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
const winner = defaults[defaults.length - 1]
if (winner) {
for (const { index } of defaults.slice(0, -1)) {
out[index].unreachable = true
out[index].shadowed_by = String(winner.r.Name ?? '')
out[index].shadowed_by_index = winner.index
out[index].shadowed_by_order = Number(winner.r.Order ?? 0)
out[index].reason =
`this rule has no conditions, so it sets the default for all traffic — but rule ` +
`"${winner.r.Name}" (order ${winner.r.Order}) has none either and comes after it, so ` +
`"${target(winner.r)}" is the default the router uses and this rule's target ` +
`"${target(rules[index])}" is never applied`
}
}
return { rules: out }
}
export async function getRulesetStatus(): Promise<RulesetStatus[]> {
await wait(90)
return RULESET_STATUS.map((r) => ({ ...r }))
+39 -6
View File
@@ -252,6 +252,43 @@
color: var(--faint);
}
/* ---- a rule that can never fire (superseded by a later condition-less rule) ----
*
* Warn semantics only: --amber, never --accent. Orange is the ACTIVE state on this
* faceplate, and a rule the router ignores is the opposite of active — painting it
* orange is what made two `default` rows look equally live. It is a dashed amber
* frame, an amber order chip, and a dimmed target, so the row reads as "wired but
* not connected" without shouting: nothing is broken, one setting is just inert. */
.rt-rule.dead {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
background: var(--panel);
box-shadow: none;
}
.rt-ord.dead {
color: var(--amber);
border-color: color-mix(in srgb, var(--amber) 45%, var(--groove));
}
.rt-badge.dead {
padding: 1px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
color: var(--amber);
}
.rt-dead-note {
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.45;
color: var(--dim);
}
/* The target is still what the operator asked for, so it stays readable — just
* quiet, because the router is not using it. */
.rt-rule.dead .rt-target {
border-style: dashed;
opacity: 0.62;
}
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
.rt-target {
display: inline-flex;
@@ -359,9 +396,6 @@
gap: 5px;
min-width: 0;
}
.rt-field-wide {
grid-column: span 2;
}
.rt-flabel {
font-family: var(--font-mono);
font-size: 9px;
@@ -478,6 +512,8 @@ select.rt-input {
border-color: var(--accent);
box-shadow: 0 1px 0 var(--edge) inset, 0 0 0 1px var(--accent-soft);
}
/* "no matchers" flag in the plate foot — shared by BOTH rule forms (add and
* edit), so the same non-blocking warning reads identically in either. */
.rt-edit-warn {
font-family: var(--font-mono);
font-size: 11.5px;
@@ -800,9 +836,6 @@ select.rt-input {
justify-content: flex-start;
align-self: start;
}
.rt-field-wide {
grid-column: auto;
}
.rt-rs-row {
grid-template-columns: 1fr;
row-gap: 10px;
+256 -122
View File
@@ -6,11 +6,12 @@ import {
apply as apiApply,
getConfig,
putConfig,
getRulesReachability,
getRulesetStatus,
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { Model, Rule, Ruleset, RulesetStatus } from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -21,9 +22,7 @@ import type { Model, Rule, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
type RRule = Rule & {
Src?: string[] | null
DstDomain?: string[] | null
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
Proto?: string
Kill?: string
@@ -34,6 +33,14 @@ type RRule = Rule & {
// Minutes east of UTC anchoring the schedule's wall-clock times; captured
// from the editing browser on save (the router has no tzdata). 0 ⇒ UTC.
SchedUTCOffset?: number
// The daemon's unmigrated-rule tripwire (model.Rule.LegacyDst), read-only here.
// Non-empty ⇒ the config STILL carries the schema-v1 `dst_domain`/`dst_ip` that
// schema v2 removed, i.e. `shaterd migrate` never ran or could not commit. Each
// element is the raw `<option>=<value>` text so the panel can quote what was
// found. The daemon holds such a rule disabled; the panel only reports it (it
// is never rendered back to UCI, so a config write from here drains it out).
// Field name is the Go one: model.Rule has no json tags.
LegacyDst?: string[] | null
}
/**
@@ -96,7 +103,6 @@ function ProtoOptions({ value }: { value: string }) {
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
const byOrder = (a: RRule, b: RRule): number => a.Order - b.Order
const csv = (s: string): string[] => s.split(',').map((x) => x.trim()).filter(Boolean)
// --- ruleset helpers --------------------------------------------------------
// A `config ruleset` (api.ts Ruleset) is a named domain/ipcidr list a rule
@@ -212,18 +218,59 @@ function everyLabel(sec: number): string {
return `every ${sec}s`
}
/** A rule with no matcher of any kind is the effective catch-all (route Final). */
/** A rule with no matcher of any kind is the effective catch-all (route Final).
* Mirrors model.IsCatchAll on the daemon side — the two must agree or the
* "never applies" badge lands on a different row than the apply warning.
*
* AN UNMIGRATED RULE IS NEVER A CATCH-ALL, and that is the first thing checked
* here, exactly as on the Go side (model/reachability.go). When LegacyDst is
* non-empty the rule's destination is still written in the schema-v1 options
* the parser no longer reads, so its lack of matchers means "the destination is
* unreadable", not "matches everything" — reading it the other way is precisely
* what turned an uncommitted `shaterd migrate` into route Final for the whole
* router. The daemon holds such a rule disabled and reports false here; if this
* copy disagreed, the panel would paint the row "default route · final" while
* the daemon routes nothing through it. */
function isCatchAll(r: RRule): boolean {
if (len(r.LegacyDst) > 0) return false
return (
len(r.Src) === 0 &&
len(r.DstDomain) === 0 &&
len(r.DstRuleset) === 0 &&
len(r.DstIP) === 0 &&
!(r.DstPort && r.DstPort.trim()) &&
!(r.Proto && r.Proto.trim())
)
}
/** isCatchAll's twin for a form still being edited: the live fields of either
* rule form with no matcher left in them.
*
* Such a rule is not "matches all" in the ordinary sense — the engine emits it
* as route.Final, and among several the LAST one in rule order owns it. So what
* saving one actually does depends on what is last right now, and there are
* three cases:
* - no rules at all, or the last rule is a conditional one → the new rule
* lands last and TAKES the default, silently retargeting every otherwise
* unmatched flow (e.g. the whole LAN to `direct`, past the tunnel);
* - the last rule is already a catch-all → nextOrder() deliberately inserts
* the new one BEFORE it (and bumps the old one up), so the existing default
* keeps route.Final and the new rule is dead on arrival — the daemon
* reports it as shadowed and the row renders as such.
* Both outcomes are worth a warning, and neither form knows which it will be
* (the add form has no rule list), so the shared text says only what is certain:
* the rule has no matchers. It is a legal configuration either way, so neither
* form blocks it — they warn, from this one predicate, so the flag cannot drift
* out of sync between add and edit. */
function formHasNoMatchers(f: {
src: string[]
port: string
rulesets: string[]
proto: string
}): boolean {
return (
f.src.length === 0 && f.port.trim() === '' && f.rulesets.length === 0 && f.proto.trim() === ''
)
}
/** Effective routing target for a rule (Target wins; a bare Egress is a target too). */
function effectiveTarget(r: RRule): string {
if (r.Target && r.Target.trim()) return r.Target.trim()
@@ -262,16 +309,19 @@ interface TargetGroups {
nodes: TargetOpt[] // node:<n> (huge — rendered last)
}
// Free-text destination matchers offered by the ADD form. Domains are NOT one of
// them (the user's call): domain matching goes through named rulesets — that's
// what they exist for. 'none' = the rule matches by rulesets/source/proto alone.
// (Legacy rules that already carry DstDomain stay editable in the edit form.)
type MatchKind = 'none' | 'ip' | 'port'
// The add form's fields. WHERE traffic is going is a ruleset choice and nothing
// else — a rule has no inline domain or address list any more, so the old
// Match-kind picker (rulesets / ip / port) collapsed into a plain Port field
// beside the ruleset picker. The cost is real and accepted: routing a single
// domain is no longer done here — you leave for the Rulesets panel, create the
// list, fill it, and come back to check it. A "create a list from here" shortcut
// was proposed and rejected (DECISIONS.md D21): a second place to author a list
// is a second place for its entry semantics and duplicate-name rules to drift,
// which is the exact thing D21 removed.
interface AddForm {
name: string
src: string[]
matchKind: MatchKind
matchValue: string
port: string
rulesets: string[]
proto: string
target: string
@@ -283,8 +333,7 @@ interface AddForm {
const EMPTY_FORM: AddForm = {
name: '',
src: [],
matchKind: 'none',
matchValue: '',
port: '',
rulesets: [],
proto: '',
target: 'direct',
@@ -334,6 +383,22 @@ export default function Routing() {
toastTimer.current = window.setTimeout(() => setToast(null), 2600)
}, [])
// Which rules can never fire, keyed by their position in the model's Rules array
// — NOT by name. The config that made this necessary had two rules both called
// `default`, which is exactly when a name-keyed verdict badges the wrong row.
const [reach, setReach] = useState<Map<number, RuleReach>>(new Map())
const loadReach = useCallback(async () => {
try {
const { rules } = await getRulesReachability()
setReach(new Map(rules.map((r) => [r.index, r])))
} catch {
// An older daemon has no such endpoint, and a stopped one answers nothing.
// Drop the verdicts rather than keep stale ones: no badge is honest, a badge
// about the previous config is not.
setReach(new Map())
}
}, [])
const load = useCallback(async () => {
try {
setLoadError(null)
@@ -341,7 +406,8 @@ export default function Routing() {
} catch (e) {
setLoadError(errMsg(e))
}
}, [])
void loadReach()
}, [loadReach])
useEffect(() => {
void load()
}, [load])
@@ -407,6 +473,36 @@ export default function Routing() {
// Every rule name, for the edit form's duplicate-name guard (it excludes self).
const ruleNames = useMemo(() => new Set(rules.map((r) => r.Name)), [rules])
// Each rule's position in the model's Rules array — the key the daemon's
// reachability verdicts use. `rules` above is a sorted COPY of the same object
// references, so identity survives the sort and this map stays valid.
const modelIndex = useMemo(() => {
const m = new Map<RRule, number>()
;((config?.Rules as RRule[] | null | undefined) ?? []).forEach((r, i) => m.set(r, i))
return m
}, [config])
/**
* The verdict for one rule, or null when it can fire.
*
* Verdicts are fetched separately from the config, so between an optimistic edit
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
* echoed name and order is what stops that window from putting a "never applies"
* badge on a working rule: a mismatch means the verdict is not about this row,
* and no badge is the honest answer.
*/
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const i = modelIndex.get(r)
if (i === undefined) return null
const v = reach.get(i)
if (!v || !v.unreachable || !v.shadowed_by) return null
if (v.name !== r.Name || v.order !== r.Order) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
},
[modelIndex, reach],
)
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
const rulesets = useMemo<Ruleset[]>(
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
@@ -458,6 +554,10 @@ export default function Routing() {
await putConfig(next)
setSavedPending(true)
flash(okMsg)
// The verdicts describe the config on disk, which just changed — re-ask.
// Adding or moving a rule is precisely what turns a working default into a
// superseded one, and vice versa.
void loadReach()
} catch (e) {
setConfig(prev)
setActionError(errMsg(e))
@@ -466,7 +566,7 @@ export default function Routing() {
setSaving(false)
}
},
[config, flash],
[config, flash, loadReach],
)
const commitRules = useCallback(
@@ -571,6 +671,10 @@ export default function Routing() {
)
// Insert a new rule just above the catch-all (so a specific rule can actually match).
// isCatchAll() is false for an unmigrated rule, which is the right answer here too:
// such a rule is held disabled and owns no default route, so there is nothing to
// insert ahead of — the new rule simply goes last, where a rule with no matchers
// does become the default.
const nextOrder = useCallback((): { order: number; bumpCatchAll?: { name: string; order: number } } => {
if (rules.length === 0) return { order: 10 }
const last = rules[rules.length - 1]
@@ -622,18 +726,15 @@ export default function Routing() {
return
}
setFormError(null)
const mv = form.matchValue.trim()
const rule: RRule = {
Name: name,
Enabled: true,
Order: 0,
Src: form.src,
// Domains are matched via rulesets only — the add form has no free-text
// domain matcher by design.
DstDomain: [],
// Destination = rulesets, always. Domains and addresses live in a
// `config ruleset` so one list serves every rule that needs it.
DstRuleset: form.rulesets,
DstIP: form.matchKind === 'ip' ? csv(mv) : [],
DstPort: form.matchKind === 'port' ? mv : '',
DstPort: form.port.trim(),
Proto: form.proto,
Target: form.target,
Egress: '',
@@ -737,14 +838,18 @@ export default function Routing() {
{rules.length === 0 ? (
<div className="rt-empty">
<p>No rules — all traffic follows the default route.</p>
<p className="rt-empty-sub">Add a rule below to steer a domain, address, or port.</p>
<p className="rt-empty-sub">Add a rule below to steer a destination list, source, or port.</p>
</div>
) : (
<ol className="rt-list" aria-label="Routing rules in first-match order">
{rules.map((r, i) =>
editingRule === r.Name ? (
<RuleEditForm
key={r.Name}
// Rule names are NOT unique in the wild — the config that prompted
// the never-applies badge had two rules called `default`, and a
// duplicate React key makes the second row shadow the first. The
// model index disambiguates without changing row identity.
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
initial={r}
names={ruleNames}
targets={targets}
@@ -755,12 +860,13 @@ export default function Routing() {
/>
) : (
<RuleRow
key={r.Name}
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
rule={r}
index={i}
total={rules.length}
busy={saving}
editingOther={editingRule !== null}
shadow={shadowOf(r)}
onEdit={onEditRule}
onToggle={onToggle}
onMove={onMove}
@@ -810,6 +916,7 @@ function RuleRow({
total,
busy,
editingOther,
shadow,
onEdit,
onToggle,
onMove,
@@ -820,6 +927,8 @@ function RuleRow({
total: number
busy: boolean
editingOther: boolean
/** Set when the daemon reports this rule can never fire; null when it can. */
shadow: { by: string; byOrder: number; reason: string } | null
onEdit: (name: string) => void
onToggle: (name: string) => void
onMove: (name: string, dir: 'up' | 'down') => void
@@ -829,13 +938,29 @@ function RuleRow({
// the first-match order can't shift under the open form. Edit itself stays live —
// clicking it just swaps which row is being edited.
const frozen = busy || editingOther
const isDefault = isCatchAll(rule)
// A rule with no conditions is the router's default — but only ONE of them can
// be, and the daemon says which. A superseded one must not wear the default's
// marks (the dashed accent frame, the "· final" order chip, the "everything not
// matched above" line): those are the claim that made two `default` rules
// indistinguishable in the first place.
const dead = shadow !== null
// The unmigrated-rule tripwire (see RRule.LegacyDst / model.IsCatchAll): this
// rule's destination is still in the removed schema-v1 options, so the daemon
// holds it disabled. It is NOT an ordinary disabled rule — nobody switched it
// off — so it gets the same "wired but not connected" amber treatment as a
// shadowed rule, plus a badge and a line saying what to run. isCatchAll()
// already refuses to call it the default, so `final` marks cannot land here.
const legacyDst = (rule.LegacyDst ?? []).filter(Boolean)
const unmigrated = legacyDst.length > 0
const inert = dead || unmigrated
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
const cls = [
'rt-rule',
rule.Enabled ? '' : 'off',
isDefault ? 'final' : '',
inert ? 'dead' : '',
]
.filter(Boolean)
.join(' ')
@@ -852,7 +977,7 @@ function RuleRow({
>
▲
</button>
<span className={isDefault ? 'rt-ord final' : 'rt-ord'}>
<span className={isDefault ? 'rt-ord final' : inert ? 'rt-ord dead' : 'rt-ord'}>
{isDefault ? '·' : rule.Order}
</span>
<button
@@ -870,9 +995,30 @@ function RuleRow({
<div className="rt-head">
<span className="rt-name">{rule.Name}</span>
{isDefault && <span className="rt-badge">default route · final</span>}
{unmigrated && <span className="rt-badge dead">held off · not migrated</span>}
{dead && <span className="rt-badge dead">never applies</span>}
</div>
<div className="rt-match">
{isDefault ? (
{unmigrated ? (
// Why the rule is off and what fixes it. Same voice as the shadow note:
// state, cause, one command. The daemon says the same thing through the
// apply warnings (model.ValidateRules); this puts it on the row it is about.
<span className="rt-dead-note">
Destination still written the old way (<span className="mono">{legacyDst.join(', ')}</span>
) — this config was never migrated, so shaterd cannot read where this rule sends
traffic and holds it disabled. Run <span className="mono">shaterd migrate</span> on the
router to turn those entries into a ruleset, then enable the rule again.
</span>
) : dead ? (
// The badge says it never fires; this line says what beat it and what to
// do. Visible text, not a tooltip — the operator has to be able to find
// the other rule, and two rows can carry the same name.
<span className="rt-dead-note" title={shadow.reason}>
“{shadow.by}” (order {shadow.byOrder}) has no conditions either and runs after this
one, so it is the default the router uses. Give this rule a condition, or delete one
of the two.
</span>
) : isDefault ? (
<span className="rt-nomatch">everything not matched above</span>
) : (
<Matchers rule={rule} />
@@ -897,11 +1043,21 @@ function RuleRow({
>
Edit
</button>
{/* An unmigrated rule cannot be switched on from here, and the switch says
so rather than pretending: the daemon holds it disabled, but a config
write from the panel DROPS the unreadable legacy options (render.go
emits neither), so enabling it here would save a live rule with no
destination left at all — the catch-all this tripwire exists to
prevent. `shaterd migrate` clears LegacyDst and the switch comes back. */}
<Toggle
pressed={rule.Enabled}
onChange={() => onToggle(rule.Name)}
label={`${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`}
disabled={frozen}
label={
unmigrated
? `Rule ${rule.Name} is held disabled until the config is migrated`
: `${rule.Enabled ? 'Disable' : 'Enable'} rule ${rule.Name}`
}
disabled={frozen || unmigrated}
/>
<button
type="button"
@@ -935,9 +1091,7 @@ function Matchers({ rule }: { rule: RRule }): ReactNode {
)
}
listChip('src', rule.Src, 'src')
listChip('dns', rule.DstDomain, 'dom')
listChip('ruleset', rule.DstRuleset, 'rs')
listChip('ip', rule.DstIP, 'ip')
if (rule.DstPort && rule.DstPort.trim()) {
chips.push(
<span className="rt-chip" key="port">
@@ -1053,7 +1207,16 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
)
}
/** The dst_ruleset checkbox group. Renders nothing when no rulesets exist. */
/**
* The destination picker: which rulesets this rule matches (dst_ruleset).
*
* Checkboxes and nothing else. This is the ONLY way a rule names a destination,
* so it renders even when the config has no lists yet — an empty picker that says
* where lists come from is the honest answer, and hiding it would leave the rule
* form with no destination control at all. Building and filling a list is the
* Rulesets panel's job, deliberately kept out of the rule editor so a list is
* created in exactly one place.
*/
function RulesetPicker({
options,
selected,
@@ -1065,24 +1228,26 @@ function RulesetPicker({
busy: boolean
onToggle: (name: string) => void
}): ReactNode {
if (options.length === 0) return null
return (
<div className="rt-rsel">
<span className="rt-flabel">Match rulesets — dst_ruleset</span>
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
<span className="rt-flabel">Destination — dst_ruleset</span>
{options.length > 0 && (
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
)}
<p className="rt-rsel-hint">
The rule also matches any traffic in the checked list(s). Combine with a domain, address, or
port, or use a ruleset on its own.
{options.length === 0
? 'No rulesets yet. Add one under Rulesets below, then come back and check it here — a rule matches a destination through a ruleset only.'
: 'The rule matches traffic in ANY checked list. Narrow it further with a source, port or protocol.'}
</p>
</div>
)
@@ -1208,7 +1373,14 @@ function AddRule({
set('rulesets', form.rulesets.includes(n) ? form.rulesets.filter((x) => x !== n) : [...form.rulesets, n])
const toggleDay = (d: string) =>
set('schedDays', form.schedDays.includes(d) ? form.schedDays.filter((x) => x !== d) : [...form.schedDays, d])
const matchPlaceholder = form.matchKind === 'ip' ? '10.0.0.0/8, 100.64.0.0/10' : '443, 8080-8090'
// Same flag as the edit form, from the same predicate: no matcher = this rule
// asks to be route.Final. Whether it wins depends on what is last — it takes
// the default when there is no catch-all yet (or the last rule is conditional),
// and lands dead when there is one, because nextOrder() inserts ahead of it.
// Both are worth flagging, so the text states the fact and not the outcome.
// Shown, never blocking: a default route is a legal thing to write.
const noMatchers = formHasNoMatchers(form)
return (
<form className="rt-add" onSubmit={onSubmit} aria-label="Add a routing rule">
@@ -1241,32 +1413,17 @@ function AddRule({
</label>
<label className="rt-field">
<span className="rt-flabel">Match</span>
<select
<span className="rt-flabel">Port(s)</span>
<input
className="rt-input mono"
value={form.matchKind}
onChange={(e) => set('matchKind', e.target.value as MatchKind)}
>
<option value="none">rulesets only</option>
<option value="ip">ip / cidr</option>
<option value="port">port</option>
</select>
value={form.port}
onChange={(e) => set('port', e.target.value)}
placeholder="443, 8080-8090"
autoComplete="off"
spellCheck={false}
/>
</label>
{form.matchKind !== 'none' && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">{form.matchKind === 'port' ? 'Port(s)' : 'Address(es)'}</span>
<input
className="rt-input mono"
value={form.matchValue}
onChange={(e) => set('matchValue', e.target.value)}
placeholder={matchPlaceholder}
autoComplete="off"
spellCheck={false}
/>
</label>
)}
<label className="rt-field">
<span className="rt-flabel">Proto</span>
<select
@@ -1313,6 +1470,9 @@ function AddRule({
<Button type="submit" variant="primary" disabled={busy}>
{busy ? 'Saving…' : 'Add rule'}
</Button>
{noMatchers && !error && (
<span className="rt-edit-warn">no matchers — matches everything</span>
)}
{error && (
<span className="rt-add-error" role="alert">
{error}
@@ -1324,11 +1484,11 @@ function AddRule({
}
// --- edit-a-rule plate (inline, replaces the row it edits) ------------------
// Unlike AddRule, editing exposes all three destination matchers at once
// (Domain(s) / IP-CIDR(s) / Port) rather than a single Match picker — a real rule
// can carry several matcher kinds simultaneously and none may be silently dropped.
// The full original rule is spread into the result on save, so Order / Enabled /
// Kill / Egress (and anything else off-form) survive untouched.
// Same fields as AddRule, on purpose: a rule carries exactly one destination
// mechanism (rulesets) plus port/proto/source, so there is nothing an edit can
// reveal that the add form hides. The full original rule is spread into the
// result on save, so Order / Enabled / Kill / Egress (and anything else off-form)
// survive untouched.
function RuleEditForm({
initial,
names,
@@ -1348,8 +1508,6 @@ function RuleEditForm({
}) {
const [name, setName] = useState(initial.Name)
const [src, setSrc] = useState<string[]>([...(initial.Src ?? [])])
const [domain, setDomain] = useState((initial.DstDomain ?? []).join(', '))
const [ip, setIp] = useState((initial.DstIP ?? []).join(', '))
const [port, setPort] = useState(initial.DstPort ?? '')
const [proto, setProto] = useState(initial.Proto ?? '')
const [target, setTarget] = useState(effectiveTarget(initial))
@@ -1365,14 +1523,15 @@ function RuleEditForm({
const toggleDay = (d: string) =>
setSchedDays((cur) => (cur.includes(d) ? cur.filter((x) => x !== d) : [...cur, d]))
// This rule's destination may still be in the schema-v1 options (RRule.LegacyDst).
// The form cannot show or edit them — they are not fields any more — and saving
// writes the model back through render.go, which does not emit them. So a save
// here silently DISCARDS that destination. Say so before it happens; the fix is
// `shaterd migrate`, which converts them into a ruleset this form can check.
const legacyDst = (initial.LegacyDst ?? []).filter(Boolean)
// A rule with no matcher of any kind is a catch-all — legal, but worth flagging.
const noMatchers =
src.length === 0 &&
csv(domain).length === 0 &&
csv(ip).length === 0 &&
port.trim() === '' &&
rulesets.length === 0 &&
proto.trim() === ''
const noMatchers = formHasNoMatchers({ src, port, rulesets, proto })
const submit = (e: FormEvent) => {
e.preventDefault()
@@ -1394,8 +1553,6 @@ function RuleEditForm({
...initial,
Name: nm,
Src: src,
DstDomain: csv(domain),
DstIP: csv(ip),
DstPort: port.trim(),
DstRuleset: rulesets,
Proto: proto,
@@ -1414,7 +1571,15 @@ function RuleEditForm({
<form className="rt-add rt-edit-form" onSubmit={submit} aria-label={`Edit rule ${initial.Name}`}>
<div className="rt-add-hd">
<span className="rt-add-title">Edit {initial.Name}</span>
<span className="rt-add-sub">Empty a field to drop that matcher. Save, then Apply.</span>
<span className="rt-add-sub">
Uncheck a list or clear a field to drop that matcher. Save, then Apply.
</span>
{legacyDst.length > 0 && (
<span className="rt-edit-warn">
not migrated — saving discards its old destination ({legacyDst.join(', ')}); run
shaterd migrate first
</span>
)}
</div>
<div className="rt-fields">
@@ -1444,37 +1609,6 @@ function RuleEditForm({
/>
</label>
{/* Domains are matched via rulesets by design — this legacy field only
appears when the rule ALREADY carries free-text domains, so they
stay visible and clearable rather than silently preserved. */}
{(initial.DstDomain ?? []).length > 0 && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">Domain(s) — legacy</span>
<input
className="rt-input mono"
value={domain}
onChange={(e) => setDomain(e.target.value)}
placeholder="youtube.com, *.googlevideo.com"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
)}
<label className="rt-field rt-field-wide">
<span className="rt-flabel">IP / CIDR(s)</span>
<input
className="rt-input mono"
value={ip}
onChange={(e) => setIp(e.target.value)}
placeholder="10.0.0.0/8, 100.64.0.0/10"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
<label className="rt-field">
<span className="rt-flabel">Port(s)</span>
<input
+96 -2
View File
@@ -16,6 +16,7 @@ import (
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/control"
E "github.com/sagernet/sing/common/exceptions" // lx:tproxy_writeback_connect
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
@@ -68,7 +69,26 @@ func (t *TProxy) Start(stage adapter.StartStage) error {
}
func (t *TProxy) Close() error {
return t.listener.Close()
err := t.listener.Close()
// lx:begin tproxy_writeback_connect
// Closing the listener stops INGRESS but leaves every live UDP NAT session in
// the cache, and each session holds a write-back socket bound to its original
// destination. Nothing else ever wakes those sessions: the cache evicts
// lazily (on the next Get/Add), and after this inbound is gone there is no
// next Get. For a DNS session answered in-engine by hijack-dns there is not
// even an outbound connection whose failure could unwind it, so its socket
// survives until a GC finalizer happens to reach it.
//
// That is invisible upstream, where an inbound is closed once at shutdown. On
// this fork the engine is REBUILT on every config apply, so each apply would
// strand another generation of sockets on the router's own LAN :53. Purge
// evicts every session; the cache's OnEvict closes the conn, which unblocks
// the session's routing goroutine and runs its onClose — the one place the
// write-back socket is actually closed. Purge AFTER the listener so a packet
// arriving mid-teardown cannot re-create a session behind us.
t.udpNat.Purge()
// lx:end tproxy_writeback_connect
return err
}
func (t *TProxy) NewConnection(ctx context.Context, conn net.Conn, metadata adapter.InboundContext, onClose N.CloseHandlerFunc) {
@@ -121,18 +141,92 @@ type tproxyPacketWriter struct {
conn *net.UDPConn
}
// lx:begin tproxy_writeback_connect
//
// The TPROXY UDP write-back socket is bound to the ORIGINAL DESTINATION, so the
// client sees the reply coming from the address it addressed. Upstream leaves
// that socket UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort), and that is
// the bug this block exists for.
//
// An unconnected socket bound to <addr>:<port> is, as far as the kernel is
// concerned, a RECEIVER for that address:port — and because the socket also
// carries SO_REUSEADDR it silently joins the UDP demultiplex set of whatever
// else is bound there. Nothing ever reads from it: this writer only sends. So
// every datagram the kernel happens to hand it is lost.
//
// On a router that transparently intercepts LAN DNS towards its OWN address
// (`nft ... udp dport 53 tproxy ...`) the original destination IS the router's
// LAN address, so each intercepted DNS session parks another silent receiver on
// <router-lan-ip>:53 right next to dnsmasq's socket — observed on the stand:
// 33 such sockets against dnsmasq's one, several with a growing Recv-Q. The
// host's own queries to that address take the loopback path, are never diverted,
// and are therefore demultiplexed among all of them: they land in one of the
// silent sockets at random and time out, permanently and unpredictably, while
// every other DNS path on the box keeps working.
//
// Connecting the socket fixes it at the root. compute_score() in the kernel's
// UDP lookup REJECTS a connected socket for any peer other than the connected
// one, so a write-back socket can no longer be handed a datagram it will not
// read. Nothing about the reply changes — same spoofed source address, same
// single peer, one Write instead of one WriteTo.
//
// The unconnected path is kept verbatim for a destination that cannot be bound
// (a domain socksaddr), so no existing case regresses.
// newTProxyWriteBack creates the connected write-back socket: local address =
// the original destination (transparent bind), peer = the client. A package
// variable so the behaviour can be tested without CAP_NET_ADMIN.
var newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
var dialer net.Dialer
dialer.LocalAddr = destination.UDPAddr()
dialer.Control = control.Append(dialer.Control, control.ReuseAddr())
dialer.Control = control.Append(dialer.Control, redir.TProxyWriteBack())
conn, err := w.listener.DialContext(dialer, w.ctx, "udp", w.source.String())
if err != nil {
return nil, err
}
udpConn, loaded := conn.(*net.UDPConn)
if !loaded {
conn.Close()
return nil, E.New("tproxy write back: unexpected connection type ", conn)
}
return udpConn, nil
}
// lx:end tproxy_writeback_connect
func (w *tproxyPacketWriter) WritePacket(buffer *buf.Buffer, destination M.Socksaddr) error {
defer buffer.Release()
if w.listener.ListenOptions().NetNs == "" {
conn := w.conn
if w.destination == destination && conn != nil {
_, err := conn.WriteToUDPAddrPort(buffer.Bytes(), w.source)
// lx:begin tproxy_writeback_connect
// The cached socket is CONNECTED to the client, so this is a plain
// Write. A failed write also closes it: upstream only dropped the
// reference, leaving the fd to the GC finalizer.
_, err := conn.Write(buffer.Bytes())
if err != nil {
conn.Close()
w.conn = nil
}
return err
// lx:end tproxy_writeback_connect
}
}
// lx:begin tproxy_writeback_connect
if destination.IsIP() {
udpConn, err := newTProxyWriteBack(w, destination)
if err != nil {
return err
}
if w.listener.ListenOptions().NetNs == "" && w.destination == destination {
w.conn = udpConn
} else {
defer udpConn.Close()
}
return common.Error(udpConn.Write(buffer.Bytes()))
}
// lx:end tproxy_writeback_connect
var listenConfig net.ListenConfig
listenConfig.Control = control.Append(listenConfig.Control, control.ReuseAddr())
listenConfig.Control = control.Append(listenConfig.Control, redir.TProxyWriteBack())
@@ -0,0 +1,234 @@
package redirect
// Regression cover for the lx:tproxy_writeback_connect block in tproxy.go.
//
// The failure it guards against, seen on a live BPi-R3 Mini: `shaterd` held 33
// UNCONNECTED UDP sockets on the router's own LAN address :53 — the write-back
// sockets of intercepted DNS sessions — alongside dnsmasq's single socket on the
// same address:port. Nothing reads a write-back socket, so every host-originated
// query that the kernel demultiplexed into one of them was silently dropped
// (Recv-Q climbing, sender timing out). Connecting the socket to the one peer it
// ever talks to removes it from the demultiplex set for everybody else.
//
// The real socket needs CAP_NET_ADMIN (IP_TRANSPARENT) and a foreign bind, so
// these tests drive the seam (newTProxyWriteBack) with an ordinary connected
// loopback socket. That is enough to pin both properties that actually broke:
//
// - the write path uses the CONNECTED form. If WritePacket ever goes back to
// WriteToUDPAddrPort, Go returns ErrWriteToConnected on a connected socket
// and these tests fail — i.e. the test cannot pass with an unconnected
// write-back socket, which is exactly the regression.
// - repeated writes to the same destination REUSE one socket instead of
// accumulating a new one per packet.
import (
"context"
"net"
"net/netip"
"testing"
"time"
"github.com/sagernet/sing-box/common/listener"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/buf"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
)
// writeBackHarness stands up a loopback "client" socket and a tproxyPacketWriter
// whose socket factory returns a plain connected UDP socket aimed at it. It
// returns the writer, the client socket and a pointer to the factory call count.
func writeBackHarness(t *testing.T) (*tproxyPacketWriter, *net.UDPConn, *int) {
t.Helper()
client, err := net.ListenUDP("udp", &net.UDPAddr{IP: net.IPv4(127, 0, 0, 1)})
if err != nil {
t.Fatalf("client socket: %v", err)
}
t.Cleanup(func() { client.Close() })
source := netip.MustParseAddrPort(client.LocalAddr().String())
w := &tproxyPacketWriter{
ctx: context.Background(),
source: source,
// A listener with empty options is enough: WritePacket only reads
// ListenOptions().NetNs, and the factory is stubbed below.
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
destination: M.SocksaddrFrom(netip.MustParseAddr("127.0.0.1"), 53),
}
calls := 0
orig := newTProxyWriteBack
t.Cleanup(func() { newTProxyWriteBack = orig })
newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
calls++
// The real implementation binds the ORIGINAL DESTINATION transparently
// and connects to w.source; here we only reproduce the connect half,
// which is the property under test.
conn, err := net.DialUDP("udp", nil, net.UDPAddrFromAddrPort(w.source))
if err != nil {
return nil, err
}
return conn, nil
}
return w, client, &calls
}
// readOne reads one datagram from the client socket with a short deadline.
func readOne(t *testing.T, client *net.UDPConn) string {
t.Helper()
_ = client.SetReadDeadline(time.Now().Add(2 * time.Second))
b := make([]byte, 512)
n, _, err := client.ReadFrom(b)
if err != nil {
t.Fatalf("client read: %v", err)
}
return string(b[:n])
}
// TestWriteBackUsesConnectedSocket: the reply must go out over a CONNECTED
// socket. On the pre-fix code the write is WriteToUDPAddrPort, which Go refuses
// on a connected socket ("use of WriteTo with pre-connected connection"), so
// this test is red for exactly the shape that caused the outage.
func TestWriteBackUsesConnectedSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("first")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
if got := readOne(t, client); got != "first" {
t.Fatalf("client got %q, want %q", got, "first")
}
if *calls != 1 {
t.Fatalf("expected one write-back socket, got %d", *calls)
}
}
// TestWriteBackReusesOneSocket: a session that keeps answering the same client
// must keep ONE socket, not open a fresh one per datagram. The live box carried
// one socket per intercepted DNS session already; one per PACKET would turn a
// nuisance into an fd exhaustion.
func TestWriteBackReusesOneSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
for i, payload := range []string{"a", "b", "c", "d"} {
if err := w.WritePacket(buf.As([]byte(payload)).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket %d: %v", i, err)
}
if got := readOne(t, client); got != payload {
t.Fatalf("packet %d: client got %q, want %q", i, got, payload)
}
}
if *calls != 1 {
t.Fatalf("four packets to one destination must share one socket, got %d sockets", *calls)
}
if w.conn == nil {
t.Fatalf("the write-back socket must be cached on the writer for reuse")
}
}
// TestWriteBackClosesSocketOnWriteFailure: upstream dropped the reference to a
// failed socket without closing it, leaving the fd to the GC finalizer. On a box
// that already parks one socket per DNS session that is the wrong direction.
func TestWriteBackClosesSocketOnWriteFailure(t *testing.T) {
w, client, _ := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("warm")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
readOne(t, client)
cached := w.conn
if cached == nil {
t.Fatalf("expected a cached socket after the first write")
}
// Close it behind WritePacket's back so the next write fails, exactly as a
// dead peer or a torn-down plane would make it fail.
cached.Close()
if err := w.WritePacket(buf.As([]byte("boom")).ToOwned(), w.destination); err == nil {
t.Fatalf("a write on a closed socket must report the failure")
}
if w.conn != nil {
t.Fatalf("a failed write must drop the cached socket")
}
// A second Close on an already-closed conn is an error, which is how we know
// WritePacket closed it rather than merely forgetting it.
if err := cached.Close(); err == nil {
t.Fatalf("WritePacket must CLOSE the failed socket, not just nil the field")
}
}
// --- Close() must release the UDP NAT sessions --------------------------------
// natSessionHandler stands in for the router: it drains the session conn until
// it errors (which is what Close does to it) and then reports onClose, exactly
// as the real routing goroutine does. onClose is where the write-back socket is
// closed, so "onClose fired" is the observable proof the socket was released.
type natSessionHandler struct {
released chan error
}
func (h *natSessionHandler) NewPacketConnectionEx(ctx context.Context, conn N.PacketConn, source M.Socksaddr, destination M.Socksaddr, onClose N.CloseHandlerFunc) {
go func() {
for {
buffer := buf.NewSize(1024)
_, err := conn.ReadPacket(buffer)
buffer.Release()
if err != nil {
if onClose != nil {
onClose(err)
}
h.released <- err
return
}
}
}()
}
// TestTProxyCloseReleasesNatSessions pins the apply-level half of the leak.
//
// Closing the inbound used to close the listener only. The NAT cache evicts
// lazily, so after the inbound is gone nothing ever touches it again and every
// live session — with the write-back socket it holds on the router's own
// LAN :53 — was left to a GC finalizer. On this fork the engine is rebuilt on
// every config apply, so that is one stranded generation of sockets per apply.
func TestTProxyCloseReleasesNatSessions(t *testing.T) {
handler := &natSessionHandler{released: make(chan error, 1)}
tp := &TProxy{
ctx: context.Background(),
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
}
prepared := 0
tp.udpNat = udpnat.New(handler, func(source M.Socksaddr, destination M.Socksaddr, userData any) (bool, context.Context, N.PacketWriter, N.CloseHandlerFunc) {
prepared++
return true, context.Background(), nil, func(error) {}
}, time.Minute, false)
tp.udpNat.NewPacket(
[][]byte{{0x00}},
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.2"), 40000),
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.1"), 53),
nil,
)
if prepared != 1 {
t.Fatalf("expected one NAT session, got %d", prepared)
}
if err := tp.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
select {
case <-handler.released:
case <-time.After(5 * time.Second):
t.Fatal("Close must release every live NAT session (and with it the write-back socket it holds)")
}
}
+9 -4
View File
@@ -27,8 +27,11 @@
#
# VERSION Version string stamped into constant.Version. Resolution order:
# 1) this positional arg, if given
# 2) $SHATER_VERSION, if set
# 3) `git describe --tags` (nearest tag + commit)
# 2) $SHATER_VERSION, if set (CI sets it from ci/version.sh)
# 3) `ci/version.sh --binary` — THE single source of truth shared
# with the package version (vX.Y.Z-rR[-g<sha>], derived from
# the git tag exactly like PKG_VERSION/PKG_RELEASE), so the
# string the panel shows always matches `apk info shaterd`
# 4) fallback: v0.2.0-dev
# --fast Skip `npm ci` when panel/node_modules already exists (dev speed-up).
#
@@ -69,8 +72,10 @@ if [ -n "$VERSION_ARG" ]; then
VERSION="$VERSION_ARG"
elif [ -n "${SHATER_VERSION:-}" ]; then
VERSION="$SHATER_VERSION"
elif VERSION="$(git -C "$REPO" describe --tags 2>/dev/null)"; then
: # git describe succeeded
elif VERSION="$(sh "$REPO/ci/version.sh" --binary)"; then
# Same computation the PACKAGE version comes from (ci/version.sh), so the
# binary's constant.Version and the .ipk/.apk version can never drift apart.
: # ci/version.sh always succeeds (it falls back to 0.0.0 without git)
else
VERSION="v0.2.0-dev"
fi
+17
View File
@@ -488,6 +488,23 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
return changed, err
}
// The plane is COMPLETE only here: table + policy routing + sysctls. ApplyNft
// already flushed the DNS conntrack when it loaded the ruleset, but that is
// one step too early — ApplyRouting is idempotent BY del-then-add, so it opens
// a window in which the fwmark rule is momentarily absent, and any DNS flow
// that crosses that window is tracked against a plane that is still being
// assembled. Flushing once more now that every piece is in place is what makes
// "no entry survives the transition" actually true. Only on a real change (the
// fast path assembled nothing), best-effort, and cheap: the :53 entry count is
// bounded by the number of clients.
if !nftCurrent {
if n, ferr := netplane.FlushDNSConntrack(); ferr != nil {
a.log.Debug("flush DNS conntrack after plane change: ", ferr)
} else if n > 0 {
a.log.Debug("plane changed: dropped ", n, " stale DNS conntrack entries")
}
}
// (4) success. Bump the effective-state generation ONLY when something really
// moved: a no-op reconcile must not invalidate an armed commit-confirm window
// (cron reconciles every minute — counting those would cancel every rollback
+172 -6
View File
@@ -20,6 +20,7 @@ import (
"net/netip"
"strings"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/netplane"
@@ -30,6 +31,110 @@ import (
// engine.RuleSetIPCIDRs).
type ruleSetCIDRs func(tag string) ([]netip.Prefix, bool)
// SCHEMA v2 MOVED A RULE'S DESTINATION OUT OF THE RULE, AND THIS FILE HAD TO FOLLOW.
//
// Until D21 a routing rule spelled its destination inline: `dst_domain` became
// DefaultRule.Domain*, `dst_ip` became DefaultRule.IPCIDR. Both were visible right
// on the rule, so matchesUntunnelable could read them off it and the ADDRESS half
// could only ever be resolved through the engine for the rare geosite/geoip case.
//
// Since v2 a rule names a `config ruleset` and generate materialises it into
// Route.RuleSet, so BOTH halves arrive by reference. That broke two things at once,
// in opposite directions:
//
// - DOMAINS WENT INVISIBLE. A v1 rule `dst_domain=[bank.ru] dst_ip=[10/8]` ANDed
// its matchers, so it could not match an ICMP packet (no domain in one) and this
// file skipped it. After migration the same rule reads `rule_set: [rule-x,
// rule-x-ip]`, whose two sets are ORed — so the address set alone would make the
// rule look applicable, and the plan would start emitting an ALLOW (target
// direct) or a DENY (target block) for 10/8 that the operator never asked for.
// The ALLOW is the dangerous half: it sends previously-tunnelled ICMP out with
// the client's real address, i.e. an upgrade silently opening a leak. So a rule
// whose rule-sets are KNOWN to carry domain matchers is treated exactly as its
// v1 form was: inapplicable to untunnelable traffic. (The OR is deliberate and
// documented in D21 for the ENGINE's TCP/UDP path; it does not follow that a
// leak-guard should widen itself during an upgrade without a word.)
// - ADDRESSES WENT ENGINE-ONLY. `rule-x-ip` is an INLINE rule-set: its prefixes
// are sitting right there in the generated options. Asking the engine for them
// — and, when the engine is not up, truncating the whole walk and denying
// everything — was needless. Inline sets are now read from the options, so an
// engine that has not started yet no longer costs the operator their ping.
//
// Both come from the same index, built once per plan: what the CONFIG ITSELF can
// say about each rule-set. Remote/local sets stay the engine's business.
// ruleSetFacts is what the generated options alone reveal about one rule-set.
//
// opaque means "not knowable from here" — a local/remote set (only the engine has
// its contents) or an inline set with a rule shape this code does not model. An
// opaque set is never used to conclude anything; it falls through to the engine
// lookup exactly as before.
type ruleSetFacts struct {
opaque bool
hasDomain bool // carries a matcher only a NAMED destination can satisfy
cidrs []netip.Prefix // addresses, when they are knowable here
}
// indexRuleSets summarises every rule-set in the generated options by tag.
func indexRuleSets(sets []option.RuleSet) map[string]ruleSetFacts {
out := make(map[string]ruleSetFacts, len(sets))
for _, rs := range sets {
var f ruleSetFacts
// option.RuleSet leaves Type empty for inline in some marshalled forms; the
// generator always sets it, and both spellings mean the same thing here.
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
f.opaque = true
out[rs.Tag] = f
continue
}
for _, hr := range rs.InlineOptions.Rules {
if hr.Type != C.RuleTypeDefault && hr.Type != "" {
// A logical headless rule, or a shape sing-box adds later. Never guessed
// at — the same rule the walk itself follows for a logical route rule.
f = ruleSetFacts{opaque: true}
break
}
d := hr.DefaultOptions
if headlessNeedsUnavailableMatcher(d) {
f.hasDomain = true
}
f.cidrs = append(f.cidrs, parsePrefixes(d.IPCIDR)...)
}
out[rs.Tag] = f
}
return out
}
// headlessNeedsUnavailableMatcher is matchesUntunnelable's predicate for the
// HEADLESS rule shape a rule-set carries. It reports whether the rule demands
// something a packet with no domain, no ports, no sniffed protocol and no process
// identity cannot supply — in which case that rule-set cannot claim such a packet.
//
// The generator only ever emits domain matchers or ip_cidr into an inline routing
// rule-set, so in practice this is the domain test; the rest is there so a future
// matcher is refused rather than silently ignored.
func headlessNeedsUnavailableMatcher(d option.DefaultHeadlessRule) bool {
switch {
case len(d.Domain) > 0, len(d.DomainSuffix) > 0, len(d.DomainKeyword) > 0,
len(d.DomainRegex) > 0, len(d.AdGuardDomain) > 0:
return true // no domain in an ICMP/ESP/GRE packet
case len(d.Port) > 0, len(d.PortRange) > 0, len(d.SourcePort) > 0, len(d.SourcePortRange) > 0:
return true
case len(d.Network) > 0, len(d.QueryType) > 0:
return true
case len(d.ProcessName) > 0, len(d.ProcessPath) > 0, len(d.ProcessPathRegex) > 0,
len(d.PackageName) > 0, len(d.PackageNameRegex) > 0:
return true
case len(d.WIFISSID) > 0, len(d.WIFIBSSID) > 0:
return true
case d.Invert:
// Same reasoning as the route-rule case: inverting flips every conservative
// assumption, so the set may only ever withhold, never grant.
return true
}
return false
}
// buildUntunnelablePlan reduces the resolved routing rules to the first-match
// walk the nft forward chain can evaluate for a packet that has no domain, no
// ports and no sniffable protocol.
@@ -57,6 +162,9 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
// No routing at all: nothing is provably direct.
return plan
}
// Since schema v2 a rule's destination lives in a rule-set, so the rule-sets
// have to be read alongside the rules. See the block above ruleSetFacts.
sets := indexRuleSets(opts.Route.RuleSet)
for _, r := range opts.Route.Rules {
if r.Type == "logical" {
@@ -78,7 +186,10 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
// and the destination logic never allowed anything. Its `protocol: dns`
// matcher makes it inapplicable to a packet with no stream to sniff, which
// is the correct and obvious reading once the checks are in this order.
if !matchesUntunnelable(d) {
if !matchesUntunnelable(d, sets) {
if note, ok := namedDestinationNote(d, sets); ok {
plan.Warnings = append(plan.Warnings, note)
}
continue // needs a domain/port/protocol: cannot apply to this traffic
}
@@ -96,21 +207,28 @@ func buildUntunnelablePlan(opts option.Options, lookup ruleSetCIDRs) *netplane.U
mt := netplane.UntunnelableMatch{Allow: verdict}
// Destination predicate: explicit ip_cidr plus every rule-set's addresses.
// An INLINE rule-set is answered straight out of the config — it is fully
// present there — so a rule whose lists are all inline no longer depends on
// the engine being up. Only local/remote sets still have to be asked for.
dst := parsePrefixes(d.IPCIDR)
resolvable := true
var unresolved string
for _, tag := range d.RuleSet {
if f, known := sets[tag]; known && !f.opaque {
dst = append(dst, f.cidrs...)
continue
}
prefixes, ok := lookup(tag)
if !ok {
resolvable = false
unresolved = tag
break
}
dst = append(dst, prefixes...)
}
if !resolvable {
if unresolved != "" {
plan.Warnings = append(plan.Warnings, fmt.Sprintf(
"list %q is not loaded yet, so ping/IPTV/VPN passthrough stays blocked for the "+
"addresses it covers (it resolves itself once the list downloads)",
strings.Join(d.RuleSet, ", ")))
unresolved))
return plan
}
hasDstMatcher := len(d.IPCIDR) > 0 || len(d.RuleSet) > 0
@@ -198,7 +316,20 @@ func classifyAction(a option.RuleAction) (isDirect bool, class int) {
// Every matcher listed here is one the packet cannot satisfy, so a rule carrying
// any of them is inapplicable rather than universally matching. Matchers ANDed
// within a rule mean one unsatisfiable matcher makes the whole rule unsatisfiable.
func matchesUntunnelable(d option.DefaultRule) bool {
//
// sets carries the same test for the matchers that are no longer written on the
// rule: since schema v2 a destination is a rule-set reference, so a rule's domains
// live in Route.RuleSet rather than in d.Domain*. A rule referencing a set that is
// KNOWN to match by name is inapplicable here for exactly the reason a d.Domain
// rule is — the packet has no name to match. Sets whose contents this code cannot
// see (local/remote) say nothing either way; the address walk below already refuses
// to conclude anything from them.
func matchesUntunnelable(d option.DefaultRule, sets map[string]ruleSetFacts) bool {
for _, tag := range d.RuleSet {
if sets[tag].hasDomain {
return false
}
}
switch {
case len(d.Domain) > 0, len(d.DomainSuffix) > 0, len(d.DomainKeyword) > 0,
len(d.DomainRegex) > 0, len(d.Geosite) > 0:
@@ -229,6 +360,41 @@ func matchesUntunnelable(d option.DefaultRule) bool {
return true
}
// namedDestinationNote explains the one skip an operator can actually be
// surprised by: a rule that names BOTH a domain list and an address list.
//
// A domain-only rule contributes nothing here and always has, so saying so would be
// noise. A mixed rule is different — its addresses look like something this plan
// could act on, and it is precisely the shape `shaterd migrate` produces from a v1
// rule that carried dst_domain and dst_ip together. The rule is skipped so the
// upgrade cannot quietly change what ping does (see the block above ruleSetFacts);
// that decision is worth one line in the operator's warning list rather than none.
//
// ok=false when there is nothing surprising to report.
func namedDestinationNote(d option.DefaultRule, sets map[string]ruleSetFacts) (string, bool) {
var named []string
addressed := len(d.IPCIDR) > 0
for _, tag := range d.RuleSet {
f, known := sets[tag]
if known && f.hasDomain {
named = append(named, tag)
continue
}
// An unknown/opaque set may well hold addresses; so may a knowable one.
if !known || f.opaque || len(f.cidrs) > 0 {
addressed = true
}
}
if len(named) == 0 || !addressed {
return "", false
}
return fmt.Sprintf(
"a routing rule matches by name (%s) as well as by address, and ping/IPTV/VPN-passthrough traffic "+
"carries no name — so that rule is left out of the ping/IPTV decision entirely and later rules "+
"decide those addresses. Split it into a name rule and an address rule if you want the addresses "+
"decided here", strings.Join(named, ", ")), true
}
// parsePrefixes converts CIDR/bare-address strings to prefixes, dropping anything
// unparseable (it cannot be matched on, so it must not silently widen a set).
func parsePrefixes(in []string) []netip.Prefix {
+156 -9
View File
@@ -22,6 +22,16 @@ func tunnelModel() *model.Model {
return m
}
// pinnedIPSet declares the inline `type=ipcidr` rule-set a rule pins its
// destination addresses with. Since schema v2 a routing rule has no dst_ip of its
// own: addresses are a `config ruleset`, so the generated rule carries a rule_set
// reference and the plan resolves the actual prefixes through the lookup below —
// i.e. through the RUNNING engine in production (engine.RuleSetIPCIDRs), not out
// of the config text. planFor's table stands in for that.
func pinnedIPSet(name string, cidrs ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: cidrs}
}
// planFor generates m and builds the untunnelable plan, resolving rule-set tags
// from the supplied table. A tag absent from the table reports "not loaded".
func planFor(t *testing.T, m *model.Model, sets map[string][]string) *netplane.UntunnelablePlan {
@@ -58,11 +68,12 @@ func renderPlan(t *testing.T, m *model.Model, plan *netplane.UntunnelablePlan) s
// only 8.8.8.8 may be un-pingable and the rest of the internet must answer.
func TestOnlyPinnedAddressIsTunnelled(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}})
if !plan.DefaultAllow {
t.Fatalf("a catch-all `direct` rule must make the default ALLOW; plan=%+v", plan)
@@ -221,14 +232,148 @@ func TestDomainRuleDoesNotAffectUntunnelable(t *testing.T) {
}
}
// TestBlockedRuleDenies: an explicitly blocked destination stays dropped.
func TestBlockedRuleDenies(t *testing.T) {
// --- schema v2: the destination moved into rule-sets, and so did the analysis ---
// migratedRule reproduces exactly what `shaterd migrate` makes of a v1 rule that
// carried BOTH lists: `dst_domain=[bank.ru] dst_ip=[203.0.113.0/24]` becomes two
// inline rule-sets, `rule-x` and `rule-x-ip`, and the rule references both.
func migratedRule(target string) *model.Model {
m := tunnelModel()
m.Rulesets = []model.Ruleset{
{Name: "rule-x", Type: "domain", Source: "inline", Entries: []string{"full:bank.ru"}},
pinnedIPSet("rule-x-ip", "203.0.113.0/24"),
}
m.Rules = []model.Rule{
{Name: "bad", Enabled: true, Order: 10, DstIP: []string{"203.0.113.0/24"}, Target: "block"},
{Name: "x", Enabled: true, Order: 10,
DstRuleset: []string{"rule-x", "rule-x-ip"}, Target: target},
{Name: "dflt", Enabled: true, Order: 99, Target: "group:auto"},
}
return m
}
// TestMigratedDomainAndIPRuleStaysOutOfTheUntunnelablePlan is the upgrade
// regression. In v1 the rule ANDed its domain and its addresses, so it could never
// claim a packet that carries no domain and this plan skipped it. After the
// migration the same rule reads `rule_set: [rule-x, rule-x-ip]`, and rule_set
// entries are ORed — so the address set alone made the rule look applicable and the
// plan started emitting a verdict for 203.0.113.0/24 that nobody asked for.
//
// With target=direct that verdict is an ALLOW, i.e. an upgrade quietly sending
// previously-tunnelled ICMP out with the client's real source address. That is the
// half that makes this a leak and not just a surprise, so it is checked first.
func TestMigratedDomainAndIPRuleStaysOutOfTheUntunnelablePlan(t *testing.T) {
// Both engine states, because they fail differently: with the engine UP the old
// code emitted a live verdict for the addresses, and with it DOWN the same
// mistake hid behind the unresolvable-list truncation. Neither may happen.
for _, target := range []string{"direct", "block", "group:auto"} {
for _, engine := range []string{"up", "down"} {
t.Run(target+"/engine-"+engine, func(t *testing.T) {
m := migratedRule(target)
var loaded map[string][]string
if engine == "up" {
loaded = map[string][]string{"rs-rule-x": {}, "rs-rule-x-ip": {"203.0.113.0/24"}}
}
plan := planFor(t, m, loaded)
if len(plan.Matches) != 0 {
t.Fatalf("a rule that matches by NAME as well as by address must contribute no "+
"address step (it cannot match a packet that carries no name): %+v", plan.Matches)
}
if plan.DefaultAllow {
t.Errorf("the catch-all routes into the tunnel, so the default must still deny")
}
fwd := renderPlan(t, m, plan)
if strings.Contains(fwd, "203.0.113.0/24") {
t.Errorf("the migrated address list must not reach the data plane through the "+
"untunnelable policy:\n%s", fwd)
}
})
}
}
}
// TestMigratedDomainAndIPRuleSaysWhyItWasSkipped: skipping is the safe answer, but
// a silently different ping after an upgrade is the complaint this whole change
// exists to answer. The mixed rule — the exact shape migration produces — is named.
func TestMigratedDomainAndIPRuleSaysWhyItWasSkipped(t *testing.T) {
plan := planFor(t, migratedRule("direct"), nil)
var seen bool
for _, w := range plan.Warnings {
if strings.Contains(w, "matches by name") && strings.Contains(w, "rs-rule-x") {
seen = true
}
}
if !seen {
t.Fatalf("the skip must be explained and the list named; warnings=%v", plan.Warnings)
}
}
// TestDomainOnlyRuleSetIsQuiet: a rule whose only destination is a NAME list has
// always contributed nothing here and is not a surprise, so it must not produce a
// note. Only the mixed shape is worth a line.
func TestDomainOnlyRuleSetIsQuiet(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{
{Name: "names", Type: "domain", Source: "inline", Entries: []string{"full:bank.ru"}},
}
m.Rules = []model.Rule{
{Name: "names", Enabled: true, Order: 10, DstRuleset: []string{"names"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
if !plan.DefaultAllow {
t.Errorf("a name-only rule must not withhold the catch-all allow; plan=%+v", plan)
}
for _, w := range plan.Warnings {
if strings.Contains(w, "matches by name") {
t.Errorf("a name-only rule is not a surprise and needs no note: %q", w)
}
}
}
// TestInlineRuleSetNeedsNoEngine: an inline rule-set's addresses are sitting in the
// generated config. Asking the engine for them — and, while it is down, truncating
// the whole walk so ping/IPTV/VPN passthrough dies everywhere — was a conservative
// answer to a question that did not have to be asked. The verdict must now be the
// same whether or not the engine is up.
func TestInlineRuleSetNeedsNoEngine(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
down := planFor(t, m, nil) // engine not up
up := planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}}) // engine up
for name, plan := range map[string]*netplane.UntunnelablePlan{"engine-down": down, "engine-up": up} {
if !plan.DefaultAllow {
t.Fatalf("%s: the catch-all direct must allow the rest of the internet; plan=%+v", name, plan)
}
if len(plan.Matches) != 1 || plan.Matches[0].Allow {
t.Fatalf("%s: the tunnelled address must produce one DENY step, got %+v", name, plan.Matches)
}
if got := plan.Matches[0].Dst4; len(got) != 1 || got[0] != "8.8.8.8/32" {
t.Fatalf("%s: step destinations = %v, want [8.8.8.8/32]", name, got)
}
}
for _, w := range down.Warnings {
if strings.Contains(w, "not loaded yet") {
t.Errorf("an inline list is never 'not loaded': %q", w)
}
}
}
// TestBlockedRuleDenies: an explicitly blocked destination stays dropped.
func TestBlockedRuleDenies(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("bad", "203.0.113.0/24")}
m.Rules = []model.Rule{
{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "block"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, map[string][]string{"rs-bad": {"203.0.113.0/24"}})
if len(plan.Matches) != 1 || plan.Matches[0].Allow {
t.Fatalf("a blocked destination must deny untunnelable traffic too: %+v", plan.Matches)
}
@@ -361,11 +506,12 @@ func TestPlanNeverAcceptsTCPOrUDP(t *testing.T) {
m := tunnelModel()
m.Globals.IPv6 = true
m.Globals.Untunnelable = policy
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
fwd := renderPlan(t, m, planFor(t, m, nil))
fwd := renderPlan(t, m, planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}}))
// Only the lines the untunnelable policy emits are in scope: the tproxy
// diverts in prerouting legitimately match TCP/UDP, which is their job.
@@ -424,12 +570,13 @@ func TestLocalPlaneSurvivesEveryPlan(t *testing.T) {
// only grant those clients, not everyone.
func TestSourceScopedRuleNarrowsTheAllow(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("lab", "198.51.100.0/24")}
m.Rules = []model.Rule{
{Name: "lab", Enabled: true, Order: 10, Src: []string{"192.168.9.0/24"},
DstIP: []string{"198.51.100.0/24"}, Target: "direct"},
DstRuleset: []string{"lab"}, Target: "direct"},
{Name: "dflt", Enabled: true, Order: 99, Target: "group:auto"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-lab": {"198.51.100.0/24"}})
if len(plan.Matches) != 1 {
t.Fatalf("expected one step, got %+v", plan.Matches)
}
+9
View File
@@ -91,6 +91,15 @@ var criticalMarkers = []string{
"not covered", // an interface outside the fail-closed guard
"fail-closed", // ''
"REJECTED", // an unusable interface name
// A condition-less rule retired by a later condition-less rule whose target is
// `direct` (generate/route.go warnUnreachableRules): the operator's default
// policy — a tunnel, or a block — is not the one the router uses, so everything
// unmatched leaves on the plain WAN. The marker is the full clause, not the
// shorter "leaves over the plain WAN" that ruleKillFallback's kill=open note
// also contains: that one is a DELIBERATE per-rule bypass the operator asked
// for, and grading it critical here would be a different decision made by
// accident.
"leaves over the plain WAN with your real IP address",
}
// protectionSections are entity kinds whose whole purpose is to block or divert
+42
View File
@@ -240,3 +240,45 @@ func blockGlobals() model.Globals {
g.Untunnelable = netplane.UntunnelableBlock
return g
}
// TestCollectWarningsGradesUnreachableRules (B1): a condition-less rule retired by
// a later condition-less one is graded by CONSEQUENCE, not by the mere fact that a
// setting is dead.
//
// - the surviving default is `direct` while the retired one wanted a tunnel:
// the operator's default policy is not in effect and everything unmatched
// leaves on the plain WAN — critical, the panel must show it loudly;
// - the surviving default is the tunnel and the dead one was `direct`: a dead
// knob, nothing is leaking — warning.
//
// Both must be attributed to section "rule" + the rule's name, so the panel can
// deep-link to the offending row instead of printing prose.
func TestCollectWarningsGradesUnreachableRules(t *testing.T) {
const leak = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "fallback" (order 100) has none either and comes after it, so "direct" wins and ` +
`this rule's target "group:auto" is never applied. Everything no other rule matches leaves ` +
`over the plain WAN with your real IP address. Delete one of the two, or give this one a condition`
const dead = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "default" (order 100) has none either and comes after it, so the default the ` +
`router uses is "group:auto" and this rule's target "direct" is never applied. Delete one ` +
`of the two, or give this one a condition so it can match something`
check := func(text, wantSeverity string) {
t.Helper()
found := false
for _, w := range collectWarnings(blockGlobals(), []string{text}, nil, nil) {
if w.Section != "rule" || w.Name != "default" {
continue
}
found = true
if w.Severity != wantSeverity {
t.Errorf("severity = %q, want %q for: %s", w.Severity, wantSeverity, text)
}
}
if !found {
t.Fatalf("never-applied warning not attributed to rule %q: %s", "default", text)
}
}
check(leak, SeverityCritical)
check(dead, SeverityWarning)
}
+27
View File
@@ -0,0 +1,27 @@
package main
// Control-plane log formatting (the IO-free half, so it is unit-testable on any
// dev host — same split as profilewatch.go).
//
// The daemon's own logger used to be built with a bare log.Formatter, whose
// DisableColors zero value is false: every control-plane line went to procd's
// stderr — i.e. straight into syslog — carrying aurora escape sequences. See
// shater/logsink/color.go for why that is a defect and not a cosmetic.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/logsink"
)
// controlLogFormatter builds the control-plane logger's formatter. out is the
// process's real stderr (the sink's syslog half): colours are emitted only when
// that is a terminal, never when it is procd/syslog or a log file.
func controlLogFormatter(baseTime time.Time, out *os.File) log.Formatter {
return log.Formatter{
BaseTime: baseTime,
DisableColors: !logsink.IsTTY(out),
}
}
+37
View File
@@ -0,0 +1,37 @@
package main
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestControlLogFormatterNoANSIOffTTY pins the control-plane half of the syslog
// colour leak: when the daemon's stderr is not a terminal — which is ALWAYS the
// case under procd, where stderr is the pipe procd relays to syslog — no line
// the daemon formats may contain an ESC (0x1b).
func TestControlLogFormatterNoANSIOffTTY(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
f := controlLogFormatter(time.Now(), w)
if !f.DisableColors {
t.Fatalf("colours enabled for a non-terminal stderr")
}
// A context ID exercises the second colouring branch of log/format.go (the
// 256-colour connection id), which is what produced ESC[38;5;193m on the router.
ctx := log.ContextWithNewID(context.Background())
for _, level := range []log.Level{log.LevelError, log.LevelWarn, log.LevelInfo, log.LevelDebug, log.LevelTrace} {
line := f.Format(ctx, level, "dns", "exchange failed for example.com. IN AAAA: unexpected EOF", time.Now())
if i := strings.IndexByte(line, 0x1b); i >= 0 {
t.Errorf("level %v: formatted line carries an ANSI escape at byte %d: %q", level, i, line)
}
}
}
+76 -8
View File
@@ -79,7 +79,7 @@ func dispatch(args []string) int {
case "mint-token":
return cmdMintToken()
case "nodes":
return cmdReadStub("nodes", "[]")
return cmdNodes()
case "stats":
return cmdReadStub("stats", "{}")
case "blocklist":
@@ -124,7 +124,7 @@ usage: shaterd <verb>
rollback roll back to the last-good config
status print daemon/data-plane status as JSON
mint-token mint a single-use panel handoff token (JSON) via the daemon
nodes print nodes as JSON
nodes print the configured nodes as JSON (manual + subscription)
stats print stats as JSON
blocklist update force a DNS-filter blocklist refresh (reconcile; no-op if down)
schedule due re-evaluate time-scheduled rules now (reconcile; no-op if down)
@@ -158,8 +158,11 @@ func cmdRun() int {
// The level starts at trace (nothing read yet) and is corrected from
// Globals.LogLevel as soon as UCI is read — the control plane now RESPECTS
// the configured level instead of the old unconditional trace.
// The formatter colours only when stderr is a real terminal
// (controlLogFormatter -> logsink.IsTTY): under procd stderr IS syslog, and
// ANSI escapes there break `logread | grep ERROR` and every log collector.
logFactory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now()}, sink, "", nil, false)
controlLogFormatter(time.Now(), os.Stderr), sink, "", nil, false)
log.SetStdLogger(logFactory.Logger())
logger := log.StdLogger()
@@ -172,10 +175,23 @@ func cmdRun() int {
logFactory.SetLevel(controlLogLevel(g.LogLevel))
}
// Refuse to start a second daemon (race-safe single-owner guard). procd's
// term_timeout ensures the previous instance exits before reload restarts us.
if pid, ok := daemonAlive(); ok && pid != os.Getpid() {
logger.Error("another shaterd is already running (pid ", pid, ") — refusing to start")
// Single-owner guard: never run two daemons at once. On a `restart` procd's
// `service delete` is ASYNCHRONOUS — the ubus call returns immediately while
// the outgoing instance is still running its honest teardown — and the
// following `service add` starts us straight away, so a predecessor being
// alive here is the NORMAL restart case, not an error.
//
// This used to exit(1) on the spot and lean on procd's `respawn ... 5 ...` to
// try again five seconds later. That is a blind retry, not synchronisation:
// it neither knows nor waits for the predecessor's teardown to finish, and it
// turns every restart into at least one logged crash plus a five-second hole
// in which the LAN has no plane at all. Waiting for the predecessor to exit
// makes `restart` behave exactly like `stop` + pause + `start`: our apply is
// then strictly ordered AFTER the previous teardown, which is the whole point
// of the guard.
if pid, waiting := waitForPredecessor(predecessorBudget, predecessorPoll, logger); waiting {
logger.Error("another shaterd is still running (pid ", pid, ") after waiting ",
predecessorBudget, " for it to exit — refusing to start")
return 1
}
if err := writePidfile(); err != nil {
@@ -832,6 +848,42 @@ func cmdStatus() int {
return 0
}
// cmdNodes prints the node inventory as a JSON array (see nodes.go for the
// shape and for why this is a view rather than the raw model.Node).
//
// Shape of the call mirrors cmdStatus: ask the RUNNING daemon first — it is the
// process that owns the engine, so its answer is the inventory the live box was
// built from — and fall back to reading the same on-disk desired state directly
// when there is no daemon (or it did not answer). Both sides call the SAME
// nodesJSON(), so the fallback cannot report something the daemon would not.
//
// The one thing it will never do is print `[]` because it could not find out:
// a read failure goes to stderr and exits non-zero, so "empty" on stdout with
// exit 0 means "no nodes are configured" and nothing else.
func cmdNodes() int {
if _, ok := daemonAlive(); ok {
resp, err := ctlRequest("nodes")
switch {
case err != nil:
fmt.Fprintf(os.Stderr, "shaterd nodes: %v — reading the on-disk config instead\n", err)
case strings.HasPrefix(strings.TrimSpace(resp), "["):
fmt.Println(strings.TrimSpace(resp))
return 0
default:
// The daemon answered, but with an error object rather than a list.
fmt.Fprintf(os.Stderr, "shaterd nodes: daemon: %s\n", strings.TrimSpace(resp))
}
}
b, err := nodesJSON()
if err != nil {
fmt.Fprintf(os.Stderr, "shaterd nodes: %v\n", err)
fmt.Println("[]") // stdout stays parseable JSON; the exit code carries the failure
return 1
}
fmt.Println(string(b))
return 0
}
// cmdMintToken asks the RUNNING daemon (over the control socket) for a single-use
// panel handoff token and prints the daemon's JSON reply verbatim — {"token":"..."}
// on success, {"error":"..."} otherwise. It is the CLI shim the LuCI/rpcd layer
@@ -869,6 +921,13 @@ func printJSONError(msg string) {
fmt.Println(string(b))
}
// cmdReadStub relays a read-only verb to the daemon and prints def when there is
// no daemon to ask. It is ONLY valid where def is the truth in the daemon-down
// case: `stats` counts what the LIVE engine saw, so with no engine there really
// is nothing counted and `{}` says exactly that. It is NOT a way to make a verb
// look implemented — `nodes` used to be routed through here with def="[]" while
// hundreds of nodes sat on disk (see nodes.go). Anything whose data outlives the
// daemon belongs in its own command that reads that data.
func cmdReadStub(cmd, def string) int {
if _, ok := daemonAlive(); ok {
if resp, err := ctlRequest(cmd); err == nil {
@@ -940,7 +999,16 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
case "rollback":
writeResult(conn, false, a.Rollback())
case "nodes":
writeLine(conn, "[]")
// Same assembly as the CLI fallback and as GET /api/config: model.ReadUCI
// (UCI + the per-subscription caches) projected onto the printable view.
b, err := nodesJSON()
if err != nil {
l.Warn("control socket: nodes: ", err)
e, _ := json.Marshal(map[string]string{"error": err.Error()})
writeLine(conn, string(e))
return
}
writeLine(conn, string(b))
case "stats":
if sa == nil {
writeLine(conn, "{}")
+132
View File
@@ -0,0 +1,132 @@
package main
// `shaterd nodes` — the node inventory, as JSON.
//
// This verb used to answer `[]` unconditionally (a `cmdReadStub("nodes", "[]")`
// on the CLI side and a hard-coded `writeLine(conn, "[]")` in the daemon's
// control-socket handler), while `shaterd --help` advertised "print nodes as
// JSON". The data was never missing: /etc/shater/subs/*.json holds every
// subscription-fetched node and `GET /api/config` reports the full merged set —
// on a live router that is hundreds of nodes answered as "none". An empty list
// is indistinguishable from a truthful "no nodes are configured", so the verb
// did not fail loudly, it lied quietly. Same class as the honesty fixes in
// 9dc954029 / aec82d444; the cure is the same: report what is actually there.
//
// # Data path (deliberately the SAME one /api/config uses)
//
// model.ReadUCI() = `uci export shater` (manual nodes + everything else) +
// model.MergeSubCaches (the per-subscription JSON caches). That single call is
// what the panel's handleConfigGet serves, what generate builds the engine from
// and what `sub update` writes back — so `shaterd nodes` cannot drift from the
// panel or from the running engine, because there is no second assembly here to
// drift. This file only PROJECTS that model onto a small, printable view.
//
// # Why a view and not the raw model.Node
//
// model.Node carries the share-link URI, and a share link is a credential
// (uuid/password in the query string). /api/config may return it — that path is
// session-authenticated and the panel needs the URI to edit a node — but a CLI
// verb whose output gets piped into support tickets, `logger`, and cron mail
// must not spray credentials. The view therefore reports the *derived* facts a
// share link answers (protocol, server, port) and drops the secret-bearing URI.
// Nodes whose URI does not parse are still listed, with the parse error in
// `parse_error`: the engine skips exactly those nodes (generate/outbound.go),
// and a node that is configured-but-unusable is precisely what an operator
// needs to see — hiding it would be the same lie in a smaller coat.
import (
"encoding/json"
"strings"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/parse"
)
// nodeView is ONE node as the `nodes` verb reports it. Field set: the model's
// own facts (name/enabled/sub/egress/stale/fingerprint) plus what the share link
// decodes to (protocol/server/port). Everything is omitempty-free where a reader
// would have to distinguish "absent" from "false"/"zero" — `enabled` and `stale`
// are always present so a consumer never has to guess.
type nodeView struct {
Name string `json:"name"`
Enabled bool `json:"enabled"`
// Sub is the subscription this node came from; "" means a manual node
// (model.Node.FromSub semantics, unchanged).
Sub string `json:"sub"`
// Protocol is the parsed protocol (vless|vmess|trojan|shadowsocks|wireguard…).
// When the URI does not parse it falls back to the bare URI scheme so the
// operator still sees what kind of thing failed; "" only when the URI is empty.
Protocol string `json:"protocol"`
Server string `json:"server,omitempty"`
Port uint16 `json:"port,omitempty"`
// Egress is the `config egress` this node's own upstream is bound to
// (multi-WAN); "" = default route.
Egress string `json:"egress,omitempty"`
// Stale marks a cached subscription node whose subscription failed to refresh.
Stale bool `json:"stale"`
Fingerprint string `json:"fingerprint,omitempty"`
// ParseError is the reason the engine will SKIP this node, verbatim from
// parse.ParseShareLink. Empty on every usable node.
ParseError string `json:"parse_error,omitempty"`
}
// readModel is the model source. A package var so the tests can exercise the
// assembly without a router's `uci` binary; production always reads the real
// merged desired state.
var readModel = model.ReadUCI
// nodesJSON reads the desired state and renders the node inventory as a JSON
// array. The array is never `null`: an empty configuration marshals to `[]`,
// which is the ONE case where `[]` is the truth.
func nodesJSON() ([]byte, error) {
m, err := readModel()
if err != nil {
return nil, err
}
return json.Marshal(nodeViews(m))
}
// nodeViews projects the merged model onto the printable view. Pure: no IO, no
// globals — the whole verb's logic is testable on a dev host.
func nodeViews(m *model.Model) []nodeView {
out := make([]nodeView, 0, len(m.Nodes)) // never nil => never `null`
for i := range m.Nodes {
n := m.Nodes[i]
v := nodeView{
Name: n.Name,
Enabled: n.Enabled,
Sub: n.FromSub,
Egress: n.Egress,
Stale: n.Stale,
Fingerprint: n.Fingerprint,
Protocol: uriScheme(n.URI),
}
// The engine reads a node through exactly this call (generate/outbound.go,
// chain.go, group.go); using it here is what makes "protocol" agree with
// what will actually be dialled, and the error agree with what will be
// skipped.
if p, err := parse.ParseShareLink(n.URI); err == nil {
if p.Protocol != "" {
v.Protocol = p.Protocol
}
v.Server = p.Server
v.Port = p.Port
} else if n.URI != "" {
v.ParseError = err.Error()
}
out = append(out, v)
}
return out
}
// uriScheme returns the lower-cased scheme of a share link ("vless://…" ->
// "vless"), or "" when there is none. It is the fallback protocol for a URI that
// ParseShareLink refuses, so an unusable node is still described rather than
// reported as a typeless blank.
func uriScheme(uri string) string {
i := strings.Index(uri, "://")
if i <= 0 {
return ""
}
return strings.ToLower(uri[:i])
}
+146
View File
@@ -0,0 +1,146 @@
package main
import (
"encoding/json"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// vlessURI is a syntactically complete share link (host, port, uuid, fragment)
// so ParseShareLink really succeeds and the derived fields are exercised.
const vlessURI = "vless://11111111-2222-3333-4444-555555555555@example.net:443?" +
"encryption=none&security=tls&sni=example.net&type=ws&path=%2Fws#NL-vless-1"
// writeSubCache points the subscription cache at a temp dir and persists one
// subscription's nodes there, the way `sub update` does on the router
// (/etc/shater/subs/<sub>.json). Returns nothing: the point is the side effect
// that model.MergeSubCaches will pick up.
func writeSubCache(t *testing.T, sub string, nodes []model.Node) {
t.Helper()
t.Setenv("SHATER_SUBS_DIR", t.TempDir())
if err := model.SaveSubCache(sub, nodes); err != nil {
t.Fatalf("SaveSubCache(%q): %v", sub, err)
}
}
// TestNodesJSONReportsCachedSubscriptionNodes is the regression for the bug this
// verb had: `shaterd nodes` answered `[]` while the subscription cache held
// hundreds of nodes. A NON-EMPTY cache must produce a NON-EMPTY list.
func TestNodesJSONReportsCachedSubscriptionNodes(t *testing.T) {
writeSubCache(t, "all-qomar", []model.Node{
{Name: "NL-vless-1", Enabled: true, URI: vlessURI, FromSub: "all-qomar", Fingerprint: "66b33"},
{Name: "NL-vless-2", Enabled: false, URI: vlessURI, FromSub: "all-qomar"},
})
// The model source stands in for `uci export shater` (absent on a dev host);
// the subscription half is the REAL model.MergeSubCaches path.
restore := readModel
readModel = func() (*model.Model, error) {
m := &model.Model{Globals: model.DefaultGlobals()}
model.MergeSubCaches(m)
return m, nil
}
t.Cleanup(func() { readModel = restore })
b, err := nodesJSON()
if err != nil {
t.Fatalf("nodesJSON: %v", err)
}
var got []nodeView
if err := json.Unmarshal(b, &got); err != nil {
t.Fatalf("nodesJSON produced invalid JSON %q: %v", b, err)
}
if len(got) == 0 {
t.Fatalf("nodes reported an EMPTY list while the subscription cache holds 2 nodes: %s", b)
}
if len(got) != 2 {
t.Fatalf("got %d nodes, want 2: %s", len(got), b)
}
byName := map[string]nodeView{}
for _, v := range got {
byName[v.Name] = v
}
first, ok := byName["NL-vless-1"]
if !ok {
t.Fatalf("cached node NL-vless-1 missing from %s", b)
}
if !first.Enabled {
t.Errorf("NL-vless-1.enabled = false, want true")
}
if first.Sub != "all-qomar" {
t.Errorf("NL-vless-1.sub = %q, want %q", first.Sub, "all-qomar")
}
if first.Protocol != "vless" {
t.Errorf("NL-vless-1.protocol = %q, want %q", first.Protocol, "vless")
}
if first.Server != "example.net" || first.Port != 443 {
t.Errorf("NL-vless-1 server:port = %s:%d, want example.net:443", first.Server, first.Port)
}
if first.ParseError != "" {
t.Errorf("NL-vless-1.parse_error = %q, want empty", first.ParseError)
}
if byName["NL-vless-2"].Enabled {
t.Errorf("NL-vless-2.enabled = true, want false (the model says disabled)")
}
// The credential-bearing share link must not be printed.
if strings.Contains(string(b), "11111111-2222-3333-4444-555555555555") {
t.Errorf("nodes output leaks the share-link uuid: %s", b)
}
}
// TestNodeViewsShape pins the projection: manual vs subscription, an unparsable
// URI still being LISTED (with the reason), and an empty model marshalling to
// `[]` rather than `null`.
func TestNodeViewsShape(t *testing.T) {
m := &model.Model{Nodes: []model.Node{
{Name: "manual-1", Enabled: true, URI: vlessURI, Egress: "wan2"},
{Name: "broken", Enabled: true, URI: "nosuch://whatever", FromSub: "qomar", Stale: true},
{Name: "no-uri", Enabled: false},
}}
got := nodeViews(m)
if len(got) != 3 {
t.Fatalf("nodeViews returned %d views, want 3", len(got))
}
if got[0].Sub != "" {
t.Errorf("manual node sub = %q, want empty", got[0].Sub)
}
if got[0].Egress != "wan2" {
t.Errorf("manual node egress = %q, want wan2", got[0].Egress)
}
if got[1].ParseError == "" {
t.Errorf("unparsable node must carry the reason the engine will skip it")
}
if got[1].Protocol != "nosuch" {
t.Errorf("unparsable node protocol = %q, want the bare scheme %q", got[1].Protocol, "nosuch")
}
if !got[1].Stale {
t.Errorf("stale flag lost")
}
if got[2].Protocol != "" || got[2].ParseError != "" {
t.Errorf("URI-less node: got protocol=%q parse_error=%q, want both empty", got[2].Protocol, got[2].ParseError)
}
b, err := json.Marshal(nodeViews(&model.Model{}))
if err != nil {
t.Fatalf("marshal empty: %v", err)
}
if string(b) != "[]" {
t.Errorf("empty model marshalled to %q, want []", b)
}
}
func TestURIScheme(t *testing.T) {
for _, tc := range []struct{ in, want string }{
{"vless://x@h:443", "vless"},
{"VMESS://payload", "vmess"},
{"", ""},
{"not-a-uri", ""},
{"://leading", ""},
} {
if got := uriScheme(tc.in); got != tc.want {
t.Errorf("uriScheme(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
+63
View File
@@ -0,0 +1,63 @@
package main
// The startup handshake that stops a restarting daemon from standing a new data
// plane up on top of the previous one's teardown.
//
// Deliberately free of build tags: the logic is pure timing and is exercised by
// the tests on any host, while the probe it polls (pidfile + kill(pid, 0)) is
// linux-only and is installed by predecessor_linux.go.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
)
// predecessorBudget / predecessorPoll bound the wait for an outgoing daemon.
//
// The budget must comfortably exceed the slowest honest teardown (engine.Close
// of a box with a few hundred outbounds plus the cache-file flush to flash, then
// the nft/ip/sysctl removal) AND the init script's procd term_timeout, which is
// the hard cap on how long a predecessor can live after its SIGTERM. Overshooting
// costs nothing on a healthy box — the wait ends the instant the predecessor is
// gone — while undershooting reintroduces the very overlap this exists to remove.
//
// Variables, not constants, so the tests can drive the loop without sleeping.
var (
predecessorBudget = 60 * time.Second
predecessorPoll = 100 * time.Millisecond
)
// aliveProbe is the seam waitForPredecessor polls. The portable default reports
// "no predecessor", so a non-linux build (where there is no daemon at all) never
// waits; predecessor_linux.go replaces it with the real pidfile probe.
var aliveProbe = func() (int, bool) { return 0, false }
// waitForPredecessor blocks until no OTHER shaterd owns the pidfile, or until the
// budget runs out.
//
// Returns (pid, true) only when a predecessor is STILL alive once the budget
// expires — the genuine "two daemons" error the caller refuses on. A predecessor
// that exits within the budget (the restart case) returns (0, false) and startup
// continues, now guaranteed to be sequenced after its teardown, exactly as it is
// after a manual `stop` + pause + `start`.
func waitForPredecessor(budget, poll time.Duration, logger log.ContextLogger) (int, bool) {
pid, ok := aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
if logger != nil {
logger.Info("a previous shaterd (pid ", pid, ") is still shutting down — ",
"waiting for its teardown to finish before applying a new data plane")
}
deadline := time.Now().Add(budget)
for time.Now().Before(deadline) {
time.Sleep(poll)
pid, ok = aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
}
return pid, true
}
+9
View File
@@ -0,0 +1,9 @@
//go:build linux
package main
// Bind the portable predecessor wait to the real single-owner probe. Split out
// of main.go so waitForPredecessor itself stays build-tag-free and testable on
// any developer host.
func init() { aliveProbe = daemonAlive }
+106
View File
@@ -0,0 +1,106 @@
package main
// B3 regression, daemon half.
//
// `/etc/init.d/shater restart` is `stop; start`, and procd's `stop` is
// asynchronous: the ubus `service delete` returns as soon as SIGTERM has been
// SENT, so `start` re-adds the instance while the outgoing `shaterd run` is
// still executing its honest teardown. The startup guard used to exit(1) the
// moment it saw a live predecessor and rely on procd's `respawn ... 5 ...` to
// try again later — a blind retry that neither knows nor waits for the teardown
// to finish, and that turns every restart into a logged crash plus a five-second
// hole with no data plane.
//
// The contract these pin: startup BLOCKS until the predecessor is gone (so our
// apply is strictly ordered after its teardown, exactly as it is after a manual
// `stop` + pause + `start`), and only refuses when the predecessor outlives the
// whole budget.
import (
"os"
"testing"
"time"
)
// scriptAlive installs an aliveProbe that reports a live predecessor for the
// first n calls and "gone" afterwards, and restores the real probe.
func scriptAlive(t *testing.T, pid, n int) *int {
t.Helper()
orig := aliveProbe
t.Cleanup(func() { aliveProbe = orig })
calls := 0
aliveProbe = func() (int, bool) {
calls++
if calls <= n {
return pid, true
}
return 0, false
}
return &calls
}
// TestWaitForPredecessorWaitsForTeardown: a predecessor that is still tearing
// the plane down must be WAITED for, not refused. Before the fix this returned
// "still alive" on the first probe and the daemon exited 1.
func TestWaitForPredecessorWaitsForTeardown(t *testing.T) {
calls := scriptAlive(t, 4242, 3)
pid, stillAlive := waitForPredecessor(2*time.Second, time.Millisecond, nil)
if stillAlive {
t.Fatalf("a predecessor that exits within the budget must not be refused (pid %d)", pid)
}
if *calls < 4 {
t.Fatalf("expected the guard to keep probing until the predecessor was gone, got %d probes", *calls)
}
}
// TestWaitForPredecessorRefusesAfterBudget: the single-owner invariant is kept —
// a predecessor that never dies still ends in a refusal, it is just no longer
// the FIRST answer.
func TestWaitForPredecessorRefusesAfterBudget(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return 4242, true }
start := time.Now()
pid, stillAlive := waitForPredecessor(30*time.Millisecond, time.Millisecond, nil)
if !stillAlive || pid != 4242 {
t.Fatalf("an immortal predecessor must still be refused, got pid=%d alive=%v", pid, stillAlive)
}
if time.Since(start) < 30*time.Millisecond {
t.Fatalf("the guard must exhaust its budget before refusing")
}
}
// TestWaitForPredecessorNoPredecessorIsFree: the boot path must pay nothing —
// one probe, no sleep. A wait that cost a tick on every clean start would show up
// as a slower boot for no reason.
func TestWaitForPredecessorNoPredecessorIsFree(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
calls := 0
aliveProbe = func() (int, bool) { calls++; return 0, false }
start := time.Now()
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("no predecessor must not be reported as alive")
}
if calls != 1 {
t.Fatalf("expected exactly one probe when nothing is running, got %d", calls)
}
if time.Since(start) > 500*time.Millisecond {
t.Fatalf("the no-predecessor path must not sleep")
}
}
// TestWaitForPredecessorIgnoresOwnPid: a pidfile naming THIS process (a crashed
// predecessor whose pid we were handed, or a re-exec) is not a predecessor.
func TestWaitForPredecessorIgnoresOwnPid(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return os.Getpid(), true }
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("our own pid must never count as a predecessor")
}
}
+37 -21
View File
@@ -610,44 +610,60 @@ func splitDomainMarker(e string) (marker, value string, ok bool) {
// reader can tell a deliberate difference from an oversight (R9.3). Verified
// against the code on both sides, not against upstream docs — this is a fork.
//
// marker domain lists (this file) routing rules (ruleMatchers, route.go)
// There are now only TWO contexts, not three. A routing rule no longer classifies
// domains at all: `dst_domain` was removed in schema v2 and a rule names its
// destination through a `config ruleset`, so every domain entry in the system —
// filter list, device list, inline rule-set — arrives at classifyDomainEntries
// below. The one caller that adds something on top is inlineRulesetRule
// (ruleset.go), which peels `regexp:` off first.
//
// marker DNS filter / device lists inline rule-sets (inlineRulesetRule)
// ---------- ---------------------------- --------------------------------------
// full: yes — exact domain yes — exact domain
// suffix: yes — domain + subdomains yes — domain + subdomains
// .example SAME AS suffix: (dot stripped) SAME AS suffix: (dot stripped)
// keyword: yes — substring yes — substring
// (bare) SUFFIX in filter/device/ EXACT domain
// inline-ruleset lists
// (bare) SUFFIX (domain + subdomains) SUFFIX (domain + subdomains)
// regexp: NO — warns yes — DomainRegex (pattern validated)
// geosite: NO — warns NO — warns, rule matcher omitted
// geosite: NO — warns NO — warns, entry dropped
//
// Two corrections this table used to get wrong, both of the "claimed a behaviour
// that does not exist" kind:
// A THIRD context exists and has NO row here on purpose: the body of a
// `source=url` plain-text list. That is a hosts/one-domain-per-line FILE parsed by
// parseDomainList (ruleset.go), not a list of typed entries, and it understands no
// marker at all — every line is a domain plus its subdomains. An entry written in
// the vocabulary above is dropped there (a domain cannot contain ":") and reported
// per list by warnListEntryVocabulary. The long block above parseDomainList states
// why unifying the two was rejected; do not read this table as covering it.
//
// - `.example.com` was documented as "subdomains only" on both sides. It is not:
// The bare-entry row is now the SAME on both sides, and that uniformity is the
// point of schema v2 — a routing rule used to read a bare entry as an EXACT host
// while every list read it as a suffix, a difference nothing in the UI showed.
// model.migrate1to2 is what preserves the old meaning of existing configs: it
// rewrites a rule's bare `dst_domain` entry as `full:` when it moves it into the
// generated rule-set.
//
// Corrections this table used to get wrong, all of the "claimed a behaviour that
// does not exist" kind:
//
// - `.example.com` was documented as "subdomains only". It is not:
// classifyDomainEntries strips the dot, so it is a synonym of `suffix:` and
// matches the apex too. See the leading-dot branch above for why the synonym
// is kept rather than the distinction implemented.
// - `geosite:` was documented as a working routing matcher. It is not: the
// route-rule geosite field was REMOVED from this engine, so route.go warns and
// omits the matcher (a rule left with no other matcher is skipped entirely).
// Use a `config ruleset` with source=geosite.
// route-rule geosite field was REMOVED from this engine. Use a `config
// ruleset` with source=geosite and category chips.
//
// The two absences in the domain-list column are DELIBERATE, not gaps:
// The two absences in the FILTER column are DELIBERATE, not gaps:
//
// - regexp: an invalid regular expression is only detected when the rule is
// built, where it aborts box.New and takes the whole config down — the exact
// fail-open violation this audit spent its time removing. Supporting it here
// would require compiling and validating every pattern at generate time.
// Warning is the honest answer until that is done.
// - geosite: filter lists and rulesets already express geosite properly, via
// - regexp: an invalid regular expression aborts box.New and takes the whole
// config down, so it may only be accepted where every pattern is compiled and
// validated at generate time. Inline rule-sets do exactly that
// (peelDomainRegexes, ruleset.go), which is why the right-hand column says
// yes; the DNS-filter path has no such validation and warns instead.
// - geosite: filter lists and rule-sets already express geosite properly, via
// `source=geosite` plus category chips, which fetches the official compiled
// .srs. A `geosite:` entry inside an inline list would be a second, weaker
// path to the same feature.
//
// The BARE-entry difference is also deliberate and long-standing: a blocklist
// entry is meant to cover subdomains, while a routing rule's bare entry is an
// exact host. Both are documented at their call sites.
// toASCIIDomain punycodes a unicode domain entry so it can match the punycoded
// names that actually arrive in DNS queries. Lenient by design: an entry idna
+2 -1
View File
@@ -289,8 +289,9 @@ func TestFailoverWarnsOnceAcrossChainCopy(t *testing.T) {
m := twoNodeGroupModel("failover")
m.Nodes = append(m.Nodes, model.Node{Name: "hop", Enabled: true, URI: ss("203.0.113.9")})
m.Chains = []model.Chain{{Name: "ch", Hops: []string{"node:hop", "group:g"}}}
m.Rulesets = []model.Ruleset{inlineDomainSet("ex", "example.com")}
m.Rules = []model.Rule{
{Name: "r", Enabled: true, DstDomain: []string{"example.com"}, Target: "chain:ch"},
{Name: "r", Enabled: true, DstRuleset: []string{"ex"}, Target: "chain:ch"},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
+43 -20
View File
@@ -67,7 +67,7 @@ package generate
import (
"fmt"
"net/netip"
"sort"
"os"
"strconv"
"strings"
"time"
@@ -75,6 +75,7 @@ import (
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/json/badoption"
"github.com/sagernet/sing-box/shater/logsink"
"github.com/sagernet/sing-box/shater/model"
)
@@ -333,13 +334,31 @@ func Warnings(m *model.Model) []string {
// it, so a level there could never be honoured and would only invite the reader to
// think it was. Note that an EMPTY Level on an ENABLED log means LevelTrace, not
// "default" — so Disabled must never be emitted speculatively.
//
// DisableColor is the OTHER half of the engine's log block, and it is not
// cosmetic. log/log.go turns it into Formatter.DisableColors, which is what
// stops aurora from painting the level word and the per-connection ID. Left at
// its zero value it painted them, and the engine's writer is the daemon's shared
// logsink whose syslog half is procd's stderr — so `logread` filled up with
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: exchange failed …
//
// and `grep ERROR` stopped matching. Colour is now emitted only when that
// destination is a terminal (an interactive `shaterd run`); see
// shater/logsink/color.go.
func logOptions(s string) *option.LogOptions {
if logSilenced(s) {
return &option.LogOptions{Disabled: true}
}
return &option.LogOptions{Level: logLevel(s)}
return &option.LogOptions{Level: logLevel(s), DisableColor: !logColorAllowed()}
}
// logColorAllowed reports whether the engine's log lines may carry ANSI colour.
// A package var so tests pin it instead of depending on how the test binary's
// stderr happens to be wired; production always asks the real stderr, which is
// the sink's syslog half.
var logColorAllowed = func() bool { return logsink.IsTTY(os.Stderr) }
// logSilenced reports whether the operator asked for no log at all.
//
// The vocabulary lives in model.SilentLogLevels rather than here because
@@ -487,16 +506,21 @@ func validPortRange(s string) bool {
// the classifier actually drops.
//
// `suffix:` belongs here. The list used to omit it on the theory that suffix:/
// regexp: are route-rule-only and are peeled off by ruleMatchers before the shared
// classifier sees them. That is true of route.go — which handles a bare `suffix:`
// in its own switch and warns there, so this list cannot double-report — but NOT
// of the classifier itself: classifyDomainEntries has a `case "suffix"`, and every
// other caller (devices.go, dnsfilter.go) reaches it directly. A lone `suffix:`
// arriving that way was dropped by add()'s empty-value guard and reported by
// nobody, which is the one outcome this pair of functions exists to prevent.
// regexp: were route-rule-only and were peeled off before the shared classifier
// saw them. classifyDomainEntries has a `case "suffix"`, and every caller
// (devices.go, dnsfilter.go, inlineRulesetRule) reaches it directly, so a lone
// `suffix:` was dropped by add()'s empty-value guard and reported by nobody —
// the one outcome this pair of functions exists to prevent. (The route-rule half
// of that old reasoning is gone entirely: `dst_domain` was removed in schema v2,
// so route.go classifies no domains at all and there is no second warner to
// double-report with.)
//
// `regexp:` is correctly absent: the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix and is reported by unrecognisedDomainPrefix.
// `regexp:` is correctly absent, for two different reasons depending on the
// caller. In a filter/device list the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix reported by unrecognisedDomainPrefix. In an
// inline rule-set peelDomainRegexes (ruleset.go) strips every `regexp:` entry
// BEFORE this predicate runs and reports a valueless one itself, so it can never
// reach here either way.
var domainMarkers = []string{"keyword:", "full:", "suffix:", "."}
// isDomainMarkerOnly reports whether an entry is a bare classification marker
@@ -564,15 +588,14 @@ func networkList(tcp, udp bool) option.NetworkList {
}
}
// sortedByOrder returns rule indices sorted by (Order, original index) so the
// sortedRuleIndices returns rule indices sorted by (Order, original index) so the
// engine's first-match order matches the model's declared order, stably.
//
// The implementation is model.SortedRuleIndices: the reachability analysis that
// decides which rules can never fire has to walk the rules in EXACTLY this order
// to be right, and it lives in model (the leaf both generate and the panel API
// import). Two copies of "what order do rules run in" is precisely the drift that
// would make the warning and the panel badge disagree.
func sortedRuleIndices(rules []model.Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
return model.SortedRuleIndices(rules)
}
+4 -2
View File
@@ -323,8 +323,9 @@ func TestEgressDPISpoofValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12366, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "spf", Type: "direct", DPI: "spoof"}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "spoof-rule", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:spf"},
{Name: "spoof-rule", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:spf"},
},
}
opts, warns, changed := applyAndClose(t, m)
@@ -346,8 +347,9 @@ func TestByedpiEgressValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12367, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: 1080}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:bd"},
{Name: "desync", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:bd"},
},
}
opts, warns, changed := applyAndClose(t, m)
+73
View File
@@ -0,0 +1,73 @@
package generate
import (
"bytes"
"context"
"strings"
"testing"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// TestLogOptionsDisableColorOffTTY pins the engine half of the syslog colour
// leak. The engine's log factory is built by box.New from these options; with
// DisableColor left false, every engine line reached procd's stderr — i.e.
// syslog — wrapped in aurora escapes.
func TestLogOptionsDisableColorOffTTY(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false } // stderr is procd's pipe
t.Cleanup(func() { logColorAllowed = restore })
for _, level := range []string{"", "info", "error", "warning"} {
got := logOptions(level)
if got.Disabled {
t.Fatalf("logOptions(%q) unexpectedly disabled the log", level)
}
if !got.DisableColor {
t.Errorf("logOptions(%q).DisableColor = false; syslog would get ANSI escapes", level)
}
}
// A terminal (an interactive `shaterd run`) keeps its colours.
logColorAllowed = func() bool { return true }
if logOptions("info").DisableColor {
t.Errorf("logOptions on a TTY disabled colour; the gate is supposed to be the destination, not a blanket off")
}
}
// TestGeneratedLogBlockProducesNoANSI is the end-to-end proof: take the log
// block Generate actually emits, build the engine's factory over it exactly the
// way box.New does (log.New with DefaultWriter), and check the bytes.
func TestGeneratedLogBlockProducesNoANSI(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false }
t.Cleanup(func() { logColorAllowed = restore })
m := &model.Model{Globals: model.DefaultGlobals()}
m.Globals.LogLevel = "info"
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("GenerateWithWarnings: %v", err)
}
if opts.Log == nil {
t.Fatal("generated options carry no log block")
}
var out bytes.Buffer
factory, err := log.New(log.Options{Options: option.LogOptions(*opts.Log), DefaultWriter: &out})
if err != nil {
t.Fatalf("log.New over the generated block: %v", err)
}
ctx := log.ContextWithNewID(context.Background())
factory.Logger().ErrorContext(ctx, "dns: exchange failed for example.com. IN AAAA: unexpected EOF")
got := out.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("engine factory wrote nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("engine log line carries an ANSI escape at byte %d: %q", i, got)
}
}
+3 -2
View File
@@ -21,9 +21,10 @@ func TestProfileAppliesCleanly(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12370, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{inlineDomainSet("ads", "ads.example")},
Rules: []model.Rule{
{Name: "lan-proxy", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "node:ss1"},
{Name: "adblock", Enabled: true, Order: 20, DstDomain: []string{"ads.example"}, Target: "block"},
{Name: "adblock", Enabled: true, Order: 20, DstRuleset: []string{"ads"}, Target: "block"},
},
Profiles: []model.Profile{
{Name: "home", Enabled: true, Priority: 1, DisableRules: []string{"adblock"}},
@@ -35,7 +36,7 @@ func TestProfileAppliesCleanly(t *testing.T) {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
// Profile-disabled 'adblock' rule must be absent.
if opts.Route == nil || hasDomainRule(opts.Route, "ads.example") {
if opts.Route == nil || hasRulesetRule(opts.Route, "ads") {
t.Fatalf("profile-disabled rule 'adblock' should not be emitted")
}
// The lan-proxy rule still routes to ss1.
+41 -23
View File
@@ -49,9 +49,13 @@ func TestNoProfilesUnchanged(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(), // kill-switch closed => Final "block"
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("a", "a.example"),
inlineDomainSet("b", "b.example"),
},
Rules: []model.Rule{
{Name: "a", Enabled: true, Order: 10, DstDomain: []string{"a.example"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstDomain: []string{"b.example"}, Target: "node:ss1"},
{Name: "a", Enabled: true, Order: 10, DstRuleset: []string{"a"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstRuleset: []string{"b"}, Target: "node:ss1"},
},
}
rt, b := buildRouteAt(m, wed12UTC)
@@ -70,7 +74,7 @@ func TestNoProfilesUnchanged(t *testing.T) {
if len(rt.Rules) != 4 {
t.Fatalf("route rule count = %d, want 4 (sniff+hijack+2 user)", len(rt.Rules))
}
if !hasDomainRule(rt, "a.example") || !hasDomainRule(rt, "b.example") {
if !hasRulesetRule(rt, "a") || !hasRulesetRule(rt, "b") {
t.Fatalf("both user rules should survive unchanged")
}
}
@@ -83,9 +87,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("blocked", "blocked.example"),
inlineDomainSet("keep", "keep.example"),
},
Rules: []model.Rule{
{Name: "blockme", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstDomain: []string{"keep.example"}, Target: "direct"},
{Name: "blockme", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstRuleset: []string{"keep"}, Target: "direct"},
},
Profiles: []model.Profile{
// Manual pin: honored regardless of conditions. Disables "blockme".
@@ -94,10 +102,10 @@ func TestActiveProfileDisablesRule(t *testing.T) {
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "blocked.example") {
if hasRulesetRule(rt, "blocked") {
t.Fatalf("profile disabled rule 'blockme' but its matcher is still emitted")
}
if !hasDomainRule(rt, "keep.example") {
if !hasRulesetRule(rt, "keep") {
t.Fatalf("rule 'keep' should be untouched")
}
if rt.Final != tagBlock {
@@ -115,9 +123,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
func TestAutoSelectHighestPriority(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
inlineDomainSet("lo", "lo.example"),
inlineDomainSet("hi", "hi.example"),
},
Rules: []model.Rule{
{Name: "rLo", Enabled: true, Order: 10, DstDomain: []string{"lo.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rLo", Enabled: true, Order: 10, DstRuleset: []string{"lo"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "lo", Enabled: true, Priority: 5, DisableRules: []string{"rLo"}},
@@ -125,10 +137,10 @@ func TestAutoSelectHighestPriority(t *testing.T) {
},
}
rt, _ := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("the highest-priority profile must win and disable rHi")
}
if !hasDomainRule(rt, "lo.example") {
if !hasRulesetRule(rt, "lo") {
t.Fatalf("only the winning profile applies (rLo should survive)")
}
}
@@ -142,9 +154,13 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
g.ActiveProfile = active
return &model.Model{
Globals: g,
Rulesets: []model.Ruleset{
inlineDomainSet("hi", "hi.example"),
inlineDomainSet("if", "if.example"),
},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstDomain: []string{"if.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstRuleset: []string{"if"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
@@ -156,19 +172,19 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
// Auto-select (no pin): iface profile skipped (the watcher owns it); 'plain' wins.
rt, _ := buildRouteAt(mk(""), wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("auto-select should apply 'plain' and disable rHi")
}
if !hasDomainRule(rt, "if.example") {
if !hasRulesetRule(rt, "if") {
t.Fatalf("iface profile 'wwan' must be skipped by auto-select (rIf should survive)")
}
// Named explicitly: iface profile honored regardless of its iface condition.
rt, _ = buildRouteAt(mk("wwan"), wed12UTC)
if hasDomainRule(rt, "if.example") {
if hasRulesetRule(rt, "if") {
t.Fatalf("explicitly named iface profile must be honored and disable rIf")
}
if !hasDomainRule(rt, "hi.example") {
if !hasRulesetRule(rt, "hi") {
t.Fatalf("only the named profile applies; rHi should survive")
}
}
@@ -179,14 +195,15 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
g := model.DefaultGlobals()
g.ActiveProfile = "ghost"
m := &model.Model{
Globals: g,
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstDomain: []string{"x.example"}, Target: "direct"}},
Globals: g,
Rulesets: []model.Ruleset{inlineDomainSet("x", "x.example")},
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstRuleset: []string{"x"}, Target: "direct"}},
Profiles: []model.Profile{
{Name: "auto", Enabled: true, Priority: 1, DisableRules: []string{"rX"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "x.example") {
if hasRulesetRule(rt, "x") {
t.Fatalf("fallback auto-select should apply 'auto' and disable rX")
}
if !hasWarning(b, "falling back to auto-select") {
@@ -208,16 +225,17 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
// suppress it or to warn about it.
func TestPlainProfileIsSelectableAfterProbeRemoval(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{inlineDomainSet("hi", "hi.example")},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("a plain profile must apply and disable rHi")
}
// No leftover diagnostic: warning about an option the parser no longer reads
+63 -126
View File
@@ -1,8 +1,6 @@
package generate
import (
"fmt"
"regexp"
"sort"
"strings"
@@ -30,6 +28,12 @@ func (b *builder) buildRoute() *option.RouteOptions {
// never aborts box.New.
b.applyProfiles()
// Report the rules that cannot fire under this config BEFORE building anything,
// so the diagnosis is about the desired state the operator wrote rather than
// about whatever survived generation. Diagnosis only — nothing is renamed,
// reordered or dropped here.
b.warnUnreachableRules()
// Kill-switch default backstop.
final := tagBlock
if strings.EqualFold(strings.TrimSpace(b.m.Globals.KillSwitch), "open") {
@@ -240,25 +244,58 @@ func (b *builder) ruleKillFallback(r model.Rule, want string) string {
// outbound — traffic left over the DEFAULT WAN while the panel (which applies the
// same "a bare Egress is a target too" rule for display) showed `egress:wan2`.
// A silent mis-route on exactly the multi-WAN setups the field exists for.
func effectiveRuleTarget(r model.Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
//
// The implementation is model.EffectiveRuleTarget — shared with the reachability
// analysis, which must resolve a rule's target identically to decide which
// default the router actually ends up using.
func effectiveRuleTarget(r model.Rule) string { return model.EffectiveRuleTarget(r) }
// isCatchAll reports whether a rule carries NO matcher of any kind (src, dst
// domain/ip/ruleset, port, proto). Such a rule is the default egress.
func (b *builder) isCatchAll(r model.Rule) bool {
return len(r.Src) == 0 &&
len(r.DstDomain) == 0 &&
len(r.DstIP) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
//
// The implementation is model.IsCatchAll — shared with the reachability analysis
// (and mirrored by the panel), because "is this rule a default?" is the single
// question both the Final assignment below and the never-fires badge turn on.
func (b *builder) isCatchAll(r model.Rule) bool { return model.IsCatchAll(r) }
// warnUnreachableRules reports every rule that CANNOT take effect under this
// config, whatever the traffic.
//
// Only one shape is certain enough to report (see model.RuleReachability): two
// condition-less rules, where the later one wins because the loop above simply
// overwrites Final. That case is silent today and looks completely healthy — the
// panel drew both as "default route · final" and the log said nothing — so a
// config with `default -> direct` at order 20 and `default -> group:auto` at
// order 100 gave no hint at all that one of the two was doing nothing.
//
// Severity is decided by CONSEQUENCE, via the wording (apply/warnings.go
// classifies generate's text): when the default that actually wins is `direct`
// while the retired rule asked for a tunnel or a block, the operator's default
// policy is not in effect and everything unmatched leaves on the plain WAN — a
// broken protection claim, i.e. critical. Any other combination (a dead `direct`
// under a live tunnel, one tunnel under another) is a dead setting, not a leak:
// a warning.
func (b *builder) warnUnreachableRules() {
for _, rr := range model.RuleReachability(b.effectiveRules) {
if !rr.Unreachable {
continue
}
dead := effectiveRuleTarget(b.effectiveRules[rr.Index])
winner := effectiveRuleTarget(b.effectiveRules[rr.ShadowedByIndex])
if winner == tagDirect && dead != tagDirect {
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so %q wins and this rule's "+
"target %q is never applied. Everything no other rule matches leaves over the plain "+
"WAN with your real IP address. Delete one of the two, or give this one a condition",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
continue
}
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so the default the router uses "+
"is %q and this rule's target %q is never applied. Delete one of the two, or give this "+
"one a condition so it can match something",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
}
}
// ruleMatchers builds the RawDefaultRule matchers the engine can evaluate.
@@ -282,114 +319,14 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
matched = true
}
// Destination domains. The route-rule-specific prefixes (geosite:/regexp:/
// suffix:) are peeled off here; everything else goes through the shared
// classifier in dnsfilter.go with bareIsSuffix=FALSE — a bare entry in a
// ROUTING rule is an exact domain, unlike the DNS/filter lists where it means
// "and all subdomains". classifyDomainEntries is also what drops marker-only
// entries ("." / "full:" / "keyword:"), which must never reach the engine:
// an empty domain/domain_suffix aborts box.New, and an empty domain_keyword
// is strings.Contains(host, "") — a silent match on EVERY host.
var explicitSuffix []string
var plain []string
for _, d := range r.DstDomain {
d = strings.TrimSpace(d)
if d == "" {
continue
}
switch {
case strings.HasPrefix(d, "geosite:"):
// A `geosite:` matcher is NOT emittable: the route-rule geosite field was
// removed in this engine and route/rule.NewDefaultRule hard-errors on it,
// which aborts box.New for the WHOLE config. Mirroring the geoip handling
// below, it is treated as inert: warned and omitted, so a legacy/UCI rule
// carrying one degrades to "this rule does nothing" instead of taking the
// entire tunnel down. Use a `config ruleset` with source=geosite instead.
b.warnf("rule %q: geosite matcher %q is removed from this engine — use a ruleset with source=geosite instead, omitted (inert)", r.Name, d)
case strings.HasPrefix(d, "regexp:"):
// route/rule.NewDomainRegexItem hard-errors on an uncompilable pattern and
// takes box.New with it; validate here and drop the bad one with a warning.
// A BARE "regexp:" compiles fine but matches every host — same silent
// match-all hazard as an empty keyword, so it is dropped too.
re := strings.TrimSpace(strings.TrimPrefix(d, "regexp:"))
if re == "" {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", r.Name, d)
break
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("rule %q: domain regexp %q is invalid (%v), omitted", r.Name, re, err)
break
}
raw.DomainRegex = append(raw.DomainRegex, re)
matched = true
case strings.HasPrefix(d, "suffix:"):
// Label-aware suffix (apex + subdomains): sing-box domain_suffix stored
// in bare form matches both "example.com" and "*.example.com" (see
// sing/common/domain matcher). A leading-dot entry, by contrast, matches
// subdomains ONLY, so the explicit `suffix:` form is how presets/rules
// ask for apex-inclusive suffix matching.
if s := strings.TrimSpace(strings.TrimPrefix(d, "suffix:")); s != "" {
explicitSuffix = append(explicitSuffix, s)
} else {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New)", r.Name, d)
}
default:
plain = append(plain, d)
}
}
domain, suffix, keyword := classifyDomainEntries(plain, false)
for _, d := range plain {
if isDomainMarkerOnly(d) {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword matches EVERY host)", r.Name, d)
}
}
// An unrecognised `word:` prefix is DROPPED by classifyDomainEntries (a domain
// cannot contain ":", so keeping it would load a provably unmatchable literal).
// It has to be reported here or the rule silently loses that destination — the
// v0.1/xray spelling `domain:example.com` is exactly what someone migrating
// writes, and it used to disappear without a trace. geosite:/regexp:/suffix:
// were already peeled off above, so `plain` carries only the shared markers.
b.warnUnrecognisedPrefixes(fmt.Sprintf("rule %q", r.Name), plain)
suffix = append(explicitSuffix, suffix...)
if len(domain) > 0 {
raw.Domain = badoption.Listable[string](domain)
matched = true
}
if len(suffix) > 0 {
raw.DomainSuffix = badoption.Listable[string](suffix)
matched = true
}
if len(keyword) > 0 {
raw.DomainKeyword = badoption.Listable[string](keyword)
matched = true
}
// Destination IPs. A `geoip:<code>` entry is NOT a CIDR: route-rule geoip was
// removed in this engine (box.New hard-errors on it), so — mirroring how the
// ruleset geoip source is handled — it is treated as inert: warned and omitted
// (never emitted as an ip_cidr, which would also abort box.New). This keeps a
// geoip-driven rule/preset (e.g. the ru-bypass pack) fail-open instead of fatal.
var ipcidr []string
for _, ip := range r.DstIP {
ip = strings.TrimSpace(ip)
if ip == "" {
continue
}
if strings.HasPrefix(strings.ToLower(ip), "geoip:") {
b.warnf("rule %q: geoip matcher %q is removed from this engine — use a ruleset with source=geoip instead, omitted (inert)", r.Name, ip)
continue
}
ipcidr = append(ipcidr, ip)
}
// Same fail-open validation as the source list: an unparseable ip_cidr aborts
// box.New for the whole config, so drop it with a warning instead.
ipcidr, badIP := validPrefixes(ipcidr)
for _, s := range badIP {
b.warnf("rule %q: destination %q is not a valid IP/CIDR, omitted", r.Name, s)
}
if len(ipcidr) > 0 {
raw.IPCIDR = badoption.Listable[string](ipcidr)
matched = true
}
// Destination: a rule's ONLY destination matcher is DstRuleset (schema v2).
// The inline `dst_domain` / `dst_ip` lists that used to be classified here are
// gone. A destination list is a `config ruleset` — compiled once into a .srs and
// shared by every rule that references it — so the prefix vocabulary
// (full:/suffix:/keyword:/regexp:, a leading dot) and the geosite/geoip sources
// live in exactly one place (generate/ruleset.go inlineRulesetRule). Existing
// configs were rewritten by model.migrate1to2, which preserves each entry's
// meaning 1:1.
// DstRuleset -> reference the rs-<name> rule-sets materialised by
// buildRoutingRuleSets. A dst_ruleset naming an UNDEFINED ruleset is warned +
+127 -72
View File
@@ -8,6 +8,7 @@ import (
"strings"
"testing"
C "github.com/sagernet/sing-box/constant"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
@@ -22,26 +23,45 @@ func killModel(kill string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "social", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "group:grp", Kill: kill},
{Name: "social", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "group:grp", Kill: kill},
{Name: "catch-tcp", Enabled: true, Order: 20, Proto: "tcp", Target: "direct"},
},
}
}
// domainRuleTarget returns the outbound the rule matching domain routes to.
func domainRuleTarget(rt *option.RouteOptions, domain string) (string, bool) {
for _, r := range generalRules(rt) {
for _, d := range r.DefaultOptions.RawDefaultRule.Domain {
if d == domain {
return r.DefaultOptions.RuleAction.RouteOptions.Outbound, true
}
}
// rulesetRuleTarget returns the outbound the rule referencing the rule-set named
// name routes to. It is the post-schema-v2 replacement for looking a rule up by
// its dst-domain matcher: a rule's destination is a rule_set reference now, so
// the matcher that identifies it is the rs-<name> tag.
func rulesetRuleTarget(rt *option.RouteOptions, name string) (string, bool) {
dr := findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name)
if dr == nil {
return "", false
}
return "", false
return dr.RuleAction.RouteOptions.Outbound, true
}
// soleInlineRule returns the single default headless rule an inline rule-set is
// expected to carry, failing the ASSERTION (rather than panicking on an index) when
// the set turns out to be remote/local or to hold a different rule shape.
func soleInlineRule(t *testing.T, rs option.RuleSet) option.DefaultHeadlessRule {
t.Helper()
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
t.Fatalf("rule-set %q is type %q, not inline — it has no rules in the config to inspect", rs.Tag, rs.Type)
}
if len(rs.InlineOptions.Rules) != 1 {
t.Fatalf("rule-set %q must carry exactly one headless rule, got %d", rs.Tag, len(rs.InlineOptions.Rules))
}
hr := rs.InlineOptions.Rules[0]
if hr.Type != C.RuleTypeDefault && hr.Type != "" {
t.Fatalf("rule-set %q carries a %q headless rule, not a default one", rs.Tag, hr.Type)
}
return hr.DefaultOptions
}
// TestRuleKillDefaultBlocks: kill="" / "default" is fail-closed — the rule is
@@ -54,7 +74,7 @@ func TestRuleKillDefaultBlocks(t *testing.T) {
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("%q: rule must be emitted routing to block, but it was dropped", kill)
}
@@ -76,7 +96,7 @@ func TestRuleKillClosedBlocksHere(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=closed must still emit the rule (warnings %v)", warns)
}
@@ -103,7 +123,7 @@ func TestRuleKillOpenGoesDirect(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=open must still emit the rule (warnings %v)", warns)
}
@@ -122,7 +142,7 @@ func TestRuleKillUnknownBlocks(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != tagBlock {
t.Fatalf("unknown kill policy must block, got %q (ok=%v)", got, ok)
}
@@ -145,7 +165,7 @@ func TestRuleKillPreservesFailClosedInvariant(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("%q: Final = %q, want block", kill, opts.Route.Final)
}
if tgt, ok := domainRuleTarget(opts.Route, "social.example"); ok && tgt == tagDirect {
if tgt, ok := rulesetRuleTarget(opts.Route, "social"); ok && tgt == tagDirect {
t.Fatalf("%q: dead group leaked direct", kill)
}
}
@@ -158,17 +178,18 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "ok", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "node:n1", Kill: kill},
{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "node:n1", Kill: kill},
},
}
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != "n1" {
t.Fatalf("%q: healthy target must win, got %q (ok=%v)", kill, got, ok)
}
@@ -176,29 +197,54 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
}
// --- R4: marker-only domain entries ------------------------------------------
//
// R4 moved with the destination list itself: a rule's domains are an inline
// `config ruleset` now (schema v2), so the marker-only guard has to hold in
// inlineRulesetRule rather than in ruleMatchers. The hazard is unchanged.
// TestRuleMarkerOnlyDomainEntriesDropped: an entry that is nothing but its marker
// must never reach the engine. An empty domain/domain_suffix makes NewDomainItem
// return "empty item is not allowed" and aborts box.New (whole LAN down from one
// stray "."); an empty domain_keyword is SILENT and matches every host.
func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
// TestRulesetMarkerOnlyDomainEntriesDropped: an entry that is nothing but its
// marker must never reach the engine. An empty domain/domain_suffix makes
// NewDomainItem return "empty item is not allowed" and aborts box.New (whole LAN
// down from one stray "."); an empty domain_keyword is SILENT and matches every
// host. As the sole entry it also leaves the rule-set with nothing usable, so the
// list is skipped and the rule referencing it is not emitted — never promoted to
// "matches everything".
func TestRulesetMarkerOnlyDomainEntriesDropped(t *testing.T) {
for _, entry := range []string{".", "full:", "keyword:", "suffix:", "regexp:", " keyword: "} {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{entry}, Target: "node:n1",
})
for _, r := range rt.Rules {
raw := r.DefaultOptions.RawDefaultRule
for _, list := range [][]string{raw.Domain, raw.DomainSuffix, raw.DomainKeyword, raw.DomainRegex} {
for _, v := range list {
if strings.TrimSpace(v) == "" {
t.Fatalf("%q: emitted an EMPTY matcher token", entry)
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m", entry)},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
// Only an INLINE rule-set has its rules in the options at all; a remote or
// local one (a geosite chip, a compiled url list) carries a URL or a path and
// an empty InlineOptions. Indexing Rules[0] unconditionally turned any such
// fixture into an index-out-of-range PANIC instead of a failed assertion, so
// the shape is checked rather than assumed.
for _, rs := range rt.RuleSet {
if rs.Type != C.RuleSetTypeInline && rs.Type != "" {
continue
}
for i, hrule := range rs.InlineOptions.Rules {
if hrule.Type != C.RuleTypeDefault && hrule.Type != "" {
continue // a logical headless rule has no matcher lists of its own
}
hr := hrule.DefaultOptions
for name, list := range map[string][]string{
"domain": hr.Domain, "domain_suffix": hr.DomainSuffix,
"domain_keyword": hr.DomainKeyword, "domain_regex": hr.DomainRegex,
} {
for _, v := range list {
if strings.TrimSpace(v) == "" {
t.Fatalf("%q: rule-set %q rule %d emitted an EMPTY %s token (an empty domain aborts box.New; an empty keyword matches every host)",
entry, rs.Tag, i, name)
}
}
}
}
}
// Sole entry => nothing left to match on => the rule must be skipped, never
// silently promoted to "matches everything".
if _, ok := ruleSetByTag(rt, "rs-m"); ok {
t.Fatalf("%q: a marker-only list must not materialise a rule-set", entry)
}
if n := len(generalRules(rt)); n != 0 {
t.Fatalf("%q: marker-only sole entry must skip the rule, got %d rules", entry, n)
}
@@ -208,46 +254,55 @@ func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
}
}
// TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must survive
// the marker-only ones.
func TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{".", "keyword:", "exact.example", "keyword:ads", ".sub.example", "suffix:apex.example"},
Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d (warnings %v)", len(gen), warns)
// TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must
// survive the marker-only ones.
func TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m",
".", "keyword:", "full:exact.example", "keyword:ads", ".sub.example", "suffix:apex.example")},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-m")
if !ok {
t.Fatalf("the usable entries must keep the list alive (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example] (bare entry in a ROUTING rule is exact)", raw.Domain)
hr := soleInlineRule(t, rs)
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example]", hr.Domain)
}
if len(raw.DomainKeyword) != 1 || raw.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", raw.DomainKeyword)
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", hr.DomainKeyword)
}
if len(raw.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", raw.DomainSuffix)
if len(hr.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", hr.DomainSuffix)
}
if findRouteRuleWithRuleSet(rt, "rs-m") == nil {
t.Fatalf("the rule must be emitted referencing rs-m; rules=%+v", rt.Rules)
}
}
// TestRuleBareEntryIsExactDomain pins the routing convention (bare == exact),
// which deliberately differs from the DNS/filter lists (bare == suffix).
func TestRuleBareEntryIsExactDomain(t *testing.T) {
rt, _ := genRules(t, model.Rule{
Name: "b", Enabled: true, Order: 10,
DstDomain: []string{"example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d", len(gen))
// TestRulesetBareEntryIsDomainSuffix pins the convention a destination list now
// follows — and it is the OPPOSITE of the one the old dst_domain used.
//
// A bare entry in a routing rule meant one EXACT domain; a bare entry in a
// rule-set means the domain AND its subdomains (the DNS/filter/device convention,
// classifyDomainEntries with bareIsSuffix=true). `full:` is how an exact match is
// written now. model.migrate1to2 rewrites old dst_domain entries accordingly, so
// this asymmetry is the thing that migration has to get right.
func TestRulesetBareEntryIsDomainSuffix(t *testing.T) {
rt, _ := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("b", "example.com", "full:exact.example")},
model.Rule{Name: "b", Enabled: true, Order: 10, DstRuleset: []string{"b"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-b")
if !ok {
t.Fatalf("rs-b not emitted; route=%+v", rt)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("bare entry must be an exact Domain, got domain=%v suffix=%v", raw.Domain, raw.DomainSuffix)
hr := soleInlineRule(t, rs)
if len(hr.DomainSuffix) != 1 || hr.DomainSuffix[0] != "example.com" {
t.Fatalf("bare entry must become a domain_suffix, got suffix=%v domain=%v", hr.DomainSuffix, hr.Domain)
}
if len(raw.DomainSuffix) != 0 {
t.Fatalf("bare entry must NOT become a suffix, got %v", raw.DomainSuffix)
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("full: must be the exact form, got domain=%v", hr.Domain)
}
}
+35 -14
View File
@@ -18,10 +18,13 @@ import (
)
// TestMalformedMatchersStillApply drives ONE model carrying every previously
// fatal matcher — a geosite: domain, a bad ip_cidr, a bad source cidr, an
// fatal matcher — a geosite: entry, a bad ip_cidr, a bad source cidr, an
// uncompilable regexp and a malformed port range — plus a healthy rule, through
// engine.Apply. It must come up, the healthy rule must survive, and the
// kill-switch backstop must stay closed.
// kill-switch backstop must stay closed. The destination lists are inline
// `config ruleset`s (schema v2), which is where the domain/IP guards live now;
// a rule-set left with no usable entry is skipped, and so is the rule whose only
// matcher it was.
func TestMalformedMatchersStillApply(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
@@ -29,13 +32,19 @@ func TestMalformedMatchersStillApply(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12395, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("geosite", "geosite:youtube"),
inlineIPSet("badip", "999.1.1.1/24"),
inlineDomainSet("badre", "regexp:*broken("),
inlineDomainSet("ok", "ok.example"),
},
Rules: []model.Rule{
{Name: "geosite", Enabled: true, Order: 1, DstDomain: []string{"geosite:youtube"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstIP: []string{"999.1.1.1/24"}, Target: "node:n1"},
{Name: "geosite", Enabled: true, Order: 1, DstRuleset: []string{"geosite"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstRuleset: []string{"badip"}, Target: "node:n1"},
{Name: "badsrc", Enabled: true, Order: 3, Src: []string{"192.168.0.0/99"}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstDomain: []string{"regexp:*broken("}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstRuleset: []string{"badre"}, Target: "node:n1"},
{Name: "badport", Enabled: true, Order: 5, DstPort: "a-b", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstDomain: []string{"ok.example"}, DstPort: "443", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstRuleset: []string{"ok"}, DstPort: "443", Target: "node:n1"},
},
}
@@ -46,14 +55,21 @@ func TestMalformedMatchersStillApply(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("Final = %q, want block", opts.Route.Final)
}
if !hasDomainRule(opts.Route, "ok.example") {
if !hasRulesetRule(opts.Route, "ok") {
t.Fatalf("the healthy rule must survive alongside the malformed ones")
}
if !hasRouteToOutbound(opts, "n1") {
t.Fatalf("expected a route rule to n1")
}
// Every malformed matcher reported itself rather than failing silently.
for _, want := range []string{"geosite matcher", "is not a valid IP/CIDR", "domain regexp", "is not a valid port/range"} {
for _, want := range []string{
"unrecognised prefix", // geosite: in a domain list
"no usable entries", // ...leaving that list empty
"bad ip_cidr entry", // 999.1.1.1/24
"is not a valid IP/CIDR", // the source cidr
"domain regexp", // regexp:*broken(
"is not a valid port/range", // a-b
} {
if !routeWarnsHave(warns, want) {
t.Fatalf("missing diagnostic %q in %v", want, warns)
}
@@ -155,10 +171,15 @@ func TestRuleKillPoliciesApply(t *testing.T) {
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12398, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "dead", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#dead"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"dead"}}},
Rulesets: []model.Ruleset{
inlineDomainSet("k-closed", "closed.example"),
inlineDomainSet("k-open", "open.example"),
inlineDomainSet("k-default", "default.example"),
},
Rules: []model.Rule{
{Name: "k-closed", Enabled: true, Order: 1, DstDomain: []string{"closed.example"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstDomain: []string{"open.example"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstDomain: []string{"default.example"}, Target: "group:grp"},
{Name: "k-closed", Enabled: true, Order: 1, DstRuleset: []string{"k-closed"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstRuleset: []string{"k-open"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstRuleset: []string{"k-default"}, Target: "group:grp"},
},
}
@@ -166,13 +187,13 @@ func TestRuleKillPoliciesApply(t *testing.T) {
if !changed {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
if got, ok := domainRuleTarget(opts.Route, "closed.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-closed"); !ok || got != tagBlock {
t.Fatalf("kill=closed must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "open.example"); !ok || got != tagDirect {
if got, ok := rulesetRuleTarget(opts.Route, "k-open"); !ok || got != tagDirect {
t.Fatalf("kill=open must emit a rule routed to direct, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "default.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-default"); !ok || got != tagBlock {
t.Fatalf("kill=default must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if opts.Route.Final != tagBlock {
+209 -85
View File
@@ -45,6 +45,40 @@ func genRules(t *testing.T, rules ...model.Rule) (*option.RouteOptions, []string
return opts.Route, warns
}
// ruleModelWithSets is ruleModel plus the `config ruleset` definitions the rules'
// DstRuleset entries point at. A rule's ONLY destination matcher is a rule-set
// (schema v2), so every case below that just needs "some destination the engine
// can match on" declares one here rather than writing an inline domain/IP list.
func ruleModelWithSets(sets []model.Ruleset, rules ...model.Rule) *model.Model {
m := ruleModel(rules...)
m.Rulesets = sets
return m
}
// genRulesWithSets is genRules for a model that also carries rule-sets.
func genRulesWithSets(t *testing.T, sets []model.Ruleset, rules ...model.Rule) (*option.RouteOptions, []string) {
t.Helper()
opts, warns, err := GenerateWithWarnings(ruleModelWithSets(sets, rules...))
if err != nil {
t.Fatalf("Generate: %v", err)
}
return opts.Route, warns
}
// inlineDomainSet builds an inline DOMAIN `config ruleset`.
//
// Mind the convention the move to rule-sets brought with it: a BARE entry here is
// a DomainSuffix (the apex AND its subdomains), whereas the routing rule's old
// dst_domain read a bare entry as an EXACT domain. `full:` is the exact form.
func inlineDomainSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "domain", Source: "inline", Entries: entries}
}
// inlineIPSet builds an inline IP-RANGE `config ruleset`.
func inlineIPSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: entries}
}
func routeWarnsHave(warns []string, substr string) bool {
for _, w := range warns {
if strings.Contains(w, substr) {
@@ -68,69 +102,117 @@ func generalRules(rt *option.RouteOptions) []option.Rule {
// --- geosite: the landmine ---------------------------------------------------
// TestGeositeMatcherIsInertNotFatal: route-rule `geosite` was REMOVED from this
// engine — route/rule.NewDefaultRule returns "geosite database is deprecated ...
// removed in sing-box 1.12.0" for a non-empty Geosite list, and that error aborts
// box.New for the whole config. A `geosite:` entry must therefore be warned and
// omitted (exactly like the geoip matcher below), never emitted.
func TestGeositeMatcherIsInertNotFatal(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "geo", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube"}, Target: "node:n1",
})
// TestGeositeEntryInRulesetIsInertNotFatal: route-rule `geosite` was REMOVED from
// this engine — route/rule.NewDefaultRule returns "geosite database is deprecated
// ... removed in sing-box 1.12.0" for a non-empty Geosite list, and that error
// aborts box.New for the whole config. Destinations now live in a `config
// ruleset`, so a `geosite:` entry lands in an inline domain list, where it is an
// unrecognised `word:` prefix: dropped by the classifier, reported, and — being
// the list's only entry — leaving the rule-set with nothing to match, so it is
// skipped and the rule that referenced it is not emitted either. Nothing about
// that path may ever put a value in RawDefaultRule.Geosite. (`source=geosite` on
// the ruleset itself is the working way to route a category; see
// TestRoutingRuleSetGeositeCategory.)
func TestGeositeEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("geo", "geosite:youtube")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.Geosite) > 0 {
t.Fatalf("geosite must never be emitted (aborts box.New), got %v", r.DefaultOptions.RawDefaultRule.Geosite)
}
}
if !routeWarnsHave(warns, "geosite matcher") {
t.Fatalf("expected an inert-geosite warning, got %v", warns)
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a rule-set with no usable entry must not be emitted (an empty list would match everything)")
}
if findRouteRuleWithRuleSet(rt, "rs-geo") != nil {
t.Fatalf("no rule may reference the skipped rule-set")
}
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("expected an unrecognised-prefix warning for geosite:, got %v", warns)
}
if !routeWarnsHave(warns, "no usable entries") {
t.Fatalf("expected a no-usable-entries warning, got %v", warns)
}
}
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule carrying BOTH a geosite entry
// and a real domain keeps the real matcher and still routes — only the geosite
// part is dropped.
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule-set carrying BOTH a geosite
// entry and a real domain keeps the real matcher, and the rule referencing it
// still routes — only the geosite entry is dropped.
func TestGeositeMixedWithRealDomainKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube", "example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d (warnings %v)", len(gen), warns)
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mixed", "geosite:youtube", "full:example.com")},
model.Rule{Name: "mixed", Enabled: true, Order: 10, DstRuleset: []string{"mixed"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-mixed")
if !ok {
t.Fatalf("rs-mixed must survive the geosite entry (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Geosite) != 0 {
t.Fatalf("geosite leaked: %v", raw.Geosite)
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", hr)
}
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", raw.Domain)
if len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.DomainRegex) != 0 {
t.Fatalf("geosite: must be dropped, not reinterpreted: %+v", hr)
}
if got := gen[0].DefaultOptions.RuleAction.RouteOptions.Outbound; got != "n1" {
dr := findRouteRuleWithRuleSet(rt, "rs-mixed")
if dr == nil {
t.Fatalf("the rule must be emitted referencing rs-mixed; rules=%+v", rt.Rules)
}
if got := dr.RuleAction.RouteOptions.Outbound; got != "n1" {
t.Fatalf("target = %q, want n1", got)
}
}
// TestGeoipEntryInRulesetIsInertNotFatal is the IP-side twin: model.migrate1to2
// moves an old `dst_ip geoip:ru` verbatim into an inline type=ipcidr rule-set
// (deliberately — see TestMigrate1to2KeepsGeoMarkersInert: it must not silently
// become a working geoip source, because the operator never asked to download
// anything). Here it is an unparseable prefix, so it must be dropped LOUDLY
// rather than reach NewIPCIDRItem, which errors and aborts box.New.
func TestGeoipEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("geo", "geoip:ru")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.GeoIP) > 0 {
t.Fatalf("geoip must never be emitted, got %v", r.DefaultOptions.RawDefaultRule.GeoIP)
}
}
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a list whose only entry is unparseable must not materialise")
}
if !routeWarnsHave(warns, `bad ip_cidr entry "geoip:ru"`) {
t.Fatalf("expected a bad-ip_cidr warning naming the entry, got %v", warns)
}
}
// --- malformed matchers that used to abort box.New ---------------------------
// TestBadDstCIDRWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. It must be dropped.
func TestBadDstCIDRWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstIP: []string{"999.1.1.1/24", "198.51.100.0/24"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadIPCIDREntryWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. Destination addresses are
// an inline `type=ipcidr` rule-set now, so the guard lives in inlineRulesetRule:
// the typo is dropped, the valid entry survives and the rule still routes.
func TestRulesetBadIPCIDREntryWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("bad", "999.1.1.1/24", "198.51.100.0/24")},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one bad entry must not take the whole list down (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.IPCIDR) != 1 || raw.IPCIDR[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", raw.IPCIDR)
got := rs.InlineOptions.Rules[0].DefaultOptions.IPCIDR
if len(got) != 1 || got[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", got)
}
if !routeWarnsHave(warns, `destination "999.1.1.1/24" is not a valid IP/CIDR`) {
t.Fatalf("expected a bad-destination warning, got %v", warns)
if !routeWarnsHave(warns, `bad ip_cidr entry "999.1.1.1/24"`) {
t.Fatalf("expected a bad-ip_cidr warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
@@ -152,24 +234,30 @@ func TestBadSrcCIDRWarnsAndSkips(t *testing.T) {
}
}
// TestBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New.
func TestBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstDomain: []string{"regexp:*broken(", `regexp:^ok\.example$`}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New. The pattern vocabulary
// moved into the inline rule-set with the rest of the destination list, so the
// validation moved with it (peelDomainRegexes): the broken pattern is dropped and
// the compilable one survives.
func TestRulesetBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("bad", "regexp:*broken(", `regexp:^ok\.example$`)},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one broken pattern must not take the whole list down (warnings %v)", warns)
}
got := gen[0].DefaultOptions.RawDefaultRule.DomainRegex
got := rs.InlineOptions.Rules[0].DefaultOptions.DomainRegex
if len(got) != 1 || got[0] != `^ok\.example$` {
t.Fatalf("domain_regex = %v, want only the compilable one", got)
}
if !routeWarnsHave(warns, "domain regexp") {
t.Fatalf("expected a bad-regexp warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
// TestBadPortRangeWarnsAndSkips: a malformed range reaches
@@ -911,10 +999,12 @@ func TestRuleProtoKnownValuesAreSilent(t *testing.T) {
// that would break existing configs either open or closed), but the widening is
// now reported with its consequence.
func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstDomain: []string{"youtube.com"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "applies to EVERY client") {
t.Fatalf("expected a rule-widening warning, got %v", warns)
}
@@ -930,11 +1020,13 @@ func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
// TestMixedSourceRuleReportsTheDroppedHalf: with one usable IP source alongside
// an unmatchable one, the rule narrows to the IP source only.
func TestMixedSourceRuleReportsTheDroppedHalf(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstDomain: []string{"youtube.com"}, Target: "block",
})
_, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "could not be used") {
t.Fatalf("expected a dropped-source warning, got %v", warns)
}
@@ -1041,41 +1133,73 @@ func TestRuleTargetKindsResolve(t *testing.T) {
}
}
// TestRuleDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped by
// the shared classifier (a domain cannot contain ":"), so the rule silently lost
// that destination. `domain:example.com` is the v0.1/xray spelling a migrating
// user writes, and it must not disappear without a trace.
func TestRuleDomainUnrecognisedPrefixWarns(t *testing.T) {
// TestRulesetDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped
// by the shared classifier (a domain cannot contain ":"), so the list silently
// lost that destination. `domain:example.com` is the v0.1/xray spelling a
// migrating user writes, and it must not disappear without a trace — the more so
// now that a destination list is ALWAYS a rule-set, i.e. the one place a typo can
// hide.
func TestRulesetDomainUnrecognisedPrefixWarns(t *testing.T) {
for _, entry := range []string{"domain:example.com", "regex:example.com", "ext:foo.dat:cn"} {
rt, warns := genRules(t, model.Rule{
Name: "mig", Enabled: true, Order: 10,
DstDomain: []string{entry, "keep.example"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mig", entry, "keep.example")},
model.Rule{Name: "mig", Enabled: true, Order: 10, DstRuleset: []string{"mig"}, Target: "block"},
)
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("%q: expected an unrecognised-prefix warning, got %v", entry, warns)
}
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("%q: want 1 rule, got %d", entry, len(gen))
rs, ok := ruleSetByTag(rt, "rs-mig")
if !ok {
t.Fatalf("%q: the usable entry must keep the rule-set alive; warnings %v", entry, warns)
}
for _, d := range gen[0].DefaultOptions.RawDefaultRule.Domain {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
hr := rs.InlineOptions.Rules[0].DefaultOptions
for _, list := range [][]string{hr.Domain, hr.DomainSuffix, hr.DomainKeyword, hr.DomainRegex} {
for _, d := range list {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
}
}
}
if findRouteRuleWithRuleSet(rt, "rs-mig") == nil {
t.Fatalf("%q: the rule must still be emitted; rules=%+v", entry, rt.Rules)
}
}
}
// TestRuleDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary routing rules really support.
func TestRuleDomainKnownPrefixesAreSilent(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "ok", Enabled: true, Order: 10, Target: "block",
DstDomain: []string{"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example"},
})
// TestRulesetDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary an inline domain rule-set really supports.
//
// `regexp:` is the load-bearing case: the shared classifier has no branch for it,
// so it WOULD be reported as an unknown prefix — inlineRulesetRule peels the
// regexes off first (peelDomainRegexes) precisely so it is not. A regression there
// would both warn about a working matcher and drop it.
func TestRulesetDomainKnownPrefixesAreSilent(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("ok",
"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example")},
model.Rule{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"ok"}, Target: "block"},
)
if routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("the supported prefixes must not warn: %v", warns)
}
rs, ok := ruleSetByTag(rt, "rs-ok")
if !ok {
t.Fatalf("rs-ok not emitted; warnings %v", warns)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "a.example" {
t.Fatalf("full: must be an exact Domain, got %v", hr.Domain)
}
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^d\.` {
t.Fatalf("regexp: must survive as a domain_regex matcher, got %v", hr.DomainRegex)
}
// suffix:, the leading dot and the BARE entry all collapse to domain_suffix.
if len(hr.DomainSuffix) != 3 {
t.Fatalf("domain_suffix = %v, want b/e/f.example (bare entry is a suffix in a rule-set)", hr.DomainSuffix)
}
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "c" {
t.Fatalf("domain_keyword = %v, want [c]", hr.DomainKeyword)
}
}
// TestDeviceDomainUnrecognisedPrefixWarns is the same guarantee in the place it
+173
View File
@@ -0,0 +1,173 @@
// B1: a routing rule that can never fire, reported instead of applied silently.
//
// The field config that prompted this had TWO rules named `default`, both with
// zero conditions — order 20 -> direct and order 100 -> group:auto. buildRoute
// points route Final at a condition-less rule and moves on, so the LAST one wins
// and the other is a dead setting. Nothing anywhere said so: the log was clean and
// the panel drew both rows identically.
//
// These tests pin the diagnosis AND the behaviour it describes, because a warning
// that disagrees with what the generator actually does is worse than none.
package generate
import (
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// warnAbout returns the warnings mentioning `rule "name"`.
func warnAbout(warns []string, name string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, `rule "`+name+`"`) {
out = append(out, w)
}
}
return out
}
func warnsMatching(warns []string, substr string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, substr) {
out = append(out, w)
}
}
return out
}
const neverApplied = "is never applied"
// twoDefaultsModel reproduces the field config: two condition-less rules sharing
// the name `default`, differing only in Order and target.
func twoDefaultsModel(lowTarget, highTarget string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: lowTarget},
{Name: "default", Enabled: true, Order: 100, Target: highTarget},
},
}
}
// TestTwoCatchAllRulesWarnAndLastWins is the B1 regression. The order-100 rule is
// the default the engine uses (route Final), and the order-20 one is reported as
// never applied — naming the rule that supersedes it.
func TestTwoCatchAllRulesWarnAndLastWins(t *testing.T) {
opts, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the LAST catch-all's target %q", got, "auto")
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
// It must name the superseding rule AND its order — with both rules called
// `default`, the order is the only thing that tells the two apart.
// Targets are quoted as the OPERATOR wrote them (`group:auto`), not as the
// engine tag they resolve to (`auto`) — the warning has to be readable next to
// the config, not next to the generated JSON.
for _, want := range []string{`rule "default"`, "order 100", `"group:auto"`, `"direct"`} {
if !strings.Contains(got[0], want) {
t.Fatalf("warning must mention %s, got: %s", want, got[0])
}
}
// Diagnosis only: neither rule is renamed, reordered or dropped.
if len(opts.Route.Rules) == 0 {
t.Fatal("route rules disappeared")
}
}
// TestTwoCatchAllsDirectWinsIsCritical: when the surviving default is `direct` and
// the retired one asked for a tunnel, everything unmatched leaves on the plain
// WAN. The warning must say so in the words apply/warnings.go grades critical —
// this is the case where a healthy-looking panel is a lie.
func TestTwoCatchAllsDirectWinsIsCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("group:auto", "direct"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
const marker = "leaves over the plain WAN with your real IP address"
if !strings.Contains(got[0], marker) {
t.Fatalf("a retired tunnel default under a live direct default must carry %q, got: %s", marker, got[0])
}
}
// TestTwoCatchAllsTunnelWinsIsNotCritical: the field config's actual shape — the
// dead rule is `direct` and the live default is the tunnel. That is a dead
// setting, not a leak, so it must NOT carry the critical marker.
func TestTwoCatchAllsTunnelWinsIsNotCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
for _, w := range warnsMatching(warns, neverApplied) {
if strings.Contains(w, "leaves over the plain WAN with your real IP address") {
t.Fatalf("a dead direct default under a live tunnel default is not a leak: %s", w)
}
}
}
// TestSingleCatchAllDoesNotWarn: the ordinary config — specific rules plus ONE
// default — must stay silent. A badge on a working rule teaches the operator to
// ignore the badge.
func TestSingleCatchAllDoesNotWarn(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "ads", Enabled: true, Order: 10, DstPort: "443", Target: "block"},
// A specific rule ordered BELOW the default: still emitted ahead of Final,
// so it is not retired either.
{Name: "default", Enabled: true, Order: 20, Target: "group:auto"},
{Name: "late", Enabled: true, Order: 900, DstPort: "8080", Target: "direct"},
},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a config with one default must not report anything never-applied, got: %q", got)
}
if got := warnAbout(warns, "late"); len(got) != 0 {
t.Fatalf("a specific rule below the default is not shadowed by it, got: %q", got)
}
}
// TestProfileDisabledCatchAllDoesNotShadow: the active profile switches the later
// default off, so the earlier one is the live default and must not be badged.
// Judging the raw config would blame the wrong rule on every profile router.
func TestProfileDisabledCatchAllDoesNotShadow(t *testing.T) {
m := twoDefaultsModel("group:auto", "direct")
m.Rules[1].Name = "fallback" // profiles address rules by name
m.Globals.ActiveProfile = "home"
m.Profiles = []model.Profile{{Name: "home", Enabled: true, DisableRules: []string{"fallback"}}}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a profile-disabled default shadows nothing, got: %q", got)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the surviving default's target %q", got, "auto")
}
}
+231 -7
View File
@@ -25,6 +25,7 @@ import (
"net/netip"
"os"
"path/filepath"
"regexp"
"runtime/debug"
"strings"
"sync"
@@ -500,6 +501,47 @@ func ruleSetURLIsEngineNative(rawURL string) (format string, native bool) {
}
}
// TWO WAYS TO WRITE A DESTINATION LIST, AND WHY THEY DO NOT SHARE A VOCABULARY.
//
// D21 promises "one destination mechanism, one vocabulary". The vocabulary half of
// that promise is about the ENTRIES AN OPERATOR TYPES, and those live in exactly
// one place: an inline `config ruleset` (inlineRulesetRule -> peelDomainRegexes +
// classifyDomainEntries), where `full:` / `suffix:` / `keyword:` / `regexp:` / a
// leading dot / a bare name all mean what docs-shater/DECISIONS.md D21 says.
//
// A `source=url` list that is not engine-native is NOT another spelling of that.
// It is a FILE FORMAT — the hosts / plain-domain / AdBlock-ish text third parties
// publish — and parseDomainList is a parser for that format, not for our entry
// vocabulary. `source=file` is a third thing again: a compiled .srs or a rule-set
// .json handed straight to the engine, which never sees shater's entry syntax at
// all. (An earlier review read this as "the same list written two ways behaves
// differently"; it is closer to "a typed list and a downloaded file are different
// artifacts". The diagnostics below exist so an operator never has to guess which
// one they are looking at.)
//
// Unifying them was considered and REJECTED, on three grounds:
//
// - The formats collide. A real AdGuard/OISD list is full of colon-bearing lines
// that are not our markers at all (`example.com##.banner:has(...)`, `$domain=`
// options, absolute URL rules). Feeding those through the marker classifier
// would either mis-import them or, if we reported every `word:`-shaped token as
// an unknown prefix, drown the operator in hundreds of warnings per list — a
// louder dishonesty than the quiet one it replaces.
// - The shapes collide. A hosts line carries SEVERAL names ("127.0.0.1 a.com
// b.com"), so this parser works per TOKEN; the entry vocabulary works per LINE
// and allows a space after the marker ("keyword: ads"). There is no split rule
// that serves both.
// - `regexp:` from a URL is regex supplied by a third party, compiled into the
// router's matcher and evaluated per query on a 512 MB box. The inline path can
// accept it because the operator typed it; a downloaded list is not that.
//
// So the difference STAYS, and is paid for in diagnostics instead: a text list that
// carries our marker vocabulary is reported per list (see listEntryMarker and
// warnListEntryVocabulary), naming the entries and where they DO work. The check is
// narrow on purpose — only the four markers D21 defines, never the general `word:`
// shape — so it fires on an operator's mistake and stays silent on ordinary filter
// syntax.
//
// parseDomainList extracts domains from the formats public blocklists ship in:
//
// - HOSTS "0.0.0.0 ads.example.com", "127.0.0.1 a.com b.com"
@@ -517,8 +559,12 @@ func ruleSetURLIsEngineNative(rawURL string) (format string, native bool) {
// de-duplication map is kept — domain.NewMatcher already de-duplicates internally
// while building the succinct set, so a second map would just double the largest
// allocation in the pipeline. See listMaxDomains for the measured budget.
func parseDomainList(content []byte) []string {
//
// The second return reports the inline-vocabulary entries seen on the way past, so
// the caller can say so instead of dropping them without a word.
func parseDomainList(content []byte) ([]string, listMarkerNote) {
out := make([]string, 0, 4096)
var note listMarkerNote
scanner := bufio.NewScanner(bytes.NewReader(content))
// Public lists are one domain per line; 64 KiB is far beyond any real line, and
// an over-long line is skipped rather than aborting the parse.
@@ -544,17 +590,123 @@ func parseDomainList(content []byte) []string {
// AdBlock-ish "||domain^" -> domain.
f = strings.TrimPrefix(f, "||")
f = strings.TrimSuffix(f, "^")
if _, isMarker := listEntryMarker(f); isMarker {
// normaliseListDomain would drop this silently (a domain cannot contain
// ":"). Record it so the caller can name it; it is the one class of junk
// in a text list that is provably an operator mistake rather than filter
// syntax we simply do not import.
note.record(f)
continue
}
d, ok := normaliseListDomain(f)
if !ok {
continue
}
out = append(out, d)
if len(out) >= listMaxDomains {
return out
return out, note
}
}
}
return out
return out, note
}
// listEntryVocabulary is EXACTLY the marker set an INLINE rule-set entry may use
// (D21). It is deliberately NOT the general `word:` shape unrecognisedDomainPrefix
// tests for: a published filter list legitimately contains hundreds of colon-
// bearing tokens, and reporting those would make the diagnostic useless. These
// four, by contrast, appear in a downloaded text list only when a human wrote them
// there expecting shater to honour them.
var listEntryVocabulary = []string{"full:", "suffix:", "keyword:", "regexp:"}
// listEntryMarker reports whether a text-list token is written in the inline entry
// vocabulary, and which marker it used.
func listEntryMarker(token string) (string, bool) {
lower := strings.ToLower(strings.TrimSpace(token))
for _, m := range listEntryVocabulary {
if strings.HasPrefix(lower, m) {
return m, true
}
}
return "", false
}
// listMarkerSamples bounds how many offending entries a warning quotes. A list is
// remote content: it must not be able to write an unbounded amount into our log.
const listMarkerSamples = 5
// listMarkerNote records the inline-vocabulary entries one plain-text list carried.
type listMarkerNote struct {
Count int
Samples []string
}
func (n *listMarkerNote) record(entry string) {
n.Count++
if len(n.Samples) < listMarkerSamples {
n.Samples = append(n.Samples, strings.TrimSpace(entry))
}
}
func (n listMarkerNote) empty() bool { return n.Count == 0 }
// listMarkerMemo remembers, per list URL, what the last COMPILATION of that list
// found. Without it the diagnostic would exist for exactly one reconcile — the one
// that happened to refresh the artifact — and then vanish for a whole
// update_interval, which is precisely the "reported to nobody" failure it is meant
// to fix. Same shape (and same reasoning) as ruleSetProbeCache above; a refresh
// that finds nothing clears the entry, so fixing the list silences it.
var (
listMarkerMu sync.Mutex
listMarkerMemo = map[string]listMarkerNote{}
)
func rememberListMarkers(url string, note listMarkerNote) {
listMarkerMu.Lock()
if note.empty() {
delete(listMarkerMemo, url)
} else {
listMarkerMemo[url] = note
}
listMarkerMu.Unlock()
}
func recallListMarkers(url string) listMarkerNote {
listMarkerMu.Lock()
defer listMarkerMu.Unlock()
return listMarkerMemo[url]
}
// resetListMarkerMemo clears the memo (tests).
func resetListMarkerMemo() {
listMarkerMu.Lock()
listMarkerMemo = map[string]listMarkerNote{}
listMarkerMu.Unlock()
}
// warnListEntryVocabulary tells the operator that entries written in the INLINE
// entry vocabulary were found in a downloaded TEXT list, where they mean nothing.
// See the parseDomainList block above for why the two vocabularies are separate
// and why saying so is the whole of the fix.
func (b *builder) warnListEntryVocabulary(diag, url string) {
note := recallListMarkers(url)
if note.empty() {
return
}
b.warnf("%s: %q is a plain-text list (a hosts file or one domain per line), but %d of its entries are written in the "+
"inline rule-set vocabulary (e.g. %s) — a plain-text list has NO markers, so every line is read as a domain plus its "+
"subdomains and anything containing \":\" is dropped, because a domain name cannot contain one. Those entries match NOTHING. "+
"Put them in a rule-set with source=inline, which is the one place full:/suffix:/keyword:/regexp: are honoured.",
diag, url, note.Count, quoteList(note.Samples))
}
// quoteList renders sample entries for a diagnostic.
func quoteList(in []string) string {
out := make([]string, 0, len(in))
for _, s := range in {
out = append(out, fmt.Sprintf("%q", s))
}
return strings.Join(out, ", ")
}
// hostsBoilerplate are the names every hosts file carries for its own bookkeeping.
@@ -675,6 +827,11 @@ func (b *builder) compiledListRuleSet(tag, url, updateInterval, diag string) (op
}
}
// Reported on EVERY generate, not only on the one that refreshed the artifact:
// an entry that matches nothing is exactly as wrong the day after it was
// compiled as the moment it was.
b.warnListEntryVocabulary(diag, url)
return option.RuleSet{
Type: C.RuleSetTypeLocal,
Tag: tag,
@@ -716,8 +873,11 @@ func (b *builder) refreshCompiledList(path, url, diag string) error {
if err != nil {
return err
}
domains := parseDomainList(body)
domains, markers := parseDomainList(body)
body = nil // release the source text before the matcher allocates
// Remember (or clear) what this compilation saw, so the diagnostic survives the
// reconciles that do no I/O at all. See warnListEntryVocabulary.
rememberListMarkers(url, markers)
if len(domains) == 0 {
return fmt.Errorf("no usable domains found at %s (fetched %s, but nothing in it parsed as a domain)", url, "the file")
@@ -1163,7 +1323,8 @@ func (b *builder) buildRoutingRuleSetRaw(rs model.Ruleset) ([]option.RuleSet, []
// its Entries, keyed by Type: an ipcidr ruleset fills ip_cidr; a domain ruleset
// (the default) is classified with inlineDomainRule — bare entry => DomainSuffix
// (so subdomains match), full: => Domain, keyword: => DomainKeyword, . => suffix
// — the same classification the DNS filter uses. ok=false when nothing usable.
// — the same classification the DNS filter uses, plus `regexp:` (see
// peelDomainRegexes). ok=false when nothing usable.
func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRule, bool) {
b.warnUnknownRuleSetType(fmt.Sprintf("ruleset %q", rs.Name), rs.Type)
if ruleSetTypeIsIPCIDR(rs.Type) {
@@ -1190,10 +1351,73 @@ func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRul
return option.DefaultHeadlessRule{IPCIDR: badoption.Listable[string](cidrs)}, true
}
// "domain" (and empty, and anything unrecognised => domain, warned above).
// `regexp:` is peeled off first: it is a routing-rule matcher the DNS-filter
// classifier does not know, and it must not be reported as an unknown prefix.
diag := fmt.Sprintf("ruleset %q", rs.Name)
rest, regexes := b.peelDomainRegexes(diag, rs.Entries)
// R4, on the path every destination list now takes. classifyDomainEntries
// DROPS an entry that is nothing but its marker (".", "full:", "keyword:"),
// silently — and the silence is the dangerous half: an empty domain token
// aborts box.New for the whole config, and an empty keyword is
// strings.Contains(host, "") i.e. EVERY host. The drop is right; not saying so
// is not. (devices.go reports the same class for a device's own lists; the bare
// `regexp:` form is reported by peelDomainRegexes above, which is why it is
// peeled off before this loop and cannot be double-reported.)
for _, e := range rest {
if isDomainMarkerOnly(e) {
b.warnf("%s: entry %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword would match EVERY host)", diag, strings.TrimSpace(e))
}
}
// Only the DOMAIN branch reports unknown `word:` prefixes — the ipcidr branch
// above is full of legitimate colons (IPv6) and must never be checked (R9.2).
b.warnUnrecognisedPrefixes(fmt.Sprintf("ruleset %q", rs.Name), rs.Entries)
return inlineDomainRule(rs.Entries)
b.warnUnrecognisedPrefixes(diag, rest)
hr, ok := inlineDomainRule(rest)
if len(regexes) > 0 {
hr.DomainRegex = badoption.Listable[string](regexes)
ok = true
}
return hr, ok
}
// peelDomainRegexes splits `regexp:<pattern>` entries out of a domain rule-set's
// entry list, returning the remaining entries and the validated patterns.
//
// It exists because a destination list is now ALWAYS a rule-set (schema v2), so
// every matcher a `dst_domain` used to express has to be expressible here —
// including the regex form, which the shared DNS-filter classifier
// (classifyDomainEntries) deliberately does not know about. Validation mirrors
// what the routing rule did before the move, and for the same reason:
// route/rule.NewDomainRegexItem returns an error for an uncompilable pattern and
// that aborts box.New for the WHOLE config, so a bad pattern must degrade to a
// warning. A BARE `regexp:` compiles fine but matches every host — the same
// silent match-all hazard as an empty keyword — so it is dropped too.
//
// SCOPE: this is the INLINE entry path only, and deliberately so. A `source=url`
// text list is a hosts/plain-domain FILE, parsed by parseDomainList, which has no
// marker vocabulary at all — see the block above parseDomainList for why the two
// are not unified and how an entry written in the wrong one is reported.
func (b *builder) peelDomainRegexes(diag string, entries []string) (rest, regexes []string) {
for _, e := range entries {
e = strings.TrimSpace(e)
if e == "" {
continue
}
if !strings.HasPrefix(strings.ToLower(e), "regexp:") {
rest = append(rest, e)
continue
}
re := strings.TrimSpace(e[len("regexp:"):])
if re == "" {
b.warnf("%s: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", diag, e)
continue
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("%s: domain regexp %q is invalid (%v), omitted", diag, re, err)
continue
}
regexes = append(regexes, re)
}
return rest, regexes
}
// Ruleset.Type — the two shapes a rule-set can match, and the accepted spellings.
+55 -3
View File
@@ -33,6 +33,14 @@ func findRouteRuleWithRuleSet(rt *option.RouteOptions, tag string) *option.Defau
return nil
}
// hasRulesetRule reports whether any emitted route rule references the rule-set
// named name (tag rs-<name>). Since schema v2 a rule's destination is ALWAYS a
// rule-set reference, so this is how a test says "that rule was emitted" — the
// former "does any rule carry this dst domain" question has no answer any more.
func hasRulesetRule(rt *option.RouteOptions, name string) bool {
return findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name) != nil
}
func ruleSetByTag(rt *option.RouteOptions, tag string) (option.RuleSet, bool) {
if rt == nil {
return option.RuleSet{}, false
@@ -117,6 +125,50 @@ func TestRoutingRuleSetInlineDomain(t *testing.T) {
}
}
// TestRoutingRuleSetInlineDomainRegex: `regexp:` is a matcher the shared domain
// classifier does NOT know — it belongs to the routing plane, and it used to be
// peeled off inside ruleMatchers, which no longer sees any domains at all. It
// therefore had to move into the inline rule-set with the rest of the destination
// vocabulary (peelDomainRegexes), or every migrated `regexp:` entry would have
// been reported as an unknown prefix and silently dropped: a routing rule that
// looks configured and matches nothing.
func TestRoutingRuleSetInlineDomainRegex(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
{Name: "ads", Type: "domain", Source: "inline", Entries: []string{`regexp:^ads\.`}},
},
Rules: []model.Rule{
{Name: "block-ads", Enabled: true, Order: 10, DstRuleset: []string{"ads"}, Target: "block"},
},
}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if len(warns) != 0 {
t.Fatalf("a valid regexp: entry must not warn: %v", warns)
}
rs, ok := ruleSetByTag(opts.Route, "rs-ads")
if !ok {
t.Fatalf("a regexp-only rule-set must still materialise; route=%+v", opts.Route)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^ads\.` {
t.Fatalf("domain_regex = %+v, want [^ads\\.]", hr.DomainRegex)
}
if len(hr.Domain)+len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.IPCIDR) != 0 {
t.Fatalf("the regexp entry must not leak into another matcher: %+v", hr)
}
dr := findRouteRuleWithRuleSet(opts.Route, "rs-ads")
if dr == nil {
t.Fatalf("no route rule references rs-ads; rules=%+v", opts.Route.Rules)
}
if dr.RuleAction.RouteOptions.Outbound != tagBlock {
t.Fatalf("route rule must route to %q, got %+v", tagBlock, dr.RuleAction)
}
}
// TestRoutingRuleSetInlineIPCIDR: an ipcidr ruleset fills ip_cidr (not domain*),
// and the route rule references it.
func TestRoutingRuleSetInlineIPCIDR(t *testing.T) {
@@ -1501,7 +1553,7 @@ func TestURLBlocklistNotRefetchedWhileFresh(t *testing.T) {
// wildcards, IPs and bare labels can never match a domain query, so importing them
// would be a silent dud (the R9 lesson applied to fetched content).
func TestParseDomainListRejectsJunk(t *testing.T) {
got := parseDomainList([]byte(strings.Join([]string{
got, _ := parseDomainList([]byte(strings.Join([]string{
"good.example.com",
"*.wildcard.example", // wildcard syntax
"/regex/", // regex rule
@@ -1532,12 +1584,12 @@ func TestParseDomainListRejectsJunk(t *testing.T) {
// TestParseDomainListHostsEdgeCases covers the messy real-world shapes.
func TestParseDomainListHostsEdgeCases(t *testing.T) {
got := parseDomainList([]byte(stevenBlackSample))
got, _ := parseDomainList([]byte(stevenBlackSample))
if len(got) != 5 {
t.Fatalf("expected 5 domains from the sample, got %d (%v)", len(got), got)
}
// Unicode is punycoded, matching what actually arrives in a DNS query.
uni := parseDomainList([]byte("0.0.0.0 реклама.рф\n"))
uni, _ := parseDomainList([]byte("0.0.0.0 реклама.рф\n"))
if len(uni) != 1 || !strings.HasPrefix(uni[0], "xn--") {
t.Fatalf("a unicode entry must be punycoded, got %v", uni)
}
+210
View File
@@ -0,0 +1,210 @@
package generate
// The destination-list vocabulary has TWO homes, not one, and the difference is
// reported rather than silent.
//
// docs-shater/DECISIONS.md D21 defines one entry vocabulary — full: / suffix: /
// keyword: / regexp: / a leading dot / a bare name — and it belongs to the entries
// an operator TYPES, i.e. an inline `config ruleset`. A `source=url` list that is
// not engine-native is a hosts/plain-domain FILE, parsed by parseDomainList, which
// has no markers at all. See the block above parseDomainList (ruleset.go) for why
// unifying the two was rejected.
//
// These tests pin the consequence that makes that acceptable: an entry written in
// the inline vocabulary inside a text list is NAMED, on every generate, and the
// check is narrow enough that an ordinary published filter list stays silent.
import (
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// urlDomainSet builds a routing `config ruleset` fed from a plain-text URL.
func urlDomainSet(name, url string) model.Ruleset {
return model.Ruleset{Name: name, Type: "domain", Source: "url", URL: url}
}
// textListModel routes one rule at a url-sourced destination list.
func textListModel(url string) (sets []model.Ruleset, rule model.Rule) {
return []model.Ruleset{urlDomainSet("dest", url)},
model.Rule{Name: "dest", Enabled: true, Order: 10, DstRuleset: []string{"dest"}, Target: "node:n1"}
}
// TestTextListInlineVocabularyIsReported is the regression: `regexp:^ads\.` (and
// every other inline marker) works in an inline rule-set and matches NOTHING in a
// plain-text list, because normaliseListDomain drops anything containing ":". The
// drop is correct — a domain name cannot contain a colon — but it used to happen
// without a word, so the same string appeared to work in one spelling of "a list of
// destinations" and to do nothing in the other, with no way to tell which.
func TestTextListInlineVocabularyIsReported(t *testing.T) {
resetListMarkerMemo()
t.Cleanup(resetListMarkerMemo)
withListFetcher(t, func(string) ([]byte, error) {
return []byte(strings.Join([]string{
"# a list someone hand-wrote in the inline vocabulary",
"0.0.0.0 ads.example.com",
`regexp:^ads\.`,
"full:exact.example",
"keyword:track",
"suffix:apex.example",
"tracker.example.org",
}, "\n")), nil
})
sets, rule := textListModel("https://lists.example/dest.txt")
rt, warns := genRulesWithSets(t, sets, rule)
// The list itself still works — the plain entries are imported as usual.
if _, ok := ruleSetByTag(rt, "rs-dest"); !ok {
t.Fatalf("the plain entries must still compile into a rule-set; warnings=%v", warns)
}
if !hasRulesetRule(rt, "dest") {
t.Fatalf("the rule referencing the list must still be emitted; warnings=%v", warns)
}
// ...and the four entries that silently vanished are named.
if !routeWarnsHave(warns, "inline rule-set vocabulary") {
t.Fatalf("marker entries in a text list must be reported, got %v", warns)
}
// %q-escaped in the message, so match the stable head of the pattern.
if !routeWarnsHave(warns, `regexp:^ads`) {
t.Fatalf("the warning must quote the offending entry, got %v", warns)
}
if !routeWarnsHave(warns, "source=inline") {
t.Fatalf("the warning must say where those markers DO work, got %v", warns)
}
if !routeWarnsHave(warns, "4 of its entries") {
t.Fatalf("the warning must count every dropped marker entry (4), got %v", warns)
}
}
// TestTextListVocabularyWarningSurvivesAFreshArtifact: the parse that can see the
// offending entries happens only when the artifact is refreshed, which is once per
// update_interval. A diagnostic that existed for exactly that one reconcile and
// then disappeared for a day would be no diagnostic at all, so the finding is
// remembered per URL and re-reported on every generate.
func TestTextListVocabularyWarningSurvivesAFreshArtifact(t *testing.T) {
resetListMarkerMemo()
t.Cleanup(resetListMarkerMemo)
var fetches int
withListFetcher(t, func(string) ([]byte, error) {
fetches++
return []byte("keyword:track\ngood.example.com\n"), nil
})
sets, rule := textListModel("https://lists.example/sticky.txt")
for pass := 1; pass <= 3; pass++ {
_, warns := genRulesWithSets(t, sets, rule)
if !routeWarnsHave(warns, "inline rule-set vocabulary") {
t.Fatalf("pass %d: the warning must persist while the list does; warnings=%v", pass, warns)
}
}
if fetches != 1 {
t.Fatalf("a fresh artifact must not be re-downloaded, fetched %d times", fetches)
}
}
// TestTextListVocabularyWarningClearsWhenTheListIsFixed: the memo is a finding
// about the list, not a sticky flag. Once a refresh sees a clean list the warning
// stops, or an operator who fixed the problem would never know they had.
func TestTextListVocabularyWarningClearsWhenTheListIsFixed(t *testing.T) {
resetListMarkerMemo()
t.Cleanup(resetListMarkerMemo)
body := "keyword:track\ngood.example.com\n"
withListFetcher(t, func(string) ([]byte, error) { return []byte(body), nil })
sets, rule := textListModel("https://lists.example/fixed.txt")
if _, warns := genRulesWithSets(t, sets, rule); !routeWarnsHave(warns, "inline rule-set vocabulary") {
t.Fatalf("setup: the first pass must report the marker entry; warnings=%v", warns)
}
body = "good.example.com\nbetter.example.org\n"
// The artifact from the first pass is fresh, so force the refresh the operator's
// next update_interval would have done anyway.
resetListMarkerMemo()
listsDirOverride = t.TempDir()
if _, warns := genRulesWithSets(t, sets, rule); routeWarnsHave(warns, "inline rule-set vocabulary") {
t.Fatalf("a clean list must not keep warning; warnings=%v", warns)
}
}
// TestPublishedFilterListDoesNotTripTheVocabularyWarning is the other half of the
// trade, and the reason the check tests only the four D21 markers instead of the
// general `word:` shape unrecognisedDomainPrefix uses. A real AdGuard/OISD list is
// full of colon-bearing tokens that are ordinary filter syntax; warning about those
// would put hundreds of lines per list in front of the operator, which is a louder
// dishonesty than the silence it replaced.
func TestPublishedFilterListDoesNotTripTheVocabularyWarning(t *testing.T) {
resetListMarkerMemo()
t.Cleanup(resetListMarkerMemo)
withListFetcher(t, func(string) ([]byte, error) {
return []byte(strings.Join([]string{
"! Title: Example filter list",
"! Homepage: https://lists.example/",
"||ads.example.com^$third-party",
"example.com##.banner:has(> .ad)",
"https://tracker.example/pixel.gif",
"@@||allowed.example.net^",
"0.0.0.0 good.example.net",
"fe80::1 ip6-localhost",
}, "\n")), nil
})
sets, rule := textListModel("https://lists.example/adguard.txt")
rt, warns := genRulesWithSets(t, sets, rule)
if _, ok := ruleSetByTag(rt, "rs-dest"); !ok {
t.Fatalf("the list must still compile; warnings=%v", warns)
}
if routeWarnsHave(warns, "inline rule-set vocabulary") {
t.Fatalf("ordinary filter syntax must not be reported as a vocabulary mistake: %v", warns)
}
}
// TestParseDomainListReportsOnlyTheInlineMarkers pins the predicate itself, away
// from the generate machinery: the four markers are recorded, everything else that
// the parser refuses stays a silent format detail.
func TestParseDomainListReportsOnlyTheInlineMarkers(t *testing.T) {
domains, note := parseDomainList([]byte(strings.Join([]string{
"0.0.0.0 kept.example.com",
"FULL:Exact.Example", // markers are case-insensitive, like splitDomainMarker
"regexp:^ads\\.",
"*.wildcard.example", // junk, but not a vocabulary mistake
"2001:db8::1", // colons, but an address — never a marker
"nodot",
}, "\n")))
if len(domains) != 1 || domains[0] != "kept.example.com" {
t.Fatalf("imported domains = %v, want [kept.example.com]", domains)
}
if note.Count != 2 {
t.Fatalf("marker count = %d, want 2 (FULL: and regexp:); samples=%v", note.Count, note.Samples)
}
for _, want := range []string{"FULL:Exact.Example", `regexp:^ads\.`} {
var seen bool
for _, s := range note.Samples {
if s == want {
seen = true
}
}
if !seen {
t.Fatalf("sample %q missing from %v", want, note.Samples)
}
}
}
// TestTextListVocabularyWarningIsBounded: the body is remote content, so it must
// not be able to write an unbounded amount into the operator's warning list.
func TestTextListVocabularyWarningIsBounded(t *testing.T) {
var lines []string
for i := 0; i < 500; i++ {
lines = append(lines, "keyword:junk")
}
_, note := parseDomainList([]byte(strings.Join(lines, "\n")))
if note.Count != 500 {
t.Fatalf("count = %d, want the true total 500", note.Count)
}
if len(note.Samples) != listMarkerSamples {
t.Fatalf("quoted %d entries, want at most %d", len(note.Samples), listMarkerSamples)
}
}
+21 -28
View File
@@ -12,47 +12,36 @@ import (
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// scheduledDomain is the distinctive dst-domain matcher used to detect whether a
// scheduled rule survived into the generated route rules.
const scheduledDomain = "sched.example"
// scheduledSet is the distinctive destination rule-set used to detect whether a
// scheduled rule survived into the generated route rules. Since schema v2 a
// rule's destination is a rule-set reference, so this doubles as a check that
// buildRoutingRuleSets honours the same schedule gate buildRoute does: outside
// the window neither the rule nor its rs- rule-set may be emitted.
const scheduledSet = "sched"
// emittedAt reports whether the given scheduled rule is present in the route
// rules when generate's clock is `now`.
func emittedAt(now time.Time, r model.Rule) bool {
b := newBuilder(&model.Model{Globals: model.DefaultGlobals(), Rules: []model.Rule{r}})
b := newBuilder(&model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{{Name: scheduledSet, Type: "domain", Source: "inline", Entries: []string{"sched.example"}}},
Rules: []model.Rule{r},
})
b.now = now
return hasDomainRule(b.buildRoute(), scheduledDomain)
return hasRulesetRule(b.buildRoute(), scheduledSet)
}
// hasDomainRule reports whether any route rule carries the given exact dst-domain
// matcher.
func hasDomainRule(rt *option.RouteOptions, domain string) bool {
if rt == nil {
return false
}
for _, r := range rt.Rules {
for _, d := range r.DefaultOptions.RawDefaultRule.Domain {
if d == domain {
return true
}
}
}
return false
}
// schedRule builds a scheduled dst-domain rule (target direct) from the schedule
// fields. It always carries the scheduledDomain matcher so emittedAt can find it.
// schedRule builds a scheduled destination rule (target direct) from the schedule
// fields. It always references the scheduledSet rule-set so emittedAt can find it.
func schedRule(days []string, start, end string) model.Rule {
return model.Rule{
Name: "sched",
Enabled: true,
Order: 10,
DstDomain: []string{scheduledDomain},
DstRuleset: []string{scheduledSet},
Target: "direct",
SchedEnabled: true,
SchedDays: days,
@@ -127,9 +116,13 @@ func TestScheduleAllDayWeekend(t *testing.T) {
// warning instead.
func TestScheduleInvalidTimeIsAlwaysOn(t *testing.T) {
rule := schedRule(nil, "9am", "17:00") // "9am" is not HH:MM
b := newBuilder(&model.Model{Globals: model.DefaultGlobals(), Rules: []model.Rule{rule}})
b := newBuilder(&model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{{Name: scheduledSet, Type: "domain", Source: "inline", Entries: []string{"sched.example"}}},
Rules: []model.Rule{rule},
})
b.now = time.Date(2026, 7, 15, 3, 0, 0, 0, time.UTC) // 03:00 — would be OUTSIDE a 09–17 window
if !hasDomainRule(b.buildRoute(), scheduledDomain) {
if !hasRulesetRule(b.buildRoute(), scheduledSet) {
t.Errorf("invalid start time should fail OPEN (rule emitted always-on)")
}
if len(b.warnings) == 0 {
+47
View File
@@ -0,0 +1,47 @@
package logsink
// Colour policy for everything that writes into the sink.
//
// Both log producers of the daemon (the control-plane factory built in
// cmd/shaterd, and the engine's own factory built by box.New from
// option.LogOptions) format with github.com/logrusorgru/aurora colours ON by
// default: log/format.go paints the level word and the per-connection ID unless
// DisableColors / LogOptions.DisableColor is set. Under procd the daemon's
// stderr is not a terminal, it is syslog — so every ERROR line landed in
// `logread` as
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m1728741629ESC[0m 70ms] dns: …
//
// Syslog is not a screen: `logread | grep ERROR` misses the coloured word
// because there are invisible bytes inside it, log collectors store the escapes
// forever, and anyone reading a captured log sees mojibake. The sink's file half
// already strips ANSI on the way out (emitLocked -> stripANSI) — the syslog half
// deliberately did not, and that is the leak.
//
// The fix is at the producer, not at the sink: colour is a property of the
// DESTINATION, so it is decided once, from whether that destination is a
// terminal, and never emitted otherwise. Stripping at the sink would keep the
// wasted formatting work and would still leak through any future path that does
// not go through the sink.
import "os"
// IsTTY reports whether w is a terminal, i.e. whether ANSI colour escapes
// written to it will be RENDERED rather than stored.
//
// The check is the portable one — a character device — so it needs no cgo, no
// termios ioctl and no new dependency (the router binary is CGO_ENABLED=0
// musl-static, and logsink also builds on the Windows/macOS dev hosts). Under
// procd, stderr is a pipe to the log daemon: not a character device, so colour
// is off, which is the case that matters. An interactive `shaterd run` from a
// shell keeps its colours.
func IsTTY(w *os.File) bool {
if w == nil {
return false
}
fi, err := w.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
}
+94
View File
@@ -0,0 +1,94 @@
package logsink
import (
"bytes"
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestIsTTYNonTerminal pins the only direction that matters in production: the
// destinations the daemon actually has under procd — a pipe (procd's stderr
// relay) and a regular file — are NOT terminals, so nothing may colour for them.
func TestIsTTYNonTerminal(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
if IsTTY(w) {
t.Errorf("IsTTY(pipe) = true; procd's stderr is a pipe and must never be coloured")
}
f, err := os.Create(filepath.Join(t.TempDir(), "log"))
if err != nil {
t.Fatalf("create: %v", err)
}
t.Cleanup(func() { _ = f.Close() })
if IsTTY(f) {
t.Errorf("IsTTY(regular file) = true, want false")
}
if IsTTY(nil) {
t.Errorf("IsTTY(nil) = true, want false")
}
}
// TestSyslogHalfHasNoANSI is the regression for the escape codes that reached
// `logread`:
//
// daemon.err shaterd[27540]: …Z ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: …
//
// It wires a factory the way the daemon does — formatter colour gated on
// IsTTY(destination), output into a Sink whose syslog half is captured — and
// asserts the captured bytes carry no ESC (0x1b). The context ID is set because
// the ID is coloured by a SEPARATE branch of log/format.go: a fix that only
// silenced the level word would still leak here.
func TestSyslogHalfHasNoANSI(t *testing.T) {
r, w, err := os.Pipe() // a non-terminal destination, exactly like procd's stderr
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
var syslog bytes.Buffer
s := New(&syslog, Config{ToSyslog: true})
factory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now(), DisableColors: !IsTTY(w)},
s, "", nil, false)
logger := factory.Logger()
ctx := log.ContextWithNewID(context.Background())
logger.ErrorContext(ctx, "dns: exchange failed for catalog.example.com. IN AAAA: unexpected EOF")
logger.WarnContext(ctx, "warn line")
logger.InfoContext(ctx, "info line")
factory.SetLevel(log.LevelTrace)
logger.DebugContext(ctx, "debug line")
logger.TraceContext(ctx, "trace line")
_ = s.Close()
got := syslog.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("the syslog half captured nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("syslog half carries an ANSI escape at byte %d: %q", i, got)
}
if !strings.Contains(got, "ERROR") {
t.Errorf("`grep ERROR` must match a plain, unbroken level word: %q", got)
}
// Teeth check: the SAME line through a colouring formatter must contain the
// escape. Without this, a future aurora that stopped colouring would make the
// assertion above pass for the wrong reason and the guard would rot silently.
coloured := log.Formatter{BaseTime: time.Now()}.
Format(ctx, log.LevelError, "", "dns: exchange failed", time.Now())
if !strings.ContainsRune(coloured, 0x1b) {
t.Fatalf("colouring formatter emitted no ANSI escape (%q) — this test can no longer detect the leak", coloured)
}
}
+365 -4
View File
@@ -14,17 +14,40 @@ import (
// CurrentSchemaVersion is the schema this build understands. Bump it when adding
// a migration step below.
const CurrentSchemaVersion = 1
const CurrentSchemaVersion = 2
// uciRunner abstracts uci get/set/delete/commit/import so migrations AND the
// config-write path (WriteUCI) are unit-testable. Import feeds `uci export`-format
// text to `uci import <pkg>` on stdin, replacing the package's staged sections.
//
// Export/Add/AddList exist for migrations that have to READ the config they are
// rewriting and GROW it. migrate1to2 needs both: it reads options the current
// Model no longer parses (`dst_domain`/`dst_ip` were removed from Rule) and adds
// the `config ruleset` sections it folds them into. Add returns the generated
// section id — the anonymous-index drift trap is real (`@ruleset[3]` means
// something different after one more add), so every write goes through the id.
type uciRunner interface {
Get(key string) (string, bool)
Set(key, val string) error
Delete(key string) error
Commit(pkg string) error
// Revert drops the package's STAGED (uncommitted) changes.
//
// A migration stages many writes and commits once, so a failure halfway
// through leaves a half-migration sitting in /tmp/.uci — and staged changes
// are not private to us: the next `uci commit shater` from ANY process (the
// panel saving one setting through WriteUCI, an operator at the shell) flushes
// them to disk, producing a config that is half v1 and half v2. Every error
// path in a migration must therefore revert before returning.
Revert(pkg string) error
Import(pkg, text string) error
// Export returns the package in `uci export` format. ok=false when the package
// does not exist (nothing to migrate), which is NOT an error.
Export(pkg string) (text string, ok bool)
// Add appends an anonymous `config <secType>` and returns its section id.
Add(pkg, secType string) (id string, err error)
// AddList appends one value to a list option (`uci add_list <key>=<val>`).
AddList(key, val string) error
}
type execUCI struct{}
@@ -39,6 +62,7 @@ func (execUCI) Get(k string) (string, bool) {
func (execUCI) Set(k, v string) error { return exec.Command("uci", "set", k+"="+v).Run() }
func (execUCI) Delete(k string) error { return exec.Command("uci", "-q", "delete", k).Run() }
func (execUCI) Commit(p string) error { return exec.Command("uci", "commit", p).Run() }
func (execUCI) Revert(p string) error { return exec.Command("uci", "revert", p).Run() }
func (execUCI) Import(pkg, text string) error {
cmd := exec.Command("uci", "import", pkg)
@@ -46,6 +70,30 @@ func (execUCI) Import(pkg, text string) error {
return cmd.Run()
}
func (execUCI) Export(pkg string) (string, bool) {
out, err := exec.Command("uci", "-q", "export", pkg).Output()
if err != nil {
return "", false
}
return string(out), true
}
func (execUCI) Add(pkg, secType string) (string, error) {
out, err := exec.Command("uci", "add", pkg, secType).Output()
if err != nil {
return "", fmt.Errorf("uci add %s %s: %w", pkg, secType, err)
}
id := strings.TrimSpace(string(out))
if id == "" {
return "", fmt.Errorf("uci add %s %s: no section id returned", pkg, secType)
}
return id, nil
}
func (execUCI) AddList(k, v string) error {
return exec.Command("uci", "add_list", k+"="+v).Run()
}
// uci is the active runner (overridable in tests).
var uci uciRunner = execUCI{}
@@ -56,6 +104,7 @@ type migration struct {
var migrations = []migration{
{from: 0, to: 1, apply: migrate0to1},
{from: 1, to: 2, apply: migrate1to2},
}
func readSchemaVersion(u uciRunner) int {
@@ -75,9 +124,23 @@ func ensureGlobals(u uciRunner) {
func setSchemaVersion(u uciRunner, v int) error {
ensureGlobals(u)
if err := u.Set("shater.globals.schema_version", strconv.Itoa(v)); err != nil {
return err
return staged(u, err)
}
return u.Commit("shater")
if err := u.Commit("shater"); err != nil {
return staged(u, err)
}
return nil
}
// staged discards the package's staged writes and returns err unchanged. It is
// the one-liner every migration error path goes through: leaving a partial
// migration staged lets somebody else's `uci commit shater` write it to disk (see
// uciRunner.Revert). A failing revert cannot be reported on top of the original
// failure without hiding it, so it is deliberately ignored — err is the one the
// operator has to act on, and the revert is best-effort cleanup.
func staged(u uciRunner, err error) error {
_ = u.Revert("shater")
return err
}
// Migrate runs pending migrations to CurrentSchemaVersion using the active runner.
@@ -121,5 +184,303 @@ func migrate0to1(u uciRunner) error {
}
_ = u.Delete("shater.globals.kill")
}
return u.Commit("shater")
if err := u.Commit("shater"); err != nil {
return staged(u, err)
}
return nil
}
// --- v1 -> v2: a rule's destination is a rule-set, never an inline list ------
//
// WHAT CHANGED. `config rule` lost `dst_domain` and `dst_ip`. A rule now names
// its destination through `dst_ruleset` only, so there is ONE destination
// mechanism, one matcher vocabulary, and one place a list is edited — and the
// list is compiled once into a .srs that every referencing rule shares.
//
// WHAT THIS STEP DOES. For every rule that still carries one of the two options
// it creates an inline `config ruleset` named `rule-<rule name>` (and
// `rule-<rule name>-ip` for the address list, because a rule-set is EITHER a
// domain list or an ip_cidr list), moves the entries into it, appends the new
// name to the rule's `dst_ruleset`, and deletes the legacy option. Nothing is
// dropped and nothing is guessed: a rule with both lists gets both rule-sets.
//
// ENTRY SEMANTICS ARE PRESERVED 1:1, and that needs one real conversion. The two
// contexts disagree about exactly one form: a BARE domain is an EXACT match in a
// routing rule (generate/route.go classified `dst_domain` with bareIsSuffix=false)
// and a SUFFIX match inside a rule-set (inlineDomainRule, bareIsSuffix=true).
// Copying `example.com` across verbatim would therefore silently widen the rule to
// every subdomain, so a bare entry is rewritten as `full:example.com`. Every other
// form already means the same thing on both sides and is copied byte-for-byte:
// `full:`, `suffix:`, `keyword:`, `regexp:` (see generate.peelDomainRegexes, added
// with this change so the regex form survives the move) and a leading dot, which
// is a synonym of `suffix:` in both. `geosite:`/`geoip:` entries are copied
// unchanged too: they are INERT in a routing rule on this engine (warned and
// omitted — the route-rule geosite/geoip fields no longer exist), and an
// unrecognised marker is equally inert inside a rule-set, so their meaning is
// unchanged and the operator's text is not thrown away. Converting them into a
// `source=geosite` rule-set would have made a dead matcher start routing traffic
// during an upgrade — a behaviour change, not a migration.
//
// ONE DELIBERATE SEMANTIC IMPROVEMENT, stated out loud: a rule that carried BOTH
// a domain list and an address list matched them with AND (an engine route rule
// ANDs its matcher fields), which is almost never what "these sites and these
// networks" was meant to say. The two generated rule-sets are ORed, because
// `rule_set: [a, b]` matches when EITHER matches. Such a rule matches more after
// the migration than before — it is called out here, in the docs, and it only
// affects configs that used both fields at once.
//
// IDEMPOTENCE. The legacy options are deleted as the last step per rule, so a
// second run finds nothing to do. A run interrupted between "create the rule-set"
// and "delete the option" is also safe: a rule that ALREADY references a rule-set
// of the expected name AND WHOSE ENTRIES ARE IN IT reuses it instead of creating
// `rule-<name>-2` (see ensureMigratedRuleset — the entry check is what tells our
// own half-finished work apart from an operator's hand-written list that happens
// to carry the same name).
//
// A LEGACY OPTION IS DELETED ONLY WHEN ITS CONTENT HAS A NEW HOME. Deleting is
// per-list and conditional on that list having produced (or confirmed) a
// dst_ruleset reference. Unconditional deletion had a hole: a list whose values
// are all blank (`list dst_domain ' '` — a space survives parseSections' empty
// check but trims away in migrateDomainEntry) produced no rule-set, and deleting
// it anyway left the rule with NO matchers at all, i.e. a catch-all that takes
// over the router's default route. Keeping the option leaves the rule visibly
// unmigrated instead, which ParseUCIExport holds disabled and ValidateRules
// reports.
//
// ERRORS ABORT AND REVERT. Every write here is staged and committed once at the
// end, so a failure that returned without reverting would leave a half-migration
// in /tmp/.uci for the next `uci commit shater` (from the panel, say) to flush.
// Every error path goes through staged(). Deletes are checked too: a delete that
// silently failed left a legacy option in the config forever, because migrateWith
// bumps schema_version to 2 on return and readSchemaVersion never asks for this
// step again.
func migrate1to2(u uciRunner) error {
text, ok := u.Export("shater")
if !ok || strings.TrimSpace(text) == "" {
return nil // no config yet (fresh install): nothing to migrate
}
secs, err := parseSections(text)
if err != nil {
return fmt.Errorf("read the current config: %w", err)
}
// Every rule-set name already in use -> its entries, so a generated name can
// never collide with a hand-written list (which would make `uci` hold two
// `config ruleset` blocks claiming the same name, and the generator drops one as
// a duplicate tag). The ENTRIES are carried, not just the name, because
// "reuse the existing list" is only correct when that list is the one a previous
// run of this migration created.
taken := map[string][]string{}
for _, s := range secs {
if s.Type == "ruleset" {
if n := firstNonEmpty(s.opt("name"), s.Name); n != "" {
taken[n] = s.list("entry")
}
}
}
ruleIdx := -1
for _, s := range secs {
if s.Type != "rule" {
continue
}
// Anonymous sections are addressed positionally and the index is PER TYPE, so
// it counts rules only. Appending `config ruleset` sections below cannot shift
// it (a new section goes to the end, and it is not a rule).
ruleIdx++
domains := s.list("dst_domain")
ips := s.list("dst_ip")
if len(domains) == 0 && len(ips) == 0 {
continue
}
rulePath := fmt.Sprintf("shater.@rule[%d]", ruleIdx)
base := rulesetBaseName(firstNonEmpty(s.opt("name"), s.Name), ruleIdx)
refs := s.list("dst_ruleset")
domainsDone, ipsDone := false, false
if len(domains) > 0 {
entries := make([]string, 0, len(domains))
for _, d := range domains {
if e := migrateDomainEntry(d); e != "" {
entries = append(entries, e)
}
}
domainsDone, err = ensureMigratedRuleset(u, rulePath, base, "domain", entries, refs, taken)
if err != nil {
return staged(u, err)
}
}
if len(ips) > 0 {
entries := make([]string, 0, len(ips))
for _, ip := range ips {
if v := strings.TrimSpace(ip); v != "" {
entries = append(entries, v)
}
}
ipsDone, err = ensureMigratedRuleset(u, rulePath, base+"-ip", "ipcidr", entries, refs, taken)
if err != nil {
return staged(u, err)
}
}
// Last, so an interrupted run still has the legacy list to redo the work from
// — and only for a list whose entries actually reached a rule-set.
if domainsDone {
if err := u.Delete(rulePath + ".dst_domain"); err != nil {
return staged(u, fmt.Errorf("drop the migrated dst_domain of rule %d: %w", ruleIdx, err))
}
}
if ipsDone {
if err := u.Delete(rulePath + ".dst_ip"); err != nil {
return staged(u, fmt.Errorf("drop the migrated dst_ip of rule %d: %w", ruleIdx, err))
}
}
}
if err := u.Commit("shater"); err != nil {
return staged(u, err)
}
return nil
}
// ensureMigratedRuleset creates the inline `config ruleset` holding entries and
// points rulePath's dst_ruleset at it. want is the preferred name; a collision
// with an existing list picks want-2, want-3, ... rsType is "domain" or "ipcidr".
// taken maps every rule-set name already in the config to its entries, and is
// updated with whatever this call adds.
//
// It reports whether the entries now have a home — which is what tells the caller
// it may delete the legacy option they came from. false with a nil error means
// "nothing was written and nothing may be deleted": the only such case is an
// entry list that came out EMPTY (a legacy list of blanks). An empty inline
// rule-set matches nothing and the generator would skip it, so creating one would
// be pure noise — but the legacy option must then survive, or the rule is left
// with no matchers at all and silently becomes the default route.
//
// RESUMING AN INTERRUPTED RUN, WITHOUT SWALLOWING A HAND-WRITTEN LIST. A rule
// that already references a rule-set of the expected name is the signature of a
// previous run that died between "create the rule-set" and "delete the option" —
// but it is ALSO what an operator's own `config ruleset name='rule-ads'` plus a
// rule referencing it looks like. Telling them apart takes one more question:
// does that rule-set actually CONTAIN these entries? Our own leftover does, by
// construction. The operator's list (a remote URL list, say) does not, and
// treating it as "already migrated" threw the legacy entries away with no trace.
// When the entries are not in it, this falls through to the ordinary collision
// path and creates want-2, so the rule ends up referencing both lists — the two
// references are ORed, so nothing the operator wrote stops matching.
func ensureMigratedRuleset(u uciRunner, rulePath, want, rsType string, entries, existingRefs []string, taken map[string][]string) (bool, error) {
if len(entries) == 0 {
return false, nil
}
if have, ok := taken[want]; ok && containsString(existingRefs, want) && containsAll(have, entries) {
return true, nil // already migrated (interrupted run); nothing to add
}
name := want
for i := 2; ; i++ {
if _, clash := taken[name]; !clash {
break
}
name = fmt.Sprintf("%s-%d", want, i)
}
taken[name] = entries
id, err := u.Add("shater", "ruleset")
if err != nil {
return false, err
}
sec := "shater." + id
if err := u.Set(sec+".name", name); err != nil {
return false, err
}
if err := u.Set(sec+".type", rsType); err != nil {
return false, err
}
if err := u.Set(sec+".source", "inline"); err != nil {
return false, err
}
for _, e := range entries {
if err := u.AddList(sec+".entry", e); err != nil {
return false, err
}
}
if containsString(existingRefs, name) {
return true, nil
}
if err := u.AddList(rulePath+".dst_ruleset", name); err != nil {
return false, err
}
return true, nil
}
// rulesetBaseName builds `rule-<name>` from a rule's name, reduced to characters
// that are safe in a rule-set name (it becomes an engine rule-set TAG, `rs-<name>`,
// and a UCI option value). An unnamed rule falls back to its position so two of
// them cannot produce the same base.
func rulesetBaseName(ruleName string, idx int) string {
var b strings.Builder
prevDash := false
for _, r := range strings.TrimSpace(ruleName) {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9', r == '_', r == '.':
b.WriteRune(r)
prevDash = false
default:
if !prevDash && b.Len() > 0 {
b.WriteByte('-')
prevDash = true
}
}
}
slug := strings.Trim(b.String(), "-.")
if slug == "" {
slug = strconv.Itoa(idx)
}
return "rule-" + slug
}
// migrateDomainEntry rewrites ONE `dst_domain` entry into the rule-set spelling
// with the same meaning. Only the bare form differs between the two contexts
// (exact in a rule, suffix in a rule-set), so only it is rewritten; see the
// migrate1to2 doc comment for the full table and the reasoning.
func migrateDomainEntry(e string) string {
v := strings.TrimSpace(e)
if v == "" {
return ""
}
// A leading dot already means `suffix:` on both sides.
if strings.HasPrefix(v, ".") {
return v
}
// Any `word:` marker — recognised (full/suffix/keyword/regexp) or not
// (geosite/geoip/typos) — carries its meaning across unchanged. A domain label
// cannot contain a colon, so this cannot misfire on a real host name.
if strings.Contains(v, ":") {
return v
}
return "full:" + v
}
// containsString reports whether list holds want (trimmed comparison).
func containsString(list []string, want string) bool {
for _, v := range list {
if strings.TrimSpace(v) == want {
return true
}
}
return false
}
// containsAll reports whether every entry in want is present in have. It answers
// exactly one question — "is this rule-set the one a previous run of this
// migration wrote?" — so it is a SUBSET test, not equality: a rule-set that also
// holds entries the operator added by hand since is still ours to reuse.
func containsAll(have, want []string) bool {
for _, w := range want {
if !containsString(have, strings.TrimSpace(w)) {
return false
}
}
return true
}
+919 -21
View File
@@ -1,53 +1,951 @@
package model
import "testing"
import (
"errors"
"fmt"
"sort"
"strconv"
"strings"
"testing"
)
// fakeUCI is an in-memory uciRunner for the migration + write tests (no router
// needed). imported records the last `uci import` text so WriteUCI can be
// asserted without a device.
// fakeUCI is an in-memory stand-in for the `uci` CLI: enough of a section model
// that a migration can EXPORT the config, rewrite it, and export it again and see
// its own writes. That is what makes the idempotence assertions below real —
// against a flat key/value map a second migration run would re-read the original
// text and "prove" nothing.
//
// Addressing mirrors uci: `shater.globals.opt` (named section), `shater.@rule[2]`
// (positional, index is PER TYPE), `shater.cfg001.opt` (the id `uci add` returns).
type fakeUCI struct {
kv map[string]string
secs []*fakeSection
imported string
commits int
reverts int
deleted []string
nextID int
// missing makes Export report "no such package", the fresh-install case.
missing bool
}
func (f *fakeUCI) Get(k string) (string, bool) { v, ok := f.kv[k]; return v, ok }
func (f *fakeUCI) Set(k, v string) error { f.kv[k] = v; return nil }
func (f *fakeUCI) Delete(k string) error { f.deleted = append(f.deleted, k); delete(f.kv, k); return nil }
func (f *fakeUCI) Commit(string) error { f.commits++; return nil }
func (f *fakeUCI) Import(pkg, text string) error { f.imported = text; return nil }
type fakeSection struct {
id string
typ string
name string // "" for an anonymous section
opts map[string]string
oKeys []string // option order, so the rendered export is deterministic
lists map[string][]string
lKeys []string
}
// loadSections copies parsed sections into the fake, in a DETERMINISTIC order.
//
// uciSection carries its options and lists in maps, and Go randomises map
// iteration, so feeding them to setOpt/addList in range order made oKeys/lKeys —
// and therefore the rendered Export — differ run to run. That is not cosmetic
// here: TestMigrateIdempotent and TestMigrate1to2IsIdempotent compare two export
// strings, so a random key order turned them into coin flips that fail a few
// percent of the time and look like a migration bug. Sorting by key restores the
// determinism the oKeys/lKeys fields were added for.
func (f *fakeUCI) loadSections(secs []uciSection) {
for _, s := range secs {
sec := f.newSection(s.Type, s.Name)
for _, k := range sortedKeys(s.Options) {
sec.setOpt(k, s.Options[k])
}
for _, k := range sortedKeys(s.Lists) {
for _, v := range s.Lists[k] {
sec.addList(k, v)
}
}
}
}
func sortedKeys[V any](m map[string]V) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
sort.Strings(out)
return out
}
func newFakeUCI(export string) *fakeUCI {
f := &fakeUCI{}
if strings.TrimSpace(export) == "" {
return f
}
secs, err := parseSections(export)
if err != nil {
panic("fakeUCI fixture: " + err.Error())
}
f.loadSections(secs)
return f
}
func (f *fakeUCI) newSection(typ, name string) *fakeSection {
sec := &fakeSection{
id: fmt.Sprintf("cfg%03d", f.nextID),
typ: typ,
name: name,
opts: map[string]string{},
lists: map[string][]string{},
}
f.nextID++
f.secs = append(f.secs, sec)
return sec
}
func (s *fakeSection) setOpt(k, v string) {
if _, seen := s.opts[k]; !seen {
s.oKeys = append(s.oKeys, k)
}
s.opts[k] = v
}
func (s *fakeSection) addList(k, v string) {
if _, seen := s.lists[k]; !seen {
s.lKeys = append(s.lKeys, k)
}
s.lists[k] = append(s.lists[k], v)
}
// resolve finds the section a `<pkg>.<sel>` selector names.
func (f *fakeUCI) resolve(sel string) *fakeSection {
if strings.HasPrefix(sel, "@") && strings.HasSuffix(sel, "]") {
open := strings.IndexByte(sel, '[')
if open < 0 {
return nil
}
typ := sel[1:open]
idx, err := strconv.Atoi(sel[open+1 : len(sel)-1])
if err != nil {
return nil
}
n := 0
for _, s := range f.secs {
if s.typ != typ {
continue
}
if n == idx {
return s
}
n++
}
return nil
}
for _, s := range f.secs {
if s.name == sel || s.id == sel {
return s
}
}
return nil
}
// splitKey cuts "shater.@rule[0].dst_domain" into ("@rule[0]", "dst_domain").
// A key with no option part yields opt == "".
func splitKey(key string) (sel, opt string) {
rest := strings.TrimPrefix(key, "shater")
rest = strings.TrimPrefix(rest, ".")
if rest == "" {
return "", ""
}
// The selector may contain a dot only inside a name, which the fixtures never
// use, so a plain LastIndex is enough — except for `@type[i]`, where the index
// brackets hold no dots either.
if i := strings.LastIndexByte(rest, '.'); i >= 0 {
return rest[:i], rest[i+1:]
}
return rest, ""
}
func (f *fakeUCI) Get(k string) (string, bool) {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return "", false
}
if opt == "" {
return sec.typ, true
}
v, ok := sec.opts[opt]
return v, ok
}
func (f *fakeUCI) Set(k, v string) error {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
if opt != "" {
return fmt.Errorf("uci set %s: no such section", k)
}
// `uci set shater.globals=globals` creates the named section.
f.newSection(v, sel)
return nil
}
if opt == "" {
sec.typ = v
return nil
}
sec.setOpt(opt, v)
return nil
}
func (f *fakeUCI) Delete(k string) error {
f.deleted = append(f.deleted, k)
if k == "shater" {
f.secs = nil
return nil
}
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return nil // `uci -q delete` on an absent key is a no-op
}
if opt == "" {
for i, s := range f.secs {
if s == sec {
f.secs = append(f.secs[:i], f.secs[i+1:]...)
break
}
}
return nil
}
delete(sec.opts, opt)
delete(sec.lists, opt)
return nil
}
func (f *fakeUCI) Commit(string) error { f.commits++; return nil }
// Revert is COUNTED, not simulated: this fake applies every write immediately
// (there is no staging area to roll back), so what a test can assert is that the
// migration ASKED for a revert on its way out of a failure. That is the property
// that matters — a real `uci` keeps the staged delta in /tmp/.uci until somebody
// commits it, and the migration's job is to not leave it there.
func (f *fakeUCI) Revert(string) error { f.reverts++; return nil }
func (f *fakeUCI) Import(_, text string) error {
f.imported = text
secs, err := parseSections(text)
if err != nil {
return err
}
f.loadSections(secs)
return nil
}
func (f *fakeUCI) Export(string) (string, bool) {
if f.missing {
return "", false
}
var b strings.Builder
b.WriteString("package shater\n")
for _, s := range f.secs {
b.WriteString("\nconfig " + s.typ)
if s.name != "" {
b.WriteString(" '" + s.name + "'")
}
b.WriteString("\n")
for _, k := range s.oKeys {
if v, ok := s.opts[k]; ok {
b.WriteString("\toption " + k + " '" + v + "'\n")
}
}
for _, k := range s.lKeys {
for _, v := range s.lists[k] {
b.WriteString("\tlist " + k + " '" + v + "'\n")
}
}
}
return b.String(), true
}
func (f *fakeUCI) Add(_, secType string) (string, error) {
return f.newSection(secType, "").id, nil
}
func (f *fakeUCI) AddList(k, v string) error {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return fmt.Errorf("uci add_list %s: no such section", k)
}
sec.addList(opt, v)
return nil
}
// ruleset returns the `config ruleset` section carrying option name == name.
func (f *fakeUCI) ruleset(name string) *fakeSection {
for _, s := range f.secs {
if s.typ == "ruleset" && s.opts["name"] == name {
return s
}
}
return nil
}
// rule returns the n-th `config rule` section.
func (f *fakeUCI) rule(idx int) *fakeSection { return f.resolve(fmt.Sprintf("@rule[%d]", idx)) }
func (f *fakeUCI) rulesetNames() []string {
var out []string
for _, s := range f.secs {
if s.typ == "ruleset" {
out = append(out, s.opts["name"])
}
}
return out
}
func eqStrings(a, b []string) bool {
if len(a) != len(b) {
return false
}
for i := range a {
if a[i] != b[i] {
return false
}
}
return true
}
func TestMigrate0to1TransformsFixture(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.kill": "open"}}
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption kill 'open'\n")
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if f.kv["shater.globals.kill_switch"] != "open" {
t.Fatalf("kill_switch = %q", f.kv["shater.globals.kill_switch"])
if v, _ := f.Get("shater.globals.kill_switch"); v != "open" {
t.Fatalf("kill_switch = %q", v)
}
if _, ok := f.kv["shater.globals.kill"]; ok {
if _, ok := f.Get("shater.globals.kill"); ok {
t.Fatal("legacy kill not removed")
}
if f.kv["shater.globals.schema_version"] != "1" {
t.Fatalf("schema_version = %q", f.kv["shater.globals.schema_version"])
if v, _ := f.Get("shater.globals.schema_version"); v != strconv.Itoa(CurrentSchemaVersion) {
t.Fatalf("schema_version = %q, want %d", v, CurrentSchemaVersion)
}
}
func TestMigrateIdempotent(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.schema_version": "1"}}
before := len(f.kv)
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption schema_version '" +
strconv.Itoa(CurrentSchemaVersion) + "'\n")
before, _ := f.Export("shater")
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if len(f.kv) != before {
t.Fatal("idempotent migrate changed state")
after, _ := f.Export("shater")
if before != after {
t.Fatalf("idempotent migrate changed state:\n--- before\n%s\n--- after\n%s", before, after)
}
}
func TestMigrateRefusesNewer(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.schema_version": "99"}}
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption schema_version '99'\n")
if err := migrateWith(f); err == nil {
t.Fatal("expected refusal of newer schema")
}
}
// --- v1 -> v2: dst_domain / dst_ip fold into generated rule-sets -------------
// legacyConfig is the shape a v1 config has: rules matching destinations inline.
// The `ru-direct` rule is copied verbatim from the live router this change was
// written against (BananaWRT 25.12.1, shater 0.2.7-r1).
const legacyConfig = `package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'ru-direct'
option enabled '1'
option order '10'
list dst_domain 'suffix:ru'
list dst_domain 'suffix:yandex.net'
list dst_domain 'full:vk.com'
list dst_domain 'keyword:sberbank'
list dst_domain '.gosuslugi.ru'
list dst_domain 'plain.example'
option target 'direct'
config rule
option name 'corp nets!'
option enabled '1'
option order '20'
list dst_ip '10.0.0.0/8'
list dst_ip '192.168.44.0/24'
option target 'group:corp'
config rule
option name 'default'
option enabled '1'
option order '100'
option target 'direct'
`
func TestMigrate1to2FoldsDomainsIntoARuleset(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
rs := f.ruleset("rule-ru-direct")
if rs == nil {
t.Fatalf("no rule-ru-direct ruleset; got %v", f.rulesetNames())
}
if rs.opts["type"] != "domain" || rs.opts["source"] != "inline" {
t.Fatalf("ruleset type/source = %q/%q, want domain/inline", rs.opts["type"], rs.opts["source"])
}
// Every entry keeps its meaning. Only the BARE one is rewritten: bare means
// EXACT in a rule and SUFFIX in a rule-set, so it becomes `full:`.
want := []string{
"suffix:ru", "suffix:yandex.net", "full:vk.com",
"keyword:sberbank", ".gosuslugi.ru", "full:plain.example",
}
if !eqStrings(rs.lists["entry"], want) {
t.Fatalf("entries = %q\nwant %q", rs.lists["entry"], want)
}
r0 := f.rule(0)
if !eqStrings(r0.lists["dst_ruleset"], []string{"rule-ru-direct"}) {
t.Fatalf("dst_ruleset = %q", r0.lists["dst_ruleset"])
}
if _, ok := r0.lists["dst_domain"]; ok {
t.Fatal("dst_domain survived the migration")
}
// Order and every other option are untouched.
if r0.opts["order"] != "10" || r0.opts["target"] != "direct" || r0.opts["name"] != "ru-direct" {
t.Fatalf("rule 0 mangled: %v", r0.opts)
}
}
func TestMigrate1to2FoldsCIDRsIntoAnIPRuleset(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
// The rule name is slugged: a rule-set name becomes an engine tag.
rs := f.ruleset("rule-corp-nets-ip")
if rs == nil {
t.Fatalf("no rule-corp-nets-ip ruleset; got %v", f.rulesetNames())
}
if rs.opts["type"] != "ipcidr" {
t.Fatalf("ruleset type = %q, want ipcidr", rs.opts["type"])
}
if !eqStrings(rs.lists["entry"], []string{"10.0.0.0/8", "192.168.44.0/24"}) {
t.Fatalf("entries = %q", rs.lists["entry"])
}
r1 := f.rule(1)
if !eqStrings(r1.lists["dst_ruleset"], []string{"rule-corp-nets-ip"}) {
t.Fatalf("dst_ruleset = %q", r1.lists["dst_ruleset"])
}
if _, ok := r1.lists["dst_ip"]; ok {
t.Fatal("dst_ip survived the migration")
}
}
// A rule with no destination at all stays a catch-all — the B1 reachability
// analysis (model.RuleReachability) keys off exactly that, so the migration must
// not hand it a rule-set it never asked for.
func TestMigrate1to2LeavesCatchAllAlone(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
r2 := f.rule(2)
if len(r2.lists["dst_ruleset"]) != 0 {
t.Fatalf("catch-all gained a ruleset: %q", r2.lists["dst_ruleset"])
}
text, _ := f.Export("shater")
m, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse migrated config: %v", err)
}
if len(m.Rules) != 3 {
t.Fatalf("rules = %d, want 3", len(m.Rules))
}
if !IsCatchAll(m.Rules[2]) {
t.Fatalf("rule %q stopped being a catch-all after the migration", m.Rules[2].Name)
}
for _, i := range []int{0, 1} {
if IsCatchAll(m.Rules[i]) {
t.Fatalf("rule %q became a catch-all — its destination was lost", m.Rules[i].Name)
}
}
reach := RuleReachability(m.Rules)
for i, v := range reach {
if v.Unreachable {
t.Fatalf("rule %d (%q) reported unreachable after the migration: %s", i, v.Name, v.Reason)
}
}
}
// Running the migration twice must not duplicate rule-sets or references. The
// second run goes through migrate1to2 directly, because migrateWith is gated by
// schema_version and would (correctly) do nothing at all.
func TestMigrate1to2IsIdempotent(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
first, _ := f.Export("shater")
if err := migrate1to2(f); err != nil {
t.Fatalf("second migrate1to2: %v", err)
}
second, _ := f.Export("shater")
if first != second {
t.Fatalf("re-running the migration changed the config:\n--- first\n%s\n--- second\n%s", first, second)
}
// And a full migrateWith re-run (schema already at CurrentSchemaVersion) is a
// no-op too.
if err := migrateWith(f); err != nil {
t.Fatalf("third migrate: %v", err)
}
third, _ := f.Export("shater")
if third != second {
t.Fatalf("re-running migrateWith changed the config:\n%s", third)
}
}
// An interrupted run — the rule-set was created and referenced, but the legacy
// option was not deleted yet — must reuse the rule-set rather than make a second.
func TestMigrate1to2ResumesAnInterruptedRun(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config ruleset
option name 'rule-half'
option type 'domain'
option source 'inline'
list entry 'full:a.example'
config rule
option name 'half'
list dst_domain 'a.example'
list dst_ruleset 'rule-half'
option target 'direct'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if names := f.rulesetNames(); !eqStrings(names, []string{"rule-half"}) {
t.Fatalf("rulesets = %q, want just rule-half", names)
}
r := f.rule(0)
if !eqStrings(r.lists["dst_ruleset"], []string{"rule-half"}) {
t.Fatalf("dst_ruleset = %q", r.lists["dst_ruleset"])
}
if _, ok := r.lists["dst_domain"]; ok {
t.Fatal("dst_domain survived")
}
}
// A hand-written rule-set already owning the generated name must not be
// clobbered: two `config ruleset` blocks with one name collide on the engine tag
// and one of them is dropped.
func TestMigrate1to2AvoidsNameCollisions(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config ruleset
option name 'rule-ads'
option type 'domain'
option source 'url'
option url 'https://example.invalid/list.txt'
config rule
option name 'ads'
list dst_domain 'ads.example'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
names := f.rulesetNames()
if !eqStrings(names, []string{"rule-ads", "rule-ads-2"}) {
t.Fatalf("rulesets = %q, want rule-ads + rule-ads-2", names)
}
if got := f.ruleset("rule-ads").opts["source"]; got != "url" {
t.Fatalf("the hand-written list was overwritten (source = %q)", got)
}
if !eqStrings(f.rule(0).lists["dst_ruleset"], []string{"rule-ads-2"}) {
t.Fatalf("dst_ruleset = %q", f.rule(0).lists["dst_ruleset"])
}
}
// A hand-written rule-set the rule ALREADY references must not swallow the legacy
// entries. `taken[want] && rule references want` alone reads an operator's own
// `rule-ads` list as our own half-finished work from an interrupted run, and the
// legacy `ads.example` was then dropped with nothing said. The entries have to
// actually BE in that rule-set for it to count as ours.
func TestMigrate1to2DoesNotFoldIntoAReferencedHandWrittenRuleset(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config ruleset
option name 'rule-ads'
option type 'domain'
option source 'url'
option url 'https://example.invalid/list.txt'
config rule
option name 'ads'
list dst_domain 'ads.example'
list dst_ruleset 'rule-ads'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
// The operator's list is untouched...
if got := f.ruleset("rule-ads").opts["source"]; got != "url" {
t.Fatalf("the hand-written list was overwritten (source = %q)", got)
}
if got := f.ruleset("rule-ads").lists["entry"]; len(got) != 0 {
t.Fatalf("entries were injected into the hand-written list: %q", got)
}
// ...and the legacy entry landed in a new list of its own.
rs := f.ruleset("rule-ads-2")
if rs == nil {
t.Fatalf("legacy entry was dropped; rulesets = %q", f.rulesetNames())
}
if !eqStrings(rs.lists["entry"], []string{"full:ads.example"}) {
t.Fatalf("entries = %q", rs.lists["entry"])
}
// Both references stand: rule_set is ORed, so the operator's list keeps
// matching everything it matched before.
if !eqStrings(f.rule(0).lists["dst_ruleset"], []string{"rule-ads", "rule-ads-2"}) {
t.Fatalf("dst_ruleset = %q", f.rule(0).lists["dst_ruleset"])
}
if _, ok := f.rule(0).lists["dst_domain"]; ok {
t.Fatal("dst_domain survived a completed migration")
}
}
// A legacy list whose values are all blank produces no rule-set — so the option
// must NOT be deleted. Deleting it left the rule with zero matchers, which is the
// spelling of a catch-all: a rule that previously matched nothing would have
// become the router's default route. Keeping it leaves the rule visibly
// unmigrated, which the parser then holds disabled.
func TestMigrate1to2KeepsALegacyListThatMigratesToNothing(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'blank'
list dst_domain ' '
option target 'direct'
config rule
option name 'default'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if names := f.rulesetNames(); len(names) != 0 {
t.Fatalf("an empty rule-set was created: %q", names)
}
if got := f.rule(0).lists["dst_domain"]; !eqStrings(got, []string{" "}) {
t.Fatalf("dst_domain = %q, want it left in place", got)
}
text, _ := f.Export("shater")
m, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse: %v", err)
}
if IsCatchAll(m.Rules[0]) {
t.Fatal("a rule with an unmigrated dst_domain became a catch-all")
}
if m.Rules[0].Enabled {
t.Fatal("a rule with an unmigrated dst_domain stayed enabled")
}
// And it does not take the default route away from the rule that owns it.
if !IsCatchAll(m.Rules[1]) || !m.Rules[1].Enabled {
t.Fatal("the real catch-all was disturbed")
}
}
// An UNMIGRATED config (the migration never ran, or could not commit) must not
// silently reroute the whole router. This is the failure the parser guard exists
// for: `list dst_domain 'bank.ru'` + `option target 'direct'` used to parse as a
// rule with no matchers at all, i.e. a catch-all, and generate points route Final
// at the LAST catch-all — so one uncommitted `uci commit` sent every packet out
// the plain WAN with the router's real address.
func TestUnmigratedRuleIsNeverACatchAll(t *testing.T) {
const v1 = `package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'bank'
option enabled '1'
option order '10'
list dst_domain 'bank.ru'
option target 'direct'
config rule
option name 'ips'
option enabled '1'
option order '20'
list dst_ip '10.0.0.0/8'
option target 'group:corp'
config rule
option name 'default'
option enabled '1'
option order '100'
option target 'group:auto'
`
m, err := ParseUCIExport(v1)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(m.Rules) != 3 {
t.Fatalf("rules = %d, want 3", len(m.Rules))
}
for _, i := range []int{0, 1} {
r := m.Rules[i]
if len(r.LegacyDst) == 0 {
t.Fatalf("rule %q: the live legacy option went unnoticed", r.Name)
}
if IsCatchAll(r) {
t.Fatalf("rule %q became a catch-all — it would take over route Final", r.Name)
}
if r.Enabled {
t.Fatalf("rule %q stayed enabled with an unreadable destination", r.Name)
}
}
if m.Rules[0].LegacyDst[0] != "dst_domain=bank.ru" || m.Rules[1].LegacyDst[0] != "dst_ip=10.0.0.0/8" {
t.Fatalf("LegacyDst = %q / %q", m.Rules[0].LegacyDst, m.Rules[1].LegacyDst)
}
// The real default is untouched, and it is the only one.
if !IsCatchAll(m.Rules[2]) || !m.Rules[2].Enabled {
t.Fatal("the genuine catch-all was disturbed")
}
// The operator is told, through the ordinary config-warning channel.
warns := ValidateRules(m.Rules, nil)
if len(warns) != 2 {
t.Fatalf("warnings = %d (%v), want one per unmigrated rule", len(warns), warns)
}
for _, w := range warns {
if w.Section != "rule" || !strings.Contains(w.Message, "shaterd migrate") {
t.Fatalf("warning does not point at the fix: %+v", w)
}
}
// A profile may not hand such a rule its Enabled bit back either.
prof := &Profile{Name: "home", EnableRules: []string{"bank"}}
got, pwarns := ApplyProfileRuleOverrides(m.Rules, prof)
if got[0].Enabled {
t.Fatal("a profile re-enabled an unmigrated rule")
}
if len(pwarns) != 1 {
t.Fatalf("profile warnings = %v, want one refusal", pwarns)
}
// And after the migration the same config is fully live again.
f := newFakeUCI(v1)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
text, _ := f.Export("shater")
mm, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse migrated: %v", err)
}
for i, r := range mm.Rules {
if len(r.LegacyDst) != 0 {
t.Fatalf("rule %d still flagged unmigrated: %q", i, r.LegacyDst)
}
if !r.Enabled {
t.Fatalf("rule %d stayed disabled after a successful migration", i)
}
}
if len(ValidateRules(mm.Rules, nil)) != 0 {
t.Fatalf("a migrated config still warns: %v", ValidateRules(mm.Rules, nil))
}
}
// failingUCI fails one named write, so the abort path can be observed.
type failingUCI struct {
*fakeUCI
failAddList bool
failDelete bool
}
var errFake = errors.New("uci: no space left on device")
func (f *failingUCI) AddList(k, v string) error {
if f.failAddList {
return errFake
}
return f.fakeUCI.AddList(k, v)
}
func (f *failingUCI) Delete(k string) error {
if f.failDelete && strings.Contains(k, "dst_") {
return errFake
}
return f.fakeUCI.Delete(k)
}
// A failed write must abort the migration AND drop the staged half-migration.
// Without the revert the partial rewrite sits in /tmp/.uci until the next
// `uci commit shater` from any process (the panel saving one setting) flushes a
// config that is half v1 and half v2 onto the disk.
func TestMigrate1to2RevertsOnWriteFailure(t *testing.T) {
f := &failingUCI{fakeUCI: newFakeUCI(legacyConfig), failAddList: true}
err := migrateWith(f)
if err == nil {
t.Fatal("expected the failed write to abort the migration")
}
if f.reverts == 0 {
t.Fatal("the staged half-migration was left behind (no revert)")
}
if f.commits != 0 {
t.Fatalf("committed %d times despite the failure", f.commits)
}
if v, _ := f.Get("shater.globals.schema_version"); v != "1" {
t.Fatalf("schema_version = %q — a failed migration must not claim v2", v)
}
}
// A delete that fails must abort too. Swallowing it left the legacy option in the
// config forever: migrateWith bumps schema_version to 2 on return, and the step
// that would have removed it never runs again.
func TestMigrate1to2FailsLoudlyWhenALegacyOptionCannotBeDeleted(t *testing.T) {
f := &failingUCI{fakeUCI: newFakeUCI(legacyConfig), failDelete: true}
if err := migrateWith(f); err == nil {
t.Fatal("a failed delete was swallowed")
}
if f.reverts == 0 {
t.Fatal("no revert after the failed delete")
}
if v, _ := f.Get("shater.globals.schema_version"); v == "2" {
t.Fatal("schema_version reached v2 with a legacy option still in the config")
}
}
// A rule carrying BOTH lists gets both rule-sets, and keeps every entry.
func TestMigrate1to2SplitsMixedRuleIntoTwoRulesets(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'mixed'
list dst_domain 'regexp:^ads\.'
list dst_ip '203.0.113.0/24'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
dom := f.ruleset("rule-mixed")
ip := f.ruleset("rule-mixed-ip")
if dom == nil || ip == nil {
t.Fatalf("want rule-mixed + rule-mixed-ip, got %q", f.rulesetNames())
}
// `regexp:` crosses over untouched (generate.peelDomainRegexes reads it).
if !eqStrings(dom.lists["entry"], []string{`regexp:^ads\.`}) {
t.Fatalf("domain entries = %q", dom.lists["entry"])
}
if !eqStrings(ip.lists["entry"], []string{"203.0.113.0/24"}) {
t.Fatalf("ip entries = %q", ip.lists["entry"])
}
if !eqStrings(f.rule(0).lists["dst_ruleset"], []string{"rule-mixed", "rule-mixed-ip"}) {
t.Fatalf("dst_ruleset = %q", f.rule(0).lists["dst_ruleset"])
}
}
// An inert `geosite:` matcher is copied verbatim rather than promoted to a
// source=geosite list: it matched nothing before the upgrade (the engine's
// route-rule geosite field is gone) and must not start routing traffic because of
// one. The text is kept so the operator can see and convert it.
func TestMigrate1to2KeepsGeoMarkersInert(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'geo'
list dst_domain 'geosite:youtube'
list dst_ip 'geoip:ru'
option target 'direct'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if got := f.ruleset("rule-geo").lists["entry"]; !eqStrings(got, []string{"geosite:youtube"}) {
t.Fatalf("domain entries = %q", got)
}
if got := f.ruleset("rule-geo-ip").lists["entry"]; !eqStrings(got, []string{"geoip:ru"}) {
t.Fatalf("ip entries = %q", got)
}
for _, s := range f.secs {
if s.typ == "ruleset" && s.opts["source"] != "inline" {
t.Fatalf("ruleset %q got source %q — a geo marker was promoted", s.opts["name"], s.opts["source"])
}
}
}
// A fresh install has no config to export; the migration must succeed silently
// rather than refuse to boot.
func TestMigrate1to2NoConfig(t *testing.T) {
f := &fakeUCI{missing: true}
if err := migrate1to2(f); err != nil {
t.Fatalf("migrate on an absent package: %v", err)
}
}
func TestRulesetBaseName(t *testing.T) {
cases := []struct{ in, want string }{
{"ru-direct", "rule-ru-direct"},
{"corp nets!", "rule-corp-nets"},
{" spaced name ", "rule-spaced-name"},
{"Ünïcode", "rule-n-code"}, // non-ASCII is not tag-safe; dropped, leaving a separator
{"", "rule-7"},
{"!!!", "rule-7"},
}
for _, c := range cases {
if got := rulesetBaseName(c.in, 7); got != c.want {
t.Errorf("rulesetBaseName(%q) = %q, want %q", c.in, got, c.want)
}
}
}
func TestMigrateDomainEntry(t *testing.T) {
cases := []struct{ in, want string }{
{"example.com", "full:example.com"}, // bare: exact in a rule, suffix in a list
{"full:example.com", "full:example.com"},
{"suffix:example.com", "suffix:example.com"},
{"keyword:ads", "keyword:ads"},
{`regexp:^a\.b$`, `regexp:^a\.b$`},
{".example.com", ".example.com"},
{"geosite:youtube", "geosite:youtube"},
{" spaced.example ", "full:spaced.example"},
{" ", ""},
}
for _, c := range cases {
if got := migrateDomainEntry(c.in); got != c.want {
t.Errorf("migrateDomainEntry(%q) = %q, want %q", c.in, got, c.want)
}
}
}
+35 -2
View File
@@ -633,19 +633,52 @@ type Ruleset struct {
}
// Rule is a `config rule` (ordered, first-match).
//
// DESTINATION IS ALWAYS A RULE-SET. A rule names WHERE traffic is going only
// through DstRuleset — there are no inline domain or IP lists on a rule any
// more (`dst_domain` / `dst_ip` were removed in schema v2; migrate1to2 folds
// every existing one into a generated `config ruleset` and rewrites the rule to
// point at it). One destination mechanism means one set of matcher semantics to
// learn, one place a list is edited, and a list that is compiled once into a
// .srs and shared by every rule that references it instead of being re-parsed
// per rule. Src (the CLIENT side), DstPort and Proto are unaffected: they are
// not lists of destinations and have no rule-set form.
type Rule struct {
Name string
Enabled bool
Order int
Src []string
DstDomain []string
DstRuleset []string
DstIP []string
DstPort string
Proto string
Target string // chain:|group:|node:|direct|block
Egress string
// LegacyDst is the UNMIGRATED-RULE TRIPWIRE, and it is a safety device, not a
// data field. It is non-empty exactly when the config STILL carries a
// schema-v1 `dst_domain`/`dst_ip` on this rule — i.e. migrate1to2 never ran,
// or ran and could not commit (a full /overlay is the documented way that
// happens). Each element is the raw `<option>=<value>` text, purely so the
// warning can quote what it found.
//
// WHY IT EXISTS. Nothing re-runs the migration on the paths that matter: the
// daemon's `run`, the SIGHUP reconcile and the panel's config write all load
// UCI directly. The parser no longer reads the two removed options, so an
// unmigrated `list dst_domain 'bank.ru'` + `option target 'direct'` parsed as
// a rule with NO matchers at all — a catch-all, which generate turns into
// route `Final`, which the LAST such rule wins. One un-committed migration
// therefore sent the WHOLE router's traffic out the plain WAN, silently.
//
// WHAT IT DOES. ParseUCIExport forces Enabled=false on such a rule (generate
// and the reachability analysis both skip disabled rules), IsCatchAll reports
// false for it whatever its matchers say (so it can never become the default
// route even if something re-enables it), ApplyProfileRuleOverrides refuses to
// enable it, and ValidateRules reports it through the normal apply-warning
// channel. It is never rendered back to UCI: render.go emits neither legacy
// option, so a panel write drains them out — with `enabled '0'` recorded, so
// the rule stays inert until an operator looks at it.
LegacyDst []string
// Kill is the per-rule policy for when the target cannot resolve at generate
// time (dead group, chain that would not assemble, missing egress/node). The
// rule is ALWAYS still emitted — its traffic never falls through to the
+16
View File
@@ -114,6 +114,22 @@ func ApplyProfileRuleOverrides(rules []Rule, prof *Profile) ([]Rule, []Warning)
continue
}
for _, i := range targets {
// A profile may switch a rule OFF freely, but it may not switch an
// UNMIGRATED one on: its destination matcher is unreadable (see
// Rule.LegacyDst), so enabling it would put a rule into force with
// fewer conditions than the operator wrote — in the worst case none
// at all, i.e. the router's default route.
if enabled && len(out[i].LegacyDst) > 0 {
warns = append(warns, Warning{
Section: "profile",
Name: prof.Name,
Message: fmt.Sprintf("cannot enable rule %q: it still carries the "+
"removed dst_domain/dst_ip options, so its destination cannot be "+
"read and it stays disabled until the config is migrated "+
"(run `shaterd migrate`)", n),
})
continue
}
out[i].Enabled = enabled
}
}
+210
View File
@@ -0,0 +1,210 @@
package model
// Rule reachability — "this rule can never fire", computed from the desired state
// alone.
//
// WHY THIS EXISTS. A routing rule with no conditions at all (no src, no dst
// domain/ip/ruleset, no port, no proto) is not a rule in the ordinary sense: it
// states a policy for EVERYTHING. generate does not emit it as a match-all route
// rule — it points the engine's route `Final` at that rule's target instead (see
// generate/route.go buildRoute). Two consequences follow, and neither is visible
// anywhere in the UI:
//
// 1. A condition-less rule can NEVER shadow a rule that has conditions. Every
// conditional rule is emitted ahead of `Final`, whatever its Order. So a
// "default -> direct" at order 20 does not hijack a specific rule at order 500.
// 2. Two condition-less rules DO shadow each other, and the LAST one in
// (Order, position) order wins, because the loop simply overwrites `Final`.
// Every earlier one is a dead setting that reads as a live one — the panel
// drew both with the same "default route · final" badge.
//
// A real config in the field had exactly that: two rules both named `default`,
// both with zero conditions, order 20 -> direct and order 100 -> group:auto. One
// of the two was doing nothing, and nothing said which.
//
// SCOPE, DELIBERATELY NARROW. This reports only the shadowing that is certain
// from the config: condition-less rule over condition-less rule. It does NOT try
// to decide whether one conditional rule's matcher set subsumes another's (is
// `dst_ruleset ru-inside` a superset of `dst_ip 5.0.0.0/8`? only the compiled
// rule-set knows) — a false "never fires" badge on a working rule would be worse
// than no badge at all.
//
// It lives in model, the stdlib-only leaf, because both consumers must agree:
// generate turns the verdict into an apply warning, and the panel API serves it
// to the Routing page. A second, drifting implementation in either is how the
// warning and the badge end up disagreeing about the same config.
import (
"fmt"
"sort"
"strings"
)
// IsCatchAll reports whether a rule carries NO matcher of any kind (src,
// dst_ruleset, port, proto). Such a rule is the default egress: generate points
// route `Final` at it rather than emitting a match-all rule.
//
// The destination side is now exactly one field — a rule names where traffic is
// going through DstRuleset alone (schema v2; see model.Rule). A rule that was
// catch-all before the migration is still catch-all after it, and a rule that
// carried `dst_domain`/`dst_ip` is not, because the migration gives it a
// DstRuleset in their place.
//
// AN UNMIGRATED RULE IS NEVER A CATCH-ALL. A rule that still carries a
// schema-v1 `dst_domain`/`dst_ip` (Rule.LegacyDst) has a destination the parser
// cannot read, so its lack of matchers here means "unreadable", not "everything".
// Calling it a catch-all is what turned an uncommitted migration into route
// `Final` for the whole router. ParseUCIExport already holds such a rule
// disabled; this is the second lock, and it is the one that holds if anything
// ever hands the rule back its Enabled bit.
//
// generate.isCatchAll and the panel's isCatchAll() are the same predicate; this
// is the one the Go side shares.
func IsCatchAll(r Rule) bool {
if len(r.LegacyDst) > 0 {
return false
}
return len(r.Src) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
}
// EffectiveRuleTarget resolves what a rule actually routes to: Target wins, and a
// rule with NO Target but a bare Egress routes to that egress outbound. "" means
// the rule states no policy at all — generate drops such a rule entirely (it does
// NOT mean "direct"), so it never becomes the default and never shadows anything.
func EffectiveRuleTarget(r Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
// SortedRuleIndices returns rule indices sorted by (Order, original index), so the
// engine's first-match order matches the model's declared order, stably. Equal
// Order keeps the configured/UCI order — the panel promises "lower Order runs
// earlier", and nothing below it may reshuffle ties.
func SortedRuleIndices(rules []Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
}
// RuleReach is one rule's reachability verdict. There is exactly one per input
// rule, at the same index, so a consumer can zip the two slices.
//
// It is the routing-rule analogue of engine.ChainHealth.Used / GroupHealth.Used —
// a quiet note about the ROUTING CONFIG, never a health or liveness signal.
type RuleReach struct {
// Index is the rule's position in the model's Rules slice (NOT the sorted
// order), so the panel can match a verdict to the row it drew.
Index int `json:"index"`
// Name and Order are echoed so a consumer holding a possibly-stale config can
// verify the verdict still describes the rule it is about to badge, instead of
// accusing the wrong row.
Name string `json:"name"`
Order int `json:"order"`
// Unreachable is true when this rule can never take effect, whatever the
// traffic. false is the normal case and carries no claim that the rule ever
// actually matches something — only that nothing in the config stops it.
Unreachable bool `json:"unreachable"`
// ShadowedBy / ShadowedByOrder name the rule that supersedes this one; empty /
// zero when Unreachable is false. ShadowedByIndex is -1 when there is none.
ShadowedBy string `json:"shadowed_by,omitempty"`
ShadowedByIndex int `json:"shadowed_by_index"`
ShadowedByOrder int `json:"shadowed_by_order,omitempty"`
// Reason is the operator-facing sentence; "" when Unreachable is false.
Reason string `json:"reason,omitempty"`
}
// RuleReachability returns one verdict per rule, in the INPUT slice's order.
//
// The input must already be the EFFECTIVE rule set — profile enable/disable
// applied (see Model.EffectiveRules). A rule the active profile switched off is
// not in force, and a rule it switched ON is, whatever /etc/config/shater says;
// judging the raw slice would badge the wrong rules on a router that uses
// profiles at all.
//
// A rule takes part in the analysis only when it is Enabled, condition-less, and
// carries a target — those are exactly the rules generate lets set `Final`.
//
// SCHEDULES. A rule that only holds inside a time window never marks anything
// permanently dead: outside its window the rule below it is the default again. So
// a SCHEDULED catch-all is skipped as a shadower (but can still BE shadowed — an
// unscheduled catch-all after it wins at every hour of every day, which makes the
// schedule pure decoration and is worth saying out loud).
func RuleReachability(rules []Rule) []RuleReach {
out := make([]RuleReach, len(rules))
for i := range rules {
out[i] = RuleReach{
Index: i, Name: rules[i].Name, Order: rules[i].Order, ShadowedByIndex: -1,
}
}
// Walk the declared order BACKWARDS and remember the last unconditional
// default seen. When we reach a rule, `winner` holds the default that is in
// force below it — i.e. the one whose target the engine actually uses.
winner := -1
for _, i := range reverse(SortedRuleIndices(rules)) {
r := rules[i]
if !r.Enabled || !IsCatchAll(r) || EffectiveRuleTarget(r) == "" {
continue
}
if winner >= 0 {
w := rules[winner]
out[i].Unreachable = true
out[i].ShadowedBy = w.Name
out[i].ShadowedByIndex = winner
out[i].ShadowedByOrder = w.Order
out[i].Reason = fmt.Sprintf(
"this rule has no conditions, so it sets the default for all traffic — but rule %q "+
"(order %d) has none either and comes after it, so %q is the default the router "+
"uses and this rule's target %q is never applied",
w.Name, w.Order, EffectiveRuleTarget(w), EffectiveRuleTarget(r))
}
// Only an UNSCHEDULED default holds at every hour, so only it can retire the
// rules above it. Keep the LAST one (highest Order): that is the one whose
// target the engine ends up with.
if !r.SchedEnabled && winner < 0 {
winner = i
}
}
return out
}
// reverse returns idx walked back to front. Small enough to inline by hand, but
// naming it keeps the loop above readable as "walk the declared order backwards".
func reverse(idx []int) []int {
out := make([]int, len(idx))
for i, v := range idx {
out[len(idx)-1-i] = v
}
return out
}
// EffectiveRules returns the rule set that is actually in force: a COPY of m.Rules
// with the active profile's enable/disable overrides applied, plus whatever
// warnings that resolution produced. m.Rules is never mutated.
//
// It is the shared front door for every consumer that must reason about "which
// rules are in force right now" without building an engine config — the panel's
// reachability endpoint today. generate keeps its own call site because it also
// needs the resolved *Profile itself (endpoint-resolver override); both go through
// ResolveActiveProfile + ApplyProfileRuleOverrides, so they cannot disagree.
func (m *Model) EffectiveRules() ([]Rule, []Warning) {
if m == nil {
return nil, nil
}
prof, warns := ResolveActiveProfile(m)
rules, owarns := ApplyProfileRuleOverrides(m.Rules, prof)
return rules, append(warns, owarns...)
}
+313
View File
@@ -0,0 +1,313 @@
package model
import "testing"
// catchAll builds a condition-less rule — the shape that becomes the engine's
// route Final.
func catchAll(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target}
}
// specific builds a rule with one matcher, so it is emitted as a real route rule
// ahead of Final and can never be retired by a default.
func specific(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target, DstPort: "443"}
}
// verdicts indexes a reachability run by rule index, asserting the contract that
// there is exactly one verdict per rule, at the same index.
func verdicts(t *testing.T, rules []Rule) []RuleReach {
t.Helper()
out := RuleReachability(rules)
if len(out) != len(rules) {
t.Fatalf("want %d verdicts, got %d", len(rules), len(out))
}
for i := range out {
if out[i].Index != i || out[i].Name != rules[i].Name {
t.Fatalf("verdict %d is about %q (index %d), want %q", i, out[i].Name, out[i].Index, rules[i].Name)
}
}
return out
}
func wantReachable(t *testing.T, v RuleReach) {
t.Helper()
if v.Unreachable {
t.Fatalf("rule %q (order %d) must be reachable, got shadowed by %q: %s",
v.Name, v.Order, v.ShadowedBy, v.Reason)
}
if v.ShadowedBy != "" || v.ShadowedByIndex != -1 || v.Reason != "" {
t.Fatalf("rule %q is reachable but carries shadow details: by=%q idx=%d reason=%q",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.Reason)
}
}
func wantShadowed(t *testing.T, v RuleReach, by string, byIndex, byOrder int) {
t.Helper()
if !v.Unreachable {
t.Fatalf("rule %q (order %d) must be unreachable, shadowed by %q", v.Name, v.Order, by)
}
if v.ShadowedBy != by || v.ShadowedByIndex != byIndex || v.ShadowedByOrder != byOrder {
t.Fatalf("rule %q: shadowed by %q(idx %d, order %d), want %q(idx %d, order %d)",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.ShadowedByOrder, by, byIndex, byOrder)
}
if v.Reason == "" {
t.Fatalf("rule %q is unreachable but carries no reason", v.Name)
}
}
// TestReachabilitySingleCatchAll: the ordinary config — some specific rules and
// ONE default. Nothing is retired, including the default itself.
func TestReachabilitySingleCatchAll(t *testing.T) {
rules := []Rule{
specific("block-ads", 10, "block"),
specific("ru-bypass", 20, "direct"),
catchAll("default-tunnel", 900, "group:auto"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityTwoCatchAllsFieldConfig is the config that prompted this
// analysis, reproduced exactly: two rules BOTH named `default`, both with zero
// conditions, order 20 -> direct and order 100 -> group:auto.
//
// The later one is the default the engine uses (buildRoute overwrites Final as it
// walks ascending Order), so it is the order-20 `direct` rule that is dead — the
// opposite of what "first match wins" would suggest, which is exactly why this
// needed saying out loud.
func TestReachabilityTwoCatchAllsFieldConfig(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
catchAll("default", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "default", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityCatchAllBelowSpecificRules: a default sorted BELOW specific
// rules retires none of them. A condition-less rule becomes route Final, which is
// evaluated after every emitted rule whatever its Order — so even a specific rule
// ordered after the default still matches first.
func TestReachabilityCatchAllBelowSpecificRules(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
specific("work-vpn", 100, "group:work"),
specific("ads", 200, "block"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityThreeCatchAlls: only the last default survives; every earlier
// one is blamed on that same last one, not on its immediate successor — that is
// the rule whose target the router actually ends up using.
func TestReachabilityThreeCatchAlls(t *testing.T) {
rules := []Rule{
catchAll("a", 10, "direct"),
catchAll("b", 20, "block"),
catchAll("c", 30, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "c", 2, 30)
wantShadowed(t, v[1], "c", 2, 30)
wantReachable(t, v[2])
}
// TestReachabilityDisabledCatchAllIsInert: a disabled default neither dies nor
// kills. It installs nothing, so badging it "never fires" would be noise, and it
// must not be credited with retiring the live default above it.
func TestReachabilityDisabledCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "parked", Enabled: false, Order: 100, Target: "direct"},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityTargetlessCatchAllIsInert: a rule with neither Target nor
// Egress states no policy, so generate drops it (it does NOT mean "direct"). It
// never becomes Final, so it cannot retire the default above it — and is not
// itself reported here, because "no target set" is its own, more useful warning.
func TestReachabilityTargetlessCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "empty", Enabled: true, Order: 100},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityBareEgressCountsAsTarget: a rule carrying only `option egress
// wan2` routes to that egress, so it IS a default and does retire the one above.
func TestReachabilityBareEgressCountsAsTarget(t *testing.T) {
rules := []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "wan2", Enabled: true, Order: 100, Egress: "wan2"},
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "wan2", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityScheduledCatchAllNeverRetires: a default that only holds inside
// a time window leaves the one above it live for the rest of the day, so it
// retires nothing.
func TestReachabilityScheduledCatchAllNeverRetires(t *testing.T) {
rules := []Rule{
catchAll("all-day", 20, "group:auto"),
{Name: "nightly", Enabled: true, Order: 100, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityScheduledCatchAllCanBeRetired: the converse — an unscheduled
// default AFTER a scheduled one wins at every hour of every day, so the schedule
// is pure decoration and the scheduled rule is dead.
func TestReachabilityScheduledCatchAllCanBeRetired(t *testing.T) {
rules := []Rule{
{Name: "nightly", Enabled: true, Order: 20, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
catchAll("all-day", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "all-day", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityEqualOrderKeepsConfiguredOrder: ties break by position in the
// slice (SortedRuleIndices is stable), so with equal Order the SECOND section in
// /etc/config/shater is the one that wins — same as generate.
func TestReachabilityEqualOrderKeepsConfiguredOrder(t *testing.T) {
rules := []Rule{
catchAll("first", 50, "direct"),
catchAll("second", 50, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "second", 1, 50)
wantReachable(t, v[1])
}
// TestReachabilityUnsortedInputIsJudgedByOrder: the model slice arrives in UCI
// order, which need not be Order order. The verdict must follow Order, and the
// returned slice must stay aligned with the INPUT indices.
func TestReachabilityUnsortedInputIsJudgedByOrder(t *testing.T) {
rules := []Rule{
catchAll("late", 100, "group:auto"),
catchAll("early", 20, "direct"),
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantShadowed(t, v[1], "late", 0, 100)
}
// TestReachabilityMatcherKindsAreNotCatchAll: every matcher field on its own is
// enough to make a rule conditional, so none of these is retired by the default
// below. A field this misses would silently badge a working rule "never fires".
func TestReachabilityMatcherKindsAreNotCatchAll(t *testing.T) {
conditional := []Rule{
{Name: "by-src", Enabled: true, Order: 10, Target: "direct", Src: []string{"192.168.1.0/24"}},
// The destination side is one field now (schema v2): a domain list and an
// address list are both `config ruleset`s a rule points dst_ruleset at.
{Name: "by-ruleset", Enabled: true, Order: 13, Target: "direct", DstRuleset: []string{"ads"}},
{Name: "by-port", Enabled: true, Order: 14, Target: "direct", DstPort: "443"},
{Name: "by-proto", Enabled: true, Order: 15, Target: "direct", Proto: "quic"},
}
for _, r := range conditional {
if IsCatchAll(r) {
t.Fatalf("rule %q carries a matcher but reads as a catch-all", r.Name)
}
}
rules := append(append([]Rule(nil), conditional...), catchAll("default", 900, "group:auto"))
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityWhitespaceOnlyMatcherIsCatchAll: a port/proto of spaces is not
// a matcher. generate trims before deciding, so this analysis must too — otherwise
// a rule the engine treats as the default reads as conditional here and its
// shadowing goes unreported.
func TestReachabilityWhitespaceOnlyMatcherIsCatchAll(t *testing.T) {
blank := Rule{Name: "blank", Enabled: true, Order: 20, Target: "direct", DstPort: " ", Proto: "\t"}
if !IsCatchAll(blank) {
t.Fatal("a rule whose only matchers are whitespace must read as a catch-all")
}
v := verdicts(t, []Rule{blank, catchAll("default", 100, "group:auto")})
wantShadowed(t, v[0], "default", 1, 100)
}
// TestReachabilityEmptyAndNil: no rules, no verdicts, no panic.
func TestReachabilityEmptyAndNil(t *testing.T) {
if got := RuleReachability(nil); len(got) != 0 {
t.Fatalf("nil rules: want no verdicts, got %d", len(got))
}
if got := RuleReachability([]Rule{}); len(got) != 0 {
t.Fatalf("empty rules: want no verdicts, got %d", len(got))
}
}
// TestEffectiveRulesAppliesActiveProfile: the analysis must judge the rules the
// active profile leaves in force. Here the profile DISABLES the later default, so
// the earlier one is live and nothing is retired — judging the raw slice would
// have badged the wrong rule.
func TestEffectiveRulesAppliesActiveProfile(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "home"},
Profiles: []Profile{
{Name: "home", Enabled: true, DisableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
catchAll("fallback", 100, "direct"),
},
}
eff, _ := m.EffectiveRules()
if eff[1].Enabled {
t.Fatal("active profile must have disabled the fallback rule")
}
if !m.Rules[1].Enabled {
t.Fatal("EffectiveRules must not mutate the model's own rules")
}
for _, v := range verdicts(t, eff) {
wantReachable(t, v)
}
// Same model, profile off: the later default is back in force and retires the
// tunnel default above it.
m.Globals.ActiveProfile = ""
m.Profiles[0].Enabled = false
eff, _ = m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
// TestEffectiveRulesProfileCanReviveADefault: a profile that force-ENABLES a
// disabled default makes it in force, and it then retires the default above it.
// The raw config would show it disabled and report nothing.
func TestEffectiveRulesProfileCanReviveADefault(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "travel"},
Profiles: []Profile{
{Name: "travel", Enabled: true, EnableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "fallback", Enabled: false, Order: 100, Target: "direct"},
},
}
eff, _ := m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
+212 -3
View File
@@ -9,7 +9,9 @@ package model
// the same uciRunner seam the migrations use).
import (
"errors"
"fmt"
"reflect"
"strconv"
"strings"
)
@@ -33,6 +35,14 @@ import (
// ParseUCIExport already carries those defaults as concrete values, so the
// panel's read→edit→write flow round-trips exactly. See render_test.go.
//
// Rule.LegacyDst IS emitted (as the `dst_domain`/`dst_ip` it names), because a
// write over an unmigrated config must not erase the operator's lists. This
// function is PURE and therefore trusts the Model it is given — so a Model whose
// LegacyDst did not come from the config on disk would have it written to disk.
// writeUCIWith is what guarantees the field's provenance (withDiskLegacyDst), and
// it is the only non-test caller; any future caller that persists the result owes
// the same guarantee.
//
// Subscription-cache nodes (FromSub != "") are NOT emitted at all: they live in
// per-subscription JSON cache files (see subcache.go) and are folded back in by
// ReadUCI's MergeSubCaches. Only manual nodes become `config node` sections, and
@@ -211,9 +221,12 @@ func RenderUCIExport(m *Model) string {
w.boolOpt("enabled", r.Enabled)
w.intOpt("order", r.Order)
w.listOpt("src", r.Src)
w.listOpt("dst_domain", r.DstDomain)
// dst_domain / dst_ip are never SYNTHESISED (they are not part of schema v2)
// — but they ARE written back when the rule still carries them, so a write
// that is allowed to proceed over an unmigrated config cannot erase them.
// See legacyDstOpts and guardUnmigratedConfig.
w.legacyDstOpts(r.LegacyDst)
w.listOpt("dst_ruleset", r.DstRuleset)
w.listOpt("dst_ip", r.DstIP)
w.strOpt("dst_port", r.DstPort)
w.strOpt("proto", r.Proto)
w.strOpt("target", r.Target)
@@ -376,6 +389,40 @@ func (w *uciWriter) listOpt(k string, vs []string) {
}
}
// legacyDstOpts re-emits the schema-v1 `dst_domain`/`dst_ip` a rule STILL carries
// on disk (Rule.LegacyDst, `<option>=<value>` text produced by legacyDstEntries).
//
// This is preservation, not support. A write over an unmigrated config is only
// ever allowed when it does not touch the rules at all (guardUnmigratedConfig) —
// a subscription refresh writing quota counters, the profile watcher switching
// the active profile. Those writers re-render the WHOLE package, so without this
// the operator's destination lists would be erased by a cron job, which is the
// same data loss the guard exists to prevent, just triggered from a different
// place. Writing them back keeps the config exactly as unmigrated as it was, so
// `shaterd migrate` can still do its job afterwards.
//
// Only the two known keys with a non-empty value are emitted: LegacyDst may be
// filled by a client (the panel PUTs back the model it GETs), and this must not
// become a way to inject arbitrary `list <anything>` lines into the config.
func (w *uciWriter) legacyDstOpts(entries []string) {
for _, e := range entries {
i := strings.IndexByte(e, '=')
if i <= 0 {
continue
}
k, v := e[:i], e[i+1:]
if v == "" {
continue
}
for _, known := range legacyDstOptions {
if k == known {
fmt.Fprintf(&w.b, "\tlist %s %s\n", k, renderQuote(v))
break
}
}
}
}
// renderQuote single-quotes a value and escapes embedded single quotes the uci
// way (`'\”`), mirroring uci.go's unquote.
//
@@ -423,8 +470,170 @@ func WriteUCI(m *Model) error {
return writeUCIWith(uci, m)
}
// ErrUnmigratedConfig is what every config write fails with when the config ON
// DISK still carries the schema-v1 rule destinations and the write would change
// the rules. Callers match it with errors.Is to tell this refusal (a config the
// operator can fix with one command) apart from a real I/O failure — the panel
// answers 409 with the message rather than a bare 500.
var ErrUnmigratedConfig = errors.New("config not migrated to schema v2")
// CheckConfigWritable reports whether persisting m would be refused, WITHOUT
// writing anything. It is the same check WriteUCI runs, exported so a caller that
// does several writes in one request can find out before the first of them lands:
// the panel's PUT /api/config writes the subscription node caches first, and
// discovering the refusal only at the UCI step would leave those caches rewritten
// for a request that was rejected.
func CheckConfigWritable(m *Model) error {
disk, ok := diskRules(uci)
return guardUnmigratedConfig(disk, ok, m)
}
// diskRules returns the rules of the config CURRENTLY on disk, and whether they
// could be read at all. ok=false covers a fresh install (no package) and an
// export that will not parse — the two cases where there is nothing on disk to
// protect and nothing to copy from.
//
// It is read ONCE per write and handed to both guardUnmigratedConfig and
// withDiskLegacyDst: they answer two halves of the same question ("may this write
// proceed?" / "what may it say about the legacy options?") and must not be able
// to see different configs.
func diskRules(u uciRunner) ([]Rule, bool) {
text, ok := u.Export("shater")
if !ok || strings.TrimSpace(text) == "" {
return nil, false
}
disk, err := ParseUCIExport(text)
if err != nil {
return nil, false
}
return disk.Rules, true
}
// withDiskLegacyDst returns m with every rule's LegacyDst forced to what the DISK
// says, so the renderer can never be told about a legacy option that is not
// really there. m is not modified: the rules are copied when (and only when)
// something actually differs.
//
// WHY. LegacyDst is a safety device, and PUT /api/config decodes the whole Model
// out of the request body — including this field. On a healthy, migrated config
// the guard above returns early (there is nothing to protect), and the renderer
// would then have written a fabricated `list dst_domain 'example.com'` straight
// into /etc/config/shater. No traffic leak — everything downstream fails closed —
// but an authenticated client could switch off any rule it named and lock the box
// out of saving its config until someone ran `shaterd migrate`, and the warning
// explaining it would have blamed a migration that never had anything to do with
// it. Taking the value from disk removes the input entirely.
//
// In the branch where the disk DOES carry legacy options, the guard has already
// established reflect.DeepEqual(disk.Rules, m.Rules), so the substitution is a
// no-op there by construction — this is purely the sanitiser for everything else.
// When the disk is unreadable there is nothing to preserve, so every LegacyDst is
// cleared: a fabricated one must never be the reason an option appears on disk.
func withDiskLegacyDst(disk []Rule, ok bool, m *Model) *Model {
if m == nil {
return m
}
want := func(i int) []string {
if !ok || i >= len(disk) {
return nil
}
return disk[i].LegacyDst
}
differs := false
for i := range m.Rules {
if !reflect.DeepEqual(m.Rules[i].LegacyDst, want(i)) {
differs = true
break
}
}
if !differs {
return m
}
out := *m
out.Rules = append([]Rule(nil), m.Rules...)
for i := range out.Rules {
out.Rules[i].LegacyDst = want(i)
}
return &out
}
// guardUnmigratedConfig refuses a write that would silently destroy schema-v1
// rule destinations still present on disk.
//
// WHY IT READS THE DISK AND NOT THE MODEL. The Model handed to WriteUCI is not
// trustworthy: PUT /api/config decodes one straight out of the request body, and
// the body belongs to the client (an older panel build, curl, a script). A check
// phrased as "refuse when m has LegacyDst" is defeated by simply omitting the
// field — and that omission is precisely the dangerous request, because
// RenderUCIExport would then write the rule with `enabled '1'` and NO destination
// at all, which the next read parses as a catch-all and generate turns into route
// `Final` for the entire router. Only the config already on disk can say whether
// there is anything to protect, so that is what is consulted.
//
// WHAT IT ALLOWS. A write whose rules are IDENTICAL to the ones on disk goes
// through, and legacyDstOpts writes the legacy options back with it. That keeps
// the non-panel writers working on an unmigrated box — a subscription refresh
// persisting quota counters and manual nodes (apply.refreshSubscription,
// `shaterd sub update`) and the profile watcher switching globals.active_profile
// — none of which has any business editing rules. They all build their Model with
// ReadUCI, so their rules ARE the disk's, byte for byte, including the LegacyDst
// the parser attached. Nothing about a migrated config changes: with no legacy
// options on disk the function returns nil before comparing anything.
//
// An unreadable/unparseable export is treated as "nothing to protect" (ok=false
// from diskRules). It cannot be distinguished from a fresh install here, and
// refusing every write on a malformed file would leave the box unconfigurable
// through its only UI.
func guardUnmigratedConfig(disk []Rule, ok bool, m *Model) error {
if !ok {
return nil // fresh install / unreadable: no config to lose
}
var stuck []string
for i, r := range disk {
if len(r.LegacyDst) == 0 {
continue
}
if r.Name != "" {
stuck = append(stuck, strconv.Quote(r.Name))
} else {
stuck = append(stuck, "#"+strconv.Itoa(i))
}
}
if len(stuck) == 0 {
return nil // migrated (the normal case): nothing to guard
}
if m != nil && reflect.DeepEqual(disk, m.Rules) {
return nil // the rules are untouched; legacyDstOpts carries them across
}
return fmt.Errorf("%w: rule %s still carr%s the removed dst_domain/dst_ip "+
"options, and this config has no place to write them — saving would erase "+
"the destination lists and leave the rule matching nothing (or, worse, "+
"everything). The write was REFUSED and nothing on disk changed. Run "+
"`shaterd migrate` on the router to fold each list into a rule-set, then "+
"reload and save again.",
ErrUnmigratedConfig, strings.Join(stuck, ", "), plural(len(stuck), "ies", "y"))
}
// plural picks the suffix for a count (carries/carry).
func plural(n int, one, many string) string {
if n == 1 {
return one
}
return many
}
func writeUCIWith(u uciRunner, m *Model) error {
text := RenderUCIExport(m)
// ONE read of the disk feeds both halves of the protection below.
disk, ok := diskRules(u)
// Refuse BEFORE the delete+import: writeUCIWith replaces the whole package, so
// by the time an import has run the legacy options are already gone.
if err := guardUnmigratedConfig(disk, ok, m); err != nil {
return err
}
// Render from the DISK's idea of which rules carry legacy options, never the
// caller's — RenderUCIExport is a pure function and will faithfully emit a
// `list dst_domain` for anything it is told about.
text := RenderUCIExport(withDiskLegacyDst(disk, ok, m))
var errs []string
// Clear the staged package first so `uci import` replaces rather than merges
// (import appends sections; without the delete a second write would duplicate
+210 -3
View File
@@ -1,6 +1,7 @@
package model
import (
"errors"
"reflect"
"strings"
"testing"
@@ -84,8 +85,7 @@ func richModel() *Model {
}},
Rules: []Rule{{
Name: "pc", Enabled: true, Order: 10,
Src: []string{"192.168.1.1/32"}, DstDomain: []string{"geosite:telegram"},
DstRuleset: []string{"ads"}, DstIP: []string{"1.1.1.1/32"},
Src: []string{"192.168.1.1/32"}, DstRuleset: []string{"ads"},
DstPort: "443", Proto: "tcp,udp", Target: "chain:triple", Egress: "frag",
Kill: "default", SchedEnabled: true, SchedDays: []string{"mon", "tue"},
SchedStart: "08:00", SchedEnd: "22:00", SchedUTCOffset: 180,
@@ -384,7 +384,7 @@ func TestRenderSkipsSubCacheNodes(t *testing.T) {
// and COMMITs — and the imported text re-parses to the original Model.
func TestWriteUCIReplaces(t *testing.T) {
m := richModel()
f := &fakeUCI{kv: map[string]string{}}
f := newFakeUCI("")
if err := writeUCIWith(f, m); err != nil {
t.Fatalf("writeUCIWith: %v", err)
}
@@ -418,3 +418,210 @@ func TestWriteUCIReplaces(t *testing.T) {
t.Fatalf("second write duplicated sections: nodes=%d rules=%d", len(got2.Nodes), len(got2.Rules))
}
}
// --- write-path guard: an unmigrated config may not be overwritten ------------
// unmigratedOnDisk is a v1 config as it sits on the router when `shaterd migrate`
// never got to commit: two rules whose destination is still an inline list, plus
// a subscription whose quota counters a refresh wants to update.
const unmigratedOnDisk = `package shater
config globals 'globals'
option enabled '1'
option schema_version '1'
config subscription
option name 'qomar'
option url 'https://example.invalid/sub'
config rule
option name 'bank'
option enabled '1'
list dst_domain 'bank.ru'
option target 'direct'
config rule
option name 'default'
option enabled '1'
option target 'group:auto'
`
// A write that would change the rules of an unmigrated config is REFUSED, and the
// on-disk config is byte-identical afterwards. Without this the renderer (which
// emits no dst_domain/dst_ip) simply dropped the operator's lists on the first
// save from the panel.
func TestWriteUCIRefusesToOverwriteAnUnmigratedConfig(t *testing.T) {
f := newFakeUCI(unmigratedOnDisk)
before, _ := f.Export("shater")
m, err := ParseUCIExport(before)
if err != nil {
t.Fatalf("parse: %v", err)
}
m.Rules[0].Order = 42 // any rule edit at all
err = writeUCIWith(f, m)
if err == nil {
t.Fatal("the write was allowed to erase the legacy destination lists")
}
if !errors.Is(err, ErrUnmigratedConfig) {
t.Fatalf("error is not ErrUnmigratedConfig: %v", err)
}
for _, want := range []string{`"bank"`, "dst_domain", "shaterd migrate", "REFUSED"} {
if !strings.Contains(err.Error(), want) {
t.Fatalf("message does not mention %q: %s", want, err)
}
}
if after, _ := f.Export("shater"); after != before {
t.Fatalf("the config on disk changed despite the refusal:\n--- before\n%s\n--- after\n%s", before, after)
}
if f.commits != 0 || len(f.deleted) != 0 {
t.Fatalf("the refused write still touched uci (commits=%d deleted=%v)", f.commits, f.deleted)
}
}
// The guard reads the DISK, not the submitted model — so a body that simply omits
// LegacyDst and sets Enabled cannot get a destination-less rule written with
// `enabled '1'`. That rule would parse back as a catch-all and become route Final
// for the whole router, which is the leak the parser guard closes on the read
// side and this closes on the write side.
func TestWriteUCIRefusesACraftedModelWithoutLegacyDst(t *testing.T) {
f := newFakeUCI(unmigratedOnDisk)
before, _ := f.Export("shater")
// Exactly what a hand-rolled PUT (or an older panel build) sends: the rule is
// there, enabled, and the field that marks it unmigrated is gone.
crafted := &Model{
Globals: DefaultGlobals(),
Rules: []Rule{
{Name: "bank", Enabled: true, Target: "direct"},
{Name: "default", Enabled: true, Target: "group:auto"},
},
}
if err := writeUCIWith(f, crafted); !errors.Is(err, ErrUnmigratedConfig) {
t.Fatalf("crafted body was accepted (err = %v)", err)
}
after, _ := f.Export("shater")
if after != before {
t.Fatalf("disk changed:\n--- before\n%s\n--- after\n%s", before, after)
}
if strings.Contains(after, "config rule\n\toption name 'bank'\n\toption enabled '1'\n\toption target") {
t.Fatal("a destination-less enabled rule reached the disk")
}
// And re-reading the disk still yields the protected shape.
m, err := ParseUCIExport(after)
if err != nil {
t.Fatalf("parse: %v", err)
}
if IsCatchAll(m.Rules[0]) || m.Rules[0].Enabled {
t.Fatal("the unmigrated rule lost its protection")
}
}
// A writer that does NOT touch the rules keeps working on an unmigrated box, and
// the legacy options survive the write. This is the subscription-refresh path
// (apply.UpdateSubscription / `shaterd sub update`) and the profile watcher: they
// build their model with ReadUCI, change a counter or globals.active_profile, and
// re-render the WHOLE package — so a blanket refusal would break a cron job, and
// a blanket allow would let that cron job erase the operator's lists.
func TestWriteUCIPreservesLegacyDstOnANonRuleWrite(t *testing.T) {
f := newFakeUCI(unmigratedOnDisk)
text, _ := f.Export("shater")
m, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse: %v", err)
}
// Exactly what subscribe.StoreUserInfo does.
m.Subscriptions[0].UserDownload = 1 << 40
m.Subscriptions[0].UserInfoAt = 1700000000
m.Globals.ActiveProfile = "home"
if err := writeUCIWith(f, m); err != nil {
t.Fatalf("a non-rule write was refused: %v", err)
}
after, _ := f.Export("shater")
got, err := ParseUCIExport(after)
if err != nil {
t.Fatalf("re-parse: %v", err)
}
if got.Subscriptions[0].UserDownload != 1<<40 || got.Globals.ActiveProfile != "home" {
t.Fatalf("the write did not land: %+v / %q", got.Subscriptions[0], got.Globals.ActiveProfile)
}
if !eqStrings(got.Rules[0].LegacyDst, []string{"dst_domain=bank.ru"}) {
t.Fatalf("the legacy destination list was erased: %q", got.Rules[0].LegacyDst)
}
if got.Rules[0].Enabled || IsCatchAll(got.Rules[0]) {
t.Fatal("the rule lost its unmigrated protection across the write")
}
// The config is still exactly as unmigrated as it was, so `shaterd migrate`
// can still fold the list into a rule-set afterwards.
if err := migrate1to2(f); err != nil {
t.Fatalf("migrate after the write: %v", err)
}
if f.ruleset("rule-bank") == nil {
t.Fatalf("migration found nothing to fold; rulesets = %q", f.rulesetNames())
}
}
// A fabricated LegacyDst in the submitted model must never reach the disk. On a
// MIGRATED config the guard has nothing to protect and returns early, so without
// withDiskLegacyDst the renderer happily wrote the client's `list dst_domain`
// into /etc/config/shater — and from there every downstream lock fired on a lie:
// the rule the client named went (and stayed) disabled, further rule writes were
// refused, and the warning blamed a migration that had never been involved.
func TestWriteUCIIgnoresAFabricatedLegacyDst(t *testing.T) {
const migratedOnDisk = `package shater
config globals 'globals'
option enabled '1'
option schema_version '2'
config rule
option name 'bank'
option enabled '1'
list dst_ruleset 'rule-bank'
option target 'direct'
`
f := newFakeUCI(migratedOnDisk)
text, _ := f.Export("shater")
m, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(m.Rules[0].LegacyDst) != 0 {
t.Fatal("fixture is not migrated")
}
// Exactly what a crafted PUT body carries.
m.Rules[0].LegacyDst = []string{"dst_domain=example.com", "dst_ip=203.0.113.0/24"}
if err := writeUCIWith(f, m); err != nil {
t.Fatalf("writeUCIWith: %v", err)
}
after, _ := f.Export("shater")
for _, forbidden := range []string{"dst_domain", "dst_ip"} {
if strings.Contains(after, forbidden) {
t.Fatalf("a fabricated %s reached the disk:\n%s", forbidden, after)
}
}
got, err := ParseUCIExport(after)
if err != nil {
t.Fatalf("re-parse: %v", err)
}
if len(got.Rules[0].LegacyDst) != 0 {
t.Fatalf("rule came back unmigrated: %q", got.Rules[0].LegacyDst)
}
// Enabled is whatever was submitted — the fabrication must not have switched
// the rule off either.
if !got.Rules[0].Enabled {
t.Fatal("the fabricated field disabled a healthy rule")
}
// And the caller's model was not mutated behind its back.
if len(m.Rules[0].LegacyDst) != 2 {
t.Fatalf("writeUCIWith mutated the caller's model: %q", m.Rules[0].LegacyDst)
}
// The config is still writable: nothing latched.
if err := writeUCIWith(f, got); err != nil {
t.Fatalf("the config became unwritable: %v", err)
}
}
@@ -74,6 +74,14 @@ func fullModel() *Model {
m.Nodes[0].FromSub = ""
m.Nodes[0].Fingerprint = ""
m.Nodes[0].Stale = false
// Rule.LegacyDst is exempt because it is not a field with a value of its own:
// each element is `<option>=<value>` naming one of exactly two removed options,
// and uciWriter.legacyDstOpts deliberately drops anything else so a client
// cannot inject arbitrary `list` lines through it. fillNonZero's generic
// "s-legacydst-1" is precisely such an anything-else, so it cannot round-trip
// by construction. The real round-trip (parse a legacy option -> render it back
// unchanged) is pinned by TestWriteUCIPreservesLegacyDstOnANonRuleWrite.
m.Rules[0].LegacyDst = nil
return m
}
+63 -4
View File
@@ -174,14 +174,39 @@ func ParseUCIExport(text string) (*Model, error) {
Entries: s.list("entry"),
})
case "rule":
// dst_domain / dst_ip were REMOVED in schema v2: a rule's destination is
// a rule-set reference and nothing else, and migrate1to2 (migrate.go)
// folds any legacy inline list into a `config ruleset` and points
// dst_ruleset at it.
//
// The migration is NOT re-run on every load — it runs from the service
// init and from uci-defaults at package install (`shaterd migrate`), and
// from nowhere else: the daemon's `run`, the SIGHUP reconcile and the
// panel's config write all come straight here. So a config that still
// carries the options is a real, reachable state — an interrupted or
// uncommittable migration (a full /overlay is the documented way that
// happens), or a hand edit after a downgrade.
//
// Ignoring them there was NOT safe. A rule whose only matcher was
// `dst_domain` parsed as a rule with NO matchers, which IS the spelling
// of a catch-all: generate points route `Final` at it and the last such
// rule wins, so `dst_domain bank.ru` + `target direct` quietly became
// "send EVERYTHING out the plain WAN". legacyDstEntries detects the live
// options; a rule that has them is held DISABLED here and is reported by
// ValidateRules, and Rule.LegacyDst keeps IsCatchAll/profile overrides
// from resurrecting it. See the field's doc comment in model.go.
legacyDst := legacyDstEntries(s)
m.Rules = append(m.Rules, Rule{
Name: firstNonEmpty(s.opt("name"), s.Name),
Enabled: s.optBool("enabled", true),
Name: firstNonEmpty(s.opt("name"), s.Name),
// An unmigrated rule is never in force. Its destination matcher is
// unreadable, so every alternative — routing it without the matcher,
// or treating the absence as "matches everything" — states a policy
// the operator did not write.
Enabled: s.optBool("enabled", true) && len(legacyDst) == 0,
Order: parseInt(s.opt("order"), 0),
Src: s.list("src"),
DstDomain: s.list("dst_domain"),
LegacyDst: legacyDst,
DstRuleset: s.list("dst_ruleset"),
DstIP: s.list("dst_ip"),
DstPort: s.opt("dst_port"),
Proto: s.opt("proto"),
Target: s.opt("target"),
@@ -335,6 +360,40 @@ func applyGlobals(g *Globals, s uciSection) {
g.StatsDiskLimitMB = parseInt(s.opt("stats_disk_limit_mb"), g.StatsDiskLimitMB)
}
// legacyDstOptions are the `config rule` options schema v2 removed. Their mere
// PRESENCE on a rule proves the config was not migrated (migrate1to2 deletes them
// as its last step per rule, and only once the replacement rule-set reference is
// in place).
var legacyDstOptions = []string{"dst_domain", "dst_ip"}
// legacyDstEntries returns the live schema-v1 destination entries of a rule
// section as `<option>=<value>` text, or nil when there are none.
//
// PRESENCE, not content, is the test. A `list dst_domain ' '` is kept as an
// entry even though it names no domain: the value is junk, but the OPTION being
// there still means the migration did not process this rule, and dropping the
// blank would hand the rule back its catch-all shape — the exact leak this
// detection exists to stop. (A truly empty `list dst_domain ”` never reaches
// here: parseSections drops empty list values, so the key is simply absent, and
// such a rule had no destination under v1 either.)
func legacyDstEntries(s uciSection) []string {
var out []string
for _, k := range legacyDstOptions {
vals, ok := s.Lists[k]
if !ok {
continue
}
if len(vals) == 0 {
out = append(out, k+"=")
continue
}
for _, v := range vals {
out = append(out, k+"="+v)
}
}
return out
}
// --- section accessors ---
func (s uciSection) opt(k string) string { return s.Options[k] }
+1 -1
View File
@@ -59,7 +59,7 @@ config rule
option enabled '1'
option order '10'
list src '192.168.11.14/32'
list dst_domain 'geosite:telegram'
list dst_ruleset 'ads'
option dst_port '443'
option proto 'tcp,udp'
option target 'chain:triple'
+22
View File
@@ -123,6 +123,28 @@ func ValidateRules(rules []Rule, classify NetClassifier) []Warning {
}
var out []Warning
for _, r := range rules {
// Checked BEFORE the Enabled gate on purpose: the parser holds an unmigrated
// rule disabled, so gating on Enabled would suppress the one warning that
// explains why the rule stopped working.
if len(r.LegacyDst) > 0 {
out = append(out, Warning{
Section: "rule",
Name: r.Name,
Message: fmt.Sprintf(
"this rule still carries the removed schema-v1 destination options "+
"(%s), so the config was never brought to schema v%d — either "+
"`shaterd migrate` never ran, or it ran and could not commit (a "+
"full /overlay is the usual cause), or the section was hand-edited. "+
"Its destination cannot be read, so the rule is HELD "+
"DISABLED: routing it without its destination would have applied "+
"target %q to far more traffic than you wrote it for, and a rule "+
"whose only matcher was one of these options would have become the "+
"router's DEFAULT ROUTE. Run `shaterd migrate` to convert them into "+
"a rule-set, then re-enable the rule.",
strings.Join(r.LegacyDst, ", "), CurrentSchemaVersion,
EffectiveRuleTarget(r)),
})
}
if !r.Enabled {
continue // a disabled rule installs nothing
}
+41 -5
View File
@@ -31,6 +31,12 @@ func ApplyNft(ruleset string) error {
if err := runNftStdin(ruleset, "-f", "-"); err != nil {
return fmt.Errorf("nft -f (load) failed: %w", err)
}
// The divert plane just changed underneath every flow that is currently
// tracked. For UDP :53 that is not self-correcting — see conntrack.go — so
// drop those entries and let the clients re-derive their path through the
// ruleset that is now loaded. Best-effort: a failed flush leaves exactly the
// behaviour we had before and must never fail an otherwise-good load.
_, _ = FlushDNSConntrack()
return nil
}
@@ -84,6 +90,11 @@ func TeardownNft() error {
if out, err := execCommand("nft", "delete", "table", "inet", "shater").CombinedOutput(); err != nil {
return fmt.Errorf("nft delete table inet shater: %v\n%s", err, out)
}
// Same reasoning as ApplyNft, mirrored: every DNS flow that was being
// delivered through the tproxy socket has just lost the rule that put it
// there. Without this, those entries survive into whatever plane comes next
// (including "no plane at all") and keep pointing at a socket that is gone.
_, _ = FlushDNSConntrack()
return nil
}
@@ -366,25 +377,50 @@ func isPointToPoint(dev string) bool {
return strings.Contains(string(out), "POINTOPOINT")
}
// RoutingPresent reports whether our fwmark ip rule is currently installed, for
// RoutingPresent reports whether our policy routing is currently installed, for
// EVERY family the model asks for. With Globals.IPv6 on, ApplyRouting installs
// both a -4 and a -6 rule, so checking only -4 was a half-truth: `ip -6 rule` is
// both a -4 and a -6 half, so checking only -4 was a half-truth: `ip -6 rule` is
// flushed independently (a `network reload` / `ifup` can drop one family and not
// the other), and the caller's idempotent fast-path would then conclude the plane
// was intact and never restore the missing v6 rule — leaving v6 clients diverted
// by nft but with nowhere to be delivered locally.
//
// With Globals.IPv6 off no v6 rule is installed BY DESIGN, so its absence must
// not be read as a missing plane; only the -4 rule is required then.
// not be read as a missing plane; only the -4 half is required then.
//
// # Both halves are checked, not just the rule
//
// ApplyRouting installs TWO things per family — the `fwmark -> table` rule AND
// the `local default dev lo` route inside that table — and they are removed by
// two INDEPENDENT commands (`ip rule del`, `ip route flush table`). Checking only
// the rule made the second half invisible: a table that had been flushed while
// its rule survived (a teardown interrupted part-way, an `ip route flush` from
// any other actor) read as "plane intact", so applyLocked's fast-path skipped
// ApplyRouting forever and NOTHING re-created the route. The visible result is a
// box that reports plane=full / engine_running=true while diverted packets are
// marked, find an empty table, fall through to the main table and are handed to
// the fail-closed forward drop — a permanent, healthy-looking outage that a
// reconcile cannot repair, because a reconcile is exactly what consults this
// function. A presence check must cover everything its Apply counterpart
// installs, or the idempotent fast-path becomes a trap.
func RoutingPresent(g model.Globals) bool {
want := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
wantRule := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
table := fmt.Sprintf("%d", effTable(g))
fams := []string{"-4"}
if g.IPv6 {
fams = append(fams, "-6")
}
for _, fam := range fams {
out, err := execCommand("ip", fam, "rule", "show").Output()
if err != nil || !strings.Contains(string(out), want) {
if err != nil || !strings.Contains(string(out), wantRule) {
return false
}
// `ip -4 route show table N` prints "local default dev lo scope host";
// the v6 form is "local default dev lo metric 1024 pref medium". Matching
// the route TYPE + destination covers both without pinning the trailing
// attributes, which differ by family and iproute2 version.
rout, rerr := execCommand("ip", fam, "route", "show", "table", table).Output()
if rerr != nil || !strings.Contains(string(rout), "local default") {
return false
}
}
+78
View File
@@ -0,0 +1,78 @@
package netplane
// Conntrack maintenance for plane transitions.
//
// # Why the data plane has to touch conntrack at all
//
// Everything else in this package is *stateless* from the kernel's point of view:
// an nft ruleset, a couple of `ip rule`/`ip route` entries and a handful of
// sysctls. Rebuilding them is atomic and idempotent, so a rebuilt plane is
// indistinguishable from a freshly-installed one — EXCEPT for one thing the
// rebuild cannot reach: the connection-tracking entries that were established
// while the previous plane (or a half-removed one) was in force.
//
// That matters here because our divert is a TPROXY divert. A `tproxy` statement
// hands the packet to a LOCAL TRANSPARENT SOCKET; when that socket is gone the
// statement evaluates to NFT_BREAK and the packet takes a completely different
// path through the ruleset. A restart necessarily walks through such a window:
// the outgoing daemon closes its engine BEFORE it removes the table (Teardown
// order — deliberately, because the reverse order would open a plaintext leak),
// and the incoming daemon starts its engine BEFORE it loads the new table. Any
// flow that crosses one of those windows keeps a conntrack entry that was formed
// against a plane that no longer exists, and — for UDP, which has no handshake to
// resynchronise on — every retry merely refreshes that entry instead of
// re-deriving the path.
//
// # What this is NOT
//
// This flush is HYGIENE, not the cure for the "DNS to the router's own LAN
// address never comes back after a restart" report (B3). That turned out to be a
// socket-level collision: the engine's TPROXY UDP write-back sockets were bound
// UNCONNECTED to the original destination — the router's own LAN address :53 —
// and so joined the kernel's demultiplex set next to dnsmasq's socket, silently
// swallowing the host's own queries. The fix for that lives in the engine
// (protocol/redirect/tproxy.go, lx:tproxy_writeback_connect); the stale
// `[UNREPLIED]` conntrack entry seen on the live box was a CONSEQUENCE of the
// unanswered query, not its cause.
//
// It is kept because it is independently correct and costs one netlink
// round-trip per plane change: a DNS flow that was mid-flight across a plane
// rebuild has a conntrack entry describing a delivery path that no longer
// exists, and dropping it makes the first query after a restart re-derive its
// path immediately instead of waiting out a retry.
//
// # Scope: :53/UDP only, deliberately
//
// A blanket `conntrack -F` would also delete the entries behind the operator's
// SSH session, the LuCI session and the admin panel — fw4's input chain accepts
// them via `ct state established,related`, so dropping their conntrack entries
// drops the sessions. Locking the admin out of the box while "fixing" DNS is not
// a trade we get to make. UDP/:53 is the narrowest cut that covers the observed
// failure class: DNS is retried by every client within a second, so deleting its
// entries costs nothing and is invisible.
//
// Best-effort by contract: a kernel without conntrack, a netlink permission
// error or a non-Linux build all report zero deletions and no error path that can
// fail an apply. Losing the flush degrades to the old behaviour; it must never
// take a working plane down.
// dnsPort is the only port whose conntrack entries we touch. See the package
// comment for why this is deliberately not "everything".
const dnsPort uint16 = 53
// flushUDPPortConntrack is the platform seam. The default is a no-op so the
// package builds (and `go vet`s) on non-Linux dev hosts; conntrack_linux.go's
// init() replaces it with the real ctnetlink delete on the router target. Tests
// substitute a recorder.
var flushUDPPortConntrack = func(port uint16) (int, error) { return 0, nil }
// FlushDNSConntrack deletes every UDP connection-tracking entry whose ORIGINAL
// destination port is 53, for both address families, and returns how many were
// removed.
//
// Called on every plane transition that can change where a DNS packet is
// delivered: after a ruleset is loaded (ApplyNft — the full plane, the holding
// plane and a rollback all go through it) and after the table is removed
// (TeardownNft). Idempotent and cheap: on an idle box the DNS entry count is a
// handful, and on a busy one it is bounded by the number of clients.
func FlushDNSConntrack() (int, error) { return flushUDPPortConntrack(dnsPort) }
+62
View File
@@ -0,0 +1,62 @@
//go:build linux
package netplane
// The real ctnetlink implementation of the conntrack seam declared in
// conntrack.go. It lives behind a build tag for the same reason apply's flock
// does: the netlink conntrack API only exists on Linux, and the control plane
// must still build and test on a developer's Windows/macOS host.
//
// Deliberately netlink and NOT the `conntrack` CLI: conntrack-tools is not a
// dependency of shater-core (and pulling it in for one call would add ~100 KiB
// of userland to a flash-constrained router), so a shell-out would silently
// no-op on every real box — the worst possible outcome for a fix whose entire
// job is to remove stale state.
import (
"github.com/sagernet/netlink"
"golang.org/x/sys/unix"
)
func init() { flushUDPPortConntrack = ctnetlinkFlushUDPPort }
// ctnetlinkFlushUDPPort deletes the UDP conntrack entries whose ORIGINAL
// destination port is `port`, in both families, and returns the total deleted.
//
// Errors are aggregated rather than short-circuited: v4 and v6 are independent
// tables and a failure on one must not hide a successful cleanup of the other.
// The first error is returned for logging; the count is still accurate for the
// families that succeeded.
func ctnetlinkFlushUDPPort(port uint16) (int, error) {
var (
total int
firstErr error
)
for _, family := range []netlink.InetFamily{
netlink.InetFamily(unix.AF_INET),
netlink.InetFamily(unix.AF_INET6),
} {
filter := &netlink.ConntrackFilter{}
// Protocol MUST be set before the port: AddPort refuses to add a port
// filter while the layer-4 protocol is unknown (a port means nothing
// without one), so the order here is load-bearing.
if err := filter.AddProtocol(unix.IPPROTO_UDP); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
if err := filter.AddPort(netlink.ConntrackOrigDstPort, port); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
n, err := netlink.ConntrackDeleteFilter(netlink.ConntrackTable, family, filter)
total += int(n)
if err != nil && firstErr == nil {
firstErr = err
}
}
return total, firstErr
}
+7
View File
@@ -449,6 +449,13 @@ func TestRoutingPresentBothFamilies(t *testing.T) {
out = c.v6
}
}
// RoutingPresent also verifies the `local default dev lo` route
// that ApplyRouting installs beside the rule (see
// TestRoutingPresentRequiresLocalDefaultRoute). This case set is
// about the RULE half, so the route half is always healthy here.
if name == "ip" && len(arg) >= 3 && arg[1] == "route" {
out = "local default dev lo scope host"
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
+210
View File
@@ -0,0 +1,210 @@
package netplane
// B3 regressions: what a restart leaves behind.
//
// The bug these pin: `/etc/init.d/shater restart` (stop immediately followed by
// start) left DNS to the router's own LAN address permanently dead, while
// `stop` + pause + `start` was fine and the status kept reporting plane=full /
// engine_running=true. Two independent defects fed it, and both live here:
//
// 1. A plane transition (load or teardown of the tproxy divert) left the
// CONNTRACK entries of flows that had crossed the transition pointing at a
// plane that no longer exists. Nothing in the tree touched conntrack, so a
// UDP flow — which has no handshake to resynchronise on and whose entry is
// refreshed by every retry — stayed wedged indefinitely.
// 2. RoutingPresent() reported "plane intact" from the ip RULE alone, ignoring
// the `local default dev lo` ROUTE that ApplyRouting installs alongside it.
// A teardown interrupted between the two (procd SIGKILL at term_timeout) or
// any other `ip route flush` therefore became invisible: applyLocked's
// idempotent fast-path skipped ApplyRouting forever and no reconcile could
// repair it.
import (
"os"
"os/exec"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// recordFlush swaps the conntrack seam for a counter and restores it after the
// test. Returns a pointer to the number of calls and the ports asked for.
func recordFlush(t *testing.T) *[]uint16 {
t.Helper()
orig := flushUDPPortConntrack
t.Cleanup(func() { flushUDPPortConntrack = orig })
var seen []uint16
flushUDPPortConntrack = func(port uint16) (int, error) {
seen = append(seen, port)
return 0, nil
}
return &seen
}
// TestApplyNftFlushesDNSConntrack: loading a ruleset must drop the DNS conntrack
// entries formed against the previous plane. Without this, a flow that crossed
// the restart window keeps being delivered by a rule set that is gone.
func TestApplyNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err != nil {
t.Fatalf("ApplyNft: %v", err)
}
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("ApplyNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
// Ordering matters: the flush is only meaningful once the NEW ruleset is in
// the kernel, otherwise the very next packet re-creates the entry against the
// old plane. The load is the last nft invocation before it.
if len(rec) != 2 || !strings.Contains(rec[0], "-c") || strings.Contains(rec[1], "-c") {
t.Fatalf("expected validate-then-load, got %v", rec)
}
}
// TestApplyNftDoesNotFlushOnFailure: a ruleset that does not load leaves the
// PREVIOUS plane in charge (nft -f is one netlink transaction). Flushing then
// would tear down live flows for nothing.
func TestApplyNftDoesNotFlushOnFailure(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err == nil {
t.Fatalf("ApplyNft must report the nft failure")
}
if len(*seen) != 0 {
t.Fatalf("a failed load must not flush conntrack, got %v", *seen)
}
}
// TestTeardownNftFlushesDNSConntrack is the mirror: removing the table strands
// every DNS flow that was being delivered through the tproxy socket, so those
// entries must go with it.
func TestTeardownNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft: %v", err)
}
// fakeExec makes every command succeed, so TableExists() is true and the
// delete runs.
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("TeardownNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
}
// TestTeardownNftAbsentTableDoesNotFlush: nothing was diverting, so nothing is
// stranded. Keeps the flush out of the hot path of an idle box.
func TestTeardownNftAbsentTableDoesNotFlush(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft with no table must be a no-op, got %v", err)
}
if len(*seen) != 0 {
t.Fatalf("absent table must not flush conntrack, got %v", *seen)
}
}
// failingExec returns an execCommand replacement whose every command exits 1.
func failingExec() func(string, ...string) *exec.Cmd {
return func(name string, arg ...string) *exec.Cmd {
cs := append([]string{"-test.run=TestRestartHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_FAIL=1")
return cmd
}
}
// TestRestartHelperProcess is the exec helper for failingExec.
func TestRestartHelperProcess(t *testing.T) {
if os.Getenv("GO_WANT_HELPER_PROCESS") != "1" {
return
}
if os.Getenv("GO_HELPER_FAIL") == "1" {
os.Exit(1)
}
os.Exit(0)
}
// TestRoutingPresentRequiresLocalDefaultRoute is the second half of B3.
//
// ApplyRouting installs TWO things per family and they are removed by two
// independent commands. A presence check that only looks at the rule declares a
// half-removed plane healthy — and because applyLocked consults exactly this
// function to decide whether to re-run ApplyRouting, the missing route is then
// never restored: marked packets find an empty table, fall through to main and
// are eaten by the fail-closed forward drop, permanently, with the status still
// saying plane=full.
//
// On the pre-fix implementation the first two cases below return true.
func TestRoutingPresentRequiresLocalDefaultRoute(t *testing.T) {
const (
rule = "32765:\tfrom all fwmark 0x2000 lookup shater"
route = "local default dev lo scope host"
)
cases := []struct {
name string
ipv6 bool
v4rule, v4rte string
v6rule, v6rte string
want bool
}{
// THE REGRESSION: rule survived, table was flushed.
{"v4 rule present, route flushed", false, rule, "", "", "", false},
{"ipv6 on, v6 route flushed", true, rule, route, rule, "", false},
// Sanity: a complete plane is still reported as present.
{"v4 complete", false, rule, route, "", "", true},
{"ipv6 on, both complete", true, rule, route, rule, route, true},
// The pre-existing rule-level contract must not regress.
{"v4 rule missing", false, "", route, "", "", false},
{"ipv6 on, v6 rule missing", true, rule, route, "", route, false},
}
orig := execCommand
defer func() { execCommand = orig }()
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
execCommand = func(name string, arg ...string) *exec.Cmd {
out := ""
if name == "ip" && len(arg) >= 2 {
v6 := arg[0] == "-6"
switch arg[1] {
case "rule":
out = c.v4rule
if v6 {
out = c.v6rule
}
case "route":
out = c.v4rte
if v6 {
out = c.v6rte
}
}
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
return cmd
}
g := model.Globals{FwmarkBase: 0x2000, TableBase: 0x2000, IPv6: c.ipv6}
if got := RoutingPresent(g); got != c.want {
t.Errorf("RoutingPresent = %v, want %v", got, c.want)
}
})
}
}
+74
View File
@@ -128,6 +128,54 @@ func (s *Server) handleConfigGet(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, m)
}
// rulesReachabilityResponse is the GET /api/rules/reachability body. `rules` is
// ALWAYS an array, never null, with one entry per rule in the SAME order as
// GET /api/config's Rules — so a client can zip the two by `index` (and check the
// echoed `name`/`order` before trusting a verdict it fetched around an edit).
type rulesReachabilityResponse struct {
Rules []model.RuleReach `json:"rules"`
}
// reachConfigRead is handleRulesReachability's test seam (same pattern as
// writeConfig / logConfigRead): production binds the real UCI read, tests
// substitute a canned model so the handler can be exercised without a `uci`
// binary on the host.
var reachConfigRead = model.ReadUCI
// handleRulesReachability → GET /api/rules/reachability: which routing rules can
// never take effect, and what supersedes each of them.
//
// It is the routing-rule analogue of the per-chain `used` flag on
// GET /api/groups/health — a quiet note about the ROUTING CONFIG that the panel
// renders as a badge, never a health or liveness signal. Separate from
// /api/config on purpose: /api/config is the desired state the panel PUTs back
// verbatim, and a derived verdict has no business travelling round-trip through
// it.
//
// The verdict is computed over the EFFECTIVE rules (active WAN profile's
// enable/disable applied via model.EffectiveRules), because a rule the active
// profile switched off is not in force and must not be blamed for retiring
// anything. Profile-resolution warnings are dropped here: this endpoint answers
// one question, and the same warnings already reach the operator through
// `shaterd status` on every apply.
func (s *Server) handleRulesReachability(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
return
}
m, err := reachConfigRead()
if err != nil {
writeError(w, http.StatusInternalServerError, "read config: "+err.Error())
return
}
rules, _ := m.EffectiveRules()
out := model.RuleReachability(rules)
if out == nil {
out = []model.RuleReach{}
}
writeJSON(w, http.StatusOK, rulesReachabilityResponse{Rules: out})
}
// handleConfigPut → PUT /api/config: decode a Model, lightly validate it, and
// persist it. The write is SPLIT on the server: the FromSub nodes carried in the
// model are reconciled into the per-subscription JSON cache files
@@ -155,6 +203,22 @@ func (s *Server) handleConfigPut(w http.ResponseWriter, r *http.Request) {
writeError(w, http.StatusBadRequest, err.Error())
return
}
// Refuse to overwrite a config that was never migrated to schema v2, BEFORE
// touching anything. The check reads the config on disk, not this body: a
// request that simply omits the LegacyDst field is exactly the dangerous one
// (it would have the renderer write a rule with `enabled '1'` and no
// destination, which the next read treats as the router's default route), so
// trusting the body would defeat the guard. Running it here, ahead of
// syncSubCaches, keeps a rejected request from rewriting the node caches.
// 409, not 500: nothing is broken, the operator has one command to run.
if err := checkWritable(&m); err != nil {
if errors.Is(err, model.ErrUnmigratedConfig) {
writeError(w, http.StatusConflict, err.Error())
return
}
writeError(w, http.StatusInternalServerError, "check config: "+err.Error())
return
}
// Caches first: they are plain files and the cheaper write; a UCI failure
// after them leaves provider-owned node caches fresh and the manual config
// untouched, which the next successful PUT converges anyway.
@@ -163,6 +227,13 @@ func (s *Server) handleConfigPut(w http.ResponseWriter, r *http.Request) {
return
}
if err := writeConfig(&m); err != nil {
// The same refusal can still surface here (WriteUCI re-checks under its own
// read, so a migration that finished between the two calls cannot slip a
// destructive write through). Keep the status honest.
if errors.Is(err, model.ErrUnmigratedConfig) {
writeError(w, http.StatusConflict, err.Error())
return
}
writeError(w, http.StatusInternalServerError, "write config: "+err.Error())
return
}
@@ -183,6 +254,9 @@ const maxConfigBytes = 4 << 20 // 4 MiB
var (
writeConfig = model.WriteUCI
syncSubCaches = model.SyncSubCaches
// checkWritable is the pre-flight refusal (see handleConfigPut); it reads the
// on-disk config and never writes.
checkWritable = model.CheckConfigWritable
)
// validateModel does light structural validation — enough to reject obviously
+79
View File
@@ -3,6 +3,7 @@ package panel
import (
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/http/httptest"
@@ -227,3 +228,81 @@ func readAll(resp *http.Response) (string, error) {
b, err := io.ReadAll(resp.Body)
return string(b), err
}
// A PUT over a config that was never migrated to schema v2 is REFUSED with 409
// and a readable body, and nothing is written — not the UCI config and not the
// subscription node caches (the guard runs ahead of both).
//
// The crafted body is the dangerous shape the coordinator flagged: the rule is
// present and Enabled, and LegacyDst — the field that marks it unmigrated — is
// simply absent. Trusting the body would have the renderer write `enabled '1'`
// with no destination at all, which the next read parses as a catch-all and
// generate turns into route Final for the entire router. The guard consults the
// config on disk instead, which is why omitting the field changes nothing.
func TestConfigPutRefusesUnmigratedConfig(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
wrote, synced := false, false
origWrite, origSync, origCheck := writeConfig, syncSubCaches, checkWritable
writeConfig = func(*model.Model) error { wrote = true; return nil }
syncSubCaches = func(*model.Model) error { synced = true; return nil }
checkWritable = func(*model.Model) error {
return fmt.Errorf("%w: rule \"bank\" still carries the removed dst_domain/dst_ip "+
"options ... The write was REFUSED and nothing on disk changed. Run "+
"`shaterd migrate` on the router", model.ErrUnmigratedConfig)
}
defer func() { writeConfig, syncSubCaches, checkWritable = origWrite, origSync, origCheck }()
body := `{
"Globals": {"Enabled": true},
"Rules": [{"Name": "bank", "Enabled": true, "Target": "direct"}]
}`
resp := doPut(t, srv, cookie, body)
defer resp.Body.Close()
if resp.StatusCode != http.StatusConflict {
b, _ := readAll(resp)
t.Fatalf("PUT over an unmigrated config: got %d, want 409 (body %s)", resp.StatusCode, b)
}
if wrote {
t.Fatal("the config was written despite the refusal")
}
if synced {
t.Fatal("the subscription caches were rewritten for a rejected request")
}
var out map[string]string
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
t.Fatalf("decode error body: %v", err)
}
for _, want := range []string{"shaterd migrate", "REFUSED", "bank"} {
if !strings.Contains(out["error"], want) {
t.Fatalf("error body does not mention %q: %q", want, out["error"])
}
}
}
// The same refusal raised by WriteUCI itself (a migration that finished between
// the pre-flight check and the write, or any caller that skipped the pre-flight)
// must still reach the client as 409 with its text, not a bare 500.
func TestConfigPutMapsWriteRefusalTo409(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
origWrite, origSync, origCheck := writeConfig, syncSubCaches, checkWritable
writeConfig = func(*model.Model) error {
return fmt.Errorf("%w: rule \"bank\" ...", model.ErrUnmigratedConfig)
}
syncSubCaches = func(*model.Model) error { return nil }
checkWritable = func(*model.Model) error { return nil }
defer func() { writeConfig, syncSubCaches, checkWritable = origWrite, origSync, origCheck }()
resp := doPut(t, srv, cookie, `{"Globals": {"Enabled": true}}`)
defer resp.Body.Close()
if resp.StatusCode != http.StatusConflict {
t.Fatalf("got %d, want 409", resp.StatusCode)
}
}
+125
View File
@@ -0,0 +1,125 @@
package panel
import (
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// getReach GETs /api/rules/reachability (authenticated) and returns the status
// code plus the decoded reply.
func getReach(t *testing.T, srv *httptest.Server, cookie *http.Cookie) (int, rulesReachabilityResponse) {
t.Helper()
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET /api/rules/reachability: %v", err)
}
defer resp.Body.Close()
var out rulesReachabilityResponse
_ = json.NewDecoder(resp.Body).Decode(&out)
return resp.StatusCode, out
}
// TestRulesReachabilityEndpoint (B1): the endpoint reports the field config's two
// condition-less `default` rules, one verdict per rule, aligned with the order
// GET /api/config returns them in — that alignment is the only thing that lets the
// panel badge the right row when both rules share a name.
func TestRulesReachabilityEndpoint(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) {
return &model.Model{Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: "direct"},
{Name: "default", Enabled: true, Order: 100, Target: "group:auto"},
}}, nil
}
code, out := getReach(t, srv, cookie)
if code != http.StatusOK {
t.Fatalf("got %d, want 200", code)
}
if len(out.Rules) != 2 {
t.Fatalf("want one verdict per rule (2), got %d: %+v", len(out.Rules), out.Rules)
}
if !out.Rules[0].Unreachable {
t.Fatalf("the order-20 default is superseded by the order-100 one: %+v", out.Rules[0])
}
if out.Rules[0].Index != 0 || out.Rules[0].ShadowedByIndex != 1 || out.Rules[0].ShadowedByOrder != 100 {
t.Fatalf("verdict must point at the superseding rule by index and order: %+v", out.Rules[0])
}
if out.Rules[0].Reason == "" {
t.Fatal("an unreachable verdict must carry a reason the panel can show")
}
if out.Rules[1].Unreachable {
t.Fatalf("the last default is the one in force: %+v", out.Rules[1])
}
}
// TestRulesReachabilityEmptyIsArray: `rules` is ALWAYS an array. Go marshals a nil
// slice as null, and a client that does `for (const r of body.rules)` breaks on it.
func TestRulesReachabilityEmptyIsArray(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) { return &model.Model{}, nil }
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET: %v", err)
}
defer resp.Body.Close()
var raw struct {
Rules *[]model.RuleReach `json:"rules"`
}
if err := json.NewDecoder(resp.Body).Decode(&raw); err != nil {
t.Fatalf("decode: %v", err)
}
if raw.Rules == nil {
t.Fatal("rules must marshal as [], never null")
}
}
// TestRulesReachabilityMethodAndAuth: it is a read endpoint behind the session
// cookie, like every other /api route except /api/session.
func TestRulesReachabilityMethodAndAuth(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
req, _ := http.NewRequest(http.MethodPost, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("POST: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusMethodNotAllowed {
t.Fatalf("POST got %d, want 405", resp.StatusCode)
}
resp, err = http.Get(srv.URL + "/api/rules/reachability")
if err != nil {
t.Fatalf("unauthenticated GET: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusUnauthorized {
t.Fatalf("unauthenticated GET got %d, want 401", resp.StatusCode)
}
}
+1
View File
@@ -178,6 +178,7 @@ func (s *Server) buildRouter() http.Handler {
mux.Handle("/api/devices", s.requireSession(http.HandlerFunc(s.handleDevices)))
mux.Handle("/api/interfaces", s.requireSession(http.HandlerFunc(s.handleInterfaces)))
mux.Handle("/api/import-wg", s.requireSession(http.HandlerFunc(s.handleImportWG)))
mux.Handle("/api/rules/reachability", s.requireSession(http.HandlerFunc(s.handleRulesReachability)))
mux.Handle("/api/ruleset/status", s.requireSession(http.HandlerFunc(s.handleRuleSetStatus)))
mux.Handle("/api/ruleset/update", s.requireSession(http.HandlerFunc(s.handleRuleSetUpdate)))
mux.Handle("/api/ruleset/check", s.requireSession(http.HandlerFunc(s.handleRuleSetCheck)))