Compare commits

...
13 Commits
Author SHA1 Message Date
omarandClaude Opus 5 244b7c4199 feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.

The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.

The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.

isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.

Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 8fd5c52488 fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.

Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.

With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.

Fix (upstream file, lx:tproxy_writeback_connect):
  * CONNECT the write-back socket to the one peer it ever talks to. The kernel's
    compute_score() rejects a connected socket for any other peer, and a
    connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
    reuseport selection outright — so it can no longer be handed a datagram it
    will not read. Nothing about the reply changes: same spoofed source, same
    single peer, Write instead of WriteTo. The unconnected path is kept verbatim
    for a destination that cannot be bound (domain socksaddr).
  * A failed cached write now CLOSES the socket instead of only dropping the
    reference (upstream left the fd to the GC finalizer).
  * TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
    but the cache evicts lazily, so after the inbound is gone nothing wakes the
    live sessions and each strands its write-back socket. Invisible upstream
    (one close at shutdown); on this fork the engine is rebuilt on every apply,
    so it was one stranded generation per apply.

Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.

The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.

Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:41:02 +03:00
omarandClaude Opus 5 32e8f8ff0b fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.

Cause: `restart` is not synchronised end to end.

  * procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
    `stop; start`, and the `service delete` ubus call returns as soon as
    SIGTERM has been SENT. `start_service` therefore re-adds the instance
    (and runs `shaterd migrate`) while the outgoing `shaterd run` is still
    executing its honest teardown.
  * The successor's only defence was `daemonAlive()` -> exit(1), leaning on
    procd's `respawn 3600 5 0` to try again five seconds later. That is a
    blind retry, not synchronisation: it neither knows nor waits for the
    teardown, and it turns every restart into a logged crash plus a
    five-second hole with no data plane.
  * `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
    engine.Close of a several-hundred-outbound box flushes cache.db to
    flash before the netplane teardown even starts — aborting the teardown
    at an arbitrary point and leaving the plane HALF removed.
  * Nothing in the tree ever touched conntrack, so flows that crossed one
    of those windows kept entries formed against a plane that no longer
    exists. For UDP there is no handshake to resynchronise on and every
    retry merely refreshes the entry, so the flow stays wedged for as long
    as the client keeps asking — a flow-scoped, permanent failure that no
    ruleset rebuild can reach.
  * RoutingPresent() reported "plane intact" from the ip RULE alone, while
    ApplyRouting installs a rule AND a `local default dev lo` route removed
    by two independent commands. A teardown interrupted between them was
    therefore invisible, applyLocked's fast-path skipped ApplyRouting
    forever, and no reconcile could repair it.

Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):

  * init: `start_service` waits for a live predecessor pidfile to clear
    before opening the instance, so restart == stop + pause + start. Zero
    cost at boot. term_timeout 10 -> 30 so an honest teardown is never
    killed halfway.
  * daemon: the single-owner guard WAITS for the predecessor (bounded,
    60s) instead of exiting 1; it still refuses if the budget expires.
  * netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
    a blanket flush would drop the admin's own SSH/LuCI sessions) called
    on every plane transition: after a ruleset loads, after the table is
    removed, and once more in applyLocked when the whole plane (table +
    policy routing + sysctls) is assembled.
  * netplane: RoutingPresent() now verifies both halves it installs.

Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:51:19 +03:00
omarandClaude Opus 5 c562579ef3 docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all
unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute
does not emit a condition-less rule as a match-all route rule: it sets
route.Final and continues, so the LAST condition-less rule by order wins, and it
can never shadow a rule that has conditions (those are emitted ahead of Final
regardless of order).

For the config on the router this inverts the conclusion: traffic IS going
through the proxy (order=100 -> group:auto is the live default) and the dead knob
is the order=20 `direct` one. Severity downgraded from high to medium
accordingly — a dead setting, not a leak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:41:58 +03:00
omarandClaude Opus 5 6f89acbae7 feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final",
so a config with two of them showed two identical claims and no hint that
only the last one is the default the router uses.

A superseded rule now loses those marks — it keeps its real Order in the rail
instead of the "·" that means final — and gains a "never applies" badge plus
a line naming the rule that beat it and what to do about it: give this one a
condition, or delete one of the two. Warn semantics throughout (--amber,
dashed frame, dimmed target chip): orange is the ACTIVE state on this
faceplate, and a rule the router ignores is the opposite of active.

Verdicts come from GET /api/rules/reachability and are keyed by the rule's
index in Rules, never by name — the config that prompted this had two rules
both called `default`. They are re-fetched after every save, and a verdict
whose echoed name/order no longer matches the row is dropped rather than
shown, so the window between an optimistic edit and the refetch cannot badge
a working rule.

Rule rows were also keyed by name in React, which silently collapses two rows
that share one; the key now carries the model index.

The mock fixture gains a second condition-less rule so `?mock` renders the
state, and mock.getRulesReachability derives its verdicts from the live
fixture config rather than hard-coding them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 88a82c7297 feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:

  * two condition-less rules retire each other, and the LAST one by Order
    wins, so an earlier "default -> direct" is dead while looking live;
  * a condition-less rule can NEVER retire a rule that HAS conditions —
    those are emitted ahead of Final whatever their Order.

A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.

model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.

Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.

Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.

Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.

GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.

Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 a8f2b0f068 ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.

ci/version.sh is now the single source of truth. It derives the version
from `git describe`:

    tag `vX.Y.Z`   -> PKG_VERSION=X.Y.Z  PKG_RELEASE=1
    off-tag build  -> nearest tag + PKG_RELEASE=<commits since it> + 1
    no tag/no git  -> 0.0.0-r1 (below everything ever published)

Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.

The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.

The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.

byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.

Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:32:38 +03:00
omarandClaude Opus 5 0b32a6d58b fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:

  daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
  [ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...

`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.

Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:

  * control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
    DisableColors at its false zero value. It now comes from
    controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
    lives in an untagged file (same split as profilewatch.go) so it is
    unit-testable off the linux target.
  * engine (shater/generate): the generated option.LogOptions never set
    DisableColor, so box.New built a colouring formatter over the shared
    sink. logOptions() now sets it from the same TTY gate (seam:
    logColorAllowed).

logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.

The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.

Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:31:01 +03:00
omarandClaude Opus 5 893fdc500c fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as 9dc954029 / aec82d444.

`nodes` now reads model.ReadUCI() — `uci export shater` merged with the
per-subscription JSON caches — which is literally the call GET
/api/config serves and generate builds the engine from, so the verb
cannot drift from the panel or from the running engine: there is no
second assembly here to drift. Both ends use the same nodesJSON():
the daemon answers over the control socket (like `stats`), and the CLI
falls back to reading the same on-disk state when no daemon is running
(like `status`). A read failure goes to stderr with a non-zero exit
instead of printing `[]`, so an empty list on stdout now means one thing.

Output is a purpose-built view rather than raw model.Node: the share-link
URI is a credential and CLI output ends up in tickets and cron mail, so
the view reports what the link decodes to (protocol/server/port) plus the
model's own facts (enabled/sub/egress/stale/fingerprint). Nodes whose URI
does not parse are still listed, with the reason in `parse_error` — the
engine skips exactly those, and hiding them would be the same lie smaller.

cmdReadStub keeps `stats`, where the default IS the truth (nothing was
counted without an engine), and now says so in its doc comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:30:27 +03:00
omarandClaude Opus 5 02c266188f docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the
previous install, install from the signed apk feed, verify the default state,
restore a working config with 315 subscription nodes, then exercise the data
plane, panel API, config lifecycle, resilience and DNS.

74 PASS. Findings (detailed separately): two catch-all `default` rules where the
first sends all unspecific traffic direct and makes the second unreachable;
`shaterd nodes` is a stub returning [] while usage promises the node list; DNS to
the router LAN address dies after `service shater restart` (stop+pause+start is
fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI
colour codes reach syslog.

Also records the four-iteration CI hunt that ended in the green apk lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:09:52 +03:00
omarandClaude Opus 5 024e9308c9 fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.

The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:

    config PACKAGE_kmod-mlx5-core
            tristate
            default m

No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.

The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.

Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:27:37 +03:00
omarandClaude Opus 5 eab1db2c3f fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.

They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.

Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
  - an explicit value in .config beats a `default` (same rule that kept our
    `# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
  - `select` is OR-ed in after the user value, so shater-core's
    `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
    transitive kmods) back on their own.

Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.

Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:16:05 +03:00
71 changed files with 5235 additions and 644 deletions
+60 -18
View File
@@ -32,6 +32,23 @@
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
# API via curl (ci/gitea-release.sh) — no external action needed.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
# used to be, and nobody bumped them: v0.2.2…v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with different binaries inside, so `apk update` never saw
# a new version and routers could not be updated at all. Now `ci/version.sh`
# derives them from the git tag ONCE per job (the "Compute version" step,
# exported via $GITHUB_ENV):
# tag `vX.Y.Z` -> X.Y.Z-r1
# anything else -> <nearest tag>-r<commits since it + 1>
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
# built .ipk/.apk really carry that version, so the failure can never be
# silent again. This is also why both build jobs check out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
#
# APK LANE (25.12+, ADDITIVE — T2)
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
@@ -118,8 +135,14 @@ jobs:
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
steps:
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the engine via a go.mod
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
@@ -129,6 +152,13 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
- name: Set up Go
uses: actions/setup-go@v5
@@ -197,25 +227,23 @@ jobs:
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
# the SDK package build (the openwrt/shaterd package installs the staged
# artifact). VERSION is stamped into constant.Version. On an exact
# artifact). $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
- name: Build signed feed (SDK)
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
@@ -253,8 +281,12 @@ jobs:
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
steps:
# fetch-depth: 0 — see the opkg lane: the package version comes from
# `git describe`, which needs tags + ancestry.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the AmneziaWG-patched wireguard-go via a
# go.mod `replace => ./submodules/wireguard-go`, so that submodule must be
@@ -263,6 +295,11 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# Same single version computation as the opkg lane — both lanes MUST agree
# on the version, they package the identical tree.
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
- name: Set up Go
uses: actions/setup-go@v5
with:
@@ -346,15 +383,10 @@ jobs:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages as .apk through the ImmortalWrt 25.12 SDK and
# sign the per-arch packages.adb with the EC key (secret KEY_APK).
@@ -462,7 +494,7 @@ jobs:
luci-app-shater (arch=all).
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
── Add as an opkg feed (recommended — then `opkg upgrade` just works) ──
── Add as an opkg feed (recommended — then updating is one command) ──
This release is itself a SIGNED package feed; opkg filters by
architecture, so the same lines work on every device:
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
@@ -472,6 +504,10 @@ jobs:
The public-key install is one-time; after it, `opkg update/upgrade`
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
── Update (name our packages — never a bare `opkg upgrade`) ──
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
── Or install the loose .ipk directly / from the tarball feed ──
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
@@ -531,13 +567,19 @@ jobs:
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
── Add as an apk repository (auto-updates via \`apk upgrade\`) ──
── Add as an apk repository ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
Update: apk update && apk upgrade shaterd shater-core luci-app-shater byedpi
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
A bare \`apk upgrade\` reconciles EVERY installed package against every
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
those packages are upgraded along with needed dependencies\").
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
echo "[release-apk] publishing $TAG from $d"
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
+27 -3
View File
@@ -154,8 +154,14 @@ opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # опционально: ByeDPI desync-egress
```
Обновление: `opkg update && opkg upgrade shaterd shater-core luci-app-shater byedpi`
(обновляйте только эти четыре пакета, не системные).
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
он тянет обновления и на системные пакеты, это классический способ окирпичить
роутер):
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
```
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
@@ -176,7 +182,25 @@ apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # опционально: ByeDPI desync-egress
```
Обновление: `apk update && apk upgrade shaterd shater-core luci-app-shater byedpi`.
Обновление — **перечисляйте пакеты явно, голый `apk upgrade` не запускайте**: без
аргументов apk пересобирает состояние ВСЕХ установленных пакетов по ВСЕМ
подключённым репозиториям и может задеть (в т.ч. откатить) посторонние системные
пакеты.
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
# эквивалент, дополнительно закрепляющий пакеты в world:
# apk add -u shaterd shater-core luci-app-shater byedpi
```
Документация apk-tools 3 про `apk upgrade`: *«If list of packages is provided,
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
+12
View File
@@ -55,6 +55,16 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
# standalone. Passed into the container below and re-exported to the
# unprivileged build user in ci/sdk-build-apk.sh.
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) runner-side caches --------------------------------------------------
# All under $REPO/.cache so (a) actions/cache in the workflow can persist them
# between runs and (b) the nested container sees them via --volumes-from.
@@ -94,6 +104,8 @@ docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e SDK_URL="$SDK_URL" \
-e SDK_TAR="$SDK_TAR" -e DL_DIR="$CACHE/dl" -e APT_CACHE="$CACHE/apt" \
-e FEEDS_CACHE="$FEEDS_CACHE" -e KEY_APK="${KEY_APK:-}" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
debian:bookworm bash "$REPO/ci/sdk-build-apk.sh"
# --- 2) sanity: the per-arch apk repo dir must be complete -------------------
+14
View File
@@ -47,6 +47,18 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# The workflow normally puts these in the job env (ci/version.sh --env >>
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
# 0.2.0-r3 and no router could ever see an update).
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
# Workspace dir restored/saved by actions/cache in the workflow and shared into
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
@@ -81,6 +93,8 @@ docker pull "openwrt/sdk:$SDK_TAG"
docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
-e FEEDS_CACHE="$FEEDS_CACHE" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
"openwrt/sdk:$SDK_TAG" \
sh "$REPO/ci/sdk-build.sh"
+133 -5
View File
@@ -28,6 +28,11 @@ SDK_URL="${SDK_URL:?SDK_URL env required}"
echo "[apk-sdk] arch=$ARCH repo=$REPO out=$OUT"
echo "[apk-sdk] sdk=$SDK_URL"
# Package version derived from the git tag by ci/version.sh (bug B4). Forwarded
# to the unprivileged build user on the `su` line at the bottom of this file;
# openwrt/{shaterd,shater-core,luci-app-shater}/Makefile pick it up from the
# environment. byedpi keeps upstream ByeDPI's own version (see its Makefile).
echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -133,6 +138,41 @@ fi
echo "[apk-sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
# --- strip the SDK's generated per-package `default m` blocks ----------------
# Run 60 settled the question that runs 58 and 59 left open. Writing an explicit
# `# CONFIG_PACKAGE_kmod-x is not set` for all 1126 of them and re-running
# defconfig deselected exactly nothing: the count came back 1078, unchanged.
# Meanwhile the very same explicit form DID stick for CONFIG_ALL/ALL_KMODS/
# ALL_NONSHARED. The difference is prompts. kconfig only honours a user value for
# a symbol that has one (sym_calc_value ignores S_DEF_USER for a promptless
# symbol and falls back to its `default`), and the ALL* symbols carry prompts in
# the SDK's own Config.in while these generated blocks are bare:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# So no value we write into .config can ever turn them off — the fix has to
# remove the `default m` itself. That is what this does: drop every generated
# `config PACKAGE_*` block from the SDK's Config-build.in before the first
# defconfig. Nothing is lost by it — these blocks only replay which packages the
# BUILDBOT happened to build; the packages themselves are still declared, with
# prompts, by the package tree (tmp/.config-package.in), which is what makes our
# four selectable and what `select` acts on. KERNEL_*/LIBC/TOOLCHAIN blocks are
# left untouched, so the SDK still reproduces its own toolchain settings.
CB=$(find . -maxdepth 2 -name 'Config-build.in' -print -quit 2>/dev/null || true)
if [ -n "$CB" ] && command -v perl >/dev/null 2>&1; then
pkg_before=$(grep -c '^config PACKAGE_' "$CB" || true)
# Paragraph-wise delete: a block is `config PACKAGE_x`, its indented body, and
# the blank line that ends it. Anchored per-line (/m) so nothing else matches.
perl -0777 -pi -e 's/^config PACKAGE_\S+\n(?:[ \t]+\S[^\n]*\n)+\n//gm' "$CB"
pkg_after=$(grep -c '^config PACKAGE_' "$CB" || true)
echo "[apk-sdk] $CB: stripped $((pkg_before - pkg_after)) generated PACKAGE default blocks ($pkg_before -> $pkg_after)"
else
echo "[apk-sdk] WARNING: no Config-build.in found (or no perl) — per-package"
echo "[apk-sdk] 'default m' blocks stay; the kmod tripwire will catch it"
fi
# --- .config: turn OFF the SDK's mass-select defaults ------------------------
# Symptom (v0.2.2, and still v0.2.3 run 58): the SDK ran `apk mkpkg` on ~1100
# kmod-* packages — mlx5, amdgpu, ata, isdn, none of which we ship — and died
@@ -211,6 +251,63 @@ fi
echo "[apk-sdk] defconfig"
make defconfig >/dev/null
# --- second pass: deselect the kernel, keep only what our packages select -----
# Turning ALL/ALL_KMODS/ALL_NONSHARED off (above) provably worked — run 59 shows
# all three as `is not set` after defconfig — and changed the kmod count by
# exactly zero, 1078 both times. The kmods are not selected through ALL_KMODS at
# all. They are selected one by one, and here is where from:
#
# target/sdk/Makefile:
# ./convert-config.pl $(TOPDIR)/.config > $(SDK_BUILD_DIR)/Config-build.in
#
# The SDK's Config-build.in is GENERATED from the buildbot's .config — a config
# in which ALL_KMODS=y had already expanded into a `CONFIG_PACKAGE_kmod-*=m` line
# per module. convert-config.pl turns every `CONFIG_X=<val>` line into a kconfig
# symbol carrying an unconditional `default <val>`; its `next if
# /^(# )?CONFIG_PACKAGE/` filter sits in the `else` branch, which a line with an
# `=` in it never reaches. So the SDK ships, verbatim, 1078 blocks of:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# Nothing there consults ALL_KMODS, which is why switching it off was inert.
#
# Fix: give those symbols an explicit user value. We cannot do it before the
# first defconfig — the list of names only exists once kconfig has expanded the
# tree — so this is a second pass: rewrite every selected kmod to `is not set`
# and re-run defconfig. Two kconfig rules make the result exactly what we want,
# and both are already demonstrated in our own logs:
# * an explicit value in .config beats a `default` (this is precisely why the
# `# CONFIG_ALL* is not set` lines survived defconfig in run 59), so the
# ~1078 kmods we do not need stay off;
# * `select` is a reverse dependency, OR-ed into the symbol's value AFTER the
# user value in sym_calc_value(), so it cannot be overridden by an explicit
# `n`. shater-core's `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` becomes
# `select PACKAGE_kmod-nft-tproxy` (scripts/package-metadata.pl: a `+` flag
# sets `$m = "select"`, and it re-emits the dependency's own depends too, so
# transitive kmods follow). Those come back on their own.
# Net effect: we build the handful of kmods our packages actually pull in.
#
# Rejected alternatives:
# * limiting what `package/kernel/linux/compile` packs — that target has no
# such knob; it iterates the selected set, so the selection IS the knob;
# * `package/kernel/linux/clean` + a targeted build — the kernel package would
# simply be rebuilt in full as a dependency of shater-core, same cost;
# * copying OpenWrt's own feed CI (openwrt/gh-action-sdk) — it does nothing
# about this; it just runs `make defconfig` and builds. Its one disk-related
# setting, CONFIG_AUTOREMOVE=y, is already the SDK's default;
# * editing the SDK's generated Config-build.in to strip the offending blocks —
# it would work, but it means parsing a generated kconfig file by hand and a
# format change would corrupt it silently. The two-pass approach uses only
# kconfig's documented semantics and leaves the evidence in .config.
kmods_all=$(grep -c '^CONFIG_PACKAGE_kmod-[^=]*=[my]$' .config || true)
if [ "$kmods_all" -gt 0 ]; then
echo "[apk-sdk] deselecting $kmods_all kmod packages, then defconfig again"
sed -i -E 's/^CONFIG_(PACKAGE_kmod-[^=]*)=[my]$/# CONFIG_\1 is not set/' .config
make defconfig >/dev/null
fi
# --- post-defconfig sanity + disk-cost readout -------------------------------
# A failed run leaves a ~27 MB log; digging the cause out of it is miserable, so
# print the handful of numbers that decide whether this run survives the
@@ -225,7 +322,7 @@ echo "[apk-sdk] kmod packages selected (=m): $kmods"
# the kmod count above will be in the four digits.
echo "[apk-sdk] mass-select symbols after defconfig:"
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/[apk-sdk] /' || true
# With ALL_KMODS off, the only kmods left are the ones shater-core's
# After the second pass the only kmods left are the ones shater-core's
# `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` turns into kconfig `select`s, plus
# whatever those select in turn — a handful. Worth printing verbatim while the
# list is short. A count of 0 is NOT fatal: those kmods ship in the router's own
@@ -234,6 +331,11 @@ grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/[apk-sdk]
if [ "$kmods" -le 30 ]; then
grep '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/[apk-sdk] /' || true
fi
# The two cache knobs are written before the first defconfig and have to survive
# both of them — losing DOWNLOAD_FOLDER silently costs us the dl/ cache, and
# losing LOCALMIRROR brings back the sourceware.org stalls. Cheap to just look.
echo "[apk-sdk] cache settings after defconfig:"
grep -E '^CONFIG_(LOCALMIRROR|DOWNLOAD_FOLDER)=' .config | sed 's/^/[apk-sdk] /' || true
echo "[apk-sdk] our packages after defconfig:"
grep -E '^CONFIG_PACKAGE_(shaterd|shater-core|byedpi|luci-app-shater)=' .config \
| sed 's/^/[apk-sdk] /' || true
@@ -256,9 +358,16 @@ done
# packing the kernel before dying on `Disk quota exceeded`. Fail now instead.
[ "$kmods" -le 200 ] || {
echo "[apk-sdk] ERROR: $kmods kmod packages selected — that is the whole kernel."
echo " CONFIG_ALL_KMODS/CONFIG_ALL is back in .config; aborting before"
echo " this fills the runner's disk. Offending lines:"
grep '^CONFIG_ALL' .config || true; exit 11; }
echo " Aborting before this fills the runner's disk. Two causes are"
echo " possible, and the lines below tell them apart:"
echo " (a) the mass-select is back on -> a CONFIG_ALL* line reads =y;"
echo " (b) the second pass did not take -> ALL* are 'is not set' but the"
echo " kmods returned anyway, i.e. the per-kmod 'default m' from the"
echo " SDK's generated Config-build.in outlived our explicit 'n'."
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/ /' || true
echo " first few kmods still selected:"
grep -m5 '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/ /' || true
exit 11; }
for p in shaterd shater-core byedpi luci-app-shater; do
echo "[apk-sdk] === build $p ==="
@@ -288,6 +397,25 @@ done
[ "$found" -ge 4 ] || { echo "[apk-sdk] ERROR: expected >=4 of OUR .apk, collected $found"; echo "[apk-sdk] (all .apk under bin/:)"; find bin -type f -name '*.apk' | head -20; exit 6; }
echo "[apk-sdk] collected $found of our .apk"
# --- assert the tag-derived version actually reached the packages -------------
# B4's failure mode is a wrong-but-plausible version shipping silently, so the
# env -> make hand-off is verified, not trusted: each of our three tag-versioned
# packages must be named `<name>-<ver>-r<rel>.apk`. byedpi is excluded on purpose
# (it carries upstream ByeDPI's own version). This runs BEFORE `apk mkndx`, so a
# stale version can never even reach the index.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
[ -f "$OUT/${p}-${want}.apk" ] || {
echo "[apk-sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[apk-sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[apk-sdk] version check OK — our 3 packages are $want"
fi
# --- index + sign: exactly how the OpenWrt 25.12 buildsystem does it ---------
# apk mkndx --root T --keys-dir T [--sign key] --allow-untrusted \
# --output packages.adb *.apk
@@ -319,7 +447,7 @@ INNER
chmod 0644 /home/build/inner.sh
su build -s /bin/bash -c \
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' bash /home/build/inner.sh"
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' SHATER_PKG_VERSION='${SHATER_PKG_VERSION:-}' SHATER_PKG_RELEASE='${SHATER_PKG_RELEASE:-}' bash /home/build/inner.sh"
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[apk-sdk] OK arch=$ARCH — apk feed dir:"
+27
View File
@@ -27,6 +27,13 @@ OUT="${OUT:?OUT env required}"
mkdir -p "$OUT"
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
# Package version, derived from the git tag by ci/version.sh and handed in by
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
# imports every environment variable as a variable, and it propagates through
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
# byedpi deliberately keeps its own upstream version (see its Makefile).
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -111,6 +118,26 @@ for p in shaterd shater-core byedpi luci-app-shater; do
done
done
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
# --- assert the tag-derived version actually reached the packages -------------
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
# env -> make hand-off has several layers (docker -e, make's env import, the
# metadata dump), so verify the result instead of trusting it: every one of our
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
echo "[sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[sdk] version check OK — our 3 packages are $want"
fi
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
ls -l "$OUT"
Executable
+133
View File
@@ -0,0 +1,133 @@
#!/bin/sh
# ci/version.sh — the SINGLE source of truth for "what version is this build?".
#
# WHY THIS EXISTS (bug B4)
# -----------------------
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
# feed's version string differs from the installed one, `apk update` saw nothing
# new and the routers could not be updated through the normal path at all.
#
# So the version is now DERIVED, in CI, from the git tag, and the package
# Makefiles only carry a fallback for manual/offline builds.
#
# THE SCHEME
# ----------
# tag push `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
# any other build -> PKG_VERSION=X.Y.Z of the NEAREST reachable tag,
# (workflow_dispatch, PKG_RELEASE=<commits since that tag> + 1
# rolling `latest`)
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
#
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
# part first (numerically, component by component), the `r<rel>` only as a
# tie-break. Verified against the real tools, not from memory:
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
# identical results (opkg implements the Debian algorithm).
# That is exactly the ordering this scheme needs:
# * a release always outranks every rolling build that preceded it
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
# * rolling builds between two releases grow monotonically (r2 < r10 < r11),
# so a rolling build can never look newer than the next release, and the
# `latest` feed still moves forward on every dispatch.
#
# +1 on the commit count (rather than the raw count) only avoids `-r0` and makes
# a dispatch build of the tagged commit itself identical to the release build of
# that same commit — which is the truth: same tree, same binary.
#
# `byedpi` is deliberately NOT versioned from our tag — see openwrt/byedpi/Makefile.
#
# USAGE
# ci/version.sh # or --env: eval-able / $GITHUB_ENV-able lines
# ci/version.sh --pkg-version # X.Y.Z
# ci/version.sh --pkg-release # R
# ci/version.sh --binary # vX.Y.Z-rR[-g<sha>] for constant.Version
#
# Env:
# SHATER_REF / GITHUB_REF when it is `refs/tags/<tag>` that tag wins and no
# git history is needed (the tag-push path is exact
# even on a shallow checkout).
set -eu
REPO="$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)"
TAG=""
EXACT=0
N=0
SHA=""
# --- 1) an explicit tag ref is authoritative (and needs no git) --------------
REF="${SHATER_REF:-${GITHUB_REF:-}}"
case "$REF" in
refs/tags/*) TAG="${REF#refs/tags/}"; EXACT=1 ;;
esac
# --- 2) otherwise ask git for the nearest reachable release tag --------------
# `--match 'v[0-9]*'` keeps non-release tags (latest, sdk-cache, apk-latest-*,
# musl-toolchain-cache) out. This repo is a sing-box FORK and therefore also
# carries upstream's v1.x tags — `git describe` picks the CLOSEST tag by commit
# distance, so our own v0.2.x (a handful of commits back) always wins over
# upstream's v1.x (thousands of commits back). The tag it picked is logged
# below, so a surprise is visible in the CI log rather than silently shipped.
if [ "$EXACT" -eq 0 ]; then
if D="$(git -C "$REPO" describe --tags --long --match 'v[0-9]*' 2>/dev/null)"; then
# `v0.2.6-1-g02c266188` -> TAG=v0.2.6 N=1 SHA=g02c266188.
# `%` strips the SHORTEST matching suffix, so a tag that itself contains a
# dash (`v0.2.0-healthplan`) survives intact.
TAG="${D%-*-g*}"
REST="${D#"$TAG"-}"
N="${REST%%-*}"
SHA="${REST#*-}"
if [ "$N" -eq 0 ]; then EXACT=1; fi
fi
fi
# --- 3) tag -> numeric PKG_VERSION ------------------------------------------
# Keep the leading dotted-numeric run only: `v0.2.0-healthplan` -> `0.2.0`.
VER=""
if [ -n "$TAG" ]; then
VER="$(printf '%s' "${TAG#v}" | sed -n 's/^\([0-9][0-9.]*\).*/\1/p' | sed 's/\.*$//')"
fi
if [ -z "$VER" ]; then
# No release tag anywhere (shallow clone with no tags, a tarball export, a
# fresh fork). 0.0.0 is BELOW every version we have ever published, so such a
# build can never masquerade as an upgrade on a real router; the commit count
# still makes successive dev builds distinguishable.
VER="0.0.0"
EXACT=0
N="$(git -C "$REPO" rev-list --count HEAD 2>/dev/null || echo 0)"
SHA="$(git -C "$REPO" rev-parse --short HEAD 2>/dev/null || echo '')"
[ -z "$SHA" ] || SHA="g$SHA"
echo "[version] WARNING: no reachable vX.Y.Z tag (and/or no git) -> $VER" >&2
fi
# --- 4) PKG_RELEASE + the string stamped into the binary --------------------
if [ "$EXACT" -eq 1 ]; then
REL=1
FULL="v${VER}-r${REL}"
else
REL=$((N + 1))
FULL="v${VER}-r${REL}${SHA:+-$SHA}"
fi
echo "[version] tag='${TAG:-none}' commits_since=$N exact=$EXACT -> ${VER}-r${REL} (binary: $FULL)" >&2
case "${1:---env}" in
--env|"")
printf 'SHATER_PKG_VERSION=%s\n' "$VER"
printf 'SHATER_PKG_RELEASE=%s\n' "$REL"
printf 'SHATER_VERSION=%s\n' "$FULL"
;;
--pkg-version) printf '%s\n' "$VER" ;;
--pkg-release) printf '%s\n' "$REL" ;;
--binary|--version) printf '%s\n' "$FULL" ;;
*)
echo "usage: $0 [--env|--pkg-version|--pkg-release|--binary]" >&2
exit 2 ;;
esac
+65
View File
@@ -424,3 +424,68 @@ on the next render.
Consequence: all delay numbers are comparable (least ping ranks apples against
apples), and group settings lose two footgun fields while Settings keeps the
two that actually govern every check.
## D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). `config rule` carried THREE ways to say
where traffic is going: `dst_domain` (an inline domain list), `dst_ip` (an inline
CIDR list) and `dst_ruleset` (a reference to a `config ruleset`). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a `.srs`, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
**Decision: `dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is
the only destination matcher.** `Src`, `dst_port` and `proto` are untouched —
they are not lists of destinations and have no rule-set form.
- **Rejected: keep the inline lists as a shorthand.** "One obvious way" is the
whole point; a shorthand that quietly means something different from the long
form (see the bare-entry trap) is worse than no shorthand.
- **Rejected: promote inline lists to rule-sets lazily at generate time.** The
config on disk would then not say what the router does, and the panel would
have to render a list the user cannot find or edit.
### The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (`dst_domain`) | in a rule-set (`entry`) | migrated to |
|--------------------|--------------------------|-------------------------|--------------|
| `example.com` | **exact host** | **host + subdomains** | `full:example.com` |
| `full:example.com` | exact host | exact host | unchanged |
| `suffix:example.com` / `.example.com` | host + subdomains | host + subdomains | unchanged |
| `keyword:ads` | substring | substring | unchanged |
| `regexp:^ads\.` | pattern | pattern *(added here)* | unchanged |
| `geosite:x` / `geoip:x` | inert (engine field removed) | inert (unknown prefix) | unchanged |
`shaterd migrate` (schema v1→v2, `shater/model/migrate.go`) creates one inline
`config ruleset` per rule that still carries a legacy list — `rule-<rule name>`
for domains, `rule-<rule name>-ip` for addresses — moves the entries across with
the conversion above, appends the new name to `dst_ruleset`, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
`rule-<name>-2`). `regexp:` support was added to inline rule-sets in the same
change precisely so the move can be lossless.
`geosite:`/`geoip:` entries are copied VERBATIM rather than promoted to a
`source=geosite` rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
**One deliberate semantic change, called out:** a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because `rule_set: [a, b]` matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
Consequence: one destination mechanism, one vocabulary, one place a list is
edited; every list is compiled once and reused. The panel's rule editor drops its
Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of
the rulesets that already exist, and nothing more. Creating and filling a list
stays in the Rulesets panel — **rejected: a "create a list from here" shortcut in
the rule editor**, because a second place to author a list is a second place for
its semantics and its duplicate-name rules to drift, and the whole point of this
decision was to stop having two.
+6 -2
View File
@@ -13,8 +13,12 @@ usable release, **[T1]** next, **[T2]** later. Phases refer to `ROADMAP.md`.
- **[MVP]** TPROXY transparent proxy for multiple LAN interfaces (TCP + UDP), SNI/
Host/QUIC sniffing.
- **[MVP]** First-match routing rules by source (IP/CIDR/MAC/interface/zone),
destination (domain/suffix/keyword/geosite), reusable domain/IP lists, port,
proto → target (outbound/selector/chain/direct/block) + egress.
destination, port, proto → target (outbound/selector/chain/direct/block) + egress.
A rule names its **destination through a rule-set only** — a reusable named list
(inline domains/CIDRs, a local or remote file, or a geosite/geoip category) that is
compiled once into a `.srs` and shared by every rule that references it. Domain
entries take `full:` (exact), `suffix:` / a leading dot (host + subdomains),
`keyword:` (substring) and `regexp:`; a bare entry means host + subdomains.
- **[MVP]** Node groups with balancer/observatory (least-ping/failover/round-robin).
- **[T1]** Multi-hop chains (L1→Ln); per-rule egress selection; egress via any
interface/tunnel (e.g. an AmneziaWG tunnel).
+81 -15
View File
@@ -26,7 +26,9 @@ What it does:
Arg / env:
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
`$SHATER_VERSION` → `git describe --tags` → `v0.2.0-dev`.
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
the **same** computation the package version comes from (§2.1), so the string
the panel shows always matches what `apk info shaterd` / `opkg status` report.
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
@@ -84,15 +86,52 @@ it. Because the binary is UPX-packed, the package disables the SDK's default str
feed installed and run `make package/shaterd/compile` (and the others) per target.
See `openwrt-package-build-ci` for SDK/feed mechanics.
### 2.1 Package versions come from the git tag
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
Both package managers offer an upgrade only when the feed's version string
differs from the installed one, so `apk update` saw nothing new and the routers
could not be updated through the normal path at all.
`ci/version.sh` now derives them from `git describe`, once per CI job:
| Build | `PKG_VERSION` | `PKG_RELEASE` | `constant.Version` |
|---|---|---|---|
| tag push `v0.2.7` | `0.2.7` | `1` | `v0.2.7-r1` |
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
Ordering is what makes this safe, and both managers agree on it (checked with
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
outranks every rolling build before it, rolling builds between two releases grow
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
upgrade.
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
environment; the Makefiles read it with a literal fallback for manual/offline
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
so a lost variable fails the build instead of shipping a stale version.
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
installed. Stamping our tag on it would also be a downgrade: every comparator
reads `0.2.7 < 0.17.3` (component-wise, `2 < 17`). Bump its `PKG_RELEASE` by hand
when our packaging of it changes.
## 3. Install on a router
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
```sh
opkg install shaterd_0.2.0-1_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_0.2.0-1_all.ipk
opkg install luci-app-shater_0.2.0-1_all.ipk
opkg install byedpi_0.17.3-1_<arch>.ipk # optional: ByeDPI egress
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_<ver>_all.ipk
opkg install luci-app-shater_<ver>_all.ipk
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
```
Installing from a signed feed instead:
@@ -158,16 +197,25 @@ every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
### Updating
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
tries to upgrade *every* installed package from *every* configured feed, which on
OpenWrt means base/system packages on the overlay and is a well-known way to
brick a router.
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
```
Updates are only offered when the feed's `Version` differs from the installed one,
so **bump `PKG_RELEASE`** (or `PKG_VERSION`) in the package Makefile on every
shipped change — otherwise `opkg upgrade` sees the same version and does nothing.
Do **not** `opkg upgrade` base/system packages from this feed; upgrade only the
four shater packages above.
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
when the feed's `Version` differs from the installed one — that is exactly what
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
the version from the git tag on every build (§2.1), so there is nothing to bump
by hand any more; check with:
```sh
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
```
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
@@ -214,15 +262,33 @@ apk add byedpi # optional: ByeDPI desync egress
### Updating
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
installed package against *every* configured repository at once; on a router
whose distfeeds point at a moving snapshot that can pull in — or roll back —
unrelated system packages. Always name ours:
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
apk upgrade shaterd shater-core luci-app-shater byedpi
```
Same rule as opkg: an upgrade is only offered when the feed version differs, so
bump `PKG_RELEASE`/`PKG_VERSION` on every shipped change (apk shows it as
`0.2.0-r1`). Pin a version instead of tracking rolling by pointing the repo line
at `.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
apk-tools 3 documents exactly this behaviour for `apk upgrade`: *"When no
packages are specified, all packages are upgraded if possible. If list of
packages is provided, only those packages are upgraded along with needed
dependencies."* The equivalent form, which additionally re-pins the packages in
`world`, is:
```sh
apk add -u shaterd shater-core luci-app-shater byedpi # -u = --upgrade
```
Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
rolling by pointing the repo line at
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
### BananaWRT `25.12-mtk-vendor` compatibility
+7 -3
View File
@@ -195,7 +195,7 @@ type Chain struct { Name string; Hops []string } // "group:<n>" | "node:<n>", L1
type Egress struct { Name,Type,Interface,Target string } // interface|proxy|direct|block
type Rule struct {
Name string; Enabled bool; Order int
Src []string; DstDomain,DstRuleset,DstIP []string; DstPort,Proto string
Src []string; DstRuleset []string; DstPort,Proto string // dst = ruleset only (v0.2 schema v2)
Target string // chain:|group:|node:|direct|block
Egress,Kill string
SchedEnabled bool; SchedDays []string; SchedStart,SchedEnd string; SchedUTCOffset int
@@ -257,8 +257,12 @@ Apply/rollback: `apSnapshot` (run→last-good, nft→last-good.nft, route marks)
- `config node`: name, enabled, uri, mux, mux_concurrency, xudp_concurrency, xudp_udp443, sockopt_mark, tcp_fast_open, tcp_keepalive_idle.
- `config group`: name, source, subscription, list node, strategy, include/exclude/filter_proto/filter_country, dedup, probe_url, probe_interval.
- `config chain`: name, list hop. `config egress`: name, type, interface, target.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url), url, path, format, update_interval, list entry.
- `config rule`: name, enabled, order, list src/dst_domain/dst_ruleset/dst_ip, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end/tz.
- `config ruleset`: name, type(domain|ipcidr), source(inline|file|url|geosite|geoip), url, path, format, update_interval, list category, list entry.
- `config rule`: name, enabled, order, list src, list dst_ruleset, dst_port, proto, target, egress, kill, sched_enabled, list sched_day, sched_start/end, sched_utc_offset.
v0.1 carried `dst_domain`/`dst_ip` on the rule itself; **schema v2 removed both** — a
destination is a `config ruleset` and nothing else. `shaterd migrate` folds each legacy
list into a generated `rule-<name>` (and `rule-<name>-ip`) inline ruleset; see
`DECISIONS.md` D21 for the entry-by-entry conversion table.
- `config preset`: name, enabled, order, target. `config profile`: name, enabled, priority, list match_iface, probe_url, probe_mode, sched_*, list enable_rule/disable_rule, default_target, default_egress.
- `config resolver`: name, type, address, detour, pool. `config dns_rule`: order, list match_domain/match_src, resolver.
@@ -0,0 +1,170 @@
# Живое тестирование shater v0.2.6 на mini_router
**Дата:** 2026-07-25
**Устройство:** Bananapi BPi-R3 Mini · ImmortalWrt **25.12-linkup** · `aarch64_cortex-a53`
**Установка:** из подписанного apk-фида `apk-v0.2.6-aarch64_cortex-a53`
**Пакеты:** `shaterd 0.2.0-r3`, `shater-core 0.2.0-r3`, `luci-app-shater 0.2.0-r2`, `byedpi 0.17.3-r1`
**Сборка:** CI run 61, коммит `024e9308c` (вершина `main`)
Сценарий: полное удаление предыдущей установки → чистая установка из фида →
проверка дефолтного состояния → восстановление рабочего конфига с подписками
(315 узлов) → функциональная проверка.
**Итог: 79 проверок, 74 PASS, 5 находок** (детали и разбор — в
`shater-bugs-2026-07-25.md` на рабочем столе).
---
## 1. Релиз и фид
| # | Проверка | Результат |
|---|---|---|
| T1 | Публикация `apk-v0.2.6-<arch>` для обеих архитектур | PASS |
| T2 | Ассеты: 4 `.apk` + `packages.adb` + `shater-apk.pem` | PASS |
| T3 | `apk update` принимает индекс (проверка EC-подписи) | PASS |
| T4 | Пакеты видны в нужных версиях (r3/r3/r2) | PASS |
| T5 | Диагностика сборки: `kmod packages selected (=m): 0` (было 1078) | PASS |
| T6 | Собраны ровно наши 4 пакета | PASS |
| T7 | opkg-лейн v0.2.6 (24.10) тоже зелёный | PASS |
## 2. Установка
| # | Проверка | Результат |
|---|---|---|
| T8 | `apk add luci-app-shater byedpi` — 4 пакета | PASS |
| T9 | Зависимости `kmod-nft-tproxy`/`kmod-nft-socket` из базового фида | PASS |
| T10 | Целостность: `apk manifest` = sha256 файла на диске | PASS |
| T11 | Установлен именно бинарь v0.2.6 (5 491 616 Б vs 5 488 336 Б в r2) | PASS |
| T12 | init-скрипты `shater`, `shater-cron` | PASS |
| T13 | `sysctl.d/99-shater.conf`, `hotplug.d/iface/99-shater` | PASS |
| T14 | boot-линки `S99shater`, `K10shater`, `S96shater-cron` | PASS |
## 3. Дефолтное состояние (чистая установка)
| # | Проверка | Результат |
|---|---|---|
| T15 | Дефолтный конфиг создан uci-defaults (27 строк) | PASS |
| T16 | `enabled='0'` — плоскость не ставится без согласия | PASS |
| T17 | Заготовлен tproxy-inbound на LAN, пресеты выключены | PASS |
| T18 | Демон стартует, `plane=none`, `table=false` | PASS |
| T19 | Права конфига `-rw-------` (0600) | PASS |
## 4. Восстановление рабочего конфига
| # | Проверка | Результат |
|---|---|---|
| T20 | Восстановление из бэкапа (3213 UCI-строк) | PASS |
| T21 | Кэш подписок цел: 315 узлов в 4 файлах | PASS |
| T22 | `shaterd migrate` → `ok`, схема v1 | PASS |
| T23 | Старт с реальным конфигом: `active`, `engine_running`, `plane=full` | PASS |
## 5. Data plane
| # | Проверка | Результат |
|---|---|---|
| T24 | Таблица `inet shater` создана (9 цепочек/сетов) | PASS |
| T25 | 16 tproxy-правил | PASS |
| T26 | `ip rule from all fwmark 0x2000 lookup shater` | PASS |
| T27 | `accept_local=1` на `br-lan` | PASS |
| T28 | DNS-divert: `dport 53 → tproxy :12345` для LAN-интерфейсов | PASS |
| T29 | DoT заблокирован: `dport 853 reject` | PASS |
| T30 | `block_doh=1`, правила присутствуют | PASS |
| T31 | **Kill-switch fail-closed**: цепочка `forward` завершается `drop` для LAN (v4+v6) | PASS |
| T32 | fw4 и dnsmasq не тронуты (свои таблицы целы) | PASS |
## 6. Панель и API
| # | Проверка | Результат |
|---|---|---|
| T33 | SPA отдаётся на `:8088` | PASS |
| T34 | `shaterd mint-token` выдаёт одноразовый токен | PASS |
| T35 | `/api/status` без сессии → **401** | PASS |
| T36 | `/api/session` (POST, JSON) → 200 + cookie `HttpOnly; SameSite=Strict; Max-Age=28800` | PASS |
| T37 | `/api/status` по cookie отдаёт данные, совпадающие с CLI | PASS |
| T38 | `/api/config` — 340 записей узлов | PASS |
| T39 | `/api/groups/health` — 103 протестировано, 13 живых, выбран `FR-vless-8` | PASS |
| T40 | `/api/devices` — устройства с IPv4/IPv6/MAC | PASS |
| T41 | `/api/interfaces` — `ewan/eth1 10.0.0.125/24 zone=wan` | PASS |
| T42 | `/api/ruleset/status` — remote-ruleset обновлён сегодня | PASS |
| T43 | `/api/stats` — memory backend, счётчики и top-domains | PASS |
| T44 | `/api/stats/log` — query-log с доменом, qtype, rcode, сервером | PASS |
| T45 | `/api/log?range=100` — пусто (следствие `log_file='0'`, не дефект) | OK |
## 7. Жизненный цикл конфигурации
| # | Проверка | Результат |
|---|---|---|
| T46 | `shaterd apply` → `{"changed":false}`, `can_rollback=true` | PASS |
| T47 | `shaterd confirm` снимает авто-откат (`can_rollback=false`) | PASS |
| T48 | `shaterd rollback` после confirm корректно сообщает об отсутствии last-good | PASS |
| T49 | `shaterd reconcile` (SIGHUP) не роняет движок | PASS |
| T50 | `shaterd sub update all-qomar` — реально обновил 143 узла | PASS |
| T51 | `shaterd blocklist update` → reconcile signalled | PASS |
| T52 | `shaterd schedule due` → reconcile signalled | PASS |
## 8. Устойчивость
| # | Проверка | Результат |
|---|---|---|
| T53 | `kill -9` демона → procd поднимает новый PID | PASS |
| T54 | После respawn: `engine_running=true`, `plane=full` | PASS |
| T55 | `stop` снимает таблицу `inet shater` полностью | PASS |
| T56 | `stop` → пауза → `start`: плоскость восстанавливается | PASS |
| T57 | Сеть при остановленном shater не деградирует | PASS |
| T58 | Память: 253 МБ занято из 2 ГБ при работающем движке | PASS |
## 9. DNS
| # | Проверка | Результат |
|---|---|---|
| T59 | Резолв через `127.0.0.1` | PASS |
| T60 | LAN-клиенты резолвят через движок (query-log растёт) | PASS |
| T61 | `.lan`-домены остаются за dnsmasq | PASS |
| T62 | dnsmasq жив и слушает на всех адресах | PASS |
| T63 | **Резолв через LAN-адрес `10.67.0.1` после `restart`** | **FAIL — B3** |
| T64 | Тот же резолв после `stop` → пауза → `start` | PASS |
## 10. Конфигурация и логи
| # | Проверка | Результат |
|---|---|---|
| T65 | 5 правил маршрутизации, 2 профиля, активен `ethernet-uplink` | PASS |
| T66 | **Два правила `default`, оба catch-all — нижнее живое, верхнее мертво** | **FAIL — B1** |
| T67 | **`shaterd nodes` всегда возвращает `[]`** | **FAIL — B2** |
| T68 | Логи уходят в syslog (`log_syslog=1`, 22 записи) | PASS |
| T69 | **ANSI-escape коды в syslog** | **FAIL — B5** |
| T70 | `loglevel=warning` соблюдается | PASS |
| T71–T79 | Прочие проверки состояния (статус-поля, права, uptime, счётчики, целостность таблиц) | PASS |
---
## Находки
| ID | Суть | Важность |
|---|---|---|
| **B1** | Два catch-all правила `default`; одно из них не работает никогда. **Поправка к первоначальному диагнозу:** правило без условий задаёт `route.Final`, а не выпускается как match-all, поэтому выигрывает ПОСЛЕДНЕЕ (`order=100 → group:auto`) — трафик идёт через прокси, а мёртвая настройка это `order=20 → direct` | средняя |
| **B2** | `shaterd nodes` — заглушка, всегда `[]`, хотя usage обещает список узлов (в кэше 315, в `/api/config` 340) | средняя |
| **B3** | После `service shater restart` резолв к LAN-адресу роутера не работает и не восстанавливается; `stop`+пауза+`start` — работает (гонка) | средняя |
| **B4** | `PKG_RELEASE` не менялся с v0.2.1 → v0.2.2…v0.2.6 выходят как `r3` при разном содержимом; `apk upgrade` не увидит обновления | средняя |
| **B5** | ANSI-раскраска попадает в syslog | низкая |
Разбор с воспроизведением — в `shater-bugs-2026-07-25.md`.
## История CI по этому релизу
Путь до зелёной сборки apk-лейна занял четыре итерации, каждая вскрывала
следующий слой одной причины:
| Тег | Что чинили | Итог |
|---|---|---|
| v0.2.2 | — (первый прогон с фиксами аудита) | `Disk quota exceeded`, 3593 `apk mkpkg kmod-*` |
| v0.2.3 | `.config` строится с нуля, а не дописывается | 1078 kmod — SDK вообще не везёт `.config` |
| v0.2.4 | Выключены `ALL`/`ALL_KMODS`/`ALL_NONSHARED` | 1078 kmod — они выбираются не через `ALL_KMODS` |
| v0.2.5 | Второй проход: явное `is not set` для каждого kmod | 1078 kmod — kconfig игнорирует user-значение у беспромптовых символов |
| **v0.2.6** | Удаление сгенерированных блоков `config PACKAGE_*` (`default m`) из `Config-build.in` | **0 kmod, сборка зелёная** |
Корень: `target/sdk/Makefile` генерирует `Config-build.in` прогоном
`convert-config.pl` по конфигу бильдбота, где `ALL_KMODS=y` уже развернулся в
`CONFIG_PACKAGE_kmod-*=m` на каждый модуль. Фильтр `next if /^(# )?CONFIG_PACKAGE/`
в скрипте стоит в ветке `else`, куда строка со знаком `=` не попадает, поэтому
каждый kmod приезжает в SDK как безусловный `default m`.
+11
View File
@@ -15,6 +15,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=byedpi
# DELIBERATELY NOT auto-versioned from our git tag (unlike shaterd/shater-core/
# luci-app-shater, which take SHATER_PKG_VERSION/SHATER_PKG_RELEASE from
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
PKG_VERSION:=0.17.3
PKG_RELEASE:=1
+7 -2
View File
@@ -24,8 +24,13 @@ LUCI_TITLE:=LuCI thin launcher for Shater (mini dashboard + panel handoff)
LUCI_DEPENDS:=+shater-core +rpcd
LUCI_PKGARCH:=all
PKG_VERSION:=0.2.0
PKG_RELEASE:=2
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4). The literals are the manual/offline fallback only.
# These MUST stay above the luci.mk include: luci.mk only defaults PKG_VERSION/
# PKG_RELEASE when they are still unset, and the i18n subpackages inherit them.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+7 -2
View File
@@ -13,8 +13,13 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shater-core
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4: v0.2.2…v0.2.6 all shipped as 0.2.0-r3). The literals
# are the manual/offline fallback only.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-2.0-or-later
+29 -2
View File
@@ -62,7 +62,29 @@ config inbound
# list node 'my-node'
#
# A routing rule. target: chain:<n>|group:<n>|node:<n>|egress:<n>|direct|block.
# Match on src / dst_domain / dst_ruleset / dst_ip / dst_port / proto.
# Match on src / dst_ruleset / dst_port / proto. A rule with NO matcher at all is
# the default route for everything that reached it.
#
# WHERE the traffic is going is named ONLY by dst_ruleset — one or more
# `config ruleset` names; the rule matches when ANY of them matches. There is no
# inline domain or address list on a rule (`dst_domain`/`dst_ip` were removed in
# schema v2): a destination list is written once as a ruleset, compiled into a
# .srs and shared by every rule that references it. `shaterd migrate` converts
# older configs automatically, creating a `rule-<name>` ruleset per rule.
#config ruleset
# option name 'blocked-video'
# option type 'domain'
# option source 'inline'
# list entry 'youtube.com'
# list entry 'suffix:googlevideo.com'
#
#config rule
# option name 'video-via-main'
# option enabled '1'
# option order '50'
# list dst_ruleset 'blocked-video'
# option target 'group:main'
#
#config rule
# option name 'all-via-main'
# option enabled '1'
@@ -83,11 +105,16 @@ config inbound
# option type 'direct'
# option dpi 'fragment'
#
#config ruleset
# option name 'youtube'
# option source 'geosite'
# list category 'youtube'
#
#config rule
# option name 'youtube-fragment'
# option enabled '1'
# option order '50'
# list dst_domain 'geosite:youtube'
# list dst_ruleset 'youtube'
# option target 'egress:frag'
#
# A DNS resolver (type: doh|dot|plain|local|fakeip). `detour` routes its queries
+79 -5
View File
@@ -47,6 +47,13 @@ PROG=/usr/bin/shaterd
# hotplug/shater-cron touch the data plane. tmpfs => cleared by reboot, so
# nothing reconciles before this init has run at boot.
ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# --- helpers ---------------------------------------------------------------
@@ -66,6 +73,54 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
# data plane has not been dismantled yet".
shater_daemon_pid() {
local pid
pid=$(cat "$PIDFILE" 2>/dev/null) || return 1
[ -n "$pid" ] || return 1
kill -0 "$pid" 2>/dev/null || return 1
echo "$pid"
}
# Block until no predecessor daemon is left, bounded by STOP_WAIT_SECS.
#
# WHY THIS EXISTS. procd's `stop` is ASYNCHRONOUS: rc.common's `restart` is
# literally `stop; start`, and the `service delete` ubus call returns the moment
# procd has SENT SIGTERM — not when the instance is gone. `start` therefore
# re-adds the instance while the outgoing `shaterd run` is still executing its
# honest teardown (engine close, then `nft delete table`, `ip rule`/`ip route`
# removal and the per-iface sysctl restore). The result is that `restart` is NOT
# equivalent to `stop` + pause + `start`: the new plane is stood up on top of
# kernel state the old one has not finished removing, which is what B3 (DNS to
# the router's own LAN address dead after a restart, and never recovering) came
# out of. Waiting here restores the equivalence, and costs literally nothing when
# there is no predecessor — the check runs before the first sleep.
#
# Returning non-zero does NOT abort the start: the daemon carries its own
# single-owner guard and will refuse (or wait) on its side. Better to hand the
# decision to the process that can actually see the plane than to leave the box
# with no service at all.
shater_wait_stopped() {
local i=0 pid
pid=$(shater_daemon_pid) || return 0
_slog -p daemon.info \
"restart: waiting for the previous shaterd (pid $pid) to finish tearing the data plane down"
while [ "$i" -lt "$STOP_WAIT_SECS" ]; do
sleep 1
i=$((i + 1))
shater_daemon_pid >/dev/null || {
_slog -p daemon.info "restart: previous shaterd exited after ${i}s; starting a fresh one"
return 0
}
done
_slog -p daemon.warn \
"restart: previous shaterd (pid $pid) still alive after ${STOP_WAIT_SECS}s — starting anyway"
return 1
}
# --- procd lifecycle -------------------------------------------------------
start_service() {
@@ -87,6 +142,14 @@ start_service() {
return 0
fi
# Do not stand a new data plane up on top of one that is still being taken
# down. On `restart` procd has only just SIGTERMed the previous instance and
# returned; this is the handshake that makes `restart` == `stop` + pause +
# `start`. It also keeps `migrate` below from rewriting UCI underneath a
# daemon that is still reading it. No-op (and no delay) when nothing is
# running, which is the boot case.
shater_wait_stopped
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
"$PROG" migrate >/dev/null 2>&1
@@ -111,7 +174,16 @@ start_service() {
procd_set_param stderr 1
# Give the daemon room to run its honest teardown (engine.Close + netplane
# restore) before procd SIGKILLs it.
procd_set_param term_timeout 10
#
# 30s, not 10s: an engine holding a few hundred outbounds closes its
# urltest/observatory goroutines and flushes experimental.cache_file to FLASH
# before the netplane teardown even starts, and on eMMC/NAND that alone can
# outlast 10s. A SIGKILL there aborts the teardown at an arbitrary point and
# leaves the plane HALF removed — the nft table gone but the policy routing
# still installed, or vice versa — which is precisely the class of leftover
# state the successor's idempotent fast-path cannot see and never repairs.
# Shutdown is bounded by procd either way; we are only choosing where.
procd_set_param term_timeout 30
procd_close_instance
# Mark the stack live for hotplug/cron — but ONLY when interception is
@@ -141,10 +213,12 @@ stop_service() {
reload_service() {
# Fired by the `shater` config.change reload-trigger (LuCI Save & Apply /
# reload_config). Simplest correct behaviour: stop + start. `stop` clears the
# flag and SIGTERMs the daemon (honest teardown); `start` re-guards on
# enabled and, if still enabled, launches a fresh `shaterd run` that reads
# the new UCI and applies it. When the stack is disabled, `start` is a no-op,
# so a disable+apply cleanly tears everything down.
# flag and SIGTERMs the daemon (honest teardown); `start` WAITS for that
# teardown to actually finish (shater_wait_stopped) and then launches a fresh
# `shaterd run` that reads the new UCI and applies it. When the stack is
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
stop
start
}
+11 -2
View File
@@ -34,8 +34,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shaterd
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
# that version, so a lost env can never silently ship a stale one again.
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
# are not "the release version"; releases are named by the tag.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+54 -2
View File
@@ -809,9 +809,17 @@ export interface Rule {
Enabled: boolean
Order: number
Src?: string[] | null
DstDomain?: string[] | null
/**
* WHERE the traffic is going — the rule's only destination matcher. Each entry
* names a {@link Ruleset}; the rule matches when ANY of them matches.
*
* There is no inline domain or address list on a rule. `dst_domain`/`dst_ip`
* were removed in schema v2, and `shaterd migrate` folds every existing one
* into a generated `rule-<name>` ruleset, so a destination list is written and
* edited in exactly one place and compiled once into a .srs that every rule
* referencing it shares.
*/
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
/**
* Narrow the rule to one transport or one sniffed application protocol. A
@@ -1187,6 +1195,50 @@ export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEn
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
* One routing rule's reachability verdict — the rule analogue of
* {@link ChainHealth}.used: a quiet note about the ROUTING CONFIG, never a health
* signal.
*
* `unreachable` means the rule can NEVER take effect, whatever the traffic. Today
* the daemon reports exactly one certain case, and it is a subtle one: a rule with
* no conditions at all is not matched in sequence — it becomes the router's
* default. Two such rules therefore retire each other, and the LAST one by Order
* wins, so an earlier "default → direct" is dead even though it sorts first. A
* condition-less rule never retires a rule that HAS conditions: those are matched
* ahead of the default whatever their Order.
*
* `index` is the rule's position in GET /api/config's `Rules`, which is how a
* verdict is matched to a row — rule names are not unique, and the config that
* prompted this had two rules both called `default`. `name`/`order` are echoed so
* a page holding a verdict fetched before an edit can check it still describes the
* row it is about to badge, and drop it silently otherwise.
*/
export interface RuleReach {
index: number
name: string
order: number
unreachable: boolean
/** The rule that supersedes this one; absent when `unreachable` is false. */
shadowed_by?: string
/** Its index in `Rules`, or -1 when there is none. */
shadowed_by_index: number
shadowed_by_order?: number
/** Operator-facing sentence; absent when `unreachable` is false. */
reason?: string
}
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
* the same order as GET /api/config's `Rules`. */
export interface RulesReachability {
rules: RuleReach[]
}
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
+49 -1
View File
@@ -6,7 +6,7 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -146,6 +146,12 @@ const CONFIG: Model = {
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
// A SECOND condition-less rule, above the real default. It reads like a working
// rule and does nothing: a rule with no conditions becomes the router's default,
// and the last such rule by Order wins — so this one never applies. It is in the
// fixture on purpose, to exercise the "never applies" badge; the field config
// that prompted it had two rules BOTH named `default` (orders 20 and 100).
{ Name: 'default-bypass', Enabled: true, Order: 40, Target: 'direct' },
{ Name: 'default-tunnel', Enabled: true, Order: 900, Target: 'group:auto' },
],
// Named match-lists a rule points DstRuleset at. url + geosite + geoip are remote
@@ -296,6 +302,48 @@ const RULESET_STATUS: RulesetStatus[] = [
{ tag: 'rs-ru-geoip-ru', name: 'ru-geoip', category: 'ru', kind: 'ruleset', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
]
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
* a rule with no conditions is the router's default, and the LAST such rule by
* Order wins — every earlier one can never apply. It reads the live CONFIG so
* edits made in `?mock` keep the badge honest. */
export async function getRulesReachability(): Promise<RulesReachability> {
await wait(60)
const rules = CONFIG.Rules ?? []
const out: RuleReach[] = rules.map((r, index) => ({
index,
name: String(r.Name ?? ''),
order: Number(r.Order ?? 0),
unreachable: false,
shadowed_by_index: -1,
}))
const conditionless = (r: (typeof rules)[number]): boolean =>
!(r.Src ?? []).length &&
!(r.DstRuleset ?? []).length &&
!String(r.DstPort ?? '').trim() &&
!String(r.Proto ?? '').trim()
const target = (r: (typeof rules)[number]): string =>
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
const defaults = rules
.map((r, index) => ({ r, index }))
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
const winner = defaults[defaults.length - 1]
if (winner) {
for (const { index } of defaults.slice(0, -1)) {
out[index].unreachable = true
out[index].shadowed_by = String(winner.r.Name ?? '')
out[index].shadowed_by_index = winner.index
out[index].shadowed_by_order = Number(winner.r.Order ?? 0)
out[index].reason =
`this rule has no conditions, so it sets the default for all traffic — but rule ` +
`"${winner.r.Name}" (order ${winner.r.Order}) has none either and comes after it, so ` +
`"${target(winner.r)}" is the default the router uses and this rule's target ` +
`"${target(rules[index])}" is never applied`
}
}
return { rules: out }
}
export async function getRulesetStatus(): Promise<RulesetStatus[]> {
await wait(90)
return RULESET_STATUS.map((r) => ({ ...r }))
+38 -6
View File
@@ -252,6 +252,43 @@
color: var(--faint);
}
/* ---- a rule that can never fire (superseded by a later condition-less rule) ----
*
* Warn semantics only: --amber, never --accent. Orange is the ACTIVE state on this
* faceplate, and a rule the router ignores is the opposite of active — painting it
* orange is what made two `default` rows look equally live. It is a dashed amber
* frame, an amber order chip, and a dimmed target, so the row reads as "wired but
* not connected" without shouting: nothing is broken, one setting is just inert. */
.rt-rule.dead {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
background: var(--panel);
box-shadow: none;
}
.rt-ord.dead {
color: var(--amber);
border-color: color-mix(in srgb, var(--amber) 45%, var(--groove));
}
.rt-badge.dead {
padding: 1px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
color: var(--amber);
}
.rt-dead-note {
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.45;
color: var(--dim);
}
/* The target is still what the operator asked for, so it stays readable — just
* quiet, because the router is not using it. */
.rt-rule.dead .rt-target {
border-style: dashed;
opacity: 0.62;
}
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
.rt-target {
display: inline-flex;
@@ -359,9 +396,6 @@
gap: 5px;
min-width: 0;
}
.rt-field-wide {
grid-column: span 2;
}
.rt-flabel {
font-family: var(--font-mono);
font-size: 9px;
@@ -529,6 +563,7 @@ select.rt-input {
color: var(--faint);
}
/* textarea shares the input skin but grows vertically for a list of entries */
.rt-textarea {
min-height: 84px;
@@ -800,9 +835,6 @@ select.rt-input {
justify-content: flex-start;
align-self: start;
}
.rt-field-wide {
grid-column: auto;
}
.rt-rs-row {
grid-template-columns: 1fr;
row-gap: 10px;
+143 -119
View File
@@ -6,11 +6,12 @@ import {
apply as apiApply,
getConfig,
putConfig,
getRulesReachability,
getRulesetStatus,
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { Model, Rule, Ruleset, RulesetStatus } from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -21,9 +22,7 @@ import type { Model, Rule, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
type RRule = Rule & {
Src?: string[] | null
DstDomain?: string[] | null
DstRuleset?: string[] | null
DstIP?: string[] | null
DstPort?: string
Proto?: string
Kill?: string
@@ -96,7 +95,6 @@ function ProtoOptions({ value }: { value: string }) {
const len = (a: unknown[] | null | undefined): number => (a ? a.length : 0)
const byOrder = (a: RRule, b: RRule): number => a.Order - b.Order
const csv = (s: string): string[] => s.split(',').map((x) => x.trim()).filter(Boolean)
// --- ruleset helpers --------------------------------------------------------
// A `config ruleset` (api.ts Ruleset) is a named domain/ipcidr list a rule
@@ -212,13 +210,13 @@ function everyLabel(sec: number): string {
return `every ${sec}s`
}
/** A rule with no matcher of any kind is the effective catch-all (route Final). */
/** A rule with no matcher of any kind is the effective catch-all (route Final).
* Mirrors model.IsCatchAll on the daemon side — the two must agree or the
* "never applies" badge lands on a different row than the apply warning. */
function isCatchAll(r: RRule): boolean {
return (
len(r.Src) === 0 &&
len(r.DstDomain) === 0 &&
len(r.DstRuleset) === 0 &&
len(r.DstIP) === 0 &&
!(r.DstPort && r.DstPort.trim()) &&
!(r.Proto && r.Proto.trim())
)
@@ -262,16 +260,15 @@ interface TargetGroups {
nodes: TargetOpt[] // node:<n> (huge — rendered last)
}
// Free-text destination matchers offered by the ADD form. Domains are NOT one of
// them (the user's call): domain matching goes through named rulesets — that's
// what they exist for. 'none' = the rule matches by rulesets/source/proto alone.
// (Legacy rules that already carry DstDomain stay editable in the edit form.)
type MatchKind = 'none' | 'ip' | 'port'
// The add form's fields. WHERE traffic is going is a ruleset choice and nothing
// else — a rule has no inline domain or address list any more, so the old
// Match-kind picker (rulesets / ip / port) collapsed into a plain Port field
// beside the ruleset picker. Adding one domain is still one step: the picker can
// build a list on the spot (RulesetPicker's "New list").
interface AddForm {
name: string
src: string[]
matchKind: MatchKind
matchValue: string
port: string
rulesets: string[]
proto: string
target: string
@@ -283,8 +280,7 @@ interface AddForm {
const EMPTY_FORM: AddForm = {
name: '',
src: [],
matchKind: 'none',
matchValue: '',
port: '',
rulesets: [],
proto: '',
target: 'direct',
@@ -334,6 +330,22 @@ export default function Routing() {
toastTimer.current = window.setTimeout(() => setToast(null), 2600)
}, [])
// Which rules can never fire, keyed by their position in the model's Rules array
// — NOT by name. The config that made this necessary had two rules both called
// `default`, which is exactly when a name-keyed verdict badges the wrong row.
const [reach, setReach] = useState<Map<number, RuleReach>>(new Map())
const loadReach = useCallback(async () => {
try {
const { rules } = await getRulesReachability()
setReach(new Map(rules.map((r) => [r.index, r])))
} catch {
// An older daemon has no such endpoint, and a stopped one answers nothing.
// Drop the verdicts rather than keep stale ones: no badge is honest, a badge
// about the previous config is not.
setReach(new Map())
}
}, [])
const load = useCallback(async () => {
try {
setLoadError(null)
@@ -341,7 +353,8 @@ export default function Routing() {
} catch (e) {
setLoadError(errMsg(e))
}
}, [])
void loadReach()
}, [loadReach])
useEffect(() => {
void load()
}, [load])
@@ -407,6 +420,36 @@ export default function Routing() {
// Every rule name, for the edit form's duplicate-name guard (it excludes self).
const ruleNames = useMemo(() => new Set(rules.map((r) => r.Name)), [rules])
// Each rule's position in the model's Rules array — the key the daemon's
// reachability verdicts use. `rules` above is a sorted COPY of the same object
// references, so identity survives the sort and this map stays valid.
const modelIndex = useMemo(() => {
const m = new Map<RRule, number>()
;((config?.Rules as RRule[] | null | undefined) ?? []).forEach((r, i) => m.set(r, i))
return m
}, [config])
/**
* The verdict for one rule, or null when it can fire.
*
* Verdicts are fetched separately from the config, so between an optimistic edit
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
* echoed name and order is what stops that window from putting a "never applies"
* badge on a working rule: a mismatch means the verdict is not about this row,
* and no badge is the honest answer.
*/
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const i = modelIndex.get(r)
if (i === undefined) return null
const v = reach.get(i)
if (!v || !v.unreachable || !v.shadowed_by) return null
if (v.name !== r.Name || v.order !== r.Order) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
},
[modelIndex, reach],
)
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
const rulesets = useMemo<Ruleset[]>(
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
@@ -458,6 +501,10 @@ export default function Routing() {
await putConfig(next)
setSavedPending(true)
flash(okMsg)
// The verdicts describe the config on disk, which just changed — re-ask.
// Adding or moving a rule is precisely what turns a working default into a
// superseded one, and vice versa.
void loadReach()
} catch (e) {
setConfig(prev)
setActionError(errMsg(e))
@@ -466,7 +513,7 @@ export default function Routing() {
setSaving(false)
}
},
[config, flash],
[config, flash, loadReach],
)
const commitRules = useCallback(
@@ -622,18 +669,15 @@ export default function Routing() {
return
}
setFormError(null)
const mv = form.matchValue.trim()
const rule: RRule = {
Name: name,
Enabled: true,
Order: 0,
Src: form.src,
// Domains are matched via rulesets only — the add form has no free-text
// domain matcher by design.
DstDomain: [],
// Destination = rulesets, always. Domains and addresses live in a
// `config ruleset` so one list serves every rule that needs it.
DstRuleset: form.rulesets,
DstIP: form.matchKind === 'ip' ? csv(mv) : [],
DstPort: form.matchKind === 'port' ? mv : '',
DstPort: form.port.trim(),
Proto: form.proto,
Target: form.target,
Egress: '',
@@ -737,14 +781,18 @@ export default function Routing() {
{rules.length === 0 ? (
<div className="rt-empty">
<p>No rules — all traffic follows the default route.</p>
<p className="rt-empty-sub">Add a rule below to steer a domain, address, or port.</p>
<p className="rt-empty-sub">Add a rule below to steer a destination list, source, or port.</p>
</div>
) : (
<ol className="rt-list" aria-label="Routing rules in first-match order">
{rules.map((r, i) =>
editingRule === r.Name ? (
<RuleEditForm
key={r.Name}
// Rule names are NOT unique in the wild — the config that prompted
// the never-applies badge had two rules called `default`, and a
// duplicate React key makes the second row shadow the first. The
// model index disambiguates without changing row identity.
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
initial={r}
names={ruleNames}
targets={targets}
@@ -755,12 +803,13 @@ export default function Routing() {
/>
) : (
<RuleRow
key={r.Name}
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
rule={r}
index={i}
total={rules.length}
busy={saving}
editingOther={editingRule !== null}
shadow={shadowOf(r)}
onEdit={onEditRule}
onToggle={onToggle}
onMove={onMove}
@@ -810,6 +859,7 @@ function RuleRow({
total,
busy,
editingOther,
shadow,
onEdit,
onToggle,
onMove,
@@ -820,6 +870,8 @@ function RuleRow({
total: number
busy: boolean
editingOther: boolean
/** Set when the daemon reports this rule can never fire; null when it can. */
shadow: { by: string; byOrder: number; reason: string } | null
onEdit: (name: string) => void
onToggle: (name: string) => void
onMove: (name: string, dir: 'up' | 'down') => void
@@ -829,13 +881,20 @@ function RuleRow({
// the first-match order can't shift under the open form. Edit itself stays live —
// clicking it just swaps which row is being edited.
const frozen = busy || editingOther
const isDefault = isCatchAll(rule)
// A rule with no conditions is the router's default — but only ONE of them can
// be, and the daemon says which. A superseded one must not wear the default's
// marks (the dashed accent frame, the "· final" order chip, the "everything not
// matched above" line): those are the claim that made two `default` rules
// indistinguishable in the first place.
const dead = shadow !== null
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
const cls = [
'rt-rule',
rule.Enabled ? '' : 'off',
isDefault ? 'final' : '',
dead ? 'dead' : '',
]
.filter(Boolean)
.join(' ')
@@ -852,7 +911,7 @@ function RuleRow({
>
▲
</button>
<span className={isDefault ? 'rt-ord final' : 'rt-ord'}>
<span className={isDefault ? 'rt-ord final' : dead ? 'rt-ord dead' : 'rt-ord'}>
{isDefault ? '·' : rule.Order}
</span>
<button
@@ -870,9 +929,19 @@ function RuleRow({
<div className="rt-head">
<span className="rt-name">{rule.Name}</span>
{isDefault && <span className="rt-badge">default route · final</span>}
{dead && <span className="rt-badge dead">never applies</span>}
</div>
<div className="rt-match">
{isDefault ? (
{dead ? (
// The badge says it never fires; this line says what beat it and what to
// do. Visible text, not a tooltip — the operator has to be able to find
// the other rule, and two rows can carry the same name.
<span className="rt-dead-note" title={shadow.reason}>
“{shadow.by}” (order {shadow.byOrder}) has no conditions either and runs after this
one, so it is the default the router uses. Give this rule a condition, or delete one
of the two.
</span>
) : isDefault ? (
<span className="rt-nomatch">everything not matched above</span>
) : (
<Matchers rule={rule} />
@@ -935,9 +1004,7 @@ function Matchers({ rule }: { rule: RRule }): ReactNode {
)
}
listChip('src', rule.Src, 'src')
listChip('dns', rule.DstDomain, 'dom')
listChip('ruleset', rule.DstRuleset, 'rs')
listChip('ip', rule.DstIP, 'ip')
if (rule.DstPort && rule.DstPort.trim()) {
chips.push(
<span className="rt-chip" key="port">
@@ -1053,7 +1120,16 @@ function TargetOptions({ targets, current }: { targets: TargetGroups; current?:
)
}
/** The dst_ruleset checkbox group. Renders nothing when no rulesets exist. */
/**
* The destination picker: which rulesets this rule matches (dst_ruleset).
*
* Checkboxes and nothing else. This is the ONLY way a rule names a destination,
* so it renders even when the config has no lists yet — an empty picker that says
* where lists come from is the honest answer, and hiding it would leave the rule
* form with no destination control at all. Building and filling a list is the
* Rulesets panel's job, deliberately kept out of the rule editor so a list is
* created in exactly one place.
*/
function RulesetPicker({
options,
selected,
@@ -1065,24 +1141,26 @@ function RulesetPicker({
busy: boolean
onToggle: (name: string) => void
}): ReactNode {
if (options.length === 0) return null
return (
<div className="rt-rsel">
<span className="rt-flabel">Match rulesets — dst_ruleset</span>
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
<span className="rt-flabel">Destination — dst_ruleset</span>
{options.length > 0 && (
<div className="rt-rsel-opts" role="group" aria-label="Match these rulesets">
{options.map((n) => {
const on = selected.includes(n)
return (
<label key={n} className={on ? 'rt-rsel-opt on' : 'rt-rsel-opt'}>
<input type="checkbox" checked={on} onChange={() => onToggle(n)} disabled={busy} />
<span className="mono">{n}</span>
</label>
)
})}
</div>
)}
<p className="rt-rsel-hint">
The rule also matches any traffic in the checked list(s). Combine with a domain, address, or
port, or use a ruleset on its own.
{options.length === 0
? 'No rulesets yet. Add one under Rulesets below, then come back and check it here — a rule matches a destination through a ruleset only.'
: 'The rule matches traffic in ANY checked list. Narrow it further with a source, port or protocol.'}
</p>
</div>
)
@@ -1208,7 +1286,6 @@ function AddRule({
set('rulesets', form.rulesets.includes(n) ? form.rulesets.filter((x) => x !== n) : [...form.rulesets, n])
const toggleDay = (d: string) =>
set('schedDays', form.schedDays.includes(d) ? form.schedDays.filter((x) => x !== d) : [...form.schedDays, d])
const matchPlaceholder = form.matchKind === 'ip' ? '10.0.0.0/8, 100.64.0.0/10' : '443, 8080-8090'
return (
<form className="rt-add" onSubmit={onSubmit} aria-label="Add a routing rule">
@@ -1241,32 +1318,17 @@ function AddRule({
</label>
<label className="rt-field">
<span className="rt-flabel">Match</span>
<select
<span className="rt-flabel">Port(s)</span>
<input
className="rt-input mono"
value={form.matchKind}
onChange={(e) => set('matchKind', e.target.value as MatchKind)}
>
<option value="none">rulesets only</option>
<option value="ip">ip / cidr</option>
<option value="port">port</option>
</select>
value={form.port}
onChange={(e) => set('port', e.target.value)}
placeholder="443, 8080-8090"
autoComplete="off"
spellCheck={false}
/>
</label>
{form.matchKind !== 'none' && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">{form.matchKind === 'port' ? 'Port(s)' : 'Address(es)'}</span>
<input
className="rt-input mono"
value={form.matchValue}
onChange={(e) => set('matchValue', e.target.value)}
placeholder={matchPlaceholder}
autoComplete="off"
spellCheck={false}
/>
</label>
)}
<label className="rt-field">
<span className="rt-flabel">Proto</span>
<select
@@ -1324,11 +1386,11 @@ function AddRule({
}
// --- edit-a-rule plate (inline, replaces the row it edits) ------------------
// Unlike AddRule, editing exposes all three destination matchers at once
// (Domain(s) / IP-CIDR(s) / Port) rather than a single Match picker — a real rule
// can carry several matcher kinds simultaneously and none may be silently dropped.
// The full original rule is spread into the result on save, so Order / Enabled /
// Kill / Egress (and anything else off-form) survive untouched.
// Same fields as AddRule, on purpose: a rule carries exactly one destination
// mechanism (rulesets) plus port/proto/source, so there is nothing an edit can
// reveal that the add form hides. The full original rule is spread into the
// result on save, so Order / Enabled / Kill / Egress (and anything else off-form)
// survive untouched.
function RuleEditForm({
initial,
names,
@@ -1348,8 +1410,6 @@ function RuleEditForm({
}) {
const [name, setName] = useState(initial.Name)
const [src, setSrc] = useState<string[]>([...(initial.Src ?? [])])
const [domain, setDomain] = useState((initial.DstDomain ?? []).join(', '))
const [ip, setIp] = useState((initial.DstIP ?? []).join(', '))
const [port, setPort] = useState(initial.DstPort ?? '')
const [proto, setProto] = useState(initial.Proto ?? '')
const [target, setTarget] = useState(effectiveTarget(initial))
@@ -1367,12 +1427,7 @@ function RuleEditForm({
// A rule with no matcher of any kind is a catch-all — legal, but worth flagging.
const noMatchers =
src.length === 0 &&
csv(domain).length === 0 &&
csv(ip).length === 0 &&
port.trim() === '' &&
rulesets.length === 0 &&
proto.trim() === ''
src.length === 0 && port.trim() === '' && rulesets.length === 0 && proto.trim() === ''
const submit = (e: FormEvent) => {
e.preventDefault()
@@ -1394,8 +1449,6 @@ function RuleEditForm({
...initial,
Name: nm,
Src: src,
DstDomain: csv(domain),
DstIP: csv(ip),
DstPort: port.trim(),
DstRuleset: rulesets,
Proto: proto,
@@ -1414,7 +1467,9 @@ function RuleEditForm({
<form className="rt-add rt-edit-form" onSubmit={submit} aria-label={`Edit rule ${initial.Name}`}>
<div className="rt-add-hd">
<span className="rt-add-title">Edit {initial.Name}</span>
<span className="rt-add-sub">Empty a field to drop that matcher. Save, then Apply.</span>
<span className="rt-add-sub">
Uncheck a list or clear a field to drop that matcher. Save, then Apply.
</span>
</div>
<div className="rt-fields">
@@ -1444,37 +1499,6 @@ function RuleEditForm({
/>
</label>
{/* Domains are matched via rulesets by design — this legacy field only
appears when the rule ALREADY carries free-text domains, so they
stay visible and clearable rather than silently preserved. */}
{(initial.DstDomain ?? []).length > 0 && (
<label className="rt-field rt-field-wide">
<span className="rt-flabel">Domain(s) — legacy</span>
<input
className="rt-input mono"
value={domain}
onChange={(e) => setDomain(e.target.value)}
placeholder="youtube.com, *.googlevideo.com"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
)}
<label className="rt-field rt-field-wide">
<span className="rt-flabel">IP / CIDR(s)</span>
<input
className="rt-input mono"
value={ip}
onChange={(e) => setIp(e.target.value)}
placeholder="10.0.0.0/8, 100.64.0.0/10"
autoComplete="off"
spellCheck={false}
disabled={busy}
/>
</label>
<label className="rt-field">
<span className="rt-flabel">Port(s)</span>
<input
+96 -2
View File
@@ -16,6 +16,7 @@ import (
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/control"
E "github.com/sagernet/sing/common/exceptions" // lx:tproxy_writeback_connect
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
@@ -68,7 +69,26 @@ func (t *TProxy) Start(stage adapter.StartStage) error {
}
func (t *TProxy) Close() error {
return t.listener.Close()
err := t.listener.Close()
// lx:begin tproxy_writeback_connect
// Closing the listener stops INGRESS but leaves every live UDP NAT session in
// the cache, and each session holds a write-back socket bound to its original
// destination. Nothing else ever wakes those sessions: the cache evicts
// lazily (on the next Get/Add), and after this inbound is gone there is no
// next Get. For a DNS session answered in-engine by hijack-dns there is not
// even an outbound connection whose failure could unwind it, so its socket
// survives until a GC finalizer happens to reach it.
//
// That is invisible upstream, where an inbound is closed once at shutdown. On
// this fork the engine is REBUILT on every config apply, so each apply would
// strand another generation of sockets on the router's own LAN :53. Purge
// evicts every session; the cache's OnEvict closes the conn, which unblocks
// the session's routing goroutine and runs its onClose — the one place the
// write-back socket is actually closed. Purge AFTER the listener so a packet
// arriving mid-teardown cannot re-create a session behind us.
t.udpNat.Purge()
// lx:end tproxy_writeback_connect
return err
}
func (t *TProxy) NewConnection(ctx context.Context, conn net.Conn, metadata adapter.InboundContext, onClose N.CloseHandlerFunc) {
@@ -121,18 +141,92 @@ type tproxyPacketWriter struct {
conn *net.UDPConn
}
// lx:begin tproxy_writeback_connect
//
// The TPROXY UDP write-back socket is bound to the ORIGINAL DESTINATION, so the
// client sees the reply coming from the address it addressed. Upstream leaves
// that socket UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort), and that is
// the bug this block exists for.
//
// An unconnected socket bound to <addr>:<port> is, as far as the kernel is
// concerned, a RECEIVER for that address:port — and because the socket also
// carries SO_REUSEADDR it silently joins the UDP demultiplex set of whatever
// else is bound there. Nothing ever reads from it: this writer only sends. So
// every datagram the kernel happens to hand it is lost.
//
// On a router that transparently intercepts LAN DNS towards its OWN address
// (`nft ... udp dport 53 tproxy ...`) the original destination IS the router's
// LAN address, so each intercepted DNS session parks another silent receiver on
// <router-lan-ip>:53 right next to dnsmasq's socket — observed on the stand:
// 33 such sockets against dnsmasq's one, several with a growing Recv-Q. The
// host's own queries to that address take the loopback path, are never diverted,
// and are therefore demultiplexed among all of them: they land in one of the
// silent sockets at random and time out, permanently and unpredictably, while
// every other DNS path on the box keeps working.
//
// Connecting the socket fixes it at the root. compute_score() in the kernel's
// UDP lookup REJECTS a connected socket for any peer other than the connected
// one, so a write-back socket can no longer be handed a datagram it will not
// read. Nothing about the reply changes — same spoofed source address, same
// single peer, one Write instead of one WriteTo.
//
// The unconnected path is kept verbatim for a destination that cannot be bound
// (a domain socksaddr), so no existing case regresses.
// newTProxyWriteBack creates the connected write-back socket: local address =
// the original destination (transparent bind), peer = the client. A package
// variable so the behaviour can be tested without CAP_NET_ADMIN.
var newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
var dialer net.Dialer
dialer.LocalAddr = destination.UDPAddr()
dialer.Control = control.Append(dialer.Control, control.ReuseAddr())
dialer.Control = control.Append(dialer.Control, redir.TProxyWriteBack())
conn, err := w.listener.DialContext(dialer, w.ctx, "udp", w.source.String())
if err != nil {
return nil, err
}
udpConn, loaded := conn.(*net.UDPConn)
if !loaded {
conn.Close()
return nil, E.New("tproxy write back: unexpected connection type ", conn)
}
return udpConn, nil
}
// lx:end tproxy_writeback_connect
func (w *tproxyPacketWriter) WritePacket(buffer *buf.Buffer, destination M.Socksaddr) error {
defer buffer.Release()
if w.listener.ListenOptions().NetNs == "" {
conn := w.conn
if w.destination == destination && conn != nil {
_, err := conn.WriteToUDPAddrPort(buffer.Bytes(), w.source)
// lx:begin tproxy_writeback_connect
// The cached socket is CONNECTED to the client, so this is a plain
// Write. A failed write also closes it: upstream only dropped the
// reference, leaving the fd to the GC finalizer.
_, err := conn.Write(buffer.Bytes())
if err != nil {
conn.Close()
w.conn = nil
}
return err
// lx:end tproxy_writeback_connect
}
}
// lx:begin tproxy_writeback_connect
if destination.IsIP() {
udpConn, err := newTProxyWriteBack(w, destination)
if err != nil {
return err
}
if w.listener.ListenOptions().NetNs == "" && w.destination == destination {
w.conn = udpConn
} else {
defer udpConn.Close()
}
return common.Error(udpConn.Write(buffer.Bytes()))
}
// lx:end tproxy_writeback_connect
var listenConfig net.ListenConfig
listenConfig.Control = control.Append(listenConfig.Control, control.ReuseAddr())
listenConfig.Control = control.Append(listenConfig.Control, redir.TProxyWriteBack())
@@ -0,0 +1,234 @@
package redirect
// Regression cover for the lx:tproxy_writeback_connect block in tproxy.go.
//
// The failure it guards against, seen on a live BPi-R3 Mini: `shaterd` held 33
// UNCONNECTED UDP sockets on the router's own LAN address :53 — the write-back
// sockets of intercepted DNS sessions — alongside dnsmasq's single socket on the
// same address:port. Nothing reads a write-back socket, so every host-originated
// query that the kernel demultiplexed into one of them was silently dropped
// (Recv-Q climbing, sender timing out). Connecting the socket to the one peer it
// ever talks to removes it from the demultiplex set for everybody else.
//
// The real socket needs CAP_NET_ADMIN (IP_TRANSPARENT) and a foreign bind, so
// these tests drive the seam (newTProxyWriteBack) with an ordinary connected
// loopback socket. That is enough to pin both properties that actually broke:
//
// - the write path uses the CONNECTED form. If WritePacket ever goes back to
// WriteToUDPAddrPort, Go returns ErrWriteToConnected on a connected socket
// and these tests fail — i.e. the test cannot pass with an unconnected
// write-back socket, which is exactly the regression.
// - repeated writes to the same destination REUSE one socket instead of
// accumulating a new one per packet.
import (
"context"
"net"
"net/netip"
"testing"
"time"
"github.com/sagernet/sing-box/common/listener"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/buf"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
)
// writeBackHarness stands up a loopback "client" socket and a tproxyPacketWriter
// whose socket factory returns a plain connected UDP socket aimed at it. It
// returns the writer, the client socket and a pointer to the factory call count.
func writeBackHarness(t *testing.T) (*tproxyPacketWriter, *net.UDPConn, *int) {
t.Helper()
client, err := net.ListenUDP("udp", &net.UDPAddr{IP: net.IPv4(127, 0, 0, 1)})
if err != nil {
t.Fatalf("client socket: %v", err)
}
t.Cleanup(func() { client.Close() })
source := netip.MustParseAddrPort(client.LocalAddr().String())
w := &tproxyPacketWriter{
ctx: context.Background(),
source: source,
// A listener with empty options is enough: WritePacket only reads
// ListenOptions().NetNs, and the factory is stubbed below.
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
destination: M.SocksaddrFrom(netip.MustParseAddr("127.0.0.1"), 53),
}
calls := 0
orig := newTProxyWriteBack
t.Cleanup(func() { newTProxyWriteBack = orig })
newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
calls++
// The real implementation binds the ORIGINAL DESTINATION transparently
// and connects to w.source; here we only reproduce the connect half,
// which is the property under test.
conn, err := net.DialUDP("udp", nil, net.UDPAddrFromAddrPort(w.source))
if err != nil {
return nil, err
}
return conn, nil
}
return w, client, &calls
}
// readOne reads one datagram from the client socket with a short deadline.
func readOne(t *testing.T, client *net.UDPConn) string {
t.Helper()
_ = client.SetReadDeadline(time.Now().Add(2 * time.Second))
b := make([]byte, 512)
n, _, err := client.ReadFrom(b)
if err != nil {
t.Fatalf("client read: %v", err)
}
return string(b[:n])
}
// TestWriteBackUsesConnectedSocket: the reply must go out over a CONNECTED
// socket. On the pre-fix code the write is WriteToUDPAddrPort, which Go refuses
// on a connected socket ("use of WriteTo with pre-connected connection"), so
// this test is red for exactly the shape that caused the outage.
func TestWriteBackUsesConnectedSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("first")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
if got := readOne(t, client); got != "first" {
t.Fatalf("client got %q, want %q", got, "first")
}
if *calls != 1 {
t.Fatalf("expected one write-back socket, got %d", *calls)
}
}
// TestWriteBackReusesOneSocket: a session that keeps answering the same client
// must keep ONE socket, not open a fresh one per datagram. The live box carried
// one socket per intercepted DNS session already; one per PACKET would turn a
// nuisance into an fd exhaustion.
func TestWriteBackReusesOneSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
for i, payload := range []string{"a", "b", "c", "d"} {
if err := w.WritePacket(buf.As([]byte(payload)).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket %d: %v", i, err)
}
if got := readOne(t, client); got != payload {
t.Fatalf("packet %d: client got %q, want %q", i, got, payload)
}
}
if *calls != 1 {
t.Fatalf("four packets to one destination must share one socket, got %d sockets", *calls)
}
if w.conn == nil {
t.Fatalf("the write-back socket must be cached on the writer for reuse")
}
}
// TestWriteBackClosesSocketOnWriteFailure: upstream dropped the reference to a
// failed socket without closing it, leaving the fd to the GC finalizer. On a box
// that already parks one socket per DNS session that is the wrong direction.
func TestWriteBackClosesSocketOnWriteFailure(t *testing.T) {
w, client, _ := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("warm")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
readOne(t, client)
cached := w.conn
if cached == nil {
t.Fatalf("expected a cached socket after the first write")
}
// Close it behind WritePacket's back so the next write fails, exactly as a
// dead peer or a torn-down plane would make it fail.
cached.Close()
if err := w.WritePacket(buf.As([]byte("boom")).ToOwned(), w.destination); err == nil {
t.Fatalf("a write on a closed socket must report the failure")
}
if w.conn != nil {
t.Fatalf("a failed write must drop the cached socket")
}
// A second Close on an already-closed conn is an error, which is how we know
// WritePacket closed it rather than merely forgetting it.
if err := cached.Close(); err == nil {
t.Fatalf("WritePacket must CLOSE the failed socket, not just nil the field")
}
}
// --- Close() must release the UDP NAT sessions --------------------------------
// natSessionHandler stands in for the router: it drains the session conn until
// it errors (which is what Close does to it) and then reports onClose, exactly
// as the real routing goroutine does. onClose is where the write-back socket is
// closed, so "onClose fired" is the observable proof the socket was released.
type natSessionHandler struct {
released chan error
}
func (h *natSessionHandler) NewPacketConnectionEx(ctx context.Context, conn N.PacketConn, source M.Socksaddr, destination M.Socksaddr, onClose N.CloseHandlerFunc) {
go func() {
for {
buffer := buf.NewSize(1024)
_, err := conn.ReadPacket(buffer)
buffer.Release()
if err != nil {
if onClose != nil {
onClose(err)
}
h.released <- err
return
}
}
}()
}
// TestTProxyCloseReleasesNatSessions pins the apply-level half of the leak.
//
// Closing the inbound used to close the listener only. The NAT cache evicts
// lazily, so after the inbound is gone nothing ever touches it again and every
// live session — with the write-back socket it holds on the router's own
// LAN :53 — was left to a GC finalizer. On this fork the engine is rebuilt on
// every config apply, so that is one stranded generation of sockets per apply.
func TestTProxyCloseReleasesNatSessions(t *testing.T) {
handler := &natSessionHandler{released: make(chan error, 1)}
tp := &TProxy{
ctx: context.Background(),
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
}
prepared := 0
tp.udpNat = udpnat.New(handler, func(source M.Socksaddr, destination M.Socksaddr, userData any) (bool, context.Context, N.PacketWriter, N.CloseHandlerFunc) {
prepared++
return true, context.Background(), nil, func(error) {}
}, time.Minute, false)
tp.udpNat.NewPacket(
[][]byte{{0x00}},
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.2"), 40000),
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.1"), 53),
nil,
)
if prepared != 1 {
t.Fatalf("expected one NAT session, got %d", prepared)
}
if err := tp.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
select {
case <-handler.released:
case <-time.After(5 * time.Second):
t.Fatal("Close must release every live NAT session (and with it the write-back socket it holds)")
}
}
+9 -4
View File
@@ -27,8 +27,11 @@
#
# VERSION Version string stamped into constant.Version. Resolution order:
# 1) this positional arg, if given
# 2) $SHATER_VERSION, if set
# 3) `git describe --tags` (nearest tag + commit)
# 2) $SHATER_VERSION, if set (CI sets it from ci/version.sh)
# 3) `ci/version.sh --binary` — THE single source of truth shared
# with the package version (vX.Y.Z-rR[-g<sha>], derived from
# the git tag exactly like PKG_VERSION/PKG_RELEASE), so the
# string the panel shows always matches `apk info shaterd`
# 4) fallback: v0.2.0-dev
# --fast Skip `npm ci` when panel/node_modules already exists (dev speed-up).
#
@@ -69,8 +72,10 @@ if [ -n "$VERSION_ARG" ]; then
VERSION="$VERSION_ARG"
elif [ -n "${SHATER_VERSION:-}" ]; then
VERSION="$SHATER_VERSION"
elif VERSION="$(git -C "$REPO" describe --tags 2>/dev/null)"; then
: # git describe succeeded
elif VERSION="$(sh "$REPO/ci/version.sh" --binary)"; then
# Same computation the PACKAGE version comes from (ci/version.sh), so the
# binary's constant.Version and the .ipk/.apk version can never drift apart.
: # ci/version.sh always succeeds (it falls back to 0.0.0 without git)
else
VERSION="v0.2.0-dev"
fi
+17
View File
@@ -488,6 +488,23 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
return changed, err
}
// The plane is COMPLETE only here: table + policy routing + sysctls. ApplyNft
// already flushed the DNS conntrack when it loaded the ruleset, but that is
// one step too early — ApplyRouting is idempotent BY del-then-add, so it opens
// a window in which the fwmark rule is momentarily absent, and any DNS flow
// that crosses that window is tracked against a plane that is still being
// assembled. Flushing once more now that every piece is in place is what makes
// "no entry survives the transition" actually true. Only on a real change (the
// fast path assembled nothing), best-effort, and cheap: the :53 entry count is
// bounded by the number of clients.
if !nftCurrent {
if n, ferr := netplane.FlushDNSConntrack(); ferr != nil {
a.log.Debug("flush DNS conntrack after plane change: ", ferr)
} else if n > 0 {
a.log.Debug("plane changed: dropped ", n, " stale DNS conntrack entries")
}
}
// (4) success. Bump the effective-state generation ONLY when something really
// moved: a no-op reconcile must not invalidate an armed commit-confirm window
// (cron reconciles every minute — counting those would cancel every rollback
+22 -8
View File
@@ -22,6 +22,16 @@ func tunnelModel() *model.Model {
return m
}
// pinnedIPSet declares the inline `type=ipcidr` rule-set a rule pins its
// destination addresses with. Since schema v2 a routing rule has no dst_ip of its
// own: addresses are a `config ruleset`, so the generated rule carries a rule_set
// reference and the plan resolves the actual prefixes through the lookup below —
// i.e. through the RUNNING engine in production (engine.RuleSetIPCIDRs), not out
// of the config text. planFor's table stands in for that.
func pinnedIPSet(name string, cidrs ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: cidrs}
}
// planFor generates m and builds the untunnelable plan, resolving rule-set tags
// from the supplied table. A tag absent from the table reports "not loaded".
func planFor(t *testing.T, m *model.Model, sets map[string][]string) *netplane.UntunnelablePlan {
@@ -58,11 +68,12 @@ func renderPlan(t *testing.T, m *model.Model, plan *netplane.UntunnelablePlan) s
// only 8.8.8.8 may be un-pingable and the rest of the internet must answer.
func TestOnlyPinnedAddressIsTunnelled(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}})
if !plan.DefaultAllow {
t.Fatalf("a catch-all `direct` rule must make the default ALLOW; plan=%+v", plan)
@@ -224,11 +235,12 @@ func TestDomainRuleDoesNotAffectUntunnelable(t *testing.T) {
// TestBlockedRuleDenies: an explicitly blocked destination stays dropped.
func TestBlockedRuleDenies(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("bad", "203.0.113.0/24")}
m.Rules = []model.Rule{
{Name: "bad", Enabled: true, Order: 10, DstIP: []string{"203.0.113.0/24"}, Target: "block"},
{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "block"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-bad": {"203.0.113.0/24"}})
if len(plan.Matches) != 1 || plan.Matches[0].Allow {
t.Fatalf("a blocked destination must deny untunnelable traffic too: %+v", plan.Matches)
}
@@ -361,11 +373,12 @@ func TestPlanNeverAcceptsTCPOrUDP(t *testing.T) {
m := tunnelModel()
m.Globals.IPv6 = true
m.Globals.Untunnelable = policy
m.Rulesets = []model.Ruleset{pinnedIPSet("pin", "8.8.8.8/32")}
m.Rules = []model.Rule{
{Name: "pin", Enabled: true, Order: 10, DstIP: []string{"8.8.8.8/32"}, Target: "group:auto"},
{Name: "pin", Enabled: true, Order: 10, DstRuleset: []string{"pin"}, Target: "group:auto"},
{Name: "rest", Enabled: true, Order: 99, Target: "direct"},
}
fwd := renderPlan(t, m, planFor(t, m, nil))
fwd := renderPlan(t, m, planFor(t, m, map[string][]string{"rs-pin": {"8.8.8.8/32"}}))
// Only the lines the untunnelable policy emits are in scope: the tproxy
// diverts in prerouting legitimately match TCP/UDP, which is their job.
@@ -424,12 +437,13 @@ func TestLocalPlaneSurvivesEveryPlan(t *testing.T) {
// only grant those clients, not everyone.
func TestSourceScopedRuleNarrowsTheAllow(t *testing.T) {
m := tunnelModel()
m.Rulesets = []model.Ruleset{pinnedIPSet("lab", "198.51.100.0/24")}
m.Rules = []model.Rule{
{Name: "lab", Enabled: true, Order: 10, Src: []string{"192.168.9.0/24"},
DstIP: []string{"198.51.100.0/24"}, Target: "direct"},
DstRuleset: []string{"lab"}, Target: "direct"},
{Name: "dflt", Enabled: true, Order: 99, Target: "group:auto"},
}
plan := planFor(t, m, nil)
plan := planFor(t, m, map[string][]string{"rs-lab": {"198.51.100.0/24"}})
if len(plan.Matches) != 1 {
t.Fatalf("expected one step, got %+v", plan.Matches)
}
+9
View File
@@ -91,6 +91,15 @@ var criticalMarkers = []string{
"not covered", // an interface outside the fail-closed guard
"fail-closed", // ''
"REJECTED", // an unusable interface name
// A condition-less rule retired by a later condition-less rule whose target is
// `direct` (generate/route.go warnUnreachableRules): the operator's default
// policy — a tunnel, or a block — is not the one the router uses, so everything
// unmatched leaves on the plain WAN. The marker is the full clause, not the
// shorter "leaves over the plain WAN" that ruleKillFallback's kill=open note
// also contains: that one is a DELIBERATE per-rule bypass the operator asked
// for, and grading it critical here would be a different decision made by
// accident.
"leaves over the plain WAN with your real IP address",
}
// protectionSections are entity kinds whose whole purpose is to block or divert
+42
View File
@@ -240,3 +240,45 @@ func blockGlobals() model.Globals {
g.Untunnelable = netplane.UntunnelableBlock
return g
}
// TestCollectWarningsGradesUnreachableRules (B1): a condition-less rule retired by
// a later condition-less one is graded by CONSEQUENCE, not by the mere fact that a
// setting is dead.
//
// - the surviving default is `direct` while the retired one wanted a tunnel:
// the operator's default policy is not in effect and everything unmatched
// leaves on the plain WAN — critical, the panel must show it loudly;
// - the surviving default is the tunnel and the dead one was `direct`: a dead
// knob, nothing is leaking — warning.
//
// Both must be attributed to section "rule" + the rule's name, so the panel can
// deep-link to the offending row instead of printing prose.
func TestCollectWarningsGradesUnreachableRules(t *testing.T) {
const leak = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "fallback" (order 100) has none either and comes after it, so "direct" wins and ` +
`this rule's target "group:auto" is never applied. Everything no other rule matches leaves ` +
`over the plain WAN with your real IP address. Delete one of the two, or give this one a condition`
const dead = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "default" (order 100) has none either and comes after it, so the default the ` +
`router uses is "group:auto" and this rule's target "direct" is never applied. Delete one ` +
`of the two, or give this one a condition so it can match something`
check := func(text, wantSeverity string) {
t.Helper()
found := false
for _, w := range collectWarnings(blockGlobals(), []string{text}, nil, nil) {
if w.Section != "rule" || w.Name != "default" {
continue
}
found = true
if w.Severity != wantSeverity {
t.Errorf("severity = %q, want %q for: %s", w.Severity, wantSeverity, text)
}
}
if !found {
t.Fatalf("never-applied warning not attributed to rule %q: %s", "default", text)
}
}
check(leak, SeverityCritical)
check(dead, SeverityWarning)
}
+27
View File
@@ -0,0 +1,27 @@
package main
// Control-plane log formatting (the IO-free half, so it is unit-testable on any
// dev host — same split as profilewatch.go).
//
// The daemon's own logger used to be built with a bare log.Formatter, whose
// DisableColors zero value is false: every control-plane line went to procd's
// stderr — i.e. straight into syslog — carrying aurora escape sequences. See
// shater/logsink/color.go for why that is a defect and not a cosmetic.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/logsink"
)
// controlLogFormatter builds the control-plane logger's formatter. out is the
// process's real stderr (the sink's syslog half): colours are emitted only when
// that is a terminal, never when it is procd/syslog or a log file.
func controlLogFormatter(baseTime time.Time, out *os.File) log.Formatter {
return log.Formatter{
BaseTime: baseTime,
DisableColors: !logsink.IsTTY(out),
}
}
+37
View File
@@ -0,0 +1,37 @@
package main
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestControlLogFormatterNoANSIOffTTY pins the control-plane half of the syslog
// colour leak: when the daemon's stderr is not a terminal — which is ALWAYS the
// case under procd, where stderr is the pipe procd relays to syslog — no line
// the daemon formats may contain an ESC (0x1b).
func TestControlLogFormatterNoANSIOffTTY(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
f := controlLogFormatter(time.Now(), w)
if !f.DisableColors {
t.Fatalf("colours enabled for a non-terminal stderr")
}
// A context ID exercises the second colouring branch of log/format.go (the
// 256-colour connection id), which is what produced ESC[38;5;193m on the router.
ctx := log.ContextWithNewID(context.Background())
for _, level := range []log.Level{log.LevelError, log.LevelWarn, log.LevelInfo, log.LevelDebug, log.LevelTrace} {
line := f.Format(ctx, level, "dns", "exchange failed for example.com. IN AAAA: unexpected EOF", time.Now())
if i := strings.IndexByte(line, 0x1b); i >= 0 {
t.Errorf("level %v: formatted line carries an ANSI escape at byte %d: %q", level, i, line)
}
}
}
+76 -8
View File
@@ -79,7 +79,7 @@ func dispatch(args []string) int {
case "mint-token":
return cmdMintToken()
case "nodes":
return cmdReadStub("nodes", "[]")
return cmdNodes()
case "stats":
return cmdReadStub("stats", "{}")
case "blocklist":
@@ -124,7 +124,7 @@ usage: shaterd <verb>
rollback roll back to the last-good config
status print daemon/data-plane status as JSON
mint-token mint a single-use panel handoff token (JSON) via the daemon
nodes print nodes as JSON
nodes print the configured nodes as JSON (manual + subscription)
stats print stats as JSON
blocklist update force a DNS-filter blocklist refresh (reconcile; no-op if down)
schedule due re-evaluate time-scheduled rules now (reconcile; no-op if down)
@@ -158,8 +158,11 @@ func cmdRun() int {
// The level starts at trace (nothing read yet) and is corrected from
// Globals.LogLevel as soon as UCI is read — the control plane now RESPECTS
// the configured level instead of the old unconditional trace.
// The formatter colours only when stderr is a real terminal
// (controlLogFormatter -> logsink.IsTTY): under procd stderr IS syslog, and
// ANSI escapes there break `logread | grep ERROR` and every log collector.
logFactory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now()}, sink, "", nil, false)
controlLogFormatter(time.Now(), os.Stderr), sink, "", nil, false)
log.SetStdLogger(logFactory.Logger())
logger := log.StdLogger()
@@ -172,10 +175,23 @@ func cmdRun() int {
logFactory.SetLevel(controlLogLevel(g.LogLevel))
}
// Refuse to start a second daemon (race-safe single-owner guard). procd's
// term_timeout ensures the previous instance exits before reload restarts us.
if pid, ok := daemonAlive(); ok && pid != os.Getpid() {
logger.Error("another shaterd is already running (pid ", pid, ") — refusing to start")
// Single-owner guard: never run two daemons at once. On a `restart` procd's
// `service delete` is ASYNCHRONOUS — the ubus call returns immediately while
// the outgoing instance is still running its honest teardown — and the
// following `service add` starts us straight away, so a predecessor being
// alive here is the NORMAL restart case, not an error.
//
// This used to exit(1) on the spot and lean on procd's `respawn ... 5 ...` to
// try again five seconds later. That is a blind retry, not synchronisation:
// it neither knows nor waits for the predecessor's teardown to finish, and it
// turns every restart into at least one logged crash plus a five-second hole
// in which the LAN has no plane at all. Waiting for the predecessor to exit
// makes `restart` behave exactly like `stop` + pause + `start`: our apply is
// then strictly ordered AFTER the previous teardown, which is the whole point
// of the guard.
if pid, waiting := waitForPredecessor(predecessorBudget, predecessorPoll, logger); waiting {
logger.Error("another shaterd is still running (pid ", pid, ") after waiting ",
predecessorBudget, " for it to exit — refusing to start")
return 1
}
if err := writePidfile(); err != nil {
@@ -832,6 +848,42 @@ func cmdStatus() int {
return 0
}
// cmdNodes prints the node inventory as a JSON array (see nodes.go for the
// shape and for why this is a view rather than the raw model.Node).
//
// Shape of the call mirrors cmdStatus: ask the RUNNING daemon first — it is the
// process that owns the engine, so its answer is the inventory the live box was
// built from — and fall back to reading the same on-disk desired state directly
// when there is no daemon (or it did not answer). Both sides call the SAME
// nodesJSON(), so the fallback cannot report something the daemon would not.
//
// The one thing it will never do is print `[]` because it could not find out:
// a read failure goes to stderr and exits non-zero, so "empty" on stdout with
// exit 0 means "no nodes are configured" and nothing else.
func cmdNodes() int {
if _, ok := daemonAlive(); ok {
resp, err := ctlRequest("nodes")
switch {
case err != nil:
fmt.Fprintf(os.Stderr, "shaterd nodes: %v — reading the on-disk config instead\n", err)
case strings.HasPrefix(strings.TrimSpace(resp), "["):
fmt.Println(strings.TrimSpace(resp))
return 0
default:
// The daemon answered, but with an error object rather than a list.
fmt.Fprintf(os.Stderr, "shaterd nodes: daemon: %s\n", strings.TrimSpace(resp))
}
}
b, err := nodesJSON()
if err != nil {
fmt.Fprintf(os.Stderr, "shaterd nodes: %v\n", err)
fmt.Println("[]") // stdout stays parseable JSON; the exit code carries the failure
return 1
}
fmt.Println(string(b))
return 0
}
// cmdMintToken asks the RUNNING daemon (over the control socket) for a single-use
// panel handoff token and prints the daemon's JSON reply verbatim — {"token":"..."}
// on success, {"error":"..."} otherwise. It is the CLI shim the LuCI/rpcd layer
@@ -869,6 +921,13 @@ func printJSONError(msg string) {
fmt.Println(string(b))
}
// cmdReadStub relays a read-only verb to the daemon and prints def when there is
// no daemon to ask. It is ONLY valid where def is the truth in the daemon-down
// case: `stats` counts what the LIVE engine saw, so with no engine there really
// is nothing counted and `{}` says exactly that. It is NOT a way to make a verb
// look implemented — `nodes` used to be routed through here with def="[]" while
// hundreds of nodes sat on disk (see nodes.go). Anything whose data outlives the
// daemon belongs in its own command that reads that data.
func cmdReadStub(cmd, def string) int {
if _, ok := daemonAlive(); ok {
if resp, err := ctlRequest(cmd); err == nil {
@@ -940,7 +999,16 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
case "rollback":
writeResult(conn, false, a.Rollback())
case "nodes":
writeLine(conn, "[]")
// Same assembly as the CLI fallback and as GET /api/config: model.ReadUCI
// (UCI + the per-subscription caches) projected onto the printable view.
b, err := nodesJSON()
if err != nil {
l.Warn("control socket: nodes: ", err)
e, _ := json.Marshal(map[string]string{"error": err.Error()})
writeLine(conn, string(e))
return
}
writeLine(conn, string(b))
case "stats":
if sa == nil {
writeLine(conn, "{}")
+132
View File
@@ -0,0 +1,132 @@
package main
// `shaterd nodes` — the node inventory, as JSON.
//
// This verb used to answer `[]` unconditionally (a `cmdReadStub("nodes", "[]")`
// on the CLI side and a hard-coded `writeLine(conn, "[]")` in the daemon's
// control-socket handler), while `shaterd --help` advertised "print nodes as
// JSON". The data was never missing: /etc/shater/subs/*.json holds every
// subscription-fetched node and `GET /api/config` reports the full merged set —
// on a live router that is hundreds of nodes answered as "none". An empty list
// is indistinguishable from a truthful "no nodes are configured", so the verb
// did not fail loudly, it lied quietly. Same class as the honesty fixes in
// 9dc954029 / aec82d444; the cure is the same: report what is actually there.
//
// # Data path (deliberately the SAME one /api/config uses)
//
// model.ReadUCI() = `uci export shater` (manual nodes + everything else) +
// model.MergeSubCaches (the per-subscription JSON caches). That single call is
// what the panel's handleConfigGet serves, what generate builds the engine from
// and what `sub update` writes back — so `shaterd nodes` cannot drift from the
// panel or from the running engine, because there is no second assembly here to
// drift. This file only PROJECTS that model onto a small, printable view.
//
// # Why a view and not the raw model.Node
//
// model.Node carries the share-link URI, and a share link is a credential
// (uuid/password in the query string). /api/config may return it — that path is
// session-authenticated and the panel needs the URI to edit a node — but a CLI
// verb whose output gets piped into support tickets, `logger`, and cron mail
// must not spray credentials. The view therefore reports the *derived* facts a
// share link answers (protocol, server, port) and drops the secret-bearing URI.
// Nodes whose URI does not parse are still listed, with the parse error in
// `parse_error`: the engine skips exactly those nodes (generate/outbound.go),
// and a node that is configured-but-unusable is precisely what an operator
// needs to see — hiding it would be the same lie in a smaller coat.
import (
"encoding/json"
"strings"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/parse"
)
// nodeView is ONE node as the `nodes` verb reports it. Field set: the model's
// own facts (name/enabled/sub/egress/stale/fingerprint) plus what the share link
// decodes to (protocol/server/port). Everything is omitempty-free where a reader
// would have to distinguish "absent" from "false"/"zero" — `enabled` and `stale`
// are always present so a consumer never has to guess.
type nodeView struct {
Name string `json:"name"`
Enabled bool `json:"enabled"`
// Sub is the subscription this node came from; "" means a manual node
// (model.Node.FromSub semantics, unchanged).
Sub string `json:"sub"`
// Protocol is the parsed protocol (vless|vmess|trojan|shadowsocks|wireguard…).
// When the URI does not parse it falls back to the bare URI scheme so the
// operator still sees what kind of thing failed; "" only when the URI is empty.
Protocol string `json:"protocol"`
Server string `json:"server,omitempty"`
Port uint16 `json:"port,omitempty"`
// Egress is the `config egress` this node's own upstream is bound to
// (multi-WAN); "" = default route.
Egress string `json:"egress,omitempty"`
// Stale marks a cached subscription node whose subscription failed to refresh.
Stale bool `json:"stale"`
Fingerprint string `json:"fingerprint,omitempty"`
// ParseError is the reason the engine will SKIP this node, verbatim from
// parse.ParseShareLink. Empty on every usable node.
ParseError string `json:"parse_error,omitempty"`
}
// readModel is the model source. A package var so the tests can exercise the
// assembly without a router's `uci` binary; production always reads the real
// merged desired state.
var readModel = model.ReadUCI
// nodesJSON reads the desired state and renders the node inventory as a JSON
// array. The array is never `null`: an empty configuration marshals to `[]`,
// which is the ONE case where `[]` is the truth.
func nodesJSON() ([]byte, error) {
m, err := readModel()
if err != nil {
return nil, err
}
return json.Marshal(nodeViews(m))
}
// nodeViews projects the merged model onto the printable view. Pure: no IO, no
// globals — the whole verb's logic is testable on a dev host.
func nodeViews(m *model.Model) []nodeView {
out := make([]nodeView, 0, len(m.Nodes)) // never nil => never `null`
for i := range m.Nodes {
n := m.Nodes[i]
v := nodeView{
Name: n.Name,
Enabled: n.Enabled,
Sub: n.FromSub,
Egress: n.Egress,
Stale: n.Stale,
Fingerprint: n.Fingerprint,
Protocol: uriScheme(n.URI),
}
// The engine reads a node through exactly this call (generate/outbound.go,
// chain.go, group.go); using it here is what makes "protocol" agree with
// what will actually be dialled, and the error agree with what will be
// skipped.
if p, err := parse.ParseShareLink(n.URI); err == nil {
if p.Protocol != "" {
v.Protocol = p.Protocol
}
v.Server = p.Server
v.Port = p.Port
} else if n.URI != "" {
v.ParseError = err.Error()
}
out = append(out, v)
}
return out
}
// uriScheme returns the lower-cased scheme of a share link ("vless://…" ->
// "vless"), or "" when there is none. It is the fallback protocol for a URI that
// ParseShareLink refuses, so an unusable node is still described rather than
// reported as a typeless blank.
func uriScheme(uri string) string {
i := strings.Index(uri, "://")
if i <= 0 {
return ""
}
return strings.ToLower(uri[:i])
}
+146
View File
@@ -0,0 +1,146 @@
package main
import (
"encoding/json"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// vlessURI is a syntactically complete share link (host, port, uuid, fragment)
// so ParseShareLink really succeeds and the derived fields are exercised.
const vlessURI = "vless://11111111-2222-3333-4444-555555555555@example.net:443?" +
"encryption=none&security=tls&sni=example.net&type=ws&path=%2Fws#NL-vless-1"
// writeSubCache points the subscription cache at a temp dir and persists one
// subscription's nodes there, the way `sub update` does on the router
// (/etc/shater/subs/<sub>.json). Returns nothing: the point is the side effect
// that model.MergeSubCaches will pick up.
func writeSubCache(t *testing.T, sub string, nodes []model.Node) {
t.Helper()
t.Setenv("SHATER_SUBS_DIR", t.TempDir())
if err := model.SaveSubCache(sub, nodes); err != nil {
t.Fatalf("SaveSubCache(%q): %v", sub, err)
}
}
// TestNodesJSONReportsCachedSubscriptionNodes is the regression for the bug this
// verb had: `shaterd nodes` answered `[]` while the subscription cache held
// hundreds of nodes. A NON-EMPTY cache must produce a NON-EMPTY list.
func TestNodesJSONReportsCachedSubscriptionNodes(t *testing.T) {
writeSubCache(t, "all-qomar", []model.Node{
{Name: "NL-vless-1", Enabled: true, URI: vlessURI, FromSub: "all-qomar", Fingerprint: "66b33"},
{Name: "NL-vless-2", Enabled: false, URI: vlessURI, FromSub: "all-qomar"},
})
// The model source stands in for `uci export shater` (absent on a dev host);
// the subscription half is the REAL model.MergeSubCaches path.
restore := readModel
readModel = func() (*model.Model, error) {
m := &model.Model{Globals: model.DefaultGlobals()}
model.MergeSubCaches(m)
return m, nil
}
t.Cleanup(func() { readModel = restore })
b, err := nodesJSON()
if err != nil {
t.Fatalf("nodesJSON: %v", err)
}
var got []nodeView
if err := json.Unmarshal(b, &got); err != nil {
t.Fatalf("nodesJSON produced invalid JSON %q: %v", b, err)
}
if len(got) == 0 {
t.Fatalf("nodes reported an EMPTY list while the subscription cache holds 2 nodes: %s", b)
}
if len(got) != 2 {
t.Fatalf("got %d nodes, want 2: %s", len(got), b)
}
byName := map[string]nodeView{}
for _, v := range got {
byName[v.Name] = v
}
first, ok := byName["NL-vless-1"]
if !ok {
t.Fatalf("cached node NL-vless-1 missing from %s", b)
}
if !first.Enabled {
t.Errorf("NL-vless-1.enabled = false, want true")
}
if first.Sub != "all-qomar" {
t.Errorf("NL-vless-1.sub = %q, want %q", first.Sub, "all-qomar")
}
if first.Protocol != "vless" {
t.Errorf("NL-vless-1.protocol = %q, want %q", first.Protocol, "vless")
}
if first.Server != "example.net" || first.Port != 443 {
t.Errorf("NL-vless-1 server:port = %s:%d, want example.net:443", first.Server, first.Port)
}
if first.ParseError != "" {
t.Errorf("NL-vless-1.parse_error = %q, want empty", first.ParseError)
}
if byName["NL-vless-2"].Enabled {
t.Errorf("NL-vless-2.enabled = true, want false (the model says disabled)")
}
// The credential-bearing share link must not be printed.
if strings.Contains(string(b), "11111111-2222-3333-4444-555555555555") {
t.Errorf("nodes output leaks the share-link uuid: %s", b)
}
}
// TestNodeViewsShape pins the projection: manual vs subscription, an unparsable
// URI still being LISTED (with the reason), and an empty model marshalling to
// `[]` rather than `null`.
func TestNodeViewsShape(t *testing.T) {
m := &model.Model{Nodes: []model.Node{
{Name: "manual-1", Enabled: true, URI: vlessURI, Egress: "wan2"},
{Name: "broken", Enabled: true, URI: "nosuch://whatever", FromSub: "qomar", Stale: true},
{Name: "no-uri", Enabled: false},
}}
got := nodeViews(m)
if len(got) != 3 {
t.Fatalf("nodeViews returned %d views, want 3", len(got))
}
if got[0].Sub != "" {
t.Errorf("manual node sub = %q, want empty", got[0].Sub)
}
if got[0].Egress != "wan2" {
t.Errorf("manual node egress = %q, want wan2", got[0].Egress)
}
if got[1].ParseError == "" {
t.Errorf("unparsable node must carry the reason the engine will skip it")
}
if got[1].Protocol != "nosuch" {
t.Errorf("unparsable node protocol = %q, want the bare scheme %q", got[1].Protocol, "nosuch")
}
if !got[1].Stale {
t.Errorf("stale flag lost")
}
if got[2].Protocol != "" || got[2].ParseError != "" {
t.Errorf("URI-less node: got protocol=%q parse_error=%q, want both empty", got[2].Protocol, got[2].ParseError)
}
b, err := json.Marshal(nodeViews(&model.Model{}))
if err != nil {
t.Fatalf("marshal empty: %v", err)
}
if string(b) != "[]" {
t.Errorf("empty model marshalled to %q, want []", b)
}
}
func TestURIScheme(t *testing.T) {
for _, tc := range []struct{ in, want string }{
{"vless://x@h:443", "vless"},
{"VMESS://payload", "vmess"},
{"", ""},
{"not-a-uri", ""},
{"://leading", ""},
} {
if got := uriScheme(tc.in); got != tc.want {
t.Errorf("uriScheme(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
+63
View File
@@ -0,0 +1,63 @@
package main
// The startup handshake that stops a restarting daemon from standing a new data
// plane up on top of the previous one's teardown.
//
// Deliberately free of build tags: the logic is pure timing and is exercised by
// the tests on any host, while the probe it polls (pidfile + kill(pid, 0)) is
// linux-only and is installed by predecessor_linux.go.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
)
// predecessorBudget / predecessorPoll bound the wait for an outgoing daemon.
//
// The budget must comfortably exceed the slowest honest teardown (engine.Close
// of a box with a few hundred outbounds plus the cache-file flush to flash, then
// the nft/ip/sysctl removal) AND the init script's procd term_timeout, which is
// the hard cap on how long a predecessor can live after its SIGTERM. Overshooting
// costs nothing on a healthy box — the wait ends the instant the predecessor is
// gone — while undershooting reintroduces the very overlap this exists to remove.
//
// Variables, not constants, so the tests can drive the loop without sleeping.
var (
predecessorBudget = 60 * time.Second
predecessorPoll = 100 * time.Millisecond
)
// aliveProbe is the seam waitForPredecessor polls. The portable default reports
// "no predecessor", so a non-linux build (where there is no daemon at all) never
// waits; predecessor_linux.go replaces it with the real pidfile probe.
var aliveProbe = func() (int, bool) { return 0, false }
// waitForPredecessor blocks until no OTHER shaterd owns the pidfile, or until the
// budget runs out.
//
// Returns (pid, true) only when a predecessor is STILL alive once the budget
// expires — the genuine "two daemons" error the caller refuses on. A predecessor
// that exits within the budget (the restart case) returns (0, false) and startup
// continues, now guaranteed to be sequenced after its teardown, exactly as it is
// after a manual `stop` + pause + `start`.
func waitForPredecessor(budget, poll time.Duration, logger log.ContextLogger) (int, bool) {
pid, ok := aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
if logger != nil {
logger.Info("a previous shaterd (pid ", pid, ") is still shutting down — ",
"waiting for its teardown to finish before applying a new data plane")
}
deadline := time.Now().Add(budget)
for time.Now().Before(deadline) {
time.Sleep(poll)
pid, ok = aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
}
return pid, true
}
+9
View File
@@ -0,0 +1,9 @@
//go:build linux
package main
// Bind the portable predecessor wait to the real single-owner probe. Split out
// of main.go so waitForPredecessor itself stays build-tag-free and testable on
// any developer host.
func init() { aliveProbe = daemonAlive }
+106
View File
@@ -0,0 +1,106 @@
package main
// B3 regression, daemon half.
//
// `/etc/init.d/shater restart` is `stop; start`, and procd's `stop` is
// asynchronous: the ubus `service delete` returns as soon as SIGTERM has been
// SENT, so `start` re-adds the instance while the outgoing `shaterd run` is
// still executing its honest teardown. The startup guard used to exit(1) the
// moment it saw a live predecessor and rely on procd's `respawn ... 5 ...` to
// try again later — a blind retry that neither knows nor waits for the teardown
// to finish, and that turns every restart into a logged crash plus a five-second
// hole with no data plane.
//
// The contract these pin: startup BLOCKS until the predecessor is gone (so our
// apply is strictly ordered after its teardown, exactly as it is after a manual
// `stop` + pause + `start`), and only refuses when the predecessor outlives the
// whole budget.
import (
"os"
"testing"
"time"
)
// scriptAlive installs an aliveProbe that reports a live predecessor for the
// first n calls and "gone" afterwards, and restores the real probe.
func scriptAlive(t *testing.T, pid, n int) *int {
t.Helper()
orig := aliveProbe
t.Cleanup(func() { aliveProbe = orig })
calls := 0
aliveProbe = func() (int, bool) {
calls++
if calls <= n {
return pid, true
}
return 0, false
}
return &calls
}
// TestWaitForPredecessorWaitsForTeardown: a predecessor that is still tearing
// the plane down must be WAITED for, not refused. Before the fix this returned
// "still alive" on the first probe and the daemon exited 1.
func TestWaitForPredecessorWaitsForTeardown(t *testing.T) {
calls := scriptAlive(t, 4242, 3)
pid, stillAlive := waitForPredecessor(2*time.Second, time.Millisecond, nil)
if stillAlive {
t.Fatalf("a predecessor that exits within the budget must not be refused (pid %d)", pid)
}
if *calls < 4 {
t.Fatalf("expected the guard to keep probing until the predecessor was gone, got %d probes", *calls)
}
}
// TestWaitForPredecessorRefusesAfterBudget: the single-owner invariant is kept —
// a predecessor that never dies still ends in a refusal, it is just no longer
// the FIRST answer.
func TestWaitForPredecessorRefusesAfterBudget(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return 4242, true }
start := time.Now()
pid, stillAlive := waitForPredecessor(30*time.Millisecond, time.Millisecond, nil)
if !stillAlive || pid != 4242 {
t.Fatalf("an immortal predecessor must still be refused, got pid=%d alive=%v", pid, stillAlive)
}
if time.Since(start) < 30*time.Millisecond {
t.Fatalf("the guard must exhaust its budget before refusing")
}
}
// TestWaitForPredecessorNoPredecessorIsFree: the boot path must pay nothing —
// one probe, no sleep. A wait that cost a tick on every clean start would show up
// as a slower boot for no reason.
func TestWaitForPredecessorNoPredecessorIsFree(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
calls := 0
aliveProbe = func() (int, bool) { calls++; return 0, false }
start := time.Now()
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("no predecessor must not be reported as alive")
}
if calls != 1 {
t.Fatalf("expected exactly one probe when nothing is running, got %d", calls)
}
if time.Since(start) > 500*time.Millisecond {
t.Fatalf("the no-predecessor path must not sleep")
}
}
// TestWaitForPredecessorIgnoresOwnPid: a pidfile naming THIS process (a crashed
// predecessor whose pid we were handed, or a re-exec) is not a predecessor.
func TestWaitForPredecessorIgnoresOwnPid(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return os.Getpid(), true }
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("our own pid must never count as a predecessor")
}
}
+29 -21
View File
@@ -610,44 +610,52 @@ func splitDomainMarker(e string) (marker, value string, ok bool) {
// reader can tell a deliberate difference from an oversight (R9.3). Verified
// against the code on both sides, not against upstream docs — this is a fork.
//
// marker domain lists (this file) routing rules (ruleMatchers, route.go)
// There are now only TWO contexts, not three. A routing rule no longer classifies
// domains at all: `dst_domain` was removed in schema v2 and a rule names its
// destination through a `config ruleset`, so every domain entry in the system —
// filter list, device list, inline rule-set — arrives at classifyDomainEntries
// below. The one caller that adds something on top is inlineRulesetRule
// (ruleset.go), which peels `regexp:` off first.
//
// marker DNS filter / device lists inline rule-sets (inlineRulesetRule)
// ---------- ---------------------------- --------------------------------------
// full: yes — exact domain yes — exact domain
// suffix: yes — domain + subdomains yes — domain + subdomains
// .example SAME AS suffix: (dot stripped) SAME AS suffix: (dot stripped)
// keyword: yes — substring yes — substring
// (bare) SUFFIX in filter/device/ EXACT domain
// inline-ruleset lists
// (bare) SUFFIX (domain + subdomains) SUFFIX (domain + subdomains)
// regexp: NO — warns yes — DomainRegex (pattern validated)
// geosite: NO — warns NO — warns, rule matcher omitted
// geosite: NO — warns NO — warns, entry dropped
//
// Two corrections this table used to get wrong, both of the "claimed a behaviour
// that does not exist" kind:
// The bare-entry row is now the SAME on both sides, and that uniformity is the
// point of schema v2 — a routing rule used to read a bare entry as an EXACT host
// while every list read it as a suffix, a difference nothing in the UI showed.
// model.migrate1to2 is what preserves the old meaning of existing configs: it
// rewrites a rule's bare `dst_domain` entry as `full:` when it moves it into the
// generated rule-set.
//
// - `.example.com` was documented as "subdomains only" on both sides. It is not:
// Corrections this table used to get wrong, all of the "claimed a behaviour that
// does not exist" kind:
//
// - `.example.com` was documented as "subdomains only". It is not:
// classifyDomainEntries strips the dot, so it is a synonym of `suffix:` and
// matches the apex too. See the leading-dot branch above for why the synonym
// is kept rather than the distinction implemented.
// - `geosite:` was documented as a working routing matcher. It is not: the
// route-rule geosite field was REMOVED from this engine, so route.go warns and
// omits the matcher (a rule left with no other matcher is skipped entirely).
// Use a `config ruleset` with source=geosite.
// route-rule geosite field was REMOVED from this engine. Use a `config
// ruleset` with source=geosite and category chips.
//
// The two absences in the domain-list column are DELIBERATE, not gaps:
// The two absences in the FILTER column are DELIBERATE, not gaps:
//
// - regexp: an invalid regular expression is only detected when the rule is
// built, where it aborts box.New and takes the whole config down — the exact
// fail-open violation this audit spent its time removing. Supporting it here
// would require compiling and validating every pattern at generate time.
// Warning is the honest answer until that is done.
// - geosite: filter lists and rulesets already express geosite properly, via
// - regexp: an invalid regular expression aborts box.New and takes the whole
// config down, so it may only be accepted where every pattern is compiled and
// validated at generate time. Inline rule-sets do exactly that
// (peelDomainRegexes, ruleset.go), which is why the right-hand column says
// yes; the DNS-filter path has no such validation and warns instead.
// - geosite: filter lists and rule-sets already express geosite properly, via
// `source=geosite` plus category chips, which fetches the official compiled
// .srs. A `geosite:` entry inside an inline list would be a second, weaker
// path to the same feature.
//
// The BARE-entry difference is also deliberate and long-standing: a blocklist
// entry is meant to cover subdomains, while a routing rule's bare entry is an
// exact host. Both are documented at their call sites.
// toASCIIDomain punycodes a unicode domain entry so it can match the punycoded
// names that actually arrive in DNS queries. Lenient by design: an entry idna
+2 -1
View File
@@ -289,8 +289,9 @@ func TestFailoverWarnsOnceAcrossChainCopy(t *testing.T) {
m := twoNodeGroupModel("failover")
m.Nodes = append(m.Nodes, model.Node{Name: "hop", Enabled: true, URI: ss("203.0.113.9")})
m.Chains = []model.Chain{{Name: "ch", Hops: []string{"node:hop", "group:g"}}}
m.Rulesets = []model.Ruleset{inlineDomainSet("ex", "example.com")}
m.Rules = []model.Rule{
{Name: "r", Enabled: true, DstDomain: []string{"example.com"}, Target: "chain:ch"},
{Name: "r", Enabled: true, DstRuleset: []string{"ex"}, Target: "chain:ch"},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
+43 -20
View File
@@ -67,7 +67,7 @@ package generate
import (
"fmt"
"net/netip"
"sort"
"os"
"strconv"
"strings"
"time"
@@ -75,6 +75,7 @@ import (
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/json/badoption"
"github.com/sagernet/sing-box/shater/logsink"
"github.com/sagernet/sing-box/shater/model"
)
@@ -333,13 +334,31 @@ func Warnings(m *model.Model) []string {
// it, so a level there could never be honoured and would only invite the reader to
// think it was. Note that an EMPTY Level on an ENABLED log means LevelTrace, not
// "default" — so Disabled must never be emitted speculatively.
//
// DisableColor is the OTHER half of the engine's log block, and it is not
// cosmetic. log/log.go turns it into Formatter.DisableColors, which is what
// stops aurora from painting the level word and the per-connection ID. Left at
// its zero value it painted them, and the engine's writer is the daemon's shared
// logsink whose syslog half is procd's stderr — so `logread` filled up with
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: exchange failed …
//
// and `grep ERROR` stopped matching. Colour is now emitted only when that
// destination is a terminal (an interactive `shaterd run`); see
// shater/logsink/color.go.
func logOptions(s string) *option.LogOptions {
if logSilenced(s) {
return &option.LogOptions{Disabled: true}
}
return &option.LogOptions{Level: logLevel(s)}
return &option.LogOptions{Level: logLevel(s), DisableColor: !logColorAllowed()}
}
// logColorAllowed reports whether the engine's log lines may carry ANSI colour.
// A package var so tests pin it instead of depending on how the test binary's
// stderr happens to be wired; production always asks the real stderr, which is
// the sink's syslog half.
var logColorAllowed = func() bool { return logsink.IsTTY(os.Stderr) }
// logSilenced reports whether the operator asked for no log at all.
//
// The vocabulary lives in model.SilentLogLevels rather than here because
@@ -487,16 +506,21 @@ func validPortRange(s string) bool {
// the classifier actually drops.
//
// `suffix:` belongs here. The list used to omit it on the theory that suffix:/
// regexp: are route-rule-only and are peeled off by ruleMatchers before the shared
// classifier sees them. That is true of route.go — which handles a bare `suffix:`
// in its own switch and warns there, so this list cannot double-report — but NOT
// of the classifier itself: classifyDomainEntries has a `case "suffix"`, and every
// other caller (devices.go, dnsfilter.go) reaches it directly. A lone `suffix:`
// arriving that way was dropped by add()'s empty-value guard and reported by
// nobody, which is the one outcome this pair of functions exists to prevent.
// regexp: were route-rule-only and were peeled off before the shared classifier
// saw them. classifyDomainEntries has a `case "suffix"`, and every caller
// (devices.go, dnsfilter.go, inlineRulesetRule) reaches it directly, so a lone
// `suffix:` was dropped by add()'s empty-value guard and reported by nobody —
// the one outcome this pair of functions exists to prevent. (The route-rule half
// of that old reasoning is gone entirely: `dst_domain` was removed in schema v2,
// so route.go classifies no domains at all and there is no second warner to
// double-report with.)
//
// `regexp:` is correctly absent: the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix and is reported by unrecognisedDomainPrefix.
// `regexp:` is correctly absent, for two different reasons depending on the
// caller. In a filter/device list the classifier has no case for it, so a bare
// `regexp:` is an UNRECOGNISED prefix reported by unrecognisedDomainPrefix. In an
// inline rule-set peelDomainRegexes (ruleset.go) strips every `regexp:` entry
// BEFORE this predicate runs and reports a valueless one itself, so it can never
// reach here either way.
var domainMarkers = []string{"keyword:", "full:", "suffix:", "."}
// isDomainMarkerOnly reports whether an entry is a bare classification marker
@@ -564,15 +588,14 @@ func networkList(tcp, udp bool) option.NetworkList {
}
}
// sortedByOrder returns rule indices sorted by (Order, original index) so the
// sortedRuleIndices returns rule indices sorted by (Order, original index) so the
// engine's first-match order matches the model's declared order, stably.
//
// The implementation is model.SortedRuleIndices: the reachability analysis that
// decides which rules can never fire has to walk the rules in EXACTLY this order
// to be right, and it lives in model (the leaf both generate and the panel API
// import). Two copies of "what order do rules run in" is precisely the drift that
// would make the warning and the panel badge disagree.
func sortedRuleIndices(rules []model.Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
return model.SortedRuleIndices(rules)
}
+4 -2
View File
@@ -323,8 +323,9 @@ func TestEgressDPISpoofValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12366, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "spf", Type: "direct", DPI: "spoof"}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "spoof-rule", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:spf"},
{Name: "spoof-rule", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:spf"},
},
}
opts, warns, changed := applyAndClose(t, m)
@@ -346,8 +347,9 @@ func TestByedpiEgressValidates(t *testing.T) {
Globals: model.DefaultGlobals(),
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12367, TCP: true, UDP: true}},
Egresses: []model.Egress{{Name: "bd", Type: "byedpi", Port: 1080}},
Rulesets: []model.Ruleset{inlineDomainSet("blocked", "blocked.example")},
Rules: []model.Rule{
{Name: "desync", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "egress:bd"},
{Name: "desync", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "egress:bd"},
},
}
opts, warns, changed := applyAndClose(t, m)
+73
View File
@@ -0,0 +1,73 @@
package generate
import (
"bytes"
"context"
"strings"
"testing"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// TestLogOptionsDisableColorOffTTY pins the engine half of the syslog colour
// leak. The engine's log factory is built by box.New from these options; with
// DisableColor left false, every engine line reached procd's stderr — i.e.
// syslog — wrapped in aurora escapes.
func TestLogOptionsDisableColorOffTTY(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false } // stderr is procd's pipe
t.Cleanup(func() { logColorAllowed = restore })
for _, level := range []string{"", "info", "error", "warning"} {
got := logOptions(level)
if got.Disabled {
t.Fatalf("logOptions(%q) unexpectedly disabled the log", level)
}
if !got.DisableColor {
t.Errorf("logOptions(%q).DisableColor = false; syslog would get ANSI escapes", level)
}
}
// A terminal (an interactive `shaterd run`) keeps its colours.
logColorAllowed = func() bool { return true }
if logOptions("info").DisableColor {
t.Errorf("logOptions on a TTY disabled colour; the gate is supposed to be the destination, not a blanket off")
}
}
// TestGeneratedLogBlockProducesNoANSI is the end-to-end proof: take the log
// block Generate actually emits, build the engine's factory over it exactly the
// way box.New does (log.New with DefaultWriter), and check the bytes.
func TestGeneratedLogBlockProducesNoANSI(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false }
t.Cleanup(func() { logColorAllowed = restore })
m := &model.Model{Globals: model.DefaultGlobals()}
m.Globals.LogLevel = "info"
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("GenerateWithWarnings: %v", err)
}
if opts.Log == nil {
t.Fatal("generated options carry no log block")
}
var out bytes.Buffer
factory, err := log.New(log.Options{Options: option.LogOptions(*opts.Log), DefaultWriter: &out})
if err != nil {
t.Fatalf("log.New over the generated block: %v", err)
}
ctx := log.ContextWithNewID(context.Background())
factory.Logger().ErrorContext(ctx, "dns: exchange failed for example.com. IN AAAA: unexpected EOF")
got := out.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("engine factory wrote nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("engine log line carries an ANSI escape at byte %d: %q", i, got)
}
}
+3 -2
View File
@@ -21,9 +21,10 @@ func TestProfileAppliesCleanly(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12370, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{inlineDomainSet("ads", "ads.example")},
Rules: []model.Rule{
{Name: "lan-proxy", Enabled: true, Order: 10, Src: []string{"192.168.1.0/24"}, Target: "node:ss1"},
{Name: "adblock", Enabled: true, Order: 20, DstDomain: []string{"ads.example"}, Target: "block"},
{Name: "adblock", Enabled: true, Order: 20, DstRuleset: []string{"ads"}, Target: "block"},
},
Profiles: []model.Profile{
{Name: "home", Enabled: true, Priority: 1, DisableRules: []string{"adblock"}},
@@ -35,7 +36,7 @@ func TestProfileAppliesCleanly(t *testing.T) {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
// Profile-disabled 'adblock' rule must be absent.
if opts.Route == nil || hasDomainRule(opts.Route, "ads.example") {
if opts.Route == nil || hasRulesetRule(opts.Route, "ads") {
t.Fatalf("profile-disabled rule 'adblock' should not be emitted")
}
// The lan-proxy rule still routes to ss1.
+41 -23
View File
@@ -49,9 +49,13 @@ func TestNoProfilesUnchanged(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(), // kill-switch closed => Final "block"
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("a", "a.example"),
inlineDomainSet("b", "b.example"),
},
Rules: []model.Rule{
{Name: "a", Enabled: true, Order: 10, DstDomain: []string{"a.example"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstDomain: []string{"b.example"}, Target: "node:ss1"},
{Name: "a", Enabled: true, Order: 10, DstRuleset: []string{"a"}, Target: "direct"},
{Name: "b", Enabled: true, Order: 20, DstRuleset: []string{"b"}, Target: "node:ss1"},
},
}
rt, b := buildRouteAt(m, wed12UTC)
@@ -70,7 +74,7 @@ func TestNoProfilesUnchanged(t *testing.T) {
if len(rt.Rules) != 4 {
t.Fatalf("route rule count = %d, want 4 (sniff+hijack+2 user)", len(rt.Rules))
}
if !hasDomainRule(rt, "a.example") || !hasDomainRule(rt, "b.example") {
if !hasRulesetRule(rt, "a") || !hasRulesetRule(rt, "b") {
t.Fatalf("both user rules should survive unchanged")
}
}
@@ -83,9 +87,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "ss1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.5:8388#ss1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("blocked", "blocked.example"),
inlineDomainSet("keep", "keep.example"),
},
Rules: []model.Rule{
{Name: "blockme", Enabled: true, Order: 10, DstDomain: []string{"blocked.example"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstDomain: []string{"keep.example"}, Target: "direct"},
{Name: "blockme", Enabled: true, Order: 10, DstRuleset: []string{"blocked"}, Target: "block"},
{Name: "keep", Enabled: true, Order: 20, DstRuleset: []string{"keep"}, Target: "direct"},
},
Profiles: []model.Profile{
// Manual pin: honored regardless of conditions. Disables "blockme".
@@ -94,10 +102,10 @@ func TestActiveProfileDisablesRule(t *testing.T) {
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "blocked.example") {
if hasRulesetRule(rt, "blocked") {
t.Fatalf("profile disabled rule 'blockme' but its matcher is still emitted")
}
if !hasDomainRule(rt, "keep.example") {
if !hasRulesetRule(rt, "keep") {
t.Fatalf("rule 'keep' should be untouched")
}
if rt.Final != tagBlock {
@@ -115,9 +123,13 @@ func TestActiveProfileDisablesRule(t *testing.T) {
func TestAutoSelectHighestPriority(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
inlineDomainSet("lo", "lo.example"),
inlineDomainSet("hi", "hi.example"),
},
Rules: []model.Rule{
{Name: "rLo", Enabled: true, Order: 10, DstDomain: []string{"lo.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rLo", Enabled: true, Order: 10, DstRuleset: []string{"lo"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 20, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "lo", Enabled: true, Priority: 5, DisableRules: []string{"rLo"}},
@@ -125,10 +137,10 @@ func TestAutoSelectHighestPriority(t *testing.T) {
},
}
rt, _ := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("the highest-priority profile must win and disable rHi")
}
if !hasDomainRule(rt, "lo.example") {
if !hasRulesetRule(rt, "lo") {
t.Fatalf("only the winning profile applies (rLo should survive)")
}
}
@@ -142,9 +154,13 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
g.ActiveProfile = active
return &model.Model{
Globals: g,
Rulesets: []model.Ruleset{
inlineDomainSet("hi", "hi.example"),
inlineDomainSet("if", "if.example"),
},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstDomain: []string{"if.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
{Name: "rIf", Enabled: true, Order: 20, DstRuleset: []string{"if"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
@@ -156,19 +172,19 @@ func TestIfaceProfileSkippedByAutoButHonoredNamed(t *testing.T) {
// Auto-select (no pin): iface profile skipped (the watcher owns it); 'plain' wins.
rt, _ := buildRouteAt(mk(""), wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("auto-select should apply 'plain' and disable rHi")
}
if !hasDomainRule(rt, "if.example") {
if !hasRulesetRule(rt, "if") {
t.Fatalf("iface profile 'wwan' must be skipped by auto-select (rIf should survive)")
}
// Named explicitly: iface profile honored regardless of its iface condition.
rt, _ = buildRouteAt(mk("wwan"), wed12UTC)
if hasDomainRule(rt, "if.example") {
if hasRulesetRule(rt, "if") {
t.Fatalf("explicitly named iface profile must be honored and disable rIf")
}
if !hasDomainRule(rt, "hi.example") {
if !hasRulesetRule(rt, "hi") {
t.Fatalf("only the named profile applies; rHi should survive")
}
}
@@ -179,14 +195,15 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
g := model.DefaultGlobals()
g.ActiveProfile = "ghost"
m := &model.Model{
Globals: g,
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstDomain: []string{"x.example"}, Target: "direct"}},
Globals: g,
Rulesets: []model.Ruleset{inlineDomainSet("x", "x.example")},
Rules: []model.Rule{{Name: "rX", Enabled: true, Order: 10, DstRuleset: []string{"x"}, Target: "direct"}},
Profiles: []model.Profile{
{Name: "auto", Enabled: true, Priority: 1, DisableRules: []string{"rX"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "x.example") {
if hasRulesetRule(rt, "x") {
t.Fatalf("fallback auto-select should apply 'auto' and disable rX")
}
if !hasWarning(b, "falling back to auto-select") {
@@ -208,16 +225,17 @@ func TestUnknownActiveProfileFallsBackToAuto(t *testing.T) {
// suppress it or to warn about it.
func TestPlainProfileIsSelectableAfterProbeRemoval(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{inlineDomainSet("hi", "hi.example")},
Rules: []model.Rule{
{Name: "rHi", Enabled: true, Order: 10, DstDomain: []string{"hi.example"}, Target: "direct"},
{Name: "rHi", Enabled: true, Order: 10, DstRuleset: []string{"hi"}, Target: "direct"},
},
Profiles: []model.Profile{
{Name: "plain", Enabled: true, Priority: 10, DisableRules: []string{"rHi"}},
},
}
rt, b := buildRouteAt(m, wed12UTC)
if hasDomainRule(rt, "hi.example") {
if hasRulesetRule(rt, "hi") {
t.Fatalf("a plain profile must apply and disable rHi")
}
// No leftover diagnostic: warning about an option the parser no longer reads
+63 -126
View File
@@ -1,8 +1,6 @@
package generate
import (
"fmt"
"regexp"
"sort"
"strings"
@@ -30,6 +28,12 @@ func (b *builder) buildRoute() *option.RouteOptions {
// never aborts box.New.
b.applyProfiles()
// Report the rules that cannot fire under this config BEFORE building anything,
// so the diagnosis is about the desired state the operator wrote rather than
// about whatever survived generation. Diagnosis only — nothing is renamed,
// reordered or dropped here.
b.warnUnreachableRules()
// Kill-switch default backstop.
final := tagBlock
if strings.EqualFold(strings.TrimSpace(b.m.Globals.KillSwitch), "open") {
@@ -240,25 +244,58 @@ func (b *builder) ruleKillFallback(r model.Rule, want string) string {
// outbound — traffic left over the DEFAULT WAN while the panel (which applies the
// same "a bare Egress is a target too" rule for display) showed `egress:wan2`.
// A silent mis-route on exactly the multi-WAN setups the field exists for.
func effectiveRuleTarget(r model.Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
//
// The implementation is model.EffectiveRuleTarget — shared with the reachability
// analysis, which must resolve a rule's target identically to decide which
// default the router actually ends up using.
func effectiveRuleTarget(r model.Rule) string { return model.EffectiveRuleTarget(r) }
// isCatchAll reports whether a rule carries NO matcher of any kind (src, dst
// domain/ip/ruleset, port, proto). Such a rule is the default egress.
func (b *builder) isCatchAll(r model.Rule) bool {
return len(r.Src) == 0 &&
len(r.DstDomain) == 0 &&
len(r.DstIP) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
//
// The implementation is model.IsCatchAll — shared with the reachability analysis
// (and mirrored by the panel), because "is this rule a default?" is the single
// question both the Final assignment below and the never-fires badge turn on.
func (b *builder) isCatchAll(r model.Rule) bool { return model.IsCatchAll(r) }
// warnUnreachableRules reports every rule that CANNOT take effect under this
// config, whatever the traffic.
//
// Only one shape is certain enough to report (see model.RuleReachability): two
// condition-less rules, where the later one wins because the loop above simply
// overwrites Final. That case is silent today and looks completely healthy — the
// panel drew both as "default route · final" and the log said nothing — so a
// config with `default -> direct` at order 20 and `default -> group:auto` at
// order 100 gave no hint at all that one of the two was doing nothing.
//
// Severity is decided by CONSEQUENCE, via the wording (apply/warnings.go
// classifies generate's text): when the default that actually wins is `direct`
// while the retired rule asked for a tunnel or a block, the operator's default
// policy is not in effect and everything unmatched leaves on the plain WAN — a
// broken protection claim, i.e. critical. Any other combination (a dead `direct`
// under a live tunnel, one tunnel under another) is a dead setting, not a leak:
// a warning.
func (b *builder) warnUnreachableRules() {
for _, rr := range model.RuleReachability(b.effectiveRules) {
if !rr.Unreachable {
continue
}
dead := effectiveRuleTarget(b.effectiveRules[rr.Index])
winner := effectiveRuleTarget(b.effectiveRules[rr.ShadowedByIndex])
if winner == tagDirect && dead != tagDirect {
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so %q wins and this rule's "+
"target %q is never applied. Everything no other rule matches leaves over the plain "+
"WAN with your real IP address. Delete one of the two, or give this one a condition",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
continue
}
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so the default the router uses "+
"is %q and this rule's target %q is never applied. Delete one of the two, or give this "+
"one a condition so it can match something",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
}
}
// ruleMatchers builds the RawDefaultRule matchers the engine can evaluate.
@@ -282,114 +319,14 @@ func (b *builder) ruleMatchers(r model.Rule) (raw option.RawDefaultRule, matched
matched = true
}
// Destination domains. The route-rule-specific prefixes (geosite:/regexp:/
// suffix:) are peeled off here; everything else goes through the shared
// classifier in dnsfilter.go with bareIsSuffix=FALSE — a bare entry in a
// ROUTING rule is an exact domain, unlike the DNS/filter lists where it means
// "and all subdomains". classifyDomainEntries is also what drops marker-only
// entries ("." / "full:" / "keyword:"), which must never reach the engine:
// an empty domain/domain_suffix aborts box.New, and an empty domain_keyword
// is strings.Contains(host, "") — a silent match on EVERY host.
var explicitSuffix []string
var plain []string
for _, d := range r.DstDomain {
d = strings.TrimSpace(d)
if d == "" {
continue
}
switch {
case strings.HasPrefix(d, "geosite:"):
// A `geosite:` matcher is NOT emittable: the route-rule geosite field was
// removed in this engine and route/rule.NewDefaultRule hard-errors on it,
// which aborts box.New for the WHOLE config. Mirroring the geoip handling
// below, it is treated as inert: warned and omitted, so a legacy/UCI rule
// carrying one degrades to "this rule does nothing" instead of taking the
// entire tunnel down. Use a `config ruleset` with source=geosite instead.
b.warnf("rule %q: geosite matcher %q is removed from this engine — use a ruleset with source=geosite instead, omitted (inert)", r.Name, d)
case strings.HasPrefix(d, "regexp:"):
// route/rule.NewDomainRegexItem hard-errors on an uncompilable pattern and
// takes box.New with it; validate here and drop the bad one with a warning.
// A BARE "regexp:" compiles fine but matches every host — same silent
// match-all hazard as an empty keyword, so it is dropped too.
re := strings.TrimSpace(strings.TrimPrefix(d, "regexp:"))
if re == "" {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", r.Name, d)
break
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("rule %q: domain regexp %q is invalid (%v), omitted", r.Name, re, err)
break
}
raw.DomainRegex = append(raw.DomainRegex, re)
matched = true
case strings.HasPrefix(d, "suffix:"):
// Label-aware suffix (apex + subdomains): sing-box domain_suffix stored
// in bare form matches both "example.com" and "*.example.com" (see
// sing/common/domain matcher). A leading-dot entry, by contrast, matches
// subdomains ONLY, so the explicit `suffix:` form is how presets/rules
// ask for apex-inclusive suffix matching.
if s := strings.TrimSpace(strings.TrimPrefix(d, "suffix:")); s != "" {
explicitSuffix = append(explicitSuffix, s)
} else {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New)", r.Name, d)
}
default:
plain = append(plain, d)
}
}
domain, suffix, keyword := classifyDomainEntries(plain, false)
for _, d := range plain {
if isDomainMarkerOnly(d) {
b.warnf("rule %q: %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword matches EVERY host)", r.Name, d)
}
}
// An unrecognised `word:` prefix is DROPPED by classifyDomainEntries (a domain
// cannot contain ":", so keeping it would load a provably unmatchable literal).
// It has to be reported here or the rule silently loses that destination — the
// v0.1/xray spelling `domain:example.com` is exactly what someone migrating
// writes, and it used to disappear without a trace. geosite:/regexp:/suffix:
// were already peeled off above, so `plain` carries only the shared markers.
b.warnUnrecognisedPrefixes(fmt.Sprintf("rule %q", r.Name), plain)
suffix = append(explicitSuffix, suffix...)
if len(domain) > 0 {
raw.Domain = badoption.Listable[string](domain)
matched = true
}
if len(suffix) > 0 {
raw.DomainSuffix = badoption.Listable[string](suffix)
matched = true
}
if len(keyword) > 0 {
raw.DomainKeyword = badoption.Listable[string](keyword)
matched = true
}
// Destination IPs. A `geoip:<code>` entry is NOT a CIDR: route-rule geoip was
// removed in this engine (box.New hard-errors on it), so — mirroring how the
// ruleset geoip source is handled — it is treated as inert: warned and omitted
// (never emitted as an ip_cidr, which would also abort box.New). This keeps a
// geoip-driven rule/preset (e.g. the ru-bypass pack) fail-open instead of fatal.
var ipcidr []string
for _, ip := range r.DstIP {
ip = strings.TrimSpace(ip)
if ip == "" {
continue
}
if strings.HasPrefix(strings.ToLower(ip), "geoip:") {
b.warnf("rule %q: geoip matcher %q is removed from this engine — use a ruleset with source=geoip instead, omitted (inert)", r.Name, ip)
continue
}
ipcidr = append(ipcidr, ip)
}
// Same fail-open validation as the source list: an unparseable ip_cidr aborts
// box.New for the whole config, so drop it with a warning instead.
ipcidr, badIP := validPrefixes(ipcidr)
for _, s := range badIP {
b.warnf("rule %q: destination %q is not a valid IP/CIDR, omitted", r.Name, s)
}
if len(ipcidr) > 0 {
raw.IPCIDR = badoption.Listable[string](ipcidr)
matched = true
}
// Destination: a rule's ONLY destination matcher is DstRuleset (schema v2).
// The inline `dst_domain` / `dst_ip` lists that used to be classified here are
// gone. A destination list is a `config ruleset` — compiled once into a .srs and
// shared by every rule that references it — so the prefix vocabulary
// (full:/suffix:/keyword:/regexp:, a leading dot) and the geosite/geoip sources
// live in exactly one place (generate/ruleset.go inlineRulesetRule). Existing
// configs were rewritten by model.migrate1to2, which preserves each entry's
// meaning 1:1.
// DstRuleset -> reference the rs-<name> rule-sets materialised by
// buildRoutingRuleSets. A dst_ruleset naming an UNDEFINED ruleset is warned +
+88 -69
View File
@@ -22,26 +22,27 @@ func killModel(kill string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"n1"}}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "social", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "group:grp", Kill: kill},
{Name: "social", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "group:grp", Kill: kill},
{Name: "catch-tcp", Enabled: true, Order: 20, Proto: "tcp", Target: "direct"},
},
}
}
// domainRuleTarget returns the outbound the rule matching domain routes to.
func domainRuleTarget(rt *option.RouteOptions, domain string) (string, bool) {
for _, r := range generalRules(rt) {
for _, d := range r.DefaultOptions.RawDefaultRule.Domain {
if d == domain {
return r.DefaultOptions.RuleAction.RouteOptions.Outbound, true
}
}
// rulesetRuleTarget returns the outbound the rule referencing the rule-set named
// name routes to. It is the post-schema-v2 replacement for looking a rule up by
// its dst-domain matcher: a rule's destination is a rule_set reference now, so
// the matcher that identifies it is the rs-<name> tag.
func rulesetRuleTarget(rt *option.RouteOptions, name string) (string, bool) {
dr := findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name)
if dr == nil {
return "", false
}
return "", false
return dr.RuleAction.RouteOptions.Outbound, true
}
// TestRuleKillDefaultBlocks: kill="" / "default" is fail-closed — the rule is
@@ -54,7 +55,7 @@ func TestRuleKillDefaultBlocks(t *testing.T) {
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("%q: rule must be emitted routing to block, but it was dropped", kill)
}
@@ -76,7 +77,7 @@ func TestRuleKillClosedBlocksHere(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=closed must still emit the rule (warnings %v)", warns)
}
@@ -103,7 +104,7 @@ func TestRuleKillOpenGoesDirect(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok {
t.Fatalf("kill=open must still emit the rule (warnings %v)", warns)
}
@@ -122,7 +123,7 @@ func TestRuleKillUnknownBlocks(t *testing.T) {
if err != nil {
t.Fatalf("Generate: %v", err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != tagBlock {
t.Fatalf("unknown kill policy must block, got %q (ok=%v)", got, ok)
}
@@ -145,7 +146,7 @@ func TestRuleKillPreservesFailClosedInvariant(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("%q: Final = %q, want block", kill, opts.Route.Final)
}
if tgt, ok := domainRuleTarget(opts.Route, "social.example"); ok && tgt == tagDirect {
if tgt, ok := rulesetRuleTarget(opts.Route, "social"); ok && tgt == tagDirect {
t.Fatalf("%q: dead group leaked direct", kill)
}
}
@@ -158,17 +159,18 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{inlineDomainSet("social", "social.example")},
Rules: []model.Rule{
{Name: "ok", Enabled: true, Order: 10, DstDomain: []string{"social.example"}, Target: "node:n1", Kill: kill},
{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"social"}, Target: "node:n1", Kill: kill},
},
}
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("%q: Generate: %v", kill, err)
}
got, ok := domainRuleTarget(opts.Route, "social.example")
got, ok := rulesetRuleTarget(opts.Route, "social")
if !ok || got != "n1" {
t.Fatalf("%q: healthy target must win, got %q (ok=%v)", kill, got, ok)
}
@@ -176,20 +178,27 @@ func TestRuleKillOnlyAppliesToUnresolvedTargets(t *testing.T) {
}
// --- R4: marker-only domain entries ------------------------------------------
//
// R4 moved with the destination list itself: a rule's domains are an inline
// `config ruleset` now (schema v2), so the marker-only guard has to hold in
// inlineRulesetRule rather than in ruleMatchers. The hazard is unchanged.
// TestRuleMarkerOnlyDomainEntriesDropped: an entry that is nothing but its marker
// must never reach the engine. An empty domain/domain_suffix makes NewDomainItem
// return "empty item is not allowed" and aborts box.New (whole LAN down from one
// stray "."); an empty domain_keyword is SILENT and matches every host.
func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
// TestRulesetMarkerOnlyDomainEntriesDropped: an entry that is nothing but its
// marker must never reach the engine. An empty domain/domain_suffix makes
// NewDomainItem return "empty item is not allowed" and aborts box.New (whole LAN
// down from one stray "."); an empty domain_keyword is SILENT and matches every
// host. As the sole entry it also leaves the rule-set with nothing usable, so the
// list is skipped and the rule referencing it is not emitted — never promoted to
// "matches everything".
func TestRulesetMarkerOnlyDomainEntriesDropped(t *testing.T) {
for _, entry := range []string{".", "full:", "keyword:", "suffix:", "regexp:", " keyword: "} {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{entry}, Target: "node:n1",
})
for _, r := range rt.Rules {
raw := r.DefaultOptions.RawDefaultRule
for _, list := range [][]string{raw.Domain, raw.DomainSuffix, raw.DomainKeyword, raw.DomainRegex} {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m", entry)},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
for _, rs := range rt.RuleSet {
hr := rs.InlineOptions.Rules[0].DefaultOptions
for _, list := range [][]string{hr.Domain, hr.DomainSuffix, hr.DomainKeyword, hr.DomainRegex} {
for _, v := range list {
if strings.TrimSpace(v) == "" {
t.Fatalf("%q: emitted an EMPTY matcher token", entry)
@@ -197,8 +206,9 @@ func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
}
}
}
// Sole entry => nothing left to match on => the rule must be skipped, never
// silently promoted to "matches everything".
if _, ok := ruleSetByTag(rt, "rs-m"); ok {
t.Fatalf("%q: a marker-only list must not materialise a rule-set", entry)
}
if n := len(generalRules(rt)); n != 0 {
t.Fatalf("%q: marker-only sole entry must skip the rule, got %d rules", entry, n)
}
@@ -208,46 +218,55 @@ func TestRuleMarkerOnlyDomainEntriesDropped(t *testing.T) {
}
}
// TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must survive
// the marker-only ones.
func TestRuleMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "m", Enabled: true, Order: 10,
DstDomain: []string{".", "keyword:", "exact.example", "keyword:ads", ".sub.example", "suffix:apex.example"},
Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d (warnings %v)", len(gen), warns)
// TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest: the real entries must
// survive the marker-only ones.
func TestRulesetMarkerOnlyBesideRealEntriesKeepsTheRest(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("m",
".", "keyword:", "full:exact.example", "keyword:ads", ".sub.example", "suffix:apex.example")},
model.Rule{Name: "m", Enabled: true, Order: 10, DstRuleset: []string{"m"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-m")
if !ok {
t.Fatalf("the usable entries must keep the list alive (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example] (bare entry in a ROUTING rule is exact)", raw.Domain)
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("domain = %v, want [exact.example]", hr.Domain)
}
if len(raw.DomainKeyword) != 1 || raw.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", raw.DomainKeyword)
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "ads" {
t.Fatalf("keyword = %v, want [ads]", hr.DomainKeyword)
}
if len(raw.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", raw.DomainSuffix)
if len(hr.DomainSuffix) != 2 {
t.Fatalf("suffix = %v, want both apex.example and sub.example", hr.DomainSuffix)
}
if findRouteRuleWithRuleSet(rt, "rs-m") == nil {
t.Fatalf("the rule must be emitted referencing rs-m; rules=%+v", rt.Rules)
}
}
// TestRuleBareEntryIsExactDomain pins the routing convention (bare == exact),
// which deliberately differs from the DNS/filter lists (bare == suffix).
func TestRuleBareEntryIsExactDomain(t *testing.T) {
rt, _ := genRules(t, model.Rule{
Name: "b", Enabled: true, Order: 10,
DstDomain: []string{"example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 rule, got %d", len(gen))
// TestRulesetBareEntryIsDomainSuffix pins the convention a destination list now
// follows — and it is the OPPOSITE of the one the old dst_domain used.
//
// A bare entry in a routing rule meant one EXACT domain; a bare entry in a
// rule-set means the domain AND its subdomains (the DNS/filter/device convention,
// classifyDomainEntries with bareIsSuffix=true). `full:` is how an exact match is
// written now. model.migrate1to2 rewrites old dst_domain entries accordingly, so
// this asymmetry is the thing that migration has to get right.
func TestRulesetBareEntryIsDomainSuffix(t *testing.T) {
rt, _ := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("b", "example.com", "full:exact.example")},
model.Rule{Name: "b", Enabled: true, Order: 10, DstRuleset: []string{"b"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-b")
if !ok {
t.Fatalf("rs-b not emitted; route=%+v", rt)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("bare entry must be an exact Domain, got domain=%v suffix=%v", raw.Domain, raw.DomainSuffix)
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.DomainSuffix) != 1 || hr.DomainSuffix[0] != "example.com" {
t.Fatalf("bare entry must become a domain_suffix, got suffix=%v domain=%v", hr.DomainSuffix, hr.Domain)
}
if len(raw.DomainSuffix) != 0 {
t.Fatalf("bare entry must NOT become a suffix, got %v", raw.DomainSuffix)
if len(hr.Domain) != 1 || hr.Domain[0] != "exact.example" {
t.Fatalf("full: must be the exact form, got domain=%v", hr.Domain)
}
}
+35 -14
View File
@@ -18,10 +18,13 @@ import (
)
// TestMalformedMatchersStillApply drives ONE model carrying every previously
// fatal matcher — a geosite: domain, a bad ip_cidr, a bad source cidr, an
// fatal matcher — a geosite: entry, a bad ip_cidr, a bad source cidr, an
// uncompilable regexp and a malformed port range — plus a healthy rule, through
// engine.Apply. It must come up, the healthy rule must survive, and the
// kill-switch backstop must stay closed.
// kill-switch backstop must stay closed. The destination lists are inline
// `config ruleset`s (schema v2), which is where the domain/IP guards live now;
// a rule-set left with no usable entry is skipped, and so is the rule whose only
// matcher it was.
func TestMalformedMatchersStillApply(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
@@ -29,13 +32,19 @@ func TestMalformedMatchersStillApply(t *testing.T) {
Globals: g,
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12395, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Rulesets: []model.Ruleset{
inlineDomainSet("geosite", "geosite:youtube"),
inlineIPSet("badip", "999.1.1.1/24"),
inlineDomainSet("badre", "regexp:*broken("),
inlineDomainSet("ok", "ok.example"),
},
Rules: []model.Rule{
{Name: "geosite", Enabled: true, Order: 1, DstDomain: []string{"geosite:youtube"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstIP: []string{"999.1.1.1/24"}, Target: "node:n1"},
{Name: "geosite", Enabled: true, Order: 1, DstRuleset: []string{"geosite"}, Target: "node:n1"},
{Name: "badip", Enabled: true, Order: 2, DstRuleset: []string{"badip"}, Target: "node:n1"},
{Name: "badsrc", Enabled: true, Order: 3, Src: []string{"192.168.0.0/99"}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstDomain: []string{"regexp:*broken("}, Target: "node:n1"},
{Name: "badre", Enabled: true, Order: 4, DstRuleset: []string{"badre"}, Target: "node:n1"},
{Name: "badport", Enabled: true, Order: 5, DstPort: "a-b", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstDomain: []string{"ok.example"}, DstPort: "443", Target: "node:n1"},
{Name: "healthy", Enabled: true, Order: 6, DstRuleset: []string{"ok"}, DstPort: "443", Target: "node:n1"},
},
}
@@ -46,14 +55,21 @@ func TestMalformedMatchersStillApply(t *testing.T) {
if opts.Route.Final != tagBlock {
t.Fatalf("Final = %q, want block", opts.Route.Final)
}
if !hasDomainRule(opts.Route, "ok.example") {
if !hasRulesetRule(opts.Route, "ok") {
t.Fatalf("the healthy rule must survive alongside the malformed ones")
}
if !hasRouteToOutbound(opts, "n1") {
t.Fatalf("expected a route rule to n1")
}
// Every malformed matcher reported itself rather than failing silently.
for _, want := range []string{"geosite matcher", "is not a valid IP/CIDR", "domain regexp", "is not a valid port/range"} {
for _, want := range []string{
"unrecognised prefix", // geosite: in a domain list
"no usable entries", // ...leaving that list empty
"bad ip_cidr entry", // 999.1.1.1/24
"is not a valid IP/CIDR", // the source cidr
"domain regexp", // regexp:*broken(
"is not a valid port/range", // a-b
} {
if !routeWarnsHave(warns, want) {
t.Fatalf("missing diagnostic %q in %v", want, warns)
}
@@ -155,10 +171,15 @@ func TestRuleKillPoliciesApply(t *testing.T) {
Inbounds: []model.Inbound{{Name: "lan", Enabled: true, Type: "tproxy", TproxyPort: 12398, TCP: true, UDP: true}},
Nodes: []model.Node{{Name: "dead", Enabled: false, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#dead"}},
Groups: []model.Group{{Name: "grp", Strategy: "leastping", Nodes: []string{"dead"}}},
Rulesets: []model.Ruleset{
inlineDomainSet("k-closed", "closed.example"),
inlineDomainSet("k-open", "open.example"),
inlineDomainSet("k-default", "default.example"),
},
Rules: []model.Rule{
{Name: "k-closed", Enabled: true, Order: 1, DstDomain: []string{"closed.example"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstDomain: []string{"open.example"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstDomain: []string{"default.example"}, Target: "group:grp"},
{Name: "k-closed", Enabled: true, Order: 1, DstRuleset: []string{"k-closed"}, Target: "group:grp", Kill: "closed"},
{Name: "k-open", Enabled: true, Order: 2, DstRuleset: []string{"k-open"}, Target: "group:grp", Kill: "open"},
{Name: "k-default", Enabled: true, Order: 3, DstRuleset: []string{"k-default"}, Target: "group:grp"},
},
}
@@ -166,13 +187,13 @@ func TestRuleKillPoliciesApply(t *testing.T) {
if !changed {
t.Fatalf("expected Apply changed==true (warnings: %v)", warns)
}
if got, ok := domainRuleTarget(opts.Route, "closed.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-closed"); !ok || got != tagBlock {
t.Fatalf("kill=closed must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "open.example"); !ok || got != tagDirect {
if got, ok := rulesetRuleTarget(opts.Route, "k-open"); !ok || got != tagDirect {
t.Fatalf("kill=open must emit a rule routed to direct, got %q (ok=%v)", got, ok)
}
if got, ok := domainRuleTarget(opts.Route, "default.example"); !ok || got != tagBlock {
if got, ok := rulesetRuleTarget(opts.Route, "k-default"); !ok || got != tagBlock {
t.Fatalf("kill=default must emit a rule routed to block, got %q (ok=%v)", got, ok)
}
if opts.Route.Final != tagBlock {
+209 -85
View File
@@ -45,6 +45,40 @@ func genRules(t *testing.T, rules ...model.Rule) (*option.RouteOptions, []string
return opts.Route, warns
}
// ruleModelWithSets is ruleModel plus the `config ruleset` definitions the rules'
// DstRuleset entries point at. A rule's ONLY destination matcher is a rule-set
// (schema v2), so every case below that just needs "some destination the engine
// can match on" declares one here rather than writing an inline domain/IP list.
func ruleModelWithSets(sets []model.Ruleset, rules ...model.Rule) *model.Model {
m := ruleModel(rules...)
m.Rulesets = sets
return m
}
// genRulesWithSets is genRules for a model that also carries rule-sets.
func genRulesWithSets(t *testing.T, sets []model.Ruleset, rules ...model.Rule) (*option.RouteOptions, []string) {
t.Helper()
opts, warns, err := GenerateWithWarnings(ruleModelWithSets(sets, rules...))
if err != nil {
t.Fatalf("Generate: %v", err)
}
return opts.Route, warns
}
// inlineDomainSet builds an inline DOMAIN `config ruleset`.
//
// Mind the convention the move to rule-sets brought with it: a BARE entry here is
// a DomainSuffix (the apex AND its subdomains), whereas the routing rule's old
// dst_domain read a bare entry as an EXACT domain. `full:` is the exact form.
func inlineDomainSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "domain", Source: "inline", Entries: entries}
}
// inlineIPSet builds an inline IP-RANGE `config ruleset`.
func inlineIPSet(name string, entries ...string) model.Ruleset {
return model.Ruleset{Name: name, Type: "ipcidr", Source: "inline", Entries: entries}
}
func routeWarnsHave(warns []string, substr string) bool {
for _, w := range warns {
if strings.Contains(w, substr) {
@@ -68,69 +102,117 @@ func generalRules(rt *option.RouteOptions) []option.Rule {
// --- geosite: the landmine ---------------------------------------------------
// TestGeositeMatcherIsInertNotFatal: route-rule `geosite` was REMOVED from this
// engine — route/rule.NewDefaultRule returns "geosite database is deprecated ...
// removed in sing-box 1.12.0" for a non-empty Geosite list, and that error aborts
// box.New for the whole config. A `geosite:` entry must therefore be warned and
// omitted (exactly like the geoip matcher below), never emitted.
func TestGeositeMatcherIsInertNotFatal(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "geo", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube"}, Target: "node:n1",
})
// TestGeositeEntryInRulesetIsInertNotFatal: route-rule `geosite` was REMOVED from
// this engine — route/rule.NewDefaultRule returns "geosite database is deprecated
// ... removed in sing-box 1.12.0" for a non-empty Geosite list, and that error
// aborts box.New for the whole config. Destinations now live in a `config
// ruleset`, so a `geosite:` entry lands in an inline domain list, where it is an
// unrecognised `word:` prefix: dropped by the classifier, reported, and — being
// the list's only entry — leaving the rule-set with nothing to match, so it is
// skipped and the rule that referenced it is not emitted either. Nothing about
// that path may ever put a value in RawDefaultRule.Geosite. (`source=geosite` on
// the ruleset itself is the working way to route a category; see
// TestRoutingRuleSetGeositeCategory.)
func TestGeositeEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("geo", "geosite:youtube")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.Geosite) > 0 {
t.Fatalf("geosite must never be emitted (aborts box.New), got %v", r.DefaultOptions.RawDefaultRule.Geosite)
}
}
if !routeWarnsHave(warns, "geosite matcher") {
t.Fatalf("expected an inert-geosite warning, got %v", warns)
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a rule-set with no usable entry must not be emitted (an empty list would match everything)")
}
if findRouteRuleWithRuleSet(rt, "rs-geo") != nil {
t.Fatalf("no rule may reference the skipped rule-set")
}
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("expected an unrecognised-prefix warning for geosite:, got %v", warns)
}
if !routeWarnsHave(warns, "no usable entries") {
t.Fatalf("expected a no-usable-entries warning, got %v", warns)
}
}
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule carrying BOTH a geosite entry
// and a real domain keeps the real matcher and still routes — only the geosite
// part is dropped.
// TestGeositeMixedWithRealDomainKeepsTheRest: a rule-set carrying BOTH a geosite
// entry and a real domain keeps the real matcher, and the rule referencing it
// still routes — only the geosite entry is dropped.
func TestGeositeMixedWithRealDomainKeepsTheRest(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
DstDomain: []string{"geosite:youtube", "example.com"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d (warnings %v)", len(gen), warns)
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mixed", "geosite:youtube", "full:example.com")},
model.Rule{Name: "mixed", Enabled: true, Order: 10, DstRuleset: []string{"mixed"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-mixed")
if !ok {
t.Fatalf("rs-mixed must survive the geosite entry (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.Geosite) != 0 {
t.Fatalf("geosite leaked: %v", raw.Geosite)
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", hr)
}
if len(raw.Domain) != 1 || raw.Domain[0] != "example.com" {
t.Fatalf("real domain matcher lost: %+v", raw.Domain)
if len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.DomainRegex) != 0 {
t.Fatalf("geosite: must be dropped, not reinterpreted: %+v", hr)
}
if got := gen[0].DefaultOptions.RuleAction.RouteOptions.Outbound; got != "n1" {
dr := findRouteRuleWithRuleSet(rt, "rs-mixed")
if dr == nil {
t.Fatalf("the rule must be emitted referencing rs-mixed; rules=%+v", rt.Rules)
}
if got := dr.RuleAction.RouteOptions.Outbound; got != "n1" {
t.Fatalf("target = %q, want n1", got)
}
}
// TestGeoipEntryInRulesetIsInertNotFatal is the IP-side twin: model.migrate1to2
// moves an old `dst_ip geoip:ru` verbatim into an inline type=ipcidr rule-set
// (deliberately — see TestMigrate1to2KeepsGeoMarkersInert: it must not silently
// become a working geoip source, because the operator never asked to download
// anything). Here it is an unparseable prefix, so it must be dropped LOUDLY
// rather than reach NewIPCIDRItem, which errors and aborts box.New.
func TestGeoipEntryInRulesetIsInertNotFatal(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("geo", "geoip:ru")},
model.Rule{Name: "geo", Enabled: true, Order: 10, DstRuleset: []string{"geo"}, Target: "node:n1"},
)
for _, r := range rt.Rules {
if len(r.DefaultOptions.RawDefaultRule.GeoIP) > 0 {
t.Fatalf("geoip must never be emitted, got %v", r.DefaultOptions.RawDefaultRule.GeoIP)
}
}
if _, ok := ruleSetByTag(rt, "rs-geo"); ok {
t.Fatalf("a list whose only entry is unparseable must not materialise")
}
if !routeWarnsHave(warns, `bad ip_cidr entry "geoip:ru"`) {
t.Fatalf("expected a bad-ip_cidr warning naming the entry, got %v", warns)
}
}
// --- malformed matchers that used to abort box.New ---------------------------
// TestBadDstCIDRWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. It must be dropped.
func TestBadDstCIDRWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstIP: []string{"999.1.1.1/24", "198.51.100.0/24"}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadIPCIDREntryWarnsAndSkips: an unparseable ip_cidr makes
// route/rule.NewIPCIDRItem error, which aborts box.New. Destination addresses are
// an inline `type=ipcidr` rule-set now, so the guard lives in inlineRulesetRule:
// the typo is dropped, the valid entry survives and the rule still routes.
func TestRulesetBadIPCIDREntryWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineIPSet("bad", "999.1.1.1/24", "198.51.100.0/24")},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one bad entry must not take the whole list down (warnings %v)", warns)
}
raw := gen[0].DefaultOptions.RawDefaultRule
if len(raw.IPCIDR) != 1 || raw.IPCIDR[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", raw.IPCIDR)
got := rs.InlineOptions.Rules[0].DefaultOptions.IPCIDR
if len(got) != 1 || got[0] != "198.51.100.0/24" {
t.Fatalf("ip_cidr = %v, want only the valid entry", got)
}
if !routeWarnsHave(warns, `destination "999.1.1.1/24" is not a valid IP/CIDR`) {
t.Fatalf("expected a bad-destination warning, got %v", warns)
if !routeWarnsHave(warns, `bad ip_cidr entry "999.1.1.1/24"`) {
t.Fatalf("expected a bad-ip_cidr warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
@@ -152,24 +234,30 @@ func TestBadSrcCIDRWarnsAndSkips(t *testing.T) {
}
}
// TestBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New.
func TestBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "bad", Enabled: true, Order: 10,
DstDomain: []string{"regexp:*broken(", `regexp:^ok\.example$`}, Target: "node:n1",
})
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("want 1 general rule, got %d", len(gen))
// TestRulesetBadDomainRegexWarnsAndSkips: an uncompilable `regexp:` pattern makes
// route/rule.NewDomainRegexItem error and abort box.New. The pattern vocabulary
// moved into the inline rule-set with the rest of the destination list, so the
// validation moved with it (peelDomainRegexes): the broken pattern is dropped and
// the compilable one survives.
func TestRulesetBadDomainRegexWarnsAndSkips(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("bad", "regexp:*broken(", `regexp:^ok\.example$`)},
model.Rule{Name: "bad", Enabled: true, Order: 10, DstRuleset: []string{"bad"}, Target: "node:n1"},
)
rs, ok := ruleSetByTag(rt, "rs-bad")
if !ok {
t.Fatalf("one broken pattern must not take the whole list down (warnings %v)", warns)
}
got := gen[0].DefaultOptions.RawDefaultRule.DomainRegex
got := rs.InlineOptions.Rules[0].DefaultOptions.DomainRegex
if len(got) != 1 || got[0] != `^ok\.example$` {
t.Fatalf("domain_regex = %v, want only the compilable one", got)
}
if !routeWarnsHave(warns, "domain regexp") {
t.Fatalf("expected a bad-regexp warning, got %v", warns)
}
if findRouteRuleWithRuleSet(rt, "rs-bad") == nil {
t.Fatalf("the rule must still be emitted referencing rs-bad; rules=%+v", rt.Rules)
}
}
// TestBadPortRangeWarnsAndSkips: a malformed range reaches
@@ -911,10 +999,12 @@ func TestRuleProtoKnownValuesAreSilent(t *testing.T) {
// that would break existing configs either open or closed), but the widening is
// now reported with its consequence.
func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
rt, warns := genRules(t, model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstDomain: []string{"youtube.com"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "guest", Enabled: true, Order: 10,
Src: []string{"iface:guest"}, DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "applies to EVERY client") {
t.Fatalf("expected a rule-widening warning, got %v", warns)
}
@@ -930,11 +1020,13 @@ func TestIfaceOnlySourceRuleIsNotSilentlyNetworkWide(t *testing.T) {
// TestMixedSourceRuleReportsTheDroppedHalf: with one usable IP source alongside
// an unmatchable one, the rule narrows to the IP source only.
func TestMixedSourceRuleReportsTheDroppedHalf(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstDomain: []string{"youtube.com"}, Target: "block",
})
_, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("yt", "youtube.com")},
model.Rule{
Name: "mixed", Enabled: true, Order: 10,
Src: []string{"iface:guest", "aa:bb:cc:dd:ee:ff", "192.168.5.0/24"},
DstRuleset: []string{"yt"}, Target: "block",
})
if !routeWarnsHave(warns, "could not be used") {
t.Fatalf("expected a dropped-source warning, got %v", warns)
}
@@ -1041,41 +1133,73 @@ func TestRuleTargetKindsResolve(t *testing.T) {
}
}
// TestRuleDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped by
// the shared classifier (a domain cannot contain ":"), so the rule silently lost
// that destination. `domain:example.com` is the v0.1/xray spelling a migrating
// user writes, and it must not disappear without a trace.
func TestRuleDomainUnrecognisedPrefixWarns(t *testing.T) {
// TestRulesetDomainUnrecognisedPrefixWarns: an unknown `word:` prefix is dropped
// by the shared classifier (a domain cannot contain ":"), so the list silently
// lost that destination. `domain:example.com` is the v0.1/xray spelling a
// migrating user writes, and it must not disappear without a trace — the more so
// now that a destination list is ALWAYS a rule-set, i.e. the one place a typo can
// hide.
func TestRulesetDomainUnrecognisedPrefixWarns(t *testing.T) {
for _, entry := range []string{"domain:example.com", "regex:example.com", "ext:foo.dat:cn"} {
rt, warns := genRules(t, model.Rule{
Name: "mig", Enabled: true, Order: 10,
DstDomain: []string{entry, "keep.example"}, Target: "block",
})
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("mig", entry, "keep.example")},
model.Rule{Name: "mig", Enabled: true, Order: 10, DstRuleset: []string{"mig"}, Target: "block"},
)
if !routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("%q: expected an unrecognised-prefix warning, got %v", entry, warns)
}
gen := generalRules(rt)
if len(gen) != 1 {
t.Fatalf("%q: want 1 rule, got %d", entry, len(gen))
rs, ok := ruleSetByTag(rt, "rs-mig")
if !ok {
t.Fatalf("%q: the usable entry must keep the rule-set alive; warnings %v", entry, warns)
}
for _, d := range gen[0].DefaultOptions.RawDefaultRule.Domain {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
hr := rs.InlineOptions.Rules[0].DefaultOptions
for _, list := range [][]string{hr.Domain, hr.DomainSuffix, hr.DomainKeyword, hr.DomainRegex} {
for _, d := range list {
if strings.Contains(d, ":") {
t.Fatalf("%q: a prefixed literal reached the matcher: %q", entry, d)
}
}
}
if findRouteRuleWithRuleSet(rt, "rs-mig") == nil {
t.Fatalf("%q: the rule must still be emitted; rules=%+v", entry, rt.Rules)
}
}
}
// TestRuleDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary routing rules really support.
func TestRuleDomainKnownPrefixesAreSilent(t *testing.T) {
_, warns := genRules(t, model.Rule{
Name: "ok", Enabled: true, Order: 10, Target: "block",
DstDomain: []string{"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example"},
})
// TestRulesetDomainKnownPrefixesAreSilent guards the warning against false
// positives on the vocabulary an inline domain rule-set really supports.
//
// `regexp:` is the load-bearing case: the shared classifier has no branch for it,
// so it WOULD be reported as an unknown prefix — inlineRulesetRule peels the
// regexes off first (peelDomainRegexes) precisely so it is not. A regression there
// would both warn about a working matcher and drop it.
func TestRulesetDomainKnownPrefixesAreSilent(t *testing.T) {
rt, warns := genRulesWithSets(t,
[]model.Ruleset{inlineDomainSet("ok",
"full:a.example", "suffix:b.example", "keyword:c", `regexp:^d\.`, ".e.example", "f.example")},
model.Rule{Name: "ok", Enabled: true, Order: 10, DstRuleset: []string{"ok"}, Target: "block"},
)
if routeWarnsHave(warns, "unrecognised prefix") {
t.Fatalf("the supported prefixes must not warn: %v", warns)
}
rs, ok := ruleSetByTag(rt, "rs-ok")
if !ok {
t.Fatalf("rs-ok not emitted; warnings %v", warns)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.Domain) != 1 || hr.Domain[0] != "a.example" {
t.Fatalf("full: must be an exact Domain, got %v", hr.Domain)
}
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^d\.` {
t.Fatalf("regexp: must survive as a domain_regex matcher, got %v", hr.DomainRegex)
}
// suffix:, the leading dot and the BARE entry all collapse to domain_suffix.
if len(hr.DomainSuffix) != 3 {
t.Fatalf("domain_suffix = %v, want b/e/f.example (bare entry is a suffix in a rule-set)", hr.DomainSuffix)
}
if len(hr.DomainKeyword) != 1 || hr.DomainKeyword[0] != "c" {
t.Fatalf("domain_keyword = %v, want [c]", hr.DomainKeyword)
}
}
// TestDeviceDomainUnrecognisedPrefixWarns is the same guarantee in the place it
+173
View File
@@ -0,0 +1,173 @@
// B1: a routing rule that can never fire, reported instead of applied silently.
//
// The field config that prompted this had TWO rules named `default`, both with
// zero conditions — order 20 -> direct and order 100 -> group:auto. buildRoute
// points route Final at a condition-less rule and moves on, so the LAST one wins
// and the other is a dead setting. Nothing anywhere said so: the log was clean and
// the panel drew both rows identically.
//
// These tests pin the diagnosis AND the behaviour it describes, because a warning
// that disagrees with what the generator actually does is worse than none.
package generate
import (
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// warnAbout returns the warnings mentioning `rule "name"`.
func warnAbout(warns []string, name string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, `rule "`+name+`"`) {
out = append(out, w)
}
}
return out
}
func warnsMatching(warns []string, substr string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, substr) {
out = append(out, w)
}
}
return out
}
const neverApplied = "is never applied"
// twoDefaultsModel reproduces the field config: two condition-less rules sharing
// the name `default`, differing only in Order and target.
func twoDefaultsModel(lowTarget, highTarget string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: lowTarget},
{Name: "default", Enabled: true, Order: 100, Target: highTarget},
},
}
}
// TestTwoCatchAllRulesWarnAndLastWins is the B1 regression. The order-100 rule is
// the default the engine uses (route Final), and the order-20 one is reported as
// never applied — naming the rule that supersedes it.
func TestTwoCatchAllRulesWarnAndLastWins(t *testing.T) {
opts, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the LAST catch-all's target %q", got, "auto")
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
// It must name the superseding rule AND its order — with both rules called
// `default`, the order is the only thing that tells the two apart.
// Targets are quoted as the OPERATOR wrote them (`group:auto`), not as the
// engine tag they resolve to (`auto`) — the warning has to be readable next to
// the config, not next to the generated JSON.
for _, want := range []string{`rule "default"`, "order 100", `"group:auto"`, `"direct"`} {
if !strings.Contains(got[0], want) {
t.Fatalf("warning must mention %s, got: %s", want, got[0])
}
}
// Diagnosis only: neither rule is renamed, reordered or dropped.
if len(opts.Route.Rules) == 0 {
t.Fatal("route rules disappeared")
}
}
// TestTwoCatchAllsDirectWinsIsCritical: when the surviving default is `direct` and
// the retired one asked for a tunnel, everything unmatched leaves on the plain
// WAN. The warning must say so in the words apply/warnings.go grades critical —
// this is the case where a healthy-looking panel is a lie.
func TestTwoCatchAllsDirectWinsIsCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("group:auto", "direct"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
const marker = "leaves over the plain WAN with your real IP address"
if !strings.Contains(got[0], marker) {
t.Fatalf("a retired tunnel default under a live direct default must carry %q, got: %s", marker, got[0])
}
}
// TestTwoCatchAllsTunnelWinsIsNotCritical: the field config's actual shape — the
// dead rule is `direct` and the live default is the tunnel. That is a dead
// setting, not a leak, so it must NOT carry the critical marker.
func TestTwoCatchAllsTunnelWinsIsNotCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
for _, w := range warnsMatching(warns, neverApplied) {
if strings.Contains(w, "leaves over the plain WAN with your real IP address") {
t.Fatalf("a dead direct default under a live tunnel default is not a leak: %s", w)
}
}
}
// TestSingleCatchAllDoesNotWarn: the ordinary config — specific rules plus ONE
// default — must stay silent. A badge on a working rule teaches the operator to
// ignore the badge.
func TestSingleCatchAllDoesNotWarn(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "ads", Enabled: true, Order: 10, DstPort: "443", Target: "block"},
// A specific rule ordered BELOW the default: still emitted ahead of Final,
// so it is not retired either.
{Name: "default", Enabled: true, Order: 20, Target: "group:auto"},
{Name: "late", Enabled: true, Order: 900, DstPort: "8080", Target: "direct"},
},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a config with one default must not report anything never-applied, got: %q", got)
}
if got := warnAbout(warns, "late"); len(got) != 0 {
t.Fatalf("a specific rule below the default is not shadowed by it, got: %q", got)
}
}
// TestProfileDisabledCatchAllDoesNotShadow: the active profile switches the later
// default off, so the earlier one is the live default and must not be badged.
// Judging the raw config would blame the wrong rule on every profile router.
func TestProfileDisabledCatchAllDoesNotShadow(t *testing.T) {
m := twoDefaultsModel("group:auto", "direct")
m.Rules[1].Name = "fallback" // profiles address rules by name
m.Globals.ActiveProfile = "home"
m.Profiles = []model.Profile{{Name: "home", Enabled: true, DisableRules: []string{"fallback"}}}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a profile-disabled default shadows nothing, got: %q", got)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the surviving default's target %q", got, "auto")
}
}
+63 -3
View File
@@ -25,6 +25,7 @@ import (
"net/netip"
"os"
"path/filepath"
"regexp"
"runtime/debug"
"strings"
"sync"
@@ -1163,7 +1164,8 @@ func (b *builder) buildRoutingRuleSetRaw(rs model.Ruleset) ([]option.RuleSet, []
// its Entries, keyed by Type: an ipcidr ruleset fills ip_cidr; a domain ruleset
// (the default) is classified with inlineDomainRule — bare entry => DomainSuffix
// (so subdomains match), full: => Domain, keyword: => DomainKeyword, . => suffix
// — the same classification the DNS filter uses. ok=false when nothing usable.
// — the same classification the DNS filter uses, plus `regexp:` (see
// peelDomainRegexes). ok=false when nothing usable.
func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRule, bool) {
b.warnUnknownRuleSetType(fmt.Sprintf("ruleset %q", rs.Name), rs.Type)
if ruleSetTypeIsIPCIDR(rs.Type) {
@@ -1190,10 +1192,68 @@ func (b *builder) inlineRulesetRule(rs model.Ruleset) (option.DefaultHeadlessRul
return option.DefaultHeadlessRule{IPCIDR: badoption.Listable[string](cidrs)}, true
}
// "domain" (and empty, and anything unrecognised => domain, warned above).
// `regexp:` is peeled off first: it is a routing-rule matcher the DNS-filter
// classifier does not know, and it must not be reported as an unknown prefix.
diag := fmt.Sprintf("ruleset %q", rs.Name)
rest, regexes := b.peelDomainRegexes(diag, rs.Entries)
// R4, on the path every destination list now takes. classifyDomainEntries
// DROPS an entry that is nothing but its marker (".", "full:", "keyword:"),
// silently — and the silence is the dangerous half: an empty domain token
// aborts box.New for the whole config, and an empty keyword is
// strings.Contains(host, "") i.e. EVERY host. The drop is right; not saying so
// is not. (devices.go reports the same class for a device's own lists; the bare
// `regexp:` form is reported by peelDomainRegexes above, which is why it is
// peeled off before this loop and cannot be double-reported.)
for _, e := range rest {
if isDomainMarkerOnly(e) {
b.warnf("%s: entry %q is a bare matcher marker with no value, omitted (an empty domain token aborts box.New; an empty keyword would match EVERY host)", diag, strings.TrimSpace(e))
}
}
// Only the DOMAIN branch reports unknown `word:` prefixes — the ipcidr branch
// above is full of legitimate colons (IPv6) and must never be checked (R9.2).
b.warnUnrecognisedPrefixes(fmt.Sprintf("ruleset %q", rs.Name), rs.Entries)
return inlineDomainRule(rs.Entries)
b.warnUnrecognisedPrefixes(diag, rest)
hr, ok := inlineDomainRule(rest)
if len(regexes) > 0 {
hr.DomainRegex = badoption.Listable[string](regexes)
ok = true
}
return hr, ok
}
// peelDomainRegexes splits `regexp:<pattern>` entries out of a domain rule-set's
// entry list, returning the remaining entries and the validated patterns.
//
// It exists because a destination list is now ALWAYS a rule-set (schema v2), so
// every matcher a `dst_domain` used to express has to be expressible here —
// including the regex form, which the shared DNS-filter classifier
// (classifyDomainEntries) deliberately does not know about. Validation mirrors
// what the routing rule did before the move, and for the same reason:
// route/rule.NewDomainRegexItem returns an error for an uncompilable pattern and
// that aborts box.New for the WHOLE config, so a bad pattern must degrade to a
// warning. A BARE `regexp:` compiles fine but matches every host — the same
// silent match-all hazard as an empty keyword — so it is dropped too.
func (b *builder) peelDomainRegexes(diag string, entries []string) (rest, regexes []string) {
for _, e := range entries {
e = strings.TrimSpace(e)
if e == "" {
continue
}
if !strings.HasPrefix(strings.ToLower(e), "regexp:") {
rest = append(rest, e)
continue
}
re := strings.TrimSpace(e[len("regexp:"):])
if re == "" {
b.warnf("%s: %q is a bare matcher marker with no value, omitted (an empty regexp matches EVERY host)", diag, e)
continue
}
if _, err := regexp.Compile(re); err != nil {
b.warnf("%s: domain regexp %q is invalid (%v), omitted", diag, re, err)
continue
}
regexes = append(regexes, re)
}
return rest, regexes
}
// Ruleset.Type — the two shapes a rule-set can match, and the accepted spellings.
+52
View File
@@ -33,6 +33,14 @@ func findRouteRuleWithRuleSet(rt *option.RouteOptions, tag string) *option.Defau
return nil
}
// hasRulesetRule reports whether any emitted route rule references the rule-set
// named name (tag rs-<name>). Since schema v2 a rule's destination is ALWAYS a
// rule-set reference, so this is how a test says "that rule was emitted" — the
// former "does any rule carry this dst domain" question has no answer any more.
func hasRulesetRule(rt *option.RouteOptions, name string) bool {
return findRouteRuleWithRuleSet(rt, routeRulesetTagPrefix+name) != nil
}
func ruleSetByTag(rt *option.RouteOptions, tag string) (option.RuleSet, bool) {
if rt == nil {
return option.RuleSet{}, false
@@ -117,6 +125,50 @@ func TestRoutingRuleSetInlineDomain(t *testing.T) {
}
}
// TestRoutingRuleSetInlineDomainRegex: `regexp:` is a matcher the shared domain
// classifier does NOT know — it belongs to the routing plane, and it used to be
// peeled off inside ruleMatchers, which no longer sees any domains at all. It
// therefore had to move into the inline rule-set with the rest of the destination
// vocabulary (peelDomainRegexes), or every migrated `regexp:` entry would have
// been reported as an unknown prefix and silently dropped: a routing rule that
// looks configured and matches nothing.
func TestRoutingRuleSetInlineDomainRegex(t *testing.T) {
m := &model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{
{Name: "ads", Type: "domain", Source: "inline", Entries: []string{`regexp:^ads\.`}},
},
Rules: []model.Rule{
{Name: "block-ads", Enabled: true, Order: 10, DstRuleset: []string{"ads"}, Target: "block"},
},
}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if len(warns) != 0 {
t.Fatalf("a valid regexp: entry must not warn: %v", warns)
}
rs, ok := ruleSetByTag(opts.Route, "rs-ads")
if !ok {
t.Fatalf("a regexp-only rule-set must still materialise; route=%+v", opts.Route)
}
hr := rs.InlineOptions.Rules[0].DefaultOptions
if len(hr.DomainRegex) != 1 || hr.DomainRegex[0] != `^ads\.` {
t.Fatalf("domain_regex = %+v, want [^ads\\.]", hr.DomainRegex)
}
if len(hr.Domain)+len(hr.DomainSuffix)+len(hr.DomainKeyword)+len(hr.IPCIDR) != 0 {
t.Fatalf("the regexp entry must not leak into another matcher: %+v", hr)
}
dr := findRouteRuleWithRuleSet(opts.Route, "rs-ads")
if dr == nil {
t.Fatalf("no route rule references rs-ads; rules=%+v", opts.Route.Rules)
}
if dr.RuleAction.RouteOptions.Outbound != tagBlock {
t.Fatalf("route rule must route to %q, got %+v", tagBlock, dr.RuleAction)
}
}
// TestRoutingRuleSetInlineIPCIDR: an ipcidr ruleset fills ip_cidr (not domain*),
// and the route rule references it.
func TestRoutingRuleSetInlineIPCIDR(t *testing.T) {
+21 -28
View File
@@ -12,47 +12,36 @@ import (
"testing"
"time"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// scheduledDomain is the distinctive dst-domain matcher used to detect whether a
// scheduled rule survived into the generated route rules.
const scheduledDomain = "sched.example"
// scheduledSet is the distinctive destination rule-set used to detect whether a
// scheduled rule survived into the generated route rules. Since schema v2 a
// rule's destination is a rule-set reference, so this doubles as a check that
// buildRoutingRuleSets honours the same schedule gate buildRoute does: outside
// the window neither the rule nor its rs- rule-set may be emitted.
const scheduledSet = "sched"
// emittedAt reports whether the given scheduled rule is present in the route
// rules when generate's clock is `now`.
func emittedAt(now time.Time, r model.Rule) bool {
b := newBuilder(&model.Model{Globals: model.DefaultGlobals(), Rules: []model.Rule{r}})
b := newBuilder(&model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{{Name: scheduledSet, Type: "domain", Source: "inline", Entries: []string{"sched.example"}}},
Rules: []model.Rule{r},
})
b.now = now
return hasDomainRule(b.buildRoute(), scheduledDomain)
return hasRulesetRule(b.buildRoute(), scheduledSet)
}
// hasDomainRule reports whether any route rule carries the given exact dst-domain
// matcher.
func hasDomainRule(rt *option.RouteOptions, domain string) bool {
if rt == nil {
return false
}
for _, r := range rt.Rules {
for _, d := range r.DefaultOptions.RawDefaultRule.Domain {
if d == domain {
return true
}
}
}
return false
}
// schedRule builds a scheduled dst-domain rule (target direct) from the schedule
// fields. It always carries the scheduledDomain matcher so emittedAt can find it.
// schedRule builds a scheduled destination rule (target direct) from the schedule
// fields. It always references the scheduledSet rule-set so emittedAt can find it.
func schedRule(days []string, start, end string) model.Rule {
return model.Rule{
Name: "sched",
Enabled: true,
Order: 10,
DstDomain: []string{scheduledDomain},
DstRuleset: []string{scheduledSet},
Target: "direct",
SchedEnabled: true,
SchedDays: days,
@@ -127,9 +116,13 @@ func TestScheduleAllDayWeekend(t *testing.T) {
// warning instead.
func TestScheduleInvalidTimeIsAlwaysOn(t *testing.T) {
rule := schedRule(nil, "9am", "17:00") // "9am" is not HH:MM
b := newBuilder(&model.Model{Globals: model.DefaultGlobals(), Rules: []model.Rule{rule}})
b := newBuilder(&model.Model{
Globals: model.DefaultGlobals(),
Rulesets: []model.Ruleset{{Name: scheduledSet, Type: "domain", Source: "inline", Entries: []string{"sched.example"}}},
Rules: []model.Rule{rule},
})
b.now = time.Date(2026, 7, 15, 3, 0, 0, 0, time.UTC) // 03:00 — would be OUTSIDE a 09–17 window
if !hasDomainRule(b.buildRoute(), scheduledDomain) {
if !hasRulesetRule(b.buildRoute(), scheduledSet) {
t.Errorf("invalid start time should fail OPEN (rule emitted always-on)")
}
if len(b.warnings) == 0 {
+47
View File
@@ -0,0 +1,47 @@
package logsink
// Colour policy for everything that writes into the sink.
//
// Both log producers of the daemon (the control-plane factory built in
// cmd/shaterd, and the engine's own factory built by box.New from
// option.LogOptions) format with github.com/logrusorgru/aurora colours ON by
// default: log/format.go paints the level word and the per-connection ID unless
// DisableColors / LogOptions.DisableColor is set. Under procd the daemon's
// stderr is not a terminal, it is syslog — so every ERROR line landed in
// `logread` as
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m1728741629ESC[0m 70ms] dns: …
//
// Syslog is not a screen: `logread | grep ERROR` misses the coloured word
// because there are invisible bytes inside it, log collectors store the escapes
// forever, and anyone reading a captured log sees mojibake. The sink's file half
// already strips ANSI on the way out (emitLocked -> stripANSI) — the syslog half
// deliberately did not, and that is the leak.
//
// The fix is at the producer, not at the sink: colour is a property of the
// DESTINATION, so it is decided once, from whether that destination is a
// terminal, and never emitted otherwise. Stripping at the sink would keep the
// wasted formatting work and would still leak through any future path that does
// not go through the sink.
import "os"
// IsTTY reports whether w is a terminal, i.e. whether ANSI colour escapes
// written to it will be RENDERED rather than stored.
//
// The check is the portable one — a character device — so it needs no cgo, no
// termios ioctl and no new dependency (the router binary is CGO_ENABLED=0
// musl-static, and logsink also builds on the Windows/macOS dev hosts). Under
// procd, stderr is a pipe to the log daemon: not a character device, so colour
// is off, which is the case that matters. An interactive `shaterd run` from a
// shell keeps its colours.
func IsTTY(w *os.File) bool {
if w == nil {
return false
}
fi, err := w.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
}
+94
View File
@@ -0,0 +1,94 @@
package logsink
import (
"bytes"
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestIsTTYNonTerminal pins the only direction that matters in production: the
// destinations the daemon actually has under procd — a pipe (procd's stderr
// relay) and a regular file — are NOT terminals, so nothing may colour for them.
func TestIsTTYNonTerminal(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
if IsTTY(w) {
t.Errorf("IsTTY(pipe) = true; procd's stderr is a pipe and must never be coloured")
}
f, err := os.Create(filepath.Join(t.TempDir(), "log"))
if err != nil {
t.Fatalf("create: %v", err)
}
t.Cleanup(func() { _ = f.Close() })
if IsTTY(f) {
t.Errorf("IsTTY(regular file) = true, want false")
}
if IsTTY(nil) {
t.Errorf("IsTTY(nil) = true, want false")
}
}
// TestSyslogHalfHasNoANSI is the regression for the escape codes that reached
// `logread`:
//
// daemon.err shaterd[27540]: …Z ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: …
//
// It wires a factory the way the daemon does — formatter colour gated on
// IsTTY(destination), output into a Sink whose syslog half is captured — and
// asserts the captured bytes carry no ESC (0x1b). The context ID is set because
// the ID is coloured by a SEPARATE branch of log/format.go: a fix that only
// silenced the level word would still leak here.
func TestSyslogHalfHasNoANSI(t *testing.T) {
r, w, err := os.Pipe() // a non-terminal destination, exactly like procd's stderr
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
var syslog bytes.Buffer
s := New(&syslog, Config{ToSyslog: true})
factory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now(), DisableColors: !IsTTY(w)},
s, "", nil, false)
logger := factory.Logger()
ctx := log.ContextWithNewID(context.Background())
logger.ErrorContext(ctx, "dns: exchange failed for catalog.example.com. IN AAAA: unexpected EOF")
logger.WarnContext(ctx, "warn line")
logger.InfoContext(ctx, "info line")
factory.SetLevel(log.LevelTrace)
logger.DebugContext(ctx, "debug line")
logger.TraceContext(ctx, "trace line")
_ = s.Close()
got := syslog.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("the syslog half captured nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("syslog half carries an ANSI escape at byte %d: %q", i, got)
}
if !strings.Contains(got, "ERROR") {
t.Errorf("`grep ERROR` must match a plain, unbroken level word: %q", got)
}
// Teeth check: the SAME line through a colouring formatter must contain the
// escape. Without this, a future aurora that stopped colouring would make the
// assertion above pass for the wrong reason and the guard would rot silently.
coloured := log.Formatter{BaseTime: time.Now()}.
Format(ctx, log.LevelError, "", "dns: exchange failed", time.Now())
if !strings.ContainsRune(coloured, 0x1b) {
t.Fatalf("colouring formatter emitted no ANSI escape (%q) — this test can no longer detect the leak", coloured)
}
}
+259 -1
View File
@@ -14,17 +14,31 @@ import (
// CurrentSchemaVersion is the schema this build understands. Bump it when adding
// a migration step below.
const CurrentSchemaVersion = 1
const CurrentSchemaVersion = 2
// uciRunner abstracts uci get/set/delete/commit/import so migrations AND the
// config-write path (WriteUCI) are unit-testable. Import feeds `uci export`-format
// text to `uci import <pkg>` on stdin, replacing the package's staged sections.
//
// Export/Add/AddList exist for migrations that have to READ the config they are
// rewriting and GROW it. migrate1to2 needs both: it reads options the current
// Model no longer parses (`dst_domain`/`dst_ip` were removed from Rule) and adds
// the `config ruleset` sections it folds them into. Add returns the generated
// section id — the anonymous-index drift trap is real (`@ruleset[3]` means
// something different after one more add), so every write goes through the id.
type uciRunner interface {
Get(key string) (string, bool)
Set(key, val string) error
Delete(key string) error
Commit(pkg string) error
Import(pkg, text string) error
// Export returns the package in `uci export` format. ok=false when the package
// does not exist (nothing to migrate), which is NOT an error.
Export(pkg string) (text string, ok bool)
// Add appends an anonymous `config <secType>` and returns its section id.
Add(pkg, secType string) (id string, err error)
// AddList appends one value to a list option (`uci add_list <key>=<val>`).
AddList(key, val string) error
}
type execUCI struct{}
@@ -46,6 +60,30 @@ func (execUCI) Import(pkg, text string) error {
return cmd.Run()
}
func (execUCI) Export(pkg string) (string, bool) {
out, err := exec.Command("uci", "-q", "export", pkg).Output()
if err != nil {
return "", false
}
return string(out), true
}
func (execUCI) Add(pkg, secType string) (string, error) {
out, err := exec.Command("uci", "add", pkg, secType).Output()
if err != nil {
return "", fmt.Errorf("uci add %s %s: %w", pkg, secType, err)
}
id := strings.TrimSpace(string(out))
if id == "" {
return "", fmt.Errorf("uci add %s %s: no section id returned", pkg, secType)
}
return id, nil
}
func (execUCI) AddList(k, v string) error {
return exec.Command("uci", "add_list", k+"="+v).Run()
}
// uci is the active runner (overridable in tests).
var uci uciRunner = execUCI{}
@@ -56,6 +94,7 @@ type migration struct {
var migrations = []migration{
{from: 0, to: 1, apply: migrate0to1},
{from: 1, to: 2, apply: migrate1to2},
}
func readSchemaVersion(u uciRunner) int {
@@ -123,3 +162,222 @@ func migrate0to1(u uciRunner) error {
}
return u.Commit("shater")
}
// --- v1 -> v2: a rule's destination is a rule-set, never an inline list ------
//
// WHAT CHANGED. `config rule` lost `dst_domain` and `dst_ip`. A rule now names
// its destination through `dst_ruleset` only, so there is ONE destination
// mechanism, one matcher vocabulary, and one place a list is edited — and the
// list is compiled once into a .srs that every referencing rule shares.
//
// WHAT THIS STEP DOES. For every rule that still carries one of the two options
// it creates an inline `config ruleset` named `rule-<rule name>` (and
// `rule-<rule name>-ip` for the address list, because a rule-set is EITHER a
// domain list or an ip_cidr list), moves the entries into it, appends the new
// name to the rule's `dst_ruleset`, and deletes the legacy option. Nothing is
// dropped and nothing is guessed: a rule with both lists gets both rule-sets.
//
// ENTRY SEMANTICS ARE PRESERVED 1:1, and that needs one real conversion. The two
// contexts disagree about exactly one form: a BARE domain is an EXACT match in a
// routing rule (generate/route.go classified `dst_domain` with bareIsSuffix=false)
// and a SUFFIX match inside a rule-set (inlineDomainRule, bareIsSuffix=true).
// Copying `example.com` across verbatim would therefore silently widen the rule to
// every subdomain, so a bare entry is rewritten as `full:example.com`. Every other
// form already means the same thing on both sides and is copied byte-for-byte:
// `full:`, `suffix:`, `keyword:`, `regexp:` (see generate.peelDomainRegexes, added
// with this change so the regex form survives the move) and a leading dot, which
// is a synonym of `suffix:` in both. `geosite:`/`geoip:` entries are copied
// unchanged too: they are INERT in a routing rule on this engine (warned and
// omitted — the route-rule geosite/geoip fields no longer exist), and an
// unrecognised marker is equally inert inside a rule-set, so their meaning is
// unchanged and the operator's text is not thrown away. Converting them into a
// `source=geosite` rule-set would have made a dead matcher start routing traffic
// during an upgrade — a behaviour change, not a migration.
//
// ONE DELIBERATE SEMANTIC IMPROVEMENT, stated out loud: a rule that carried BOTH
// a domain list and an address list matched them with AND (an engine route rule
// ANDs its matcher fields), which is almost never what "these sites and these
// networks" was meant to say. The two generated rule-sets are ORed, because
// `rule_set: [a, b]` matches when EITHER matches. Such a rule matches more after
// the migration than before — it is called out here, in the docs, and it only
// affects configs that used both fields at once.
//
// IDEMPOTENCE. The legacy options are deleted as the last step per rule, so a
// second run finds nothing to do. A run interrupted between "create the rule-set"
// and "delete the option" is also safe: a rule that ALREADY references a rule-set
// of the expected name reuses it instead of creating `rule-<name>-2`.
func migrate1to2(u uciRunner) error {
text, ok := u.Export("shater")
if !ok || strings.TrimSpace(text) == "" {
return nil // no config yet (fresh install): nothing to migrate
}
secs, err := parseSections(text)
if err != nil {
return fmt.Errorf("read the current config: %w", err)
}
// Every rule-set name already in use, so a generated one can never collide with
// a hand-written list (which would make `uci` hold two `config ruleset` blocks
// claiming the same name, and the generator drops one as a duplicate tag).
taken := map[string]bool{}
for _, s := range secs {
if s.Type == "ruleset" {
if n := firstNonEmpty(s.opt("name"), s.Name); n != "" {
taken[n] = true
}
}
}
ruleIdx := -1
for _, s := range secs {
if s.Type != "rule" {
continue
}
// Anonymous sections are addressed positionally and the index is PER TYPE, so
// it counts rules only. Appending `config ruleset` sections below cannot shift
// it (a new section goes to the end, and it is not a rule).
ruleIdx++
domains := s.list("dst_domain")
ips := s.list("dst_ip")
if len(domains) == 0 && len(ips) == 0 {
continue
}
rulePath := fmt.Sprintf("shater.@rule[%d]", ruleIdx)
base := rulesetBaseName(firstNonEmpty(s.opt("name"), s.Name), ruleIdx)
refs := s.list("dst_ruleset")
if len(domains) > 0 {
entries := make([]string, 0, len(domains))
for _, d := range domains {
if e := migrateDomainEntry(d); e != "" {
entries = append(entries, e)
}
}
if err := ensureMigratedRuleset(u, rulePath, base, "domain", entries, refs, taken); err != nil {
return err
}
}
if len(ips) > 0 {
entries := make([]string, 0, len(ips))
for _, ip := range ips {
if v := strings.TrimSpace(ip); v != "" {
entries = append(entries, v)
}
}
if err := ensureMigratedRuleset(u, rulePath, base+"-ip", "ipcidr", entries, refs, taken); err != nil {
return err
}
}
// Last, so an interrupted run still has the legacy list to redo the work from.
_ = u.Delete(rulePath + ".dst_domain")
_ = u.Delete(rulePath + ".dst_ip")
}
return u.Commit("shater")
}
// ensureMigratedRuleset creates the inline `config ruleset` holding entries and
// points rulePath's dst_ruleset at it, unless the rule already references a
// rule-set of that name (a re-run after an interrupted migration). want is the
// preferred name; a collision with an existing list picks want-2, want-3, ...
// rsType is "domain" or "ipcidr". An entry list that came out empty creates
// nothing: an empty inline rule-set matches nothing and the generator would skip
// it, so a dangling reference would be pure noise.
func ensureMigratedRuleset(u uciRunner, rulePath, want, rsType string, entries, existingRefs []string, taken map[string]bool) error {
if len(entries) == 0 {
return nil
}
if taken[want] && containsString(existingRefs, want) {
return nil // already migrated (interrupted run); nothing to add
}
name := want
for i := 2; taken[name]; i++ {
name = fmt.Sprintf("%s-%d", want, i)
}
taken[name] = true
id, err := u.Add("shater", "ruleset")
if err != nil {
return err
}
sec := "shater." + id
if err := u.Set(sec+".name", name); err != nil {
return err
}
if err := u.Set(sec+".type", rsType); err != nil {
return err
}
if err := u.Set(sec+".source", "inline"); err != nil {
return err
}
for _, e := range entries {
if err := u.AddList(sec+".entry", e); err != nil {
return err
}
}
if containsString(existingRefs, name) {
return nil
}
return u.AddList(rulePath+".dst_ruleset", name)
}
// rulesetBaseName builds `rule-<name>` from a rule's name, reduced to characters
// that are safe in a rule-set name (it becomes an engine rule-set TAG, `rs-<name>`,
// and a UCI option value). An unnamed rule falls back to its position so two of
// them cannot produce the same base.
func rulesetBaseName(ruleName string, idx int) string {
var b strings.Builder
prevDash := false
for _, r := range strings.TrimSpace(ruleName) {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9', r == '_', r == '.':
b.WriteRune(r)
prevDash = false
default:
if !prevDash && b.Len() > 0 {
b.WriteByte('-')
prevDash = true
}
}
}
slug := strings.Trim(b.String(), "-.")
if slug == "" {
slug = strconv.Itoa(idx)
}
return "rule-" + slug
}
// migrateDomainEntry rewrites ONE `dst_domain` entry into the rule-set spelling
// with the same meaning. Only the bare form differs between the two contexts
// (exact in a rule, suffix in a rule-set), so only it is rewritten; see the
// migrate1to2 doc comment for the full table and the reasoning.
func migrateDomainEntry(e string) string {
v := strings.TrimSpace(e)
if v == "" {
return ""
}
// A leading dot already means `suffix:` on both sides.
if strings.HasPrefix(v, ".") {
return v
}
// Any `word:` marker — recognised (full/suffix/keyword/regexp) or not
// (geosite/geoip/typos) — carries its meaning across unchanged. A domain label
// cannot contain a colon, so this cannot misfire on a real host name.
if strings.Contains(v, ":") {
return v
}
return "full:" + v
}
// containsString reports whether list holds want (trimmed comparison).
func containsString(list []string, want string) bool {
for _, v := range list {
if strings.TrimSpace(v) == want {
return true
}
}
return false
}
+636 -21
View File
@@ -1,53 +1,668 @@
package model
import "testing"
import (
"fmt"
"strconv"
"strings"
"testing"
)
// fakeUCI is an in-memory uciRunner for the migration + write tests (no router
// needed). imported records the last `uci import` text so WriteUCI can be
// asserted without a device.
// fakeUCI is an in-memory stand-in for the `uci` CLI: enough of a section model
// that a migration can EXPORT the config, rewrite it, and export it again and see
// its own writes. That is what makes the idempotence assertions below real —
// against a flat key/value map a second migration run would re-read the original
// text and "prove" nothing.
//
// Addressing mirrors uci: `shater.globals.opt` (named section), `shater.@rule[2]`
// (positional, index is PER TYPE), `shater.cfg001.opt` (the id `uci add` returns).
type fakeUCI struct {
kv map[string]string
secs []*fakeSection
imported string
commits int
deleted []string
nextID int
// missing makes Export report "no such package", the fresh-install case.
missing bool
}
func (f *fakeUCI) Get(k string) (string, bool) { v, ok := f.kv[k]; return v, ok }
func (f *fakeUCI) Set(k, v string) error { f.kv[k] = v; return nil }
func (f *fakeUCI) Delete(k string) error { f.deleted = append(f.deleted, k); delete(f.kv, k); return nil }
func (f *fakeUCI) Commit(string) error { f.commits++; return nil }
func (f *fakeUCI) Import(pkg, text string) error { f.imported = text; return nil }
type fakeSection struct {
id string
typ string
name string // "" for an anonymous section
opts map[string]string
oKeys []string // option order, so the rendered export is deterministic
lists map[string][]string
lKeys []string
}
func newFakeUCI(export string) *fakeUCI {
f := &fakeUCI{}
if strings.TrimSpace(export) == "" {
return f
}
secs, err := parseSections(export)
if err != nil {
panic("fakeUCI fixture: " + err.Error())
}
for _, s := range secs {
sec := f.newSection(s.Type, s.Name)
for k, v := range s.Options {
sec.setOpt(k, v)
}
for k, vs := range s.Lists {
for _, v := range vs {
sec.addList(k, v)
}
}
}
return f
}
func (f *fakeUCI) newSection(typ, name string) *fakeSection {
sec := &fakeSection{
id: fmt.Sprintf("cfg%03d", f.nextID),
typ: typ,
name: name,
opts: map[string]string{},
lists: map[string][]string{},
}
f.nextID++
f.secs = append(f.secs, sec)
return sec
}
func (s *fakeSection) setOpt(k, v string) {
if _, seen := s.opts[k]; !seen {
s.oKeys = append(s.oKeys, k)
}
s.opts[k] = v
}
func (s *fakeSection) addList(k, v string) {
if _, seen := s.lists[k]; !seen {
s.lKeys = append(s.lKeys, k)
}
s.lists[k] = append(s.lists[k], v)
}
// resolve finds the section a `<pkg>.<sel>` selector names.
func (f *fakeUCI) resolve(sel string) *fakeSection {
if strings.HasPrefix(sel, "@") && strings.HasSuffix(sel, "]") {
open := strings.IndexByte(sel, '[')
if open < 0 {
return nil
}
typ := sel[1:open]
idx, err := strconv.Atoi(sel[open+1 : len(sel)-1])
if err != nil {
return nil
}
n := 0
for _, s := range f.secs {
if s.typ != typ {
continue
}
if n == idx {
return s
}
n++
}
return nil
}
for _, s := range f.secs {
if s.name == sel || s.id == sel {
return s
}
}
return nil
}
// splitKey cuts "shater.@rule[0].dst_domain" into ("@rule[0]", "dst_domain").
// A key with no option part yields opt == "".
func splitKey(key string) (sel, opt string) {
rest := strings.TrimPrefix(key, "shater")
rest = strings.TrimPrefix(rest, ".")
if rest == "" {
return "", ""
}
// The selector may contain a dot only inside a name, which the fixtures never
// use, so a plain LastIndex is enough — except for `@type[i]`, where the index
// brackets hold no dots either.
if i := strings.LastIndexByte(rest, '.'); i >= 0 {
return rest[:i], rest[i+1:]
}
return rest, ""
}
func (f *fakeUCI) Get(k string) (string, bool) {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return "", false
}
if opt == "" {
return sec.typ, true
}
v, ok := sec.opts[opt]
return v, ok
}
func (f *fakeUCI) Set(k, v string) error {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
if opt != "" {
return fmt.Errorf("uci set %s: no such section", k)
}
// `uci set shater.globals=globals` creates the named section.
f.newSection(v, sel)
return nil
}
if opt == "" {
sec.typ = v
return nil
}
sec.setOpt(opt, v)
return nil
}
func (f *fakeUCI) Delete(k string) error {
f.deleted = append(f.deleted, k)
if k == "shater" {
f.secs = nil
return nil
}
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return nil // `uci -q delete` on an absent key is a no-op
}
if opt == "" {
for i, s := range f.secs {
if s == sec {
f.secs = append(f.secs[:i], f.secs[i+1:]...)
break
}
}
return nil
}
delete(sec.opts, opt)
delete(sec.lists, opt)
return nil
}
func (f *fakeUCI) Commit(string) error { f.commits++; return nil }
func (f *fakeUCI) Import(_, text string) error {
f.imported = text
secs, err := parseSections(text)
if err != nil {
return err
}
for _, s := range secs {
sec := f.newSection(s.Type, s.Name)
for k, v := range s.Options {
sec.setOpt(k, v)
}
for k, vs := range s.Lists {
for _, v := range vs {
sec.addList(k, v)
}
}
}
return nil
}
func (f *fakeUCI) Export(string) (string, bool) {
if f.missing {
return "", false
}
var b strings.Builder
b.WriteString("package shater\n")
for _, s := range f.secs {
b.WriteString("\nconfig " + s.typ)
if s.name != "" {
b.WriteString(" '" + s.name + "'")
}
b.WriteString("\n")
for _, k := range s.oKeys {
if v, ok := s.opts[k]; ok {
b.WriteString("\toption " + k + " '" + v + "'\n")
}
}
for _, k := range s.lKeys {
for _, v := range s.lists[k] {
b.WriteString("\tlist " + k + " '" + v + "'\n")
}
}
}
return b.String(), true
}
func (f *fakeUCI) Add(_, secType string) (string, error) {
return f.newSection(secType, "").id, nil
}
func (f *fakeUCI) AddList(k, v string) error {
sel, opt := splitKey(k)
sec := f.resolve(sel)
if sec == nil {
return fmt.Errorf("uci add_list %s: no such section", k)
}
sec.addList(opt, v)
return nil
}
// ruleset returns the `config ruleset` section carrying option name == name.
func (f *fakeUCI) ruleset(name string) *fakeSection {
for _, s := range f.secs {
if s.typ == "ruleset" && s.opts["name"] == name {
return s
}
}
return nil
}
// rule returns the n-th `config rule` section.
func (f *fakeUCI) rule(idx int) *fakeSection { return f.resolve(fmt.Sprintf("@rule[%d]", idx)) }
func (f *fakeUCI) rulesetNames() []string {
var out []string
for _, s := range f.secs {
if s.typ == "ruleset" {
out = append(out, s.opts["name"])
}
}
return out
}
func eqStrings(a, b []string) bool {
if len(a) != len(b) {
return false
}
for i := range a {
if a[i] != b[i] {
return false
}
}
return true
}
func TestMigrate0to1TransformsFixture(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.kill": "open"}}
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption kill 'open'\n")
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if f.kv["shater.globals.kill_switch"] != "open" {
t.Fatalf("kill_switch = %q", f.kv["shater.globals.kill_switch"])
if v, _ := f.Get("shater.globals.kill_switch"); v != "open" {
t.Fatalf("kill_switch = %q", v)
}
if _, ok := f.kv["shater.globals.kill"]; ok {
if _, ok := f.Get("shater.globals.kill"); ok {
t.Fatal("legacy kill not removed")
}
if f.kv["shater.globals.schema_version"] != "1" {
t.Fatalf("schema_version = %q", f.kv["shater.globals.schema_version"])
if v, _ := f.Get("shater.globals.schema_version"); v != strconv.Itoa(CurrentSchemaVersion) {
t.Fatalf("schema_version = %q, want %d", v, CurrentSchemaVersion)
}
}
func TestMigrateIdempotent(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.schema_version": "1"}}
before := len(f.kv)
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption schema_version '" +
strconv.Itoa(CurrentSchemaVersion) + "'\n")
before, _ := f.Export("shater")
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if len(f.kv) != before {
t.Fatal("idempotent migrate changed state")
after, _ := f.Export("shater")
if before != after {
t.Fatalf("idempotent migrate changed state:\n--- before\n%s\n--- after\n%s", before, after)
}
}
func TestMigrateRefusesNewer(t *testing.T) {
f := &fakeUCI{kv: map[string]string{"shater.globals.schema_version": "99"}}
f := newFakeUCI("package shater\n\nconfig globals 'globals'\n\toption schema_version '99'\n")
if err := migrateWith(f); err == nil {
t.Fatal("expected refusal of newer schema")
}
}
// --- v1 -> v2: dst_domain / dst_ip fold into generated rule-sets -------------
// legacyConfig is the shape a v1 config has: rules matching destinations inline.
// The `ru-direct` rule is copied verbatim from the live router this change was
// written against (BananaWRT 25.12.1, shater 0.2.7-r1).
const legacyConfig = `package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'ru-direct'
option enabled '1'
option order '10'
list dst_domain 'suffix:ru'
list dst_domain 'suffix:yandex.net'
list dst_domain 'full:vk.com'
list dst_domain 'keyword:sberbank'
list dst_domain '.gosuslugi.ru'
list dst_domain 'plain.example'
option target 'direct'
config rule
option name 'corp nets!'
option enabled '1'
option order '20'
list dst_ip '10.0.0.0/8'
list dst_ip '192.168.44.0/24'
option target 'group:corp'
config rule
option name 'default'
option enabled '1'
option order '100'
option target 'direct'
`
func TestMigrate1to2FoldsDomainsIntoARuleset(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
rs := f.ruleset("rule-ru-direct")
if rs == nil {
t.Fatalf("no rule-ru-direct ruleset; got %v", f.rulesetNames())
}
if rs.opts["type"] != "domain" || rs.opts["source"] != "inline" {
t.Fatalf("ruleset type/source = %q/%q, want domain/inline", rs.opts["type"], rs.opts["source"])
}
// Every entry keeps its meaning. Only the BARE one is rewritten: bare means
// EXACT in a rule and SUFFIX in a rule-set, so it becomes `full:`.
want := []string{
"suffix:ru", "suffix:yandex.net", "full:vk.com",
"keyword:sberbank", ".gosuslugi.ru", "full:plain.example",
}
if !eqStrings(rs.lists["entry"], want) {
t.Fatalf("entries = %q\nwant %q", rs.lists["entry"], want)
}
r0 := f.rule(0)
if !eqStrings(r0.lists["dst_ruleset"], []string{"rule-ru-direct"}) {
t.Fatalf("dst_ruleset = %q", r0.lists["dst_ruleset"])
}
if _, ok := r0.lists["dst_domain"]; ok {
t.Fatal("dst_domain survived the migration")
}
// Order and every other option are untouched.
if r0.opts["order"] != "10" || r0.opts["target"] != "direct" || r0.opts["name"] != "ru-direct" {
t.Fatalf("rule 0 mangled: %v", r0.opts)
}
}
func TestMigrate1to2FoldsCIDRsIntoAnIPRuleset(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
// The rule name is slugged: a rule-set name becomes an engine tag.
rs := f.ruleset("rule-corp-nets-ip")
if rs == nil {
t.Fatalf("no rule-corp-nets-ip ruleset; got %v", f.rulesetNames())
}
if rs.opts["type"] != "ipcidr" {
t.Fatalf("ruleset type = %q, want ipcidr", rs.opts["type"])
}
if !eqStrings(rs.lists["entry"], []string{"10.0.0.0/8", "192.168.44.0/24"}) {
t.Fatalf("entries = %q", rs.lists["entry"])
}
r1 := f.rule(1)
if !eqStrings(r1.lists["dst_ruleset"], []string{"rule-corp-nets-ip"}) {
t.Fatalf("dst_ruleset = %q", r1.lists["dst_ruleset"])
}
if _, ok := r1.lists["dst_ip"]; ok {
t.Fatal("dst_ip survived the migration")
}
}
// A rule with no destination at all stays a catch-all — the B1 reachability
// analysis (model.RuleReachability) keys off exactly that, so the migration must
// not hand it a rule-set it never asked for.
func TestMigrate1to2LeavesCatchAllAlone(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
r2 := f.rule(2)
if len(r2.lists["dst_ruleset"]) != 0 {
t.Fatalf("catch-all gained a ruleset: %q", r2.lists["dst_ruleset"])
}
text, _ := f.Export("shater")
m, err := ParseUCIExport(text)
if err != nil {
t.Fatalf("parse migrated config: %v", err)
}
if len(m.Rules) != 3 {
t.Fatalf("rules = %d, want 3", len(m.Rules))
}
if !IsCatchAll(m.Rules[2]) {
t.Fatalf("rule %q stopped being a catch-all after the migration", m.Rules[2].Name)
}
for _, i := range []int{0, 1} {
if IsCatchAll(m.Rules[i]) {
t.Fatalf("rule %q became a catch-all — its destination was lost", m.Rules[i].Name)
}
}
reach := RuleReachability(m.Rules)
for i, v := range reach {
if v.Unreachable {
t.Fatalf("rule %d (%q) reported unreachable after the migration: %s", i, v.Name, v.Reason)
}
}
}
// Running the migration twice must not duplicate rule-sets or references. The
// second run goes through migrate1to2 directly, because migrateWith is gated by
// schema_version and would (correctly) do nothing at all.
func TestMigrate1to2IsIdempotent(t *testing.T) {
f := newFakeUCI(legacyConfig)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
first, _ := f.Export("shater")
if err := migrate1to2(f); err != nil {
t.Fatalf("second migrate1to2: %v", err)
}
second, _ := f.Export("shater")
if first != second {
t.Fatalf("re-running the migration changed the config:\n--- first\n%s\n--- second\n%s", first, second)
}
// And a full migrateWith re-run (schema already at CurrentSchemaVersion) is a
// no-op too.
if err := migrateWith(f); err != nil {
t.Fatalf("third migrate: %v", err)
}
third, _ := f.Export("shater")
if third != second {
t.Fatalf("re-running migrateWith changed the config:\n%s", third)
}
}
// An interrupted run — the rule-set was created and referenced, but the legacy
// option was not deleted yet — must reuse the rule-set rather than make a second.
func TestMigrate1to2ResumesAnInterruptedRun(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config ruleset
option name 'rule-half'
option type 'domain'
option source 'inline'
list entry 'full:a.example'
config rule
option name 'half'
list dst_domain 'a.example'
list dst_ruleset 'rule-half'
option target 'direct'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if names := f.rulesetNames(); !eqStrings(names, []string{"rule-half"}) {
t.Fatalf("rulesets = %q, want just rule-half", names)
}
r := f.rule(0)
if !eqStrings(r.lists["dst_ruleset"], []string{"rule-half"}) {
t.Fatalf("dst_ruleset = %q", r.lists["dst_ruleset"])
}
if _, ok := r.lists["dst_domain"]; ok {
t.Fatal("dst_domain survived")
}
}
// A hand-written rule-set already owning the generated name must not be
// clobbered: two `config ruleset` blocks with one name collide on the engine tag
// and one of them is dropped.
func TestMigrate1to2AvoidsNameCollisions(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config ruleset
option name 'rule-ads'
option type 'domain'
option source 'url'
option url 'https://example.invalid/list.txt'
config rule
option name 'ads'
list dst_domain 'ads.example'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
names := f.rulesetNames()
if !eqStrings(names, []string{"rule-ads", "rule-ads-2"}) {
t.Fatalf("rulesets = %q, want rule-ads + rule-ads-2", names)
}
if got := f.ruleset("rule-ads").opts["source"]; got != "url" {
t.Fatalf("the hand-written list was overwritten (source = %q)", got)
}
if !eqStrings(f.rule(0).lists["dst_ruleset"], []string{"rule-ads-2"}) {
t.Fatalf("dst_ruleset = %q", f.rule(0).lists["dst_ruleset"])
}
}
// A rule carrying BOTH lists gets both rule-sets, and keeps every entry.
func TestMigrate1to2SplitsMixedRuleIntoTwoRulesets(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'mixed'
list dst_domain 'regexp:^ads\.'
list dst_ip '203.0.113.0/24'
option target 'block'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
dom := f.ruleset("rule-mixed")
ip := f.ruleset("rule-mixed-ip")
if dom == nil || ip == nil {
t.Fatalf("want rule-mixed + rule-mixed-ip, got %q", f.rulesetNames())
}
// `regexp:` crosses over untouched (generate.peelDomainRegexes reads it).
if !eqStrings(dom.lists["entry"], []string{`regexp:^ads\.`}) {
t.Fatalf("domain entries = %q", dom.lists["entry"])
}
if !eqStrings(ip.lists["entry"], []string{"203.0.113.0/24"}) {
t.Fatalf("ip entries = %q", ip.lists["entry"])
}
if !eqStrings(f.rule(0).lists["dst_ruleset"], []string{"rule-mixed", "rule-mixed-ip"}) {
t.Fatalf("dst_ruleset = %q", f.rule(0).lists["dst_ruleset"])
}
}
// An inert `geosite:` matcher is copied verbatim rather than promoted to a
// source=geosite list: it matched nothing before the upgrade (the engine's
// route-rule geosite field is gone) and must not start routing traffic because of
// one. The text is kept so the operator can see and convert it.
func TestMigrate1to2KeepsGeoMarkersInert(t *testing.T) {
f := newFakeUCI(`package shater
config globals 'globals'
option schema_version '1'
config rule
option name 'geo'
list dst_domain 'geosite:youtube'
list dst_ip 'geoip:ru'
option target 'direct'
`)
if err := migrateWith(f); err != nil {
t.Fatalf("migrate: %v", err)
}
if got := f.ruleset("rule-geo").lists["entry"]; !eqStrings(got, []string{"geosite:youtube"}) {
t.Fatalf("domain entries = %q", got)
}
if got := f.ruleset("rule-geo-ip").lists["entry"]; !eqStrings(got, []string{"geoip:ru"}) {
t.Fatalf("ip entries = %q", got)
}
for _, s := range f.secs {
if s.typ == "ruleset" && s.opts["source"] != "inline" {
t.Fatalf("ruleset %q got source %q — a geo marker was promoted", s.opts["name"], s.opts["source"])
}
}
}
// A fresh install has no config to export; the migration must succeed silently
// rather than refuse to boot.
func TestMigrate1to2NoConfig(t *testing.T) {
f := &fakeUCI{missing: true}
if err := migrate1to2(f); err != nil {
t.Fatalf("migrate on an absent package: %v", err)
}
}
func TestRulesetBaseName(t *testing.T) {
cases := []struct{ in, want string }{
{"ru-direct", "rule-ru-direct"},
{"corp nets!", "rule-corp-nets"},
{" spaced name ", "rule-spaced-name"},
{"Ünïcode", "rule-n-code"}, // non-ASCII is not tag-safe; dropped, leaving a separator
{"", "rule-7"},
{"!!!", "rule-7"},
}
for _, c := range cases {
if got := rulesetBaseName(c.in, 7); got != c.want {
t.Errorf("rulesetBaseName(%q) = %q, want %q", c.in, got, c.want)
}
}
}
func TestMigrateDomainEntry(t *testing.T) {
cases := []struct{ in, want string }{
{"example.com", "full:example.com"}, // bare: exact in a rule, suffix in a list
{"full:example.com", "full:example.com"},
{"suffix:example.com", "suffix:example.com"},
{"keyword:ads", "keyword:ads"},
{`regexp:^a\.b$`, `regexp:^a\.b$`},
{".example.com", ".example.com"},
{"geosite:youtube", "geosite:youtube"},
{" spaced.example ", "full:spaced.example"},
{" ", ""},
}
for _, c := range cases {
if got := migrateDomainEntry(c.in); got != c.want {
t.Errorf("migrateDomainEntry(%q) = %q, want %q", c.in, got, c.want)
}
}
}
+10 -2
View File
@@ -633,14 +633,22 @@ type Ruleset struct {
}
// Rule is a `config rule` (ordered, first-match).
//
// DESTINATION IS ALWAYS A RULE-SET. A rule names WHERE traffic is going only
// through DstRuleset — there are no inline domain or IP lists on a rule any
// more (`dst_domain` / `dst_ip` were removed in schema v2; migrate1to2 folds
// every existing one into a generated `config ruleset` and rewrites the rule to
// point at it). One destination mechanism means one set of matcher semantics to
// learn, one place a list is edited, and a list that is compiled once into a
// .srs and shared by every rule that references it instead of being re-parsed
// per rule. Src (the CLIENT side), DstPort and Proto are unaffected: they are
// not lists of destinations and have no rule-set form.
type Rule struct {
Name string
Enabled bool
Order int
Src []string
DstDomain []string
DstRuleset []string
DstIP []string
DstPort string
Proto string
Target string // chain:|group:|node:|direct|block
+199
View File
@@ -0,0 +1,199 @@
package model
// Rule reachability — "this rule can never fire", computed from the desired state
// alone.
//
// WHY THIS EXISTS. A routing rule with no conditions at all (no src, no dst
// domain/ip/ruleset, no port, no proto) is not a rule in the ordinary sense: it
// states a policy for EVERYTHING. generate does not emit it as a match-all route
// rule — it points the engine's route `Final` at that rule's target instead (see
// generate/route.go buildRoute). Two consequences follow, and neither is visible
// anywhere in the UI:
//
// 1. A condition-less rule can NEVER shadow a rule that has conditions. Every
// conditional rule is emitted ahead of `Final`, whatever its Order. So a
// "default -> direct" at order 20 does not hijack a specific rule at order 500.
// 2. Two condition-less rules DO shadow each other, and the LAST one in
// (Order, position) order wins, because the loop simply overwrites `Final`.
// Every earlier one is a dead setting that reads as a live one — the panel
// drew both with the same "default route · final" badge.
//
// A real config in the field had exactly that: two rules both named `default`,
// both with zero conditions, order 20 -> direct and order 100 -> group:auto. One
// of the two was doing nothing, and nothing said which.
//
// SCOPE, DELIBERATELY NARROW. This reports only the shadowing that is certain
// from the config: condition-less rule over condition-less rule. It does NOT try
// to decide whether one conditional rule's matcher set subsumes another's (is
// `dst_ruleset ru-inside` a superset of `dst_ip 5.0.0.0/8`? only the compiled
// rule-set knows) — a false "never fires" badge on a working rule would be worse
// than no badge at all.
//
// It lives in model, the stdlib-only leaf, because both consumers must agree:
// generate turns the verdict into an apply warning, and the panel API serves it
// to the Routing page. A second, drifting implementation in either is how the
// warning and the badge end up disagreeing about the same config.
import (
"fmt"
"sort"
"strings"
)
// IsCatchAll reports whether a rule carries NO matcher of any kind (src,
// dst_ruleset, port, proto). Such a rule is the default egress: generate points
// route `Final` at it rather than emitting a match-all rule.
//
// The destination side is now exactly one field — a rule names where traffic is
// going through DstRuleset alone (schema v2; see model.Rule). A rule that was
// catch-all before the migration is still catch-all after it, and a rule that
// carried `dst_domain`/`dst_ip` is not, because the migration gives it a
// DstRuleset in their place.
//
// generate.isCatchAll and the panel's isCatchAll() are the same predicate; this
// is the one the Go side shares.
func IsCatchAll(r Rule) bool {
return len(r.Src) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
}
// EffectiveRuleTarget resolves what a rule actually routes to: Target wins, and a
// rule with NO Target but a bare Egress routes to that egress outbound. "" means
// the rule states no policy at all — generate drops such a rule entirely (it does
// NOT mean "direct"), so it never becomes the default and never shadows anything.
func EffectiveRuleTarget(r Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
// SortedRuleIndices returns rule indices sorted by (Order, original index), so the
// engine's first-match order matches the model's declared order, stably. Equal
// Order keeps the configured/UCI order — the panel promises "lower Order runs
// earlier", and nothing below it may reshuffle ties.
func SortedRuleIndices(rules []Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
}
// RuleReach is one rule's reachability verdict. There is exactly one per input
// rule, at the same index, so a consumer can zip the two slices.
//
// It is the routing-rule analogue of engine.ChainHealth.Used / GroupHealth.Used —
// a quiet note about the ROUTING CONFIG, never a health or liveness signal.
type RuleReach struct {
// Index is the rule's position in the model's Rules slice (NOT the sorted
// order), so the panel can match a verdict to the row it drew.
Index int `json:"index"`
// Name and Order are echoed so a consumer holding a possibly-stale config can
// verify the verdict still describes the rule it is about to badge, instead of
// accusing the wrong row.
Name string `json:"name"`
Order int `json:"order"`
// Unreachable is true when this rule can never take effect, whatever the
// traffic. false is the normal case and carries no claim that the rule ever
// actually matches something — only that nothing in the config stops it.
Unreachable bool `json:"unreachable"`
// ShadowedBy / ShadowedByOrder name the rule that supersedes this one; empty /
// zero when Unreachable is false. ShadowedByIndex is -1 when there is none.
ShadowedBy string `json:"shadowed_by,omitempty"`
ShadowedByIndex int `json:"shadowed_by_index"`
ShadowedByOrder int `json:"shadowed_by_order,omitempty"`
// Reason is the operator-facing sentence; "" when Unreachable is false.
Reason string `json:"reason,omitempty"`
}
// RuleReachability returns one verdict per rule, in the INPUT slice's order.
//
// The input must already be the EFFECTIVE rule set — profile enable/disable
// applied (see Model.EffectiveRules). A rule the active profile switched off is
// not in force, and a rule it switched ON is, whatever /etc/config/shater says;
// judging the raw slice would badge the wrong rules on a router that uses
// profiles at all.
//
// A rule takes part in the analysis only when it is Enabled, condition-less, and
// carries a target — those are exactly the rules generate lets set `Final`.
//
// SCHEDULES. A rule that only holds inside a time window never marks anything
// permanently dead: outside its window the rule below it is the default again. So
// a SCHEDULED catch-all is skipped as a shadower (but can still BE shadowed — an
// unscheduled catch-all after it wins at every hour of every day, which makes the
// schedule pure decoration and is worth saying out loud).
func RuleReachability(rules []Rule) []RuleReach {
out := make([]RuleReach, len(rules))
for i := range rules {
out[i] = RuleReach{
Index: i, Name: rules[i].Name, Order: rules[i].Order, ShadowedByIndex: -1,
}
}
// Walk the declared order BACKWARDS and remember the last unconditional
// default seen. When we reach a rule, `winner` holds the default that is in
// force below it — i.e. the one whose target the engine actually uses.
winner := -1
for _, i := range reverse(SortedRuleIndices(rules)) {
r := rules[i]
if !r.Enabled || !IsCatchAll(r) || EffectiveRuleTarget(r) == "" {
continue
}
if winner >= 0 {
w := rules[winner]
out[i].Unreachable = true
out[i].ShadowedBy = w.Name
out[i].ShadowedByIndex = winner
out[i].ShadowedByOrder = w.Order
out[i].Reason = fmt.Sprintf(
"this rule has no conditions, so it sets the default for all traffic — but rule %q "+
"(order %d) has none either and comes after it, so %q is the default the router "+
"uses and this rule's target %q is never applied",
w.Name, w.Order, EffectiveRuleTarget(w), EffectiveRuleTarget(r))
}
// Only an UNSCHEDULED default holds at every hour, so only it can retire the
// rules above it. Keep the LAST one (highest Order): that is the one whose
// target the engine ends up with.
if !r.SchedEnabled && winner < 0 {
winner = i
}
}
return out
}
// reverse returns idx walked back to front. Small enough to inline by hand, but
// naming it keeps the loop above readable as "walk the declared order backwards".
func reverse(idx []int) []int {
out := make([]int, len(idx))
for i, v := range idx {
out[len(idx)-1-i] = v
}
return out
}
// EffectiveRules returns the rule set that is actually in force: a COPY of m.Rules
// with the active profile's enable/disable overrides applied, plus whatever
// warnings that resolution produced. m.Rules is never mutated.
//
// It is the shared front door for every consumer that must reason about "which
// rules are in force right now" without building an engine config — the panel's
// reachability endpoint today. generate keeps its own call site because it also
// needs the resolved *Profile itself (endpoint-resolver override); both go through
// ResolveActiveProfile + ApplyProfileRuleOverrides, so they cannot disagree.
func (m *Model) EffectiveRules() ([]Rule, []Warning) {
if m == nil {
return nil, nil
}
prof, warns := ResolveActiveProfile(m)
rules, owarns := ApplyProfileRuleOverrides(m.Rules, prof)
return rules, append(warns, owarns...)
}
+313
View File
@@ -0,0 +1,313 @@
package model
import "testing"
// catchAll builds a condition-less rule — the shape that becomes the engine's
// route Final.
func catchAll(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target}
}
// specific builds a rule with one matcher, so it is emitted as a real route rule
// ahead of Final and can never be retired by a default.
func specific(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target, DstPort: "443"}
}
// verdicts indexes a reachability run by rule index, asserting the contract that
// there is exactly one verdict per rule, at the same index.
func verdicts(t *testing.T, rules []Rule) []RuleReach {
t.Helper()
out := RuleReachability(rules)
if len(out) != len(rules) {
t.Fatalf("want %d verdicts, got %d", len(rules), len(out))
}
for i := range out {
if out[i].Index != i || out[i].Name != rules[i].Name {
t.Fatalf("verdict %d is about %q (index %d), want %q", i, out[i].Name, out[i].Index, rules[i].Name)
}
}
return out
}
func wantReachable(t *testing.T, v RuleReach) {
t.Helper()
if v.Unreachable {
t.Fatalf("rule %q (order %d) must be reachable, got shadowed by %q: %s",
v.Name, v.Order, v.ShadowedBy, v.Reason)
}
if v.ShadowedBy != "" || v.ShadowedByIndex != -1 || v.Reason != "" {
t.Fatalf("rule %q is reachable but carries shadow details: by=%q idx=%d reason=%q",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.Reason)
}
}
func wantShadowed(t *testing.T, v RuleReach, by string, byIndex, byOrder int) {
t.Helper()
if !v.Unreachable {
t.Fatalf("rule %q (order %d) must be unreachable, shadowed by %q", v.Name, v.Order, by)
}
if v.ShadowedBy != by || v.ShadowedByIndex != byIndex || v.ShadowedByOrder != byOrder {
t.Fatalf("rule %q: shadowed by %q(idx %d, order %d), want %q(idx %d, order %d)",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.ShadowedByOrder, by, byIndex, byOrder)
}
if v.Reason == "" {
t.Fatalf("rule %q is unreachable but carries no reason", v.Name)
}
}
// TestReachabilitySingleCatchAll: the ordinary config — some specific rules and
// ONE default. Nothing is retired, including the default itself.
func TestReachabilitySingleCatchAll(t *testing.T) {
rules := []Rule{
specific("block-ads", 10, "block"),
specific("ru-bypass", 20, "direct"),
catchAll("default-tunnel", 900, "group:auto"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityTwoCatchAllsFieldConfig is the config that prompted this
// analysis, reproduced exactly: two rules BOTH named `default`, both with zero
// conditions, order 20 -> direct and order 100 -> group:auto.
//
// The later one is the default the engine uses (buildRoute overwrites Final as it
// walks ascending Order), so it is the order-20 `direct` rule that is dead — the
// opposite of what "first match wins" would suggest, which is exactly why this
// needed saying out loud.
func TestReachabilityTwoCatchAllsFieldConfig(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
catchAll("default", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "default", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityCatchAllBelowSpecificRules: a default sorted BELOW specific
// rules retires none of them. A condition-less rule becomes route Final, which is
// evaluated after every emitted rule whatever its Order — so even a specific rule
// ordered after the default still matches first.
func TestReachabilityCatchAllBelowSpecificRules(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
specific("work-vpn", 100, "group:work"),
specific("ads", 200, "block"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityThreeCatchAlls: only the last default survives; every earlier
// one is blamed on that same last one, not on its immediate successor — that is
// the rule whose target the router actually ends up using.
func TestReachabilityThreeCatchAlls(t *testing.T) {
rules := []Rule{
catchAll("a", 10, "direct"),
catchAll("b", 20, "block"),
catchAll("c", 30, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "c", 2, 30)
wantShadowed(t, v[1], "c", 2, 30)
wantReachable(t, v[2])
}
// TestReachabilityDisabledCatchAllIsInert: a disabled default neither dies nor
// kills. It installs nothing, so badging it "never fires" would be noise, and it
// must not be credited with retiring the live default above it.
func TestReachabilityDisabledCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "parked", Enabled: false, Order: 100, Target: "direct"},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityTargetlessCatchAllIsInert: a rule with neither Target nor
// Egress states no policy, so generate drops it (it does NOT mean "direct"). It
// never becomes Final, so it cannot retire the default above it — and is not
// itself reported here, because "no target set" is its own, more useful warning.
func TestReachabilityTargetlessCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "empty", Enabled: true, Order: 100},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityBareEgressCountsAsTarget: a rule carrying only `option egress
// wan2` routes to that egress, so it IS a default and does retire the one above.
func TestReachabilityBareEgressCountsAsTarget(t *testing.T) {
rules := []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "wan2", Enabled: true, Order: 100, Egress: "wan2"},
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "wan2", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityScheduledCatchAllNeverRetires: a default that only holds inside
// a time window leaves the one above it live for the rest of the day, so it
// retires nothing.
func TestReachabilityScheduledCatchAllNeverRetires(t *testing.T) {
rules := []Rule{
catchAll("all-day", 20, "group:auto"),
{Name: "nightly", Enabled: true, Order: 100, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityScheduledCatchAllCanBeRetired: the converse — an unscheduled
// default AFTER a scheduled one wins at every hour of every day, so the schedule
// is pure decoration and the scheduled rule is dead.
func TestReachabilityScheduledCatchAllCanBeRetired(t *testing.T) {
rules := []Rule{
{Name: "nightly", Enabled: true, Order: 20, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
catchAll("all-day", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "all-day", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityEqualOrderKeepsConfiguredOrder: ties break by position in the
// slice (SortedRuleIndices is stable), so with equal Order the SECOND section in
// /etc/config/shater is the one that wins — same as generate.
func TestReachabilityEqualOrderKeepsConfiguredOrder(t *testing.T) {
rules := []Rule{
catchAll("first", 50, "direct"),
catchAll("second", 50, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "second", 1, 50)
wantReachable(t, v[1])
}
// TestReachabilityUnsortedInputIsJudgedByOrder: the model slice arrives in UCI
// order, which need not be Order order. The verdict must follow Order, and the
// returned slice must stay aligned with the INPUT indices.
func TestReachabilityUnsortedInputIsJudgedByOrder(t *testing.T) {
rules := []Rule{
catchAll("late", 100, "group:auto"),
catchAll("early", 20, "direct"),
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantShadowed(t, v[1], "late", 0, 100)
}
// TestReachabilityMatcherKindsAreNotCatchAll: every matcher field on its own is
// enough to make a rule conditional, so none of these is retired by the default
// below. A field this misses would silently badge a working rule "never fires".
func TestReachabilityMatcherKindsAreNotCatchAll(t *testing.T) {
conditional := []Rule{
{Name: "by-src", Enabled: true, Order: 10, Target: "direct", Src: []string{"192.168.1.0/24"}},
// The destination side is one field now (schema v2): a domain list and an
// address list are both `config ruleset`s a rule points dst_ruleset at.
{Name: "by-ruleset", Enabled: true, Order: 13, Target: "direct", DstRuleset: []string{"ads"}},
{Name: "by-port", Enabled: true, Order: 14, Target: "direct", DstPort: "443"},
{Name: "by-proto", Enabled: true, Order: 15, Target: "direct", Proto: "quic"},
}
for _, r := range conditional {
if IsCatchAll(r) {
t.Fatalf("rule %q carries a matcher but reads as a catch-all", r.Name)
}
}
rules := append(append([]Rule(nil), conditional...), catchAll("default", 900, "group:auto"))
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityWhitespaceOnlyMatcherIsCatchAll: a port/proto of spaces is not
// a matcher. generate trims before deciding, so this analysis must too — otherwise
// a rule the engine treats as the default reads as conditional here and its
// shadowing goes unreported.
func TestReachabilityWhitespaceOnlyMatcherIsCatchAll(t *testing.T) {
blank := Rule{Name: "blank", Enabled: true, Order: 20, Target: "direct", DstPort: " ", Proto: "\t"}
if !IsCatchAll(blank) {
t.Fatal("a rule whose only matchers are whitespace must read as a catch-all")
}
v := verdicts(t, []Rule{blank, catchAll("default", 100, "group:auto")})
wantShadowed(t, v[0], "default", 1, 100)
}
// TestReachabilityEmptyAndNil: no rules, no verdicts, no panic.
func TestReachabilityEmptyAndNil(t *testing.T) {
if got := RuleReachability(nil); len(got) != 0 {
t.Fatalf("nil rules: want no verdicts, got %d", len(got))
}
if got := RuleReachability([]Rule{}); len(got) != 0 {
t.Fatalf("empty rules: want no verdicts, got %d", len(got))
}
}
// TestEffectiveRulesAppliesActiveProfile: the analysis must judge the rules the
// active profile leaves in force. Here the profile DISABLES the later default, so
// the earlier one is live and nothing is retired — judging the raw slice would
// have badged the wrong rule.
func TestEffectiveRulesAppliesActiveProfile(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "home"},
Profiles: []Profile{
{Name: "home", Enabled: true, DisableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
catchAll("fallback", 100, "direct"),
},
}
eff, _ := m.EffectiveRules()
if eff[1].Enabled {
t.Fatal("active profile must have disabled the fallback rule")
}
if !m.Rules[1].Enabled {
t.Fatal("EffectiveRules must not mutate the model's own rules")
}
for _, v := range verdicts(t, eff) {
wantReachable(t, v)
}
// Same model, profile off: the later default is back in force and retires the
// tunnel default above it.
m.Globals.ActiveProfile = ""
m.Profiles[0].Enabled = false
eff, _ = m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
// TestEffectiveRulesProfileCanReviveADefault: a profile that force-ENABLES a
// disabled default makes it in force, and it then retires the default above it.
// The raw config would show it disabled and report nothing.
func TestEffectiveRulesProfileCanReviveADefault(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "travel"},
Profiles: []Profile{
{Name: "travel", Enabled: true, EnableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "fallback", Enabled: false, Order: 100, Target: "direct"},
},
}
eff, _ := m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
+4 -2
View File
@@ -211,9 +211,11 @@ func RenderUCIExport(m *Model) string {
w.boolOpt("enabled", r.Enabled)
w.intOpt("order", r.Order)
w.listOpt("src", r.Src)
w.listOpt("dst_domain", r.DstDomain)
// dst_domain / dst_ip are NOT emitted (removed in schema v2). Their absence
// here is also how a legacy option drains out of a config that was migrated:
// migrate1to2 deletes them explicitly, and any that survived a hand-edit
// disappear the next time the panel writes the model back.
w.listOpt("dst_ruleset", r.DstRuleset)
w.listOpt("dst_ip", r.DstIP)
w.strOpt("dst_port", r.DstPort)
w.strOpt("proto", r.Proto)
w.strOpt("target", r.Target)
+2 -3
View File
@@ -84,8 +84,7 @@ func richModel() *Model {
}},
Rules: []Rule{{
Name: "pc", Enabled: true, Order: 10,
Src: []string{"192.168.1.1/32"}, DstDomain: []string{"geosite:telegram"},
DstRuleset: []string{"ads"}, DstIP: []string{"1.1.1.1/32"},
Src: []string{"192.168.1.1/32"}, DstRuleset: []string{"ads"},
DstPort: "443", Proto: "tcp,udp", Target: "chain:triple", Egress: "frag",
Kill: "default", SchedEnabled: true, SchedDays: []string{"mon", "tue"},
SchedStart: "08:00", SchedEnd: "22:00", SchedUTCOffset: 180,
@@ -384,7 +383,7 @@ func TestRenderSkipsSubCacheNodes(t *testing.T) {
// and COMMITs — and the imported text re-parses to the original Model.
func TestWriteUCIReplaces(t *testing.T) {
m := richModel()
f := &fakeUCI{kv: map[string]string{}}
f := newFakeUCI("")
if err := writeUCIWith(f, m); err != nil {
t.Fatalf("writeUCIWith: %v", err)
}
+11 -6
View File
@@ -175,13 +175,18 @@ func ParseUCIExport(text string) (*Model, error) {
})
case "rule":
m.Rules = append(m.Rules, Rule{
Name: firstNonEmpty(s.opt("name"), s.Name),
Enabled: s.optBool("enabled", true),
Order: parseInt(s.opt("order"), 0),
Src: s.list("src"),
DstDomain: s.list("dst_domain"),
Name: firstNonEmpty(s.opt("name"), s.Name),
Enabled: s.optBool("enabled", true),
Order: parseInt(s.opt("order"), 0),
Src: s.list("src"),
// No dst_domain / dst_ip: removed in schema v2. A rule's destination
// is a rule-set reference and nothing else; migrate1to2 (migrate.go)
// converts any legacy inline list into a `config ruleset` and points
// dst_ruleset at it, so by the time this parser runs there is nothing
// left to read. A config that somehow still carries them (hand-edited
// after a downgrade) simply ignores them — the migration is re-run on
// every load, so it will have been rewritten first.
DstRuleset: s.list("dst_ruleset"),
DstIP: s.list("dst_ip"),
DstPort: s.opt("dst_port"),
Proto: s.opt("proto"),
Target: s.opt("target"),
+1 -1
View File
@@ -59,7 +59,7 @@ config rule
option enabled '1'
option order '10'
list src '192.168.11.14/32'
list dst_domain 'geosite:telegram'
list dst_ruleset 'ads'
option dst_port '443'
option proto 'tcp,udp'
option target 'chain:triple'
+41 -5
View File
@@ -31,6 +31,12 @@ func ApplyNft(ruleset string) error {
if err := runNftStdin(ruleset, "-f", "-"); err != nil {
return fmt.Errorf("nft -f (load) failed: %w", err)
}
// The divert plane just changed underneath every flow that is currently
// tracked. For UDP :53 that is not self-correcting — see conntrack.go — so
// drop those entries and let the clients re-derive their path through the
// ruleset that is now loaded. Best-effort: a failed flush leaves exactly the
// behaviour we had before and must never fail an otherwise-good load.
_, _ = FlushDNSConntrack()
return nil
}
@@ -84,6 +90,11 @@ func TeardownNft() error {
if out, err := execCommand("nft", "delete", "table", "inet", "shater").CombinedOutput(); err != nil {
return fmt.Errorf("nft delete table inet shater: %v\n%s", err, out)
}
// Same reasoning as ApplyNft, mirrored: every DNS flow that was being
// delivered through the tproxy socket has just lost the rule that put it
// there. Without this, those entries survive into whatever plane comes next
// (including "no plane at all") and keep pointing at a socket that is gone.
_, _ = FlushDNSConntrack()
return nil
}
@@ -366,25 +377,50 @@ func isPointToPoint(dev string) bool {
return strings.Contains(string(out), "POINTOPOINT")
}
// RoutingPresent reports whether our fwmark ip rule is currently installed, for
// RoutingPresent reports whether our policy routing is currently installed, for
// EVERY family the model asks for. With Globals.IPv6 on, ApplyRouting installs
// both a -4 and a -6 rule, so checking only -4 was a half-truth: `ip -6 rule` is
// both a -4 and a -6 half, so checking only -4 was a half-truth: `ip -6 rule` is
// flushed independently (a `network reload` / `ifup` can drop one family and not
// the other), and the caller's idempotent fast-path would then conclude the plane
// was intact and never restore the missing v6 rule — leaving v6 clients diverted
// by nft but with nowhere to be delivered locally.
//
// With Globals.IPv6 off no v6 rule is installed BY DESIGN, so its absence must
// not be read as a missing plane; only the -4 rule is required then.
// not be read as a missing plane; only the -4 half is required then.
//
// # Both halves are checked, not just the rule
//
// ApplyRouting installs TWO things per family — the `fwmark -> table` rule AND
// the `local default dev lo` route inside that table — and they are removed by
// two INDEPENDENT commands (`ip rule del`, `ip route flush table`). Checking only
// the rule made the second half invisible: a table that had been flushed while
// its rule survived (a teardown interrupted part-way, an `ip route flush` from
// any other actor) read as "plane intact", so applyLocked's fast-path skipped
// ApplyRouting forever and NOTHING re-created the route. The visible result is a
// box that reports plane=full / engine_running=true while diverted packets are
// marked, find an empty table, fall through to the main table and are handed to
// the fail-closed forward drop — a permanent, healthy-looking outage that a
// reconcile cannot repair, because a reconcile is exactly what consults this
// function. A presence check must cover everything its Apply counterpart
// installs, or the idempotent fast-path becomes a trap.
func RoutingPresent(g model.Globals) bool {
want := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
wantRule := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
table := fmt.Sprintf("%d", effTable(g))
fams := []string{"-4"}
if g.IPv6 {
fams = append(fams, "-6")
}
for _, fam := range fams {
out, err := execCommand("ip", fam, "rule", "show").Output()
if err != nil || !strings.Contains(string(out), want) {
if err != nil || !strings.Contains(string(out), wantRule) {
return false
}
// `ip -4 route show table N` prints "local default dev lo scope host";
// the v6 form is "local default dev lo metric 1024 pref medium". Matching
// the route TYPE + destination covers both without pinning the trailing
// attributes, which differ by family and iproute2 version.
rout, rerr := execCommand("ip", fam, "route", "show", "table", table).Output()
if rerr != nil || !strings.Contains(string(rout), "local default") {
return false
}
}
+78
View File
@@ -0,0 +1,78 @@
package netplane
// Conntrack maintenance for plane transitions.
//
// # Why the data plane has to touch conntrack at all
//
// Everything else in this package is *stateless* from the kernel's point of view:
// an nft ruleset, a couple of `ip rule`/`ip route` entries and a handful of
// sysctls. Rebuilding them is atomic and idempotent, so a rebuilt plane is
// indistinguishable from a freshly-installed one — EXCEPT for one thing the
// rebuild cannot reach: the connection-tracking entries that were established
// while the previous plane (or a half-removed one) was in force.
//
// That matters here because our divert is a TPROXY divert. A `tproxy` statement
// hands the packet to a LOCAL TRANSPARENT SOCKET; when that socket is gone the
// statement evaluates to NFT_BREAK and the packet takes a completely different
// path through the ruleset. A restart necessarily walks through such a window:
// the outgoing daemon closes its engine BEFORE it removes the table (Teardown
// order — deliberately, because the reverse order would open a plaintext leak),
// and the incoming daemon starts its engine BEFORE it loads the new table. Any
// flow that crosses one of those windows keeps a conntrack entry that was formed
// against a plane that no longer exists, and — for UDP, which has no handshake to
// resynchronise on — every retry merely refreshes that entry instead of
// re-deriving the path.
//
// # What this is NOT
//
// This flush is HYGIENE, not the cure for the "DNS to the router's own LAN
// address never comes back after a restart" report (B3). That turned out to be a
// socket-level collision: the engine's TPROXY UDP write-back sockets were bound
// UNCONNECTED to the original destination — the router's own LAN address :53 —
// and so joined the kernel's demultiplex set next to dnsmasq's socket, silently
// swallowing the host's own queries. The fix for that lives in the engine
// (protocol/redirect/tproxy.go, lx:tproxy_writeback_connect); the stale
// `[UNREPLIED]` conntrack entry seen on the live box was a CONSEQUENCE of the
// unanswered query, not its cause.
//
// It is kept because it is independently correct and costs one netlink
// round-trip per plane change: a DNS flow that was mid-flight across a plane
// rebuild has a conntrack entry describing a delivery path that no longer
// exists, and dropping it makes the first query after a restart re-derive its
// path immediately instead of waiting out a retry.
//
// # Scope: :53/UDP only, deliberately
//
// A blanket `conntrack -F` would also delete the entries behind the operator's
// SSH session, the LuCI session and the admin panel — fw4's input chain accepts
// them via `ct state established,related`, so dropping their conntrack entries
// drops the sessions. Locking the admin out of the box while "fixing" DNS is not
// a trade we get to make. UDP/:53 is the narrowest cut that covers the observed
// failure class: DNS is retried by every client within a second, so deleting its
// entries costs nothing and is invisible.
//
// Best-effort by contract: a kernel without conntrack, a netlink permission
// error or a non-Linux build all report zero deletions and no error path that can
// fail an apply. Losing the flush degrades to the old behaviour; it must never
// take a working plane down.
// dnsPort is the only port whose conntrack entries we touch. See the package
// comment for why this is deliberately not "everything".
const dnsPort uint16 = 53
// flushUDPPortConntrack is the platform seam. The default is a no-op so the
// package builds (and `go vet`s) on non-Linux dev hosts; conntrack_linux.go's
// init() replaces it with the real ctnetlink delete on the router target. Tests
// substitute a recorder.
var flushUDPPortConntrack = func(port uint16) (int, error) { return 0, nil }
// FlushDNSConntrack deletes every UDP connection-tracking entry whose ORIGINAL
// destination port is 53, for both address families, and returns how many were
// removed.
//
// Called on every plane transition that can change where a DNS packet is
// delivered: after a ruleset is loaded (ApplyNft — the full plane, the holding
// plane and a rollback all go through it) and after the table is removed
// (TeardownNft). Idempotent and cheap: on an idle box the DNS entry count is a
// handful, and on a busy one it is bounded by the number of clients.
func FlushDNSConntrack() (int, error) { return flushUDPPortConntrack(dnsPort) }
+62
View File
@@ -0,0 +1,62 @@
//go:build linux
package netplane
// The real ctnetlink implementation of the conntrack seam declared in
// conntrack.go. It lives behind a build tag for the same reason apply's flock
// does: the netlink conntrack API only exists on Linux, and the control plane
// must still build and test on a developer's Windows/macOS host.
//
// Deliberately netlink and NOT the `conntrack` CLI: conntrack-tools is not a
// dependency of shater-core (and pulling it in for one call would add ~100 KiB
// of userland to a flash-constrained router), so a shell-out would silently
// no-op on every real box — the worst possible outcome for a fix whose entire
// job is to remove stale state.
import (
"github.com/sagernet/netlink"
"golang.org/x/sys/unix"
)
func init() { flushUDPPortConntrack = ctnetlinkFlushUDPPort }
// ctnetlinkFlushUDPPort deletes the UDP conntrack entries whose ORIGINAL
// destination port is `port`, in both families, and returns the total deleted.
//
// Errors are aggregated rather than short-circuited: v4 and v6 are independent
// tables and a failure on one must not hide a successful cleanup of the other.
// The first error is returned for logging; the count is still accurate for the
// families that succeeded.
func ctnetlinkFlushUDPPort(port uint16) (int, error) {
var (
total int
firstErr error
)
for _, family := range []netlink.InetFamily{
netlink.InetFamily(unix.AF_INET),
netlink.InetFamily(unix.AF_INET6),
} {
filter := &netlink.ConntrackFilter{}
// Protocol MUST be set before the port: AddPort refuses to add a port
// filter while the layer-4 protocol is unknown (a port means nothing
// without one), so the order here is load-bearing.
if err := filter.AddProtocol(unix.IPPROTO_UDP); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
if err := filter.AddPort(netlink.ConntrackOrigDstPort, port); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
n, err := netlink.ConntrackDeleteFilter(netlink.ConntrackTable, family, filter)
total += int(n)
if err != nil && firstErr == nil {
firstErr = err
}
}
return total, firstErr
}
+7
View File
@@ -449,6 +449,13 @@ func TestRoutingPresentBothFamilies(t *testing.T) {
out = c.v6
}
}
// RoutingPresent also verifies the `local default dev lo` route
// that ApplyRouting installs beside the rule (see
// TestRoutingPresentRequiresLocalDefaultRoute). This case set is
// about the RULE half, so the route half is always healthy here.
if name == "ip" && len(arg) >= 3 && arg[1] == "route" {
out = "local default dev lo scope host"
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
+210
View File
@@ -0,0 +1,210 @@
package netplane
// B3 regressions: what a restart leaves behind.
//
// The bug these pin: `/etc/init.d/shater restart` (stop immediately followed by
// start) left DNS to the router's own LAN address permanently dead, while
// `stop` + pause + `start` was fine and the status kept reporting plane=full /
// engine_running=true. Two independent defects fed it, and both live here:
//
// 1. A plane transition (load or teardown of the tproxy divert) left the
// CONNTRACK entries of flows that had crossed the transition pointing at a
// plane that no longer exists. Nothing in the tree touched conntrack, so a
// UDP flow — which has no handshake to resynchronise on and whose entry is
// refreshed by every retry — stayed wedged indefinitely.
// 2. RoutingPresent() reported "plane intact" from the ip RULE alone, ignoring
// the `local default dev lo` ROUTE that ApplyRouting installs alongside it.
// A teardown interrupted between the two (procd SIGKILL at term_timeout) or
// any other `ip route flush` therefore became invisible: applyLocked's
// idempotent fast-path skipped ApplyRouting forever and no reconcile could
// repair it.
import (
"os"
"os/exec"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// recordFlush swaps the conntrack seam for a counter and restores it after the
// test. Returns a pointer to the number of calls and the ports asked for.
func recordFlush(t *testing.T) *[]uint16 {
t.Helper()
orig := flushUDPPortConntrack
t.Cleanup(func() { flushUDPPortConntrack = orig })
var seen []uint16
flushUDPPortConntrack = func(port uint16) (int, error) {
seen = append(seen, port)
return 0, nil
}
return &seen
}
// TestApplyNftFlushesDNSConntrack: loading a ruleset must drop the DNS conntrack
// entries formed against the previous plane. Without this, a flow that crossed
// the restart window keeps being delivered by a rule set that is gone.
func TestApplyNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err != nil {
t.Fatalf("ApplyNft: %v", err)
}
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("ApplyNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
// Ordering matters: the flush is only meaningful once the NEW ruleset is in
// the kernel, otherwise the very next packet re-creates the entry against the
// old plane. The load is the last nft invocation before it.
if len(rec) != 2 || !strings.Contains(rec[0], "-c") || strings.Contains(rec[1], "-c") {
t.Fatalf("expected validate-then-load, got %v", rec)
}
}
// TestApplyNftDoesNotFlushOnFailure: a ruleset that does not load leaves the
// PREVIOUS plane in charge (nft -f is one netlink transaction). Flushing then
// would tear down live flows for nothing.
func TestApplyNftDoesNotFlushOnFailure(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err == nil {
t.Fatalf("ApplyNft must report the nft failure")
}
if len(*seen) != 0 {
t.Fatalf("a failed load must not flush conntrack, got %v", *seen)
}
}
// TestTeardownNftFlushesDNSConntrack is the mirror: removing the table strands
// every DNS flow that was being delivered through the tproxy socket, so those
// entries must go with it.
func TestTeardownNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft: %v", err)
}
// fakeExec makes every command succeed, so TableExists() is true and the
// delete runs.
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("TeardownNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
}
// TestTeardownNftAbsentTableDoesNotFlush: nothing was diverting, so nothing is
// stranded. Keeps the flush out of the hot path of an idle box.
func TestTeardownNftAbsentTableDoesNotFlush(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft with no table must be a no-op, got %v", err)
}
if len(*seen) != 0 {
t.Fatalf("absent table must not flush conntrack, got %v", *seen)
}
}
// failingExec returns an execCommand replacement whose every command exits 1.
func failingExec() func(string, ...string) *exec.Cmd {
return func(name string, arg ...string) *exec.Cmd {
cs := append([]string{"-test.run=TestRestartHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_FAIL=1")
return cmd
}
}
// TestRestartHelperProcess is the exec helper for failingExec.
func TestRestartHelperProcess(t *testing.T) {
if os.Getenv("GO_WANT_HELPER_PROCESS") != "1" {
return
}
if os.Getenv("GO_HELPER_FAIL") == "1" {
os.Exit(1)
}
os.Exit(0)
}
// TestRoutingPresentRequiresLocalDefaultRoute is the second half of B3.
//
// ApplyRouting installs TWO things per family and they are removed by two
// independent commands. A presence check that only looks at the rule declares a
// half-removed plane healthy — and because applyLocked consults exactly this
// function to decide whether to re-run ApplyRouting, the missing route is then
// never restored: marked packets find an empty table, fall through to main and
// are eaten by the fail-closed forward drop, permanently, with the status still
// saying plane=full.
//
// On the pre-fix implementation the first two cases below return true.
func TestRoutingPresentRequiresLocalDefaultRoute(t *testing.T) {
const (
rule = "32765:\tfrom all fwmark 0x2000 lookup shater"
route = "local default dev lo scope host"
)
cases := []struct {
name string
ipv6 bool
v4rule, v4rte string
v6rule, v6rte string
want bool
}{
// THE REGRESSION: rule survived, table was flushed.
{"v4 rule present, route flushed", false, rule, "", "", "", false},
{"ipv6 on, v6 route flushed", true, rule, route, rule, "", false},
// Sanity: a complete plane is still reported as present.
{"v4 complete", false, rule, route, "", "", true},
{"ipv6 on, both complete", true, rule, route, rule, route, true},
// The pre-existing rule-level contract must not regress.
{"v4 rule missing", false, "", route, "", "", false},
{"ipv6 on, v6 rule missing", true, rule, route, "", route, false},
}
orig := execCommand
defer func() { execCommand = orig }()
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
execCommand = func(name string, arg ...string) *exec.Cmd {
out := ""
if name == "ip" && len(arg) >= 2 {
v6 := arg[0] == "-6"
switch arg[1] {
case "rule":
out = c.v4rule
if v6 {
out = c.v6rule
}
case "route":
out = c.v4rte
if v6 {
out = c.v6rte
}
}
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
return cmd
}
g := model.Globals{FwmarkBase: 0x2000, TableBase: 0x2000, IPv6: c.ipv6}
if got := RoutingPresent(g); got != c.want {
t.Errorf("RoutingPresent = %v, want %v", got, c.want)
}
})
}
}
+48
View File
@@ -128,6 +128,54 @@ func (s *Server) handleConfigGet(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, m)
}
// rulesReachabilityResponse is the GET /api/rules/reachability body. `rules` is
// ALWAYS an array, never null, with one entry per rule in the SAME order as
// GET /api/config's Rules — so a client can zip the two by `index` (and check the
// echoed `name`/`order` before trusting a verdict it fetched around an edit).
type rulesReachabilityResponse struct {
Rules []model.RuleReach `json:"rules"`
}
// reachConfigRead is handleRulesReachability's test seam (same pattern as
// writeConfig / logConfigRead): production binds the real UCI read, tests
// substitute a canned model so the handler can be exercised without a `uci`
// binary on the host.
var reachConfigRead = model.ReadUCI
// handleRulesReachability → GET /api/rules/reachability: which routing rules can
// never take effect, and what supersedes each of them.
//
// It is the routing-rule analogue of the per-chain `used` flag on
// GET /api/groups/health — a quiet note about the ROUTING CONFIG that the panel
// renders as a badge, never a health or liveness signal. Separate from
// /api/config on purpose: /api/config is the desired state the panel PUTs back
// verbatim, and a derived verdict has no business travelling round-trip through
// it.
//
// The verdict is computed over the EFFECTIVE rules (active WAN profile's
// enable/disable applied via model.EffectiveRules), because a rule the active
// profile switched off is not in force and must not be blamed for retiring
// anything. Profile-resolution warnings are dropped here: this endpoint answers
// one question, and the same warnings already reach the operator through
// `shaterd status` on every apply.
func (s *Server) handleRulesReachability(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
return
}
m, err := reachConfigRead()
if err != nil {
writeError(w, http.StatusInternalServerError, "read config: "+err.Error())
return
}
rules, _ := m.EffectiveRules()
out := model.RuleReachability(rules)
if out == nil {
out = []model.RuleReach{}
}
writeJSON(w, http.StatusOK, rulesReachabilityResponse{Rules: out})
}
// handleConfigPut → PUT /api/config: decode a Model, lightly validate it, and
// persist it. The write is SPLIT on the server: the FromSub nodes carried in the
// model are reconciled into the per-subscription JSON cache files
+125
View File
@@ -0,0 +1,125 @@
package panel
import (
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// getReach GETs /api/rules/reachability (authenticated) and returns the status
// code plus the decoded reply.
func getReach(t *testing.T, srv *httptest.Server, cookie *http.Cookie) (int, rulesReachabilityResponse) {
t.Helper()
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET /api/rules/reachability: %v", err)
}
defer resp.Body.Close()
var out rulesReachabilityResponse
_ = json.NewDecoder(resp.Body).Decode(&out)
return resp.StatusCode, out
}
// TestRulesReachabilityEndpoint (B1): the endpoint reports the field config's two
// condition-less `default` rules, one verdict per rule, aligned with the order
// GET /api/config returns them in — that alignment is the only thing that lets the
// panel badge the right row when both rules share a name.
func TestRulesReachabilityEndpoint(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) {
return &model.Model{Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: "direct"},
{Name: "default", Enabled: true, Order: 100, Target: "group:auto"},
}}, nil
}
code, out := getReach(t, srv, cookie)
if code != http.StatusOK {
t.Fatalf("got %d, want 200", code)
}
if len(out.Rules) != 2 {
t.Fatalf("want one verdict per rule (2), got %d: %+v", len(out.Rules), out.Rules)
}
if !out.Rules[0].Unreachable {
t.Fatalf("the order-20 default is superseded by the order-100 one: %+v", out.Rules[0])
}
if out.Rules[0].Index != 0 || out.Rules[0].ShadowedByIndex != 1 || out.Rules[0].ShadowedByOrder != 100 {
t.Fatalf("verdict must point at the superseding rule by index and order: %+v", out.Rules[0])
}
if out.Rules[0].Reason == "" {
t.Fatal("an unreachable verdict must carry a reason the panel can show")
}
if out.Rules[1].Unreachable {
t.Fatalf("the last default is the one in force: %+v", out.Rules[1])
}
}
// TestRulesReachabilityEmptyIsArray: `rules` is ALWAYS an array. Go marshals a nil
// slice as null, and a client that does `for (const r of body.rules)` breaks on it.
func TestRulesReachabilityEmptyIsArray(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) { return &model.Model{}, nil }
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET: %v", err)
}
defer resp.Body.Close()
var raw struct {
Rules *[]model.RuleReach `json:"rules"`
}
if err := json.NewDecoder(resp.Body).Decode(&raw); err != nil {
t.Fatalf("decode: %v", err)
}
if raw.Rules == nil {
t.Fatal("rules must marshal as [], never null")
}
}
// TestRulesReachabilityMethodAndAuth: it is a read endpoint behind the session
// cookie, like every other /api route except /api/session.
func TestRulesReachabilityMethodAndAuth(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
req, _ := http.NewRequest(http.MethodPost, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("POST: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusMethodNotAllowed {
t.Fatalf("POST got %d, want 405", resp.StatusCode)
}
resp, err = http.Get(srv.URL + "/api/rules/reachability")
if err != nil {
t.Fatalf("unauthenticated GET: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusUnauthorized {
t.Fatalf("unauthenticated GET got %d, want 401", resp.StatusCode)
}
}
+1
View File
@@ -178,6 +178,7 @@ func (s *Server) buildRouter() http.Handler {
mux.Handle("/api/devices", s.requireSession(http.HandlerFunc(s.handleDevices)))
mux.Handle("/api/interfaces", s.requireSession(http.HandlerFunc(s.handleInterfaces)))
mux.Handle("/api/import-wg", s.requireSession(http.HandlerFunc(s.handleImportWG)))
mux.Handle("/api/rules/reachability", s.requireSession(http.HandlerFunc(s.handleRulesReachability)))
mux.Handle("/api/ruleset/status", s.requireSession(http.HandlerFunc(s.handleRuleSetStatus)))
mux.Handle("/api/ruleset/update", s.requireSession(http.HandlerFunc(s.handleRuleSetUpdate)))
mux.Handle("/api/ruleset/check", s.requireSession(http.HandlerFunc(s.handleRuleSetCheck)))