Compare commits

..
Author SHA1 Message Date
omarandClaude Opus 5 8fd5c52488 fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.

Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.

With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.

Fix (upstream file, lx:tproxy_writeback_connect):
  * CONNECT the write-back socket to the one peer it ever talks to. The kernel's
    compute_score() rejects a connected socket for any other peer, and a
    connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
    reuseport selection outright — so it can no longer be handed a datagram it
    will not read. Nothing about the reply changes: same spoofed source, same
    single peer, Write instead of WriteTo. The unconnected path is kept verbatim
    for a destination that cannot be bound (domain socksaddr).
  * A failed cached write now CLOSES the socket instead of only dropping the
    reference (upstream left the fd to the GC finalizer).
  * TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
    but the cache evicts lazily, so after the inbound is gone nothing wakes the
    live sessions and each strands its write-back socket. Invisible upstream
    (one close at shutdown); on this fork the engine is rebuilt on every apply,
    so it was one stranded generation per apply.

Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.

The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.

Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:41:02 +03:00
omarandClaude Opus 5 32e8f8ff0b fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.

Cause: `restart` is not synchronised end to end.

  * procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
    `stop; start`, and the `service delete` ubus call returns as soon as
    SIGTERM has been SENT. `start_service` therefore re-adds the instance
    (and runs `shaterd migrate`) while the outgoing `shaterd run` is still
    executing its honest teardown.
  * The successor's only defence was `daemonAlive()` -> exit(1), leaning on
    procd's `respawn 3600 5 0` to try again five seconds later. That is a
    blind retry, not synchronisation: it neither knows nor waits for the
    teardown, and it turns every restart into a logged crash plus a
    five-second hole with no data plane.
  * `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
    engine.Close of a several-hundred-outbound box flushes cache.db to
    flash before the netplane teardown even starts — aborting the teardown
    at an arbitrary point and leaving the plane HALF removed.
  * Nothing in the tree ever touched conntrack, so flows that crossed one
    of those windows kept entries formed against a plane that no longer
    exists. For UDP there is no handshake to resynchronise on and every
    retry merely refreshes the entry, so the flow stays wedged for as long
    as the client keeps asking — a flow-scoped, permanent failure that no
    ruleset rebuild can reach.
  * RoutingPresent() reported "plane intact" from the ip RULE alone, while
    ApplyRouting installs a rule AND a `local default dev lo` route removed
    by two independent commands. A teardown interrupted between them was
    therefore invisible, applyLocked's fast-path skipped ApplyRouting
    forever, and no reconcile could repair it.

Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):

  * init: `start_service` waits for a live predecessor pidfile to clear
    before opening the instance, so restart == stop + pause + start. Zero
    cost at boot. term_timeout 10 -> 30 so an honest teardown is never
    killed halfway.
  * daemon: the single-owner guard WAITS for the predecessor (bounded,
    60s) instead of exiting 1; it still refuses if the budget expires.
  * netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
    a blanket flush would drop the admin's own SSH/LuCI sessions) called
    on every plane transition: after a ruleset loads, after the table is
    removed, and once more in applyLocked when the whole plane (table +
    policy routing + sysctls) is assembled.
  * netplane: RoutingPresent() now verifies both halves it installs.

Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:51:19 +03:00
omarandClaude Opus 5 c562579ef3 docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all
unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute
does not emit a condition-less rule as a match-all route rule: it sets
route.Final and continues, so the LAST condition-less rule by order wins, and it
can never shadow a rule that has conditions (those are emitted ahead of Final
regardless of order).

For the config on the router this inverts the conclusion: traffic IS going
through the proxy (order=100 -> group:auto is the live default) and the dead knob
is the order=20 `direct` one. Severity downgraded from high to medium
accordingly — a dead setting, not a leak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:41:58 +03:00
omarandClaude Opus 5 6f89acbae7 feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final",
so a config with two of them showed two identical claims and no hint that
only the last one is the default the router uses.

A superseded rule now loses those marks — it keeps its real Order in the rail
instead of the "·" that means final — and gains a "never applies" badge plus
a line naming the rule that beat it and what to do about it: give this one a
condition, or delete one of the two. Warn semantics throughout (--amber,
dashed frame, dimmed target chip): orange is the ACTIVE state on this
faceplate, and a rule the router ignores is the opposite of active.

Verdicts come from GET /api/rules/reachability and are keyed by the rule's
index in Rules, never by name — the config that prompted this had two rules
both called `default`. They are re-fetched after every save, and a verdict
whose echoed name/order no longer matches the row is dropped rather than
shown, so the window between an optimistic edit and the refetch cannot badge
a working rule.

Rule rows were also keyed by name in React, which silently collapses two rows
that share one; the key now carries the model index.

The mock fixture gains a second condition-less rule so `?mock` renders the
state, and mock.getRulesReachability derives its verdicts from the live
fixture config rather than hard-coding them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 88a82c7297 feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:

  * two condition-less rules retire each other, and the LAST one by Order
    wins, so an earlier "default -> direct" is dead while looking live;
  * a condition-less rule can NEVER retire a rule that HAS conditions —
    those are emitted ahead of Final whatever their Order.

A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.

model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.

Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.

Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.

Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.

GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.

Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 a8f2b0f068 ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.

ci/version.sh is now the single source of truth. It derives the version
from `git describe`:

    tag `vX.Y.Z`   -> PKG_VERSION=X.Y.Z  PKG_RELEASE=1
    off-tag build  -> nearest tag + PKG_RELEASE=<commits since it> + 1
    no tag/no git  -> 0.0.0-r1 (below everything ever published)

Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.

The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.

The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.

byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.

Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:32:38 +03:00
omarandClaude Opus 5 0b32a6d58b fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:

  daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
  [ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...

`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.

Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:

  * control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
    DisableColors at its false zero value. It now comes from
    controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
    lives in an untagged file (same split as profilewatch.go) so it is
    unit-testable off the linux target.
  * engine (shater/generate): the generated option.LogOptions never set
    DisableColor, so box.New built a colouring formatter over the shared
    sink. logOptions() now sets it from the same TTY gate (seam:
    logColorAllowed).

logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.

The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.

Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:31:01 +03:00
omarandClaude Opus 5 893fdc500c fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as 9dc954029 / aec82d444.

`nodes` now reads model.ReadUCI() — `uci export shater` merged with the
per-subscription JSON caches — which is literally the call GET
/api/config serves and generate builds the engine from, so the verb
cannot drift from the panel or from the running engine: there is no
second assembly here to drift. Both ends use the same nodesJSON():
the daemon answers over the control socket (like `stats`), and the CLI
falls back to reading the same on-disk state when no daemon is running
(like `status`). A read failure goes to stderr with a non-zero exit
instead of printing `[]`, so an empty list on stdout now means one thing.

Output is a purpose-built view rather than raw model.Node: the share-link
URI is a credential and CLI output ends up in tickets and cron mail, so
the view reports what the link decodes to (protocol/server/port) plus the
model's own facts (enabled/sub/egress/stale/fingerprint). Nodes whose URI
does not parse are still listed, with the reason in `parse_error` — the
engine skips exactly those, and hiding them would be the same lie smaller.

cmdReadStub keeps `stats`, where the default IS the truth (nothing was
counted without an engine), and now says so in its doc comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:30:27 +03:00
omarandClaude Opus 5 02c266188f docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the
previous install, install from the signed apk feed, verify the default state,
restore a working config with 315 subscription nodes, then exercise the data
plane, panel API, config lifecycle, resilience and DNS.

74 PASS. Findings (detailed separately): two catch-all `default` rules where the
first sends all unspecific traffic direct and makes the second unreachable;
`shaterd nodes` is a stub returning [] while usage promises the node list; DNS to
the router LAN address dies after `service shater restart` (stop+pause+start is
fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI
colour codes reach syslog.

Also records the four-iteration CI hunt that ended in the green apk lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:09:52 +03:00
omarandClaude Opus 5 024e9308c9 fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.

The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:

    config PACKAGE_kmod-mlx5-core
            tristate
            default m

No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.

The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.

Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:27:37 +03:00
omarandClaude Opus 5 eab1db2c3f fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.

They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.

Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
  - an explicit value in .config beats a `default` (same rule that kept our
    `# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
  - `select` is OR-ed in after the user value, so shater-core's
    `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
    transitive kmods) back on their own.

Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.

Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:16:05 +03:00
omarandClaude Opus 5 22d7161c08 fix(ci/apk): disable the SDK's ALL/ALL_KMODS mass-select
release / aarch64_cortex-a53 (push) Successful in 3m25s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m3s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.

The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:

    config ALL_NONSHARED ... default ALL
    config ALL_KMODS     ... default ALL
    config ALL           ... default y

In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.

Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.

Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.

Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:57:19 +03:00
omarandClaude Opus 5 24c5a1615d fix(ci/apk): build .config from scratch — stop packing all 3593 kmods
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m18s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m2s
release / release (push) Successful in 7s
release / release apk (push) Successful in 5s
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was
running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn —
none of which we ship) before it ever got near our four.

Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the
.config that ships inside the ImmortalWrt SDK tarball. That file is the
buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus
CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so
`make defconfig` re-selected every kernel module of the target as =m and
package/kernel/linux/compile — pulled in via shater-core's nft kmod
deps — packed the lot.

Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds":
start the .config EMPTY so kconfig can only pull in what our packages
actually select. Carried over from the SDK's .config, nothing more:
the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a
wrong guess here means silently cross-compiling for another arch),
CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and
CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one
makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel).

Also adds the diagnostics this lane never had, since a failed run leaves
a 27 MB log: the carried-over identity lines, the post-defconfig kmod
count and target readout, a hard check that all four of our packages
survived defconfig, an abort if the kmod count is back in the hundreds,
and du/df after compile.

opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR,
CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 00:31:28 +03:00
omarandClaude Fable 5 4492f0599c ci: harden feed/packaging shell scripts
release / aarch64_cortex-a53 (push) Successful in 7m3s
release / x86_64 (push) Successful in 3m13s
release / apk aarch64_cortex-a53 (push) Failing after 4m34s
release / apk x86_64 (push) Failing after 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in
  the signed Packages index can't mask an empty SHA256. Script survives -u
  (all vars use :? or :- defaults).
- .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the
  copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a
  trap-based tmpdir cleanup, and set -euo pipefail.

bash -n clean on both.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:34:50 +03:00
omarandClaude Fable 5 bc2b53069a fix(security): audit remediation — file perms, const-time auth, leaks, CSPRNG
Backend audit fixes (upstream-file edits wrapped in // lx: markers):

- experimental/libbox oom_report.go/report.go: OOM reports + configuration.json
  (server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files
  → 0o700 / 0o600. [sec-perms]
- daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret
  compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare.
  [sec-consttime]
- service/oomkiller/timer.go: network-extension cleanupTriggered logic was
  inverted, so FreeOSMemory was never called after a trigger; flip both
  assignments so a trigger schedules the deferred free and the next poll runs +
  clears it. [sec-oomcleanup]
- transport/v2rayxhttp/client.go (lx-native file): session id used math/rand →
  crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface.
- daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a
  goroutine + the ssh-agent fd on every closed session (second io.Copy blocked
  on an idle agent Read forever); tie both copies + the session ctx to a
  cancel that closes both ends. [sec-sshagent]
- daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to
  1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate]
- route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in
  select while stopIdleSuspend niled it after close (race + goroutine leak on
  Close-during-tick); pass the stop channel to the loop by value.

go build ./... (default) and the D9 shaterd linux build (tags
with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,
with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer
warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/...
./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend
and v2rayxhttp with with_xhttp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:34:50 +03:00
omarandClaude Fable 5 c257d6c5cc docs(readme): rewrite root README as Russian shater product face
- README.md: new Russian product README (what/features/architecture
  mermaid/install both feeds/build/repo layout/CI/upstream/docs/license)
- README.en.md: concise English mirror (root readme was previously English)
- README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme,
  a competing Russian README) -> points to README.md + engine-fork docs
- docs-shater/README.md: folder index

Install commands copied verbatim from docs-shater/INSTALL.md; all links
verified against existing files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:31:06 +03:00
omarandClaude Fable 5 225cce5397 chore: purge working-session junk from tree; gitignore recurrence
Remove untracked-quality artifacts accidentally committed during work
sessions (all authored downstream, unreferenced anywhere in code/docs/CI):
- 5 session screenshots in repo root (devices-after-copy-fix.png,
  live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png)
- tmp/gen_linux_test (29 MB throwaway traffic-gen binary)

Guard against repeats: ignore /*.png (root screenshots) and /tmp/.
Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are
left untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:23:04 +03:00
69 changed files with 4206 additions and 418 deletions
+60 -18
View File
@@ -32,6 +32,23 @@
# -> rolling `latest` pre-release (always-fresh feed). Publish uses the Gitea
# API via curl (ci/gitea-release.sh) — no external action needed.
#
# PACKAGE VERSIONING (bug B4)
# PKG_VERSION/PKG_RELEASE are NOT hand-written in the Makefiles any more. They
# used to be, and nobody bumped them: v0.2.2…v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with different binaries inside, so `apk update` never saw
# a new version and routers could not be updated at all. Now `ci/version.sh`
# derives them from the git tag ONCE per job (the "Compute version" step,
# exported via $GITHUB_ENV):
# tag `vX.Y.Z` -> X.Y.Z-r1
# anything else -> <nearest tag>-r<commits since it + 1>
# and hands them to the SDK builds as SHATER_PKG_VERSION/SHATER_PKG_RELEASE;
# $SHATER_VERSION (the same numbers, plus the short sha off-tag) is stamped
# into the binary's constant.Version. ci/sdk-build*.sh then ASSERT that the
# built .ipk/.apk really carry that version, so the failure can never be
# silent again. This is also why both build jobs check out with fetch-depth: 0
# — `git describe` needs tags and ancestry. `byedpi` is excluded: it keeps
# upstream ByeDPI's own PKG_VERSION (see openwrt/byedpi/Makefile).
#
# APK LANE (25.12+, ADDITIVE — T2)
# The fleet is migrating to BananaWRT 25.12-mtk-vendor (= ImmortalWrt 25.12
# base), where opkg is replaced by Alpine apk (.apk, binary packages.adb
@@ -118,8 +135,14 @@ jobs:
- { arch: x86_64, sdk: x86_64-24.10.4 } # testbed VM (generic x86-64)
- { arch: aarch64_cortex-a53, sdk: mediatek-filogic-24.10.4 } # BPI-R3 + BPI-R4 (mediatek/filogic)
steps:
# fetch-depth: 0 — the package version is DERIVED from the git tag
# (ci/version.sh: nearest `vX.Y.Z` + commits since it). The default
# shallow checkout has neither tags nor ancestry, so `git describe` would
# fail and every dispatch build would fall back to 0.0.0.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the engine via a go.mod
# `replace => ./submodules/wireguard-go` (AmneziaWG fork), so that submodule
@@ -129,6 +152,13 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# THE version step (bug B4). One computation, used by both the binary
# (constant.Version) and the three tag-versioned packages, exported to
# every later step of this job:
# tag vX.Y.Z -> X.Y.Z-r1 ; off-tag -> <last tag>-r<commits+1>
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
# Toolchain for scripts/build-shaterd.sh: Go (daemon), Node (Vite SPA), UPX.
- name: Set up Go
uses: actions/setup-go@v5
@@ -197,25 +227,23 @@ jobs:
# Build the SPA-embedded, static-musl, UPX'd shaterd for BOTH arches and
# stage dist/shaterd-<a>.upx into openwrt/shaterd/files/. MUST run before
# the SDK package build (the openwrt/shaterd package installs the staged
# artifact). VERSION is stamped into constant.Version. On an exact
# artifact). $SHATER_VERSION (from the version step above) is stamped into
# constant.Version, so the binary and the package agree. On an exact
# node_modules cache hit, --fast skips the redundant `npm ci`.
- name: Build & stage shaterd artifact
env:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages through the arch-matched OpenWrt SDK and produce a
# signed per-arch opkg feed (Packages + Packages.gz + Packages.sig + .ipk).
# SHATER_PKG_VERSION/SHATER_PKG_RELEASE reach the package Makefiles through
# the SDK container; ci/sdk-build.sh asserts the .ipk really carry them.
- name: Build signed feed (SDK)
env:
KEY_BUILD: ${{ secrets.KEY_BUILD }}
@@ -253,8 +281,12 @@ jobs:
- arch: aarch64_cortex-a53 # BPI-R3 mini (BananaWRT 25.12-mtk-vendor) + BPI-R4
sdk_url: https://downloads.immortalwrt.org/releases/25.12.1/targets/mediatek/filogic/immortalwrt-sdk-25.12.1-mediatek-filogic_gcc-14.3.0_musl.Linux-x86_64.tar.zst
steps:
# fetch-depth: 0 — see the opkg lane: the package version comes from
# `git describe`, which needs tags + ancestry.
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# scripts/build-shaterd.sh builds the AmneziaWG-patched wireguard-go via a
# go.mod `replace => ./submodules/wireguard-go`, so that submodule must be
@@ -263,6 +295,11 @@ jobs:
- name: Init wireguard-go submodule (awg)
run: git submodule update --init --depth 1 submodules/wireguard-go
# Same single version computation as the opkg lane — both lanes MUST agree
# on the version, they package the identical tree.
- name: Compute version from git tag
run: bash ci/version.sh --env >> "$GITHUB_ENV"
- name: Set up Go
uses: actions/setup-go@v5
with:
@@ -346,15 +383,10 @@ jobs:
NPM_CACHE_HIT: ${{ steps.npm-cache.outputs.cache-hit }}
run: |
set -eu
if [ "${GITHUB_REF#refs/tags/}" != "$GITHUB_REF" ]; then
V="${GITHUB_REF#refs/tags/}"
else
V="v0.2.0-dev"
fi
FAST=""
if [ "${NPM_CACHE_HIT:-}" = "true" ]; then FAST="--fast"; fi
echo "shaterd version: $V (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh "$V" $FAST
echo "shaterd version: $SHATER_VERSION / package ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE} (npm cache hit: ${NPM_CACHE_HIT:-false})"
bash scripts/build-shaterd.sh $FAST
# Compile the 4 packages as .apk through the ImmortalWrt 25.12 SDK and
# sign the per-arch packages.adb with the EC key (secret KEY_APK).
@@ -462,7 +494,7 @@ jobs:
luci-app-shater (arch=all).
Targets: x86_64 (testbed) and aarch64_cortex-a53 (BPI-R3 + BPI-R4, mediatek/filogic).
── Add as an opkg feed (recommended — then `opkg upgrade` just works) ──
── Add as an opkg feed (recommended — then updating is one command) ──
This release is itself a SIGNED package feed; opkg filters by
architecture, so the same lines work on every device:
wget -O /etc/opkg/keys/5ac4b177689cb8e0 https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
@@ -472,6 +504,10 @@ jobs:
The public-key install is one-time; after it, `opkg update/upgrade`
verify the signature with check_signature left on. Full guide: docs-shater/INSTALL.md.
── Update (name our packages — never a bare `opkg upgrade`) ──
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
── Or install the loose .ipk directly / from the tarball feed ──
wget -O /tmp/f.tgz <this release>/shater-feed-aarch64_cortex-a53.tar.gz
mkdir -p /tmp/shater && tar -C /tmp/shater -xzf /tmp/f.tgz
@@ -531,13 +567,19 @@ jobs:
Packages: shaterd + byedpi (per-arch), shater-core + luci-app-shater (arch=all).
The index \`packages.adb\` is EC-signed; trust anchor \`shater-apk.pem\` (also in \`dist/\`).
── Add as an apk repository (auto-updates via \`apk upgrade\`) ──
── Add as an apk repository ──
wget -O /etc/apk/keys/shater-apk.pem https://git.qomar.pw/omar/shater/releases/download/$TAG/shater-apk.pem
echo \"https://git.qomar.pw/omar/shater/releases/download/apk-latest-\$(cat /etc/apk/arch)/packages.adb\" > /etc/apk/repositories.d/shater.list
apk update
apk add luci-app-shater # pulls shater-core + shaterd too
apk add byedpi # optional: ByeDPI desync egress
Update: apk update && apk upgrade shaterd shater-core luci-app-shater byedpi
── Update — ALWAYS name the packages, NEVER a bare \`apk upgrade\` ──
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
A bare \`apk upgrade\` reconciles EVERY installed package against every
configured repo and can downgrade unrelated system packages; naming them
upgrades only those (apk-tools 3: \"If list of packages is provided, only
those packages are upgraded along with needed dependencies\").
Full guide: docs-shater/INSTALL.md §6. The opkg/24.10 feed lives in the \`latest\` release."
echo "[release-apk] publishing $TAG from $d"
TAG="$TAG" NAME="shater apk $VER ($arch)" BODY="$BODY" \
+20 -14
View File
@@ -1,28 +1,34 @@
#!/usr/bin/env bash
# mod from https://gist.github.com/pldubouilh/c5703052986bfdd404005951dee54683
set -e -o pipefail
set -euo pipefail
ARCH=$1
DEB_SRC=$2
OUT_IPK=$3
PROJECT=$(dirname "$0")/../..
TMP_PATH=`mktemp -d`
cp $2 $TMP_PATH
pushd $TMP_PATH
TMP_PATH=$(mktemp -d)
trap 'rm -rf "$TMP_PATH"' EXIT
DEB_NAME=`ls *.deb`
ar x $DEB_NAME
cp "$DEB_SRC" "$TMP_PATH"/
pushd "$TMP_PATH" >/dev/null
# Derive the name from the file we copied — do not glob-parse `ls *.deb`.
DEB_NAME=$(basename "$DEB_SRC")
ar x "$DEB_NAME"
mkdir control
pushd control
pushd control >/dev/null
tar xf ../control.tar.gz
rm md5sums
sed "s/Architecture:\\ \w*/Architecture:\\ $1/g" ./control -i
rm -f md5sums
sed "s/Architecture:\\ \w*/Architecture:\\ $ARCH/g" ./control -i
cat control
tar czf ../control.tar.gz ./*
popd
popd >/dev/null
DEB_NAME=${DEB_NAME%.deb}
tar czf $DEB_NAME.ipk control.tar.gz data.tar.gz debian-binary
popd
tar czf "$DEB_NAME.ipk" control.tar.gz data.tar.gz debian-binary
popd >/dev/null
cp $TMP_PATH/$DEB_NAME.ipk $3
rm -r $TMP_PATH
cp "$TMP_PATH/$DEB_NAME.ipk" "$OUT_IPK"
+6
View File
@@ -36,6 +36,12 @@ nul
# playwright MCP screenshots/snapshots
.playwright-mcp/
# working-session screenshots dropped in the repo root (not shipped docs)
/*.png
# throwaway build binaries / scratch staged under tmp/
/tmp/
# -- upstream sing-box-lx -----------------------------------------
/.idea/
.idea/
+109
View File
@@ -0,0 +1,109 @@
<!-- Language: [Русский](README.md) · **English** -->
# shater
**A self-hosted internet-control appliance for OpenWrt routers.** One box turns a
home or office network into a transparent VPN gateway, a network-wide
ad/tracker/malware blocker, per-device parental control, and a live traffic
dashboard — all local, all configured from a rich built-in web panel.
> The primary README is Russian — [README.md](README.md). This is a condensed
> English mirror.
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue.svg)](LICENSE)
![targets: x86_64 · aarch64_cortex-a53](https://img.shields.io/badge/targets-x86__64%20%C2%B7%20aarch64__cortex--a53-brightgreen.svg)
## What it is
shater is a network proxy stack for **OpenWrt / ImmortalWrt / BananaWRT** routers
(Banana Pi BPI-R3, BPI-R4 and compatible). It transparently routes all LAN traffic
through a proxy (split by domain/geo/client), filters DNS, gathers statistics, and
is managed from a built-in web panel.
The engine is a **fork of [sing-box](https://github.com/SagerNet/sing-box) via
[sing-box-lx](https://github.com/Leadaxe/sing-box-lx)**, compiled into a single Go
binary `shaterd` together with the control plane, DNS filter, stats aggregator and
the web panel itself. Broad protocol set: VLESS/VMess/Trojan/Shadowsocks,
Reality/XTLS, WireGuard, **AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
A thin **LuCI launcher** (mini-dashboard + "Open panel" button) hands the browser a
single-use token into the standalone SPA the daemon serves on its own port
(default `:8088`).
## Highlights
- Transparent **TPROXY** data plane (TCP + UDP), SNI/Host/QUIC sniffing, no DNS leaks.
- First-match routing by source / destination / list / geo / client → outbound /
selector / chain / direct / block; node groups with balancer/observatory;
multi-hop chains; per-rule egress.
- **Fail-closed kill-switch** (dead group → block, never a silent direct leak); own
`inet shater` nft table; atomic apply with `nft -c` validation and commit-confirm
auto-rollback.
- **DNS filtering & blocklists** with flexible sources (inline / file / url /
geosite), compiled `.srs` matcher; Block-DoH/DoT to stop filter bypass.
- Subscriptions (Clash / sing-box / Xray-JSON) and manual nodes; node health board.
- Per-device control (proxy/blocklist toggles, exit country, per-device block/allow,
schedules) and per-domain/client/device statistics from in-process DNS events.
Full list with MVP/T1/T2 tags — [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md).
## Install
Two signed feeds. Pick by the router's OpenWrt version. Verbatim commands and the
manual `.ipk`/`.apk` install are in [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
**opkg (OpenWrt 24.10):**
```sh
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
opkg update && opkg install luci-app-shater # -> shater-core -> shaterd
```
**apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+):**
```sh
wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
apk update && apk add luci-app-shater # -> shater-core -> shaterd
```
shater ships **inert** (globals off) so install never breaks connectivity. After
configuring nodes/rules: `uci set shater.globals.enabled=1 && uci commit shater`,
then `shaterd apply` and `shaterd confirm`.
## Build from source
`scripts/build-shaterd.sh [VERSION] [--fast]` builds the SPA (Vite), embeds it via
`//go:embed`, cross-builds musl-static `{amd64, arm64}` and UPX-packs the artifact
into `openwrt/shaterd/files/`. Details in
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
## Repository layout
| Path | What |
|------|------|
| `shater/` | Go control plane, DNS filter, stats aggregator, engine host |
| `panel/` | Admin SPA (Vite + React + TS) and its Go server |
| `openwrt/` | Packages: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Product documentation |
| `scripts/`, `ci/`, `.gitea/workflows/` | Build script, feed/release scripts, CI |
| `SPECS/`, `docs-lx/` | Engine-fork constitution/specs and feature-config reference |
| `docs/`, `mkdocs.yml` | **Upstream** sing-box docs (mkdocs) — kept as-is |
| `adapter/ cmd/ dns/ route/ option/ protocol/ transport/ …` | sing-box-lx engine tree |
## CI, upstream & license
CI (`.gitea/workflows/release.yml`) builds all 4 packages and publishes signed
feeds: opkg (usign, key `5ac4b177689cb8e0`) and apk (EC key `shater-apk.pem`). A
`vX.Y.Z` tag → versioned release; `workflow_dispatch` → rolling `latest`.
The engine is the **sing-box-lx** fork — a thin downstream of upstream sing-box that
lives by **rebase, never merge**; its constitution is
[`SPECS/CONSTITUTION.md`](SPECS/CONSTITUTION.md). Licensed under
[GPL-3.0](LICENSE), like upstream sing-box. Unofficial fork, not affiliated with
SagerNet.
+314 -51
View File
@@ -1,71 +1,334 @@
<!-- Язык: **Русский** · [English](README.en.md) -->
# shater
**A self-hosted internet-control appliance for OpenWrt.** One box turns your
network into a transparent VPN gateway, a network-wide ad/tracker/malware blocker,
per-device parental control, and a live traffic dashboard — configured from a rich
web admin panel, all local.
**Управляемый интернет-шлюз для роутеров на OpenWrt.** Одна коробка превращает
домашнюю или офисную сеть в прозрачный VPN-шлюз, сетевой блокировщик рекламы,
трекеров и вредоносных доменов, средство родительского контроля по устройствам
и живую панель аналитики трафика — всё локально, всё self-hosted, всё
настраивается из богатой веб-панели.
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue.svg)](LICENSE)
![status: v0.2 in development](https://img.shields.io/badge/status-v0.2%20in%20development-orange.svg)
![targets: x86_64 · aarch64_cortex-a53](https://img.shields.io/badge/targets-x86__64%20%C2%B7%20aarch64__cortex--a53-brightgreen.svg)
![feeds: opkg 24.10 · apk 25.12](https://img.shields.io/badge/feeds-opkg%2024.10%20%C2%B7%20apk%2025.12-orange.svg)
> ⚠️ **v0.2 is under active development on a new foundation.** The previous,
> complete and VM-verified xray-based version lives on the **[`v0.1`](../../src/branch/v0.1)**
> branch and still installs from the signed feed.
---
## What v0.2 is
## Что это
shater v0.2 is built as a **fork of [sing-box](https://github.com/SagerNet/sing-box)
(via [sing-box-lx](https://github.com/Leadaxe/sing-box-lx))** with our whole
product embedded in the one binary: the proxy engine, a control plane, a DNS
filter, and a full admin panel. Riding sing-box gives a broad, up-to-date protocol
set — VLESS/VMess/Trojan/Shadowsocks, Reality, **AmneziaWG 2.0**, Hysteria2, TUIC —
without reinventing the anti-DPI arms race.
**shater** — это сетевой прокси-стек для роутеров на **OpenWrt / ImmortalWrt /
BananaWRT** (Banana Pi BPI-R3, BPI-R4 и совместимые). Он прозрачно (без настройки
клиентов) заворачивает весь LAN-трафик через прокси с маршрутизацией по домену,
гео и клиенту, фильтрует DNS, собирает статистику и управляется из встроенной
веб-панели.
The UI is split for both integration and a great experience: a **thin LuCI app**
(a small dashboard + an "Open panel" button) hands a short-lived token to a
**standalone admin panel** the daemon serves on its own port — so panel auth is
bootstrapped from LuCI's existing login, and the real UX is a modern SPA we fully
own.
Ядро — **форк движка [sing-box](https://github.com/SagerNet/sing-box) через
[sing-box-lx](https://github.com/Leadaxe/sing-box-lx)** — вкомпилировано в один
Go-бинарь `shaterd` вместе с control-plane, DNS-фильтром, агрегатором статистики и
самой веб-панелью. За счёт sing-box поддерживается широкий и актуальный набор
протоколов: VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, WireGuard,
**AmneziaWG 2.0**, Hysteria2, TUIC, XHTTP, MASQUE/CONNECT-IP.
## Highlights (planned)
Интеграция в OpenWrt — тонкий **LuCI-лаунчер**: мини-дашборд и кнопка «Открыть
панель», которая по одноразовому токену передаёт браузер в полноценную SPA-панель,
поднятую демоном на собственном порту (по умолчанию `:8088`).
- Transparent TPROXY proxy (TCP+UDP), split by domain/geo/client, no DNS leaks.
- Broad protocols incl. **AmneziaWG 2.0**, Reality, Hysteria2, TUIC.
- Network-wide **DNS blocklists** with flexible sources (inline / file / url /
geosite) and an efficient matcher for million-entry lists.
- **Per-domain, per-client, per-device statistics** — fed by the engine's DNS
events in-process (no log scraping).
- **Per-device control**: block a site for one device or everyone; per-device
exit/proxy toggles; schedules; alerts.
- Fail-closed kill-switch, atomic apply with commit-confirm rollback, signed opkg
feed.
---
See **[`docs-shater/FEATURES.md`](docs-shater/FEATURES.md)** for the full list.
## Ключевые возможности
## Documentation
**Прозрачный прокси и маршрутизация**
- TPROXY data-plane для нескольких LAN-интерфейсов (TCP + UDP), сниффинг
SNI/Host/QUIC, без утечек DNS.
- Правила маршрутизации по источнику (IP/CIDR/MAC/интерфейс/зона), назначению
(domain/suffix/keyword/geosite), спискам, порту, протоколу →
outbound / selector / chain / direct / block.
- Группы узлов с балансировщиком/обсерваторией (least-ping / failover /
round-robin), **мульти-хоп цепочки** и выбор egress по правилу.
| Doc | What |
|-----|------|
| [`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md) | **Start here** — project context, v0.1→v0.2 history, decisions in brief, testbed/infra |
| [`docs-shater/ROADMAP.md`](docs-shater/ROADMAP.md) | Phased plan (Phase 1 = fork + embedding prototype) |
| [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md) | Full feature list with MVP/T1/T2 tags |
| [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md) | One-binary design, auth handoff, data/DNS/apply flow (diagrams) |
| [`docs-shater/DECISIONS.md`](docs-shater/DECISIONS.md) | Why sing-box, why fork, why the panel split, license, etc. |
| [`docs-shater/DESIGN.md`](docs-shater/DESIGN.md) | Admin-panel visual system — the "Faceplate" direction, tokens, components, north-star prototype |
**Надёжность («железно»)**
- **Fail-closed kill-switch**: мёртвая группа → block, а не тихая утечка мимо
прокси; собственная nft-таблица `inet shater` и свои марки/таблицы, fw4 не
трогаем.
- Атомарный apply с валидацией движком и `nft -c`, **commit-confirm** с
авто-откатом к последней рабочей конфигурации.
- Идемпотентный reconcile из hotplug/boot под flock; management-bypass
(SSH/LuCI/LAN) всегда в обход.
## Status
**DNS, фильтрация, блокировки**
- Перехват `:53`, DNS движка sing-box в процессе; резолверы DoH/DoT/plain/FakeIP,
выбор резолвера по домену.
- **Блок-листы с гибкими источниками**: `inline` / `file` / `url` (авто-обновление) /
категория `geosite`; hosts-файл, plain-список или AdBlock-стиль `||domain^`
компилируются в локальный `.srs`. Эффективный компилированный матчер вместо
dnsmasq-мегасписков.
- **Block-DoH/DoT** — не даёт устройствам обходить фильтр через свой шифрованный DNS.
Foundation reset complete: v0.1 preserved on its branch, `main` reset for v0.2.
Next is **Phase 1** — fork sing-box-lx into `main` and stand up the embedding
prototype (prove AmneziaWG 2.0, measure binary size). Follow `docs-shater/ROADMAP.md`.
**Подписки и узлы**
- Подписки (VLESS/VMess/Trojan/SS/WG/AmneziaWG), форматы Clash/sing-box/Xray-JSON,
интервал обновления + вручную + на загрузке; стабильная идентичность узла между
обновлениями; квоты/срок из `subscription-userinfo`.
- Ручные узлы: share-ссылки, импорт файла, `wg-quick`/AmneziaWG `.conf`.
- Health board: TCP + реальная проба через прокси-путь, exit-IP, «протестировать
все».
## Hardware
**Контроль по устройствам**
- Авто-обнаружение устройств (dhcp.leases + `ip neigh`), имена, живой статус/трафик.
- Тумблеры на устройство: прокси on/off, блок-листы on/off, страна/узел выхода;
блок/allow домена для одного устройства или для всех; расписания.
`aarch64_cortex-a53` covers Banana Pi **BPI-R3** (MT7986/Filogic 830) and **BPI-R4**
(MT7988/Filogic 880), both the OpenWrt `mediatek/filogic` target. `x86_64` is the
QEMU test VM.
**Статистика и видимость**
- Топ доменов (запрошенные/заблокированные), allowed-vs-blocked, разбивка по
устройствам, таймлайны — из DNS-событий движка в процессе (без скрейпинга логов).
- Трафик по клиенту/узлу/правилу (байты) из nft-счётчиков; живой query-log.
## License
**Панель и профили**
- Встроенная SPA-панель (собственный порт, вшита в бинарь): overview, узлы и
подписки, правила маршрутизации, DNS/блок-листы, устройства, apply/rollback.
- Именованные профили/сцены и WAN-профили (условные оверрайды).
[GPL-3.0](LICENSE) (sing-box is GPL-3.0). See `docs-shater/DECISIONS.md` D6.
Полный список с тегами MVP/T1/T2 — [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md).
---
## Архитектура
Один бинарь `shaterd` держит движок, control-plane, DNS-фильтр и веб-сервер панели
в одном процессе; OpenWrt-обвязка (тонкий LuCI + procd/system glue) оборачивает его.
Конфиг — UCI desired-state; демон рендерит его в конфиг движка и применяет;
телеметрия течёт обратно в панель.
```mermaid
flowchart TB
subgraph BIN["shaterd — один бинарь (форк sing-box-lx)"]
ENG["движок sing-box\nпротоколы · Reality · AmneziaWG 2.0 · DNS · routing · stats"]
CTRL["control-plane (shater/)\nUCI-модель · генерация конфига · apply/rollback · nft/routing"]
FILT["DNS-фильтр + блок-листы + политика по устройствам (shater/)"]
STAT["агрегатор статистики (shater/)"]
PANEL["веб-сервер панели + вшитая SPA (свой порт, токен-auth)"]
end
subgraph WRT["OpenWrt-обвязка (openwrt/)"]
LUCI["тонкий LuCI — мини-дашборд + кнопка «Открыть панель»"]
PROCD["procd init · hotplug · uci-defaults · fw4/routing"]
end
LUCI -->|"ubus: mint token"| PANEL
PROCD --> BIN
CTRL --> ENG
FILT --> ENG
ENG --> STAT
STAT --> PANEL
```
Путь трафика: LAN-клиент → `nft tproxy` (mark → tproxy-порт) → tproxy-inbound
sing-box (сниффинг SNI/Host/QUIC) → маршрут по правилу → outbound/selector/chain
(проксировано) · direct (flow-offload) · block. Подробные диаграммы (auth-handoff,
data-plane, DNS-flow, apply-flow) — в [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md).
---
## Установка
shater поставляется двумя подписанными фидами. Выберите по версии OpenWrt на роутере:
- **OpenWrt 24.10** → фид **opkg** (`.ipk`, `Packages.gz`, ключ usign).
- **OpenWrt / ImmortalWrt / BananaWRT 25.12+** → фид **apk** (`.apk`, `packages.adb`,
EC-ключ).
Пакеты ставятся по зависимостям: `shaterd` → `shater-core` → `luci-app-shater`
(+ опциональный `byedpi`). `shaterd` подтягивается автоматически как зависимость.
### Путь A — фид opkg (OpenWrt 24.10)
```sh
# 1) доверяем ключу фида — ИМЯ файла обязано равняться отпечатку usign-ключа.
wget -O /etc/opkg/keys/5ac4b177689cb8e0 \
https://git.qomar.pw/omar/shater/releases/download/latest/shater-feed.pub
# 2) добавляем фид (один URL обслуживает все арки).
echo "src/gz shater https://git.qomar.pw/omar/shater/releases/download/latest" \
>> /etc/opkg/customfeeds.conf
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
opkg update
opkg install luci-app-shater # -> shater-core -> shaterd
opkg install byedpi # опционально: ByeDPI desync-egress
```
Обновление — **только наши пакеты, никогда голый `opkg upgrade`** (без аргументов
он тянет обновления и на системные пакеты, это классический способ окирпичить
роутер):
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi
```
### Путь B — фид apk (OpenWrt / ImmortalWrt / BananaWRT 25.12+)
`/etc/apk/arch` сам выбирает нужный per-arch релиз (apk-релизы раздельны по арке):
```sh
# 1) доверяем ключу apk-фида (любое имя *.pem под /etc/apk/keys подходит).
wget -O /etc/apk/keys/shater-apk.pem \
"https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/shater-apk.pem"
# 2) добавляем репозиторий — строка указывает на сам ФАЙЛ-ИНДЕКС packages.adb.
echo "https://git.qomar.pw/omar/shater/releases/download/apk-latest-$(cat /etc/apk/arch)/packages.adb" \
> /etc/apk/repositories.d/shater.list
# 3) обновляемся и ставим (shaterd подтянется как зависимость).
apk update
apk add luci-app-shater # -> shater-core -> shaterd
apk add byedpi # опционально: ByeDPI desync-egress
```
Обновление — **перечисляйте пакеты явно, голый `apk upgrade` не запускайте**: без
аргументов apk пересобирает состояние ВСЕХ установленных пакетов по ВСЕМ
подключённым репозиториям и может задеть (в т.ч. откатить) посторонние системные
пакеты.
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi
# эквивалент, дополнительно закрепляющий пакеты в world:
# apk add -u shaterd shater-core luci-app-shater byedpi
```
Документация apk-tools 3 про `apk upgrade`: *«If list of packages is provided,
only those packages are upgraded along with needed dependencies»*. Проверить
установленные версии: `apk list -I shaterd shater-core luci-app-shater byedpi`.
> Версии пакетов CI берёт из git-тега (`vX.Y.Z` → `X.Y.Z-r1`, сборка вне тега →
> `X.Y.Z-r<коммитов+1>`), поэтому каждая новая сборка действительно видна
> менеджеру пакетов как новая. Подробности — `docs-shater/INSTALL.md` §2.1.
> Полные инструкции — раздельная установка из `.ipk`/`.apk` вручную, закрепление
> версии (`vX.Y.Z` / `apk-vX.Y.Z-<arch>`), совместимость с BananaWRT
> `25.12-mtk-vendor` — в [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
### Включение
shater ставится **инертным** (globals выключены), чтобы установка не рвала связь.
Настройте узлы/правила (через панель или `uci`), затем включите и примените:
```sh
uci set shater.globals.enabled=1
uci commit shater
shaterd apply # apply + вооружить commit-confirm на живом демоне
shaterd confirm # подтвердить (отменяет авто-откат)
```
`/etc/init.d/shater enable && /etc/init.d/shater start` поднимает демона под procd.
Кнопка «Открыть панель» в LuCI чеканит одноразовый токен и передаёт браузер в
панель (`:8088` по умолчанию).
---
## Сборка из исходников
Ship-артефакт — бинарь `shaterd` со вшитой SPA. Собирается вне дерева SDK скриптом
`scripts/build-shaterd.sh`:
```sh
scripts/build-shaterd.sh [VERSION] [--fast]
```
Что он делает: (1) собирает панель — `cd panel && npm ci && npm run build` (Vite →
`panel/dist`); (2) копирует `panel/dist/*` в `shater/panel/webroot/`, откуда
`//go:embed` вшивает **реальную** SPA в бинарь; (3) кросс-собирает под `{amd64,
arm64}` с musl-static набором тегов (`CGO_ENABLED=0 GOOS=linux`), stripped/trimmed;
(4) прогоняет UPX `--lzma --best` (~42 МБ → ~8–11 МБ); (5) стейджит артефакт в
`openwrt/shaterd/files/` для пакета.
Затем OpenWrt-пакеты из `openwrt/` собираются каноническим путём SDK. Детали
(набор build-тегов, почему `shaterd` — prebuilt-пакет, порядок CI) — в
[`docs-shater/INSTALL.md`](docs-shater/INSTALL.md).
---
## Структура репозитория
Репозиторий — это оверлей продукта **shater** поверх дерева форка движка
**sing-box-lx** (конфликт-фри: движок в апстрим-каталогах, продукт в своих).
| Путь | Что это |
|------|---------|
| `shater/` | Go: control-plane, DNS-фильтр, агрегатор статистики, хост движка |
| `panel/` | Админ-SPA (Vite + React + TS) и её Go-сервер |
| `openwrt/` | Пакеты: `shaterd`, `shater-core`, `luci-app-shater`, `byedpi` |
| `docs-shater/` | Документация продукта (см. таблицу ниже) |
| `scripts/` | `build-shaterd.sh` — сборка ship-артефакта |
| `ci/` | Скрипты сборки фидов и релизов (SDK, usign/EC, Gitea API) |
| `.gitea/workflows/` | `release.yml` — CI: сборка пакетов + подписанные фиды opkg/apk |
| `SPECS/` | Конституция форка движка и спеки (Spec Kit) |
| `docs-lx/` | Справочник конфигурации фич движка (`lx-config.md`, `.ru.md`) |
| `lx-test/`, `submodules/` | Примеры конфигов движка и submodule AmneziaWG-рантайма |
| `docs/`, `mkdocs.yml` | **Апстрим** документация sing-box (mkdocs) — как есть |
| `adapter/ cmd/ dns/ route/ option/ protocol/ transport/ …` | Дерево движка sing-box-lx |
---
## CI и релизы
CI на **Gitea Actions** (`.gitea/workflows/release.yml`) собирает все 4 пакета и
публикует **подписанные фиды**:
- **opkg (24.10):** один комбинированный релиз, подписан usign-ключом (публичный
`dist/shater-feed.pub`, отпечаток `5ac4b177689cb8e0`; секрет — в Gitea-secret
`KEY_BUILD`).
- **apk (25.12+):** параллельная линия, **по релизу на арку**, подписан EC-ключом
(`dist/shater-apk.pem`; секрет — `KEY_APK`).
Триггеры: push тега **`vX.Y.Z`** → версионный релиз; `workflow_dispatch` →
плавающий `latest`/`apk-latest-<arch>` (всегда свежий фид). Публикация — через
Gitea API (`ci/gitea-release.sh`). Ключи **никогда не перегенерируются** — это
инвалидировало бы доверие на всех развёрнутых роутерах.
---
## Связь с upstream и движок
shater вкомпилирует **форк движка sing-box-lx** — тонкий downstream апстрима
[SagerNet/sing-box](https://github.com/SagerNet/sing-box), добавляющий набор
клиентских фич (XHTTP, AmneziaWG 2.0, MASQUE, расширения наблюдаемости) за
build-тегами и живущий **ребейзом на каждый upstream-тег, а не merge**. Форк
разрабатывается по Spec Kit; неизменяемые принципы — в
[`SPECS/CONSTITUTION.md`](SPECS/CONSTITUTION.md), справочник фич движка — в
[`docs-lx/lx-config.ru.md`](docs-lx/lx-config.ru.md).
История: **v0.1** (движок на xray-core, полностью рабочая и VM-проверенная версия)
сохранена на ветке **[`v0.1`](../../src/branch/v0.1)**. v0.2 схлопнула runtime в
один форкнутый бинарь.
---
## Документация
| Документ | О чём |
|----------|-------|
| [`docs-shater/CONTEXT.md`](docs-shater/CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, testbed/инфра |
| [`docs-shater/INSTALL.md`](docs-shater/INSTALL.md) | Сборка ship-артефакта и установка обоих фидов (opkg/apk) |
| [`docs-shater/ARCHITECTURE.md`](docs-shater/ARCHITECTURE.md) | One-binary дизайн, auth-handoff, data/DNS/apply-потоки (диаграммы) |
| [`docs-shater/FEATURES.md`](docs-shater/FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [`docs-shater/ROADMAP.md`](docs-shater/ROADMAP.md) | Фазовый план |
| [`docs-shater/DECISIONS.md`](docs-shater/DECISIONS.md) | Почему sing-box, почему форк, split панели, лицензия |
| [`docs-shater/DESIGN.md`](docs-shater/DESIGN.md) | Визуальная система панели — «Faceplate», токены, компоненты |
| [`docs-shater/PORTING.md`](docs-shater/PORTING.md) | Порт проверенных кусков из v0.1 |
Индекс папки — [`docs-shater/README.md`](docs-shater/README.md).
---
## Оборудование
Арка `aarch64_cortex-a53` покрывает Banana Pi **BPI-R3** (MT7986/Filogic 830) и
**BPI-R4** (MT7988/Filogic 880) — оба таргет OpenWrt `mediatek/filogic`. `x86_64` —
QEMU-стенд для тестов.
---
## Лицензия
[GPL-3.0](LICENSE) — как у upstream sing-box. Подробности — в
[`docs-shater/DECISIONS.md`](docs-shater/DECISIONS.md) (D6). Неофициальный форк, не
аффилирован с SagerNet.
+15 -235
View File
@@ -1,239 +1,19 @@
[English](README.md) · **Русский**
# shater — этот файл переехал
# sing-box-lx
Лицо этого репозитория — продукт **shater** (управляемый интернет-шлюз для
роутеров на OpenWrt). Основной README на русском — **[README.md](README.md)**;
краткая английская версия — **[README.en.md](README.en.md)**.
> **Тонкий downstream-форк [SagerNet/sing-box](https://github.com/SagerNet/sing-box).**
> Небольшой набор клиентских фич поверх upstream — транспорт **XHTTP**, **AmneziaWG 2.0**, **MASQUE** (CONNECT-IP / Cloudflare WARP), расширения **наблюдаемости** (CommandClient) и балансировка нагрузки **round_robin** — каждая за своим build-tag.
> Набор может расти, философия — нет: жить ребейзом на каждый upstream-тег, а не отдельной жизнью.
Раньше здесь лежал README форка движка **sing-box-lx**, который shater
вкомпилирует в свой бинарь. Документация именно движка-форка живёт в его слое:
> 📄 README самого upstream sing-box — **[на GitHub](https://github.com/SagerNet/sing-box/blob/main/README.md)** (всегда актуальный).
- **[docs-lx/lx-config.ru.md](docs-lx/lx-config.ru.md)** — справочник конфигурации
фич движка (XHTTP, AmneziaWG 2.0, MASQUE).
- **[SPECS/CONSTITUTION.md](SPECS/CONSTITUTION.md)** — конституция тонкого форка
(принципы, build-tag изоляция, ребейз-модель).
- **[SPECS/README.md](SPECS/README.md)** — формат задач Spec Kit.
- Апстрим-README самого sing-box —
[на GitHub](https://github.com/Leadaxe/sing-box-lx).
Это не отдельный проект и не «улучшенный sing-box». Это upstream sing-box **плюс несколько фич**, реализованных так, чтобы их можно было переносить на новые версии sing-box годами, почти без конфликтов. Со временем фич может становиться больше — другие протоколы, новые возможности, — но каждая обязана жить по тем же правилам тонкого форка ([CONSTITUTION](SPECS/CONSTITUTION.md)).
---
## Уникальное позиционирование
В экосистеме sing-box форки, добавляющие XHTTP/AmneziaWG, делятся на два лагеря — и `sing-box-lx` не входит ни в один:
| Форк | Фичи | Подход | Синк с upstream |
|------|------|--------|-----------------|
| **SagerNet/sing-box** (upstream) | базовый | — | — |
| **shtorm-7/sing-box-extended** | десятки (WARP, MASQUE, MTProxy, XHTTP, AWG2, …) | «комбайн», правки повсюду | отдельная ветка, без ребейза на теги |
| **amnezia-vpn/amnezia-box**, **hoaxisr/amnezia-box** | только AWG | толстый форк, правки in-place | синк по веткам (`dev-next`/`stable-next`) |
| **➡ sing-box-lx** (этот репозиторий) | **малый набор (XHTTP, AWG2, наблюдаемость, round_robin)** | **тонкий: новые файлы за build-tag, минимум касаний upstream** | **ребейз атомарных `// lx`-коммитов на upstream-теги** |
**Чем мы отличаемся:**
- **Минимальная дивергенция.** Новый код живёт в новых файлах. Существующие upstream-файлы трогаются только в крошечных помеченных швах `// lx:begin … // lx:end`. → дешёвые ребейзы.
- **Изоляция за build-tag.** Фичи включаются тегами `with_xhttp` / `with_awg`. Сборка **без** них байт-в-байт повторяет поведение upstream — фичи ничего не ломают по умолчанию.
- **Идентичность сохранена.** Go-модуль остаётся `github.com/sagernet/sing-box`, бинарь называется `sing-box`. Суффикс `-lx` есть только в строке версии (`1.13.13-lx.N`).
- **Build-tag — родная конвенция sing-box**, а не наше изобретение (`with_quic`, `with_wireguard`, …). Мы просто применяем её с максимальной дисциплиной.
> Готовые форки-комбайны мы **не тянем как зависимость**, а используем только как референс wire-протокола.
---
## Фичи и статус
| # | Фича | Что это | Статус |
|---|------|---------|--------|
| **XHTTP** | клиентский транспорт | Xray-совместимый «splithttp» (режимы `auto`/`packet-up`/`stream-up`/`stream-one`) поверх Reality/TLS/h2c | ✅ **проверен живым Xray (3x-ui) сервером** (packet-up/auto): handshake + DNS + HTTPS + скачивание. `stream-one` — известный баг framing |
| **AmneziaWG 2.0** | клиентский endpoint | обфускация WireGuard: `Jc/Jmin/Jmax`, `S1–S4`, `H1–H4` + **2.0**: `I1–I5` (CPS — кастомные пакеты-приманки) | ✅ собирается, проходит `check`; зависимость **активирована** ([Leadaxe/wireguard-go-awg2-lx](https://github.com/Leadaxe/wireguard-go-awg2-lx) — sagernet-база + обфускация); **проверено живым AWG2-сервером**: handshake + keepalive + трафик наружу |
| **Маскировка `id/ip/ib`** | сахар над AWG | WireSock-стиль: декларативная маскировка поверх `I1` — домен (`id`) + протокол (`ip`: `quic`/`dns`/`stun`/`sip`) + браузер (`ib`), ядро строит клиент-инициированную `I1`-приманку: `quic` = out-of-order фрагментированный Initial (i1+i2), `dns`/`stun`/`sip` = query/Binding-Request/INVITE | ✅ **`ip=quic` device-проверен на реальном LTE/WARP DPI** (~330 мс, упрощает Cloudflare WARP); `dns`/`stun`/`sip` собираются и проходят `check`, но режутся как класс протокола к WARP-edge — для других провайдеров |
| **Наблюдаемость** (расширения CommandClient) | live-стрим для UI | нативные расширения libbox gRPC за `with_lx_command`: `URLTestOutbound`, `GetRules`, `GetGroups`, `GetOutbounds`, `GetPool`, плюс `Connection.detourList` (хвост detour'а отдельным полем, SPEC 017) и `SubscribeDNSQueries` — структурный live-поток DNS (домен, qtype, rcode `-1`=ошибка, CNAME-цепочка, привязка к процессу, `dnsServer`/`dnsServerType`/`outbound`, SPEC 018) | ✅ в rc-серии, потребляется **LxBox**. SPEC 014–018: [`014`](SPECS/014-CLASH_API_TO_COMMANDCLIENT_MIGRATION/SPEC.md) · [`015`](SPECS/015-COMMAND_PROTOCOL_RPC_EXTENSIONS/SPEC.md) · [`017`](SPECS/017-CONNECTION_DETOUR_CHAIN/SPEC.md) · [`018`](SPECS/018-DNS_QUERY_STREAM/SPEC.md) |
| **round_robin** (балансировка нагрузки) | режим `urltest` | пул-балансировка на `urltest` за `with_lx_command` (для `GetPool`): `mode` `least_test` (дефолт) \| `round_robin`; `balancer{pool (дефолт 3), pool_tolerance (0=держать живые / >0=топ по задержке), sticky_hash}`. Sticky-ключ: пропущен/`[]` → дефолт `["process","domain"]`, `["none"]` → выкл; компоненты `process`/`domain`/`source_ip`/`dest_ip`/`dest_port`. Фиксированные слоты `slot[hash(key)%pool]` (FNV-64a), замена в слоте; `GetPool` отдаёт слоты | ✅ локально равномерно (10/10/10, sticky off); rc.15 починил схлопывание `domain`-ключа (теперь читается `metadata.Domain`, переживающий resolve домен→IP, а не пустой `destination.Fqdn`) — на устройстве равномерность 0.27 → 0.95+. SPEC [`019`](SPECS/019-URLTEST_MODE_STICKY/SPEC.md), конфиг — [docs/.../urltest.md](docs/configuration/outbound/urltest.md) |
| **MASQUE** (`type: masque`) | клиентский outbound | CONNECT-IP (RFC 9484) поверх HTTP/3 **или** HTTP/2 для **Cloudflare WARP** (SPEC 021): туннелирует целые IP-пакеты через userspace gVisor-стек; `profile` (`cloudflare`/`standard`), `network` (`h3`/`h2`), pinning ECDSA public key, idle-suspend + самовосстановление. h2 — ручной фреймер поверх `x/net/http2` (без доп. зависимостей); `connect-ip-go` вкопан | ✅ **device-verified на Wi-Fi и LTE** (`warp=on`, реальный трафик на `h3` и `h2`); на сетях, режущих входящий UDP:443, `h3`-handshake виснет — там `network: h2` (TCP:443) |
Подробные отчёты — в [`SPECS/002-…`](SPECS/002-XHTTP_CLIENT_TRANSPORT/IMPLEMENTATION_REPORT.md), [`SPECS/003-…`](SPECS/003-AWG2_CLIENT_ENDPOINT/IMPLEMENTATION_REPORT.md) и [`SPECS/009-…`](SPECS/009-WIRESOCK_MASQUERADE_PROFILES/IMPLEMENTATION_REPORT.md). Полный справочник конфига — **[docs-lx/lx-config.ru.md](docs-lx/lx-config.ru.md)**.
> **Не поддерживается (слой Reality, отложено):** post-quantum Reality (`pqv` / ML-DSA-65) и `spiderX` из Xray. Это Xray-специфичные фичи Reality, которых нет в sing-box, а Reality — upstream-слой TLS, который мы держим нетронутым (это не одна из наших фич). Классический X25519 Reality работает; сервер, который **требует** post-quantum Reality, не подключится. Это ограничение sing-box — правильнее решать в upstream (получим на ребейзе).
---
## Сборка
Сборка идёт через отдельный **`Makefile.lx`** (upstream `Makefile` не трогаем):
```bash
git clone --recurse-submodules https://github.com/Leadaxe/sing-box-lx
make -f Makefile.lx lx-build
# → бинарь ./sing-box с версией вида 1.13.13-lx.1
```
> `--recurse-submodules` обязателен для `with_awg`: рантайм AmneziaWG подключён submodule'ом `submodules/wireguard-go` → [Leadaxe/wireguard-go-awg2-lx](https://github.com/Leadaxe/wireguard-go-awg2-lx).
Под капотом — стандартный `go build` с набором тегов (единственный источник истины — `make -f Makefile.lx lx-print-tags`):
```
with_gvisor,with_quic,with_dhcp,with_wireguard,with_utls,with_clash_api,with_naive_outbound,with_purego,badlinkname,tfogo_checklinkname0,with_xhttp,with_awg
```
Это клиентский feature-set upstream **минус** серверные/нерелевантные теги — `with_acme` (серверный выпуск сертов), `with_tailscale`, `with_ccm`/`with_ocm` (AI-прокси) — **плюс** `with_purego` (CGO-free кросс-сборка, чтобы `with_naive_outbound`/cronet собирался при `CGO=0` на любом desktop-таргете, кроме Windows 7 / 32-бит legacy-сборки, где naive выкинут — у `cronet-go` нет windows/386) и наши фичи `with_xhttp` / `with_awg`. Всё остальное — ровно как upstream.
Проверка конфигов:
```bash
./sing-box check -c lx-test/config/xhttp_reality.json
./sing-box check -c lx-test/config/awg2_basic.json
```
> `lx-test/config/` — наши примеры (upstream `test/` — отдельный Go-модуль, его не используем).
**Android (`libbox.aar`).** `make lib_install && make lib_android` собирает gomobile-AAR — `libbox.aar` (SDK 23) + `libbox-legacy.aar` (SDK 21) — с зашитыми `with_xhttp`/`with_awg` (и без `tailscale`), для встраивания в Android-приложение-потребитель (нужны NDK r28 + OpenJDK 17). `Libbox.version()` отдаёт `…-lx.N`.
---
## Конфигурация фич
> Полные таблицы полей, дефолты и `awg-quick`→JSON маппинг — **[docs-lx/lx-config.ru.md](docs-lx/lx-config.ru.md)**. Здесь — кратко.
### XHTTP (outbound transport)
```jsonc
"transport": {
"type": "xhttp",
"host": "example.com",
"path": "/xhttp",
"mode": "auto" // auto | packet-up | stream-up | stream-one
}
```
### AmneziaWG 2.0 (endpoint)
Поля AWG промотированы прямо в `WireGuardEndpointOptions`:
```jsonc
{
"type": "wireguard",
// … стандартные поля wireguard (private_key, address, peers, …) …
"jc": 10, "jmin": 50, "jmax": 100,
"s1": 20, "s2": 20, "s3": 60, "s4": 60,
"h1": 1, "h2": 2, "h3": 3, "h4": 4,
"i1": "<b 0x...><r 12>", "i2": "", "i3": "", "i4": "", "i5": "" // 2.0 CPS
}
```
> `I1–I5` — это конфиг (не согласуется по сети), значения должны **совпадать на клиенте и сервере**, регистрозависимы.
**Сахар-маскировка (`id`/`ip`/`ib`).** Вместо ручного `i1` задаёшь домен, протокол и
браузер — ядро само собирает `I1`-приманку (стиль WireSock). Удобно для упрощения
коннекта к **Cloudflare WARP**:
```jsonc
{
"type": "wireguard",
// … стандартные поля wireguard …
"id": "www.google.com", "ip": "quic", "ib": "chrome" // quic: id идёт как SNI в ClientHello
// или: "ip": "dns", "id": "www.google.com" // dns/sip: id идёт как QNAME/host
}
```
`ip` ∈ `quic|dns|stun|sip`; `id` обязателен только для `quic` (SNI); для `dns`/`sip` опционален (без него генерится псевдо-имя), `stun` игнорирует. Где задан — идёт на провод (SNI / QNAME / host)
и опционален для `sip` (без него генерится псевдо-host) и `stun`; `ib` ∈ `chrome|firefox|curl`
(только quic, эффект минимальный — без JA3-fingerprint). Взаимоисключается с явным `i1`.
Для **`quic`** ядро генерит out-of-order фрагментированный QUIC Initial (RFC 9001) — реальный
ClientHello, нарезанный на CRYPTO-фреймы в перемешанном порядке, так что line-rate DPI парсит
мусор и пропускает. Раскладка рандомизируется на каждый вызов (нет межюзерной сигнатуры), и
`ip=quic` теперь шлёт **два** независимых Initial (i1+i2) — поток читается как развивающаяся
QUIC-сессия. Это **единственный профиль, device-проверенный на реальном LTE/WARP DPI** (~330 мс).
`dns`/`stun`/`sip` реализованы как корректные клиент-инициированные запросы, но режутся как класс
протокола к WARP-edge (raw DNS/STUN/SIP к дата-центровому IP сам по себе аномален) — сохранены
для других провайдеров, чей DPI проверяет лишь корректность пакета. См.
[docs-lx/lx-config.ru.md](docs-lx/lx-config.ru.md) и [примеры SPECS/009](SPECS/009-WIRESOCK_MASQUERADE_PROFILES/EXAMPLES.md).
### MASQUE (outbound — Cloudflare WARP)
Outbound `masque` туннелирует целые IP-пакеты через **CONNECT-IP (RFC 9484)**, HTTP/3 или HTTP/2,
к **Cloudflare WARP**. Не путать с AWG-сахаром *masquerade* `id/ip/ib` выше — разные фичи, одно слово.
```jsonc
{
"type": "masque",
"tag": "warp",
"server": "162.159.198.2",
"server_port": 443,
"profile": "cloudflare", // cloudflare (WARP) | standard (RFC 9484)
"network": "h3", // ТРАНСПОРТ: h3 (QUIC) | h2 (HTTP/2). НЕ tcp/udp — это network_list
"sni": "www.microsoft.com", // domain-fronting; endpoint аутентифицируется пиннингом public key, не по SNI
"private_key": "<base64 DER EC>",
"public_key": "<base64 DER PKIX>",
"ip": "172.16.0.2/32", "ipv6": "2606:4700:110:...::/128"
}
```
Ключевой материал (`private_key`/`public_key`/`ip`/`ipv6`) берётся готовым из конфига — регистрацию
устройства в WARP делает клиент. На сетях, режущих входящий UDP:443, `h3`-handshake виснет —
переключите узел на `network: h2` (TCP:443). Полный справочник —
[docs-lx/lx-config.ru.md §4](docs-lx/lx-config.ru.md) и [SPECS/021](SPECS/021-MASQUE_CONNECT_IP_OUTBOUND/CONFIG.md).
---
## Модель сопровождения
```
upstream tag (vX.Y.Z)
│
└─► ветка lx = upstream + N атомарных // lx-коммитов
├─ FORK_BOOTSTRAP (Makefile.lx, CI, версия)
├─ XHTTP client transport
├─ AWG2 client endpoint
└─ … (новые фичи — такими же атомарными // lx-коммитами)
```
- **Только ребейз, никогда merge.** На новый upstream-тег ветка `lx` ребейзится поверх него.
- Каждая фича — атомарный коммит(ы), помеченный `// lx`. Новые файлы конфликтов не дают; швы в upstream-файлах малы и переносятся вручную.
- Разработка ведётся по **Spec Kit** (`SPECS/NNN-T-S-NAME/`: SPEC → PLAN → TASKS → IMPLEMENTATION_REPORT).
### Remotes
```bash
origin git@github.com:Leadaxe/sing-box-lx.git # ветка по умолчанию: lx
upstream https://github.com/SagerNet/sing-box.git
```
---
## Структура lx-специфики
| Путь | Назначение |
|------|------------|
| `Makefile.lx` | сборка с lx-тегами и версией `-lx` |
| `.github/workflows/lx-ci.yml` | CI: матрица фич (baseline/xhttp/awg/full) + negative-check + кросс-платформа + android AAR |
| `.github/workflows/lx-release.yml` | релиз на `v*-lx.*`: desktop ×6 + `libbox.aar` → GitHub Release |
| `SPECS/` | Spec Kit (конституция, задачи, отчёты) |
| `lx-test/config/` | примеры конфигов для `sing-box check` |
| `transport/v2rayxhttp/` | XHTTP-клиент (новый пакет) |
| `transport/wireguard/device_awg.go` | AWG IpcSet-параметры (за `with_awg`) |
| `submodules/wireguard-go` | submodule: merged-форк AmneziaWG-рантайма ([Leadaxe/wireguard-go-awg2-lx](https://github.com/Leadaxe/wireguard-go-awg2-lx)) |
| `option/v2ray_xhttp.go`, `option/wireguard_awg.go` | опции фич |
| `include/v2rayxhttp.go` | регистрация транспорта за build-tag |
Поиск всех правок upstream-файлов: `grep -rn "// lx"`.
---
## Потребитель
Ядро собирается для десктоп-лаунчера **singbox-launcher** (бандлит `bin/sing-box`). На Android потребитель встраивает **`libbox.aar`** (gomobile) вместо бинаря — конфиг-JSON тот же. Маппинг `type=xhttp` и AWG-полей в визарде — задачи на стороне потребителя, не здесь.
---
## Ссылки
| | |
|---|---|
| Upstream | [SagerNet/sing-box](https://github.com/SagerNet/sing-box) · [документация](https://sing-box.sagernet.org/) |
| Этот форк | [Leadaxe/sing-box-lx](https://github.com/Leadaxe/sing-box-lx) |
| AmneziaWG-рантайм | [Leadaxe/wireguard-go-awg2-lx](https://github.com/Leadaxe/wireguard-go-awg2-lx) — sagernet-база + обфускация (3-way merge) |
| AmneziaWG upstream | [amnezia-vpn/amneziawg-go](https://github.com/amnezia-vpn/amneziawg-go) · [docs.amnezia.org](https://docs.amnezia.org/documentation/amnezia-wg/) |
| XHTTP (исток) | [XTLS/Xray-core](https://github.com/XTLS/Xray-core) — `transport/internet/splithttp` |
| Конфиг фич | [docs-lx/lx-config.ru.md](docs-lx/lx-config.ru.md) |
| Spec Kit | [SPECS/](SPECS/) — [README](SPECS/README.md) · [CONSTITUTION](SPECS/CONSTITUTION.md) · [IMPLEMENTATION_PROMPT](SPECS/IMPLEMENTATION_PROMPT.md) |
---
## Лицензия
Наследует лицензию upstream sing-box (**GPL-3.0**). Все правки помечены `// lx` и распространяются под той же лицензией. Это неофициальный форк, не аффилирован с SagerNet.
> Файл оставлен как указатель, чтобы у репозитория был один основной русский
> README (`README.md`), а не два конкурирующих.
+12
View File
@@ -55,6 +55,16 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# Same contract as the opkg lane (ci/build-feed.sh): the workflow puts these in
# the job env via `ci/version.sh --env >> $GITHUB_ENV`; recompute here when run
# standalone. Passed into the container below and re-exported to the
# unprivileged build user in ci/sdk-build-apk.sh.
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[apk-feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) runner-side caches --------------------------------------------------
# All under $REPO/.cache so (a) actions/cache in the workflow can persist them
# between runs and (b) the nested container sees them via --volumes-from.
@@ -94,6 +104,8 @@ docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e SDK_URL="$SDK_URL" \
-e SDK_TAR="$SDK_TAR" -e DL_DIR="$CACHE/dl" -e APT_CACHE="$CACHE/apt" \
-e FEEDS_CACHE="$FEEDS_CACHE" -e KEY_APK="${KEY_APK:-}" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
debian:bookworm bash "$REPO/ci/sdk-build-apk.sh"
# --- 2) sanity: the per-arch apk repo dir must be complete -------------------
+14
View File
@@ -47,6 +47,18 @@ fi
chmod +x "$REPO"/ci/*.sh 2>/dev/null || true
# --- 0.4) package version from the git tag ------------------------------------
# The workflow normally puts these in the job env (ci/version.sh --env >>
# $GITHUB_ENV); recompute here when this script is run standalone so a manual
# `ci/build-feed.sh ...` produces the same versions as CI. They are handed to the
# SDK container below and read by openwrt/*/Makefile (bug B4 — versions used to
# be hand-written literals that nobody bumped, so v0.2.2…v0.2.6 all shipped as
# 0.2.0-r3 and no router could ever see an update).
if [ -z "${SHATER_PKG_VERSION:-}" ] || [ -z "${SHATER_PKG_RELEASE:-}" ]; then
eval "$(sh "$REPO/ci/version.sh" --env)"
fi
echo "[feed] package version: ${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
# --- 0.5) persistent dl/ (package source tarballs) ----------------------------
# Workspace dir restored/saved by actions/cache in the workflow and shared into
# the nested SDK container via --volumes-from; becomes CONFIG_DOWNLOAD_FOLDER
@@ -81,6 +93,8 @@ docker pull "openwrt/sdk:$SDK_TAG"
docker run --rm --volumes-from "$(hostname)" \
-e ARCH="$ARCH" -e REPO="$REPO" -e OUT="$OUT" -e DL_DIR="$DL_DIR" \
-e FEEDS_CACHE="$FEEDS_CACHE" \
-e SHATER_PKG_VERSION="$SHATER_PKG_VERSION" \
-e SHATER_PKG_RELEASE="$SHATER_PKG_RELEASE" \
"openwrt/sdk:$SDK_TAG" \
sh "$REPO/ci/sdk-build.sh"
+1 -1
View File
@@ -13,7 +13,7 @@
# Feed format: opkg `src/gz` (.ipk + text Packages index, usign signature).
# OpenWrt 24.10 (our SDK) still uses opkg; apk arrives at 25.12. The committed
# trust anchor dist/shater-feed.pub is a usign (Ed25519) key, matching this.
set -e
set -euo pipefail
OUT="${1:?feed dir required}"; cd "$OUT"
: > Packages
for ipk in *.ipk; do
+244 -1
View File
@@ -28,6 +28,11 @@ SDK_URL="${SDK_URL:?SDK_URL env required}"
echo "[apk-sdk] arch=$ARCH repo=$REPO out=$OUT"
echo "[apk-sdk] sdk=$SDK_URL"
# Package version derived from the git tag by ci/version.sh (bug B4). Forwarded
# to the unprivileged build user on the `su` line at the bottom of this file;
# openwrt/{shaterd,shater-core,luci-app-shater}/Makefile pick it up from the
# environment. byedpi keeps upstream ByeDPI's own version (see its Makefile).
echo "[apk-sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[apk-sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -133,6 +138,101 @@ fi
echo "[apk-sdk] feeds install (prefer shater feed)"
./scripts/feeds install -p shater shaterd shater-core byedpi luci-app-shater
# --- strip the SDK's generated per-package `default m` blocks ----------------
# Run 60 settled the question that runs 58 and 59 left open. Writing an explicit
# `# CONFIG_PACKAGE_kmod-x is not set` for all 1126 of them and re-running
# defconfig deselected exactly nothing: the count came back 1078, unchanged.
# Meanwhile the very same explicit form DID stick for CONFIG_ALL/ALL_KMODS/
# ALL_NONSHARED. The difference is prompts. kconfig only honours a user value for
# a symbol that has one (sym_calc_value ignores S_DEF_USER for a promptless
# symbol and falls back to its `default`), and the ALL* symbols carry prompts in
# the SDK's own Config.in while these generated blocks are bare:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# So no value we write into .config can ever turn them off — the fix has to
# remove the `default m` itself. That is what this does: drop every generated
# `config PACKAGE_*` block from the SDK's Config-build.in before the first
# defconfig. Nothing is lost by it — these blocks only replay which packages the
# BUILDBOT happened to build; the packages themselves are still declared, with
# prompts, by the package tree (tmp/.config-package.in), which is what makes our
# four selectable and what `select` acts on. KERNEL_*/LIBC/TOOLCHAIN blocks are
# left untouched, so the SDK still reproduces its own toolchain settings.
CB=$(find . -maxdepth 2 -name 'Config-build.in' -print -quit 2>/dev/null || true)
if [ -n "$CB" ] && command -v perl >/dev/null 2>&1; then
pkg_before=$(grep -c '^config PACKAGE_' "$CB" || true)
# Paragraph-wise delete: a block is `config PACKAGE_x`, its indented body, and
# the blank line that ends it. Anchored per-line (/m) so nothing else matches.
perl -0777 -pi -e 's/^config PACKAGE_\S+\n(?:[ \t]+\S[^\n]*\n)+\n//gm' "$CB"
pkg_after=$(grep -c '^config PACKAGE_' "$CB" || true)
echo "[apk-sdk] $CB: stripped $((pkg_before - pkg_after)) generated PACKAGE default blocks ($pkg_before -> $pkg_after)"
else
echo "[apk-sdk] WARNING: no Config-build.in found (or no perl) — per-package"
echo "[apk-sdk] 'default m' blocks stay; the kmod tripwire will catch it"
fi
# --- .config: turn OFF the SDK's mass-select defaults ------------------------
# Symptom (v0.2.2, and still v0.2.3 run 58): the SDK ran `apk mkpkg` on ~1100
# kmod-* packages — mlx5, amdgpu, ata, isdn, none of which we ship — and died
# with `Disk quota exceeded` on the runner's 64 GB ZFS quota. Our kmod deps pull
# in `package/kernel/linux/compile`, which packs every module marked =m.
#
# Why they are =m has nothing to do with anything we write here. An OpenWrt SDK
# carries its OWN top-level Config.in (target/sdk/files/Config.in), and it reads:
#
# config ALL_NONSHARED
# bool "Select all target specific packages by default"
# default ALL
# config ALL_KMODS
# bool "Select all kernel module packages by default"
# default ALL
# config ALL
# bool "Select all userspace packages by default"
# default y <-- y, not n, and ONLY inside the SDK
#
# In the main tree those three default to n; the SDK flips ALL to y so that
# `make world` in a bare SDK builds something useful. So `make defconfig` on ANY
# .config — empty or not — selects the entire kernel. This is stock OpenWrt, not
# an ImmortalWrt quirk: openwrt/openwrt's target/sdk/files/Config.in is identical.
# (It also means the reference we copied, Slava-Shchipunov/awg-openwrt, builds
# every kmod too — it just never hits a disk quota on GitHub's runners.)
#
# Fix: state all three explicitly. They carry prompts in the SDK's Config.in, so
# they are user-settable and an explicit value beats the `default`. Note the FORM:
# kconfig writes a false bool as `# CONFIG_X is not set` and `CONFIG_X=n` is not
# reliably honoured, so `is not set` is the only form used here. All three are set
# rather than just the root `ALL`, so this keeps working whichever symbol a future
# SDK makes the root of the chain.
# Stash anything the SDK shipped (see below — today there is nothing) and start
# from a known-empty file, so what we build here is exactly what we intended.
if [ -s .config ]; then mv -f .config .config.sdk; fi
: > .config
for s in ALL ALL_KMODS ALL_NONSHARED; do
echo "# CONFIG_$s is not set" >> .config
done
# About that stash: an SDK tarball ships NO top-level .config (run 58 logged
# `grep: .config: No such file or directory` — the only `.config` inside the
# tarball is the prebuilt KERNEL's, under the linux dir). This is also why the
# first version of this fix was aimed at the wrong thing: there was never a
# buildbot .config here to append to. Nothing needs carrying over from it either,
# because
# target/sdk/Makefile bakes the buildbot's non-package settings — every
# CONFIG_KERNEL_* included — into the SDK's generated Config-build.in as kconfig
# `default`s (target/sdk/convert-config.pl). defconfig therefore reproduces the
# exact toolchain/kernel settings the SDK was built with, on its own; an earlier
# attempt to copy those lines by hand was redundant and is gone.
# Should a future SDK start shipping a .config, this keeps the two things that
# would then be worth honouring — the target identity and the package format —
# and still lets the lines above override the mass-select.
if [ -s .config.sdk ]; then
echo "[apk-sdk] SDK shipped a .config — carrying over target identity + format:"
grep -E '^CONFIG_TARGET_[a-z0-9_]+=y$|^CONFIG_TARGET_(BOARD|SUBTARGET|ARCH_PACKAGES)=|^CONFIG_USE_APK=' \
.config.sdk | tee -a .config | sed 's/^/[apk-sdk] /' || true
fi
for p in shaterd shater-core byedpi luci-app-shater; do
echo "CONFIG_PACKAGE_$p=m" >> .config
done
@@ -151,11 +251,135 @@ fi
echo "[apk-sdk] defconfig"
make defconfig >/dev/null
# --- second pass: deselect the kernel, keep only what our packages select -----
# Turning ALL/ALL_KMODS/ALL_NONSHARED off (above) provably worked — run 59 shows
# all three as `is not set` after defconfig — and changed the kmod count by
# exactly zero, 1078 both times. The kmods are not selected through ALL_KMODS at
# all. They are selected one by one, and here is where from:
#
# target/sdk/Makefile:
# ./convert-config.pl $(TOPDIR)/.config > $(SDK_BUILD_DIR)/Config-build.in
#
# The SDK's Config-build.in is GENERATED from the buildbot's .config — a config
# in which ALL_KMODS=y had already expanded into a `CONFIG_PACKAGE_kmod-*=m` line
# per module. convert-config.pl turns every `CONFIG_X=<val>` line into a kconfig
# symbol carrying an unconditional `default <val>`; its `next if
# /^(# )?CONFIG_PACKAGE/` filter sits in the `else` branch, which a line with an
# `=` in it never reaches. So the SDK ships, verbatim, 1078 blocks of:
#
# config PACKAGE_kmod-mlx5-core
# tristate
# default m
#
# Nothing there consults ALL_KMODS, which is why switching it off was inert.
#
# Fix: give those symbols an explicit user value. We cannot do it before the
# first defconfig — the list of names only exists once kconfig has expanded the
# tree — so this is a second pass: rewrite every selected kmod to `is not set`
# and re-run defconfig. Two kconfig rules make the result exactly what we want,
# and both are already demonstrated in our own logs:
# * an explicit value in .config beats a `default` (this is precisely why the
# `# CONFIG_ALL* is not set` lines survived defconfig in run 59), so the
# ~1078 kmods we do not need stay off;
# * `select` is a reverse dependency, OR-ed into the symbol's value AFTER the
# user value in sym_calc_value(), so it cannot be overridden by an explicit
# `n`. shater-core's `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` becomes
# `select PACKAGE_kmod-nft-tproxy` (scripts/package-metadata.pl: a `+` flag
# sets `$m = "select"`, and it re-emits the dependency's own depends too, so
# transitive kmods follow). Those come back on their own.
# Net effect: we build the handful of kmods our packages actually pull in.
#
# Rejected alternatives:
# * limiting what `package/kernel/linux/compile` packs — that target has no
# such knob; it iterates the selected set, so the selection IS the knob;
# * `package/kernel/linux/clean` + a targeted build — the kernel package would
# simply be rebuilt in full as a dependency of shater-core, same cost;
# * copying OpenWrt's own feed CI (openwrt/gh-action-sdk) — it does nothing
# about this; it just runs `make defconfig` and builds. Its one disk-related
# setting, CONFIG_AUTOREMOVE=y, is already the SDK's default;
# * editing the SDK's generated Config-build.in to strip the offending blocks —
# it would work, but it means parsing a generated kconfig file by hand and a
# format change would corrupt it silently. The two-pass approach uses only
# kconfig's documented semantics and leaves the evidence in .config.
kmods_all=$(grep -c '^CONFIG_PACKAGE_kmod-[^=]*=[my]$' .config || true)
if [ "$kmods_all" -gt 0 ]; then
echo "[apk-sdk] deselecting $kmods_all kmod packages, then defconfig again"
sed -i -E 's/^CONFIG_(PACKAGE_kmod-[^=]*)=[my]$/# CONFIG_\1 is not set/' .config
make defconfig >/dev/null
fi
# --- post-defconfig sanity + disk-cost readout -------------------------------
# A failed run leaves a ~27 MB log; digging the cause out of it is miserable, so
# print the handful of numbers that decide whether this run survives the
# runner's disk quota BEFORE anything is compiled.
kmods=$(grep -c '^CONFIG_PACKAGE_kmod.*=m' .config || true)
echo "[apk-sdk] target: board=$(sed -n 's/^CONFIG_TARGET_BOARD=//p' .config)" \
"subtarget=$(sed -n 's/^CONFIG_TARGET_SUBTARGET=//p' .config)" \
"arch_packages=$(sed -n 's/^CONFIG_TARGET_ARCH_PACKAGES=//p' .config)"
echo "[apk-sdk] kmod packages selected (=m): $kmods"
# Proof the mass-select stayed off: these three must come back out of defconfig
# as `is not set`. If any reads `=y`, the SDK's `default ALL`/`default y` won and
# the kmod count above will be in the four digits.
echo "[apk-sdk] mass-select symbols after defconfig:"
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/[apk-sdk] /' || true
# After the second pass the only kmods left are the ones shater-core's
# `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` turns into kconfig `select`s, plus
# whatever those select in turn — a handful. Worth printing verbatim while the
# list is short. A count of 0 is NOT fatal: those kmods ship in the router's own
# base feed, so apk resolves them there; but it would mean the selects did not
# fire, and that is something we want to see in the log rather than guess at.
if [ "$kmods" -le 30 ]; then
grep '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/[apk-sdk] /' || true
fi
# The two cache knobs are written before the first defconfig and have to survive
# both of them — losing DOWNLOAD_FOLDER silently costs us the dl/ cache, and
# losing LOCALMIRROR brings back the sourceware.org stalls. Cheap to just look.
echo "[apk-sdk] cache settings after defconfig:"
grep -E '^CONFIG_(LOCALMIRROR|DOWNLOAD_FOLDER)=' .config | sed 's/^/[apk-sdk] /' || true
echo "[apk-sdk] our packages after defconfig:"
grep -E '^CONFIG_PACKAGE_(shaterd|shater-core|byedpi|luci-app-shater)=' .config \
| sed 's/^/[apk-sdk] /' || true
# Each of our 4 must have SURVIVED defconfig. If kconfig dropped one, it is
# because a symbol it `select`s (a DEPENDS entry) does not exist in the installed
# feeds — with the old append-everything .config that was masked by the SDK
# pre-selecting half the distro. `make package/<p>/compile` would then die with a
# cryptic "No rule to make target", far from the real cause.
for p in shaterd shater-core byedpi luci-app-shater; do
grep -q "^CONFIG_PACKAGE_$p=m" .config || {
echo "[apk-sdk] ERROR: $p is NOT selected after defconfig."
echo " kconfig dropped it -> one of its DEPENDS is missing from the"
echo " installed feeds (check the 'feeds install' step above)."; exit 10; }
done
# Only our two nft kmods (+ whatever they themselves depend on) have any business
# being selected here — a dozen at the very most. A count in the hundreds means an
# ALL_KMODS-style mass-select crept back in, and the run would spend ~40 min
# packing the kernel before dying on `Disk quota exceeded`. Fail now instead.
[ "$kmods" -le 200 ] || {
echo "[apk-sdk] ERROR: $kmods kmod packages selected — that is the whole kernel."
echo " Aborting before this fills the runner's disk. Two causes are"
echo " possible, and the lines below tell them apart:"
echo " (a) the mass-select is back on -> a CONFIG_ALL* line reads =y;"
echo " (b) the second pass did not take -> ALL* are 'is not set' but the"
echo " kmods returned anyway, i.e. the per-kmod 'default m' from the"
echo " SDK's generated Config-build.in outlived our explicit 'n'."
grep -E '^(# )?CONFIG_ALL(_KMODS|_NONSHARED)?[ =]' .config | sed 's/^/ /' || true
echo " first few kmods still selected:"
grep -m5 '^CONFIG_PACKAGE_kmod.*=m' .config | sed 's/^/ /' || true
exit 11; }
for p in shaterd shater-core byedpi luci-app-shater; do
echo "[apk-sdk] === build $p ==="
make "package/$p/compile" V=s -j"$(nproc)"
done
# What the build actually cost on disk. The runner's 64 GB ZFS quota is the
# binding constraint on this lane, so record it while the tree still exists.
echo "[apk-sdk] disk usage after compile:"
du -sh build_dir staging_dir bin 2>/dev/null || true
df -h /home/build || true
# A 25.12 apk-SDK must emit .apk — finding only .ipk means a wrong SDK was fed in.
anyapk=$(find bin -type f -name '*.apk' | wc -l)
[ "$anyapk" -gt 0 ] || {
@@ -173,6 +397,25 @@ done
[ "$found" -ge 4 ] || { echo "[apk-sdk] ERROR: expected >=4 of OUR .apk, collected $found"; echo "[apk-sdk] (all .apk under bin/:)"; find bin -type f -name '*.apk' | head -20; exit 6; }
echo "[apk-sdk] collected $found of our .apk"
# --- assert the tag-derived version actually reached the packages -------------
# B4's failure mode is a wrong-but-plausible version shipping silently, so the
# env -> make hand-off is verified, not trusted: each of our three tag-versioned
# packages must be named `<name>-<ver>-r<rel>.apk`. byedpi is excluded on purpose
# (it carries upstream ByeDPI's own version). This runs BEFORE `apk mkndx`, so a
# stale version can never even reach the index.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
[ -f "$OUT/${p}-${want}.apk" ] || {
echo "[apk-sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[apk-sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[apk-sdk] version check OK — our 3 packages are $want"
fi
# --- index + sign: exactly how the OpenWrt 25.12 buildsystem does it ---------
# apk mkndx --root T --keys-dir T [--sign key] --allow-untrusted \
# --output packages.adb *.apk
@@ -204,7 +447,7 @@ INNER
chmod 0644 /home/build/inner.sh
su build -s /bin/bash -c \
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' bash /home/build/inner.sh"
"ARCH='$ARCH' REPO='$REPO' OUT='$OUT' SDKDIR='$SDKDIR' KEYFILE='${KEYFILE:-}' DL_DIR='${DL_DIR:-}' FEEDS_CACHE='${FEEDS_CACHE:-}' SHATER_PKG_VERSION='${SHATER_PKG_VERSION:-}' SHATER_PKG_RELEASE='${SHATER_PKG_RELEASE:-}' bash /home/build/inner.sh"
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[apk-sdk] OK arch=$ARCH — apk feed dir:"
+27
View File
@@ -27,6 +27,13 @@ OUT="${OUT:?OUT env required}"
mkdir -p "$OUT"
echo "[sdk] arch=$ARCH repo=$REPO out=$OUT"
# Package version, derived from the git tag by ci/version.sh and handed in by
# ci/build-feed.sh. openwrt/{shaterd,shater-core,luci-app-shater}/Makefile read
# these straight out of the environment ($(if $(SHATER_PKG_VERSION),...)); make
# imports every environment variable as a variable, and it propagates through
# `make package/<p>/compile`, the metadata dump and the sub-makes alike.
# byedpi deliberately keeps its own upstream version (see its Makefile).
echo "[sdk] package version: ${SHATER_PKG_VERSION:-<unset -> Makefile fallback>}-r${SHATER_PKG_RELEASE:-?}"
test -f "$REPO/openwrt/shaterd/Makefile" || {
echo "[sdk] ERROR: feed not mounted ($REPO/openwrt/shaterd/Makefile missing)"; ls -la "$REPO" || true; exit 9; }
@@ -111,6 +118,26 @@ for p in shaterd shater-core byedpi luci-app-shater; do
done
done
[ "$found" -ge 4 ] || { echo "[sdk] ERROR: expected >=4 of OUR .ipk, collected $found"; echo "[sdk] (all .ipk under bin/:)"; find bin -type f -name '*.ipk' | head -20; exit 4; }
# --- assert the tag-derived version actually reached the packages -------------
# The whole point of B4 is that a WRONG-but-plausible version ships silently. The
# env -> make hand-off has several layers (docker -e, make's env import, the
# metadata dump), so verify the result instead of trusting it: every one of our
# three tag-versioned packages must be named `<name>_<ver>-r<rel>_<arch>.ipk`.
# byedpi is excluded on purpose — it keeps upstream ByeDPI's own version.
if [ -n "${SHATER_PKG_VERSION:-}" ] && [ -n "${SHATER_PKG_RELEASE:-}" ]; then
want="${SHATER_PKG_VERSION}-r${SHATER_PKG_RELEASE}"
for p in shaterd shater-core luci-app-shater; do
ls "$OUT/${p}_${want}_"*.ipk >/dev/null 2>&1 || {
echo "[sdk] ERROR: $p was not built as version '$want'."
echo " SHATER_PKG_VERSION/SHATER_PKG_RELEASE did not reach the package"
echo " Makefile — the build would have shipped a stale version (bug B4)."
echo "[sdk] collected:"; ls -1 "$OUT" | sed 's/^/ /'
exit 12; }
done
echo "[sdk] version check OK — our 3 packages are $want"
fi
chmod -R a+rwX "$OUT" 2>/dev/null || true
echo "[sdk] OK arch=$ARCH — collected $found of our .ipk:"
ls -l "$OUT"
Executable
+133
View File
@@ -0,0 +1,133 @@
#!/bin/sh
# ci/version.sh — the SINGLE source of truth for "what version is this build?".
#
# WHY THIS EXISTS (bug B4)
# -----------------------
# PKG_VERSION/PKG_RELEASE used to be hand-written literals in the four package
# Makefiles, and nobody remembered to bump them: v0.2.2 … v0.2.6 all shipped as
# `shaterd 0.2.0-r3` with DIFFERENT binaries inside (v0.2.6's ELF is 5 491 616 B
# vs r2's 5 488 336 B). Since both opkg and apk offer an upgrade only when the
# feed's version string differs from the installed one, `apk update` saw nothing
# new and the routers could not be updated through the normal path at all.
#
# So the version is now DERIVED, in CI, from the git tag, and the package
# Makefiles only carry a fallback for manual/offline builds.
#
# THE SCHEME
# ----------
# tag push `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
# any other build -> PKG_VERSION=X.Y.Z of the NEAREST reachable tag,
# (workflow_dispatch, PKG_RELEASE=<commits since that tag> + 1
# rolling `latest`)
# no tag / no git at all -> PKG_VERSION=0.0.0 PKG_RELEASE=1 (+ warning)
#
# Both managers compare `<upstream>-r<rel>` the same way: the dotted upstream
# part first (numerically, component by component), the `r<rel>` only as a
# tie-break. Verified against the real tools, not from memory:
# apk-tools 3.0.3 (`apk version -t`) and apk-tools 2.14.6:
# 0.2.6-r1 > 0.2.0-r3 0.2.6-r12 > 0.2.6-r1
# 0.2.7-r1 > 0.2.6-r12 0.0.0-r1 < 0.2.0-r3
# opkg 38eccbb1 from openwrt/rootfs:x86-64-24.10.4 (`opkg compare-versions`):
# identical results (opkg implements the Debian algorithm).
# That is exactly the ordering this scheme needs:
# * a release always outranks every rolling build that preceded it
# (0.2.7-r1 > 0.2.6-rN for any N — the dotted part decides), and
# * rolling builds between two releases grow monotonically (r2 < r10 < r11),
# so a rolling build can never look newer than the next release, and the
# `latest` feed still moves forward on every dispatch.
#
# +1 on the commit count (rather than the raw count) only avoids `-r0` and makes
# a dispatch build of the tagged commit itself identical to the release build of
# that same commit — which is the truth: same tree, same binary.
#
# `byedpi` is deliberately NOT versioned from our tag — see openwrt/byedpi/Makefile.
#
# USAGE
# ci/version.sh # or --env: eval-able / $GITHUB_ENV-able lines
# ci/version.sh --pkg-version # X.Y.Z
# ci/version.sh --pkg-release # R
# ci/version.sh --binary # vX.Y.Z-rR[-g<sha>] for constant.Version
#
# Env:
# SHATER_REF / GITHUB_REF when it is `refs/tags/<tag>` that tag wins and no
# git history is needed (the tag-push path is exact
# even on a shallow checkout).
set -eu
REPO="$(CDPATH='' cd -- "$(dirname -- "$0")/.." && pwd)"
TAG=""
EXACT=0
N=0
SHA=""
# --- 1) an explicit tag ref is authoritative (and needs no git) --------------
REF="${SHATER_REF:-${GITHUB_REF:-}}"
case "$REF" in
refs/tags/*) TAG="${REF#refs/tags/}"; EXACT=1 ;;
esac
# --- 2) otherwise ask git for the nearest reachable release tag --------------
# `--match 'v[0-9]*'` keeps non-release tags (latest, sdk-cache, apk-latest-*,
# musl-toolchain-cache) out. This repo is a sing-box FORK and therefore also
# carries upstream's v1.x tags — `git describe` picks the CLOSEST tag by commit
# distance, so our own v0.2.x (a handful of commits back) always wins over
# upstream's v1.x (thousands of commits back). The tag it picked is logged
# below, so a surprise is visible in the CI log rather than silently shipped.
if [ "$EXACT" -eq 0 ]; then
if D="$(git -C "$REPO" describe --tags --long --match 'v[0-9]*' 2>/dev/null)"; then
# `v0.2.6-1-g02c266188` -> TAG=v0.2.6 N=1 SHA=g02c266188.
# `%` strips the SHORTEST matching suffix, so a tag that itself contains a
# dash (`v0.2.0-healthplan`) survives intact.
TAG="${D%-*-g*}"
REST="${D#"$TAG"-}"
N="${REST%%-*}"
SHA="${REST#*-}"
if [ "$N" -eq 0 ]; then EXACT=1; fi
fi
fi
# --- 3) tag -> numeric PKG_VERSION ------------------------------------------
# Keep the leading dotted-numeric run only: `v0.2.0-healthplan` -> `0.2.0`.
VER=""
if [ -n "$TAG" ]; then
VER="$(printf '%s' "${TAG#v}" | sed -n 's/^\([0-9][0-9.]*\).*/\1/p' | sed 's/\.*$//')"
fi
if [ -z "$VER" ]; then
# No release tag anywhere (shallow clone with no tags, a tarball export, a
# fresh fork). 0.0.0 is BELOW every version we have ever published, so such a
# build can never masquerade as an upgrade on a real router; the commit count
# still makes successive dev builds distinguishable.
VER="0.0.0"
EXACT=0
N="$(git -C "$REPO" rev-list --count HEAD 2>/dev/null || echo 0)"
SHA="$(git -C "$REPO" rev-parse --short HEAD 2>/dev/null || echo '')"
[ -z "$SHA" ] || SHA="g$SHA"
echo "[version] WARNING: no reachable vX.Y.Z tag (and/or no git) -> $VER" >&2
fi
# --- 4) PKG_RELEASE + the string stamped into the binary --------------------
if [ "$EXACT" -eq 1 ]; then
REL=1
FULL="v${VER}-r${REL}"
else
REL=$((N + 1))
FULL="v${VER}-r${REL}${SHA:+-$SHA}"
fi
echo "[version] tag='${TAG:-none}' commits_since=$N exact=$EXACT -> ${VER}-r${REL} (binary: $FULL)" >&2
case "${1:---env}" in
--env|"")
printf 'SHATER_PKG_VERSION=%s\n' "$VER"
printf 'SHATER_PKG_RELEASE=%s\n' "$REL"
printf 'SHATER_VERSION=%s\n' "$FULL"
;;
--pkg-version) printf '%s\n' "$VER" ;;
--pkg-release) printf '%s\n' "$REL" ;;
--binary|--version) printf '%s\n' "$FULL" ;;
*)
echo "usage: $0 [--env|--pkg-version|--pkg-release|--binary]" >&2
exit 2 ;;
esac
+19
View File
@@ -2,6 +2,9 @@ package daemon
import (
"context"
// lx:begin sec-oomgate
"sync"
// lx:end sec-oomgate
"time"
"unsafe"
@@ -19,6 +22,10 @@ type ManagedService struct {
handler ManagedHandler
debug bool
oomReporter oomkiller.OOMReporter
// lx:begin sec-oomgate
oomReportMu sync.Mutex
oomReportLast time.Time
// lx:end sec-oomgate
}
type ManagedServiceOptions struct {
@@ -90,6 +97,18 @@ func (s *ManagedService) TriggerOOMReport(ctx context.Context, _ *emptypb.Empty)
if s.oomReporter == nil {
return nil, status.Error(codes.Unavailable, "OOM reporter not available")
}
// lx:begin sec-oomgate
// Rate-limit operator-triggered reports to at most one per minute: each write
// dumps process state + the config snapshot (secrets) to disk, so an
// authenticated client must not be able to spin it in a tight loop.
s.oomReportMu.Lock()
if !s.oomReportLast.IsZero() && time.Since(s.oomReportLast) < time.Minute {
s.oomReportMu.Unlock()
return nil, status.Error(codes.ResourceExhausted, "OOM report rate-limited (max 1/min)")
}
s.oomReportLast = time.Now()
s.oomReportMu.Unlock()
// lx:end sec-oomgate
return &emptypb.Empty{}, s.oomReporter.WriteReport(memory.Total())
}
+7 -1
View File
@@ -2,6 +2,9 @@ package daemon
import (
"context"
// lx:begin sec-consttime
"crypto/subtle"
// lx:end sec-consttime
"strings"
"google.golang.org/grpc"
@@ -59,8 +62,11 @@ func authenticate(ctx context.Context, secret string) error {
return status.Error(codes.Unauthenticated, "missing authorization")
}
token, isBearer := strings.CutPrefix(values[0], "Bearer ")
if !isBearer || token != secret {
// lx:begin sec-consttime
// Constant-time compare: a plain != leaks the secret via response timing.
if !isBearer || subtle.ConstantTimeCompare([]byte(token), []byte(secret)) != 1 {
return status.Error(codes.Unauthenticated, "invalid authorization")
}
// lx:end sec-consttime
return nil
}
+23 -2
View File
@@ -131,7 +131,9 @@ func (s *StartedService) StartTailscaleSSHSession(
continue
}
go ssh.DiscardRequests(reqs)
go s.forwardSSHAgentChannel(channel)
// lx:begin sec-sshagent
go s.forwardSSHAgentChannel(sessionCtx, channel)
// lx:end sec-sshagent
}
}()
}
@@ -313,7 +315,8 @@ func (s *StartedService) StartTailscaleSSHSession(
return nil
}
func (s *StartedService) forwardSSHAgentChannel(channel ssh.Channel) {
// lx:begin sec-sshagent
func (s *StartedService) forwardSSHAgentChannel(ctx context.Context, channel ssh.Channel) {
defer channel.Close()
fd, err := s.handler.ConnectSSHAgent()
if err != nil {
@@ -326,15 +329,33 @@ func (s *StartedService) forwardSSHAgentChannel(channel ssh.Channel) {
return
}
defer conn.Close()
// The ssh-agent conn stays blocked in Read while idle, so io.Copy(channel,
// conn) never returns on its own — without this it leaks a goroutine + the
// agent fd for every closed session. Cancelling on either copy finishing (or
// on the session ctx) closes both ends, unblocking the peer copy. Both Close
// calls are idempotent with the deferred ones above.
ctx, cancel := context.WithCancel(ctx)
defer cancel()
go func() {
<-ctx.Done()
conn.Close()
channel.Close()
}()
var wg sync.WaitGroup
wg.Add(2)
go func() {
defer wg.Done()
io.Copy(conn, channel)
cancel()
}()
go func() {
defer wg.Done()
io.Copy(channel, conn)
cancel()
}()
wg.Wait()
}
// lx:end sec-sshagent
Binary file not shown.

Before

Width:  |  Height:  |  Size: 419 KiB

+81 -15
View File
@@ -26,7 +26,9 @@ What it does:
Arg / env:
- `VERSION` — stamped into `constant.Version`. Resolution: positional arg →
`$SHATER_VERSION` → `git describe --tags` → `v0.2.0-dev`.
`$SHATER_VERSION` → `ci/version.sh --binary` → `v0.2.0-dev`. `ci/version.sh` is
the **same** computation the package version comes from (§2.1), so the string
the panel shows always matches what `apk info shaterd` / `opkg status` report.
- `--fast` — skip `npm ci` when `panel/node_modules` already exists.
- `UPX=/path/to/upx` — override the UPX binary (default `upx` on `PATH`). UPX is
cross-arch, so one host packs both the amd64 and aarch64 ELFs. (Note: UPX also
@@ -84,15 +86,52 @@ it. Because the binary is UPX-packed, the package disables the SDK's default str
feed installed and run `make package/shaterd/compile` (and the others) per target.
See `openwrt-package-build-ci` for SDK/feed mechanics.
### 2.1 Package versions come from the git tag
`PKG_VERSION`/`PKG_RELEASE` are **not** maintained by hand. They used to be, and
nobody bumped them: **v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3`** with
different binaries inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B).
Both package managers offer an upgrade only when the feed's version string
differs from the installed one, so `apk update` saw nothing new and the routers
could not be updated through the normal path at all.
`ci/version.sh` now derives them from `git describe`, once per CI job:
| Build | `PKG_VERSION` | `PKG_RELEASE` | `constant.Version` |
|---|---|---|---|
| tag push `v0.2.7` | `0.2.7` | `1` | `v0.2.7-r1` |
| dispatch, 3 commits past `v0.2.7` | `0.2.7` | `4` | `v0.2.7-r4-g<sha>` |
| no reachable tag / no git | `0.0.0` | `1` | `v0.0.0-r1` |
Ordering is what makes this safe, and both managers agree on it (checked with
`apk version -t` on apk-tools 3.0.3 and `opkg compare-versions` on opkg
38eccbb1): the dotted part decides first, `-rN` only breaks ties — so
`0.2.7-r1 > 0.2.6-r12 > 0.2.6-r1 > 0.2.0-r3`. A release therefore always
outranks every rolling build before it, rolling builds between two releases grow
monotonically, and an untagged build (`0.0.0`) can never masquerade as an
upgrade.
The value travels as `SHATER_PKG_VERSION`/`SHATER_PKG_RELEASE` in the SDK build
environment; the Makefiles read it with a literal fallback for manual/offline
builds. Both lanes then **assert** the produced `.ipk`/`.apk` really carries it,
so a lost variable fails the build instead of shipping a stale version.
`byedpi` is deliberately excluded — `PKG_VERSION:=0.17.3` is *upstream ByeDPI's*
version, which is what `PKG_HASH` pins and what tells you which ByeDPI is
installed. Stamping our tag on it would also be a downgrade: every comparator
reads `0.2.7 < 0.17.3` (component-wise, `2 < 17`). Bump its `PKG_RELEASE` by hand
when our packaging of it changes.
## 3. Install on a router
Install order follows the deps (`shaterd` → `shater-core` → `luci-app-shater`):
```sh
opkg install shaterd_0.2.0-1_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_0.2.0-1_all.ipk
opkg install luci-app-shater_0.2.0-1_all.ipk
opkg install byedpi_0.17.3-1_<arch>.ipk # optional: ByeDPI egress
# <ver> = the release version, e.g. 0.2.7-r1 (§2.1 — it comes from the git tag)
opkg install shaterd_<ver>_<arch>.ipk # or: apk add shaterd (25.12+)
opkg install shater-core_<ver>_all.ipk
opkg install luci-app-shater_<ver>_all.ipk
opkg install byedpi_0.17.3-r1_<arch>.ipk # optional: ByeDPI egress
```
Installing from a signed feed instead:
@@ -158,16 +197,25 @@ every `opkg update`; no `--nocheck-signature` needed. A **tagged** release
### Updating
Name the packages. **Never run a bare `opkg upgrade`** — with no arguments it
tries to upgrade *every* installed package from *every* configured feed, which on
OpenWrt means base/system packages on the overlay and is a well-known way to
brick a router.
```sh
opkg update
opkg upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
```
Updates are only offered when the feed's `Version` differs from the installed one,
so **bump `PKG_RELEASE`** (or `PKG_VERSION`) in the package Makefile on every
shipped change — otherwise `opkg upgrade` sees the same version and does nothing.
Do **not** `opkg upgrade` base/system packages from this feed; upgrade only the
four shater packages above.
Drop `byedpi` from the list if you never installed it. An upgrade is offered only
when the feed's `Version` differs from the installed one — that is exactly what
bug B4 broke (v0.2.2…v0.2.6 all published as `0.2.0-r3`). Since then CI derives
the version from the git tag on every build (§2.1), so there is nothing to bump
by hand any more; check with:
```sh
opkg list-installed | grep -E 'shaterd|shater-core|luci-app-shater|byedpi'
```
## 6. apk feed (OpenWrt/ImmortalWrt 25.12+ — incl. BananaWRT 25.12-mtk-vendor)
@@ -214,15 +262,33 @@ apk add byedpi # optional: ByeDPI desync egress
### Updating
**Never run a bare `apk upgrade`.** With no arguments apk reconciles *every*
installed package against *every* configured repository at once; on a router
whose distfeeds point at a moving snapshot that can pull in — or roll back —
unrelated system packages. Always name ours:
```sh
apk update
apk upgrade shaterd shater-core luci-app-shater byedpi # only our own packages
apk upgrade shaterd shater-core luci-app-shater byedpi
```
Same rule as opkg: an upgrade is only offered when the feed version differs, so
bump `PKG_RELEASE`/`PKG_VERSION` on every shipped change (apk shows it as
`0.2.0-r1`). Pin a version instead of tracking rolling by pointing the repo line
at `.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
apk-tools 3 documents exactly this behaviour for `apk upgrade`: *"When no
packages are specified, all packages are upgraded if possible. If list of
packages is provided, only those packages are upgraded along with needed
dependencies."* The equivalent form, which additionally re-pins the packages in
`world`, is:
```sh
apk add -u shaterd shater-core luci-app-shater byedpi # -u = --upgrade
```
Drop `byedpi` from either list if you never installed it. Check what you are on
with `apk list -I shaterd shater-core luci-app-shater byedpi` — the version reads
`0.2.7-r1` (§2.1: `PKG_VERSION-rPKG_RELEASE`, derived from the git tag by CI, so
every build really is a new version; before that fix v0.2.2…v0.2.6 all published
as `0.2.0-r3` and `apk update` offered nothing). Pin a version instead of tracking
rolling by pointing the repo line at
`.../download/apk-vX.Y.Z-$(cat /etc/apk/arch)/packages.adb`.
### BananaWRT `25.12-mtk-vendor` compatibility
+18
View File
@@ -0,0 +1,18 @@
# Документация shater
Документация продукта **shater** (управляемый интернет-шлюз для роутеров на
OpenWrt). Лицо репозитория и быстрый старт — в корневом [`../README.md`](../README.md).
| Документ | О чём |
|----------|-------|
| [CONTEXT.md](CONTEXT.md) | **Начните здесь** — контекст проекта, история v0.1→v0.2, решения в кратце, testbed/инфра |
| [INSTALL.md](INSTALL.md) | Сборка ship-артефакта (`shaterd`) и установка обоих фидов — opkg (24.10) и apk (25.12+) |
| [ARCHITECTURE.md](ARCHITECTURE.md) | One-binary дизайн, auth-handoff LuCI→панель, data/DNS/apply-потоки (диаграммы) |
| [FEATURES.md](FEATURES.md) | Полный список фич с тегами MVP/T1/T2 |
| [ROADMAP.md](ROADMAP.md) | Фазовый план |
| [DECISIONS.md](DECISIONS.md) | Почему sing-box, почему форк, split панели, лицензия и т.д. |
| [DESIGN.md](DESIGN.md) | Визуальная система панели — направление «Faceplate», токены, компоненты |
| [PORTING.md](PORTING.md) | Порт проверенных кусков из v0.1 |
Документация движка-форка (sing-box-lx) — в его слое: [`../docs-lx/`](../docs-lx/)
и [`../SPECS/`](../SPECS/).
@@ -0,0 +1,170 @@
# Живое тестирование shater v0.2.6 на mini_router
**Дата:** 2026-07-25
**Устройство:** Bananapi BPi-R3 Mini · ImmortalWrt **25.12-linkup** · `aarch64_cortex-a53`
**Установка:** из подписанного apk-фида `apk-v0.2.6-aarch64_cortex-a53`
**Пакеты:** `shaterd 0.2.0-r3`, `shater-core 0.2.0-r3`, `luci-app-shater 0.2.0-r2`, `byedpi 0.17.3-r1`
**Сборка:** CI run 61, коммит `024e9308c` (вершина `main`)
Сценарий: полное удаление предыдущей установки → чистая установка из фида →
проверка дефолтного состояния → восстановление рабочего конфига с подписками
(315 узлов) → функциональная проверка.
**Итог: 79 проверок, 74 PASS, 5 находок** (детали и разбор — в
`shater-bugs-2026-07-25.md` на рабочем столе).
---
## 1. Релиз и фид
| # | Проверка | Результат |
|---|---|---|
| T1 | Публикация `apk-v0.2.6-<arch>` для обеих архитектур | PASS |
| T2 | Ассеты: 4 `.apk` + `packages.adb` + `shater-apk.pem` | PASS |
| T3 | `apk update` принимает индекс (проверка EC-подписи) | PASS |
| T4 | Пакеты видны в нужных версиях (r3/r3/r2) | PASS |
| T5 | Диагностика сборки: `kmod packages selected (=m): 0` (было 1078) | PASS |
| T6 | Собраны ровно наши 4 пакета | PASS |
| T7 | opkg-лейн v0.2.6 (24.10) тоже зелёный | PASS |
## 2. Установка
| # | Проверка | Результат |
|---|---|---|
| T8 | `apk add luci-app-shater byedpi` — 4 пакета | PASS |
| T9 | Зависимости `kmod-nft-tproxy`/`kmod-nft-socket` из базового фида | PASS |
| T10 | Целостность: `apk manifest` = sha256 файла на диске | PASS |
| T11 | Установлен именно бинарь v0.2.6 (5 491 616 Б vs 5 488 336 Б в r2) | PASS |
| T12 | init-скрипты `shater`, `shater-cron` | PASS |
| T13 | `sysctl.d/99-shater.conf`, `hotplug.d/iface/99-shater` | PASS |
| T14 | boot-линки `S99shater`, `K10shater`, `S96shater-cron` | PASS |
## 3. Дефолтное состояние (чистая установка)
| # | Проверка | Результат |
|---|---|---|
| T15 | Дефолтный конфиг создан uci-defaults (27 строк) | PASS |
| T16 | `enabled='0'` — плоскость не ставится без согласия | PASS |
| T17 | Заготовлен tproxy-inbound на LAN, пресеты выключены | PASS |
| T18 | Демон стартует, `plane=none`, `table=false` | PASS |
| T19 | Права конфига `-rw-------` (0600) | PASS |
## 4. Восстановление рабочего конфига
| # | Проверка | Результат |
|---|---|---|
| T20 | Восстановление из бэкапа (3213 UCI-строк) | PASS |
| T21 | Кэш подписок цел: 315 узлов в 4 файлах | PASS |
| T22 | `shaterd migrate` → `ok`, схема v1 | PASS |
| T23 | Старт с реальным конфигом: `active`, `engine_running`, `plane=full` | PASS |
## 5. Data plane
| # | Проверка | Результат |
|---|---|---|
| T24 | Таблица `inet shater` создана (9 цепочек/сетов) | PASS |
| T25 | 16 tproxy-правил | PASS |
| T26 | `ip rule from all fwmark 0x2000 lookup shater` | PASS |
| T27 | `accept_local=1` на `br-lan` | PASS |
| T28 | DNS-divert: `dport 53 → tproxy :12345` для LAN-интерфейсов | PASS |
| T29 | DoT заблокирован: `dport 853 reject` | PASS |
| T30 | `block_doh=1`, правила присутствуют | PASS |
| T31 | **Kill-switch fail-closed**: цепочка `forward` завершается `drop` для LAN (v4+v6) | PASS |
| T32 | fw4 и dnsmasq не тронуты (свои таблицы целы) | PASS |
## 6. Панель и API
| # | Проверка | Результат |
|---|---|---|
| T33 | SPA отдаётся на `:8088` | PASS |
| T34 | `shaterd mint-token` выдаёт одноразовый токен | PASS |
| T35 | `/api/status` без сессии → **401** | PASS |
| T36 | `/api/session` (POST, JSON) → 200 + cookie `HttpOnly; SameSite=Strict; Max-Age=28800` | PASS |
| T37 | `/api/status` по cookie отдаёт данные, совпадающие с CLI | PASS |
| T38 | `/api/config` — 340 записей узлов | PASS |
| T39 | `/api/groups/health` — 103 протестировано, 13 живых, выбран `FR-vless-8` | PASS |
| T40 | `/api/devices` — устройства с IPv4/IPv6/MAC | PASS |
| T41 | `/api/interfaces` — `ewan/eth1 10.0.0.125/24 zone=wan` | PASS |
| T42 | `/api/ruleset/status` — remote-ruleset обновлён сегодня | PASS |
| T43 | `/api/stats` — memory backend, счётчики и top-domains | PASS |
| T44 | `/api/stats/log` — query-log с доменом, qtype, rcode, сервером | PASS |
| T45 | `/api/log?range=100` — пусто (следствие `log_file='0'`, не дефект) | OK |
## 7. Жизненный цикл конфигурации
| # | Проверка | Результат |
|---|---|---|
| T46 | `shaterd apply` → `{"changed":false}`, `can_rollback=true` | PASS |
| T47 | `shaterd confirm` снимает авто-откат (`can_rollback=false`) | PASS |
| T48 | `shaterd rollback` после confirm корректно сообщает об отсутствии last-good | PASS |
| T49 | `shaterd reconcile` (SIGHUP) не роняет движок | PASS |
| T50 | `shaterd sub update all-qomar` — реально обновил 143 узла | PASS |
| T51 | `shaterd blocklist update` → reconcile signalled | PASS |
| T52 | `shaterd schedule due` → reconcile signalled | PASS |
## 8. Устойчивость
| # | Проверка | Результат |
|---|---|---|
| T53 | `kill -9` демона → procd поднимает новый PID | PASS |
| T54 | После respawn: `engine_running=true`, `plane=full` | PASS |
| T55 | `stop` снимает таблицу `inet shater` полностью | PASS |
| T56 | `stop` → пауза → `start`: плоскость восстанавливается | PASS |
| T57 | Сеть при остановленном shater не деградирует | PASS |
| T58 | Память: 253 МБ занято из 2 ГБ при работающем движке | PASS |
## 9. DNS
| # | Проверка | Результат |
|---|---|---|
| T59 | Резолв через `127.0.0.1` | PASS |
| T60 | LAN-клиенты резолвят через движок (query-log растёт) | PASS |
| T61 | `.lan`-домены остаются за dnsmasq | PASS |
| T62 | dnsmasq жив и слушает на всех адресах | PASS |
| T63 | **Резолв через LAN-адрес `10.67.0.1` после `restart`** | **FAIL — B3** |
| T64 | Тот же резолв после `stop` → пауза → `start` | PASS |
## 10. Конфигурация и логи
| # | Проверка | Результат |
|---|---|---|
| T65 | 5 правил маршрутизации, 2 профиля, активен `ethernet-uplink` | PASS |
| T66 | **Два правила `default`, оба catch-all — нижнее живое, верхнее мертво** | **FAIL — B1** |
| T67 | **`shaterd nodes` всегда возвращает `[]`** | **FAIL — B2** |
| T68 | Логи уходят в syslog (`log_syslog=1`, 22 записи) | PASS |
| T69 | **ANSI-escape коды в syslog** | **FAIL — B5** |
| T70 | `loglevel=warning` соблюдается | PASS |
| T71–T79 | Прочие проверки состояния (статус-поля, права, uptime, счётчики, целостность таблиц) | PASS |
---
## Находки
| ID | Суть | Важность |
|---|---|---|
| **B1** | Два catch-all правила `default`; одно из них не работает никогда. **Поправка к первоначальному диагнозу:** правило без условий задаёт `route.Final`, а не выпускается как match-all, поэтому выигрывает ПОСЛЕДНЕЕ (`order=100 → group:auto`) — трафик идёт через прокси, а мёртвая настройка это `order=20 → direct` | средняя |
| **B2** | `shaterd nodes` — заглушка, всегда `[]`, хотя usage обещает список узлов (в кэше 315, в `/api/config` 340) | средняя |
| **B3** | После `service shater restart` резолв к LAN-адресу роутера не работает и не восстанавливается; `stop`+пауза+`start` — работает (гонка) | средняя |
| **B4** | `PKG_RELEASE` не менялся с v0.2.1 → v0.2.2…v0.2.6 выходят как `r3` при разном содержимом; `apk upgrade` не увидит обновления | средняя |
| **B5** | ANSI-раскраска попадает в syslog | низкая |
Разбор с воспроизведением — в `shater-bugs-2026-07-25.md`.
## История CI по этому релизу
Путь до зелёной сборки apk-лейна занял четыре итерации, каждая вскрывала
следующий слой одной причины:
| Тег | Что чинили | Итог |
|---|---|---|
| v0.2.2 | — (первый прогон с фиксами аудита) | `Disk quota exceeded`, 3593 `apk mkpkg kmod-*` |
| v0.2.3 | `.config` строится с нуля, а не дописывается | 1078 kmod — SDK вообще не везёт `.config` |
| v0.2.4 | Выключены `ALL`/`ALL_KMODS`/`ALL_NONSHARED` | 1078 kmod — они выбираются не через `ALL_KMODS` |
| v0.2.5 | Второй проход: явное `is not set` для каждого kmod | 1078 kmod — kconfig игнорирует user-значение у беспромптовых символов |
| **v0.2.6** | Удаление сгенерированных блоков `config PACKAGE_*` (`default m`) из `Config-build.in` | **0 kmod, сборка зелёная** |
Корень: `target/sdk/Makefile` генерирует `Config-build.in` прогоном
`convert-config.pl` по конфигу бильдбота, где `ALL_KMODS=y` уже развернулся в
`CONFIG_PACKAGE_kmod-*=m` на каждый модуль. Фильтр `next if /^(# )?CONFIG_PACKAGE/`
в скрипте стоит в ветке `else`, куда строка со знаком `=` не попадает, поэтому
каждый kmod приезжает в SDK как безусловный `default m`.
+9 -2
View File
@@ -2,6 +2,9 @@ package libbox
import (
"context"
// lx:begin sec-consttime
"crypto/subtle"
// lx:end sec-consttime
"errors"
"net"
"os"
@@ -97,9 +100,11 @@ func unaryAuthInterceptor(ctx context.Context, req any, info *grpc.UnaryServerIn
if len(values) == 0 {
return nil, status.Error(codes.Unauthenticated, "missing authentication secret")
}
if values[0] != sCommandServerSecret {
// lx:begin sec-consttime
if subtle.ConstantTimeCompare([]byte(values[0]), []byte(sCommandServerSecret)) != 1 {
return nil, status.Error(codes.Unauthenticated, "invalid authentication secret")
}
// lx:end sec-consttime
return handler(ctx, req)
}
@@ -115,9 +120,11 @@ func streamAuthInterceptor(srv any, ss grpc.ServerStream, info *grpc.StreamServe
if len(values) == 0 {
return status.Error(codes.Unauthenticated, "missing authentication secret")
}
if values[0] != sCommandServerSecret {
// lx:begin sec-consttime
if subtle.ConstantTimeCompare([]byte(values[0]), []byte(sCommandServerSecret)) != 1 {
return status.Error(codes.Unauthenticated, "invalid authentication secret")
}
// lx:end sec-consttime
return handler(srv, ss)
}
+8 -2
View File
@@ -74,7 +74,11 @@ func (r *oomReporter) WriteReport(memoryUsage uint64) error {
draftInfo = nil
}
reportsDir := filepath.Join(sWorkingPath, "oom_reports")
err = os.MkdirAll(reportsDir, 0o777)
// lx:begin sec-perms
// OOM reports embed the config snapshot (server secrets, keys) and logs;
// keep the tree owner-only (0700 dirs / 0600 files) instead of 0777/0666.
err = os.MkdirAll(reportsDir, 0o700)
// lx:end sec-perms
if err != nil {
return err
}
@@ -121,7 +125,9 @@ func discardDraftIfCurrent(draftPath string, draftInfo os.FileInfo) error {
func (r *oomReporter) writeSnapshot(destPath string, memoryUsage uint64) error {
now := time.Now().UTC()
err := os.MkdirAll(destPath, 0o777)
// lx:begin sec-perms
err := os.MkdirAll(destPath, 0o700)
// lx:end sec-perms
if err != nil {
return err
}
+7 -2
View File
@@ -44,7 +44,10 @@ func baseReportMetadata() reportMetadata {
func writeReportFile(destPath string, name string, content []byte) {
filePath := filepath.Join(destPath, name)
os.WriteFile(filePath, content, 0o666)
// lx:begin sec-perms
// Report files may carry the config snapshot (secrets) — owner-only.
os.WriteFile(filePath, content, 0o600)
// lx:end sec-perms
chownReport(filePath)
}
@@ -69,7 +72,9 @@ func copyConfigSnapshot(destPath string) {
}
func initReportDir(path string) {
os.MkdirAll(path, 0o777)
// lx:begin sec-perms
os.MkdirAll(path, 0o700)
// lx:end sec-perms
chownReport(path)
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 204 KiB

+11
View File
@@ -15,6 +15,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=byedpi
# DELIBERATELY NOT auto-versioned from our git tag (unlike shaterd/shater-core/
# luci-app-shater, which take SHATER_PKG_VERSION/SHATER_PKG_RELEASE from
# ci/version.sh). PKG_VERSION here is THIRD-PARTY UPSTREAM's version — it is what
# PKG_SOURCE_URL/PKG_HASH pin, and what tells an operator which ByeDPI is
# actually installed. Stamping our tag on it would be both a lie and a
# regression: our tags are 0.2.x, and every version comparator (apk-tools 3 and
# opkg alike, verified) reads 0.2.7 < 0.17.3 — component-wise numerically, 2 < 17
# — so the "new" package would be a DOWNGRADE and routers would refuse it.
# Bump PKG_RELEASE BY HAND when *our packaging* of it changes (init script, uci
# defaults, build flags); bump PKG_VERSION+PKG_HASH when upstream releases.
PKG_VERSION:=0.17.3
PKG_RELEASE:=1
+7 -2
View File
@@ -24,8 +24,13 @@ LUCI_TITLE:=LuCI thin launcher for Shater (mini dashboard + panel handoff)
LUCI_DEPENDS:=+shater-core +rpcd
LUCI_PKGARCH:=all
PKG_VERSION:=0.2.0
PKG_RELEASE:=2
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4). The literals are the manual/offline fallback only.
# These MUST stay above the luci.mk include: luci.mk only defaults PKG_VERSION/
# PKG_RELEASE when they are still unset, and the i18n subpackages inherit them.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+7 -2
View File
@@ -13,8 +13,13 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shater-core
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# Version comes from the git tag via ci/version.sh -> SHATER_PKG_VERSION /
# SHATER_PKG_RELEASE in the SDK build env (see openwrt/shaterd/Makefile for the
# full rationale — bug B4: v0.2.2…v0.2.6 all shipped as 0.2.0-r3). The literals
# are the manual/offline fallback only.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-2.0-or-later
+79 -5
View File
@@ -47,6 +47,13 @@ PROG=/usr/bin/shaterd
# hotplug/shater-cron touch the data plane. tmpfs => cleared by reboot, so
# nothing reconciles before this init has run at boot.
ACTIVE_FLAG=/var/run/shater.active
# Written by `shaterd run`; the single-owner token this init waits on so a
# restart never overlaps a new data plane with the previous one's teardown.
PIDFILE=/var/run/shaterd.pid
# Seconds `start` will wait for a predecessor to finish its teardown. Must be
# >= term_timeout below (procd's hard cap on a predecessor's life after SIGTERM)
# so we never give up while procd is still letting it shut down cleanly.
STOP_WAIT_SECS=40
# --- helpers ---------------------------------------------------------------
@@ -66,6 +73,54 @@ _slog() {
[ "$(uci -q get shater.globals.log_syslog)" = "0" ] || logger -t shater "$@"
}
# Echo the pid of a LIVE `shaterd run`, or fail. The pidfile is written by the
# daemon itself and removed only by the daemon that owns it, AFTER its teardown
# has completed — so "pidfile names a live process" is precisely "the previous
# data plane has not been dismantled yet".
shater_daemon_pid() {
local pid
pid=$(cat "$PIDFILE" 2>/dev/null) || return 1
[ -n "$pid" ] || return 1
kill -0 "$pid" 2>/dev/null || return 1
echo "$pid"
}
# Block until no predecessor daemon is left, bounded by STOP_WAIT_SECS.
#
# WHY THIS EXISTS. procd's `stop` is ASYNCHRONOUS: rc.common's `restart` is
# literally `stop; start`, and the `service delete` ubus call returns the moment
# procd has SENT SIGTERM — not when the instance is gone. `start` therefore
# re-adds the instance while the outgoing `shaterd run` is still executing its
# honest teardown (engine close, then `nft delete table`, `ip rule`/`ip route`
# removal and the per-iface sysctl restore). The result is that `restart` is NOT
# equivalent to `stop` + pause + `start`: the new plane is stood up on top of
# kernel state the old one has not finished removing, which is what B3 (DNS to
# the router's own LAN address dead after a restart, and never recovering) came
# out of. Waiting here restores the equivalence, and costs literally nothing when
# there is no predecessor — the check runs before the first sleep.
#
# Returning non-zero does NOT abort the start: the daemon carries its own
# single-owner guard and will refuse (or wait) on its side. Better to hand the
# decision to the process that can actually see the plane than to leave the box
# with no service at all.
shater_wait_stopped() {
local i=0 pid
pid=$(shater_daemon_pid) || return 0
_slog -p daemon.info \
"restart: waiting for the previous shaterd (pid $pid) to finish tearing the data plane down"
while [ "$i" -lt "$STOP_WAIT_SECS" ]; do
sleep 1
i=$((i + 1))
shater_daemon_pid >/dev/null || {
_slog -p daemon.info "restart: previous shaterd exited after ${i}s; starting a fresh one"
return 0
}
done
_slog -p daemon.warn \
"restart: previous shaterd (pid $pid) still alive after ${STOP_WAIT_SECS}s — starting anyway"
return 1
}
# --- procd lifecycle -------------------------------------------------------
start_service() {
@@ -87,6 +142,14 @@ start_service() {
return 0
fi
# Do not stand a new data plane up on top of one that is still being taken
# down. On `restart` procd has only just SIGTERMed the previous instance and
# returned; this is the handshake that makes `restart` == `stop` + pause +
# `start`. It also keeps `migrate` below from rewriting UCI underneath a
# daemon that is still reading it. No-op (and no delay) when nothing is
# running, which is the boot case.
shater_wait_stopped
# Bring the UCI schema forward before the daemon reads it (idempotent;
# refuses a newer schema) so an upgraded package never applies a stale config.
"$PROG" migrate >/dev/null 2>&1
@@ -111,7 +174,16 @@ start_service() {
procd_set_param stderr 1
# Give the daemon room to run its honest teardown (engine.Close + netplane
# restore) before procd SIGKILLs it.
procd_set_param term_timeout 10
#
# 30s, not 10s: an engine holding a few hundred outbounds closes its
# urltest/observatory goroutines and flushes experimental.cache_file to FLASH
# before the netplane teardown even starts, and on eMMC/NAND that alone can
# outlast 10s. A SIGKILL there aborts the teardown at an arbitrary point and
# leaves the plane HALF removed — the nft table gone but the policy routing
# still installed, or vice versa — which is precisely the class of leftover
# state the successor's idempotent fast-path cannot see and never repairs.
# Shutdown is bounded by procd either way; we are only choosing where.
procd_set_param term_timeout 30
procd_close_instance
# Mark the stack live for hotplug/cron — but ONLY when interception is
@@ -141,10 +213,12 @@ stop_service() {
reload_service() {
# Fired by the `shater` config.change reload-trigger (LuCI Save & Apply /
# reload_config). Simplest correct behaviour: stop + start. `stop` clears the
# flag and SIGTERMs the daemon (honest teardown); `start` re-guards on
# enabled and, if still enabled, launches a fresh `shaterd run` that reads
# the new UCI and applies it. When the stack is disabled, `start` is a no-op,
# so a disable+apply cleanly tears everything down.
# flag and SIGTERMs the daemon (honest teardown); `start` WAITS for that
# teardown to actually finish (shater_wait_stopped) and then launches a fresh
# `shaterd run` that reads the new UCI and applies it. When the stack is
# disabled, `start` is a no-op, so a disable+apply cleanly tears everything
# down. Because the wait lives in start_service, this path gets the same
# stop-then-start ordering guarantee as `restart`.
stop
start
}
+11 -2
View File
@@ -34,8 +34,17 @@
include $(TOPDIR)/rules.mk
PKG_NAME:=shaterd
PKG_VERSION:=0.2.0
PKG_RELEASE:=3
# VERSIONING — derived from the git tag, NOT hand-maintained here (bug B4).
# ci/version.sh turns `git describe` into SHATER_PKG_VERSION/SHATER_PKG_RELEASE
# (tag vX.Y.Z -> X.Y.Z + r1; off-tag -> last tag + r<commits+1>), and
# ci/build-feed.sh / ci/build-feed-apk.sh export them into the SDK build env of
# both lanes. Both lanes then ASSERT that the produced .ipk/.apk really carries
# that version, so a lost env can never silently ship a stale one again.
# The literals below are ONLY the manual/offline fallback (no CI, no git) — they
# are not "the release version"; releases are named by the tag.
PKG_VERSION:=$(if $(SHATER_PKG_VERSION),$(SHATER_PKG_VERSION),0.2.0)
PKG_RELEASE:=$(if $(SHATER_PKG_RELEASE),$(SHATER_PKG_RELEASE),1)
PKG_MAINTAINER:=Shater <maqrota@icloud.com>
PKG_LICENSE:=GPL-3.0-or-later
+44
View File
@@ -1187,6 +1187,50 @@ export function getStatsConns(q: number | StatsLogQuery = {}): Promise<ConnLogEn
return MOCK ? mock.getStatsConns(o) : req<ConnLogEntry[]>(`api/stats/conns${statsLogQS(o)}`)
}
/**
* One routing rule's reachability verdict — the rule analogue of
* {@link ChainHealth}.used: a quiet note about the ROUTING CONFIG, never a health
* signal.
*
* `unreachable` means the rule can NEVER take effect, whatever the traffic. Today
* the daemon reports exactly one certain case, and it is a subtle one: a rule with
* no conditions at all is not matched in sequence — it becomes the router's
* default. Two such rules therefore retire each other, and the LAST one by Order
* wins, so an earlier "default → direct" is dead even though it sorts first. A
* condition-less rule never retires a rule that HAS conditions: those are matched
* ahead of the default whatever their Order.
*
* `index` is the rule's position in GET /api/config's `Rules`, which is how a
* verdict is matched to a row — rule names are not unique, and the config that
* prompted this had two rules both called `default`. `name`/`order` are echoed so
* a page holding a verdict fetched before an edit can check it still describes the
* row it is about to badge, and drop it silently otherwise.
*/
export interface RuleReach {
index: number
name: string
order: number
unreachable: boolean
/** The rule that supersedes this one; absent when `unreachable` is false. */
shadowed_by?: string
/** Its index in `Rules`, or -1 when there is none. */
shadowed_by_index: number
shadowed_by_order?: number
/** Operator-facing sentence; absent when `unreachable` is false. */
reason?: string
}
/** GET /api/rules/reachability. `rules` is ALWAYS an array, one entry per rule in
* the same order as GET /api/config's `Rules`. */
export interface RulesReachability {
rules: RuleReach[]
}
/** GET /api/rules/reachability — which routing rules can never fire, and why. */
export function getRulesReachability(): Promise<RulesReachability> {
return MOCK ? mock.getRulesReachability() : req<RulesReachability>('api/rules/reachability')
}
/** GET /api/ruleset/status — remote rule-set / blocklist freshness + rule counts. */
export function getRulesetStatus(): Promise<RulesetStatus[]> {
return MOCK ? mock.getRulesetStatus() : req<RulesetStatus[]>('api/ruleset/status')
+51 -1
View File
@@ -6,7 +6,7 @@
// state mutates in-memory so the Apply / Confirm / Rollback flow is exercisable.
//
// Type-only imports from api.ts (erased at build) keep this free of a runtime cycle.
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
import type { ApplyResult, ChainHealth, ConnLogEntry, DiscoveredDevice, GroupHealth, GroupMemberHealth, GroupsHealth, GroupTestResult, GroupTestStart, GroupTestStatus, Interface, Model, QueryLogEntry, RuleReach, RulesReachability, RulesetCategories, RulesetCheck, RulesetStatus, Stats, StatsLogPage, StatsLogQuery, Status, StatusWarning } from './api'
let armed = false // a pending commit-confirm auto-rollback
let hasLastGood = false // a predecessor config exists to roll back to (post-apply)
@@ -146,6 +146,12 @@ const CONFIG: Model = {
{ Name: 'block-ads', Enabled: true, Order: 10, DstRuleset: ['ad-hosts'], Target: 'block' },
{ Name: 'ru-bypass', Enabled: true, Order: 20, DstRuleset: ['ru-inside'], Target: 'direct' },
{ Name: 'private-direct', Enabled: true, Order: 30, DstRuleset: ['private-nets'], Target: 'direct' },
// A SECOND condition-less rule, above the real default. It reads like a working
// rule and does nothing: a rule with no conditions becomes the router's default,
// and the last such rule by Order wins — so this one never applies. It is in the
// fixture on purpose, to exercise the "never applies" badge; the field config
// that prompted it had two rules BOTH named `default` (orders 20 and 100).
{ Name: 'default-bypass', Enabled: true, Order: 40, Target: 'direct' },
{ Name: 'default-tunnel', Enabled: true, Order: 900, Target: 'group:auto' },
],
// Named match-lists a rule points DstRuleset at. url + geosite + geoip are remote
@@ -296,6 +302,50 @@ const RULESET_STATUS: RulesetStatus[] = [
{ tag: 'rs-ru-geoip-ru', name: 'ru-geoip', category: 'ru', kind: 'ruleset', remote: true, last_updated: '', interval_seconds: 86_400, rule_count: 0 },
]
/** GET /api/rules/reachability. Mirrors the daemon's analysis over CONFIG.Rules:
* a rule with no conditions is the router's default, and the LAST such rule by
* Order wins — every earlier one can never apply. It reads the live CONFIG so
* edits made in `?mock` keep the badge honest. */
export async function getRulesReachability(): Promise<RulesReachability> {
await wait(60)
const rules = CONFIG.Rules ?? []
const out: RuleReach[] = rules.map((r, index) => ({
index,
name: String(r.Name ?? ''),
order: Number(r.Order ?? 0),
unreachable: false,
shadowed_by_index: -1,
}))
const conditionless = (r: (typeof rules)[number]): boolean =>
!(r.Src ?? []).length &&
!(r.DstDomain ?? []).length &&
!(r.DstRuleset ?? []).length &&
!(r.DstIP ?? []).length &&
!String(r.DstPort ?? '').trim() &&
!String(r.Proto ?? '').trim()
const target = (r: (typeof rules)[number]): string =>
String(r.Target ?? '').trim() || (r.Egress ? `egress:${String(r.Egress).trim()}` : '')
const defaults = rules
.map((r, index) => ({ r, index }))
.filter(({ r }) => r.Enabled && conditionless(r) && target(r))
.sort((a, b) => Number(a.r.Order ?? 0) - Number(b.r.Order ?? 0) || a.index - b.index)
const winner = defaults[defaults.length - 1]
if (winner) {
for (const { index } of defaults.slice(0, -1)) {
out[index].unreachable = true
out[index].shadowed_by = String(winner.r.Name ?? '')
out[index].shadowed_by_index = winner.index
out[index].shadowed_by_order = Number(winner.r.Order ?? 0)
out[index].reason =
`this rule has no conditions, so it sets the default for all traffic — but rule ` +
`"${winner.r.Name}" (order ${winner.r.Order}) has none either and comes after it, so ` +
`"${target(winner.r)}" is the default the router uses and this rule's target ` +
`"${target(rules[index])}" is never applied`
}
}
return { rules: out }
}
export async function getRulesetStatus(): Promise<RulesetStatus[]> {
await wait(90)
return RULESET_STATUS.map((r) => ({ ...r }))
+37
View File
@@ -252,6 +252,43 @@
color: var(--faint);
}
/* ---- a rule that can never fire (superseded by a later condition-less rule) ----
*
* Warn semantics only: --amber, never --accent. Orange is the ACTIVE state on this
* faceplate, and a rule the router ignores is the opposite of active — painting it
* orange is what made two `default` rows look equally live. It is a dashed amber
* frame, an amber order chip, and a dimmed target, so the row reads as "wired but
* not connected" without shouting: nothing is broken, one setting is just inert. */
.rt-rule.dead {
border-style: dashed;
border-color: color-mix(in srgb, var(--amber) 55%, var(--groove));
background: var(--panel);
box-shadow: none;
}
.rt-ord.dead {
color: var(--amber);
border-color: color-mix(in srgb, var(--amber) 45%, var(--groove));
}
.rt-badge.dead {
padding: 1px 7px;
border: 1px solid color-mix(in srgb, var(--amber) 55%, var(--groove));
border-radius: 999px;
background: color-mix(in srgb, var(--amber) 12%, transparent);
color: var(--amber);
}
.rt-dead-note {
font-family: var(--font-sans);
font-size: 11.5px;
line-height: 1.45;
color: var(--dim);
}
/* The target is still what the operator asked for, so it stays readable — just
* quiet, because the router is not using it. */
.rt-rule.dead .rt-target {
border-style: dashed;
opacity: 0.62;
}
/* ---- target chip (styled like the artifact's group:auto mono chips) ---- */
.rt-target {
display: inline-flex;
+85 -8
View File
@@ -6,11 +6,12 @@ import {
apply as apiApply,
getConfig,
putConfig,
getRulesReachability,
getRulesetStatus,
updateRuleset as apiUpdateRuleset,
ApiError,
} from '../api'
import type { Model, Rule, Ruleset, RulesetStatus } from '../api'
import type { Model, Rule, RuleReach, Ruleset, RulesetStatus } from '../api'
// ---------------------------------------------------------------------------
// The api.ts `Rule` is a deliberately thin subset (Name/Enabled/Order/Target/
@@ -334,6 +335,22 @@ export default function Routing() {
toastTimer.current = window.setTimeout(() => setToast(null), 2600)
}, [])
// Which rules can never fire, keyed by their position in the model's Rules array
// — NOT by name. The config that made this necessary had two rules both called
// `default`, which is exactly when a name-keyed verdict badges the wrong row.
const [reach, setReach] = useState<Map<number, RuleReach>>(new Map())
const loadReach = useCallback(async () => {
try {
const { rules } = await getRulesReachability()
setReach(new Map(rules.map((r) => [r.index, r])))
} catch {
// An older daemon has no such endpoint, and a stopped one answers nothing.
// Drop the verdicts rather than keep stale ones: no badge is honest, a badge
// about the previous config is not.
setReach(new Map())
}
}, [])
const load = useCallback(async () => {
try {
setLoadError(null)
@@ -341,7 +358,8 @@ export default function Routing() {
} catch (e) {
setLoadError(errMsg(e))
}
}, [])
void loadReach()
}, [loadReach])
useEffect(() => {
void load()
}, [load])
@@ -407,6 +425,36 @@ export default function Routing() {
// Every rule name, for the edit form's duplicate-name guard (it excludes self).
const ruleNames = useMemo(() => new Set(rules.map((r) => r.Name)), [rules])
// Each rule's position in the model's Rules array — the key the daemon's
// reachability verdicts use. `rules` above is a sorted COPY of the same object
// references, so identity survives the sort and this map stays valid.
const modelIndex = useMemo(() => {
const m = new Map<RRule, number>()
;((config?.Rules as RRule[] | null | undefined) ?? []).forEach((r, i) => m.set(r, i))
return m
}, [config])
/**
* The verdict for one rule, or null when it can fire.
*
* Verdicts are fetched separately from the config, so between an optimistic edit
* and the refetch they can describe the PREVIOUS rule list. Re-checking the
* echoed name and order is what stops that window from putting a "never applies"
* badge on a working rule: a mismatch means the verdict is not about this row,
* and no badge is the honest answer.
*/
const shadowOf = useCallback(
(r: RRule): { by: string; byOrder: number; reason: string } | null => {
const i = modelIndex.get(r)
if (i === undefined) return null
const v = reach.get(i)
if (!v || !v.unreachable || !v.shadowed_by) return null
if (v.name !== r.Name || v.order !== r.Order) return null
return { by: v.shadowed_by, byOrder: v.shadowed_by_order ?? 0, reason: v.reason ?? '' }
},
[modelIndex, reach],
)
// Rulesets are named domain/IP lists rules match against (rule.DstRuleset).
const rulesets = useMemo<Ruleset[]>(
() => [...((config?.Rulesets as Ruleset[] | null | undefined) ?? [])],
@@ -458,6 +506,10 @@ export default function Routing() {
await putConfig(next)
setSavedPending(true)
flash(okMsg)
// The verdicts describe the config on disk, which just changed — re-ask.
// Adding or moving a rule is precisely what turns a working default into a
// superseded one, and vice versa.
void loadReach()
} catch (e) {
setConfig(prev)
setActionError(errMsg(e))
@@ -466,7 +518,7 @@ export default function Routing() {
setSaving(false)
}
},
[config, flash],
[config, flash, loadReach],
)
const commitRules = useCallback(
@@ -744,7 +796,11 @@ export default function Routing() {
{rules.map((r, i) =>
editingRule === r.Name ? (
<RuleEditForm
key={r.Name}
// Rule names are NOT unique in the wild — the config that prompted
// the never-applies badge had two rules called `default`, and a
// duplicate React key makes the second row shadow the first. The
// model index disambiguates without changing row identity.
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
initial={r}
names={ruleNames}
targets={targets}
@@ -755,12 +811,13 @@ export default function Routing() {
/>
) : (
<RuleRow
key={r.Name}
key={`${modelIndex.get(r) ?? i}:${r.Name}`}
rule={r}
index={i}
total={rules.length}
busy={saving}
editingOther={editingRule !== null}
shadow={shadowOf(r)}
onEdit={onEditRule}
onToggle={onToggle}
onMove={onMove}
@@ -810,6 +867,7 @@ function RuleRow({
total,
busy,
editingOther,
shadow,
onEdit,
onToggle,
onMove,
@@ -820,6 +878,8 @@ function RuleRow({
total: number
busy: boolean
editingOther: boolean
/** Set when the daemon reports this rule can never fire; null when it can. */
shadow: { by: string; byOrder: number; reason: string } | null
onEdit: (name: string) => void
onToggle: (name: string) => void
onMove: (name: string, dir: 'up' | 'down') => void
@@ -829,13 +889,20 @@ function RuleRow({
// the first-match order can't shift under the open form. Edit itself stays live —
// clicking it just swaps which row is being edited.
const frozen = busy || editingOther
const isDefault = isCatchAll(rule)
// A rule with no conditions is the router's default — but only ONE of them can
// be, and the daemon says which. A superseded one must not wear the default's
// marks (the dashed accent frame, the "· final" order chip, the "everything not
// matched above" line): those are the claim that made two `default` rules
// indistinguishable in the first place.
const dead = shadow !== null
const isDefault = isCatchAll(rule) && !dead
const target = effectiveTarget(rule)
const tone = targetTone(target)
const cls = [
'rt-rule',
rule.Enabled ? '' : 'off',
isDefault ? 'final' : '',
dead ? 'dead' : '',
]
.filter(Boolean)
.join(' ')
@@ -852,7 +919,7 @@ function RuleRow({
>
▲
</button>
<span className={isDefault ? 'rt-ord final' : 'rt-ord'}>
<span className={isDefault ? 'rt-ord final' : dead ? 'rt-ord dead' : 'rt-ord'}>
{isDefault ? '·' : rule.Order}
</span>
<button
@@ -870,9 +937,19 @@ function RuleRow({
<div className="rt-head">
<span className="rt-name">{rule.Name}</span>
{isDefault && <span className="rt-badge">default route · final</span>}
{dead && <span className="rt-badge dead">never applies</span>}
</div>
<div className="rt-match">
{isDefault ? (
{dead ? (
// The badge says it never fires; this line says what beat it and what to
// do. Visible text, not a tooltip — the operator has to be able to find
// the other rule, and two rows can carry the same name.
<span className="rt-dead-note" title={shadow.reason}>
“{shadow.by}” (order {shadow.byOrder}) has no conditions either and runs after this
one, so it is the default the router uses. Give this rule a condition, or delete one
of the two.
</span>
) : isDefault ? (
<span className="rt-nomatch">everything not matched above</span>
) : (
<Matchers rule={rule} />
Binary file not shown.

Before

Width:  |  Height:  |  Size: 288 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 280 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 279 KiB

+96 -2
View File
@@ -16,6 +16,7 @@ import (
"github.com/sagernet/sing/common"
"github.com/sagernet/sing/common/buf"
"github.com/sagernet/sing/common/control"
E "github.com/sagernet/sing/common/exceptions" // lx:tproxy_writeback_connect
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
@@ -68,7 +69,26 @@ func (t *TProxy) Start(stage adapter.StartStage) error {
}
func (t *TProxy) Close() error {
return t.listener.Close()
err := t.listener.Close()
// lx:begin tproxy_writeback_connect
// Closing the listener stops INGRESS but leaves every live UDP NAT session in
// the cache, and each session holds a write-back socket bound to its original
// destination. Nothing else ever wakes those sessions: the cache evicts
// lazily (on the next Get/Add), and after this inbound is gone there is no
// next Get. For a DNS session answered in-engine by hijack-dns there is not
// even an outbound connection whose failure could unwind it, so its socket
// survives until a GC finalizer happens to reach it.
//
// That is invisible upstream, where an inbound is closed once at shutdown. On
// this fork the engine is REBUILT on every config apply, so each apply would
// strand another generation of sockets on the router's own LAN :53. Purge
// evicts every session; the cache's OnEvict closes the conn, which unblocks
// the session's routing goroutine and runs its onClose — the one place the
// write-back socket is actually closed. Purge AFTER the listener so a packet
// arriving mid-teardown cannot re-create a session behind us.
t.udpNat.Purge()
// lx:end tproxy_writeback_connect
return err
}
func (t *TProxy) NewConnection(ctx context.Context, conn net.Conn, metadata adapter.InboundContext, onClose N.CloseHandlerFunc) {
@@ -121,18 +141,92 @@ type tproxyPacketWriter struct {
conn *net.UDPConn
}
// lx:begin tproxy_writeback_connect
//
// The TPROXY UDP write-back socket is bound to the ORIGINAL DESTINATION, so the
// client sees the reply coming from the address it addressed. Upstream leaves
// that socket UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort), and that is
// the bug this block exists for.
//
// An unconnected socket bound to <addr>:<port> is, as far as the kernel is
// concerned, a RECEIVER for that address:port — and because the socket also
// carries SO_REUSEADDR it silently joins the UDP demultiplex set of whatever
// else is bound there. Nothing ever reads from it: this writer only sends. So
// every datagram the kernel happens to hand it is lost.
//
// On a router that transparently intercepts LAN DNS towards its OWN address
// (`nft ... udp dport 53 tproxy ...`) the original destination IS the router's
// LAN address, so each intercepted DNS session parks another silent receiver on
// <router-lan-ip>:53 right next to dnsmasq's socket — observed on the stand:
// 33 such sockets against dnsmasq's one, several with a growing Recv-Q. The
// host's own queries to that address take the loopback path, are never diverted,
// and are therefore demultiplexed among all of them: they land in one of the
// silent sockets at random and time out, permanently and unpredictably, while
// every other DNS path on the box keeps working.
//
// Connecting the socket fixes it at the root. compute_score() in the kernel's
// UDP lookup REJECTS a connected socket for any peer other than the connected
// one, so a write-back socket can no longer be handed a datagram it will not
// read. Nothing about the reply changes — same spoofed source address, same
// single peer, one Write instead of one WriteTo.
//
// The unconnected path is kept verbatim for a destination that cannot be bound
// (a domain socksaddr), so no existing case regresses.
// newTProxyWriteBack creates the connected write-back socket: local address =
// the original destination (transparent bind), peer = the client. A package
// variable so the behaviour can be tested without CAP_NET_ADMIN.
var newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
var dialer net.Dialer
dialer.LocalAddr = destination.UDPAddr()
dialer.Control = control.Append(dialer.Control, control.ReuseAddr())
dialer.Control = control.Append(dialer.Control, redir.TProxyWriteBack())
conn, err := w.listener.DialContext(dialer, w.ctx, "udp", w.source.String())
if err != nil {
return nil, err
}
udpConn, loaded := conn.(*net.UDPConn)
if !loaded {
conn.Close()
return nil, E.New("tproxy write back: unexpected connection type ", conn)
}
return udpConn, nil
}
// lx:end tproxy_writeback_connect
func (w *tproxyPacketWriter) WritePacket(buffer *buf.Buffer, destination M.Socksaddr) error {
defer buffer.Release()
if w.listener.ListenOptions().NetNs == "" {
conn := w.conn
if w.destination == destination && conn != nil {
_, err := conn.WriteToUDPAddrPort(buffer.Bytes(), w.source)
// lx:begin tproxy_writeback_connect
// The cached socket is CONNECTED to the client, so this is a plain
// Write. A failed write also closes it: upstream only dropped the
// reference, leaving the fd to the GC finalizer.
_, err := conn.Write(buffer.Bytes())
if err != nil {
conn.Close()
w.conn = nil
}
return err
// lx:end tproxy_writeback_connect
}
}
// lx:begin tproxy_writeback_connect
if destination.IsIP() {
udpConn, err := newTProxyWriteBack(w, destination)
if err != nil {
return err
}
if w.listener.ListenOptions().NetNs == "" && w.destination == destination {
w.conn = udpConn
} else {
defer udpConn.Close()
}
return common.Error(udpConn.Write(buffer.Bytes()))
}
// lx:end tproxy_writeback_connect
var listenConfig net.ListenConfig
listenConfig.Control = control.Append(listenConfig.Control, control.ReuseAddr())
listenConfig.Control = control.Append(listenConfig.Control, redir.TProxyWriteBack())
@@ -0,0 +1,234 @@
package redirect
// Regression cover for the lx:tproxy_writeback_connect block in tproxy.go.
//
// The failure it guards against, seen on a live BPi-R3 Mini: `shaterd` held 33
// UNCONNECTED UDP sockets on the router's own LAN address :53 — the write-back
// sockets of intercepted DNS sessions — alongside dnsmasq's single socket on the
// same address:port. Nothing reads a write-back socket, so every host-originated
// query that the kernel demultiplexed into one of them was silently dropped
// (Recv-Q climbing, sender timing out). Connecting the socket to the one peer it
// ever talks to removes it from the demultiplex set for everybody else.
//
// The real socket needs CAP_NET_ADMIN (IP_TRANSPARENT) and a foreign bind, so
// these tests drive the seam (newTProxyWriteBack) with an ordinary connected
// loopback socket. That is enough to pin both properties that actually broke:
//
// - the write path uses the CONNECTED form. If WritePacket ever goes back to
// WriteToUDPAddrPort, Go returns ErrWriteToConnected on a connected socket
// and these tests fail — i.e. the test cannot pass with an unconnected
// write-back socket, which is exactly the regression.
// - repeated writes to the same destination REUSE one socket instead of
// accumulating a new one per packet.
import (
"context"
"net"
"net/netip"
"testing"
"time"
"github.com/sagernet/sing-box/common/listener"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/buf"
M "github.com/sagernet/sing/common/metadata"
N "github.com/sagernet/sing/common/network"
"github.com/sagernet/sing/common/udpnat2"
)
// writeBackHarness stands up a loopback "client" socket and a tproxyPacketWriter
// whose socket factory returns a plain connected UDP socket aimed at it. It
// returns the writer, the client socket and a pointer to the factory call count.
func writeBackHarness(t *testing.T) (*tproxyPacketWriter, *net.UDPConn, *int) {
t.Helper()
client, err := net.ListenUDP("udp", &net.UDPAddr{IP: net.IPv4(127, 0, 0, 1)})
if err != nil {
t.Fatalf("client socket: %v", err)
}
t.Cleanup(func() { client.Close() })
source := netip.MustParseAddrPort(client.LocalAddr().String())
w := &tproxyPacketWriter{
ctx: context.Background(),
source: source,
// A listener with empty options is enough: WritePacket only reads
// ListenOptions().NetNs, and the factory is stubbed below.
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
destination: M.SocksaddrFrom(netip.MustParseAddr("127.0.0.1"), 53),
}
calls := 0
orig := newTProxyWriteBack
t.Cleanup(func() { newTProxyWriteBack = orig })
newTProxyWriteBack = func(w *tproxyPacketWriter, destination M.Socksaddr) (*net.UDPConn, error) {
calls++
// The real implementation binds the ORIGINAL DESTINATION transparently
// and connects to w.source; here we only reproduce the connect half,
// which is the property under test.
conn, err := net.DialUDP("udp", nil, net.UDPAddrFromAddrPort(w.source))
if err != nil {
return nil, err
}
return conn, nil
}
return w, client, &calls
}
// readOne reads one datagram from the client socket with a short deadline.
func readOne(t *testing.T, client *net.UDPConn) string {
t.Helper()
_ = client.SetReadDeadline(time.Now().Add(2 * time.Second))
b := make([]byte, 512)
n, _, err := client.ReadFrom(b)
if err != nil {
t.Fatalf("client read: %v", err)
}
return string(b[:n])
}
// TestWriteBackUsesConnectedSocket: the reply must go out over a CONNECTED
// socket. On the pre-fix code the write is WriteToUDPAddrPort, which Go refuses
// on a connected socket ("use of WriteTo with pre-connected connection"), so
// this test is red for exactly the shape that caused the outage.
func TestWriteBackUsesConnectedSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("first")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
if got := readOne(t, client); got != "first" {
t.Fatalf("client got %q, want %q", got, "first")
}
if *calls != 1 {
t.Fatalf("expected one write-back socket, got %d", *calls)
}
}
// TestWriteBackReusesOneSocket: a session that keeps answering the same client
// must keep ONE socket, not open a fresh one per datagram. The live box carried
// one socket per intercepted DNS session already; one per PACKET would turn a
// nuisance into an fd exhaustion.
func TestWriteBackReusesOneSocket(t *testing.T) {
w, client, calls := writeBackHarness(t)
for i, payload := range []string{"a", "b", "c", "d"} {
if err := w.WritePacket(buf.As([]byte(payload)).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket %d: %v", i, err)
}
if got := readOne(t, client); got != payload {
t.Fatalf("packet %d: client got %q, want %q", i, got, payload)
}
}
if *calls != 1 {
t.Fatalf("four packets to one destination must share one socket, got %d sockets", *calls)
}
if w.conn == nil {
t.Fatalf("the write-back socket must be cached on the writer for reuse")
}
}
// TestWriteBackClosesSocketOnWriteFailure: upstream dropped the reference to a
// failed socket without closing it, leaving the fd to the GC finalizer. On a box
// that already parks one socket per DNS session that is the wrong direction.
func TestWriteBackClosesSocketOnWriteFailure(t *testing.T) {
w, client, _ := writeBackHarness(t)
if err := w.WritePacket(buf.As([]byte("warm")).ToOwned(), w.destination); err != nil {
t.Fatalf("WritePacket: %v", err)
}
readOne(t, client)
cached := w.conn
if cached == nil {
t.Fatalf("expected a cached socket after the first write")
}
// Close it behind WritePacket's back so the next write fails, exactly as a
// dead peer or a torn-down plane would make it fail.
cached.Close()
if err := w.WritePacket(buf.As([]byte("boom")).ToOwned(), w.destination); err == nil {
t.Fatalf("a write on a closed socket must report the failure")
}
if w.conn != nil {
t.Fatalf("a failed write must drop the cached socket")
}
// A second Close on an already-closed conn is an error, which is how we know
// WritePacket closed it rather than merely forgetting it.
if err := cached.Close(); err == nil {
t.Fatalf("WritePacket must CLOSE the failed socket, not just nil the field")
}
}
// --- Close() must release the UDP NAT sessions --------------------------------
// natSessionHandler stands in for the router: it drains the session conn until
// it errors (which is what Close does to it) and then reports onClose, exactly
// as the real routing goroutine does. onClose is where the write-back socket is
// closed, so "onClose fired" is the observable proof the socket was released.
type natSessionHandler struct {
released chan error
}
func (h *natSessionHandler) NewPacketConnectionEx(ctx context.Context, conn N.PacketConn, source M.Socksaddr, destination M.Socksaddr, onClose N.CloseHandlerFunc) {
go func() {
for {
buffer := buf.NewSize(1024)
_, err := conn.ReadPacket(buffer)
buffer.Release()
if err != nil {
if onClose != nil {
onClose(err)
}
h.released <- err
return
}
}
}()
}
// TestTProxyCloseReleasesNatSessions pins the apply-level half of the leak.
//
// Closing the inbound used to close the listener only. The NAT cache evicts
// lazily, so after the inbound is gone nothing ever touches it again and every
// live session — with the write-back socket it holds on the router's own
// LAN :53 — was left to a GC finalizer. On this fork the engine is rebuilt on
// every config apply, so that is one stranded generation of sockets per apply.
func TestTProxyCloseReleasesNatSessions(t *testing.T) {
handler := &natSessionHandler{released: make(chan error, 1)}
tp := &TProxy{
ctx: context.Background(),
listener: listener.New(listener.Options{
Context: context.Background(),
Listen: option.ListenOptions{},
}),
}
prepared := 0
tp.udpNat = udpnat.New(handler, func(source M.Socksaddr, destination M.Socksaddr, userData any) (bool, context.Context, N.PacketWriter, N.CloseHandlerFunc) {
prepared++
return true, context.Background(), nil, func(error) {}
}, time.Minute, false)
tp.udpNat.NewPacket(
[][]byte{{0x00}},
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.2"), 40000),
M.SocksaddrFrom(netip.MustParseAddr("10.67.0.1"), 53),
nil,
)
if prepared != 1 {
t.Fatalf("expected one NAT session, got %d", prepared)
}
if err := tp.Close(); err != nil {
t.Fatalf("Close: %v", err)
}
select {
case <-handler.released:
case <-time.After(5 * time.Second):
t.Fatal("Close must release every live NAT session (and with it the write-back socket it holds)")
}
}
+6 -3
View File
@@ -159,7 +159,10 @@ func (r *Router) startIdleSuspend() error {
if period < idleTickFloor {
period = idleTickFloor
}
go r.idleSuspendLoop(period)
// Pass the stop channel by value so the loop never reads the r.idleStop field
// (stopIdleSuspend nils it right after close): a closed local channel stays
// ready in select, so the loop exits even when Close races an in-flight tick.
go r.idleSuspendLoop(period, r.idleStop)
return nil
}
@@ -176,12 +179,12 @@ func (r *Router) stopIdleSuspend() {
// asks each WG/AWG endpoint to suspend itself if it is unreachable AND idle past
// the threshold. Per endpoint this is a single map lookup + an atomic idle
// comparison; the endpoint owns the suspend/CAS/log decision.
func (r *Router) idleSuspendLoop(period time.Duration) {
func (r *Router) idleSuspendLoop(period time.Duration, stop <-chan struct{}) {
ticker := time.NewTicker(period)
defer ticker.Stop()
for {
select {
case <-r.idleStop:
case <-stop:
return
case <-ticker.C:
r.suspendIdleEndpoints(r.reachableOutbounds())
+9 -4
View File
@@ -27,8 +27,11 @@
#
# VERSION Version string stamped into constant.Version. Resolution order:
# 1) this positional arg, if given
# 2) $SHATER_VERSION, if set
# 3) `git describe --tags` (nearest tag + commit)
# 2) $SHATER_VERSION, if set (CI sets it from ci/version.sh)
# 3) `ci/version.sh --binary` — THE single source of truth shared
# with the package version (vX.Y.Z-rR[-g<sha>], derived from
# the git tag exactly like PKG_VERSION/PKG_RELEASE), so the
# string the panel shows always matches `apk info shaterd`
# 4) fallback: v0.2.0-dev
# --fast Skip `npm ci` when panel/node_modules already exists (dev speed-up).
#
@@ -69,8 +72,10 @@ if [ -n "$VERSION_ARG" ]; then
VERSION="$VERSION_ARG"
elif [ -n "${SHATER_VERSION:-}" ]; then
VERSION="$SHATER_VERSION"
elif VERSION="$(git -C "$REPO" describe --tags 2>/dev/null)"; then
: # git describe succeeded
elif VERSION="$(sh "$REPO/ci/version.sh" --binary)"; then
# Same computation the PACKAGE version comes from (ci/version.sh), so the
# binary's constant.Version and the .ipk/.apk version can never drift apart.
: # ci/version.sh always succeeds (it falls back to 0.0.0 without git)
else
VERSION="v0.2.0-dev"
fi
+10 -2
View File
@@ -163,10 +163,15 @@ func (t *adaptiveTimer) poll() {
return
}
if t.timerConfig.policyMode == policyModeNetworkExtension {
// lx:begin sec-oomcleanup
// A trigger schedules a deferred FreeOSMemory for the next poll; run it
// here and clear the flag (was inverted: it re-set true and so the free
// after a trigger never happened).
if t.cleanupTriggered {
runtimeDebug.FreeOSMemory()
t.cleanupTriggered = true
t.cleanupTriggered = false
}
// lx:end sec-oomcleanup
}
if t.pendingPressureBaseline {
t.pressureBaseline = sample
@@ -202,7 +207,10 @@ func (t *adaptiveTimer) poll() {
if !triggered {
return
}
t.cleanupTriggered = false
// lx:begin sec-oomcleanup
// Schedule the deferred FreeOSMemory for the next network-extension poll.
t.cleanupTriggered = true
// lx:end sec-oomcleanup
t.onTriggered(sample.usage)
if rateTriggered {
if t.killerDisabled {
+17
View File
@@ -488,6 +488,23 @@ func (a *Applier) applyLocked(m *model.Model) (bool, error) {
return changed, err
}
// The plane is COMPLETE only here: table + policy routing + sysctls. ApplyNft
// already flushed the DNS conntrack when it loaded the ruleset, but that is
// one step too early — ApplyRouting is idempotent BY del-then-add, so it opens
// a window in which the fwmark rule is momentarily absent, and any DNS flow
// that crosses that window is tracked against a plane that is still being
// assembled. Flushing once more now that every piece is in place is what makes
// "no entry survives the transition" actually true. Only on a real change (the
// fast path assembled nothing), best-effort, and cheap: the :53 entry count is
// bounded by the number of clients.
if !nftCurrent {
if n, ferr := netplane.FlushDNSConntrack(); ferr != nil {
a.log.Debug("flush DNS conntrack after plane change: ", ferr)
} else if n > 0 {
a.log.Debug("plane changed: dropped ", n, " stale DNS conntrack entries")
}
}
// (4) success. Bump the effective-state generation ONLY when something really
// moved: a no-op reconcile must not invalidate an armed commit-confirm window
// (cron reconciles every minute — counting those would cancel every rollback
+9
View File
@@ -91,6 +91,15 @@ var criticalMarkers = []string{
"not covered", // an interface outside the fail-closed guard
"fail-closed", // ''
"REJECTED", // an unusable interface name
// A condition-less rule retired by a later condition-less rule whose target is
// `direct` (generate/route.go warnUnreachableRules): the operator's default
// policy — a tunnel, or a block — is not the one the router uses, so everything
// unmatched leaves on the plain WAN. The marker is the full clause, not the
// shorter "leaves over the plain WAN" that ruleKillFallback's kill=open note
// also contains: that one is a DELIBERATE per-rule bypass the operator asked
// for, and grading it critical here would be a different decision made by
// accident.
"leaves over the plain WAN with your real IP address",
}
// protectionSections are entity kinds whose whole purpose is to block or divert
+42
View File
@@ -240,3 +240,45 @@ func blockGlobals() model.Globals {
g.Untunnelable = netplane.UntunnelableBlock
return g
}
// TestCollectWarningsGradesUnreachableRules (B1): a condition-less rule retired by
// a later condition-less one is graded by CONSEQUENCE, not by the mere fact that a
// setting is dead.
//
// - the surviving default is `direct` while the retired one wanted a tunnel:
// the operator's default policy is not in effect and everything unmatched
// leaves on the plain WAN — critical, the panel must show it loudly;
// - the surviving default is the tunnel and the dead one was `direct`: a dead
// knob, nothing is leaking — warning.
//
// Both must be attributed to section "rule" + the rule's name, so the panel can
// deep-link to the offending row instead of printing prose.
func TestCollectWarningsGradesUnreachableRules(t *testing.T) {
const leak = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "fallback" (order 100) has none either and comes after it, so "direct" wins and ` +
`this rule's target "group:auto" is never applied. Everything no other rule matches leaves ` +
`over the plain WAN with your real IP address. Delete one of the two, or give this one a condition`
const dead = `rule "default": it has no conditions, so it sets the default for ALL traffic — ` +
`but rule "default" (order 100) has none either and comes after it, so the default the ` +
`router uses is "group:auto" and this rule's target "direct" is never applied. Delete one ` +
`of the two, or give this one a condition so it can match something`
check := func(text, wantSeverity string) {
t.Helper()
found := false
for _, w := range collectWarnings(blockGlobals(), []string{text}, nil, nil) {
if w.Section != "rule" || w.Name != "default" {
continue
}
found = true
if w.Severity != wantSeverity {
t.Errorf("severity = %q, want %q for: %s", w.Severity, wantSeverity, text)
}
}
if !found {
t.Fatalf("never-applied warning not attributed to rule %q: %s", "default", text)
}
}
check(leak, SeverityCritical)
check(dead, SeverityWarning)
}
+27
View File
@@ -0,0 +1,27 @@
package main
// Control-plane log formatting (the IO-free half, so it is unit-testable on any
// dev host — same split as profilewatch.go).
//
// The daemon's own logger used to be built with a bare log.Formatter, whose
// DisableColors zero value is false: every control-plane line went to procd's
// stderr — i.e. straight into syslog — carrying aurora escape sequences. See
// shater/logsink/color.go for why that is a defect and not a cosmetic.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/shater/logsink"
)
// controlLogFormatter builds the control-plane logger's formatter. out is the
// process's real stderr (the sink's syslog half): colours are emitted only when
// that is a terminal, never when it is procd/syslog or a log file.
func controlLogFormatter(baseTime time.Time, out *os.File) log.Formatter {
return log.Formatter{
BaseTime: baseTime,
DisableColors: !logsink.IsTTY(out),
}
}
+37
View File
@@ -0,0 +1,37 @@
package main
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestControlLogFormatterNoANSIOffTTY pins the control-plane half of the syslog
// colour leak: when the daemon's stderr is not a terminal — which is ALWAYS the
// case under procd, where stderr is the pipe procd relays to syslog — no line
// the daemon formats may contain an ESC (0x1b).
func TestControlLogFormatterNoANSIOffTTY(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
f := controlLogFormatter(time.Now(), w)
if !f.DisableColors {
t.Fatalf("colours enabled for a non-terminal stderr")
}
// A context ID exercises the second colouring branch of log/format.go (the
// 256-colour connection id), which is what produced ESC[38;5;193m on the router.
ctx := log.ContextWithNewID(context.Background())
for _, level := range []log.Level{log.LevelError, log.LevelWarn, log.LevelInfo, log.LevelDebug, log.LevelTrace} {
line := f.Format(ctx, level, "dns", "exchange failed for example.com. IN AAAA: unexpected EOF", time.Now())
if i := strings.IndexByte(line, 0x1b); i >= 0 {
t.Errorf("level %v: formatted line carries an ANSI escape at byte %d: %q", level, i, line)
}
}
}
+76 -8
View File
@@ -79,7 +79,7 @@ func dispatch(args []string) int {
case "mint-token":
return cmdMintToken()
case "nodes":
return cmdReadStub("nodes", "[]")
return cmdNodes()
case "stats":
return cmdReadStub("stats", "{}")
case "blocklist":
@@ -124,7 +124,7 @@ usage: shaterd <verb>
rollback roll back to the last-good config
status print daemon/data-plane status as JSON
mint-token mint a single-use panel handoff token (JSON) via the daemon
nodes print nodes as JSON
nodes print the configured nodes as JSON (manual + subscription)
stats print stats as JSON
blocklist update force a DNS-filter blocklist refresh (reconcile; no-op if down)
schedule due re-evaluate time-scheduled rules now (reconcile; no-op if down)
@@ -158,8 +158,11 @@ func cmdRun() int {
// The level starts at trace (nothing read yet) and is corrected from
// Globals.LogLevel as soon as UCI is read — the control plane now RESPECTS
// the configured level instead of the old unconditional trace.
// The formatter colours only when stderr is a real terminal
// (controlLogFormatter -> logsink.IsTTY): under procd stderr IS syslog, and
// ANSI escapes there break `logread | grep ERROR` and every log collector.
logFactory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now()}, sink, "", nil, false)
controlLogFormatter(time.Now(), os.Stderr), sink, "", nil, false)
log.SetStdLogger(logFactory.Logger())
logger := log.StdLogger()
@@ -172,10 +175,23 @@ func cmdRun() int {
logFactory.SetLevel(controlLogLevel(g.LogLevel))
}
// Refuse to start a second daemon (race-safe single-owner guard). procd's
// term_timeout ensures the previous instance exits before reload restarts us.
if pid, ok := daemonAlive(); ok && pid != os.Getpid() {
logger.Error("another shaterd is already running (pid ", pid, ") — refusing to start")
// Single-owner guard: never run two daemons at once. On a `restart` procd's
// `service delete` is ASYNCHRONOUS — the ubus call returns immediately while
// the outgoing instance is still running its honest teardown — and the
// following `service add` starts us straight away, so a predecessor being
// alive here is the NORMAL restart case, not an error.
//
// This used to exit(1) on the spot and lean on procd's `respawn ... 5 ...` to
// try again five seconds later. That is a blind retry, not synchronisation:
// it neither knows nor waits for the predecessor's teardown to finish, and it
// turns every restart into at least one logged crash plus a five-second hole
// in which the LAN has no plane at all. Waiting for the predecessor to exit
// makes `restart` behave exactly like `stop` + pause + `start`: our apply is
// then strictly ordered AFTER the previous teardown, which is the whole point
// of the guard.
if pid, waiting := waitForPredecessor(predecessorBudget, predecessorPoll, logger); waiting {
logger.Error("another shaterd is still running (pid ", pid, ") after waiting ",
predecessorBudget, " for it to exit — refusing to start")
return 1
}
if err := writePidfile(); err != nil {
@@ -832,6 +848,42 @@ func cmdStatus() int {
return 0
}
// cmdNodes prints the node inventory as a JSON array (see nodes.go for the
// shape and for why this is a view rather than the raw model.Node).
//
// Shape of the call mirrors cmdStatus: ask the RUNNING daemon first — it is the
// process that owns the engine, so its answer is the inventory the live box was
// built from — and fall back to reading the same on-disk desired state directly
// when there is no daemon (or it did not answer). Both sides call the SAME
// nodesJSON(), so the fallback cannot report something the daemon would not.
//
// The one thing it will never do is print `[]` because it could not find out:
// a read failure goes to stderr and exits non-zero, so "empty" on stdout with
// exit 0 means "no nodes are configured" and nothing else.
func cmdNodes() int {
if _, ok := daemonAlive(); ok {
resp, err := ctlRequest("nodes")
switch {
case err != nil:
fmt.Fprintf(os.Stderr, "shaterd nodes: %v — reading the on-disk config instead\n", err)
case strings.HasPrefix(strings.TrimSpace(resp), "["):
fmt.Println(strings.TrimSpace(resp))
return 0
default:
// The daemon answered, but with an error object rather than a list.
fmt.Fprintf(os.Stderr, "shaterd nodes: daemon: %s\n", strings.TrimSpace(resp))
}
}
b, err := nodesJSON()
if err != nil {
fmt.Fprintf(os.Stderr, "shaterd nodes: %v\n", err)
fmt.Println("[]") // stdout stays parseable JSON; the exit code carries the failure
return 1
}
fmt.Println(string(b))
return 0
}
// cmdMintToken asks the RUNNING daemon (over the control socket) for a single-use
// panel handoff token and prints the daemon's JSON reply verbatim — {"token":"..."}
// on success, {"error":"..."} otherwise. It is the CLI shim the LuCI/rpcd layer
@@ -869,6 +921,13 @@ func printJSONError(msg string) {
fmt.Println(string(b))
}
// cmdReadStub relays a read-only verb to the daemon and prints def when there is
// no daemon to ask. It is ONLY valid where def is the truth in the daemon-down
// case: `stats` counts what the LIVE engine saw, so with no engine there really
// is nothing counted and `{}` says exactly that. It is NOT a way to make a verb
// look implemented — `nodes` used to be routed through here with def="[]" while
// hundreds of nodes sat on disk (see nodes.go). Anything whose data outlives the
// daemon belongs in its own command that reads that data.
func cmdReadStub(cmd, def string) int {
if _, ok := daemonAlive(); ok {
if resp, err := ctlRequest(cmd); err == nil {
@@ -940,7 +999,16 @@ func handleCtl(conn net.Conn, a *apply.Applier, ps *panel.Server, sa stats.Stats
case "rollback":
writeResult(conn, false, a.Rollback())
case "nodes":
writeLine(conn, "[]")
// Same assembly as the CLI fallback and as GET /api/config: model.ReadUCI
// (UCI + the per-subscription caches) projected onto the printable view.
b, err := nodesJSON()
if err != nil {
l.Warn("control socket: nodes: ", err)
e, _ := json.Marshal(map[string]string{"error": err.Error()})
writeLine(conn, string(e))
return
}
writeLine(conn, string(b))
case "stats":
if sa == nil {
writeLine(conn, "{}")
+132
View File
@@ -0,0 +1,132 @@
package main
// `shaterd nodes` — the node inventory, as JSON.
//
// This verb used to answer `[]` unconditionally (a `cmdReadStub("nodes", "[]")`
// on the CLI side and a hard-coded `writeLine(conn, "[]")` in the daemon's
// control-socket handler), while `shaterd --help` advertised "print nodes as
// JSON". The data was never missing: /etc/shater/subs/*.json holds every
// subscription-fetched node and `GET /api/config` reports the full merged set —
// on a live router that is hundreds of nodes answered as "none". An empty list
// is indistinguishable from a truthful "no nodes are configured", so the verb
// did not fail loudly, it lied quietly. Same class as the honesty fixes in
// 9dc954029 / aec82d444; the cure is the same: report what is actually there.
//
// # Data path (deliberately the SAME one /api/config uses)
//
// model.ReadUCI() = `uci export shater` (manual nodes + everything else) +
// model.MergeSubCaches (the per-subscription JSON caches). That single call is
// what the panel's handleConfigGet serves, what generate builds the engine from
// and what `sub update` writes back — so `shaterd nodes` cannot drift from the
// panel or from the running engine, because there is no second assembly here to
// drift. This file only PROJECTS that model onto a small, printable view.
//
// # Why a view and not the raw model.Node
//
// model.Node carries the share-link URI, and a share link is a credential
// (uuid/password in the query string). /api/config may return it — that path is
// session-authenticated and the panel needs the URI to edit a node — but a CLI
// verb whose output gets piped into support tickets, `logger`, and cron mail
// must not spray credentials. The view therefore reports the *derived* facts a
// share link answers (protocol, server, port) and drops the secret-bearing URI.
// Nodes whose URI does not parse are still listed, with the parse error in
// `parse_error`: the engine skips exactly those nodes (generate/outbound.go),
// and a node that is configured-but-unusable is precisely what an operator
// needs to see — hiding it would be the same lie in a smaller coat.
import (
"encoding/json"
"strings"
"github.com/sagernet/sing-box/shater/model"
"github.com/sagernet/sing-box/shater/parse"
)
// nodeView is ONE node as the `nodes` verb reports it. Field set: the model's
// own facts (name/enabled/sub/egress/stale/fingerprint) plus what the share link
// decodes to (protocol/server/port). Everything is omitempty-free where a reader
// would have to distinguish "absent" from "false"/"zero" — `enabled` and `stale`
// are always present so a consumer never has to guess.
type nodeView struct {
Name string `json:"name"`
Enabled bool `json:"enabled"`
// Sub is the subscription this node came from; "" means a manual node
// (model.Node.FromSub semantics, unchanged).
Sub string `json:"sub"`
// Protocol is the parsed protocol (vless|vmess|trojan|shadowsocks|wireguard…).
// When the URI does not parse it falls back to the bare URI scheme so the
// operator still sees what kind of thing failed; "" only when the URI is empty.
Protocol string `json:"protocol"`
Server string `json:"server,omitempty"`
Port uint16 `json:"port,omitempty"`
// Egress is the `config egress` this node's own upstream is bound to
// (multi-WAN); "" = default route.
Egress string `json:"egress,omitempty"`
// Stale marks a cached subscription node whose subscription failed to refresh.
Stale bool `json:"stale"`
Fingerprint string `json:"fingerprint,omitempty"`
// ParseError is the reason the engine will SKIP this node, verbatim from
// parse.ParseShareLink. Empty on every usable node.
ParseError string `json:"parse_error,omitempty"`
}
// readModel is the model source. A package var so the tests can exercise the
// assembly without a router's `uci` binary; production always reads the real
// merged desired state.
var readModel = model.ReadUCI
// nodesJSON reads the desired state and renders the node inventory as a JSON
// array. The array is never `null`: an empty configuration marshals to `[]`,
// which is the ONE case where `[]` is the truth.
func nodesJSON() ([]byte, error) {
m, err := readModel()
if err != nil {
return nil, err
}
return json.Marshal(nodeViews(m))
}
// nodeViews projects the merged model onto the printable view. Pure: no IO, no
// globals — the whole verb's logic is testable on a dev host.
func nodeViews(m *model.Model) []nodeView {
out := make([]nodeView, 0, len(m.Nodes)) // never nil => never `null`
for i := range m.Nodes {
n := m.Nodes[i]
v := nodeView{
Name: n.Name,
Enabled: n.Enabled,
Sub: n.FromSub,
Egress: n.Egress,
Stale: n.Stale,
Fingerprint: n.Fingerprint,
Protocol: uriScheme(n.URI),
}
// The engine reads a node through exactly this call (generate/outbound.go,
// chain.go, group.go); using it here is what makes "protocol" agree with
// what will actually be dialled, and the error agree with what will be
// skipped.
if p, err := parse.ParseShareLink(n.URI); err == nil {
if p.Protocol != "" {
v.Protocol = p.Protocol
}
v.Server = p.Server
v.Port = p.Port
} else if n.URI != "" {
v.ParseError = err.Error()
}
out = append(out, v)
}
return out
}
// uriScheme returns the lower-cased scheme of a share link ("vless://…" ->
// "vless"), or "" when there is none. It is the fallback protocol for a URI that
// ParseShareLink refuses, so an unusable node is still described rather than
// reported as a typeless blank.
func uriScheme(uri string) string {
i := strings.Index(uri, "://")
if i <= 0 {
return ""
}
return strings.ToLower(uri[:i])
}
+146
View File
@@ -0,0 +1,146 @@
package main
import (
"encoding/json"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// vlessURI is a syntactically complete share link (host, port, uuid, fragment)
// so ParseShareLink really succeeds and the derived fields are exercised.
const vlessURI = "vless://11111111-2222-3333-4444-555555555555@example.net:443?" +
"encryption=none&security=tls&sni=example.net&type=ws&path=%2Fws#NL-vless-1"
// writeSubCache points the subscription cache at a temp dir and persists one
// subscription's nodes there, the way `sub update` does on the router
// (/etc/shater/subs/<sub>.json). Returns nothing: the point is the side effect
// that model.MergeSubCaches will pick up.
func writeSubCache(t *testing.T, sub string, nodes []model.Node) {
t.Helper()
t.Setenv("SHATER_SUBS_DIR", t.TempDir())
if err := model.SaveSubCache(sub, nodes); err != nil {
t.Fatalf("SaveSubCache(%q): %v", sub, err)
}
}
// TestNodesJSONReportsCachedSubscriptionNodes is the regression for the bug this
// verb had: `shaterd nodes` answered `[]` while the subscription cache held
// hundreds of nodes. A NON-EMPTY cache must produce a NON-EMPTY list.
func TestNodesJSONReportsCachedSubscriptionNodes(t *testing.T) {
writeSubCache(t, "all-qomar", []model.Node{
{Name: "NL-vless-1", Enabled: true, URI: vlessURI, FromSub: "all-qomar", Fingerprint: "66b33"},
{Name: "NL-vless-2", Enabled: false, URI: vlessURI, FromSub: "all-qomar"},
})
// The model source stands in for `uci export shater` (absent on a dev host);
// the subscription half is the REAL model.MergeSubCaches path.
restore := readModel
readModel = func() (*model.Model, error) {
m := &model.Model{Globals: model.DefaultGlobals()}
model.MergeSubCaches(m)
return m, nil
}
t.Cleanup(func() { readModel = restore })
b, err := nodesJSON()
if err != nil {
t.Fatalf("nodesJSON: %v", err)
}
var got []nodeView
if err := json.Unmarshal(b, &got); err != nil {
t.Fatalf("nodesJSON produced invalid JSON %q: %v", b, err)
}
if len(got) == 0 {
t.Fatalf("nodes reported an EMPTY list while the subscription cache holds 2 nodes: %s", b)
}
if len(got) != 2 {
t.Fatalf("got %d nodes, want 2: %s", len(got), b)
}
byName := map[string]nodeView{}
for _, v := range got {
byName[v.Name] = v
}
first, ok := byName["NL-vless-1"]
if !ok {
t.Fatalf("cached node NL-vless-1 missing from %s", b)
}
if !first.Enabled {
t.Errorf("NL-vless-1.enabled = false, want true")
}
if first.Sub != "all-qomar" {
t.Errorf("NL-vless-1.sub = %q, want %q", first.Sub, "all-qomar")
}
if first.Protocol != "vless" {
t.Errorf("NL-vless-1.protocol = %q, want %q", first.Protocol, "vless")
}
if first.Server != "example.net" || first.Port != 443 {
t.Errorf("NL-vless-1 server:port = %s:%d, want example.net:443", first.Server, first.Port)
}
if first.ParseError != "" {
t.Errorf("NL-vless-1.parse_error = %q, want empty", first.ParseError)
}
if byName["NL-vless-2"].Enabled {
t.Errorf("NL-vless-2.enabled = true, want false (the model says disabled)")
}
// The credential-bearing share link must not be printed.
if strings.Contains(string(b), "11111111-2222-3333-4444-555555555555") {
t.Errorf("nodes output leaks the share-link uuid: %s", b)
}
}
// TestNodeViewsShape pins the projection: manual vs subscription, an unparsable
// URI still being LISTED (with the reason), and an empty model marshalling to
// `[]` rather than `null`.
func TestNodeViewsShape(t *testing.T) {
m := &model.Model{Nodes: []model.Node{
{Name: "manual-1", Enabled: true, URI: vlessURI, Egress: "wan2"},
{Name: "broken", Enabled: true, URI: "nosuch://whatever", FromSub: "qomar", Stale: true},
{Name: "no-uri", Enabled: false},
}}
got := nodeViews(m)
if len(got) != 3 {
t.Fatalf("nodeViews returned %d views, want 3", len(got))
}
if got[0].Sub != "" {
t.Errorf("manual node sub = %q, want empty", got[0].Sub)
}
if got[0].Egress != "wan2" {
t.Errorf("manual node egress = %q, want wan2", got[0].Egress)
}
if got[1].ParseError == "" {
t.Errorf("unparsable node must carry the reason the engine will skip it")
}
if got[1].Protocol != "nosuch" {
t.Errorf("unparsable node protocol = %q, want the bare scheme %q", got[1].Protocol, "nosuch")
}
if !got[1].Stale {
t.Errorf("stale flag lost")
}
if got[2].Protocol != "" || got[2].ParseError != "" {
t.Errorf("URI-less node: got protocol=%q parse_error=%q, want both empty", got[2].Protocol, got[2].ParseError)
}
b, err := json.Marshal(nodeViews(&model.Model{}))
if err != nil {
t.Fatalf("marshal empty: %v", err)
}
if string(b) != "[]" {
t.Errorf("empty model marshalled to %q, want []", b)
}
}
func TestURIScheme(t *testing.T) {
for _, tc := range []struct{ in, want string }{
{"vless://x@h:443", "vless"},
{"VMESS://payload", "vmess"},
{"", ""},
{"not-a-uri", ""},
{"://leading", ""},
} {
if got := uriScheme(tc.in); got != tc.want {
t.Errorf("uriScheme(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
+63
View File
@@ -0,0 +1,63 @@
package main
// The startup handshake that stops a restarting daemon from standing a new data
// plane up on top of the previous one's teardown.
//
// Deliberately free of build tags: the logic is pure timing and is exercised by
// the tests on any host, while the probe it polls (pidfile + kill(pid, 0)) is
// linux-only and is installed by predecessor_linux.go.
import (
"os"
"time"
"github.com/sagernet/sing-box/log"
)
// predecessorBudget / predecessorPoll bound the wait for an outgoing daemon.
//
// The budget must comfortably exceed the slowest honest teardown (engine.Close
// of a box with a few hundred outbounds plus the cache-file flush to flash, then
// the nft/ip/sysctl removal) AND the init script's procd term_timeout, which is
// the hard cap on how long a predecessor can live after its SIGTERM. Overshooting
// costs nothing on a healthy box — the wait ends the instant the predecessor is
// gone — while undershooting reintroduces the very overlap this exists to remove.
//
// Variables, not constants, so the tests can drive the loop without sleeping.
var (
predecessorBudget = 60 * time.Second
predecessorPoll = 100 * time.Millisecond
)
// aliveProbe is the seam waitForPredecessor polls. The portable default reports
// "no predecessor", so a non-linux build (where there is no daemon at all) never
// waits; predecessor_linux.go replaces it with the real pidfile probe.
var aliveProbe = func() (int, bool) { return 0, false }
// waitForPredecessor blocks until no OTHER shaterd owns the pidfile, or until the
// budget runs out.
//
// Returns (pid, true) only when a predecessor is STILL alive once the budget
// expires — the genuine "two daemons" error the caller refuses on. A predecessor
// that exits within the budget (the restart case) returns (0, false) and startup
// continues, now guaranteed to be sequenced after its teardown, exactly as it is
// after a manual `stop` + pause + `start`.
func waitForPredecessor(budget, poll time.Duration, logger log.ContextLogger) (int, bool) {
pid, ok := aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
if logger != nil {
logger.Info("a previous shaterd (pid ", pid, ") is still shutting down — ",
"waiting for its teardown to finish before applying a new data plane")
}
deadline := time.Now().Add(budget)
for time.Now().Before(deadline) {
time.Sleep(poll)
pid, ok = aliveProbe()
if !ok || pid == os.Getpid() {
return 0, false
}
}
return pid, true
}
+9
View File
@@ -0,0 +1,9 @@
//go:build linux
package main
// Bind the portable predecessor wait to the real single-owner probe. Split out
// of main.go so waitForPredecessor itself stays build-tag-free and testable on
// any developer host.
func init() { aliveProbe = daemonAlive }
+106
View File
@@ -0,0 +1,106 @@
package main
// B3 regression, daemon half.
//
// `/etc/init.d/shater restart` is `stop; start`, and procd's `stop` is
// asynchronous: the ubus `service delete` returns as soon as SIGTERM has been
// SENT, so `start` re-adds the instance while the outgoing `shaterd run` is
// still executing its honest teardown. The startup guard used to exit(1) the
// moment it saw a live predecessor and rely on procd's `respawn ... 5 ...` to
// try again later — a blind retry that neither knows nor waits for the teardown
// to finish, and that turns every restart into a logged crash plus a five-second
// hole with no data plane.
//
// The contract these pin: startup BLOCKS until the predecessor is gone (so our
// apply is strictly ordered after its teardown, exactly as it is after a manual
// `stop` + pause + `start`), and only refuses when the predecessor outlives the
// whole budget.
import (
"os"
"testing"
"time"
)
// scriptAlive installs an aliveProbe that reports a live predecessor for the
// first n calls and "gone" afterwards, and restores the real probe.
func scriptAlive(t *testing.T, pid, n int) *int {
t.Helper()
orig := aliveProbe
t.Cleanup(func() { aliveProbe = orig })
calls := 0
aliveProbe = func() (int, bool) {
calls++
if calls <= n {
return pid, true
}
return 0, false
}
return &calls
}
// TestWaitForPredecessorWaitsForTeardown: a predecessor that is still tearing
// the plane down must be WAITED for, not refused. Before the fix this returned
// "still alive" on the first probe and the daemon exited 1.
func TestWaitForPredecessorWaitsForTeardown(t *testing.T) {
calls := scriptAlive(t, 4242, 3)
pid, stillAlive := waitForPredecessor(2*time.Second, time.Millisecond, nil)
if stillAlive {
t.Fatalf("a predecessor that exits within the budget must not be refused (pid %d)", pid)
}
if *calls < 4 {
t.Fatalf("expected the guard to keep probing until the predecessor was gone, got %d probes", *calls)
}
}
// TestWaitForPredecessorRefusesAfterBudget: the single-owner invariant is kept —
// a predecessor that never dies still ends in a refusal, it is just no longer
// the FIRST answer.
func TestWaitForPredecessorRefusesAfterBudget(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return 4242, true }
start := time.Now()
pid, stillAlive := waitForPredecessor(30*time.Millisecond, time.Millisecond, nil)
if !stillAlive || pid != 4242 {
t.Fatalf("an immortal predecessor must still be refused, got pid=%d alive=%v", pid, stillAlive)
}
if time.Since(start) < 30*time.Millisecond {
t.Fatalf("the guard must exhaust its budget before refusing")
}
}
// TestWaitForPredecessorNoPredecessorIsFree: the boot path must pay nothing —
// one probe, no sleep. A wait that cost a tick on every clean start would show up
// as a slower boot for no reason.
func TestWaitForPredecessorNoPredecessorIsFree(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
calls := 0
aliveProbe = func() (int, bool) { calls++; return 0, false }
start := time.Now()
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("no predecessor must not be reported as alive")
}
if calls != 1 {
t.Fatalf("expected exactly one probe when nothing is running, got %d", calls)
}
if time.Since(start) > 500*time.Millisecond {
t.Fatalf("the no-predecessor path must not sleep")
}
}
// TestWaitForPredecessorIgnoresOwnPid: a pidfile naming THIS process (a crashed
// predecessor whose pid we were handed, or a re-exec) is not a predecessor.
func TestWaitForPredecessorIgnoresOwnPid(t *testing.T) {
orig := aliveProbe
defer func() { aliveProbe = orig }()
aliveProbe = func() (int, bool) { return os.Getpid(), true }
if _, stillAlive := waitForPredecessor(time.Minute, time.Second, nil); stillAlive {
t.Fatalf("our own pid must never count as a predecessor")
}
}
+29 -11
View File
@@ -67,7 +67,7 @@ package generate
import (
"fmt"
"net/netip"
"sort"
"os"
"strconv"
"strings"
"time"
@@ -75,6 +75,7 @@ import (
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing/common/json/badoption"
"github.com/sagernet/sing-box/shater/logsink"
"github.com/sagernet/sing-box/shater/model"
)
@@ -333,13 +334,31 @@ func Warnings(m *model.Model) []string {
// it, so a level there could never be honoured and would only invite the reader to
// think it was. Note that an EMPTY Level on an ENABLED log means LevelTrace, not
// "default" — so Disabled must never be emitted speculatively.
//
// DisableColor is the OTHER half of the engine's log block, and it is not
// cosmetic. log/log.go turns it into Formatter.DisableColors, which is what
// stops aurora from painting the level word and the per-connection ID. Left at
// its zero value it painted them, and the engine's writer is the daemon's shared
// logsink whose syslog half is procd's stderr — so `logread` filled up with
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: exchange failed …
//
// and `grep ERROR` stopped matching. Colour is now emitted only when that
// destination is a terminal (an interactive `shaterd run`); see
// shater/logsink/color.go.
func logOptions(s string) *option.LogOptions {
if logSilenced(s) {
return &option.LogOptions{Disabled: true}
}
return &option.LogOptions{Level: logLevel(s)}
return &option.LogOptions{Level: logLevel(s), DisableColor: !logColorAllowed()}
}
// logColorAllowed reports whether the engine's log lines may carry ANSI colour.
// A package var so tests pin it instead of depending on how the test binary's
// stderr happens to be wired; production always asks the real stderr, which is
// the sink's syslog half.
var logColorAllowed = func() bool { return logsink.IsTTY(os.Stderr) }
// logSilenced reports whether the operator asked for no log at all.
//
// The vocabulary lives in model.SilentLogLevels rather than here because
@@ -564,15 +583,14 @@ func networkList(tcp, udp bool) option.NetworkList {
}
}
// sortedByOrder returns rule indices sorted by (Order, original index) so the
// sortedRuleIndices returns rule indices sorted by (Order, original index) so the
// engine's first-match order matches the model's declared order, stably.
//
// The implementation is model.SortedRuleIndices: the reachability analysis that
// decides which rules can never fire has to walk the rules in EXACTLY this order
// to be right, and it lives in model (the leaf both generate and the panel API
// import). Two copies of "what order do rules run in" is precisely the drift that
// would make the warning and the panel badge disagree.
func sortedRuleIndices(rules []model.Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
return model.SortedRuleIndices(rules)
}
+73
View File
@@ -0,0 +1,73 @@
package generate
import (
"bytes"
"context"
"strings"
"testing"
"github.com/sagernet/sing-box/log"
"github.com/sagernet/sing-box/option"
"github.com/sagernet/sing-box/shater/model"
)
// TestLogOptionsDisableColorOffTTY pins the engine half of the syslog colour
// leak. The engine's log factory is built by box.New from these options; with
// DisableColor left false, every engine line reached procd's stderr — i.e.
// syslog — wrapped in aurora escapes.
func TestLogOptionsDisableColorOffTTY(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false } // stderr is procd's pipe
t.Cleanup(func() { logColorAllowed = restore })
for _, level := range []string{"", "info", "error", "warning"} {
got := logOptions(level)
if got.Disabled {
t.Fatalf("logOptions(%q) unexpectedly disabled the log", level)
}
if !got.DisableColor {
t.Errorf("logOptions(%q).DisableColor = false; syslog would get ANSI escapes", level)
}
}
// A terminal (an interactive `shaterd run`) keeps its colours.
logColorAllowed = func() bool { return true }
if logOptions("info").DisableColor {
t.Errorf("logOptions on a TTY disabled colour; the gate is supposed to be the destination, not a blanket off")
}
}
// TestGeneratedLogBlockProducesNoANSI is the end-to-end proof: take the log
// block Generate actually emits, build the engine's factory over it exactly the
// way box.New does (log.New with DefaultWriter), and check the bytes.
func TestGeneratedLogBlockProducesNoANSI(t *testing.T) {
restore := logColorAllowed
logColorAllowed = func() bool { return false }
t.Cleanup(func() { logColorAllowed = restore })
m := &model.Model{Globals: model.DefaultGlobals()}
m.Globals.LogLevel = "info"
opts, _, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("GenerateWithWarnings: %v", err)
}
if opts.Log == nil {
t.Fatal("generated options carry no log block")
}
var out bytes.Buffer
factory, err := log.New(log.Options{Options: option.LogOptions(*opts.Log), DefaultWriter: &out})
if err != nil {
t.Fatalf("log.New over the generated block: %v", err)
}
ctx := log.ContextWithNewID(context.Background())
factory.Logger().ErrorContext(ctx, "dns: exchange failed for example.com. IN AAAA: unexpected EOF")
got := out.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("engine factory wrote nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("engine log line carries an ANSI escape at byte %d: %q", i, got)
}
}
+55 -16
View File
@@ -30,6 +30,12 @@ func (b *builder) buildRoute() *option.RouteOptions {
// never aborts box.New.
b.applyProfiles()
// Report the rules that cannot fire under this config BEFORE building anything,
// so the diagnosis is about the desired state the operator wrote rather than
// about whatever survived generation. Diagnosis only — nothing is renamed,
// reordered or dropped here.
b.warnUnreachableRules()
// Kill-switch default backstop.
final := tagBlock
if strings.EqualFold(strings.TrimSpace(b.m.Globals.KillSwitch), "open") {
@@ -240,25 +246,58 @@ func (b *builder) ruleKillFallback(r model.Rule, want string) string {
// outbound — traffic left over the DEFAULT WAN while the panel (which applies the
// same "a bare Egress is a target too" rule for display) showed `egress:wan2`.
// A silent mis-route on exactly the multi-WAN setups the field exists for.
func effectiveRuleTarget(r model.Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
//
// The implementation is model.EffectiveRuleTarget — shared with the reachability
// analysis, which must resolve a rule's target identically to decide which
// default the router actually ends up using.
func effectiveRuleTarget(r model.Rule) string { return model.EffectiveRuleTarget(r) }
// isCatchAll reports whether a rule carries NO matcher of any kind (src, dst
// domain/ip/ruleset, port, proto). Such a rule is the default egress.
func (b *builder) isCatchAll(r model.Rule) bool {
return len(r.Src) == 0 &&
len(r.DstDomain) == 0 &&
len(r.DstIP) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
//
// The implementation is model.IsCatchAll — shared with the reachability analysis
// (and mirrored by the panel), because "is this rule a default?" is the single
// question both the Final assignment below and the never-fires badge turn on.
func (b *builder) isCatchAll(r model.Rule) bool { return model.IsCatchAll(r) }
// warnUnreachableRules reports every rule that CANNOT take effect under this
// config, whatever the traffic.
//
// Only one shape is certain enough to report (see model.RuleReachability): two
// condition-less rules, where the later one wins because the loop above simply
// overwrites Final. That case is silent today and looks completely healthy — the
// panel drew both as "default route · final" and the log said nothing — so a
// config with `default -> direct` at order 20 and `default -> group:auto` at
// order 100 gave no hint at all that one of the two was doing nothing.
//
// Severity is decided by CONSEQUENCE, via the wording (apply/warnings.go
// classifies generate's text): when the default that actually wins is `direct`
// while the retired rule asked for a tunnel or a block, the operator's default
// policy is not in effect and everything unmatched leaves on the plain WAN — a
// broken protection claim, i.e. critical. Any other combination (a dead `direct`
// under a live tunnel, one tunnel under another) is a dead setting, not a leak:
// a warning.
func (b *builder) warnUnreachableRules() {
for _, rr := range model.RuleReachability(b.effectiveRules) {
if !rr.Unreachable {
continue
}
dead := effectiveRuleTarget(b.effectiveRules[rr.Index])
winner := effectiveRuleTarget(b.effectiveRules[rr.ShadowedByIndex])
if winner == tagDirect && dead != tagDirect {
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so %q wins and this rule's "+
"target %q is never applied. Everything no other rule matches leaves over the plain "+
"WAN with your real IP address. Delete one of the two, or give this one a condition",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
continue
}
b.warnf("rule %q: it has no conditions, so it sets the default for ALL traffic — but "+
"rule %q (order %d) has none either and comes after it, so the default the router uses "+
"is %q and this rule's target %q is never applied. Delete one of the two, or give this "+
"one a condition so it can match something",
rr.Name, rr.ShadowedBy, rr.ShadowedByOrder, winner, dead)
}
}
// ruleMatchers builds the RawDefaultRule matchers the engine can evaluate.
+173
View File
@@ -0,0 +1,173 @@
// B1: a routing rule that can never fire, reported instead of applied silently.
//
// The field config that prompted this had TWO rules named `default`, both with
// zero conditions — order 20 -> direct and order 100 -> group:auto. buildRoute
// points route Final at a condition-less rule and moves on, so the LAST one wins
// and the other is a dead setting. Nothing anywhere said so: the log was clean and
// the panel drew both rows identically.
//
// These tests pin the diagnosis AND the behaviour it describes, because a warning
// that disagrees with what the generator actually does is worse than none.
package generate
import (
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// warnAbout returns the warnings mentioning `rule "name"`.
func warnAbout(warns []string, name string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, `rule "`+name+`"`) {
out = append(out, w)
}
}
return out
}
func warnsMatching(warns []string, substr string) []string {
var out []string
for _, w := range warns {
if strings.Contains(w, substr) {
out = append(out, w)
}
}
return out
}
const neverApplied = "is never applied"
// twoDefaultsModel reproduces the field config: two condition-less rules sharing
// the name `default`, differing only in Order and target.
func twoDefaultsModel(lowTarget, highTarget string) *model.Model {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
return &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: lowTarget},
{Name: "default", Enabled: true, Order: 100, Target: highTarget},
},
}
}
// TestTwoCatchAllRulesWarnAndLastWins is the B1 regression. The order-100 rule is
// the default the engine uses (route Final), and the order-20 one is reported as
// never applied — naming the rule that supersedes it.
func TestTwoCatchAllRulesWarnAndLastWins(t *testing.T) {
opts, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the LAST catch-all's target %q", got, "auto")
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
// It must name the superseding rule AND its order — with both rules called
// `default`, the order is the only thing that tells the two apart.
// Targets are quoted as the OPERATOR wrote them (`group:auto`), not as the
// engine tag they resolve to (`auto`) — the warning has to be readable next to
// the config, not next to the generated JSON.
for _, want := range []string{`rule "default"`, "order 100", `"group:auto"`, `"direct"`} {
if !strings.Contains(got[0], want) {
t.Fatalf("warning must mention %s, got: %s", want, got[0])
}
}
// Diagnosis only: neither rule is renamed, reordered or dropped.
if len(opts.Route.Rules) == 0 {
t.Fatal("route rules disappeared")
}
}
// TestTwoCatchAllsDirectWinsIsCritical: when the surviving default is `direct` and
// the retired one asked for a tunnel, everything unmatched leaves on the plain
// WAN. The warning must say so in the words apply/warnings.go grades critical —
// this is the case where a healthy-looking panel is a lie.
func TestTwoCatchAllsDirectWinsIsCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("group:auto", "direct"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
got := warnsMatching(warns, neverApplied)
if len(got) != 1 {
t.Fatalf("want exactly one never-applied warning, got %d: %q", len(got), warns)
}
const marker = "leaves over the plain WAN with your real IP address"
if !strings.Contains(got[0], marker) {
t.Fatalf("a retired tunnel default under a live direct default must carry %q, got: %s", marker, got[0])
}
}
// TestTwoCatchAllsTunnelWinsIsNotCritical: the field config's actual shape — the
// dead rule is `direct` and the live default is the tunnel. That is a dead
// setting, not a leak, so it must NOT carry the critical marker.
func TestTwoCatchAllsTunnelWinsIsNotCritical(t *testing.T) {
_, warns, err := GenerateWithWarnings(twoDefaultsModel("direct", "group:auto"))
if err != nil {
t.Fatalf("Generate: %v", err)
}
for _, w := range warnsMatching(warns, neverApplied) {
if strings.Contains(w, "leaves over the plain WAN with your real IP address") {
t.Fatalf("a dead direct default under a live tunnel default is not a leak: %s", w)
}
}
}
// TestSingleCatchAllDoesNotWarn: the ordinary config — specific rules plus ONE
// default — must stay silent. A badge on a working rule teaches the operator to
// ignore the badge.
func TestSingleCatchAllDoesNotWarn(t *testing.T) {
g := model.DefaultGlobals()
g.KillSwitch = "closed"
m := &model.Model{
Globals: g,
Nodes: []model.Node{{Name: "n1", Enabled: true, URI: "ss://aes-256-gcm:secret@203.0.113.1:8388#n1"}},
Groups: []model.Group{{Name: "auto", Strategy: "leastping", Nodes: []string{"n1"}}},
Rules: []model.Rule{
{Name: "ads", Enabled: true, Order: 10, DstPort: "443", Target: "block"},
// A specific rule ordered BELOW the default: still emitted ahead of Final,
// so it is not retired either.
{Name: "default", Enabled: true, Order: 20, Target: "group:auto"},
{Name: "late", Enabled: true, Order: 900, DstPort: "8080", Target: "direct"},
},
}
_, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a config with one default must not report anything never-applied, got: %q", got)
}
if got := warnAbout(warns, "late"); len(got) != 0 {
t.Fatalf("a specific rule below the default is not shadowed by it, got: %q", got)
}
}
// TestProfileDisabledCatchAllDoesNotShadow: the active profile switches the later
// default off, so the earlier one is the live default and must not be badged.
// Judging the raw config would blame the wrong rule on every profile router.
func TestProfileDisabledCatchAllDoesNotShadow(t *testing.T) {
m := twoDefaultsModel("group:auto", "direct")
m.Rules[1].Name = "fallback" // profiles address rules by name
m.Globals.ActiveProfile = "home"
m.Profiles = []model.Profile{{Name: "home", Enabled: true, DisableRules: []string{"fallback"}}}
opts, warns, err := GenerateWithWarnings(m)
if err != nil {
t.Fatalf("Generate: %v", err)
}
if got := warnsMatching(warns, neverApplied); len(got) != 0 {
t.Fatalf("a profile-disabled default shadows nothing, got: %q", got)
}
if got := opts.Route.Final; got != "auto" {
t.Fatalf("route Final = %q, want the surviving default's target %q", got, "auto")
}
}
+47
View File
@@ -0,0 +1,47 @@
package logsink
// Colour policy for everything that writes into the sink.
//
// Both log producers of the daemon (the control-plane factory built in
// cmd/shaterd, and the engine's own factory built by box.New from
// option.LogOptions) format with github.com/logrusorgru/aurora colours ON by
// default: log/format.go paints the level word and the per-connection ID unless
// DisableColors / LogOptions.DisableColor is set. Under procd the daemon's
// stderr is not a terminal, it is syslog — so every ERROR line landed in
// `logread` as
//
// ESC[31mERRORESC[0m[0026] [ESC[38;5;193m1728741629ESC[0m 70ms] dns: …
//
// Syslog is not a screen: `logread | grep ERROR` misses the coloured word
// because there are invisible bytes inside it, log collectors store the escapes
// forever, and anyone reading a captured log sees mojibake. The sink's file half
// already strips ANSI on the way out (emitLocked -> stripANSI) — the syslog half
// deliberately did not, and that is the leak.
//
// The fix is at the producer, not at the sink: colour is a property of the
// DESTINATION, so it is decided once, from whether that destination is a
// terminal, and never emitted otherwise. Stripping at the sink would keep the
// wasted formatting work and would still leak through any future path that does
// not go through the sink.
import "os"
// IsTTY reports whether w is a terminal, i.e. whether ANSI colour escapes
// written to it will be RENDERED rather than stored.
//
// The check is the portable one — a character device — so it needs no cgo, no
// termios ioctl and no new dependency (the router binary is CGO_ENABLED=0
// musl-static, and logsink also builds on the Windows/macOS dev hosts). Under
// procd, stderr is a pipe to the log daemon: not a character device, so colour
// is off, which is the case that matters. An interactive `shaterd run` from a
// shell keeps its colours.
func IsTTY(w *os.File) bool {
if w == nil {
return false
}
fi, err := w.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
}
+94
View File
@@ -0,0 +1,94 @@
package logsink
import (
"bytes"
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/sagernet/sing-box/log"
)
// TestIsTTYNonTerminal pins the only direction that matters in production: the
// destinations the daemon actually has under procd — a pipe (procd's stderr
// relay) and a regular file — are NOT terminals, so nothing may colour for them.
func TestIsTTYNonTerminal(t *testing.T) {
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
if IsTTY(w) {
t.Errorf("IsTTY(pipe) = true; procd's stderr is a pipe and must never be coloured")
}
f, err := os.Create(filepath.Join(t.TempDir(), "log"))
if err != nil {
t.Fatalf("create: %v", err)
}
t.Cleanup(func() { _ = f.Close() })
if IsTTY(f) {
t.Errorf("IsTTY(regular file) = true, want false")
}
if IsTTY(nil) {
t.Errorf("IsTTY(nil) = true, want false")
}
}
// TestSyslogHalfHasNoANSI is the regression for the escape codes that reached
// `logread`:
//
// daemon.err shaterd[27540]: …Z ESC[31mERRORESC[0m[0026] [ESC[38;5;193m…ESC[0m 70ms] dns: …
//
// It wires a factory the way the daemon does — formatter colour gated on
// IsTTY(destination), output into a Sink whose syslog half is captured — and
// asserts the captured bytes carry no ESC (0x1b). The context ID is set because
// the ID is coloured by a SEPARATE branch of log/format.go: a fix that only
// silenced the level word would still leak here.
func TestSyslogHalfHasNoANSI(t *testing.T) {
r, w, err := os.Pipe() // a non-terminal destination, exactly like procd's stderr
if err != nil {
t.Fatalf("pipe: %v", err)
}
t.Cleanup(func() { _ = r.Close(); _ = w.Close() })
var syslog bytes.Buffer
s := New(&syslog, Config{ToSyslog: true})
factory := log.NewDefaultFactory(context.Background(),
log.Formatter{BaseTime: time.Now(), DisableColors: !IsTTY(w)},
s, "", nil, false)
logger := factory.Logger()
ctx := log.ContextWithNewID(context.Background())
logger.ErrorContext(ctx, "dns: exchange failed for catalog.example.com. IN AAAA: unexpected EOF")
logger.WarnContext(ctx, "warn line")
logger.InfoContext(ctx, "info line")
factory.SetLevel(log.LevelTrace)
logger.DebugContext(ctx, "debug line")
logger.TraceContext(ctx, "trace line")
_ = s.Close()
got := syslog.String()
if !strings.Contains(got, "unexpected EOF") {
t.Fatalf("the syslog half captured nothing usable: %q", got)
}
if i := strings.IndexByte(got, 0x1b); i >= 0 {
t.Fatalf("syslog half carries an ANSI escape at byte %d: %q", i, got)
}
if !strings.Contains(got, "ERROR") {
t.Errorf("`grep ERROR` must match a plain, unbroken level word: %q", got)
}
// Teeth check: the SAME line through a colouring formatter must contain the
// escape. Without this, a future aurora that stopped colouring would make the
// assertion above pass for the wrong reason and the guard would rot silently.
coloured := log.Formatter{BaseTime: time.Now()}.
Format(ctx, log.LevelError, "", "dns: exchange failed", time.Now())
if !strings.ContainsRune(coloured, 0x1b) {
t.Fatalf("colouring formatter emitted no ANSI escape (%q) — this test can no longer detect the leak", coloured)
}
}
+195
View File
@@ -0,0 +1,195 @@
package model
// Rule reachability — "this rule can never fire", computed from the desired state
// alone.
//
// WHY THIS EXISTS. A routing rule with no conditions at all (no src, no dst
// domain/ip/ruleset, no port, no proto) is not a rule in the ordinary sense: it
// states a policy for EVERYTHING. generate does not emit it as a match-all route
// rule — it points the engine's route `Final` at that rule's target instead (see
// generate/route.go buildRoute). Two consequences follow, and neither is visible
// anywhere in the UI:
//
// 1. A condition-less rule can NEVER shadow a rule that has conditions. Every
// conditional rule is emitted ahead of `Final`, whatever its Order. So a
// "default -> direct" at order 20 does not hijack a specific rule at order 500.
// 2. Two condition-less rules DO shadow each other, and the LAST one in
// (Order, position) order wins, because the loop simply overwrites `Final`.
// Every earlier one is a dead setting that reads as a live one — the panel
// drew both with the same "default route · final" badge.
//
// A real config in the field had exactly that: two rules both named `default`,
// both with zero conditions, order 20 -> direct and order 100 -> group:auto. One
// of the two was doing nothing, and nothing said which.
//
// SCOPE, DELIBERATELY NARROW. This reports only the shadowing that is certain
// from the config: condition-less rule over condition-less rule. It does NOT try
// to decide whether one conditional rule's matcher set subsumes another's (is
// `dst_ruleset ru-inside` a superset of `dst_ip 5.0.0.0/8`? only the compiled
// rule-set knows) — a false "never fires" badge on a working rule would be worse
// than no badge at all.
//
// It lives in model, the stdlib-only leaf, because both consumers must agree:
// generate turns the verdict into an apply warning, and the panel API serves it
// to the Routing page. A second, drifting implementation in either is how the
// warning and the badge end up disagreeing about the same config.
import (
"fmt"
"sort"
"strings"
)
// IsCatchAll reports whether a rule carries NO matcher of any kind (src, dst
// domain/ip/ruleset, port, proto). Such a rule is the default egress: generate
// points route `Final` at it rather than emitting a match-all rule.
//
// generate.isCatchAll and the panel's isCatchAll() are the same predicate; this
// is the one the Go side shares.
func IsCatchAll(r Rule) bool {
return len(r.Src) == 0 &&
len(r.DstDomain) == 0 &&
len(r.DstIP) == 0 &&
len(r.DstRuleset) == 0 &&
strings.TrimSpace(r.DstPort) == "" &&
strings.TrimSpace(r.Proto) == ""
}
// EffectiveRuleTarget resolves what a rule actually routes to: Target wins, and a
// rule with NO Target but a bare Egress routes to that egress outbound. "" means
// the rule states no policy at all — generate drops such a rule entirely (it does
// NOT mean "direct"), so it never becomes the default and never shadows anything.
func EffectiveRuleTarget(r Rule) string {
if t := strings.TrimSpace(r.Target); t != "" {
return t
}
if e := strings.TrimSpace(r.Egress); e != "" {
return "egress:" + e
}
return ""
}
// SortedRuleIndices returns rule indices sorted by (Order, original index), so the
// engine's first-match order matches the model's declared order, stably. Equal
// Order keeps the configured/UCI order — the panel promises "lower Order runs
// earlier", and nothing below it may reshuffle ties.
func SortedRuleIndices(rules []Rule) []int {
idx := make([]int, len(rules))
for i := range rules {
idx[i] = i
}
sort.SliceStable(idx, func(a, b int) bool {
return rules[idx[a]].Order < rules[idx[b]].Order
})
return idx
}
// RuleReach is one rule's reachability verdict. There is exactly one per input
// rule, at the same index, so a consumer can zip the two slices.
//
// It is the routing-rule analogue of engine.ChainHealth.Used / GroupHealth.Used —
// a quiet note about the ROUTING CONFIG, never a health or liveness signal.
type RuleReach struct {
// Index is the rule's position in the model's Rules slice (NOT the sorted
// order), so the panel can match a verdict to the row it drew.
Index int `json:"index"`
// Name and Order are echoed so a consumer holding a possibly-stale config can
// verify the verdict still describes the rule it is about to badge, instead of
// accusing the wrong row.
Name string `json:"name"`
Order int `json:"order"`
// Unreachable is true when this rule can never take effect, whatever the
// traffic. false is the normal case and carries no claim that the rule ever
// actually matches something — only that nothing in the config stops it.
Unreachable bool `json:"unreachable"`
// ShadowedBy / ShadowedByOrder name the rule that supersedes this one; empty /
// zero when Unreachable is false. ShadowedByIndex is -1 when there is none.
ShadowedBy string `json:"shadowed_by,omitempty"`
ShadowedByIndex int `json:"shadowed_by_index"`
ShadowedByOrder int `json:"shadowed_by_order,omitempty"`
// Reason is the operator-facing sentence; "" when Unreachable is false.
Reason string `json:"reason,omitempty"`
}
// RuleReachability returns one verdict per rule, in the INPUT slice's order.
//
// The input must already be the EFFECTIVE rule set — profile enable/disable
// applied (see Model.EffectiveRules). A rule the active profile switched off is
// not in force, and a rule it switched ON is, whatever /etc/config/shater says;
// judging the raw slice would badge the wrong rules on a router that uses
// profiles at all.
//
// A rule takes part in the analysis only when it is Enabled, condition-less, and
// carries a target — those are exactly the rules generate lets set `Final`.
//
// SCHEDULES. A rule that only holds inside a time window never marks anything
// permanently dead: outside its window the rule below it is the default again. So
// a SCHEDULED catch-all is skipped as a shadower (but can still BE shadowed — an
// unscheduled catch-all after it wins at every hour of every day, which makes the
// schedule pure decoration and is worth saying out loud).
func RuleReachability(rules []Rule) []RuleReach {
out := make([]RuleReach, len(rules))
for i := range rules {
out[i] = RuleReach{
Index: i, Name: rules[i].Name, Order: rules[i].Order, ShadowedByIndex: -1,
}
}
// Walk the declared order BACKWARDS and remember the last unconditional
// default seen. When we reach a rule, `winner` holds the default that is in
// force below it — i.e. the one whose target the engine actually uses.
winner := -1
for _, i := range reverse(SortedRuleIndices(rules)) {
r := rules[i]
if !r.Enabled || !IsCatchAll(r) || EffectiveRuleTarget(r) == "" {
continue
}
if winner >= 0 {
w := rules[winner]
out[i].Unreachable = true
out[i].ShadowedBy = w.Name
out[i].ShadowedByIndex = winner
out[i].ShadowedByOrder = w.Order
out[i].Reason = fmt.Sprintf(
"this rule has no conditions, so it sets the default for all traffic — but rule %q "+
"(order %d) has none either and comes after it, so %q is the default the router "+
"uses and this rule's target %q is never applied",
w.Name, w.Order, EffectiveRuleTarget(w), EffectiveRuleTarget(r))
}
// Only an UNSCHEDULED default holds at every hour, so only it can retire the
// rules above it. Keep the LAST one (highest Order): that is the one whose
// target the engine ends up with.
if !r.SchedEnabled && winner < 0 {
winner = i
}
}
return out
}
// reverse returns idx walked back to front. Small enough to inline by hand, but
// naming it keeps the loop above readable as "walk the declared order backwards".
func reverse(idx []int) []int {
out := make([]int, len(idx))
for i, v := range idx {
out[len(idx)-1-i] = v
}
return out
}
// EffectiveRules returns the rule set that is actually in force: a COPY of m.Rules
// with the active profile's enable/disable overrides applied, plus whatever
// warnings that resolution produced. m.Rules is never mutated.
//
// It is the shared front door for every consumer that must reason about "which
// rules are in force right now" without building an engine config — the panel's
// reachability endpoint today. generate keeps its own call site because it also
// needs the resolved *Profile itself (endpoint-resolver override); both go through
// ResolveActiveProfile + ApplyProfileRuleOverrides, so they cannot disagree.
func (m *Model) EffectiveRules() ([]Rule, []Warning) {
if m == nil {
return nil, nil
}
prof, warns := ResolveActiveProfile(m)
rules, owarns := ApplyProfileRuleOverrides(m.Rules, prof)
return rules, append(warns, owarns...)
}
+313
View File
@@ -0,0 +1,313 @@
package model
import "testing"
// catchAll builds a condition-less rule — the shape that becomes the engine's
// route Final.
func catchAll(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target}
}
// specific builds a rule with one matcher, so it is emitted as a real route rule
// ahead of Final and can never be retired by a default.
func specific(name string, order int, target string) Rule {
return Rule{Name: name, Enabled: true, Order: order, Target: target, DstPort: "443"}
}
// verdicts indexes a reachability run by rule index, asserting the contract that
// there is exactly one verdict per rule, at the same index.
func verdicts(t *testing.T, rules []Rule) []RuleReach {
t.Helper()
out := RuleReachability(rules)
if len(out) != len(rules) {
t.Fatalf("want %d verdicts, got %d", len(rules), len(out))
}
for i := range out {
if out[i].Index != i || out[i].Name != rules[i].Name {
t.Fatalf("verdict %d is about %q (index %d), want %q", i, out[i].Name, out[i].Index, rules[i].Name)
}
}
return out
}
func wantReachable(t *testing.T, v RuleReach) {
t.Helper()
if v.Unreachable {
t.Fatalf("rule %q (order %d) must be reachable, got shadowed by %q: %s",
v.Name, v.Order, v.ShadowedBy, v.Reason)
}
if v.ShadowedBy != "" || v.ShadowedByIndex != -1 || v.Reason != "" {
t.Fatalf("rule %q is reachable but carries shadow details: by=%q idx=%d reason=%q",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.Reason)
}
}
func wantShadowed(t *testing.T, v RuleReach, by string, byIndex, byOrder int) {
t.Helper()
if !v.Unreachable {
t.Fatalf("rule %q (order %d) must be unreachable, shadowed by %q", v.Name, v.Order, by)
}
if v.ShadowedBy != by || v.ShadowedByIndex != byIndex || v.ShadowedByOrder != byOrder {
t.Fatalf("rule %q: shadowed by %q(idx %d, order %d), want %q(idx %d, order %d)",
v.Name, v.ShadowedBy, v.ShadowedByIndex, v.ShadowedByOrder, by, byIndex, byOrder)
}
if v.Reason == "" {
t.Fatalf("rule %q is unreachable but carries no reason", v.Name)
}
}
// TestReachabilitySingleCatchAll: the ordinary config — some specific rules and
// ONE default. Nothing is retired, including the default itself.
func TestReachabilitySingleCatchAll(t *testing.T) {
rules := []Rule{
specific("block-ads", 10, "block"),
specific("ru-bypass", 20, "direct"),
catchAll("default-tunnel", 900, "group:auto"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityTwoCatchAllsFieldConfig is the config that prompted this
// analysis, reproduced exactly: two rules BOTH named `default`, both with zero
// conditions, order 20 -> direct and order 100 -> group:auto.
//
// The later one is the default the engine uses (buildRoute overwrites Final as it
// walks ascending Order), so it is the order-20 `direct` rule that is dead — the
// opposite of what "first match wins" would suggest, which is exactly why this
// needed saying out loud.
func TestReachabilityTwoCatchAllsFieldConfig(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
catchAll("default", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "default", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityCatchAllBelowSpecificRules: a default sorted BELOW specific
// rules retires none of them. A condition-less rule becomes route Final, which is
// evaluated after every emitted rule whatever its Order — so even a specific rule
// ordered after the default still matches first.
func TestReachabilityCatchAllBelowSpecificRules(t *testing.T) {
rules := []Rule{
catchAll("default", 20, "direct"),
specific("work-vpn", 100, "group:work"),
specific("ads", 200, "block"),
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityThreeCatchAlls: only the last default survives; every earlier
// one is blamed on that same last one, not on its immediate successor — that is
// the rule whose target the router actually ends up using.
func TestReachabilityThreeCatchAlls(t *testing.T) {
rules := []Rule{
catchAll("a", 10, "direct"),
catchAll("b", 20, "block"),
catchAll("c", 30, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "c", 2, 30)
wantShadowed(t, v[1], "c", 2, 30)
wantReachable(t, v[2])
}
// TestReachabilityDisabledCatchAllIsInert: a disabled default neither dies nor
// kills. It installs nothing, so badging it "never fires" would be noise, and it
// must not be credited with retiring the live default above it.
func TestReachabilityDisabledCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "parked", Enabled: false, Order: 100, Target: "direct"},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityTargetlessCatchAllIsInert: a rule with neither Target nor
// Egress states no policy, so generate drops it (it does NOT mean "direct"). It
// never becomes Final, so it cannot retire the default above it — and is not
// itself reported here, because "no target set" is its own, more useful warning.
func TestReachabilityTargetlessCatchAllIsInert(t *testing.T) {
rules := []Rule{
catchAll("live", 20, "group:auto"),
{Name: "empty", Enabled: true, Order: 100},
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantReachable(t, v[1])
}
// TestReachabilityBareEgressCountsAsTarget: a rule carrying only `option egress
// wan2` routes to that egress, so it IS a default and does retire the one above.
func TestReachabilityBareEgressCountsAsTarget(t *testing.T) {
rules := []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "wan2", Enabled: true, Order: 100, Egress: "wan2"},
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "wan2", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityScheduledCatchAllNeverRetires: a default that only holds inside
// a time window leaves the one above it live for the rest of the day, so it
// retires nothing.
func TestReachabilityScheduledCatchAllNeverRetires(t *testing.T) {
rules := []Rule{
catchAll("all-day", 20, "group:auto"),
{Name: "nightly", Enabled: true, Order: 100, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
}
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityScheduledCatchAllCanBeRetired: the converse — an unscheduled
// default AFTER a scheduled one wins at every hour of every day, so the schedule
// is pure decoration and the scheduled rule is dead.
func TestReachabilityScheduledCatchAllCanBeRetired(t *testing.T) {
rules := []Rule{
{Name: "nightly", Enabled: true, Order: 20, Target: "direct",
SchedEnabled: true, SchedStart: "01:00", SchedEnd: "05:00"},
catchAll("all-day", 100, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "all-day", 1, 100)
wantReachable(t, v[1])
}
// TestReachabilityEqualOrderKeepsConfiguredOrder: ties break by position in the
// slice (SortedRuleIndices is stable), so with equal Order the SECOND section in
// /etc/config/shater is the one that wins — same as generate.
func TestReachabilityEqualOrderKeepsConfiguredOrder(t *testing.T) {
rules := []Rule{
catchAll("first", 50, "direct"),
catchAll("second", 50, "group:auto"),
}
v := verdicts(t, rules)
wantShadowed(t, v[0], "second", 1, 50)
wantReachable(t, v[1])
}
// TestReachabilityUnsortedInputIsJudgedByOrder: the model slice arrives in UCI
// order, which need not be Order order. The verdict must follow Order, and the
// returned slice must stay aligned with the INPUT indices.
func TestReachabilityUnsortedInputIsJudgedByOrder(t *testing.T) {
rules := []Rule{
catchAll("late", 100, "group:auto"),
catchAll("early", 20, "direct"),
}
v := verdicts(t, rules)
wantReachable(t, v[0])
wantShadowed(t, v[1], "late", 0, 100)
}
// TestReachabilityMatcherKindsAreNotCatchAll: every matcher field on its own is
// enough to make a rule conditional, so none of these is retired by the default
// below. A field this misses would silently badge a working rule "never fires".
func TestReachabilityMatcherKindsAreNotCatchAll(t *testing.T) {
conditional := []Rule{
{Name: "by-src", Enabled: true, Order: 10, Target: "direct", Src: []string{"192.168.1.0/24"}},
{Name: "by-domain", Enabled: true, Order: 11, Target: "direct", DstDomain: []string{"example.com"}},
{Name: "by-ip", Enabled: true, Order: 12, Target: "direct", DstIP: []string{"1.1.1.1/32"}},
{Name: "by-ruleset", Enabled: true, Order: 13, Target: "direct", DstRuleset: []string{"ads"}},
{Name: "by-port", Enabled: true, Order: 14, Target: "direct", DstPort: "443"},
{Name: "by-proto", Enabled: true, Order: 15, Target: "direct", Proto: "quic"},
}
for _, r := range conditional {
if IsCatchAll(r) {
t.Fatalf("rule %q carries a matcher but reads as a catch-all", r.Name)
}
}
rules := append(append([]Rule(nil), conditional...), catchAll("default", 900, "group:auto"))
for _, v := range verdicts(t, rules) {
wantReachable(t, v)
}
}
// TestReachabilityWhitespaceOnlyMatcherIsCatchAll: a port/proto of spaces is not
// a matcher. generate trims before deciding, so this analysis must too — otherwise
// a rule the engine treats as the default reads as conditional here and its
// shadowing goes unreported.
func TestReachabilityWhitespaceOnlyMatcherIsCatchAll(t *testing.T) {
blank := Rule{Name: "blank", Enabled: true, Order: 20, Target: "direct", DstPort: " ", Proto: "\t"}
if !IsCatchAll(blank) {
t.Fatal("a rule whose only matchers are whitespace must read as a catch-all")
}
v := verdicts(t, []Rule{blank, catchAll("default", 100, "group:auto")})
wantShadowed(t, v[0], "default", 1, 100)
}
// TestReachabilityEmptyAndNil: no rules, no verdicts, no panic.
func TestReachabilityEmptyAndNil(t *testing.T) {
if got := RuleReachability(nil); len(got) != 0 {
t.Fatalf("nil rules: want no verdicts, got %d", len(got))
}
if got := RuleReachability([]Rule{}); len(got) != 0 {
t.Fatalf("empty rules: want no verdicts, got %d", len(got))
}
}
// TestEffectiveRulesAppliesActiveProfile: the analysis must judge the rules the
// active profile leaves in force. Here the profile DISABLES the later default, so
// the earlier one is live and nothing is retired — judging the raw slice would
// have badged the wrong rule.
func TestEffectiveRulesAppliesActiveProfile(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "home"},
Profiles: []Profile{
{Name: "home", Enabled: true, DisableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
catchAll("fallback", 100, "direct"),
},
}
eff, _ := m.EffectiveRules()
if eff[1].Enabled {
t.Fatal("active profile must have disabled the fallback rule")
}
if !m.Rules[1].Enabled {
t.Fatal("EffectiveRules must not mutate the model's own rules")
}
for _, v := range verdicts(t, eff) {
wantReachable(t, v)
}
// Same model, profile off: the later default is back in force and retires the
// tunnel default above it.
m.Globals.ActiveProfile = ""
m.Profiles[0].Enabled = false
eff, _ = m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
// TestEffectiveRulesProfileCanReviveADefault: a profile that force-ENABLES a
// disabled default makes it in force, and it then retires the default above it.
// The raw config would show it disabled and report nothing.
func TestEffectiveRulesProfileCanReviveADefault(t *testing.T) {
m := &Model{
Globals: Globals{ActiveProfile: "travel"},
Profiles: []Profile{
{Name: "travel", Enabled: true, EnableRules: []string{"fallback"}},
},
Rules: []Rule{
catchAll("tunnel", 20, "group:auto"),
{Name: "fallback", Enabled: false, Order: 100, Target: "direct"},
},
}
eff, _ := m.EffectiveRules()
v := verdicts(t, eff)
wantShadowed(t, v[0], "fallback", 1, 100)
wantReachable(t, v[1])
}
+41 -5
View File
@@ -31,6 +31,12 @@ func ApplyNft(ruleset string) error {
if err := runNftStdin(ruleset, "-f", "-"); err != nil {
return fmt.Errorf("nft -f (load) failed: %w", err)
}
// The divert plane just changed underneath every flow that is currently
// tracked. For UDP :53 that is not self-correcting — see conntrack.go — so
// drop those entries and let the clients re-derive their path through the
// ruleset that is now loaded. Best-effort: a failed flush leaves exactly the
// behaviour we had before and must never fail an otherwise-good load.
_, _ = FlushDNSConntrack()
return nil
}
@@ -84,6 +90,11 @@ func TeardownNft() error {
if out, err := execCommand("nft", "delete", "table", "inet", "shater").CombinedOutput(); err != nil {
return fmt.Errorf("nft delete table inet shater: %v\n%s", err, out)
}
// Same reasoning as ApplyNft, mirrored: every DNS flow that was being
// delivered through the tproxy socket has just lost the rule that put it
// there. Without this, those entries survive into whatever plane comes next
// (including "no plane at all") and keep pointing at a socket that is gone.
_, _ = FlushDNSConntrack()
return nil
}
@@ -366,25 +377,50 @@ func isPointToPoint(dev string) bool {
return strings.Contains(string(out), "POINTOPOINT")
}
// RoutingPresent reports whether our fwmark ip rule is currently installed, for
// RoutingPresent reports whether our policy routing is currently installed, for
// EVERY family the model asks for. With Globals.IPv6 on, ApplyRouting installs
// both a -4 and a -6 rule, so checking only -4 was a half-truth: `ip -6 rule` is
// both a -4 and a -6 half, so checking only -4 was a half-truth: `ip -6 rule` is
// flushed independently (a `network reload` / `ifup` can drop one family and not
// the other), and the caller's idempotent fast-path would then conclude the plane
// was intact and never restore the missing v6 rule — leaving v6 clients diverted
// by nft but with nowhere to be delivered locally.
//
// With Globals.IPv6 off no v6 rule is installed BY DESIGN, so its absence must
// not be read as a missing plane; only the -4 rule is required then.
// not be read as a missing plane; only the -4 half is required then.
//
// # Both halves are checked, not just the rule
//
// ApplyRouting installs TWO things per family — the `fwmark -> table` rule AND
// the `local default dev lo` route inside that table — and they are removed by
// two INDEPENDENT commands (`ip rule del`, `ip route flush table`). Checking only
// the rule made the second half invisible: a table that had been flushed while
// its rule survived (a teardown interrupted part-way, an `ip route flush` from
// any other actor) read as "plane intact", so applyLocked's fast-path skipped
// ApplyRouting forever and NOTHING re-created the route. The visible result is a
// box that reports plane=full / engine_running=true while diverted packets are
// marked, find an empty table, fall through to the main table and are handed to
// the fail-closed forward drop — a permanent, healthy-looking outage that a
// reconcile cannot repair, because a reconcile is exactly what consults this
// function. A presence check must cover everything its Apply counterpart
// installs, or the idempotent fast-path becomes a trap.
func RoutingPresent(g model.Globals) bool {
want := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
wantRule := fmt.Sprintf("fwmark 0x%x", effFwmark(g))
table := fmt.Sprintf("%d", effTable(g))
fams := []string{"-4"}
if g.IPv6 {
fams = append(fams, "-6")
}
for _, fam := range fams {
out, err := execCommand("ip", fam, "rule", "show").Output()
if err != nil || !strings.Contains(string(out), want) {
if err != nil || !strings.Contains(string(out), wantRule) {
return false
}
// `ip -4 route show table N` prints "local default dev lo scope host";
// the v6 form is "local default dev lo metric 1024 pref medium". Matching
// the route TYPE + destination covers both without pinning the trailing
// attributes, which differ by family and iproute2 version.
rout, rerr := execCommand("ip", fam, "route", "show", "table", table).Output()
if rerr != nil || !strings.Contains(string(rout), "local default") {
return false
}
}
+78
View File
@@ -0,0 +1,78 @@
package netplane
// Conntrack maintenance for plane transitions.
//
// # Why the data plane has to touch conntrack at all
//
// Everything else in this package is *stateless* from the kernel's point of view:
// an nft ruleset, a couple of `ip rule`/`ip route` entries and a handful of
// sysctls. Rebuilding them is atomic and idempotent, so a rebuilt plane is
// indistinguishable from a freshly-installed one — EXCEPT for one thing the
// rebuild cannot reach: the connection-tracking entries that were established
// while the previous plane (or a half-removed one) was in force.
//
// That matters here because our divert is a TPROXY divert. A `tproxy` statement
// hands the packet to a LOCAL TRANSPARENT SOCKET; when that socket is gone the
// statement evaluates to NFT_BREAK and the packet takes a completely different
// path through the ruleset. A restart necessarily walks through such a window:
// the outgoing daemon closes its engine BEFORE it removes the table (Teardown
// order — deliberately, because the reverse order would open a plaintext leak),
// and the incoming daemon starts its engine BEFORE it loads the new table. Any
// flow that crosses one of those windows keeps a conntrack entry that was formed
// against a plane that no longer exists, and — for UDP, which has no handshake to
// resynchronise on — every retry merely refreshes that entry instead of
// re-deriving the path.
//
// # What this is NOT
//
// This flush is HYGIENE, not the cure for the "DNS to the router's own LAN
// address never comes back after a restart" report (B3). That turned out to be a
// socket-level collision: the engine's TPROXY UDP write-back sockets were bound
// UNCONNECTED to the original destination — the router's own LAN address :53 —
// and so joined the kernel's demultiplex set next to dnsmasq's socket, silently
// swallowing the host's own queries. The fix for that lives in the engine
// (protocol/redirect/tproxy.go, lx:tproxy_writeback_connect); the stale
// `[UNREPLIED]` conntrack entry seen on the live box was a CONSEQUENCE of the
// unanswered query, not its cause.
//
// It is kept because it is independently correct and costs one netlink
// round-trip per plane change: a DNS flow that was mid-flight across a plane
// rebuild has a conntrack entry describing a delivery path that no longer
// exists, and dropping it makes the first query after a restart re-derive its
// path immediately instead of waiting out a retry.
//
// # Scope: :53/UDP only, deliberately
//
// A blanket `conntrack -F` would also delete the entries behind the operator's
// SSH session, the LuCI session and the admin panel — fw4's input chain accepts
// them via `ct state established,related`, so dropping their conntrack entries
// drops the sessions. Locking the admin out of the box while "fixing" DNS is not
// a trade we get to make. UDP/:53 is the narrowest cut that covers the observed
// failure class: DNS is retried by every client within a second, so deleting its
// entries costs nothing and is invisible.
//
// Best-effort by contract: a kernel without conntrack, a netlink permission
// error or a non-Linux build all report zero deletions and no error path that can
// fail an apply. Losing the flush degrades to the old behaviour; it must never
// take a working plane down.
// dnsPort is the only port whose conntrack entries we touch. See the package
// comment for why this is deliberately not "everything".
const dnsPort uint16 = 53
// flushUDPPortConntrack is the platform seam. The default is a no-op so the
// package builds (and `go vet`s) on non-Linux dev hosts; conntrack_linux.go's
// init() replaces it with the real ctnetlink delete on the router target. Tests
// substitute a recorder.
var flushUDPPortConntrack = func(port uint16) (int, error) { return 0, nil }
// FlushDNSConntrack deletes every UDP connection-tracking entry whose ORIGINAL
// destination port is 53, for both address families, and returns how many were
// removed.
//
// Called on every plane transition that can change where a DNS packet is
// delivered: after a ruleset is loaded (ApplyNft — the full plane, the holding
// plane and a rollback all go through it) and after the table is removed
// (TeardownNft). Idempotent and cheap: on an idle box the DNS entry count is a
// handful, and on a busy one it is bounded by the number of clients.
func FlushDNSConntrack() (int, error) { return flushUDPPortConntrack(dnsPort) }
+62
View File
@@ -0,0 +1,62 @@
//go:build linux
package netplane
// The real ctnetlink implementation of the conntrack seam declared in
// conntrack.go. It lives behind a build tag for the same reason apply's flock
// does: the netlink conntrack API only exists on Linux, and the control plane
// must still build and test on a developer's Windows/macOS host.
//
// Deliberately netlink and NOT the `conntrack` CLI: conntrack-tools is not a
// dependency of shater-core (and pulling it in for one call would add ~100 KiB
// of userland to a flash-constrained router), so a shell-out would silently
// no-op on every real box — the worst possible outcome for a fix whose entire
// job is to remove stale state.
import (
"github.com/sagernet/netlink"
"golang.org/x/sys/unix"
)
func init() { flushUDPPortConntrack = ctnetlinkFlushUDPPort }
// ctnetlinkFlushUDPPort deletes the UDP conntrack entries whose ORIGINAL
// destination port is `port`, in both families, and returns the total deleted.
//
// Errors are aggregated rather than short-circuited: v4 and v6 are independent
// tables and a failure on one must not hide a successful cleanup of the other.
// The first error is returned for logging; the count is still accurate for the
// families that succeeded.
func ctnetlinkFlushUDPPort(port uint16) (int, error) {
var (
total int
firstErr error
)
for _, family := range []netlink.InetFamily{
netlink.InetFamily(unix.AF_INET),
netlink.InetFamily(unix.AF_INET6),
} {
filter := &netlink.ConntrackFilter{}
// Protocol MUST be set before the port: AddPort refuses to add a port
// filter while the layer-4 protocol is unknown (a port means nothing
// without one), so the order here is load-bearing.
if err := filter.AddProtocol(unix.IPPROTO_UDP); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
if err := filter.AddPort(netlink.ConntrackOrigDstPort, port); err != nil {
if firstErr == nil {
firstErr = err
}
continue
}
n, err := netlink.ConntrackDeleteFilter(netlink.ConntrackTable, family, filter)
total += int(n)
if err != nil && firstErr == nil {
firstErr = err
}
}
return total, firstErr
}
+7
View File
@@ -449,6 +449,13 @@ func TestRoutingPresentBothFamilies(t *testing.T) {
out = c.v6
}
}
// RoutingPresent also verifies the `local default dev lo` route
// that ApplyRouting installs beside the rule (see
// TestRoutingPresentRequiresLocalDefaultRoute). This case set is
// about the RULE half, so the route half is always healthy here.
if name == "ip" && len(arg) >= 3 && arg[1] == "route" {
out = "local default dev lo scope host"
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
+210
View File
@@ -0,0 +1,210 @@
package netplane
// B3 regressions: what a restart leaves behind.
//
// The bug these pin: `/etc/init.d/shater restart` (stop immediately followed by
// start) left DNS to the router's own LAN address permanently dead, while
// `stop` + pause + `start` was fine and the status kept reporting plane=full /
// engine_running=true. Two independent defects fed it, and both live here:
//
// 1. A plane transition (load or teardown of the tproxy divert) left the
// CONNTRACK entries of flows that had crossed the transition pointing at a
// plane that no longer exists. Nothing in the tree touched conntrack, so a
// UDP flow — which has no handshake to resynchronise on and whose entry is
// refreshed by every retry — stayed wedged indefinitely.
// 2. RoutingPresent() reported "plane intact" from the ip RULE alone, ignoring
// the `local default dev lo` ROUTE that ApplyRouting installs alongside it.
// A teardown interrupted between the two (procd SIGKILL at term_timeout) or
// any other `ip route flush` therefore became invisible: applyLocked's
// idempotent fast-path skipped ApplyRouting forever and no reconcile could
// repair it.
import (
"os"
"os/exec"
"strings"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// recordFlush swaps the conntrack seam for a counter and restores it after the
// test. Returns a pointer to the number of calls and the ports asked for.
func recordFlush(t *testing.T) *[]uint16 {
t.Helper()
orig := flushUDPPortConntrack
t.Cleanup(func() { flushUDPPortConntrack = orig })
var seen []uint16
flushUDPPortConntrack = func(port uint16) (int, error) {
seen = append(seen, port)
return 0, nil
}
return &seen
}
// TestApplyNftFlushesDNSConntrack: loading a ruleset must drop the DNS conntrack
// entries formed against the previous plane. Without this, a flow that crossed
// the restart window keeps being delivered by a rule set that is gone.
func TestApplyNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err != nil {
t.Fatalf("ApplyNft: %v", err)
}
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("ApplyNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
// Ordering matters: the flush is only meaningful once the NEW ruleset is in
// the kernel, otherwise the very next packet re-creates the entry against the
// old plane. The load is the last nft invocation before it.
if len(rec) != 2 || !strings.Contains(rec[0], "-c") || strings.Contains(rec[1], "-c") {
t.Fatalf("expected validate-then-load, got %v", rec)
}
}
// TestApplyNftDoesNotFlushOnFailure: a ruleset that does not load leaves the
// PREVIOUS plane in charge (nft -f is one netlink transaction). Flushing then
// would tear down live flows for nothing.
func TestApplyNftDoesNotFlushOnFailure(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := ApplyNft("table inet shater {}\n"); err == nil {
t.Fatalf("ApplyNft must report the nft failure")
}
if len(*seen) != 0 {
t.Fatalf("a failed load must not flush conntrack, got %v", *seen)
}
}
// TestTeardownNftFlushesDNSConntrack is the mirror: removing the table strands
// every DNS flow that was being delivered through the tproxy socket, so those
// entries must go with it.
func TestTeardownNftFlushesDNSConntrack(t *testing.T) {
seen := recordFlush(t)
var rec []string
orig := execCommand
execCommand = fakeExec(&rec)
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft: %v", err)
}
// fakeExec makes every command succeed, so TableExists() is true and the
// delete runs.
if len(*seen) != 1 || (*seen)[0] != dnsPort {
t.Fatalf("TeardownNft must flush UDP :%d conntrack exactly once, got %v", dnsPort, *seen)
}
}
// TestTeardownNftAbsentTableDoesNotFlush: nothing was diverting, so nothing is
// stranded. Keeps the flush out of the hot path of an idle box.
func TestTeardownNftAbsentTableDoesNotFlush(t *testing.T) {
seen := recordFlush(t)
orig := execCommand
execCommand = failingExec()
defer func() { execCommand = orig }()
if err := TeardownNft(); err != nil {
t.Fatalf("TeardownNft with no table must be a no-op, got %v", err)
}
if len(*seen) != 0 {
t.Fatalf("absent table must not flush conntrack, got %v", *seen)
}
}
// failingExec returns an execCommand replacement whose every command exits 1.
func failingExec() func(string, ...string) *exec.Cmd {
return func(name string, arg ...string) *exec.Cmd {
cs := append([]string{"-test.run=TestRestartHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_FAIL=1")
return cmd
}
}
// TestRestartHelperProcess is the exec helper for failingExec.
func TestRestartHelperProcess(t *testing.T) {
if os.Getenv("GO_WANT_HELPER_PROCESS") != "1" {
return
}
if os.Getenv("GO_HELPER_FAIL") == "1" {
os.Exit(1)
}
os.Exit(0)
}
// TestRoutingPresentRequiresLocalDefaultRoute is the second half of B3.
//
// ApplyRouting installs TWO things per family and they are removed by two
// independent commands. A presence check that only looks at the rule declares a
// half-removed plane healthy — and because applyLocked consults exactly this
// function to decide whether to re-run ApplyRouting, the missing route is then
// never restored: marked packets find an empty table, fall through to main and
// are eaten by the fail-closed forward drop, permanently, with the status still
// saying plane=full.
//
// On the pre-fix implementation the first two cases below return true.
func TestRoutingPresentRequiresLocalDefaultRoute(t *testing.T) {
const (
rule = "32765:\tfrom all fwmark 0x2000 lookup shater"
route = "local default dev lo scope host"
)
cases := []struct {
name string
ipv6 bool
v4rule, v4rte string
v6rule, v6rte string
want bool
}{
// THE REGRESSION: rule survived, table was flushed.
{"v4 rule present, route flushed", false, rule, "", "", "", false},
{"ipv6 on, v6 route flushed", true, rule, route, rule, "", false},
// Sanity: a complete plane is still reported as present.
{"v4 complete", false, rule, route, "", "", true},
{"ipv6 on, both complete", true, rule, route, rule, route, true},
// The pre-existing rule-level contract must not regress.
{"v4 rule missing", false, "", route, "", "", false},
{"ipv6 on, v6 rule missing", true, rule, route, "", route, false},
}
orig := execCommand
defer func() { execCommand = orig }()
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
execCommand = func(name string, arg ...string) *exec.Cmd {
out := ""
if name == "ip" && len(arg) >= 2 {
v6 := arg[0] == "-6"
switch arg[1] {
case "rule":
out = c.v4rule
if v6 {
out = c.v6rule
}
case "route":
out = c.v4rte
if v6 {
out = c.v6rte
}
}
}
cs := append([]string{"-test.run=TestSysctlRevertHelperProcess", "--", name}, arg...)
cmd := exec.Command(os.Args[0], cs...)
cmd.Env = append(os.Environ(), "GO_WANT_HELPER_PROCESS=1", "GO_HELPER_STDOUT="+out)
return cmd
}
g := model.Globals{FwmarkBase: 0x2000, TableBase: 0x2000, IPv6: c.ipv6}
if got := RoutingPresent(g); got != c.want {
t.Errorf("RoutingPresent = %v, want %v", got, c.want)
}
})
}
}
+48
View File
@@ -128,6 +128,54 @@ func (s *Server) handleConfigGet(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, m)
}
// rulesReachabilityResponse is the GET /api/rules/reachability body. `rules` is
// ALWAYS an array, never null, with one entry per rule in the SAME order as
// GET /api/config's Rules — so a client can zip the two by `index` (and check the
// echoed `name`/`order` before trusting a verdict it fetched around an edit).
type rulesReachabilityResponse struct {
Rules []model.RuleReach `json:"rules"`
}
// reachConfigRead is handleRulesReachability's test seam (same pattern as
// writeConfig / logConfigRead): production binds the real UCI read, tests
// substitute a canned model so the handler can be exercised without a `uci`
// binary on the host.
var reachConfigRead = model.ReadUCI
// handleRulesReachability → GET /api/rules/reachability: which routing rules can
// never take effect, and what supersedes each of them.
//
// It is the routing-rule analogue of the per-chain `used` flag on
// GET /api/groups/health — a quiet note about the ROUTING CONFIG that the panel
// renders as a badge, never a health or liveness signal. Separate from
// /api/config on purpose: /api/config is the desired state the panel PUTs back
// verbatim, and a derived verdict has no business travelling round-trip through
// it.
//
// The verdict is computed over the EFFECTIVE rules (active WAN profile's
// enable/disable applied via model.EffectiveRules), because a rule the active
// profile switched off is not in force and must not be blamed for retiring
// anything. Profile-resolution warnings are dropped here: this endpoint answers
// one question, and the same warnings already reach the operator through
// `shaterd status` on every apply.
func (s *Server) handleRulesReachability(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
writeError(w, http.StatusMethodNotAllowed, "method not allowed")
return
}
m, err := reachConfigRead()
if err != nil {
writeError(w, http.StatusInternalServerError, "read config: "+err.Error())
return
}
rules, _ := m.EffectiveRules()
out := model.RuleReachability(rules)
if out == nil {
out = []model.RuleReach{}
}
writeJSON(w, http.StatusOK, rulesReachabilityResponse{Rules: out})
}
// handleConfigPut → PUT /api/config: decode a Model, lightly validate it, and
// persist it. The write is SPLIT on the server: the FromSub nodes carried in the
// model are reconciled into the per-subscription JSON cache files
+125
View File
@@ -0,0 +1,125 @@
package panel
import (
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"github.com/sagernet/sing-box/shater/model"
)
// getReach GETs /api/rules/reachability (authenticated) and returns the status
// code plus the decoded reply.
func getReach(t *testing.T, srv *httptest.Server, cookie *http.Cookie) (int, rulesReachabilityResponse) {
t.Helper()
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET /api/rules/reachability: %v", err)
}
defer resp.Body.Close()
var out rulesReachabilityResponse
_ = json.NewDecoder(resp.Body).Decode(&out)
return resp.StatusCode, out
}
// TestRulesReachabilityEndpoint (B1): the endpoint reports the field config's two
// condition-less `default` rules, one verdict per rule, aligned with the order
// GET /api/config returns them in — that alignment is the only thing that lets the
// panel badge the right row when both rules share a name.
func TestRulesReachabilityEndpoint(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) {
return &model.Model{Rules: []model.Rule{
{Name: "default", Enabled: true, Order: 20, Target: "direct"},
{Name: "default", Enabled: true, Order: 100, Target: "group:auto"},
}}, nil
}
code, out := getReach(t, srv, cookie)
if code != http.StatusOK {
t.Fatalf("got %d, want 200", code)
}
if len(out.Rules) != 2 {
t.Fatalf("want one verdict per rule (2), got %d: %+v", len(out.Rules), out.Rules)
}
if !out.Rules[0].Unreachable {
t.Fatalf("the order-20 default is superseded by the order-100 one: %+v", out.Rules[0])
}
if out.Rules[0].Index != 0 || out.Rules[0].ShadowedByIndex != 1 || out.Rules[0].ShadowedByOrder != 100 {
t.Fatalf("verdict must point at the superseding rule by index and order: %+v", out.Rules[0])
}
if out.Rules[0].Reason == "" {
t.Fatal("an unreachable verdict must carry a reason the panel can show")
}
if out.Rules[1].Unreachable {
t.Fatalf("the last default is the one in force: %+v", out.Rules[1])
}
}
// TestRulesReachabilityEmptyIsArray: `rules` is ALWAYS an array. Go marshals a nil
// slice as null, and a client that does `for (const r of body.rules)` breaks on it.
func TestRulesReachabilityEmptyIsArray(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
orig := reachConfigRead
defer func() { reachConfigRead = orig }()
reachConfigRead = func() (*model.Model, error) { return &model.Model{}, nil }
req, _ := http.NewRequest(http.MethodGet, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET: %v", err)
}
defer resp.Body.Close()
var raw struct {
Rules *[]model.RuleReach `json:"rules"`
}
if err := json.NewDecoder(resp.Body).Decode(&raw); err != nil {
t.Fatalf("decode: %v", err)
}
if raw.Rules == nil {
t.Fatal("rules must marshal as [], never null")
}
}
// TestRulesReachabilityMethodAndAuth: it is a read endpoint behind the session
// cookie, like every other /api route except /api/session.
func TestRulesReachabilityMethodAndAuth(t *testing.T) {
s := newTestServer(t)
srv := httptest.NewServer(s.Handler())
defer srv.Close()
cookie := login(t, srv, s)
req, _ := http.NewRequest(http.MethodPost, srv.URL+"/api/rules/reachability", nil)
req.AddCookie(cookie)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("POST: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusMethodNotAllowed {
t.Fatalf("POST got %d, want 405", resp.StatusCode)
}
resp, err = http.Get(srv.URL + "/api/rules/reachability")
if err != nil {
t.Fatalf("unauthenticated GET: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusUnauthorized {
t.Fatalf("unauthenticated GET got %d, want 401", resp.StatusCode)
}
}
+1
View File
@@ -178,6 +178,7 @@ func (s *Server) buildRouter() http.Handler {
mux.Handle("/api/devices", s.requireSession(http.HandlerFunc(s.handleDevices)))
mux.Handle("/api/interfaces", s.requireSession(http.HandlerFunc(s.handleInterfaces)))
mux.Handle("/api/import-wg", s.requireSession(http.HandlerFunc(s.handleImportWG)))
mux.Handle("/api/rules/reachability", s.requireSession(http.HandlerFunc(s.handleRulesReachability)))
mux.Handle("/api/ruleset/status", s.requireSession(http.HandlerFunc(s.handleRuleSetStatus)))
mux.Handle("/api/ruleset/update", s.requireSession(http.HandlerFunc(s.handleRuleSetUpdate)))
mux.Handle("/api/ruleset/check", s.requireSession(http.HandlerFunc(s.handleRuleSetCheck)))
Binary file not shown.
+7 -3
View File
@@ -23,7 +23,7 @@ package v2rayxhttp
import (
"context"
"math/rand"
"crypto/rand"
"net"
"net/http"
"net/url"
@@ -255,8 +255,12 @@ func (c *Client) newRequest(ctx context.Context, method, sessionID, seqStr strin
// an opaque grouping key; the dashed format keeps it interchangeable with Xray.
func newSessionID() string {
var b [16]byte
for i := range b {
b[i] = byte(rand.Intn(256))
// crypto/rand.Read fills b fully or returns an error; on every platform we
// target it does not fail (it is the same entropy source behind Xray's
// uuid.New()). A failure here means the OS CSPRNG is broken — there is no
// safe fallback for a session id, so we intentionally do not swallow it.
if _, err := rand.Read(b[:]); err != nil {
panic(err)
}
const hexdigits = "0123456789abcdef"
var h [32]byte