The posture was inverted. A client using the DHCP-supplied resolver — the router itself — was NOT intercepted: dnsmasq answered and forwarded to the ISP in the clear, so the filter, the blocklists, the per-device rules and BlockDoH were all inert for exactly the clients that did nothing wrong. A client that hardcoded 8.8.8.8 to route around us WAS intercepted, by the catch-all. Meanwhile the docs promised no DNS leaks. The default now matches the promise. Turning it on crosses a threshold that was already dangerous for anyone with two resolvers. Above one transport, a node's domain server address stops being resolved by the transport directly and goes through the client DNS plane instead — so a blocklist entry, a block_doh NXDOMAIN or any dns_rule can answer your own node's hostname, and one sloppy line in an ad list stops being an ad that got through and becomes a tunnel that never comes up. So the fix is gated on having two or more transports, not on the intercept toggle: resolver_default plus resolver_fallback always reached that threshold, long before this change. When no endpoint_resolver is configured the plane now carries a bootstrap server — the default resolver cloned with its detour dropped, keeping its type, so a DoH default stays DoH and only the tunnel hop goes. An explicit endpoint_resolver still wins. This is not a restore of the previous behaviour and the comment says so: at one transport the dialer used the default resolver WITH its detour, so a lone DoH-through-the-tunnel resolver was already a bootstrap loop. It is strictly better than what came before. Existing installs keep whatever they set — the config file is a conffile and is never replaced — and an explicit dns_intercept '0' survives the render-parse round trip, which a default-true bool otherwise makes easy to lose. The no-resolver warning stays, and no default resolver is shipped to silence it: a placeholder would remove the sentence without moving a single query, and the panel would then say a resolver was configured while nothing was filtered. Its wording is corrected instead — .lan keeps working through the built-in local transport, which the old text denied. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
50 KiB
Decisions (ADR log)
Key architectural decisions and why, so nobody re-litigates them (we spent a
whole discussion converging — see also CONTEXT.md).
D1 — Do NOT write a proxy engine from scratch
The proxy protocols (VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, AmneziaWG, Hysteria2/TUIC, transports, TLS fingerprinting) are a byte-precise, adversarial anti-DPI arms race maintained by large communities and changing monthly. A from-scratch engine would be slower, buggier, less secure, and more detectable — the opposite of "optimized." Reuse a real engine.
D2 — Engine = sing-box (via sing-box-lx), not xray-core
- AmneziaWG 2.0 (I1–I5 CPS decoy packets) is a hard requirement; sing-box has
first-class AmneziaWG, xray's is weaker. sing-box also has broader protocol
coverage (Hysteria2, TUIC, ShadowTLS, MASQUE) and is library-first
(
libbox), which suits embedding. sing-box-lx(github.com/Leadaxe/sing-box-lx) is a thin, actively-rebased downstream fork adding AmneziaWG 2.0, XHTTP, MASQUE/WARP and — usefully for us — gRPC observability of DNS queries / rules / outbounds, which feeds our stats.- License GPL-3.0, compatible with (and cleaner than) our prior GPL-2.0-or-later. See D6.
- Cost accepted: our whole control plane / generator / share-link handling is re-based from xray-shaped (v0.1) to sing-box-shaped.
D3 — FORK sing-box-lx and embed the whole product inside it (not "depend as a library")
Considered: (a) consume as a pinned Go module + Gitea mirror; (b) fork. Chosen: (b) fork, so we can embed control-plane + admin panel + DNS filter directly and integrate tightly with engine internals (DNS, routing, stats). This was a deliberate call favouring maximum integration over minimum maintenance.
To keep the fork from rotting, a mandatory discipline:
- Our code lives in new top-level dirs (
shater/,panel/,openwrt/) that never collide with upstream files on rebase. - Edits to upstream files are minimal and marked
// shater. - We rebase/merge onto sing-box-lx tags on a cadence (mirroring how sing-box-lx rebases onto sing-box). Fork = additive overlay tracking upstream, not a divergent rewrite.
mainof theshaterrepo is the fork;v0.1branch keeps the old project.
D4 — UI = thin LuCI launcher + a separate embedded admin panel
LuCI stays a minimal, pretty mini-dashboard with an "Open panel" button. The button mints a short-lived token in the authenticated LuCI/ubus session and redirects to our admin panel on its own port, served by the daemon; the panel validates the token and opens a session.
- Why not do everything in LuCI: LuCI's form/view model is limiting for the rich stats/config UX we want.
- Why not a standalone panel with its own login: a router-facing panel with its own auth is a serious security surface to get right. Bootstrapping the token from LuCI's existing, hardened auth avoids a second login system.
- The panel is a modern SPA, embedded in and served by the forked binary.
D5 — Long blocklists: don't push megalists into dnsmasq; use an efficient matcher
Naive address=/domain/# in dnsmasq holds every domain in RAM and reloads slowly
(100k–1M+ entry lists are common). Instead the engine's own domain matcher
(sing-box already compiles geosite-scale lists) or our compact matcher
(suffix hash-set + optional bloom prefilter, compiled/cached blob, dedup +
subdomain-collapse, streaming parse, refresh by content-hash) does the filtering.
Blocklist sources are flexible: inline / file / url / geosite
(geosite only when geodata is present, else that source is inert / fail-open).
D6 — License = GPL-3.0
sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files may stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is GPL-3.0 for clarity.
D7 — Keep the v0.1 feed signing identity (SUPERSEDED by D22)
The usign feed key 5ac4b177689cb8e0 (public key in dist/shater-feed.pub,
secret in Gitea secret KEY_BUILD) carries over, so routers that already trust it
keep verifying v0.2 packages. Do not regenerate it without a documented rotation.
Superseded 2026-07-25 (D22). The opkg feed this identity signed no longer exists, so there is nothing left for the key to verify. It was never rotated or compromised — it is simply unused.
dist/shater-feed.pubwas deleted from the tree; the reasoning, and how to resurrect the identity if it is ever needed again, is in D22.
D8 — Preserve, don't destroy: v0.1 lives on its branch
The reset moved the full working xray-based project to the v0.1 branch and
cleaned main. Nothing is lost; reusable logic (reliability layer, nft/routing,
sub-fetch, CI/signing, design system) is ported forward, not rewritten.
D9 — Router build uses a musl-static tag set (drop with_naive_outbound,with_purego)
Verified on the x86_64 OpenWrt VM (2026-07-14): the canonical Makefile.lx
LX_TAGS produces a binary dynamically linked to glibc and it will NOT run on
musl OpenWrt. Root cause (proven three ways — ldd, ELF PT_INTERP, static
hello-world control): with_naive_outbound pulls in cronet-go, which needs
with_purego; purego uses //go:cgo_import_dynamic to dlopen libcronet at
runtime, forcing an ELF with PT_INTERP=/lib64/ld-linux-x86-64.so.2 + PT_DYNAMIC
even with CGO_ENABLED=0. musl has no glibc loader → not found (rc 127).
- Decision: the OpenWrt/router binary uses a router-specific tag set =
LX_TAGSminuswith_naive_outbound,with_purego. That variant is fully static (ET_EXEC, noPT_INTERP) and runs directly on musl. naiveproxy outbound is not in our feature set (FEATURES.md), so nothing we ship is lost.with_awgis independent of naive/purego and stays. - Keep
badlinkname,tfogo_checklinkname0+-checklinkname=0(needed by badtls); they build fine static. - The desktop/CLI
LX_TAGSstays as upstream for any non-router use. Do not reuse the desktop tag set for the router package — defineSHATER_ROUTER_TAGS. - 2026-07-23:
with_gvisoralso dropped from the router set. The shater data plane is tproxy/redirect (netplane); generate never emits a tun inbound, so the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever appears it falls back to the system stack — re-add the tag then. REVERTED 2026-07-25 — that reasoning was wrong and shipped a dead feature. gVisor is not only the tun stack: it is the netstack of the WireGuard endpoint, which we do emit and do declare [MVP]. See D23; the tag is back and is now held there by a test. - 2026-07-23:
with_clash_apialso dropped. The admin panel is shater's own web server and generate never emits aclash_apiservice; the desktop/CLILX_TAGSkeeps the tag for external dashboards. - 2026-07-23:
with_dhcpalso dropped. shater resolver types areudp/tcp/doh/dot/local/fakeip; adhcp://DNS transport is never generated, and the slimshater/registrynever registers it.
D11 — The engine runs IN-PROCESS (box.New), not as a subprocess
v0.1 generated an xray JSON file and ran xray as a separate process. v0.2 embeds:
the daemon builds option.Options in Go and drives the engine via
box.New(box.Options{Context, Options}) + Start()/Close() in the same process
(proven in Phase 1, shater/cmd/shater-proto). This is the whole point of the
fork — it gives us in-process access to DNS query events, routing, and stats
(D5/observability) without log-scraping, and one binary to ship.
- Apply = atomic instance swap. sing-box box instances are built-then-started;
a config change rebuilds
option.Options, validates by constructing a freshbox.New(which runs every adapter's constructor), and on success swaps: start the new instance, then close the old (or close→start within a short window under the kill-switch). Commit-confirm rollback keeps the last-goodoption.Options. - The generator (v0.1
generate.go) is rewritten to emit sing-box option structs instead of xray JSON. The parsers (share-link/subscription/awg-conf) and the nft/policy-routing layer port with little change (engine-agnostic).
D12 — v0.2 binary = shaterd, our own entrypoint under shater/
The router ships a single binary shaterd = shater/cmd/shaterd: our main that
imports the engine as a library (D11) and also runs the control-plane, the DNS
filter, the stats aggregator, and the admin-panel web server. cmd/sing-box
(upstream) stays untouched for CLI/debug use. The router build uses the musl tag
set (D9); shaterd links the same engine packages cmd/sing-box does.
D10 — Ship the router binary UPX-lzma compressed
Raw static binary is ~40–43 MB; the stock OpenWrt ext4 rootfs is ~100 MB. UPX
--lzma --best takes it to ~9–11 MB (measured), a comfortable flash budget
with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep an
uncompressed artifact for debugging.
D13 — External DPI-bypass tool = ByeDPI (a SOCKS egress), NOT zapret
Decided 2026-07-14. We evaluated exactly two external desync tools — zapret
(nfqws/tpws, NFQUEUE packet plane) vs ByeDPI/ciadpi (a local SOCKS5 desync
proxy) — and picked one: ByeDPI. DPI-bypass stays a per-ruleset egress
choice (a routing rule sends selected domains to it, exactly like a WG node or
direct), never a global switch.
Why ByeDPI over zapret — decided by shater's architecture, not by raw method strength. In shater all LAN traffic is already TPROXY-intercepted into the in-process engine, and routing rules pick the egress:
- ByeDPI is an egress. It's a SOCKS5 proxy on
127.0.0.1:<port>; a rule →socksoutbound → ByeDPI desyncs (split/disorder/fake/oob/autottl, fake-TTL) → goes direct to the target. No tunnel, no exit-node bandwidth, low latency. Zero new nft rules; it never touches our packet plane. Ships as a ~100 KB musl-static binary behind one procd service. - zapret is a competing packet plane. nfqws hooks packets via its OWN
nft/NFQUEUE rules, which would fight our verified fail-closed
inet shaterTPROXY table for the same packets/marks — a second packet plane on a fail-closed router multiplies kill-switch/leak surface, and it cannot be expressed as a sing-box egress. It also entanglesnetplane(a core verified subsystem), violating the additive-overlay discipline (D3). - zapret's real edge is QUIC / bad-checksum fakes at the raw-packet level.
ByeDPI is TCP/TLS-centric and weak on QUIC. We accept that gap and close it by
routing, not by tool: a rule dropping
udp/443for desync-domains forces them to fall back to TCP+TLS, which ByeDPI handles. That stays inside our rule model — no parallel packet plane.
Native fragment is free and complementary (not a second tool). The fork already
compiles in, at the route-action level, tls_fragment +
tls_fragment_fallback_delay, tls_record_fragment, tls_spoof/tls_spoof_method,
and dialer udp_fragment. These are just engine flags on a direct egress and
cover the light "just fragment the ClientHello" case with no external binary.
ByeDPI is what we add for the stronger methods the engine lacks (fake/disorder/oob/
autottl).
Decision / implementation shape:
- Egress gains an optional
dpipreset. Values map to escalating power:off→ plain;fragment/record/spoof→ native route-action flags on adirectoutbound (in current Wave, ~free);byedpi→ asocksegress pointed at a supervised local ByeDPI (ciadpi) instance (later phase: its ownopenwrt/procd package + musl-static cross-build). shater/generatemaps the native presets now; thebyedpiegress kind is reserved inmodeland wired when its package lands.- zapret is explicitly rejected for shater and should not be revisited unless in-place QUIC packet-desync becomes a hard requirement that routing cannot cover.
When zapret would have won: a pure DPI-bypass box with no tunnel engine (then nfqws's packet power is the whole product). shater already has a full tunnel engine
- native fragment, so it wants the desync tool that composes with routing — ByeDPI.
D14 — LAN DNS anti-leak = sing-box hijack-dns route action, NOT a :53 listener
Decided 2026-07-15 after the Phase-2 gate exposed the gap (see memory
shater-dns-no-engine-listener). The netplane dnsnat chain redirected LAN :53
to the router's :53, assuming the engine answers there — but generate/engine
create NO :53 DNS server (the engine only binds the tproxy port). On the VM the
redirect landed on stock dnsmasq, which resolved via its WAN upstream outside
the tunnel = a DNS leak, and the engine's own resolvers (DoH-via-WARP, already
built by generate.buildDNS with anti-leak detours) went unused. On a real box
(dnsmasq removed) client DNS would break entirely.
Decision: use the sing-box-idiomatic hijack, not a listener. sing-box has no
simple ":53 DNS inbound"; its DNS module is designed to be fed via a hijack-dns
route action (C.RuleActionTypeHijackDNS). The tproxy inbound already sniffs
(the leading sniffRule() in generate/route.go), so:
- generate: after the sniff rule, add a route rule
protocol: dns → action hijack-dns, so every sniffed DNS query is answered by the engine's internal resolver (which routes each query through its configured detour — the anti-leak). - netplane: STOP special-casing
:53. Remove the prerouting:53 acceptand the wholednsnatchain so LAN:53is DIVERTED by the normal tproxy catch-all into the engine, wherehijack-dnscatches it. Keep the:853DoT reject (forces clients off DoT onto:53, which is now hijacked). DoH (:443) stays SNI-routed through the proxy. - dnsmasq: the tproxy divert steals LAN
:53before local delivery, so dnsmasq never sees LAN queries; it keeps serving ROUTER-local:53only. No dnsmasq removal needed for correctness, though shipping images may still drop it.
Rejected alternative (a real engine :53 listener + keep dnsnat): non-idiomatic
for sing-box, needs a bespoke DNS server, and duplicates what hijack-dns gives for
free. generate + netplane MUST stay in agreement on this — they are the two
halves of the same DNS plane.
D15 — DNS filter = sing-box rule-sets + reject DNS rules, NOT a custom matcher
Decided 2026-07-15 (Phase 4). D5 said "don't push megalists into dnsmasq; reuse the
engine's matcher OR a compact custom matcher." sing-box already ships a compiled,
memory-mapped rule-set matcher (geosite-scale, the .srs binary format + inline /
local / remote rule-sets with auto-fetch+cache), and DNS rules support
action: reject. So the DNS filter is built ON that, not a bespoke bloom matcher:
- model: a
blocklistsection (name, enabled, sourceinline|file|url|geosite, url/path/entries, responsenxdomain|zero|refuse) + anallowlist(overrides). A globaldns_filterenable. - generate/dns.go: materialize each blocklist into a sing-box rule-set (inline
rule for small inline lists;
localrule-set for a compiled/file.srs;remoterule-set for a url, letting sing-box fetch+cache with a download detour through the proxy) and emit DNS rules: allowlist-accept FIRST (higher priority), then blocklist →reject(methoddefault=NXDOMAIN, or a0.0.0.0/::predefined answer forzero). geosite sources are inert/fail-open when geodata is absent (D5). - daemon: a
blocklist updateverb (likesub update) to (re)compile file sources and warm remote caches; sing-box remote rule-sets self-refresh, so the daemon mostly compiles local sources + seeds well-known lists (StevenBlack / OISD / AdGuard) as disabled presets. Rationale: reuses a battle-tested, RAM-efficient matcher (no custom megalist code to get wrong), and it composes with the hijack-dns plane already in place (D14) — the filter rules sit on the same in-engine DNS path. Only build a custom matcher if a concrete sing-box rule-set limitation is hit (none known).
D16 — Enable sing-box cache_file; engine apply-swap must go close-first when it's shared
Decided 2026-07-15 (Phase 4). Remote rule-sets (url blocklists) and urltest/rdrc
state need sing-box's experimental.cache_file to persist across restarts,
auto-update on their interval, and keep RAM sane for megalists. So generate now
emits cache_file (enabled, /etc/shater/cache.db, falling back to
/tmp/shater-cache.db).
Consequence discovered on the VM: bbolt opens the cache DB under an EXCLUSIVE
file lock and retries ~10s before failing timeout. The engine's default
apply-swap starts the NEW box before closing the old (D11) — but the new box can't
acquire the cache lock the old still holds, so EVERY live reconcile/apply stalled
~10s then failed (allowlist toggle, blocklist update, any edit silently didn't
take). Fix (extends the D-fix for EADDRINUSE): engine.Apply now (a) treats the
cache-file lock timeout as a swap conflict, and (b) proactively takes the
close-old-then-start-new path when the incoming config shares an enabled
cache_file with the running one (sharesCacheFileLock), avoiding the ~10s stall
entirely. The fail-closed nft kill-switch covers the brief gap. Regression test
TestEngineCacheFileSwap. See shater-engine-applyswap-portconflict (same
close-first machinery, now also for the cache lock).
D17 — Audit pass: fail-closed must survive engine-start failure; several knobs were fiction
Decided 2026-07-20. A full-stack audit (8 parallel agents, every finding reproduced on the QEMU testbed) changed a few architectural positions. Recording the ones that constrain future work.
The data plane must never be absent while kill_switch=closed. Reproduced live:
a remote rule-set that could not be fetched aborted box.Start, so apply failed
and no nft table was installed at all. The router silently degraded to plain
routing — no tunnel, no filtering, no kill-switch — while the panel looked healthy.
Fail-closed had been implemented as "the engine died after starting"; it now also
covers "the engine never started", via a holding plane (RenderHoldNft): forward
blocked for the diverted interfaces, LAN-to-LAN + link-local + ND preserved, and the
daemon's own egress marks accepted so it can still reach the network to recover.
Management (SSH/LuCI/panel) is safe by construction — the holding chain hooks
forward only, never input. Status now reports plane: full|hold|none.
Rejected: "keep the previous table". After a cold boot there IS no previous table,
which is precisely the reproduced scenario.
Engine start must not depend on the network. RemoteRuleSet.StartContext returns
an error when its first fetch fails, and router.Start fast-fails the group — so a
router that boots before its ISP link comes up cannot start the engine. Remote lists
are now reachability-preflighted at generate time and omitted (loudly) rather than
handed to the engine. Accepted cost, documented in ruleset.go: a brief source
outage disables the list for up to one reconcile + probe TTL even when a good cached
copy exists, because the cache cannot be consulted (the running engine holds an
exclusive bbolt flock — see D16 — and a lock a process already holds on one fd denies
its own second open). The structurally better fix (seed an empty cache entry between
closing the old box and starting the new one, so StartContext never fetches at
start) is deliberately deferred, not overlooked.
BlockDoH does not need to exclude upstream resolvers on the route plane. The
original design punched a hole for every configured resolver IP so the engine could
reach it — which also handed every LAN client an un-blocked public DoH endpoint, i.e.
the more reputable the upstream, the wider the bypass. This was unnecessary:
engine-originated DNS dials go through DetourDialer straight to the outbound and
never traverse route rules, and they additionally carry LoopMark 0xff, which nft
accepts before the tproxy divert. The exclusion now applies to the DNS plane only
(a hostname upstream must still resolve), and the :443 rejects use the full list.
Globals.DNSMode is unimplementable and its control was removed. nftset named
the v0.1 xray architecture, which does not exist here; fakeip is a resolver type
in sing-box 1.14 (the top-level dns.fakeip block was removed upstream), already
expressible as config resolver with type=fakeip + pool, which is strictly more
capable (per-domain via a dns_rule). Wiring the toggle would have created a second
configuration path with undefined precedence. The field stays in the model for config
compatibility; the panel now documents the resolver route instead.
TPROXY cannot carry ICMP/IGMP/ESP/GRE, so this is now an explicit policy.
Globals.Untunnelable = block (default, unchanged behaviour) | icmp | direct.
Three values, not two, because the leaks differ in kind: an ICMP echo is ephemeral,
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
Fail-open degradations must be visible in the panel, not only in logread. The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
a rule whose matchers were all invalid, a rejected interface name). apply now
collects warnings from generate + netplane + model validation, classifies them
(info|warning|critical, plus section+name for deep-linking) and exposes them on
GET /api/status. Classification is by section, not by text markers: the wording
of "list skipped" varies across six sites in generate and a marker list would silently
miss the seventh.
blocklist source=url now consumes what public lists actually publish. It had
accepted only compiled .srs, so every list D15 names (StevenBlack / OISD / AdGuard,
all hosts-format) failed at apply with invalid sing-box rule-set file — a headline
feature that broke on its own documented sources. The daemon now fetches (size-capped),
parses hosts / plain-domain / AdBlock-comment formats, and compiles to .srs under
/etc/shater/lists/, which the engine loads as a local rule-set. Compiling is what
makes it affordable on the target hardware: measured on the testbed, StevenBlack's
2.4 MB of text became 80873 domains in a 491 KB .srs, costing ~0.5 MB of a rootfs
with ~33 MB free — so it ships in the production posture rather than being traded away
for the 8 KB geosite ads list.
D18 — Node health = a board with computed verdicts, NOT delete-on-failure + a dead-overlay
Decided 2026-07-24 (proxy-health plan §5.A). Upstream common/urltest recorded
only success ({Time, Delay}) and a failed check deleted the entry — a dead
member was indistinguishable from a never-measured one. On top of that, deaths
found by shater's own probing went to a private engine.dead overlay that only
the panel read: selection never saw them, and an old success in history counted
as "alive forever". Net effect, reproduced in the field: YouTube dead on a real
device while the panel showed the group healthy.
Decision: one health board, one source of truth. The history entry becomes
{LastOK, Delay, LastFail}; a failure marks (MarkFailed), never deletes —
deletion is reserved for removing a node from the config. The verdict is a
method computed on read, not a stored field: alive when the success is
fresher than the failure and younger than TTL; dead while a fresher failure is
itself fresh; untested otherwise, with TTL = max(3 × global probe interval,
10 min). Everyone who learns of a death — the group's native checker, the
observatory, a failed user dial — writes the same board; selection, balancer
slots and the panel read the same board. The engine.dead overlay is deleted.
- Rejected: keep delete-on-failure and widen the overlay to selection. Two stores of the same truth with undefined precedence; the overlay would need its own TTL/pruning and every reader would have to merge — exactly the divergence ("green panel, dead path") this decision exists to kill.
- Rejected: persist the board across reboots. Measurements go stale faster
than an overlay write is worth; after reboot everything is
untestedfor seconds until the observatory's immediate first pass. In-memory, deliberately. - Deferred, not rejected: Xray-style sliding-window stats / hysteresis. The delay of the last success is enough to rank alive members in v1; the board keys and record shape leave room to widen without migration. Flapping is visible instead through the alive↔dead flip log (info) — the single diagnostic trail. Consequence: a stale success decays (alive → untested after TTL) instead of reading as alive forever, and a dead group blocks fail-closed — visible and alertable — rather than silently falling back.
D19 — Background probing = an observatory driven by rule reachability, NOT a population sweep
Decided 2026-07-24 (proxy-health plan §5.C, replaces shater/engine/sweep.go).
The sweep probed the whole population round-robin — all nodes plus all
per-group copies, used or not — yet chain copies were excluded entirely, so the
one path users actually complained about ("rule → chain") was never measured.
The manual "Test all nodes" run duplicated the same full-population walk on
demand. On router budgets that is the wrong shape twice: work grows with the
subscription size (~200 nodes), not with what the config uses.
Decision: probe only what the rules can reach. Following the Xray model
(central observatory + balancers reading observations, §3 of the plan), a plan
is built from the applied option.Options by reachability: enabled-rule
targets (plus Final and DNS detours) → groups → members / egress copies; a
chain is probed end-to-end by dialing its exit tag through the whole hop
path (the Xray observatory-through-proxySettings equivalent); chain group-hop
members are probed through their path prefix. Everything unreferenced stays
untested and its group/chain is badged "unused" in the panel — so untested
never looks like a health problem. A freshness gate skips tags an active
group's own checker already measures; the cursor survives no-op reconciles
(the sweep's release-blocker: cron reconciles must not restart the cycle);
the manual probe-all is removed with the sweep.
- Rejected: keep the sweep. Probes hundreds of unused nodes on a router budget and still misses chains; its "coverage" is what made per-node health look authoritative while the used path went unmeasured.
- Rejected: probe every chain hop individually. Multiplied probe traffic
for diagnostics that never affects selection — no choice depends on a middle
hop's individual health. A dead middle hop makes the exit verdict honestly
dead; localization is served by the verdict flip log. - Rejected: auto-reroute rules when a group dies. A dead group blocks fail-closed — visible and alertable. Silent rerouting would hide the outage and change routing semantics behind the user's back. Consequence: the probe budget is bounded by the config, not the subscription; cold start converges in seconds (immediate first pass after apply); wanting numbers for an unused group has one honest answer — reference it from a rule.
D20 — Probe URL / Interval are global-only; per-group overrides deleted
Decided 2026-07-24 (proxy-health plan §5.D). Groups carried optional
ProbeURL/ProbeInterval overrides. That bred a documented ambiguity — one
node shared by two groups with different URLs yields incomparable delays and
needs "whose URL wins" dedup machinery (the (dial, URL) plan key) — and it
breaks the D18 board: a verdict is a (record, TTL) pair with TTL derived from
the probe interval, so per-group intervals would make the same record mean
different things to different readers. Xray's observatory has exactly one
global probeURL/probeInterval; that is the model we mapped onto (§3).
Decision: only Globals.ProbeURL/Globals.ProbeInterval. They drive the
native group checker, the observatory and the manual exit test alike (failover
keeps its 30 s default interval). Probe-plan dedup collapses to the dial path.
Migration is the standard dead-option drainage: old UCI configs carrying
probe_url/probe_interval on a group parse silently and the options vanish
on the next render.
- Rejected: keep per-group overrides. Incomparable measurements across groups, undefined semantics for shared members, and a per-group TTL that fractures the single-board verdict.
- Deferred: per-tier probe cadence (rule-critical tags more often). If ever
needed it is a field on the observatory's
ProbeJob— a scheduling knob, not a return of per-group configuration. Consequence: all delay numbers are comparable (least ping ranks apples against apples), and group settings lose two footgun fields while Settings keeps the two that actually govern every check.
D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). config rule carried THREE ways to say
where traffic is going: dst_domain (an inline domain list), dst_ip (an inline
CIDR list) and dst_ruleset (a reference to a config ruleset). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a .srs, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
Decision: dst_domain and dst_ip are removed (schema v2). dst_ruleset is
the only destination matcher. Src, dst_port and proto are untouched —
they are not lists of destinations and have no rule-set form.
- Rejected: keep the inline lists as a shorthand. "One obvious way" is the whole point; a shorthand that quietly means something different from the long form (see the bare-entry trap) is worse than no shorthand.
- Rejected: promote inline lists to rule-sets lazily at generate time. The config on disk would then not say what the router does, and the panel would have to render a list the user cannot find or edit.
The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (dst_domain) |
in a rule-set (entry) |
migrated to |
|---|---|---|---|
example.com |
exact host | host + subdomains | full:example.com |
full:example.com |
exact host | exact host | unchanged |
suffix:example.com / .example.com |
host + subdomains | host + subdomains | unchanged |
keyword:ads |
substring | substring | unchanged |
regexp:^ads\. |
pattern | pattern (added here) | unchanged |
geosite:x / geoip:x |
inert (engine field removed) | inert (unknown prefix) | unchanged |
shaterd migrate (schema v1→v2, shater/model/migrate.go) creates one inline
config ruleset per rule that still carries a legacy list — rule-<rule name>
for domains, rule-<rule name>-ip for addresses — moves the entries across with
the conversion above, appends the new name to dst_ruleset, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
rule-<name>-2). regexp: support was added to inline rule-sets in the same
change precisely so the move can be lossless.
geosite:/geoip: entries are copied VERBATIM rather than promoted to a
source=geosite rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
One deliberate semantic change, called out: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because rule_set: [a, b] matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
That AND→OR change is about the ENGINE's TCP/UDP path, and it deliberately does
not extend to the untunnelable-protocol plane (shater/apply/untunnelable.go,
the ping / IPTV / VPN-passthrough policy in nftables). There, a v1
dst_domain + dst_ip rule could never claim a packet that carries no domain, so
the plan skipped it; reading the migrated form as OR would have made the address
half suddenly decisive, and with target=direct that means an upgrade quietly
sending previously-tunnelled ICMP out with the client's real source address. A
rule whose rule-sets are known to match by NAME is therefore still skipped by that
plan, and the skip is reported ("a routing rule matches by name … as well as by
address"). Split the rule in two if you want the addresses decided there.
The vocabulary is about ENTRIES YOU TYPE, not about every list body
The table above is the vocabulary of an inline rule-set's entry values (and
of the DNS-filter/device lists, which share the classifier). The other two rule-set
sources are not other spellings of it:
| source | what it is | vocabulary |
|---|---|---|
inline |
entries you type | the table above |
url → .srs / .json |
a compiled rule-set, engine-owned | the engine's, not ours |
url → anything else |
a hosts / one-domain-per-line / AdBlock TEXT FILE | none — every line is a domain plus its subdomains |
file |
a local .srs / .json |
the engine's, not ours |
Rejected: run text lists through the entry classifier too. A published
AdGuard/OISD list is full of colon-bearing tokens that are ordinary filter syntax
(##…:has(…), $domain=, absolute URLs); classifying them would either mis-import
them or bury the operator under hundreds of "unrecognised prefix" warnings per
list. The formats also disagree structurally — a hosts line carries several names,
so the text parser works per token, while an entry is a whole line. And regexp:
arriving from a third-party URL is a pattern compiled into the router's matcher and
evaluated per query, which is a very different proposition from one the operator
typed.
So the difference stands and is paid for in diagnostics instead: a text list
containing full: / suffix: / keyword: / regexp: is reported per list, on
every generate, naming the entries and pointing at source=inline where they work
(warnListEntryVocabulary, shater/generate/ruleset.go). The check tests only
those four markers, never the general word: shape, so it fires on a human's
mistake and stays quiet on published filter syntax.
Consequence: one destination mechanism, one vocabulary, one place a list is edited; every list is compiled once and reused. The panel's rule editor drops its Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of the rulesets that already exist, and nothing more. Creating and filling a list stays in the Rulesets panel — rejected: a "create a list from here" shortcut in the rule editor, because a second place to author a list is a second place for its semantics and its duplicate-name rules to drift, and the whole point of this decision was to stop having two.
D22 — One packaging lane: apk. The opkg/.ipk lane is deleted, not disabled
Decided 2026-07-25 (product owner). CI built and published TWO signed feeds from
every run: opkg/usign (.ipk + Packages.gz, OpenWrt 24.10) and apk/EC (.apk +
packages.adb, OpenWrt/ImmortalWrt 25.12). The opkg half served nobody. Checked
on the actual hardware, not inferred:
| Device | Firmware | pkg arch | package manager |
|---|---|---|---|
mini_router (BPi-R3 Mini) |
ImmortalWrt 25.12.1 | aarch64_cortex-a53 |
apk-tools 3.0.5 |
main_router (BPi-R4) |
OpenWrt 25.12.0 | aarch64_cortex-a53 |
apk-tools 3.0.5 — no opkg binary on the system at all |
Decision: delete the opkg lane outright. Removed: the build + release
jobs from .gitea/workflows/release.yml; ci/build-feed.sh, ci/sdk-build.sh,
ci/make-index.sh, ci/install-usign.sh; and the trust anchor
dist/shater-feed.pub. The Gitea secret KEY_BUILD is now referenced by
nothing and can be deleted from the repo settings. ci/version.sh,
ci/gitea-release.sh and ci/fetch-sdk.sh are shared or apk-only and stay.
- Rejected: keep the lane but stop triggering it (comment it out / gate it on a dispatch input). Dead code in CI is worse than no code: it keeps a second SDK matrix, a second signing key and a second feed layout alive in everyone's head and in every future edit, and it silently rots because nothing runs it. The 24.10 SDK images it pins are themselves a frozen dependency.
- Rejected: keep
dist/shater-feed.pubas a historical artifact. A committed trust anchor is an instruction — it invites someone to follow the old install path for a feed that is no longer produced. Nothing is lost by removing it: git history still holds the file, the SECRET half is untouched inKEY_BUILD, and a usign secret key blob contains its own public half, so the identity can be reconstructed if a 24.10 device ever has to be served again. Deleting the file is reversible; a stale trust anchor pointing at an unmaintained feed is the thing that quietly misleads. - Not done: revoking or rotating the usign key. There is no incident. It is retired, not burned (D7).
Consequence: one SDK, one key, one feed layout, one set of install instructions.
It also makes the rolling release apk-latest-<arch> the only install path
that does not require hand-editing a file per release — which is why the same
change fixed it: publishing was an either/or (apk-latest-<arch> on dispatch,
ELSE apk-vX.Y.Z-<arch> on a tag), so once releases moved to tag pushes the
rolling pointer stopped being written and froze at 0.2.0 while v0.2.9/v0.2.10
shipped — routers on the rolling URL got a successful, silent apk update with
nothing new. release-apk now writes the rolling pointer on every run and
asserts, by reading the published release back over the Gitea API, that it holds
our three tag-versioned packages at exactly the version just built and no asset
at any other version.
D23 — The router tag set is a checked contract, not a string literal
with_gvisor was trimmed from the router set on 2026-07-23 (D9) as "unreachable
code: we never emit a tun inbound". True about tun — and irrelevant, because
gVisor is also the netstack of the WireGuard endpoint, which shater emits and
FEATURES.md declares [MVP] (AmneziaWG is called "a driving requirement"). Every
binary shipped between then and 2026-07-25 answered a configured WireGuard node
with:
create instance: initialize endpoint[0]: create WireGuard device:
gVisor is not included in this build, rebuild with -tags with_gvisor
transport/wireguard/device_stack_stub.go (//go:build !with_gvisor) returns
tun.ErrGVisorNotIncluded from both device constructors, so
system_interface: true is not an escape hatch either: WireGuard was 100% dead
in the shipped artifact while the panel offered it, the parser accepted wg://,
awg:// and wg-quick .conf imports, and the owner had 7 WireGuard sections in
UCI on a production router.
- Decision:
with_gvisoris part of the router tag set and stays there for as long as we ship WireGuard. It costs ~2.8 MB raw / ~0.65 MB UPX per arch (measured 2026-07-25, both arches;/overlayon the production router is 6.9 GB with 205 MB used). A tag whose absence turns a declared feature into a runtime error is not "dead weight" — it is the feature.
Why the bug was invisible, and what now makes it visible
The defect was not a typo in a tag list. It was that nothing connected the tag
list to the feature list, and the shipped tag combination was the one build
configuration nothing exercised: the whole test suite compiles with the FULL
upstream set (with_gvisor included), so TestAmneziaWGEndpoint passed happily
while the artifact it was supposed to vouch for could not create a WireGuard
device. Tests proved the code was right; they never proved the build was.
Three pieces now hold it together:
- One definition of the set —
scripts/router-tags.sh(SHATER_ROUTER_TAGSSHATER_ROUTER_LDFLAGS), sourced byscripts/build-shaterd.shand by the checker. The tag list used to live as a literal inside the build script, i.e. in a file no test reads. A second copy is a second truth.
- A declared-feature table —
shater/buildtags: every tag-gated capability we promise, with the exact tags it needs to run and why (the code anchor).TestRouterTagSetCoversDeclaredFeaturesparses the shell file and fails if a declared feature lost a tag. It needs no build tags, no Linux, no network and no privileges, so it runs in every plaingo test ./...— including on the Windows dev host, where nothing else can see the shipped configuration. - A construction test under the shipped tags —
shater/generate.TestShippedTagSetConstructsDeclaredProtocolsdrives one node of every declared protocol (ss/vmess/trojan/vless ws-grpc-httpupgrade-quic- xhttp/REALITY/uTLS-fp/hysteria2/tuic/wg/awg) throughbox.New+Start.scripts/check-router-tags.shruns it withSHATER_ROUTER_TAGS, and CI runs that script (.gitea/workflows/release.yml) before the artifact is built. In a router-tag-set run nothing may be skipped: a protocol that is not compiled in fails the run instead of quietly disappearing from it.
(2) catches a trim the moment it is made and names the feature it kills; (3)
catches what a list comparison cannot — a tag that is present but insufficient.
Neither is a substitute for the other. A new protocol in shater/parse +
shater/generate means a new row in buildtags.Features and a new probe case;
TestEveryTagGatedFeatureIsProbed fails until both exist.
- Rejected: "just add the tag". The one-line fix restores WireGuard and leaves the mechanism that hid it fully intact — the next size-driven trim is equally invisible. The tag is the smallest part of this decision.
- Rejected: run the WHOLE test suite with the router tag set in CI. It is the obvious move and it does not work: parts of the suite legitimately depend on upstream-only tags, and the run costs a second full compile of a 25 MB binary's worth of packages on every release. A focused, unprivileged construction test buys the same evidence for ~10 s and, unlike a full run, can be required to skip nothing.
- Rejected: assert the tag set against upstream's
DEFAULT_BUILD_TAGS. That makes any trim a failure, which turns the check into noise and re-litigates D9 on every upstream rebase. The contract is with our own feature list, not with upstream's. - Not done: dropping
with_lx_command. It is inert forshaterd— nothing undershater/importssing-box/daemonorexperimental/libbox, andgo list -deps ./shater/cmd/shaterdlinks neither, so it costs zero bytes. It stays only so the router set remains a subset of the lx desktop set. Noted because "a tag that buys nothing" is the mirror image of this bug and should be removed deliberately, not silently.
D24 — DNS interception is the DEFAULT (dns_intercept=1), not an opt-in
Decided 2026-07-26. Globals.DNSIntercept shipped as opt-in (default false, and
absent from both DefaultGlobals and the shipped /etc/config/shater). The result
was an inverted posture, which is the reason this is a decision and not a
preference:
- a client with standard settings — DNS = the router's address, exactly what
DHCP hands out — sent its queries to the router. The nft
:53divert was behind the flag (netplane/nft.go), and the rule right after it is an unconditionalfib daddr type local accept, so the query was delivered locally to dnsmasq and forwarded to the ISP in the clear: no blocklists, no per-device DNS rules, no Block-DoH, no resolver detour, nothing; - a client that hard-coded
8.8.8.8"to bypass the router" was addressing a non-local IP and was caught by the ordinary tproxy catch-all.
The obedient client leaked; the evader did not. Meanwhile FEATURES.md, README.md
and D14 all promised "no DNS leaks" and "dnsmasq never sees LAN queries" — true only
for the traffic pattern the default did not cover. dns_intercept appeared nowhere
in docs-shater/ at all.
Decision: DNSIntercept is seeded ON in model.DefaultGlobals, and the shipped
/etc/config/shater carries an explicit option dns_intercept '1'. Nothing about
the interception MECHANISM changed — only which side of the switch is the default.
.lan and the private PTR zones keep working, and that is a pre-existing part of
the mechanism, not something bolted on for this flip. generate/dns.go adds a
synthetic DNS server (shater-local-dns, plain UDP to 127.0.0.1:53, detour
direct, so the daemon's own loop-mark keeps it out of the divert) and PREPENDS a
domain_suffix rule for lan + the RFC6303 private reverse zones, ahead of every
device/filter rule. Two honest limitations: it hardcodes lan (a router whose
dnsmasq domain was changed needs a config dns_rule for the new suffix), and it
only exists when the model has at least one config resolver — with none, buildDNS
emits no DNS plane at all and the engine falls back to its built-in local
transport, which reads /etc/resolv.conf (127.0.0.1 → dnsmasq), so local names
still resolve but nothing is filtered.
A dead engine does NOT black out the LAN's DNS. This was the first thing checked, because "intercept everything" invites the reading "engine down = no DNS anywhere", and that is not what happens:
- the fail-closed holding plane (D17,
RenderHoldNft) hooksforwardONLY. A query addressed to the router is INPUT-hook traffic, so dnsmasq answers it as it always did — unfiltered and plaintext to the ISP. Deliberate: blocking it would also cut the daemon's own name resolution and with it any chance of self-recovery; - with the FULL plane loaded and the engine's tproxy socket gone, the
tproxystatement returnsNFT_BREAK, which aborts its own rule; the packet continues down the chain into the samefib daddr type local acceptand reaches dnsmasq.
So the failure mode is a DNS fail-open (working, unfiltered) while client TRAFFIC stays fail-closed — and a query aimed at an EXTERNAL resolver is dropped with the rest of the forwarded traffic. Operators must know this: "the tunnel is down" does not mean "DNS is private".
Existing installs. /etc/config/shater is a conffile
(openwrt/shater-core/Makefile), so an upgrade never replaces it:
- a config that never mentioned the option (all of them, before this change) now parses over the ON seed and starts intercepting on the next apply. That is the intended behaviour change, and the only one this decision makes;
- an explicit
option dns_intercept '0'keeps winning. It survives theWriteUCI→ReadUCIround-trip becauserender.goemits booleans ALWAYS — the trap a default-true bool has and a default-false one does not: a value omitted at false would come back as the seed and silently re-enable itself.shater/model/dnsintercept_test.gopins both directions, plus the shipped file.
Not done: silencing the "no resolvers configured" warning by shipping a resolver.
With interception on and no config resolver, generate warns — and it is right to:
every client query now lands in an engine that has no resolver plane, so it is
answered by the system resolver (dnsmasq → the ISP, in the clear) with filtering and
anti-leak inert. Shipping a type local resolver would make the warning disappear
while changing nothing about where the queries go: the panel would show a configured
resolver and the operator would believe DNS was handled. That is the inverted lie
this project keeps deleting. The warning stays; what it needs is the accurate
wording (it currently claims .lan breaks, which the fallback above disproves), not
a workaround. Note also that a fresh install ships INERT (enabled '0') and
Reconcile tears down instead of generating, so the warning cannot appear before the
operator has enabled the stack — at which point it describes their live config.
OPEN, and it gates shipping this default: the synthetic local server changes how proxy-endpoint DOMAINS are resolved. Found while landing D24, reproduced on Linux with one resolver and a node addressed by a hostname:
common/dialer/dialer.goresolves a domain server address throughroute.default_domain_resolver; when that is unset it usesdnsTransport.Default()— the engine's built-inlocaltransport, i.e. a bootstrap-DIRECT lookup — but only while fewer than two DNS transports exist. With two or more and no default, it reports themissing-domain-resolverdeprecation and leaves the query transport nil, sodns.Router.Lookupfalls back tolookupWithRules: the CLIENT DNS plane.dns_interceptaddsshater-local-dns, which takes a single-resolver config from one transport to two. So a config whose only resolver is DoH-through-the-tunnel — the recommended anti-leak setup — would start resolving its own node's hostname through that same tunnel: a bootstrap loop where there was none.- Evidence: the same model emits no deprecation notice with
dns_intercept=0and twomissing-domain-resolvernotices withdns_intercept=1;generate.TestDNSFilterRemoteBlocklistHTTPClient(Linux-only) fails on exactly that notice and is deliberately left failing rather than relaxed.
The fix belongs in generate (route.go:160 already sets
route.default_domain_resolver from endpointResolver(), which is opt-in and unset
by default): when buildDNS emits the synthetic local server and no endpoint resolver
is configured, default_domain_resolver must be pointed at a bootstrap-direct
server, which restores exactly the pre-D24 behaviour and clears the notice. Until
that lands, an operator can get the same result by setting endpoint_resolver to a
direct resolver. Note the hazard is not created by D24 — any config with two
resolvers has it today; the default merely makes it universal.