Files
shater/docs-shater/DECISIONS.md
T
omarandClaude Opus 5 a0de597d69
test / go + panel tests (push) Successful in 4m56s
feat(dns): intercept by default, and bootstrap node addresses off the tunnel
The posture was inverted. A client using the DHCP-supplied resolver — the router
itself — was NOT intercepted: dnsmasq answered and forwarded to the ISP in the
clear, so the filter, the blocklists, the per-device rules and BlockDoH were all
inert for exactly the clients that did nothing wrong. A client that hardcoded
8.8.8.8 to route around us WAS intercepted, by the catch-all. Meanwhile the
docs promised no DNS leaks. The default now matches the promise.

Turning it on crosses a threshold that was already dangerous for anyone with two
resolvers. Above one transport, a node's domain server address stops being
resolved by the transport directly and goes through the client DNS plane
instead — so a blocklist entry, a block_doh NXDOMAIN or any dns_rule can answer
your own node's hostname, and one sloppy line in an ad list stops being an ad
that got through and becomes a tunnel that never comes up.

So the fix is gated on having two or more transports, not on the intercept
toggle: resolver_default plus resolver_fallback always reached that threshold,
long before this change. When no endpoint_resolver is configured the plane now
carries a bootstrap server — the default resolver cloned with its detour
dropped, keeping its type, so a DoH default stays DoH and only the tunnel hop
goes. An explicit endpoint_resolver still wins.

This is not a restore of the previous behaviour and the comment says so: at one
transport the dialer used the default resolver WITH its detour, so a lone
DoH-through-the-tunnel resolver was already a bootstrap loop. It is strictly
better than what came before.

Existing installs keep whatever they set — the config file is a conffile and is
never replaced — and an explicit dns_intercept '0' survives the render-parse
round trip, which a default-true bool otherwise makes easy to lose.

The no-resolver warning stays, and no default resolver is shipped to silence it:
a placeholder would remove the sentence without moving a single query, and the
panel would then say a resolver was configured while nothing was filtered. Its
wording is corrected instead — .lan keeps working through the built-in local
transport, which the old text denied.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:54:15 +03:00

50 KiB
Raw Blame History

Decisions (ADR log)

Key architectural decisions and why, so nobody re-litigates them (we spent a whole discussion converging — see also CONTEXT.md).

D1 — Do NOT write a proxy engine from scratch

The proxy protocols (VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, AmneziaWG, Hysteria2/TUIC, transports, TLS fingerprinting) are a byte-precise, adversarial anti-DPI arms race maintained by large communities and changing monthly. A from-scratch engine would be slower, buggier, less secure, and more detectable — the opposite of "optimized." Reuse a real engine.

D2 — Engine = sing-box (via sing-box-lx), not xray-core

  • AmneziaWG 2.0 (I1–I5 CPS decoy packets) is a hard requirement; sing-box has first-class AmneziaWG, xray's is weaker. sing-box also has broader protocol coverage (Hysteria2, TUIC, ShadowTLS, MASQUE) and is library-first (libbox), which suits embedding.
  • sing-box-lx (github.com/Leadaxe/sing-box-lx) is a thin, actively-rebased downstream fork adding AmneziaWG 2.0, XHTTP, MASQUE/WARP and — usefully for us — gRPC observability of DNS queries / rules / outbounds, which feeds our stats.
  • License GPL-3.0, compatible with (and cleaner than) our prior GPL-2.0-or-later. See D6.
  • Cost accepted: our whole control plane / generator / share-link handling is re-based from xray-shaped (v0.1) to sing-box-shaped.

D3 — FORK sing-box-lx and embed the whole product inside it (not "depend as a library")

Considered: (a) consume as a pinned Go module + Gitea mirror; (b) fork. Chosen: (b) fork, so we can embed control-plane + admin panel + DNS filter directly and integrate tightly with engine internals (DNS, routing, stats). This was a deliberate call favouring maximum integration over minimum maintenance.

To keep the fork from rotting, a mandatory discipline:

  • Our code lives in new top-level dirs (shater/, panel/, openwrt/) that never collide with upstream files on rebase.
  • Edits to upstream files are minimal and marked // shater.
  • We rebase/merge onto sing-box-lx tags on a cadence (mirroring how sing-box-lx rebases onto sing-box). Fork = additive overlay tracking upstream, not a divergent rewrite.
  • main of the shater repo is the fork; v0.1 branch keeps the old project.

D4 — UI = thin LuCI launcher + a separate embedded admin panel

LuCI stays a minimal, pretty mini-dashboard with an "Open panel" button. The button mints a short-lived token in the authenticated LuCI/ubus session and redirects to our admin panel on its own port, served by the daemon; the panel validates the token and opens a session.

  • Why not do everything in LuCI: LuCI's form/view model is limiting for the rich stats/config UX we want.
  • Why not a standalone panel with its own login: a router-facing panel with its own auth is a serious security surface to get right. Bootstrapping the token from LuCI's existing, hardened auth avoids a second login system.
  • The panel is a modern SPA, embedded in and served by the forked binary.

D5 — Long blocklists: don't push megalists into dnsmasq; use an efficient matcher

Naive address=/domain/# in dnsmasq holds every domain in RAM and reloads slowly (100k–1M+ entry lists are common). Instead the engine's own domain matcher (sing-box already compiles geosite-scale lists) or our compact matcher (suffix hash-set + optional bloom prefilter, compiled/cached blob, dedup + subdomain-collapse, streaming parse, refresh by content-hash) does the filtering. Blocklist sources are flexible: inline / file / url / geosite (geosite only when geodata is present, else that source is inert / fail-open).

D6 — License = GPL-3.0

sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files may stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is GPL-3.0 for clarity.

D7 — Keep the v0.1 feed signing identity (SUPERSEDED by D22)

The usign feed key 5ac4b177689cb8e0 (public key in dist/shater-feed.pub, secret in Gitea secret KEY_BUILD) carries over, so routers that already trust it keep verifying v0.2 packages. Do not regenerate it without a documented rotation.

Superseded 2026-07-25 (D22). The opkg feed this identity signed no longer exists, so there is nothing left for the key to verify. It was never rotated or compromised — it is simply unused. dist/shater-feed.pub was deleted from the tree; the reasoning, and how to resurrect the identity if it is ever needed again, is in D22.

D8 — Preserve, don't destroy: v0.1 lives on its branch

The reset moved the full working xray-based project to the v0.1 branch and cleaned main. Nothing is lost; reusable logic (reliability layer, nft/routing, sub-fetch, CI/signing, design system) is ported forward, not rewritten.

D9 — Router build uses a musl-static tag set (drop with_naive_outbound,with_purego)

Verified on the x86_64 OpenWrt VM (2026-07-14): the canonical Makefile.lx LX_TAGS produces a binary dynamically linked to glibc and it will NOT run on musl OpenWrt. Root cause (proven three ways — ldd, ELF PT_INTERP, static hello-world control): with_naive_outbound pulls in cronet-go, which needs with_purego; purego uses //go:cgo_import_dynamic to dlopen libcronet at runtime, forcing an ELF with PT_INTERP=/lib64/ld-linux-x86-64.so.2 + PT_DYNAMIC even with CGO_ENABLED=0. musl has no glibc loader → not found (rc 127).

  • Decision: the OpenWrt/router binary uses a router-specific tag set = LX_TAGS minus with_naive_outbound,with_purego. That variant is fully static (ET_EXEC, no PT_INTERP) and runs directly on musl. naiveproxy outbound is not in our feature set (FEATURES.md), so nothing we ship is lost. with_awg is independent of naive/purego and stays.
  • Keep badlinkname,tfogo_checklinkname0 + -checklinkname=0 (needed by badtls); they build fine static.
  • The desktop/CLI LX_TAGS stays as upstream for any non-router use. Do not reuse the desktop tag set for the router package — define SHATER_ROUTER_TAGS.
  • 2026-07-23: with_gvisor also dropped from the router set. The shater data plane is tproxy/redirect (netplane); generate never emits a tun inbound, so the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever appears it falls back to the system stack — re-add the tag then. REVERTED 2026-07-25 — that reasoning was wrong and shipped a dead feature. gVisor is not only the tun stack: it is the netstack of the WireGuard endpoint, which we do emit and do declare [MVP]. See D23; the tag is back and is now held there by a test.
  • 2026-07-23: with_clash_api also dropped. The admin panel is shater's own web server and generate never emits a clash_api service; the desktop/CLI LX_TAGS keeps the tag for external dashboards.
  • 2026-07-23: with_dhcp also dropped. shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS transport is never generated, and the slim shater/registry never registers it.

D11 — The engine runs IN-PROCESS (box.New), not as a subprocess

v0.1 generated an xray JSON file and ran xray as a separate process. v0.2 embeds: the daemon builds option.Options in Go and drives the engine via box.New(box.Options{Context, Options}) + Start()/Close() in the same process (proven in Phase 1, shater/cmd/shater-proto). This is the whole point of the fork — it gives us in-process access to DNS query events, routing, and stats (D5/observability) without log-scraping, and one binary to ship.

  • Apply = atomic instance swap. sing-box box instances are built-then-started; a config change rebuilds option.Options, validates by constructing a fresh box.New (which runs every adapter's constructor), and on success swaps: start the new instance, then close the old (or close→start within a short window under the kill-switch). Commit-confirm rollback keeps the last-good option.Options.
  • The generator (v0.1 generate.go) is rewritten to emit sing-box option structs instead of xray JSON. The parsers (share-link/subscription/awg-conf) and the nft/policy-routing layer port with little change (engine-agnostic).

D12 — v0.2 binary = shaterd, our own entrypoint under shater/

The router ships a single binary shaterd = shater/cmd/shaterd: our main that imports the engine as a library (D11) and also runs the control-plane, the DNS filter, the stats aggregator, and the admin-panel web server. cmd/sing-box (upstream) stays untouched for CLI/debug use. The router build uses the musl tag set (D9); shaterd links the same engine packages cmd/sing-box does.

D10 — Ship the router binary UPX-lzma compressed

Raw static binary is ~40–43 MB; the stock OpenWrt ext4 rootfs is ~100 MB. UPX --lzma --best takes it to ~9–11 MB (measured), a comfortable flash budget with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep an uncompressed artifact for debugging.

D13 — External DPI-bypass tool = ByeDPI (a SOCKS egress), NOT zapret

Decided 2026-07-14. We evaluated exactly two external desync tools — zapret (nfqws/tpws, NFQUEUE packet plane) vs ByeDPI/ciadpi (a local SOCKS5 desync proxy) — and picked one: ByeDPI. DPI-bypass stays a per-ruleset egress choice (a routing rule sends selected domains to it, exactly like a WG node or direct), never a global switch.

Why ByeDPI over zapret — decided by shater's architecture, not by raw method strength. In shater all LAN traffic is already TPROXY-intercepted into the in-process engine, and routing rules pick the egress:

  • ByeDPI is an egress. It's a SOCKS5 proxy on 127.0.0.1:<port>; a rule → socks outbound → ByeDPI desyncs (split/disorder/fake/oob/autottl, fake-TTL) → goes direct to the target. No tunnel, no exit-node bandwidth, low latency. Zero new nft rules; it never touches our packet plane. Ships as a ~100 KB musl-static binary behind one procd service.
  • zapret is a competing packet plane. nfqws hooks packets via its OWN nft/NFQUEUE rules, which would fight our verified fail-closed inet shater TPROXY table for the same packets/marks — a second packet plane on a fail-closed router multiplies kill-switch/leak surface, and it cannot be expressed as a sing-box egress. It also entangles netplane (a core verified subsystem), violating the additive-overlay discipline (D3).
  • zapret's real edge is QUIC / bad-checksum fakes at the raw-packet level. ByeDPI is TCP/TLS-centric and weak on QUIC. We accept that gap and close it by routing, not by tool: a rule dropping udp/443 for desync-domains forces them to fall back to TCP+TLS, which ByeDPI handles. That stays inside our rule model — no parallel packet plane.

Native fragment is free and complementary (not a second tool). The fork already compiles in, at the route-action level, tls_fragment + tls_fragment_fallback_delay, tls_record_fragment, tls_spoof/tls_spoof_method, and dialer udp_fragment. These are just engine flags on a direct egress and cover the light "just fragment the ClientHello" case with no external binary. ByeDPI is what we add for the stronger methods the engine lacks (fake/disorder/oob/ autottl).

Decision / implementation shape:

  1. Egress gains an optional dpi preset. Values map to escalating power: off → plain; fragment/record/spoof → native route-action flags on a direct outbound (in current Wave, ~free); byedpi → a socks egress pointed at a supervised local ByeDPI (ciadpi) instance (later phase: its own openwrt/ procd package + musl-static cross-build).
  2. shater/generate maps the native presets now; the byedpi egress kind is reserved in model and wired when its package lands.
  3. zapret is explicitly rejected for shater and should not be revisited unless in-place QUIC packet-desync becomes a hard requirement that routing cannot cover.

When zapret would have won: a pure DPI-bypass box with no tunnel engine (then nfqws's packet power is the whole product). shater already has a full tunnel engine

  • native fragment, so it wants the desync tool that composes with routing — ByeDPI.

D14 — LAN DNS anti-leak = sing-box hijack-dns route action, NOT a :53 listener

Decided 2026-07-15 after the Phase-2 gate exposed the gap (see memory shater-dns-no-engine-listener). The netplane dnsnat chain redirected LAN :53 to the router's :53, assuming the engine answers there — but generate/engine create NO :53 DNS server (the engine only binds the tproxy port). On the VM the redirect landed on stock dnsmasq, which resolved via its WAN upstream outside the tunnel = a DNS leak, and the engine's own resolvers (DoH-via-WARP, already built by generate.buildDNS with anti-leak detours) went unused. On a real box (dnsmasq removed) client DNS would break entirely.

Decision: use the sing-box-idiomatic hijack, not a listener. sing-box has no simple ":53 DNS inbound"; its DNS module is designed to be fed via a hijack-dns route action (C.RuleActionTypeHijackDNS). The tproxy inbound already sniffs (the leading sniffRule() in generate/route.go), so:

  • generate: after the sniff rule, add a route rule protocol: dns → action hijack-dns, so every sniffed DNS query is answered by the engine's internal resolver (which routes each query through its configured detour — the anti-leak).
  • netplane: STOP special-casing :53. Remove the prerouting :53 accept and the whole dnsnat chain so LAN :53 is DIVERTED by the normal tproxy catch-all into the engine, where hijack-dns catches it. Keep the :853 DoT reject (forces clients off DoT onto :53, which is now hijacked). DoH (:443) stays SNI-routed through the proxy.
  • dnsmasq: the tproxy divert steals LAN :53 before local delivery, so dnsmasq never sees LAN queries; it keeps serving ROUTER-local :53 only. No dnsmasq removal needed for correctness, though shipping images may still drop it.

Rejected alternative (a real engine :53 listener + keep dnsnat): non-idiomatic for sing-box, needs a bespoke DNS server, and duplicates what hijack-dns gives for free. generate + netplane MUST stay in agreement on this — they are the two halves of the same DNS plane.

D15 — DNS filter = sing-box rule-sets + reject DNS rules, NOT a custom matcher

Decided 2026-07-15 (Phase 4). D5 said "don't push megalists into dnsmasq; reuse the engine's matcher OR a compact custom matcher." sing-box already ships a compiled, memory-mapped rule-set matcher (geosite-scale, the .srs binary format + inline / local / remote rule-sets with auto-fetch+cache), and DNS rules support action: reject. So the DNS filter is built ON that, not a bespoke bloom matcher:

  • model: a blocklist section (name, enabled, source inline|file|url|geosite, url/path/entries, response nxdomain|zero|refuse) + an allowlist (overrides). A global dns_filter enable.
  • generate/dns.go: materialize each blocklist into a sing-box rule-set (inline rule for small inline lists; local rule-set for a compiled/file .srs; remote rule-set for a url, letting sing-box fetch+cache with a download detour through the proxy) and emit DNS rules: allowlist-accept FIRST (higher priority), then blocklist → reject (method default=NXDOMAIN, or a 0.0.0.0/:: predefined answer for zero). geosite sources are inert/fail-open when geodata is absent (D5).
  • daemon: a blocklist update verb (like sub update) to (re)compile file sources and warm remote caches; sing-box remote rule-sets self-refresh, so the daemon mostly compiles local sources + seeds well-known lists (StevenBlack / OISD / AdGuard) as disabled presets. Rationale: reuses a battle-tested, RAM-efficient matcher (no custom megalist code to get wrong), and it composes with the hijack-dns plane already in place (D14) — the filter rules sit on the same in-engine DNS path. Only build a custom matcher if a concrete sing-box rule-set limitation is hit (none known).

D16 — Enable sing-box cache_file; engine apply-swap must go close-first when it's shared

Decided 2026-07-15 (Phase 4). Remote rule-sets (url blocklists) and urltest/rdrc state need sing-box's experimental.cache_file to persist across restarts, auto-update on their interval, and keep RAM sane for megalists. So generate now emits cache_file (enabled, /etc/shater/cache.db, falling back to /tmp/shater-cache.db).

Consequence discovered on the VM: bbolt opens the cache DB under an EXCLUSIVE file lock and retries ~10s before failing timeout. The engine's default apply-swap starts the NEW box before closing the old (D11) — but the new box can't acquire the cache lock the old still holds, so EVERY live reconcile/apply stalled ~10s then failed (allowlist toggle, blocklist update, any edit silently didn't take). Fix (extends the D-fix for EADDRINUSE): engine.Apply now (a) treats the cache-file lock timeout as a swap conflict, and (b) proactively takes the close-old-then-start-new path when the incoming config shares an enabled cache_file with the running one (sharesCacheFileLock), avoiding the ~10s stall entirely. The fail-closed nft kill-switch covers the brief gap. Regression test TestEngineCacheFileSwap. See shater-engine-applyswap-portconflict (same close-first machinery, now also for the cache lock).

D17 — Audit pass: fail-closed must survive engine-start failure; several knobs were fiction

Decided 2026-07-20. A full-stack audit (8 parallel agents, every finding reproduced on the QEMU testbed) changed a few architectural positions. Recording the ones that constrain future work.

The data plane must never be absent while kill_switch=closed. Reproduced live: a remote rule-set that could not be fetched aborted box.Start, so apply failed and no nft table was installed at all. The router silently degraded to plain routing — no tunnel, no filtering, no kill-switch — while the panel looked healthy. Fail-closed had been implemented as "the engine died after starting"; it now also covers "the engine never started", via a holding plane (RenderHoldNft): forward blocked for the diverted interfaces, LAN-to-LAN + link-local + ND preserved, and the daemon's own egress marks accepted so it can still reach the network to recover. Management (SSH/LuCI/panel) is safe by construction — the holding chain hooks forward only, never input. Status now reports plane: full|hold|none. Rejected: "keep the previous table". After a cold boot there IS no previous table, which is precisely the reproduced scenario.

Engine start must not depend on the network. RemoteRuleSet.StartContext returns an error when its first fetch fails, and router.Start fast-fails the group — so a router that boots before its ISP link comes up cannot start the engine. Remote lists are now reachability-preflighted at generate time and omitted (loudly) rather than handed to the engine. Accepted cost, documented in ruleset.go: a brief source outage disables the list for up to one reconcile + probe TTL even when a good cached copy exists, because the cache cannot be consulted (the running engine holds an exclusive bbolt flock — see D16 — and a lock a process already holds on one fd denies its own second open). The structurally better fix (seed an empty cache entry between closing the old box and starting the new one, so StartContext never fetches at start) is deliberately deferred, not overlooked.

BlockDoH does not need to exclude upstream resolvers on the route plane. The original design punched a hole for every configured resolver IP so the engine could reach it — which also handed every LAN client an un-blocked public DoH endpoint, i.e. the more reputable the upstream, the wider the bypass. This was unnecessary: engine-originated DNS dials go through DetourDialer straight to the outbound and never traverse route rules, and they additionally carry LoopMark 0xff, which nft accepts before the tproxy divert. The exclusion now applies to the DNS plane only (a hostname upstream must still resolve), and the :443 rejects use the full list.

Globals.DNSMode is unimplementable and its control was removed. nftset named the v0.1 xray architecture, which does not exist here; fakeip is a resolver type in sing-box 1.14 (the top-level dns.fakeip block was removed upstream), already expressible as config resolver with type=fakeip + pool, which is strictly more capable (per-domain via a dns_rule). Wiring the toggle would have created a second configuration path with undefined precedence. The field stays in the model for config compatibility; the panel now documents the resolver route instead.

TPROXY cannot carry ICMP/IGMP/ESP/GRE, so this is now an explicit policy. Globals.Untunnelable = block (default, unchanged behaviour) | icmp | direct. Three values, not two, because the leaks differ in kind: an ICMP echo is ephemeral, user-initiated and reveals the address only to a host the user deliberately contacted, whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".

Fail-open degradations must be visible in the panel, not only in logread. The audit deliberately converted many aborts into warn-and-continue (an unfetchable list, a rule whose matchers were all invalid, a rejected interface name). apply now collects warnings from generate + netplane + model validation, classifies them (info|warning|critical, plus section+name for deep-linking) and exposes them on GET /api/status. Classification is by section, not by text markers: the wording of "list skipped" varies across six sites in generate and a marker list would silently miss the seventh.

blocklist source=url now consumes what public lists actually publish. It had accepted only compiled .srs, so every list D15 names (StevenBlack / OISD / AdGuard, all hosts-format) failed at apply with invalid sing-box rule-set file — a headline feature that broke on its own documented sources. The daemon now fetches (size-capped), parses hosts / plain-domain / AdBlock-comment formats, and compiles to .srs under /etc/shater/lists/, which the engine loads as a local rule-set. Compiling is what makes it affordable on the target hardware: measured on the testbed, StevenBlack's 2.4 MB of text became 80873 domains in a 491 KB .srs, costing ~0.5 MB of a rootfs with ~33 MB free — so it ships in the production posture rather than being traded away for the 8 KB geosite ads list.

D18 — Node health = a board with computed verdicts, NOT delete-on-failure + a dead-overlay

Decided 2026-07-24 (proxy-health plan §5.A). Upstream common/urltest recorded only success ({Time, Delay}) and a failed check deleted the entry — a dead member was indistinguishable from a never-measured one. On top of that, deaths found by shater's own probing went to a private engine.dead overlay that only the panel read: selection never saw them, and an old success in history counted as "alive forever". Net effect, reproduced in the field: YouTube dead on a real device while the panel showed the group healthy.

Decision: one health board, one source of truth. The history entry becomes {LastOK, Delay, LastFail}; a failure marks (MarkFailed), never deletes — deletion is reserved for removing a node from the config. The verdict is a method computed on read, not a stored field: alive when the success is fresher than the failure and younger than TTL; dead while a fresher failure is itself fresh; untested otherwise, with TTL = max(3 × global probe interval, 10 min). Everyone who learns of a death — the group's native checker, the observatory, a failed user dial — writes the same board; selection, balancer slots and the panel read the same board. The engine.dead overlay is deleted.

  • Rejected: keep delete-on-failure and widen the overlay to selection. Two stores of the same truth with undefined precedence; the overlay would need its own TTL/pruning and every reader would have to merge — exactly the divergence ("green panel, dead path") this decision exists to kill.
  • Rejected: persist the board across reboots. Measurements go stale faster than an overlay write is worth; after reboot everything is untested for seconds until the observatory's immediate first pass. In-memory, deliberately.
  • Deferred, not rejected: Xray-style sliding-window stats / hysteresis. The delay of the last success is enough to rank alive members in v1; the board keys and record shape leave room to widen without migration. Flapping is visible instead through the alive↔dead flip log (info) — the single diagnostic trail. Consequence: a stale success decays (alive → untested after TTL) instead of reading as alive forever, and a dead group blocks fail-closed — visible and alertable — rather than silently falling back.

D19 — Background probing = an observatory driven by rule reachability, NOT a population sweep

Decided 2026-07-24 (proxy-health plan §5.C, replaces shater/engine/sweep.go). The sweep probed the whole population round-robin — all nodes plus all per-group copies, used or not — yet chain copies were excluded entirely, so the one path users actually complained about ("rule → chain") was never measured. The manual "Test all nodes" run duplicated the same full-population walk on demand. On router budgets that is the wrong shape twice: work grows with the subscription size (~200 nodes), not with what the config uses.

Decision: probe only what the rules can reach. Following the Xray model (central observatory + balancers reading observations, §3 of the plan), a plan is built from the applied option.Options by reachability: enabled-rule targets (plus Final and DNS detours) → groups → members / egress copies; a chain is probed end-to-end by dialing its exit tag through the whole hop path (the Xray observatory-through-proxySettings equivalent); chain group-hop members are probed through their path prefix. Everything unreferenced stays untested and its group/chain is badged "unused" in the panel — so untested never looks like a health problem. A freshness gate skips tags an active group's own checker already measures; the cursor survives no-op reconciles (the sweep's release-blocker: cron reconciles must not restart the cycle); the manual probe-all is removed with the sweep.

  • Rejected: keep the sweep. Probes hundreds of unused nodes on a router budget and still misses chains; its "coverage" is what made per-node health look authoritative while the used path went unmeasured.
  • Rejected: probe every chain hop individually. Multiplied probe traffic for diagnostics that never affects selection — no choice depends on a middle hop's individual health. A dead middle hop makes the exit verdict honestly dead; localization is served by the verdict flip log.
  • Rejected: auto-reroute rules when a group dies. A dead group blocks fail-closed — visible and alertable. Silent rerouting would hide the outage and change routing semantics behind the user's back. Consequence: the probe budget is bounded by the config, not the subscription; cold start converges in seconds (immediate first pass after apply); wanting numbers for an unused group has one honest answer — reference it from a rule.

D20 — Probe URL / Interval are global-only; per-group overrides deleted

Decided 2026-07-24 (proxy-health plan §5.D). Groups carried optional ProbeURL/ProbeInterval overrides. That bred a documented ambiguity — one node shared by two groups with different URLs yields incomparable delays and needs "whose URL wins" dedup machinery (the (dial, URL) plan key) — and it breaks the D18 board: a verdict is a (record, TTL) pair with TTL derived from the probe interval, so per-group intervals would make the same record mean different things to different readers. Xray's observatory has exactly one global probeURL/probeInterval; that is the model we mapped onto (§3).

Decision: only Globals.ProbeURL/Globals.ProbeInterval. They drive the native group checker, the observatory and the manual exit test alike (failover keeps its 30 s default interval). Probe-plan dedup collapses to the dial path. Migration is the standard dead-option drainage: old UCI configs carrying probe_url/probe_interval on a group parse silently and the options vanish on the next render.

  • Rejected: keep per-group overrides. Incomparable measurements across groups, undefined semantics for shared members, and a per-group TTL that fractures the single-board verdict.
  • Deferred: per-tier probe cadence (rule-critical tags more often). If ever needed it is a field on the observatory's ProbeJob — a scheduling knob, not a return of per-group configuration. Consequence: all delay numbers are comparable (least ping ranks apples against apples), and group settings lose two footgun fields while Settings keeps the two that actually govern every check.

D21 — A rule's destination is a rule-set, and nothing else

Decided 2026-07-25 (product owner). config rule carried THREE ways to say where traffic is going: dst_domain (an inline domain list), dst_ip (an inline CIDR list) and dst_ruleset (a reference to a config ruleset). Three mechanisms meant three sets of semantics to learn and keep straight, and the inline ones were the worse half of the trade: they are re-parsed per rule instead of being compiled once into a .srs, they cannot be shared between rules, and their matcher vocabulary had drifted from the rule-set one in a way nobody could see (below).

Decision: dst_domain and dst_ip are removed (schema v2). dst_ruleset is the only destination matcher. Src, dst_port and proto are untouched — they are not lists of destinations and have no rule-set form.

  • Rejected: keep the inline lists as a shorthand. "One obvious way" is the whole point; a shorthand that quietly means something different from the long form (see the bare-entry trap) is worse than no shorthand.
  • Rejected: promote inline lists to rule-sets lazily at generate time. The config on disk would then not say what the router does, and the panel would have to render a list the user cannot find or edit.

The bare-entry trap, and how the migration handles it

The two contexts already disagreed about exactly one spelling, silently:

entry in a rule (dst_domain) in a rule-set (entry) migrated to
example.com exact host host + subdomains full:example.com
full:example.com exact host exact host unchanged
suffix:example.com / .example.com host + subdomains host + subdomains unchanged
keyword:ads substring substring unchanged
regexp:^ads\. pattern pattern (added here) unchanged
geosite:x / geoip:x inert (engine field removed) inert (unknown prefix) unchanged

shaterd migrate (schema v1→v2, shater/model/migrate.go) creates one inline config ruleset per rule that still carries a legacy list — rule-<rule name> for domains, rule-<rule name>-ip for addresses — moves the entries across with the conversion above, appends the new name to dst_ruleset, and deletes the old option. It is idempotent, it resumes an interrupted run, and it never overwrites a hand-written rule-set that already owns the generated name (it picks rule-<name>-2). regexp: support was added to inline rule-sets in the same change precisely so the move can be lossless.

geosite:/geoip: entries are copied VERBATIM rather than promoted to a source=geosite rule-set: those matchers have been inert since the engine dropped the route-rule geosite/geoip fields, and turning a dead matcher live during an upgrade would be a behaviour change, not a migration. The text is kept so the operator can see it and convert it deliberately.

One deliberate semantic change, called out: a rule that used BOTH lists matched them with AND (an engine route rule ANDs its matcher fields), which is almost never what "these sites and these networks" meant. The two generated rule-sets are ORed, because rule_set: [a, b] matches when either matches. Such a rule matches more after the migration than before; it affects only configs that used both fields at once.

That AND→OR change is about the ENGINE's TCP/UDP path, and it deliberately does not extend to the untunnelable-protocol plane (shater/apply/untunnelable.go, the ping / IPTV / VPN-passthrough policy in nftables). There, a v1 dst_domain + dst_ip rule could never claim a packet that carries no domain, so the plan skipped it; reading the migrated form as OR would have made the address half suddenly decisive, and with target=direct that means an upgrade quietly sending previously-tunnelled ICMP out with the client's real source address. A rule whose rule-sets are known to match by NAME is therefore still skipped by that plan, and the skip is reported ("a routing rule matches by name … as well as by address"). Split the rule in two if you want the addresses decided there.

The vocabulary is about ENTRIES YOU TYPE, not about every list body

The table above is the vocabulary of an inline rule-set's entry values (and of the DNS-filter/device lists, which share the classifier). The other two rule-set sources are not other spellings of it:

source what it is vocabulary
inline entries you type the table above
url → .srs / .json a compiled rule-set, engine-owned the engine's, not ours
url → anything else a hosts / one-domain-per-line / AdBlock TEXT FILE none — every line is a domain plus its subdomains
file a local .srs / .json the engine's, not ours

Rejected: run text lists through the entry classifier too. A published AdGuard/OISD list is full of colon-bearing tokens that are ordinary filter syntax (##…:has(…), $domain=, absolute URLs); classifying them would either mis-import them or bury the operator under hundreds of "unrecognised prefix" warnings per list. The formats also disagree structurally — a hosts line carries several names, so the text parser works per token, while an entry is a whole line. And regexp: arriving from a third-party URL is a pattern compiled into the router's matcher and evaluated per query, which is a very different proposition from one the operator typed.

So the difference stands and is paid for in diagnostics instead: a text list containing full: / suffix: / keyword: / regexp: is reported per list, on every generate, naming the entries and pointing at source=inline where they work (warnListEntryVocabulary, shater/generate/ruleset.go). The check tests only those four markers, never the general word: shape, so it fires on a human's mistake and stays quiet on published filter syntax.

Consequence: one destination mechanism, one vocabulary, one place a list is edited; every list is compiled once and reused. The panel's rule editor drops its Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of the rulesets that already exist, and nothing more. Creating and filling a list stays in the Rulesets panel — rejected: a "create a list from here" shortcut in the rule editor, because a second place to author a list is a second place for its semantics and its duplicate-name rules to drift, and the whole point of this decision was to stop having two.

D22 — One packaging lane: apk. The opkg/.ipk lane is deleted, not disabled

Decided 2026-07-25 (product owner). CI built and published TWO signed feeds from every run: opkg/usign (.ipk + Packages.gz, OpenWrt 24.10) and apk/EC (.apk + packages.adb, OpenWrt/ImmortalWrt 25.12). The opkg half served nobody. Checked on the actual hardware, not inferred:

Device Firmware pkg arch package manager
mini_router (BPi-R3 Mini) ImmortalWrt 25.12.1 aarch64_cortex-a53 apk-tools 3.0.5
main_router (BPi-R4) OpenWrt 25.12.0 aarch64_cortex-a53 apk-tools 3.0.5 — no opkg binary on the system at all

Decision: delete the opkg lane outright. Removed: the build + release jobs from .gitea/workflows/release.yml; ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh; and the trust anchor dist/shater-feed.pub. The Gitea secret KEY_BUILD is now referenced by nothing and can be deleted from the repo settings. ci/version.sh, ci/gitea-release.sh and ci/fetch-sdk.sh are shared or apk-only and stay.

  • Rejected: keep the lane but stop triggering it (comment it out / gate it on a dispatch input). Dead code in CI is worse than no code: it keeps a second SDK matrix, a second signing key and a second feed layout alive in everyone's head and in every future edit, and it silently rots because nothing runs it. The 24.10 SDK images it pins are themselves a frozen dependency.
  • Rejected: keep dist/shater-feed.pub as a historical artifact. A committed trust anchor is an instruction — it invites someone to follow the old install path for a feed that is no longer produced. Nothing is lost by removing it: git history still holds the file, the SECRET half is untouched in KEY_BUILD, and a usign secret key blob contains its own public half, so the identity can be reconstructed if a 24.10 device ever has to be served again. Deleting the file is reversible; a stale trust anchor pointing at an unmaintained feed is the thing that quietly misleads.
  • Not done: revoking or rotating the usign key. There is no incident. It is retired, not burned (D7).

Consequence: one SDK, one key, one feed layout, one set of install instructions. It also makes the rolling release apk-latest-<arch> the only install path that does not require hand-editing a file per release — which is why the same change fixed it: publishing was an either/or (apk-latest-<arch> on dispatch, ELSE apk-vX.Y.Z-<arch> on a tag), so once releases moved to tag pushes the rolling pointer stopped being written and froze at 0.2.0 while v0.2.9/v0.2.10 shipped — routers on the rolling URL got a successful, silent apk update with nothing new. release-apk now writes the rolling pointer on every run and asserts, by reading the published release back over the Gitea API, that it holds our three tag-versioned packages at exactly the version just built and no asset at any other version.

D23 — The router tag set is a checked contract, not a string literal

with_gvisor was trimmed from the router set on 2026-07-23 (D9) as "unreachable code: we never emit a tun inbound". True about tun — and irrelevant, because gVisor is also the netstack of the WireGuard endpoint, which shater emits and FEATURES.md declares [MVP] (AmneziaWG is called "a driving requirement"). Every binary shipped between then and 2026-07-25 answered a configured WireGuard node with:

create instance: initialize endpoint[0]: create WireGuard device:
gVisor is not included in this build, rebuild with -tags with_gvisor

transport/wireguard/device_stack_stub.go (//go:build !with_gvisor) returns tun.ErrGVisorNotIncluded from both device constructors, so system_interface: true is not an escape hatch either: WireGuard was 100% dead in the shipped artifact while the panel offered it, the parser accepted wg://, awg:// and wg-quick .conf imports, and the owner had 7 WireGuard sections in UCI on a production router.

  • Decision: with_gvisor is part of the router tag set and stays there for as long as we ship WireGuard. It costs ~2.8 MB raw / ~0.65 MB UPX per arch (measured 2026-07-25, both arches; /overlay on the production router is 6.9 GB with 205 MB used). A tag whose absence turns a declared feature into a runtime error is not "dead weight" — it is the feature.

Why the bug was invisible, and what now makes it visible

The defect was not a typo in a tag list. It was that nothing connected the tag list to the feature list, and the shipped tag combination was the one build configuration nothing exercised: the whole test suite compiles with the FULL upstream set (with_gvisor included), so TestAmneziaWGEndpoint passed happily while the artifact it was supposed to vouch for could not create a WireGuard device. Tests proved the code was right; they never proved the build was.

Three pieces now hold it together:

  1. One definition of the set — scripts/router-tags.sh (SHATER_ROUTER_TAGS
    • SHATER_ROUTER_LDFLAGS), sourced by scripts/build-shaterd.sh and by the checker. The tag list used to live as a literal inside the build script, i.e. in a file no test reads. A second copy is a second truth.
  2. A declared-feature table — shater/buildtags: every tag-gated capability we promise, with the exact tags it needs to run and why (the code anchor). TestRouterTagSetCoversDeclaredFeatures parses the shell file and fails if a declared feature lost a tag. It needs no build tags, no Linux, no network and no privileges, so it runs in every plain go test ./... — including on the Windows dev host, where nothing else can see the shipped configuration.
  3. A construction test under the shipped tags — shater/generate.TestShippedTagSetConstructsDeclaredProtocols drives one node of every declared protocol (ss/vmess/trojan/vless ws-grpc-httpupgrade-quic- xhttp/REALITY/uTLS-fp/hysteria2/tuic/wg/awg) through box.New+Start. scripts/check-router-tags.sh runs it with SHATER_ROUTER_TAGS, and CI runs that script (.gitea/workflows/release.yml) before the artifact is built. In a router-tag-set run nothing may be skipped: a protocol that is not compiled in fails the run instead of quietly disappearing from it.

(2) catches a trim the moment it is made and names the feature it kills; (3) catches what a list comparison cannot — a tag that is present but insufficient. Neither is a substitute for the other. A new protocol in shater/parse + shater/generate means a new row in buildtags.Features and a new probe case; TestEveryTagGatedFeatureIsProbed fails until both exist.

  • Rejected: "just add the tag". The one-line fix restores WireGuard and leaves the mechanism that hid it fully intact — the next size-driven trim is equally invisible. The tag is the smallest part of this decision.
  • Rejected: run the WHOLE test suite with the router tag set in CI. It is the obvious move and it does not work: parts of the suite legitimately depend on upstream-only tags, and the run costs a second full compile of a 25 MB binary's worth of packages on every release. A focused, unprivileged construction test buys the same evidence for ~10 s and, unlike a full run, can be required to skip nothing.
  • Rejected: assert the tag set against upstream's DEFAULT_BUILD_TAGS. That makes any trim a failure, which turns the check into noise and re-litigates D9 on every upstream rebase. The contract is with our own feature list, not with upstream's.
  • Not done: dropping with_lx_command. It is inert for shaterd — nothing under shater/ imports sing-box/daemon or experimental/libbox, and go list -deps ./shater/cmd/shaterd links neither, so it costs zero bytes. It stays only so the router set remains a subset of the lx desktop set. Noted because "a tag that buys nothing" is the mirror image of this bug and should be removed deliberately, not silently.

D24 — DNS interception is the DEFAULT (dns_intercept=1), not an opt-in

Decided 2026-07-26. Globals.DNSIntercept shipped as opt-in (default false, and absent from both DefaultGlobals and the shipped /etc/config/shater). The result was an inverted posture, which is the reason this is a decision and not a preference:

  • a client with standard settings — DNS = the router's address, exactly what DHCP hands out — sent its queries to the router. The nft :53 divert was behind the flag (netplane/nft.go), and the rule right after it is an unconditional fib daddr type local accept, so the query was delivered locally to dnsmasq and forwarded to the ISP in the clear: no blocklists, no per-device DNS rules, no Block-DoH, no resolver detour, nothing;
  • a client that hard-coded 8.8.8.8 "to bypass the router" was addressing a non-local IP and was caught by the ordinary tproxy catch-all.

The obedient client leaked; the evader did not. Meanwhile FEATURES.md, README.md and D14 all promised "no DNS leaks" and "dnsmasq never sees LAN queries" — true only for the traffic pattern the default did not cover. dns_intercept appeared nowhere in docs-shater/ at all.

Decision: DNSIntercept is seeded ON in model.DefaultGlobals, and the shipped /etc/config/shater carries an explicit option dns_intercept '1'. Nothing about the interception MECHANISM changed — only which side of the switch is the default.

.lan and the private PTR zones keep working, and that is a pre-existing part of the mechanism, not something bolted on for this flip. generate/dns.go adds a synthetic DNS server (shater-local-dns, plain UDP to 127.0.0.1:53, detour direct, so the daemon's own loop-mark keeps it out of the divert) and PREPENDS a domain_suffix rule for lan + the RFC6303 private reverse zones, ahead of every device/filter rule. Two honest limitations: it hardcodes lan (a router whose dnsmasq domain was changed needs a config dns_rule for the new suffix), and it only exists when the model has at least one config resolver — with none, buildDNS emits no DNS plane at all and the engine falls back to its built-in local transport, which reads /etc/resolv.conf (127.0.0.1 → dnsmasq), so local names still resolve but nothing is filtered.

A dead engine does NOT black out the LAN's DNS. This was the first thing checked, because "intercept everything" invites the reading "engine down = no DNS anywhere", and that is not what happens:

  • the fail-closed holding plane (D17, RenderHoldNft) hooks forward ONLY. A query addressed to the router is INPUT-hook traffic, so dnsmasq answers it as it always did — unfiltered and plaintext to the ISP. Deliberate: blocking it would also cut the daemon's own name resolution and with it any chance of self-recovery;
  • with the FULL plane loaded and the engine's tproxy socket gone, the tproxy statement returns NFT_BREAK, which aborts its own rule; the packet continues down the chain into the same fib daddr type local accept and reaches dnsmasq.

So the failure mode is a DNS fail-open (working, unfiltered) while client TRAFFIC stays fail-closed — and a query aimed at an EXTERNAL resolver is dropped with the rest of the forwarded traffic. Operators must know this: "the tunnel is down" does not mean "DNS is private".

Existing installs. /etc/config/shater is a conffile (openwrt/shater-core/Makefile), so an upgrade never replaces it:

  • a config that never mentioned the option (all of them, before this change) now parses over the ON seed and starts intercepting on the next apply. That is the intended behaviour change, and the only one this decision makes;
  • an explicit option dns_intercept '0' keeps winning. It survives the WriteUCI→ReadUCI round-trip because render.go emits booleans ALWAYS — the trap a default-true bool has and a default-false one does not: a value omitted at false would come back as the seed and silently re-enable itself. shater/model/dnsintercept_test.go pins both directions, plus the shipped file.

Not done: silencing the "no resolvers configured" warning by shipping a resolver. With interception on and no config resolver, generate warns — and it is right to: every client query now lands in an engine that has no resolver plane, so it is answered by the system resolver (dnsmasq → the ISP, in the clear) with filtering and anti-leak inert. Shipping a type local resolver would make the warning disappear while changing nothing about where the queries go: the panel would show a configured resolver and the operator would believe DNS was handled. That is the inverted lie this project keeps deleting. The warning stays; what it needs is the accurate wording (it currently claims .lan breaks, which the fallback above disproves), not a workaround. Note also that a fresh install ships INERT (enabled '0') and Reconcile tears down instead of generating, so the warning cannot appear before the operator has enabled the stack — at which point it describes their live config.

OPEN, and it gates shipping this default: the synthetic local server changes how proxy-endpoint DOMAINS are resolved. Found while landing D24, reproduced on Linux with one resolver and a node addressed by a hostname:

  • common/dialer/dialer.go resolves a domain server address through route.default_domain_resolver; when that is unset it uses dnsTransport.Default() — the engine's built-in local transport, i.e. a bootstrap-DIRECT lookup — but only while fewer than two DNS transports exist. With two or more and no default, it reports the missing-domain-resolver deprecation and leaves the query transport nil, so dns.Router.Lookup falls back to lookupWithRules: the CLIENT DNS plane.
  • dns_intercept adds shater-local-dns, which takes a single-resolver config from one transport to two. So a config whose only resolver is DoH-through-the-tunnel — the recommended anti-leak setup — would start resolving its own node's hostname through that same tunnel: a bootstrap loop where there was none.
  • Evidence: the same model emits no deprecation notice with dns_intercept=0 and two missing-domain-resolver notices with dns_intercept=1; generate.TestDNSFilterRemoteBlocklistHTTPClient (Linux-only) fails on exactly that notice and is deliberately left failing rather than relaxed.

The fix belongs in generate (route.go:160 already sets route.default_domain_resolver from endpointResolver(), which is opt-in and unset by default): when buildDNS emits the synthetic local server and no endpoint resolver is configured, default_domain_resolver must be pointed at a bootstrap-direct server, which restores exactly the pre-D24 behaviour and clears the notice. Until that lands, an operator can get the same result by setting endpoint_resolver to a direct resolver. Note the hazard is not created by D24 — any config with two resolvers has it today; the default merely makes it universal.