Files
shater/docs-shater/DECISIONS.md
T
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00

32 KiB
Raw Permalink Blame History

Decisions (ADR log)

Key architectural decisions and why, so nobody re-litigates them (we spent a whole discussion converging — see also CONTEXT.md).

D1 — Do NOT write a proxy engine from scratch

The proxy protocols (VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, AmneziaWG, Hysteria2/TUIC, transports, TLS fingerprinting) are a byte-precise, adversarial anti-DPI arms race maintained by large communities and changing monthly. A from-scratch engine would be slower, buggier, less secure, and more detectable — the opposite of "optimized." Reuse a real engine.

D2 — Engine = sing-box (via sing-box-lx), not xray-core

  • AmneziaWG 2.0 (I1–I5 CPS decoy packets) is a hard requirement; sing-box has first-class AmneziaWG, xray's is weaker. sing-box also has broader protocol coverage (Hysteria2, TUIC, ShadowTLS, MASQUE) and is library-first (libbox), which suits embedding.
  • sing-box-lx (github.com/Leadaxe/sing-box-lx) is a thin, actively-rebased downstream fork adding AmneziaWG 2.0, XHTTP, MASQUE/WARP and — usefully for us — gRPC observability of DNS queries / rules / outbounds, which feeds our stats.
  • License GPL-3.0, compatible with (and cleaner than) our prior GPL-2.0-or-later. See D6.
  • Cost accepted: our whole control plane / generator / share-link handling is re-based from xray-shaped (v0.1) to sing-box-shaped.

D3 — FORK sing-box-lx and embed the whole product inside it (not "depend as a library")

Considered: (a) consume as a pinned Go module + Gitea mirror; (b) fork. Chosen: (b) fork, so we can embed control-plane + admin panel + DNS filter directly and integrate tightly with engine internals (DNS, routing, stats). This was a deliberate call favouring maximum integration over minimum maintenance.

To keep the fork from rotting, a mandatory discipline:

  • Our code lives in new top-level dirs (shater/, panel/, openwrt/) that never collide with upstream files on rebase.
  • Edits to upstream files are minimal and marked // shater.
  • We rebase/merge onto sing-box-lx tags on a cadence (mirroring how sing-box-lx rebases onto sing-box). Fork = additive overlay tracking upstream, not a divergent rewrite.
  • main of the shater repo is the fork; v0.1 branch keeps the old project.

D4 — UI = thin LuCI launcher + a separate embedded admin panel

LuCI stays a minimal, pretty mini-dashboard with an "Open panel" button. The button mints a short-lived token in the authenticated LuCI/ubus session and redirects to our admin panel on its own port, served by the daemon; the panel validates the token and opens a session.

  • Why not do everything in LuCI: LuCI's form/view model is limiting for the rich stats/config UX we want.
  • Why not a standalone panel with its own login: a router-facing panel with its own auth is a serious security surface to get right. Bootstrapping the token from LuCI's existing, hardened auth avoids a second login system.
  • The panel is a modern SPA, embedded in and served by the forked binary.

D5 — Long blocklists: don't push megalists into dnsmasq; use an efficient matcher

Naive address=/domain/# in dnsmasq holds every domain in RAM and reloads slowly (100k–1M+ entry lists are common). Instead the engine's own domain matcher (sing-box already compiles geosite-scale lists) or our compact matcher (suffix hash-set + optional bloom prefilter, compiled/cached blob, dedup + subdomain-collapse, streaming parse, refresh by content-hash) does the filtering. Blocklist sources are flexible: inline / file / url / geosite (geosite only when geodata is present, else that source is inert / fail-open).

D6 — License = GPL-3.0

sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files may stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is GPL-3.0 for clarity.

D7 — Keep the v0.1 feed signing identity

The usign feed key 5ac4b177689cb8e0 (public key in dist/shater-feed.pub, secret in Gitea secret KEY_BUILD) carries over, so routers that already trust it keep verifying v0.2 packages. Do not regenerate it without a documented rotation.

D8 — Preserve, don't destroy: v0.1 lives on its branch

The reset moved the full working xray-based project to the v0.1 branch and cleaned main. Nothing is lost; reusable logic (reliability layer, nft/routing, sub-fetch, CI/signing, design system) is ported forward, not rewritten.

D9 — Router build uses a musl-static tag set (drop with_naive_outbound,with_purego)

Verified on the x86_64 OpenWrt VM (2026-07-14): the canonical Makefile.lx LX_TAGS produces a binary dynamically linked to glibc and it will NOT run on musl OpenWrt. Root cause (proven three ways — ldd, ELF PT_INTERP, static hello-world control): with_naive_outbound pulls in cronet-go, which needs with_purego; purego uses //go:cgo_import_dynamic to dlopen libcronet at runtime, forcing an ELF with PT_INTERP=/lib64/ld-linux-x86-64.so.2 + PT_DYNAMIC even with CGO_ENABLED=0. musl has no glibc loader → not found (rc 127).

  • Decision: the OpenWrt/router binary uses a router-specific tag set = LX_TAGS minus with_naive_outbound,with_purego. That variant is fully static (ET_EXEC, no PT_INTERP) and runs directly on musl. naiveproxy outbound is not in our feature set (FEATURES.md), so nothing we ship is lost. with_awg is independent of naive/purego and stays.
  • Keep badlinkname,tfogo_checklinkname0 + -checklinkname=0 (needed by badtls); they build fine static.
  • The desktop/CLI LX_TAGS stays as upstream for any non-router use. Do not reuse the desktop tag set for the router package — define SHATER_ROUTER_TAGS.
  • 2026-07-23: with_gvisor also dropped from the router set. The shater data plane is tproxy/redirect (netplane); generate never emits a tun inbound, so the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever appears it falls back to the system stack — re-add the tag then.
  • 2026-07-23: with_clash_api also dropped. The admin panel is shater's own web server and generate never emits a clash_api service; the desktop/CLI LX_TAGS keeps the tag for external dashboards.
  • 2026-07-23: with_dhcp also dropped. shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS transport is never generated, and the slim shater/registry never registers it.

D11 — The engine runs IN-PROCESS (box.New), not as a subprocess

v0.1 generated an xray JSON file and ran xray as a separate process. v0.2 embeds: the daemon builds option.Options in Go and drives the engine via box.New(box.Options{Context, Options}) + Start()/Close() in the same process (proven in Phase 1, shater/cmd/shater-proto). This is the whole point of the fork — it gives us in-process access to DNS query events, routing, and stats (D5/observability) without log-scraping, and one binary to ship.

  • Apply = atomic instance swap. sing-box box instances are built-then-started; a config change rebuilds option.Options, validates by constructing a fresh box.New (which runs every adapter's constructor), and on success swaps: start the new instance, then close the old (or close→start within a short window under the kill-switch). Commit-confirm rollback keeps the last-good option.Options.
  • The generator (v0.1 generate.go) is rewritten to emit sing-box option structs instead of xray JSON. The parsers (share-link/subscription/awg-conf) and the nft/policy-routing layer port with little change (engine-agnostic).

D12 — v0.2 binary = shaterd, our own entrypoint under shater/

The router ships a single binary shaterd = shater/cmd/shaterd: our main that imports the engine as a library (D11) and also runs the control-plane, the DNS filter, the stats aggregator, and the admin-panel web server. cmd/sing-box (upstream) stays untouched for CLI/debug use. The router build uses the musl tag set (D9); shaterd links the same engine packages cmd/sing-box does.

D10 — Ship the router binary UPX-lzma compressed

Raw static binary is ~40–43 MB; the stock OpenWrt ext4 rootfs is ~100 MB. UPX --lzma --best takes it to ~9–11 MB (measured), a comfortable flash budget with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep an uncompressed artifact for debugging.

D13 — External DPI-bypass tool = ByeDPI (a SOCKS egress), NOT zapret

Decided 2026-07-14. We evaluated exactly two external desync tools — zapret (nfqws/tpws, NFQUEUE packet plane) vs ByeDPI/ciadpi (a local SOCKS5 desync proxy) — and picked one: ByeDPI. DPI-bypass stays a per-ruleset egress choice (a routing rule sends selected domains to it, exactly like a WG node or direct), never a global switch.

Why ByeDPI over zapret — decided by shater's architecture, not by raw method strength. In shater all LAN traffic is already TPROXY-intercepted into the in-process engine, and routing rules pick the egress:

  • ByeDPI is an egress. It's a SOCKS5 proxy on 127.0.0.1:<port>; a rule → socks outbound → ByeDPI desyncs (split/disorder/fake/oob/autottl, fake-TTL) → goes direct to the target. No tunnel, no exit-node bandwidth, low latency. Zero new nft rules; it never touches our packet plane. Ships as a ~100 KB musl-static binary behind one procd service.
  • zapret is a competing packet plane. nfqws hooks packets via its OWN nft/NFQUEUE rules, which would fight our verified fail-closed inet shater TPROXY table for the same packets/marks — a second packet plane on a fail-closed router multiplies kill-switch/leak surface, and it cannot be expressed as a sing-box egress. It also entangles netplane (a core verified subsystem), violating the additive-overlay discipline (D3).
  • zapret's real edge is QUIC / bad-checksum fakes at the raw-packet level. ByeDPI is TCP/TLS-centric and weak on QUIC. We accept that gap and close it by routing, not by tool: a rule dropping udp/443 for desync-domains forces them to fall back to TCP+TLS, which ByeDPI handles. That stays inside our rule model — no parallel packet plane.

Native fragment is free and complementary (not a second tool). The fork already compiles in, at the route-action level, tls_fragment + tls_fragment_fallback_delay, tls_record_fragment, tls_spoof/tls_spoof_method, and dialer udp_fragment. These are just engine flags on a direct egress and cover the light "just fragment the ClientHello" case with no external binary. ByeDPI is what we add for the stronger methods the engine lacks (fake/disorder/oob/ autottl).

Decision / implementation shape:

  1. Egress gains an optional dpi preset. Values map to escalating power: off → plain; fragment/record/spoof → native route-action flags on a direct outbound (in current Wave, ~free); byedpi → a socks egress pointed at a supervised local ByeDPI (ciadpi) instance (later phase: its own openwrt/ procd package + musl-static cross-build).
  2. shater/generate maps the native presets now; the byedpi egress kind is reserved in model and wired when its package lands.
  3. zapret is explicitly rejected for shater and should not be revisited unless in-place QUIC packet-desync becomes a hard requirement that routing cannot cover.

When zapret would have won: a pure DPI-bypass box with no tunnel engine (then nfqws's packet power is the whole product). shater already has a full tunnel engine

  • native fragment, so it wants the desync tool that composes with routing — ByeDPI.

D14 — LAN DNS anti-leak = sing-box hijack-dns route action, NOT a :53 listener

Decided 2026-07-15 after the Phase-2 gate exposed the gap (see memory shater-dns-no-engine-listener). The netplane dnsnat chain redirected LAN :53 to the router's :53, assuming the engine answers there — but generate/engine create NO :53 DNS server (the engine only binds the tproxy port). On the VM the redirect landed on stock dnsmasq, which resolved via its WAN upstream outside the tunnel = a DNS leak, and the engine's own resolvers (DoH-via-WARP, already built by generate.buildDNS with anti-leak detours) went unused. On a real box (dnsmasq removed) client DNS would break entirely.

Decision: use the sing-box-idiomatic hijack, not a listener. sing-box has no simple ":53 DNS inbound"; its DNS module is designed to be fed via a hijack-dns route action (C.RuleActionTypeHijackDNS). The tproxy inbound already sniffs (the leading sniffRule() in generate/route.go), so:

  • generate: after the sniff rule, add a route rule protocol: dns → action hijack-dns, so every sniffed DNS query is answered by the engine's internal resolver (which routes each query through its configured detour — the anti-leak).
  • netplane: STOP special-casing :53. Remove the prerouting :53 accept and the whole dnsnat chain so LAN :53 is DIVERTED by the normal tproxy catch-all into the engine, where hijack-dns catches it. Keep the :853 DoT reject (forces clients off DoT onto :53, which is now hijacked). DoH (:443) stays SNI-routed through the proxy.
  • dnsmasq: the tproxy divert steals LAN :53 before local delivery, so dnsmasq never sees LAN queries; it keeps serving ROUTER-local :53 only. No dnsmasq removal needed for correctness, though shipping images may still drop it.

Rejected alternative (a real engine :53 listener + keep dnsnat): non-idiomatic for sing-box, needs a bespoke DNS server, and duplicates what hijack-dns gives for free. generate + netplane MUST stay in agreement on this — they are the two halves of the same DNS plane.

D15 — DNS filter = sing-box rule-sets + reject DNS rules, NOT a custom matcher

Decided 2026-07-15 (Phase 4). D5 said "don't push megalists into dnsmasq; reuse the engine's matcher OR a compact custom matcher." sing-box already ships a compiled, memory-mapped rule-set matcher (geosite-scale, the .srs binary format + inline / local / remote rule-sets with auto-fetch+cache), and DNS rules support action: reject. So the DNS filter is built ON that, not a bespoke bloom matcher:

  • model: a blocklist section (name, enabled, source inline|file|url|geosite, url/path/entries, response nxdomain|zero|refuse) + an allowlist (overrides). A global dns_filter enable.
  • generate/dns.go: materialize each blocklist into a sing-box rule-set (inline rule for small inline lists; local rule-set for a compiled/file .srs; remote rule-set for a url, letting sing-box fetch+cache with a download detour through the proxy) and emit DNS rules: allowlist-accept FIRST (higher priority), then blocklist → reject (method default=NXDOMAIN, or a 0.0.0.0/:: predefined answer for zero). geosite sources are inert/fail-open when geodata is absent (D5).
  • daemon: a blocklist update verb (like sub update) to (re)compile file sources and warm remote caches; sing-box remote rule-sets self-refresh, so the daemon mostly compiles local sources + seeds well-known lists (StevenBlack / OISD / AdGuard) as disabled presets. Rationale: reuses a battle-tested, RAM-efficient matcher (no custom megalist code to get wrong), and it composes with the hijack-dns plane already in place (D14) — the filter rules sit on the same in-engine DNS path. Only build a custom matcher if a concrete sing-box rule-set limitation is hit (none known).

D16 — Enable sing-box cache_file; engine apply-swap must go close-first when it's shared

Decided 2026-07-15 (Phase 4). Remote rule-sets (url blocklists) and urltest/rdrc state need sing-box's experimental.cache_file to persist across restarts, auto-update on their interval, and keep RAM sane for megalists. So generate now emits cache_file (enabled, /etc/shater/cache.db, falling back to /tmp/shater-cache.db).

Consequence discovered on the VM: bbolt opens the cache DB under an EXCLUSIVE file lock and retries ~10s before failing timeout. The engine's default apply-swap starts the NEW box before closing the old (D11) — but the new box can't acquire the cache lock the old still holds, so EVERY live reconcile/apply stalled ~10s then failed (allowlist toggle, blocklist update, any edit silently didn't take). Fix (extends the D-fix for EADDRINUSE): engine.Apply now (a) treats the cache-file lock timeout as a swap conflict, and (b) proactively takes the close-old-then-start-new path when the incoming config shares an enabled cache_file with the running one (sharesCacheFileLock), avoiding the ~10s stall entirely. The fail-closed nft kill-switch covers the brief gap. Regression test TestEngineCacheFileSwap. See shater-engine-applyswap-portconflict (same close-first machinery, now also for the cache lock).

D17 — Audit pass: fail-closed must survive engine-start failure; several knobs were fiction

Decided 2026-07-20. A full-stack audit (8 parallel agents, every finding reproduced on the QEMU testbed) changed a few architectural positions. Recording the ones that constrain future work.

The data plane must never be absent while kill_switch=closed. Reproduced live: a remote rule-set that could not be fetched aborted box.Start, so apply failed and no nft table was installed at all. The router silently degraded to plain routing — no tunnel, no filtering, no kill-switch — while the panel looked healthy. Fail-closed had been implemented as "the engine died after starting"; it now also covers "the engine never started", via a holding plane (RenderHoldNft): forward blocked for the diverted interfaces, LAN-to-LAN + link-local + ND preserved, and the daemon's own egress marks accepted so it can still reach the network to recover. Management (SSH/LuCI/panel) is safe by construction — the holding chain hooks forward only, never input. Status now reports plane: full|hold|none. Rejected: "keep the previous table". After a cold boot there IS no previous table, which is precisely the reproduced scenario.

Engine start must not depend on the network. RemoteRuleSet.StartContext returns an error when its first fetch fails, and router.Start fast-fails the group — so a router that boots before its ISP link comes up cannot start the engine. Remote lists are now reachability-preflighted at generate time and omitted (loudly) rather than handed to the engine. Accepted cost, documented in ruleset.go: a brief source outage disables the list for up to one reconcile + probe TTL even when a good cached copy exists, because the cache cannot be consulted (the running engine holds an exclusive bbolt flock — see D16 — and a lock a process already holds on one fd denies its own second open). The structurally better fix (seed an empty cache entry between closing the old box and starting the new one, so StartContext never fetches at start) is deliberately deferred, not overlooked.

BlockDoH does not need to exclude upstream resolvers on the route plane. The original design punched a hole for every configured resolver IP so the engine could reach it — which also handed every LAN client an un-blocked public DoH endpoint, i.e. the more reputable the upstream, the wider the bypass. This was unnecessary: engine-originated DNS dials go through DetourDialer straight to the outbound and never traverse route rules, and they additionally carry LoopMark 0xff, which nft accepts before the tproxy divert. The exclusion now applies to the DNS plane only (a hostname upstream must still resolve), and the :443 rejects use the full list.

Globals.DNSMode is unimplementable and its control was removed. nftset named the v0.1 xray architecture, which does not exist here; fakeip is a resolver type in sing-box 1.14 (the top-level dns.fakeip block was removed upstream), already expressible as config resolver with type=fakeip + pool, which is strictly more capable (per-domain via a dns_rule). Wiring the toggle would have created a second configuration path with undefined precedence. The field stays in the model for config compatibility; the panel now documents the resolver route instead.

TPROXY cannot carry ICMP/IGMP/ESP/GRE, so this is now an explicit policy. Globals.Untunnelable = block (default, unchanged behaviour) | icmp | direct. Three values, not two, because the leaks differ in kind: an ICMP echo is ephemeral, user-initiated and reveals the address only to a host the user deliberately contacted, whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".

Fail-open degradations must be visible in the panel, not only in logread. The audit deliberately converted many aborts into warn-and-continue (an unfetchable list, a rule whose matchers were all invalid, a rejected interface name). apply now collects warnings from generate + netplane + model validation, classifies them (info|warning|critical, plus section+name for deep-linking) and exposes them on GET /api/status. Classification is by section, not by text markers: the wording of "list skipped" varies across six sites in generate and a marker list would silently miss the seventh.

blocklist source=url now consumes what public lists actually publish. It had accepted only compiled .srs, so every list D15 names (StevenBlack / OISD / AdGuard, all hosts-format) failed at apply with invalid sing-box rule-set file — a headline feature that broke on its own documented sources. The daemon now fetches (size-capped), parses hosts / plain-domain / AdBlock-comment formats, and compiles to .srs under /etc/shater/lists/, which the engine loads as a local rule-set. Compiling is what makes it affordable on the target hardware: measured on the testbed, StevenBlack's 2.4 MB of text became 80873 domains in a 491 KB .srs, costing ~0.5 MB of a rootfs with ~33 MB free — so it ships in the production posture rather than being traded away for the 8 KB geosite ads list.

D18 — Node health = a board with computed verdicts, NOT delete-on-failure + a dead-overlay

Decided 2026-07-24 (proxy-health plan §5.A). Upstream common/urltest recorded only success ({Time, Delay}) and a failed check deleted the entry — a dead member was indistinguishable from a never-measured one. On top of that, deaths found by shater's own probing went to a private engine.dead overlay that only the panel read: selection never saw them, and an old success in history counted as "alive forever". Net effect, reproduced in the field: YouTube dead on a real device while the panel showed the group healthy.

Decision: one health board, one source of truth. The history entry becomes {LastOK, Delay, LastFail}; a failure marks (MarkFailed), never deletes — deletion is reserved for removing a node from the config. The verdict is a method computed on read, not a stored field: alive when the success is fresher than the failure and younger than TTL; dead while a fresher failure is itself fresh; untested otherwise, with TTL = max(3 × global probe interval, 10 min). Everyone who learns of a death — the group's native checker, the observatory, a failed user dial — writes the same board; selection, balancer slots and the panel read the same board. The engine.dead overlay is deleted.

  • Rejected: keep delete-on-failure and widen the overlay to selection. Two stores of the same truth with undefined precedence; the overlay would need its own TTL/pruning and every reader would have to merge — exactly the divergence ("green panel, dead path") this decision exists to kill.
  • Rejected: persist the board across reboots. Measurements go stale faster than an overlay write is worth; after reboot everything is untested for seconds until the observatory's immediate first pass. In-memory, deliberately.
  • Deferred, not rejected: Xray-style sliding-window stats / hysteresis. The delay of the last success is enough to rank alive members in v1; the board keys and record shape leave room to widen without migration. Flapping is visible instead through the alive↔dead flip log (info) — the single diagnostic trail. Consequence: a stale success decays (alive → untested after TTL) instead of reading as alive forever, and a dead group blocks fail-closed — visible and alertable — rather than silently falling back.

D19 — Background probing = an observatory driven by rule reachability, NOT a population sweep

Decided 2026-07-24 (proxy-health plan §5.C, replaces shater/engine/sweep.go). The sweep probed the whole population round-robin — all nodes plus all per-group copies, used or not — yet chain copies were excluded entirely, so the one path users actually complained about ("rule → chain") was never measured. The manual "Test all nodes" run duplicated the same full-population walk on demand. On router budgets that is the wrong shape twice: work grows with the subscription size (~200 nodes), not with what the config uses.

Decision: probe only what the rules can reach. Following the Xray model (central observatory + balancers reading observations, §3 of the plan), a plan is built from the applied option.Options by reachability: enabled-rule targets (plus Final and DNS detours) → groups → members / egress copies; a chain is probed end-to-end by dialing its exit tag through the whole hop path (the Xray observatory-through-proxySettings equivalent); chain group-hop members are probed through their path prefix. Everything unreferenced stays untested and its group/chain is badged "unused" in the panel — so untested never looks like a health problem. A freshness gate skips tags an active group's own checker already measures; the cursor survives no-op reconciles (the sweep's release-blocker: cron reconciles must not restart the cycle); the manual probe-all is removed with the sweep.

  • Rejected: keep the sweep. Probes hundreds of unused nodes on a router budget and still misses chains; its "coverage" is what made per-node health look authoritative while the used path went unmeasured.
  • Rejected: probe every chain hop individually. Multiplied probe traffic for diagnostics that never affects selection — no choice depends on a middle hop's individual health. A dead middle hop makes the exit verdict honestly dead; localization is served by the verdict flip log.
  • Rejected: auto-reroute rules when a group dies. A dead group blocks fail-closed — visible and alertable. Silent rerouting would hide the outage and change routing semantics behind the user's back. Consequence: the probe budget is bounded by the config, not the subscription; cold start converges in seconds (immediate first pass after apply); wanting numbers for an unused group has one honest answer — reference it from a rule.

D20 — Probe URL / Interval are global-only; per-group overrides deleted

Decided 2026-07-24 (proxy-health plan §5.D). Groups carried optional ProbeURL/ProbeInterval overrides. That bred a documented ambiguity — one node shared by two groups with different URLs yields incomparable delays and needs "whose URL wins" dedup machinery (the (dial, URL) plan key) — and it breaks the D18 board: a verdict is a (record, TTL) pair with TTL derived from the probe interval, so per-group intervals would make the same record mean different things to different readers. Xray's observatory has exactly one global probeURL/probeInterval; that is the model we mapped onto (§3).

Decision: only Globals.ProbeURL/Globals.ProbeInterval. They drive the native group checker, the observatory and the manual exit test alike (failover keeps its 30 s default interval). Probe-plan dedup collapses to the dial path. Migration is the standard dead-option drainage: old UCI configs carrying probe_url/probe_interval on a group parse silently and the options vanish on the next render.

  • Rejected: keep per-group overrides. Incomparable measurements across groups, undefined semantics for shared members, and a per-group TTL that fractures the single-board verdict.
  • Deferred: per-tier probe cadence (rule-critical tags more often). If ever needed it is a field on the observatory's ProbeJob — a scheduling knob, not a return of per-group configuration. Consequence: all delay numbers are comparable (least ping ranks apples against apples), and group settings lose two footgun fields while Settings keeps the two that actually govern every check.

D21 — A rule's destination is a rule-set, and nothing else

Decided 2026-07-25 (product owner). config rule carried THREE ways to say where traffic is going: dst_domain (an inline domain list), dst_ip (an inline CIDR list) and dst_ruleset (a reference to a config ruleset). Three mechanisms meant three sets of semantics to learn and keep straight, and the inline ones were the worse half of the trade: they are re-parsed per rule instead of being compiled once into a .srs, they cannot be shared between rules, and their matcher vocabulary had drifted from the rule-set one in a way nobody could see (below).

Decision: dst_domain and dst_ip are removed (schema v2). dst_ruleset is the only destination matcher. Src, dst_port and proto are untouched — they are not lists of destinations and have no rule-set form.

  • Rejected: keep the inline lists as a shorthand. "One obvious way" is the whole point; a shorthand that quietly means something different from the long form (see the bare-entry trap) is worse than no shorthand.
  • Rejected: promote inline lists to rule-sets lazily at generate time. The config on disk would then not say what the router does, and the panel would have to render a list the user cannot find or edit.

The bare-entry trap, and how the migration handles it

The two contexts already disagreed about exactly one spelling, silently:

entry in a rule (dst_domain) in a rule-set (entry) migrated to
example.com exact host host + subdomains full:example.com
full:example.com exact host exact host unchanged
suffix:example.com / .example.com host + subdomains host + subdomains unchanged
keyword:ads substring substring unchanged
regexp:^ads\. pattern pattern (added here) unchanged
geosite:x / geoip:x inert (engine field removed) inert (unknown prefix) unchanged

shaterd migrate (schema v1→v2, shater/model/migrate.go) creates one inline config ruleset per rule that still carries a legacy list — rule-<rule name> for domains, rule-<rule name>-ip for addresses — moves the entries across with the conversion above, appends the new name to dst_ruleset, and deletes the old option. It is idempotent, it resumes an interrupted run, and it never overwrites a hand-written rule-set that already owns the generated name (it picks rule-<name>-2). regexp: support was added to inline rule-sets in the same change precisely so the move can be lossless.

geosite:/geoip: entries are copied VERBATIM rather than promoted to a source=geosite rule-set: those matchers have been inert since the engine dropped the route-rule geosite/geoip fields, and turning a dead matcher live during an upgrade would be a behaviour change, not a migration. The text is kept so the operator can see it and convert it deliberately.

One deliberate semantic change, called out: a rule that used BOTH lists matched them with AND (an engine route rule ANDs its matcher fields), which is almost never what "these sites and these networks" meant. The two generated rule-sets are ORed, because rule_set: [a, b] matches when either matches. Such a rule matches more after the migration than before; it affects only configs that used both fields at once.

Consequence: one destination mechanism, one vocabulary, one place a list is edited; every list is compiled once and reused. The panel's rule editor drops its Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of the rulesets that already exist, and nothing more. Creating and filling a list stays in the Rulesets panel — rejected: a "create a list from here" shortcut in the rule editor, because a second place to author a list is a second place for its semantics and its duplicate-name rules to drift, and the whole point of this decision was to stop having two.