`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.
`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.
THE BARE-ENTRY TRAP, and why the migration is not a copy
A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).
`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.
`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.
THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)
Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.
Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.
ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.
Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.
untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.
Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
32 KiB
Decisions (ADR log)
Key architectural decisions and why, so nobody re-litigates them (we spent a
whole discussion converging — see also CONTEXT.md).
D1 — Do NOT write a proxy engine from scratch
The proxy protocols (VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, AmneziaWG, Hysteria2/TUIC, transports, TLS fingerprinting) are a byte-precise, adversarial anti-DPI arms race maintained by large communities and changing monthly. A from-scratch engine would be slower, buggier, less secure, and more detectable — the opposite of "optimized." Reuse a real engine.
D2 — Engine = sing-box (via sing-box-lx), not xray-core
- AmneziaWG 2.0 (I1–I5 CPS decoy packets) is a hard requirement; sing-box has
first-class AmneziaWG, xray's is weaker. sing-box also has broader protocol
coverage (Hysteria2, TUIC, ShadowTLS, MASQUE) and is library-first
(
libbox), which suits embedding. sing-box-lx(github.com/Leadaxe/sing-box-lx) is a thin, actively-rebased downstream fork adding AmneziaWG 2.0, XHTTP, MASQUE/WARP and — usefully for us — gRPC observability of DNS queries / rules / outbounds, which feeds our stats.- License GPL-3.0, compatible with (and cleaner than) our prior GPL-2.0-or-later. See D6.
- Cost accepted: our whole control plane / generator / share-link handling is re-based from xray-shaped (v0.1) to sing-box-shaped.
D3 — FORK sing-box-lx and embed the whole product inside it (not "depend as a library")
Considered: (a) consume as a pinned Go module + Gitea mirror; (b) fork. Chosen: (b) fork, so we can embed control-plane + admin panel + DNS filter directly and integrate tightly with engine internals (DNS, routing, stats). This was a deliberate call favouring maximum integration over minimum maintenance.
To keep the fork from rotting, a mandatory discipline:
- Our code lives in new top-level dirs (
shater/,panel/,openwrt/) that never collide with upstream files on rebase. - Edits to upstream files are minimal and marked
// shater. - We rebase/merge onto sing-box-lx tags on a cadence (mirroring how sing-box-lx rebases onto sing-box). Fork = additive overlay tracking upstream, not a divergent rewrite.
mainof theshaterrepo is the fork;v0.1branch keeps the old project.
D4 — UI = thin LuCI launcher + a separate embedded admin panel
LuCI stays a minimal, pretty mini-dashboard with an "Open panel" button. The button mints a short-lived token in the authenticated LuCI/ubus session and redirects to our admin panel on its own port, served by the daemon; the panel validates the token and opens a session.
- Why not do everything in LuCI: LuCI's form/view model is limiting for the rich stats/config UX we want.
- Why not a standalone panel with its own login: a router-facing panel with its own auth is a serious security surface to get right. Bootstrapping the token from LuCI's existing, hardened auth avoids a second login system.
- The panel is a modern SPA, embedded in and served by the forked binary.
D5 — Long blocklists: don't push megalists into dnsmasq; use an efficient matcher
Naive address=/domain/# in dnsmasq holds every domain in RAM and reloads slowly
(100k–1M+ entry lists are common). Instead the engine's own domain matcher
(sing-box already compiles geosite-scale lists) or our compact matcher
(suffix hash-set + optional bloom prefilter, compiled/cached blob, dedup +
subdomain-collapse, streaming parse, refresh by content-hash) does the filtering.
Blocklist sources are flexible: inline / file / url / geosite
(geosite only when geodata is present, else that source is inert / fail-open).
D6 — License = GPL-3.0
sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files may stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is GPL-3.0 for clarity.
D7 — Keep the v0.1 feed signing identity
The usign feed key 5ac4b177689cb8e0 (public key in dist/shater-feed.pub,
secret in Gitea secret KEY_BUILD) carries over, so routers that already trust it
keep verifying v0.2 packages. Do not regenerate it without a documented rotation.
D8 — Preserve, don't destroy: v0.1 lives on its branch
The reset moved the full working xray-based project to the v0.1 branch and
cleaned main. Nothing is lost; reusable logic (reliability layer, nft/routing,
sub-fetch, CI/signing, design system) is ported forward, not rewritten.
D9 — Router build uses a musl-static tag set (drop with_naive_outbound,with_purego)
Verified on the x86_64 OpenWrt VM (2026-07-14): the canonical Makefile.lx
LX_TAGS produces a binary dynamically linked to glibc and it will NOT run on
musl OpenWrt. Root cause (proven three ways — ldd, ELF PT_INTERP, static
hello-world control): with_naive_outbound pulls in cronet-go, which needs
with_purego; purego uses //go:cgo_import_dynamic to dlopen libcronet at
runtime, forcing an ELF with PT_INTERP=/lib64/ld-linux-x86-64.so.2 + PT_DYNAMIC
even with CGO_ENABLED=0. musl has no glibc loader → not found (rc 127).
- Decision: the OpenWrt/router binary uses a router-specific tag set =
LX_TAGSminuswith_naive_outbound,with_purego. That variant is fully static (ET_EXEC, noPT_INTERP) and runs directly on musl. naiveproxy outbound is not in our feature set (FEATURES.md), so nothing we ship is lost.with_awgis independent of naive/purego and stays. - Keep
badlinkname,tfogo_checklinkname0+-checklinkname=0(needed by badtls); they build fine static. - The desktop/CLI
LX_TAGSstays as upstream for any non-router use. Do not reuse the desktop tag set for the router package — defineSHATER_ROUTER_TAGS. - 2026-07-23:
with_gvisoralso dropped from the router set. The shater data plane is tproxy/redirect (netplane); generate never emits a tun inbound, so the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever appears it falls back to the system stack — re-add the tag then. - 2026-07-23:
with_clash_apialso dropped. The admin panel is shater's own web server and generate never emits aclash_apiservice; the desktop/CLILX_TAGSkeeps the tag for external dashboards. - 2026-07-23:
with_dhcpalso dropped. shater resolver types areudp/tcp/doh/dot/local/fakeip; adhcp://DNS transport is never generated, and the slimshater/registrynever registers it.
D11 — The engine runs IN-PROCESS (box.New), not as a subprocess
v0.1 generated an xray JSON file and ran xray as a separate process. v0.2 embeds:
the daemon builds option.Options in Go and drives the engine via
box.New(box.Options{Context, Options}) + Start()/Close() in the same process
(proven in Phase 1, shater/cmd/shater-proto). This is the whole point of the
fork — it gives us in-process access to DNS query events, routing, and stats
(D5/observability) without log-scraping, and one binary to ship.
- Apply = atomic instance swap. sing-box box instances are built-then-started;
a config change rebuilds
option.Options, validates by constructing a freshbox.New(which runs every adapter's constructor), and on success swaps: start the new instance, then close the old (or close→start within a short window under the kill-switch). Commit-confirm rollback keeps the last-goodoption.Options. - The generator (v0.1
generate.go) is rewritten to emit sing-box option structs instead of xray JSON. The parsers (share-link/subscription/awg-conf) and the nft/policy-routing layer port with little change (engine-agnostic).
D12 — v0.2 binary = shaterd, our own entrypoint under shater/
The router ships a single binary shaterd = shater/cmd/shaterd: our main that
imports the engine as a library (D11) and also runs the control-plane, the DNS
filter, the stats aggregator, and the admin-panel web server. cmd/sing-box
(upstream) stays untouched for CLI/debug use. The router build uses the musl tag
set (D9); shaterd links the same engine packages cmd/sing-box does.
D10 — Ship the router binary UPX-lzma compressed
Raw static binary is ~40–43 MB; the stock OpenWrt ext4 rootfs is ~100 MB. UPX
--lzma --best takes it to ~9–11 MB (measured), a comfortable flash budget
with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep an
uncompressed artifact for debugging.
D13 — External DPI-bypass tool = ByeDPI (a SOCKS egress), NOT zapret
Decided 2026-07-14. We evaluated exactly two external desync tools — zapret
(nfqws/tpws, NFQUEUE packet plane) vs ByeDPI/ciadpi (a local SOCKS5 desync
proxy) — and picked one: ByeDPI. DPI-bypass stays a per-ruleset egress
choice (a routing rule sends selected domains to it, exactly like a WG node or
direct), never a global switch.
Why ByeDPI over zapret — decided by shater's architecture, not by raw method strength. In shater all LAN traffic is already TPROXY-intercepted into the in-process engine, and routing rules pick the egress:
- ByeDPI is an egress. It's a SOCKS5 proxy on
127.0.0.1:<port>; a rule →socksoutbound → ByeDPI desyncs (split/disorder/fake/oob/autottl, fake-TTL) → goes direct to the target. No tunnel, no exit-node bandwidth, low latency. Zero new nft rules; it never touches our packet plane. Ships as a ~100 KB musl-static binary behind one procd service. - zapret is a competing packet plane. nfqws hooks packets via its OWN
nft/NFQUEUE rules, which would fight our verified fail-closed
inet shaterTPROXY table for the same packets/marks — a second packet plane on a fail-closed router multiplies kill-switch/leak surface, and it cannot be expressed as a sing-box egress. It also entanglesnetplane(a core verified subsystem), violating the additive-overlay discipline (D3). - zapret's real edge is QUIC / bad-checksum fakes at the raw-packet level.
ByeDPI is TCP/TLS-centric and weak on QUIC. We accept that gap and close it by
routing, not by tool: a rule dropping
udp/443for desync-domains forces them to fall back to TCP+TLS, which ByeDPI handles. That stays inside our rule model — no parallel packet plane.
Native fragment is free and complementary (not a second tool). The fork already
compiles in, at the route-action level, tls_fragment +
tls_fragment_fallback_delay, tls_record_fragment, tls_spoof/tls_spoof_method,
and dialer udp_fragment. These are just engine flags on a direct egress and
cover the light "just fragment the ClientHello" case with no external binary.
ByeDPI is what we add for the stronger methods the engine lacks (fake/disorder/oob/
autottl).
Decision / implementation shape:
- Egress gains an optional
dpipreset. Values map to escalating power:off→ plain;fragment/record/spoof→ native route-action flags on adirectoutbound (in current Wave, ~free);byedpi→ asocksegress pointed at a supervised local ByeDPI (ciadpi) instance (later phase: its ownopenwrt/procd package + musl-static cross-build). shater/generatemaps the native presets now; thebyedpiegress kind is reserved inmodeland wired when its package lands.- zapret is explicitly rejected for shater and should not be revisited unless in-place QUIC packet-desync becomes a hard requirement that routing cannot cover.
When zapret would have won: a pure DPI-bypass box with no tunnel engine (then nfqws's packet power is the whole product). shater already has a full tunnel engine
- native fragment, so it wants the desync tool that composes with routing — ByeDPI.
D14 — LAN DNS anti-leak = sing-box hijack-dns route action, NOT a :53 listener
Decided 2026-07-15 after the Phase-2 gate exposed the gap (see memory
shater-dns-no-engine-listener). The netplane dnsnat chain redirected LAN :53
to the router's :53, assuming the engine answers there — but generate/engine
create NO :53 DNS server (the engine only binds the tproxy port). On the VM the
redirect landed on stock dnsmasq, which resolved via its WAN upstream outside
the tunnel = a DNS leak, and the engine's own resolvers (DoH-via-WARP, already
built by generate.buildDNS with anti-leak detours) went unused. On a real box
(dnsmasq removed) client DNS would break entirely.
Decision: use the sing-box-idiomatic hijack, not a listener. sing-box has no
simple ":53 DNS inbound"; its DNS module is designed to be fed via a hijack-dns
route action (C.RuleActionTypeHijackDNS). The tproxy inbound already sniffs
(the leading sniffRule() in generate/route.go), so:
- generate: after the sniff rule, add a route rule
protocol: dns → action hijack-dns, so every sniffed DNS query is answered by the engine's internal resolver (which routes each query through its configured detour — the anti-leak). - netplane: STOP special-casing
:53. Remove the prerouting:53 acceptand the wholednsnatchain so LAN:53is DIVERTED by the normal tproxy catch-all into the engine, wherehijack-dnscatches it. Keep the:853DoT reject (forces clients off DoT onto:53, which is now hijacked). DoH (:443) stays SNI-routed through the proxy. - dnsmasq: the tproxy divert steals LAN
:53before local delivery, so dnsmasq never sees LAN queries; it keeps serving ROUTER-local:53only. No dnsmasq removal needed for correctness, though shipping images may still drop it.
Rejected alternative (a real engine :53 listener + keep dnsnat): non-idiomatic
for sing-box, needs a bespoke DNS server, and duplicates what hijack-dns gives for
free. generate + netplane MUST stay in agreement on this — they are the two
halves of the same DNS plane.
D15 — DNS filter = sing-box rule-sets + reject DNS rules, NOT a custom matcher
Decided 2026-07-15 (Phase 4). D5 said "don't push megalists into dnsmasq; reuse the
engine's matcher OR a compact custom matcher." sing-box already ships a compiled,
memory-mapped rule-set matcher (geosite-scale, the .srs binary format + inline /
local / remote rule-sets with auto-fetch+cache), and DNS rules support
action: reject. So the DNS filter is built ON that, not a bespoke bloom matcher:
- model: a
blocklistsection (name, enabled, sourceinline|file|url|geosite, url/path/entries, responsenxdomain|zero|refuse) + anallowlist(overrides). A globaldns_filterenable. - generate/dns.go: materialize each blocklist into a sing-box rule-set (inline
rule for small inline lists;
localrule-set for a compiled/file.srs;remoterule-set for a url, letting sing-box fetch+cache with a download detour through the proxy) and emit DNS rules: allowlist-accept FIRST (higher priority), then blocklist →reject(methoddefault=NXDOMAIN, or a0.0.0.0/::predefined answer forzero). geosite sources are inert/fail-open when geodata is absent (D5). - daemon: a
blocklist updateverb (likesub update) to (re)compile file sources and warm remote caches; sing-box remote rule-sets self-refresh, so the daemon mostly compiles local sources + seeds well-known lists (StevenBlack / OISD / AdGuard) as disabled presets. Rationale: reuses a battle-tested, RAM-efficient matcher (no custom megalist code to get wrong), and it composes with the hijack-dns plane already in place (D14) — the filter rules sit on the same in-engine DNS path. Only build a custom matcher if a concrete sing-box rule-set limitation is hit (none known).
D16 — Enable sing-box cache_file; engine apply-swap must go close-first when it's shared
Decided 2026-07-15 (Phase 4). Remote rule-sets (url blocklists) and urltest/rdrc
state need sing-box's experimental.cache_file to persist across restarts,
auto-update on their interval, and keep RAM sane for megalists. So generate now
emits cache_file (enabled, /etc/shater/cache.db, falling back to
/tmp/shater-cache.db).
Consequence discovered on the VM: bbolt opens the cache DB under an EXCLUSIVE
file lock and retries ~10s before failing timeout. The engine's default
apply-swap starts the NEW box before closing the old (D11) — but the new box can't
acquire the cache lock the old still holds, so EVERY live reconcile/apply stalled
~10s then failed (allowlist toggle, blocklist update, any edit silently didn't
take). Fix (extends the D-fix for EADDRINUSE): engine.Apply now (a) treats the
cache-file lock timeout as a swap conflict, and (b) proactively takes the
close-old-then-start-new path when the incoming config shares an enabled
cache_file with the running one (sharesCacheFileLock), avoiding the ~10s stall
entirely. The fail-closed nft kill-switch covers the brief gap. Regression test
TestEngineCacheFileSwap. See shater-engine-applyswap-portconflict (same
close-first machinery, now also for the cache lock).
D17 — Audit pass: fail-closed must survive engine-start failure; several knobs were fiction
Decided 2026-07-20. A full-stack audit (8 parallel agents, every finding reproduced on the QEMU testbed) changed a few architectural positions. Recording the ones that constrain future work.
The data plane must never be absent while kill_switch=closed. Reproduced live:
a remote rule-set that could not be fetched aborted box.Start, so apply failed
and no nft table was installed at all. The router silently degraded to plain
routing — no tunnel, no filtering, no kill-switch — while the panel looked healthy.
Fail-closed had been implemented as "the engine died after starting"; it now also
covers "the engine never started", via a holding plane (RenderHoldNft): forward
blocked for the diverted interfaces, LAN-to-LAN + link-local + ND preserved, and the
daemon's own egress marks accepted so it can still reach the network to recover.
Management (SSH/LuCI/panel) is safe by construction — the holding chain hooks
forward only, never input. Status now reports plane: full|hold|none.
Rejected: "keep the previous table". After a cold boot there IS no previous table,
which is precisely the reproduced scenario.
Engine start must not depend on the network. RemoteRuleSet.StartContext returns
an error when its first fetch fails, and router.Start fast-fails the group — so a
router that boots before its ISP link comes up cannot start the engine. Remote lists
are now reachability-preflighted at generate time and omitted (loudly) rather than
handed to the engine. Accepted cost, documented in ruleset.go: a brief source
outage disables the list for up to one reconcile + probe TTL even when a good cached
copy exists, because the cache cannot be consulted (the running engine holds an
exclusive bbolt flock — see D16 — and a lock a process already holds on one fd denies
its own second open). The structurally better fix (seed an empty cache entry between
closing the old box and starting the new one, so StartContext never fetches at
start) is deliberately deferred, not overlooked.
BlockDoH does not need to exclude upstream resolvers on the route plane. The
original design punched a hole for every configured resolver IP so the engine could
reach it — which also handed every LAN client an un-blocked public DoH endpoint, i.e.
the more reputable the upstream, the wider the bypass. This was unnecessary:
engine-originated DNS dials go through DetourDialer straight to the outbound and
never traverse route rules, and they additionally carry LoopMark 0xff, which nft
accepts before the tproxy divert. The exclusion now applies to the DNS plane only
(a hostname upstream must still resolve), and the :443 rejects use the full list.
Globals.DNSMode is unimplementable and its control was removed. nftset named
the v0.1 xray architecture, which does not exist here; fakeip is a resolver type
in sing-box 1.14 (the top-level dns.fakeip block was removed upstream), already
expressible as config resolver with type=fakeip + pool, which is strictly more
capable (per-domain via a dns_rule). Wiring the toggle would have created a second
configuration path with undefined precedence. The field stays in the model for config
compatibility; the panel now documents the resolver route instead.
TPROXY cannot carry ICMP/IGMP/ESP/GRE, so this is now an explicit policy.
Globals.Untunnelable = block (default, unchanged behaviour) | icmp | direct.
Three values, not two, because the leaks differ in kind: an ICMP echo is ephemeral,
user-initiated and reveals the address only to a host the user deliberately contacted,
whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A
single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".
Fail-open degradations must be visible in the panel, not only in logread. The
audit deliberately converted many aborts into warn-and-continue (an unfetchable list,
a rule whose matchers were all invalid, a rejected interface name). apply now
collects warnings from generate + netplane + model validation, classifies them
(info|warning|critical, plus section+name for deep-linking) and exposes them on
GET /api/status. Classification is by section, not by text markers: the wording
of "list skipped" varies across six sites in generate and a marker list would silently
miss the seventh.
blocklist source=url now consumes what public lists actually publish. It had
accepted only compiled .srs, so every list D15 names (StevenBlack / OISD / AdGuard,
all hosts-format) failed at apply with invalid sing-box rule-set file — a headline
feature that broke on its own documented sources. The daemon now fetches (size-capped),
parses hosts / plain-domain / AdBlock-comment formats, and compiles to .srs under
/etc/shater/lists/, which the engine loads as a local rule-set. Compiling is what
makes it affordable on the target hardware: measured on the testbed, StevenBlack's
2.4 MB of text became 80873 domains in a 491 KB .srs, costing ~0.5 MB of a rootfs
with ~33 MB free — so it ships in the production posture rather than being traded away
for the 8 KB geosite ads list.
D18 — Node health = a board with computed verdicts, NOT delete-on-failure + a dead-overlay
Decided 2026-07-24 (proxy-health plan §5.A). Upstream common/urltest recorded
only success ({Time, Delay}) and a failed check deleted the entry — a dead
member was indistinguishable from a never-measured one. On top of that, deaths
found by shater's own probing went to a private engine.dead overlay that only
the panel read: selection never saw them, and an old success in history counted
as "alive forever". Net effect, reproduced in the field: YouTube dead on a real
device while the panel showed the group healthy.
Decision: one health board, one source of truth. The history entry becomes
{LastOK, Delay, LastFail}; a failure marks (MarkFailed), never deletes —
deletion is reserved for removing a node from the config. The verdict is a
method computed on read, not a stored field: alive when the success is
fresher than the failure and younger than TTL; dead while a fresher failure is
itself fresh; untested otherwise, with TTL = max(3 × global probe interval,
10 min). Everyone who learns of a death — the group's native checker, the
observatory, a failed user dial — writes the same board; selection, balancer
slots and the panel read the same board. The engine.dead overlay is deleted.
- Rejected: keep delete-on-failure and widen the overlay to selection. Two stores of the same truth with undefined precedence; the overlay would need its own TTL/pruning and every reader would have to merge — exactly the divergence ("green panel, dead path") this decision exists to kill.
- Rejected: persist the board across reboots. Measurements go stale faster
than an overlay write is worth; after reboot everything is
untestedfor seconds until the observatory's immediate first pass. In-memory, deliberately. - Deferred, not rejected: Xray-style sliding-window stats / hysteresis. The delay of the last success is enough to rank alive members in v1; the board keys and record shape leave room to widen without migration. Flapping is visible instead through the alive↔dead flip log (info) — the single diagnostic trail. Consequence: a stale success decays (alive → untested after TTL) instead of reading as alive forever, and a dead group blocks fail-closed — visible and alertable — rather than silently falling back.
D19 — Background probing = an observatory driven by rule reachability, NOT a population sweep
Decided 2026-07-24 (proxy-health plan §5.C, replaces shater/engine/sweep.go).
The sweep probed the whole population round-robin — all nodes plus all
per-group copies, used or not — yet chain copies were excluded entirely, so the
one path users actually complained about ("rule → chain") was never measured.
The manual "Test all nodes" run duplicated the same full-population walk on
demand. On router budgets that is the wrong shape twice: work grows with the
subscription size (~200 nodes), not with what the config uses.
Decision: probe only what the rules can reach. Following the Xray model
(central observatory + balancers reading observations, §3 of the plan), a plan
is built from the applied option.Options by reachability: enabled-rule
targets (plus Final and DNS detours) → groups → members / egress copies; a
chain is probed end-to-end by dialing its exit tag through the whole hop
path (the Xray observatory-through-proxySettings equivalent); chain group-hop
members are probed through their path prefix. Everything unreferenced stays
untested and its group/chain is badged "unused" in the panel — so untested
never looks like a health problem. A freshness gate skips tags an active
group's own checker already measures; the cursor survives no-op reconciles
(the sweep's release-blocker: cron reconciles must not restart the cycle);
the manual probe-all is removed with the sweep.
- Rejected: keep the sweep. Probes hundreds of unused nodes on a router budget and still misses chains; its "coverage" is what made per-node health look authoritative while the used path went unmeasured.
- Rejected: probe every chain hop individually. Multiplied probe traffic
for diagnostics that never affects selection — no choice depends on a middle
hop's individual health. A dead middle hop makes the exit verdict honestly
dead; localization is served by the verdict flip log. - Rejected: auto-reroute rules when a group dies. A dead group blocks fail-closed — visible and alertable. Silent rerouting would hide the outage and change routing semantics behind the user's back. Consequence: the probe budget is bounded by the config, not the subscription; cold start converges in seconds (immediate first pass after apply); wanting numbers for an unused group has one honest answer — reference it from a rule.
D20 — Probe URL / Interval are global-only; per-group overrides deleted
Decided 2026-07-24 (proxy-health plan §5.D). Groups carried optional
ProbeURL/ProbeInterval overrides. That bred a documented ambiguity — one
node shared by two groups with different URLs yields incomparable delays and
needs "whose URL wins" dedup machinery (the (dial, URL) plan key) — and it
breaks the D18 board: a verdict is a (record, TTL) pair with TTL derived from
the probe interval, so per-group intervals would make the same record mean
different things to different readers. Xray's observatory has exactly one
global probeURL/probeInterval; that is the model we mapped onto (§3).
Decision: only Globals.ProbeURL/Globals.ProbeInterval. They drive the
native group checker, the observatory and the manual exit test alike (failover
keeps its 30 s default interval). Probe-plan dedup collapses to the dial path.
Migration is the standard dead-option drainage: old UCI configs carrying
probe_url/probe_interval on a group parse silently and the options vanish
on the next render.
- Rejected: keep per-group overrides. Incomparable measurements across groups, undefined semantics for shared members, and a per-group TTL that fractures the single-board verdict.
- Deferred: per-tier probe cadence (rule-critical tags more often). If ever
needed it is a field on the observatory's
ProbeJob— a scheduling knob, not a return of per-group configuration. Consequence: all delay numbers are comparable (least ping ranks apples against apples), and group settings lose two footgun fields while Settings keeps the two that actually govern every check.
D21 — A rule's destination is a rule-set, and nothing else
Decided 2026-07-25 (product owner). config rule carried THREE ways to say
where traffic is going: dst_domain (an inline domain list), dst_ip (an inline
CIDR list) and dst_ruleset (a reference to a config ruleset). Three
mechanisms meant three sets of semantics to learn and keep straight, and the
inline ones were the worse half of the trade: they are re-parsed per rule instead
of being compiled once into a .srs, they cannot be shared between rules, and
their matcher vocabulary had drifted from the rule-set one in a way nobody could
see (below).
Decision: dst_domain and dst_ip are removed (schema v2). dst_ruleset is
the only destination matcher. Src, dst_port and proto are untouched —
they are not lists of destinations and have no rule-set form.
- Rejected: keep the inline lists as a shorthand. "One obvious way" is the whole point; a shorthand that quietly means something different from the long form (see the bare-entry trap) is worse than no shorthand.
- Rejected: promote inline lists to rule-sets lazily at generate time. The config on disk would then not say what the router does, and the panel would have to render a list the user cannot find or edit.
The bare-entry trap, and how the migration handles it
The two contexts already disagreed about exactly one spelling, silently:
| entry | in a rule (dst_domain) |
in a rule-set (entry) |
migrated to |
|---|---|---|---|
example.com |
exact host | host + subdomains | full:example.com |
full:example.com |
exact host | exact host | unchanged |
suffix:example.com / .example.com |
host + subdomains | host + subdomains | unchanged |
keyword:ads |
substring | substring | unchanged |
regexp:^ads\. |
pattern | pattern (added here) | unchanged |
geosite:x / geoip:x |
inert (engine field removed) | inert (unknown prefix) | unchanged |
shaterd migrate (schema v1→v2, shater/model/migrate.go) creates one inline
config ruleset per rule that still carries a legacy list — rule-<rule name>
for domains, rule-<rule name>-ip for addresses — moves the entries across with
the conversion above, appends the new name to dst_ruleset, and deletes the old
option. It is idempotent, it resumes an interrupted run, and it never overwrites
a hand-written rule-set that already owns the generated name (it picks
rule-<name>-2). regexp: support was added to inline rule-sets in the same
change precisely so the move can be lossless.
geosite:/geoip: entries are copied VERBATIM rather than promoted to a
source=geosite rule-set: those matchers have been inert since the engine
dropped the route-rule geosite/geoip fields, and turning a dead matcher live
during an upgrade would be a behaviour change, not a migration. The text is kept
so the operator can see it and convert it deliberately.
One deliberate semantic change, called out: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which is
almost never what "these sites and these networks" meant. The two generated
rule-sets are ORed, because rule_set: [a, b] matches when either matches. Such
a rule matches more after the migration than before; it affects only configs that
used both fields at once.
Consequence: one destination mechanism, one vocabulary, one place a list is edited; every list is compiled once and reused. The panel's rule editor drops its Domain(s) and IP/CIDR(s) fields; its destination control is a checkbox list of the rulesets that already exist, and nothing more. Creating and filling a list stays in the Rulesets panel — rejected: a "create a list from here" shortcut in the rule editor, because a second place to author a list is a second place for its semantics and its duplicate-name rules to drift, and the whole point of this decision was to stop having two.