Files
shater/docs-shater/DECISIONS.md
T
omar a57717dabb
release / aarch64_cortex-a53 (push) Successful in 3m43s
release / x86_64 (push) Successful in 3m30s
release / apk aarch64_cortex-a53 (push) Successful in 5m11s
release / apk x86_64 (push) Failing after 5m8s
release / release apk (push) Has been skipped
release / release (push) Successful in 12s
health plan S7: docs, contract comments, SPEC 019 update
- lx-changelog: health board + observatory + global probe + sub cache entry
- DECISIONS.md: D18 board vs delete-and-overlay, D19 observatory vs sweep, D20 global probe settings
- contract comments: urltest.go CheckOutbounds (fast circuit) + observatory.go loop (background circuit + freshness gate) document the two-circuit split
- SPEC 019: dial-error section updated - slots still not moved, but board verdict demotes dead slot on next pick + retry (§5.B); sticky/replace-in-slot/never-shrink invariants preserved
2026-07-24 18:30:48 +03:00

28 KiB
Raw Permalink Blame History

Decisions (ADR log)

Key architectural decisions and why, so nobody re-litigates them (we spent a whole discussion converging — see also CONTEXT.md).

D1 — Do NOT write a proxy engine from scratch

The proxy protocols (VLESS/VMess/Trojan/Shadowsocks, Reality/XTLS, AmneziaWG, Hysteria2/TUIC, transports, TLS fingerprinting) are a byte-precise, adversarial anti-DPI arms race maintained by large communities and changing monthly. A from-scratch engine would be slower, buggier, less secure, and more detectable — the opposite of "optimized." Reuse a real engine.

D2 — Engine = sing-box (via sing-box-lx), not xray-core

  • AmneziaWG 2.0 (I1–I5 CPS decoy packets) is a hard requirement; sing-box has first-class AmneziaWG, xray's is weaker. sing-box also has broader protocol coverage (Hysteria2, TUIC, ShadowTLS, MASQUE) and is library-first (libbox), which suits embedding.
  • sing-box-lx (github.com/Leadaxe/sing-box-lx) is a thin, actively-rebased downstream fork adding AmneziaWG 2.0, XHTTP, MASQUE/WARP and — usefully for us — gRPC observability of DNS queries / rules / outbounds, which feeds our stats.
  • License GPL-3.0, compatible with (and cleaner than) our prior GPL-2.0-or-later. See D6.
  • Cost accepted: our whole control plane / generator / share-link handling is re-based from xray-shaped (v0.1) to sing-box-shaped.

D3 — FORK sing-box-lx and embed the whole product inside it (not "depend as a library")

Considered: (a) consume as a pinned Go module + Gitea mirror; (b) fork. Chosen: (b) fork, so we can embed control-plane + admin panel + DNS filter directly and integrate tightly with engine internals (DNS, routing, stats). This was a deliberate call favouring maximum integration over minimum maintenance.

To keep the fork from rotting, a mandatory discipline:

  • Our code lives in new top-level dirs (shater/, panel/, openwrt/) that never collide with upstream files on rebase.
  • Edits to upstream files are minimal and marked // shater.
  • We rebase/merge onto sing-box-lx tags on a cadence (mirroring how sing-box-lx rebases onto sing-box). Fork = additive overlay tracking upstream, not a divergent rewrite.
  • main of the shater repo is the fork; v0.1 branch keeps the old project.

D4 — UI = thin LuCI launcher + a separate embedded admin panel

LuCI stays a minimal, pretty mini-dashboard with an "Open panel" button. The button mints a short-lived token in the authenticated LuCI/ubus session and redirects to our admin panel on its own port, served by the daemon; the panel validates the token and opens a session.

  • Why not do everything in LuCI: LuCI's form/view model is limiting for the rich stats/config UX we want.
  • Why not a standalone panel with its own login: a router-facing panel with its own auth is a serious security surface to get right. Bootstrapping the token from LuCI's existing, hardened auth avoids a second login system.
  • The panel is a modern SPA, embedded in and served by the forked binary.

D5 — Long blocklists: don't push megalists into dnsmasq; use an efficient matcher

Naive address=/domain/# in dnsmasq holds every domain in RAM and reloads slowly (100k–1M+ entry lists are common). Instead the engine's own domain matcher (sing-box already compiles geosite-scale lists) or our compact matcher (suffix hash-set + optional bloom prefilter, compiled/cached blob, dedup + subdomain-collapse, streaming parse, refresh by content-hash) does the filtering. Blocklist sources are flexible: inline / file / url / geosite (geosite only when geodata is present, else that source is inert / fail-open).

D6 — License = GPL-3.0

sing-box is GPL-3.0; linking it makes the combined work GPL-3.0. Our own files may stay GPL-2.0-or-later (which permits the upgrade), but the project LICENSE is GPL-3.0 for clarity.

D7 — Keep the v0.1 feed signing identity

The usign feed key 5ac4b177689cb8e0 (public key in dist/shater-feed.pub, secret in Gitea secret KEY_BUILD) carries over, so routers that already trust it keep verifying v0.2 packages. Do not regenerate it without a documented rotation.

D8 — Preserve, don't destroy: v0.1 lives on its branch

The reset moved the full working xray-based project to the v0.1 branch and cleaned main. Nothing is lost; reusable logic (reliability layer, nft/routing, sub-fetch, CI/signing, design system) is ported forward, not rewritten.

D9 — Router build uses a musl-static tag set (drop with_naive_outbound,with_purego)

Verified on the x86_64 OpenWrt VM (2026-07-14): the canonical Makefile.lx LX_TAGS produces a binary dynamically linked to glibc and it will NOT run on musl OpenWrt. Root cause (proven three ways — ldd, ELF PT_INTERP, static hello-world control): with_naive_outbound pulls in cronet-go, which needs with_purego; purego uses //go:cgo_import_dynamic to dlopen libcronet at runtime, forcing an ELF with PT_INTERP=/lib64/ld-linux-x86-64.so.2 + PT_DYNAMIC even with CGO_ENABLED=0. musl has no glibc loader → not found (rc 127).

  • Decision: the OpenWrt/router binary uses a router-specific tag set = LX_TAGS minus with_naive_outbound,with_purego. That variant is fully static (ET_EXEC, no PT_INTERP) and runs directly on musl. naiveproxy outbound is not in our feature set (FEATURES.md), so nothing we ship is lost. with_awg is independent of naive/purego and stays.
  • Keep badlinkname,tfogo_checklinkname0 + -checklinkname=0 (needed by badtls); they build fine static.
  • The desktop/CLI LX_TAGS stays as upstream for any non-router use. Do not reuse the desktop tag set for the router package — define SHATER_ROUTER_TAGS.
  • 2026-07-23: with_gvisor also dropped from the router set. The shater data plane is tproxy/redirect (netplane); generate never emits a tun inbound, so the userspace gvisor netstack (~3.6 MB) is unreachable. If a tun inbound ever appears it falls back to the system stack — re-add the tag then.
  • 2026-07-23: with_clash_api also dropped. The admin panel is shater's own web server and generate never emits a clash_api service; the desktop/CLI LX_TAGS keeps the tag for external dashboards.
  • 2026-07-23: with_dhcp also dropped. shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS transport is never generated, and the slim shater/registry never registers it.

D11 — The engine runs IN-PROCESS (box.New), not as a subprocess

v0.1 generated an xray JSON file and ran xray as a separate process. v0.2 embeds: the daemon builds option.Options in Go and drives the engine via box.New(box.Options{Context, Options}) + Start()/Close() in the same process (proven in Phase 1, shater/cmd/shater-proto). This is the whole point of the fork — it gives us in-process access to DNS query events, routing, and stats (D5/observability) without log-scraping, and one binary to ship.

  • Apply = atomic instance swap. sing-box box instances are built-then-started; a config change rebuilds option.Options, validates by constructing a fresh box.New (which runs every adapter's constructor), and on success swaps: start the new instance, then close the old (or close→start within a short window under the kill-switch). Commit-confirm rollback keeps the last-good option.Options.
  • The generator (v0.1 generate.go) is rewritten to emit sing-box option structs instead of xray JSON. The parsers (share-link/subscription/awg-conf) and the nft/policy-routing layer port with little change (engine-agnostic).

D12 — v0.2 binary = shaterd, our own entrypoint under shater/

The router ships a single binary shaterd = shater/cmd/shaterd: our main that imports the engine as a library (D11) and also runs the control-plane, the DNS filter, the stats aggregator, and the admin-panel web server. cmd/sing-box (upstream) stays untouched for CLI/debug use. The router build uses the musl tag set (D9); shaterd links the same engine packages cmd/sing-box does.

D10 — Ship the router binary UPX-lzma compressed

Raw static binary is ~40–43 MB; the stock OpenWrt ext4 rootfs is ~100 MB. UPX --lzma --best takes it to ~9–11 MB (measured), a comfortable flash budget with a modest one-time decompress-into-RAM cost at start. Ship compressed; keep an uncompressed artifact for debugging.

D13 — External DPI-bypass tool = ByeDPI (a SOCKS egress), NOT zapret

Decided 2026-07-14. We evaluated exactly two external desync tools — zapret (nfqws/tpws, NFQUEUE packet plane) vs ByeDPI/ciadpi (a local SOCKS5 desync proxy) — and picked one: ByeDPI. DPI-bypass stays a per-ruleset egress choice (a routing rule sends selected domains to it, exactly like a WG node or direct), never a global switch.

Why ByeDPI over zapret — decided by shater's architecture, not by raw method strength. In shater all LAN traffic is already TPROXY-intercepted into the in-process engine, and routing rules pick the egress:

  • ByeDPI is an egress. It's a SOCKS5 proxy on 127.0.0.1:<port>; a rule → socks outbound → ByeDPI desyncs (split/disorder/fake/oob/autottl, fake-TTL) → goes direct to the target. No tunnel, no exit-node bandwidth, low latency. Zero new nft rules; it never touches our packet plane. Ships as a ~100 KB musl-static binary behind one procd service.
  • zapret is a competing packet plane. nfqws hooks packets via its OWN nft/NFQUEUE rules, which would fight our verified fail-closed inet shater TPROXY table for the same packets/marks — a second packet plane on a fail-closed router multiplies kill-switch/leak surface, and it cannot be expressed as a sing-box egress. It also entangles netplane (a core verified subsystem), violating the additive-overlay discipline (D3).
  • zapret's real edge is QUIC / bad-checksum fakes at the raw-packet level. ByeDPI is TCP/TLS-centric and weak on QUIC. We accept that gap and close it by routing, not by tool: a rule dropping udp/443 for desync-domains forces them to fall back to TCP+TLS, which ByeDPI handles. That stays inside our rule model — no parallel packet plane.

Native fragment is free and complementary (not a second tool). The fork already compiles in, at the route-action level, tls_fragment + tls_fragment_fallback_delay, tls_record_fragment, tls_spoof/tls_spoof_method, and dialer udp_fragment. These are just engine flags on a direct egress and cover the light "just fragment the ClientHello" case with no external binary. ByeDPI is what we add for the stronger methods the engine lacks (fake/disorder/oob/ autottl).

Decision / implementation shape:

  1. Egress gains an optional dpi preset. Values map to escalating power: off → plain; fragment/record/spoof → native route-action flags on a direct outbound (in current Wave, ~free); byedpi → a socks egress pointed at a supervised local ByeDPI (ciadpi) instance (later phase: its own openwrt/ procd package + musl-static cross-build).
  2. shater/generate maps the native presets now; the byedpi egress kind is reserved in model and wired when its package lands.
  3. zapret is explicitly rejected for shater and should not be revisited unless in-place QUIC packet-desync becomes a hard requirement that routing cannot cover.

When zapret would have won: a pure DPI-bypass box with no tunnel engine (then nfqws's packet power is the whole product). shater already has a full tunnel engine

  • native fragment, so it wants the desync tool that composes with routing — ByeDPI.

D14 — LAN DNS anti-leak = sing-box hijack-dns route action, NOT a :53 listener

Decided 2026-07-15 after the Phase-2 gate exposed the gap (see memory shater-dns-no-engine-listener). The netplane dnsnat chain redirected LAN :53 to the router's :53, assuming the engine answers there — but generate/engine create NO :53 DNS server (the engine only binds the tproxy port). On the VM the redirect landed on stock dnsmasq, which resolved via its WAN upstream outside the tunnel = a DNS leak, and the engine's own resolvers (DoH-via-WARP, already built by generate.buildDNS with anti-leak detours) went unused. On a real box (dnsmasq removed) client DNS would break entirely.

Decision: use the sing-box-idiomatic hijack, not a listener. sing-box has no simple ":53 DNS inbound"; its DNS module is designed to be fed via a hijack-dns route action (C.RuleActionTypeHijackDNS). The tproxy inbound already sniffs (the leading sniffRule() in generate/route.go), so:

  • generate: after the sniff rule, add a route rule protocol: dns → action hijack-dns, so every sniffed DNS query is answered by the engine's internal resolver (which routes each query through its configured detour — the anti-leak).
  • netplane: STOP special-casing :53. Remove the prerouting :53 accept and the whole dnsnat chain so LAN :53 is DIVERTED by the normal tproxy catch-all into the engine, where hijack-dns catches it. Keep the :853 DoT reject (forces clients off DoT onto :53, which is now hijacked). DoH (:443) stays SNI-routed through the proxy.
  • dnsmasq: the tproxy divert steals LAN :53 before local delivery, so dnsmasq never sees LAN queries; it keeps serving ROUTER-local :53 only. No dnsmasq removal needed for correctness, though shipping images may still drop it.

Rejected alternative (a real engine :53 listener + keep dnsnat): non-idiomatic for sing-box, needs a bespoke DNS server, and duplicates what hijack-dns gives for free. generate + netplane MUST stay in agreement on this — they are the two halves of the same DNS plane.

D15 — DNS filter = sing-box rule-sets + reject DNS rules, NOT a custom matcher

Decided 2026-07-15 (Phase 4). D5 said "don't push megalists into dnsmasq; reuse the engine's matcher OR a compact custom matcher." sing-box already ships a compiled, memory-mapped rule-set matcher (geosite-scale, the .srs binary format + inline / local / remote rule-sets with auto-fetch+cache), and DNS rules support action: reject. So the DNS filter is built ON that, not a bespoke bloom matcher:

  • model: a blocklist section (name, enabled, source inline|file|url|geosite, url/path/entries, response nxdomain|zero|refuse) + an allowlist (overrides). A global dns_filter enable.
  • generate/dns.go: materialize each blocklist into a sing-box rule-set (inline rule for small inline lists; local rule-set for a compiled/file .srs; remote rule-set for a url, letting sing-box fetch+cache with a download detour through the proxy) and emit DNS rules: allowlist-accept FIRST (higher priority), then blocklist → reject (method default=NXDOMAIN, or a 0.0.0.0/:: predefined answer for zero). geosite sources are inert/fail-open when geodata is absent (D5).
  • daemon: a blocklist update verb (like sub update) to (re)compile file sources and warm remote caches; sing-box remote rule-sets self-refresh, so the daemon mostly compiles local sources + seeds well-known lists (StevenBlack / OISD / AdGuard) as disabled presets. Rationale: reuses a battle-tested, RAM-efficient matcher (no custom megalist code to get wrong), and it composes with the hijack-dns plane already in place (D14) — the filter rules sit on the same in-engine DNS path. Only build a custom matcher if a concrete sing-box rule-set limitation is hit (none known).

D16 — Enable sing-box cache_file; engine apply-swap must go close-first when it's shared

Decided 2026-07-15 (Phase 4). Remote rule-sets (url blocklists) and urltest/rdrc state need sing-box's experimental.cache_file to persist across restarts, auto-update on their interval, and keep RAM sane for megalists. So generate now emits cache_file (enabled, /etc/shater/cache.db, falling back to /tmp/shater-cache.db).

Consequence discovered on the VM: bbolt opens the cache DB under an EXCLUSIVE file lock and retries ~10s before failing timeout. The engine's default apply-swap starts the NEW box before closing the old (D11) — but the new box can't acquire the cache lock the old still holds, so EVERY live reconcile/apply stalled ~10s then failed (allowlist toggle, blocklist update, any edit silently didn't take). Fix (extends the D-fix for EADDRINUSE): engine.Apply now (a) treats the cache-file lock timeout as a swap conflict, and (b) proactively takes the close-old-then-start-new path when the incoming config shares an enabled cache_file with the running one (sharesCacheFileLock), avoiding the ~10s stall entirely. The fail-closed nft kill-switch covers the brief gap. Regression test TestEngineCacheFileSwap. See shater-engine-applyswap-portconflict (same close-first machinery, now also for the cache lock).

D17 — Audit pass: fail-closed must survive engine-start failure; several knobs were fiction

Decided 2026-07-20. A full-stack audit (8 parallel agents, every finding reproduced on the QEMU testbed) changed a few architectural positions. Recording the ones that constrain future work.

The data plane must never be absent while kill_switch=closed. Reproduced live: a remote rule-set that could not be fetched aborted box.Start, so apply failed and no nft table was installed at all. The router silently degraded to plain routing — no tunnel, no filtering, no kill-switch — while the panel looked healthy. Fail-closed had been implemented as "the engine died after starting"; it now also covers "the engine never started", via a holding plane (RenderHoldNft): forward blocked for the diverted interfaces, LAN-to-LAN + link-local + ND preserved, and the daemon's own egress marks accepted so it can still reach the network to recover. Management (SSH/LuCI/panel) is safe by construction — the holding chain hooks forward only, never input. Status now reports plane: full|hold|none. Rejected: "keep the previous table". After a cold boot there IS no previous table, which is precisely the reproduced scenario.

Engine start must not depend on the network. RemoteRuleSet.StartContext returns an error when its first fetch fails, and router.Start fast-fails the group — so a router that boots before its ISP link comes up cannot start the engine. Remote lists are now reachability-preflighted at generate time and omitted (loudly) rather than handed to the engine. Accepted cost, documented in ruleset.go: a brief source outage disables the list for up to one reconcile + probe TTL even when a good cached copy exists, because the cache cannot be consulted (the running engine holds an exclusive bbolt flock — see D16 — and a lock a process already holds on one fd denies its own second open). The structurally better fix (seed an empty cache entry between closing the old box and starting the new one, so StartContext never fetches at start) is deliberately deferred, not overlooked.

BlockDoH does not need to exclude upstream resolvers on the route plane. The original design punched a hole for every configured resolver IP so the engine could reach it — which also handed every LAN client an un-blocked public DoH endpoint, i.e. the more reputable the upstream, the wider the bypass. This was unnecessary: engine-originated DNS dials go through DetourDialer straight to the outbound and never traverse route rules, and they additionally carry LoopMark 0xff, which nft accepts before the tproxy divert. The exclusion now applies to the DNS plane only (a hostname upstream must still resolve), and the :443 rejects use the full list.

Globals.DNSMode is unimplementable and its control was removed. nftset named the v0.1 xray architecture, which does not exist here; fakeip is a resolver type in sing-box 1.14 (the top-level dns.fakeip block was removed upstream), already expressible as config resolver with type=fakeip + pool, which is strictly more capable (per-domain via a dns_rule). Wiring the toggle would have created a second configuration path with undefined precedence. The field stays in the model for config compatibility; the panel now documents the resolver route instead.

TPROXY cannot carry ICMP/IGMP/ESP/GRE, so this is now an explicit policy. Globals.Untunnelable = block (default, unchanged behaviour) | icmp | direct. Three values, not two, because the leaks differ in kind: an ICMP echo is ephemeral, user-initiated and reveals the address only to a host the user deliberately contacted, whereas ESP/GRE is a standing second tunnel carrying arbitrary traffic beside ours. A single toggle would make "I want ping to work" mean "I allow a parallel VPN bypass".

Fail-open degradations must be visible in the panel, not only in logread. The audit deliberately converted many aborts into warn-and-continue (an unfetchable list, a rule whose matchers were all invalid, a rejected interface name). apply now collects warnings from generate + netplane + model validation, classifies them (info|warning|critical, plus section+name for deep-linking) and exposes them on GET /api/status. Classification is by section, not by text markers: the wording of "list skipped" varies across six sites in generate and a marker list would silently miss the seventh.

blocklist source=url now consumes what public lists actually publish. It had accepted only compiled .srs, so every list D15 names (StevenBlack / OISD / AdGuard, all hosts-format) failed at apply with invalid sing-box rule-set file — a headline feature that broke on its own documented sources. The daemon now fetches (size-capped), parses hosts / plain-domain / AdBlock-comment formats, and compiles to .srs under /etc/shater/lists/, which the engine loads as a local rule-set. Compiling is what makes it affordable on the target hardware: measured on the testbed, StevenBlack's 2.4 MB of text became 80873 domains in a 491 KB .srs, costing ~0.5 MB of a rootfs with ~33 MB free — so it ships in the production posture rather than being traded away for the 8 KB geosite ads list.

D18 — Node health = a board with computed verdicts, NOT delete-on-failure + a dead-overlay

Decided 2026-07-24 (proxy-health plan §5.A). Upstream common/urltest recorded only success ({Time, Delay}) and a failed check deleted the entry — a dead member was indistinguishable from a never-measured one. On top of that, deaths found by shater's own probing went to a private engine.dead overlay that only the panel read: selection never saw them, and an old success in history counted as "alive forever". Net effect, reproduced in the field: YouTube dead on a real device while the panel showed the group healthy.

Decision: one health board, one source of truth. The history entry becomes {LastOK, Delay, LastFail}; a failure marks (MarkFailed), never deletes — deletion is reserved for removing a node from the config. The verdict is a method computed on read, not a stored field: alive when the success is fresher than the failure and younger than TTL; dead while a fresher failure is itself fresh; untested otherwise, with TTL = max(3 × global probe interval, 10 min). Everyone who learns of a death — the group's native checker, the observatory, a failed user dial — writes the same board; selection, balancer slots and the panel read the same board. The engine.dead overlay is deleted.

  • Rejected: keep delete-on-failure and widen the overlay to selection. Two stores of the same truth with undefined precedence; the overlay would need its own TTL/pruning and every reader would have to merge — exactly the divergence ("green panel, dead path") this decision exists to kill.
  • Rejected: persist the board across reboots. Measurements go stale faster than an overlay write is worth; after reboot everything is untested for seconds until the observatory's immediate first pass. In-memory, deliberately.
  • Deferred, not rejected: Xray-style sliding-window stats / hysteresis. The delay of the last success is enough to rank alive members in v1; the board keys and record shape leave room to widen without migration. Flapping is visible instead through the alive↔dead flip log (info) — the single diagnostic trail. Consequence: a stale success decays (alive → untested after TTL) instead of reading as alive forever, and a dead group blocks fail-closed — visible and alertable — rather than silently falling back.

D19 — Background probing = an observatory driven by rule reachability, NOT a population sweep

Decided 2026-07-24 (proxy-health plan §5.C, replaces shater/engine/sweep.go). The sweep probed the whole population round-robin — all nodes plus all per-group copies, used or not — yet chain copies were excluded entirely, so the one path users actually complained about ("rule → chain") was never measured. The manual "Test all nodes" run duplicated the same full-population walk on demand. On router budgets that is the wrong shape twice: work grows with the subscription size (~200 nodes), not with what the config uses.

Decision: probe only what the rules can reach. Following the Xray model (central observatory + balancers reading observations, §3 of the plan), a plan is built from the applied option.Options by reachability: enabled-rule targets (plus Final and DNS detours) → groups → members / egress copies; a chain is probed end-to-end by dialing its exit tag through the whole hop path (the Xray observatory-through-proxySettings equivalent); chain group-hop members are probed through their path prefix. Everything unreferenced stays untested and its group/chain is badged "unused" in the panel — so untested never looks like a health problem. A freshness gate skips tags an active group's own checker already measures; the cursor survives no-op reconciles (the sweep's release-blocker: cron reconciles must not restart the cycle); the manual probe-all is removed with the sweep.

  • Rejected: keep the sweep. Probes hundreds of unused nodes on a router budget and still misses chains; its "coverage" is what made per-node health look authoritative while the used path went unmeasured.
  • Rejected: probe every chain hop individually. Multiplied probe traffic for diagnostics that never affects selection — no choice depends on a middle hop's individual health. A dead middle hop makes the exit verdict honestly dead; localization is served by the verdict flip log.
  • Rejected: auto-reroute rules when a group dies. A dead group blocks fail-closed — visible and alertable. Silent rerouting would hide the outage and change routing semantics behind the user's back. Consequence: the probe budget is bounded by the config, not the subscription; cold start converges in seconds (immediate first pass after apply); wanting numbers for an unused group has one honest answer — reference it from a rule.

D20 — Probe URL / Interval are global-only; per-group overrides deleted

Decided 2026-07-24 (proxy-health plan §5.D). Groups carried optional ProbeURL/ProbeInterval overrides. That bred a documented ambiguity — one node shared by two groups with different URLs yields incomparable delays and needs "whose URL wins" dedup machinery (the (dial, URL) plan key) — and it breaks the D18 board: a verdict is a (record, TTL) pair with TTL derived from the probe interval, so per-group intervals would make the same record mean different things to different readers. Xray's observatory has exactly one global probeURL/probeInterval; that is the model we mapped onto (§3).

Decision: only Globals.ProbeURL/Globals.ProbeInterval. They drive the native group checker, the observatory and the manual exit test alike (failover keeps its 30 s default interval). Probe-plan dedup collapses to the dial path. Migration is the standard dead-option drainage: old UCI configs carrying probe_url/probe_interval on a group parse silently and the options vanish on the next render.

  • Rejected: keep per-group overrides. Incomparable measurements across groups, undefined semantics for shared members, and a per-group TTL that fractures the single-board verdict.
  • Deferred: per-tier probe cadence (rule-critical tags more often). If ever needed it is a field on the observatory's ProbeJob — a scheduling knob, not a return of per-group configuration. Consequence: all delay numbers are comparable (least ping ranks apples against apples), and group settings lose two footgun fields while Settings keeps the two that actually govern every check.