- drop c/Users/.../gen_linux_test (28MB binary accidentally committed in 129e31fbd)
- .gitignore: ignore .idea/ at any depth (shater/.idea from IDE)
- CLAUDE.md: orchestrator delegates to model fable
- bump shaterd/shater-core r2->r3, luci-app-shater r1->r2 for v0.2.1 release
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
release-apk now runs with if: !cancelled() so an unrelated arch build
failure (e.g. x86_64) does not block publishing the aarch64 apk feed.
download-artifact only fetches existing artifacts and the publish loop
already skips missing apkfeed-* dirs.
The filtered-query-log block (SegMeter + top blocked + live QueryLog) is
redundant with Insights. The stats poll stays - it still feeds the DNS
filtering and Groups modules.
The subscription form now carries exactly: name, URL, update interval,
fetch via (+detour when proxied), User-Agent, HWID, and extra headers.
Format, device identity, regex/proto/country filters, dedup and expiry-alert
knobs are gone from the form (still honoured from UCI; a save carries them
through untouched). Headers are edited as key-value rows and serialize to
the existing `Headers: []string` "Key: value" contract. Name is editable:
a rename rewrites FromSub on the sub's cached nodes and refuses collisions.
Rule.Kill ""/"default" used to drop the rule, letting its traffic fall
through to the broader rules below and finally the default route - a silent
leak of exactly the traffic the operator singled out. ruleKillFallback now
always returns an outbound: ""/"default"/"closed"/unrecognised block the
rule's traffic in place; only an explicit kill=open goes direct. The default
route exists solely for traffic no rule matched.
One physical device with several addresses (v4+v6, multiple leases) used to
show as several devices. Discover now folds addresses sharing a MAC into a
single row: new `ips` field lists every address primary-first, `ip` stays
the primary (most recent lease), state is the best among addresses.
MAC-less hosts remain one-per-IP. The panel shows the extra addresses as
secondary chips; naming keys the config entry by MAC whenever it is known.
The engine resolves domains for itself (node server names, urltest probes,
subscription/DoH fetches). Those queries carried an invalid client address
and still landed in every insights surface. Gate them out at the single
ingestion point (Aggregator.handleEvent): an event with an invalid or
loopback client is dropped before totals, top domains, per-server counts,
the timeline, and the query-log rings. Only LAN-client traffic is collected.
The built-in block-ads / ru-bypass / private rule bundles are gone:
model.Preset, Model.Presets, the `config preset` UCI section, its render,
the panel Preset type, and every fixture. The generate-side expansion was
already removed with the profile rewrite in the previous commit.
A profile is now a pure uplink-conditional rule switch: Name/Enabled/
Priority/MatchIface/Enable-DisableRules/EndpointResolver. The per-profile
DefaultTarget/DefaultEgress overrides and the profile-level schedule window
(SchedDays/SchedStart/SchedEnd/SchedUTCOffset) are removed from the model,
UCI parse/render, the generator, the WAN watcher, and the panel. Rule-level
scheduling is untouched.
The panel's uplink condition is now picked from a dropdown of the router's
UCI interfaces (GET /api/interfaces, same source as the egress picker);
stored interfaces missing from the live list render as stale chips.
generate/profile.go is rewritten here (applyProfilesAndPresets ->
applyProfiles), which also drops the generate-side preset-pack expansion;
the preset model/UCI/panel surface is removed in the next commit.
Audit of the run-51 logs showed actions/cache@v3.3.2 works on the act_runner
(cold: "Cache saved" x4; next job: "Cache restored" in ~2s, npm --fast skip,
usign/dl reused) and the sdk-cache mirror seeds correctly — but the single
biggest recurring cost was NOT cached: `scripts/feeds update -a` re-cloned
base+packages+luci+routing+telephony every run (~7.8 min warm x 4 SDK jobs on
the serial runner ≈ ~28 min/run wasted; github ~1 MB/s from this host).
Cache .cache/feeds/{opkg,apk} (workspace dir, actions/cache-persisted, visible
in the SDK container via --volumes-from) symlinked over the SDK's empty feeds/:
`feeds update` now git-fetches deltas (seconds) instead of full clones, always
checking out feeds.conf's pins. Fail-safe: any error on the cached checkouts
wipes the cache and clones fresh. Key by SDK release (feeds-opkg-24.10.4 /
feeds-apk-25.12.1) — stable across runs, invalidates on an SDK bump; both arch
jobs of a lane share one entry (identical pins, serial runner).
Steady-state warm run: ~60+ min -> ~20-22 min. Also documented in the workflow
header: never key a cache on github.sha — each cache SAVE stalls the act_runner
~3 min, so per-run-changing keys would add +3 min/entry every run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Builds were dominated by re-fetching the ImmortalWrt 25.12 SDK tarball
(~300 MB) every run, and a stalled downloads.immortalwrt.org transfer wedged
the apk job for 40+ min (plain `wget -q`, no timeout — same class as the
elfutils hang).
- New ci/fetch-sdk.sh (runner-side): cache -> our durable `sdk-cache` release
mirror -> upstream with a stall-kill (curl --speed-limit 64K --speed-time 60
--max-time 1800) + 3 retries + zstd-magic/size validation; seeds the mirror
best-effort (github.token, non-fatal) so cold runs never touch upstream again.
A 40-min hang is now impossible; the in-container fallback wget also gets
--timeout=60 --tries=3.
- actions/cache@v3.3.2 (last release on the OLD cache API that Gitea act_runner
implements; v4/v3.4.x use the new GitHub cache service) for: SDK tarball, SDK
dl/ sources (hash of package Makefiles; PKG_HASH re-verified so a stale cache
can't leak a wrong source), Go mod+build (go.sum), npm node_modules
(package-lock.json) with build-shaterd.sh --fast, apt archives, built usign.
Degrades safely if the cache server is off — the SDK mirror is independent.
- concurrency group release-${github.ref} cancel-in-progress so a re-dispatch
cancels the stale run instead of piling up (tags stay isolated).
Signing (usign/apk), both keys, per-arch publish, manual triggers, LOCALMIRROR
and the scoped 4-package collection are unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
byedpi (ciadpi) ships as a separate optional package; the panel offered the
`byedpi` egress type regardless, so selecting it without the package installed
created a dead, fail-closed egress. Now GET /api/status reports
`byedpi_installed` (exec.LookPath("ciadpi"), os.Stat fallback), and the egress
type picker disables the ByeDPI option with a hint when it's absent. Existing
byedpi egresses are never hidden or rewritten (config is sacred) — shown with an
amber warning and still round-trip on save; only NEW selection is blocked.
Unknown status (older daemon / fetch fail) => no gating.
Bump shaterd PKG_RELEASE 1 -> 2 (the SPA is embedded in the daemon binary).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On a live `apk add` / `opkg install`, shater-core's post-install hung forever
(observed on BananaWRT 25.12 at "Executing shater-core...post-install", child
`flock 1000` in locks_lock_inode_wait). Root cause: a USE_PROCD init sources
/lib/functions/procd.sh on every rc.common action, whose procd_lock takes a
BLOCKING exclusive flock on /var/lock/procd_<svc>.lock held until the process
exits. shater-cron re-execs itself as the eternal `loop`, so it held that lock
forever; base-files' default_postinst then ran `/etc/init.d/shater-cron enable`
synchronously inside the transaction, blocking on the flock while the package
manager waited on the postinst — a permanent deadlock.
Fix (two layers):
- shater-cron `loop()`: `exec 1000>&-` closes fd 1000 up front so the eternal
loop never holds the rc.common flock (no-op when procd_lock is absent).
- 30_shater-core: defer enable/restart into a detached (setsid + bounded)
background block that waits for apk/opkg to finish before touching init.d,
with all fds to /dev/null (a held stdout pipe would hang apk on EOF too).
Bump PKG_RELEASE 1 -> 2 so existing installs pick up the fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk lane runs the SDK on a bare debian:bookworm host, and the ImmortalWrt
25.12 SDK prerequisite check requires python3-distutils ("Checking
'python3-distutils'... failed. Prerequisite check failed." ->
.prereq-build Error 1), aborting before any package built. The opkg lane was
unaffected because the openwrt/sdk image ships the prereqs. Add
python3-distutils (and python3-setuptools defensively) to the host deps. apk
lane only; opkg untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two build-harness bugs surfaced once the SDK builds actually ran:
1. Permission denied writing the feed. ci/build-feed.sh creates $OUT as root on
the runner, but the openwrt/sdk container runs as the unprivileged `buildbot`
(uid 1000) — so `cp` of the .ipk into $OUT failed ("Permission denied"),
yielding 0 packages and then "usign signing failed" (nothing to sign). Set
`chmod 0777 "$OUT"` on the runner before docker run (a chmod from inside the
container, as buildbot, cannot fix a root-owned dir). The apk lane already
chmods $OUT from its root debian container, so it was unaffected.
2. Collecting the whole SDK. ci/sdk-build.sh did `find bin -name '*.ipk'`, which
swept up the hundreds of prebuilt kmod/base .ipk shipped in the SDK image —
bloating the feed and signing foreign kmods under our key. Collect strictly
our four by name (`<pkg>_*.ipk`) and require >=4. Applied the same narrowing
to ci/sdk-build-apk.sh (apk names carry no arch: `<pkg>-*.apk`), keeping the
"wrong SDK produced only .ipk" guard.
No change to the feed format/signing (usign/KEY_BUILD/shater-feed.pub, apk EC
key), the package set, or triggers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The apk (and opkg) SDK builds intermittently hung fetching build-time
sources like elfutils-0.192.tar.bz2 from sourceware.org: curl's
--connect-timeout covers only the TCP handshake, not a stalled mid-transfer,
so a slow upstream hangs the whole job (no --max-time in OpenWrt download.mk).
Set CONFIG_LOCALMIRROR=https://sources.cdn.openwrt.org in .config before
`make defconfig` in both ci/sdk-build-apk.sh and ci/sdk-build.sh so the SDK
tries the fast OpenWrt source CDN before each package's own PKG_SOURCE_URL —
fixes elfutils and any other flaky upstream. Mirror verified to hold the file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Same root cause as the apk jobs: scripts/build-shaterd.sh builds through a
go.mod `replace => ./submodules/wireguard-go` (AmneziaWG fork, bumped in
16a47b596), and actions/checkout does not fetch submodules by default, so
`go build` died with "reading submodules/wireguard-go/go.mod: no such file or
directory" in the opkg build jobs (x86_64 + aarch64_cortex-a53) as well. Init
only that one submodule — build-harness only, no change to the opkg feed
format/signing (usign/KEY_BUILD/shater-feed.pub) or package set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/build-shaterd.sh builds via a go.mod `replace => ./submodules/
wireguard-go` (the AmneziaWG-patched fork), so that submodule must exist or
`go build` dies with "reading submodules/wireguard-go/go.mod: no such file or
directory". actions/checkout does not fetch submodules by default. Init only
that one submodule (public GitHub URL; clients/apple+android are large and
unused) in the additive build-apk jobs — the opkg build jobs are left untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Additive next to the opkg/24.10 lane — nothing existing changed. The same 4
packages (shaterd, shater-core, luci-app-shater, byedpi) are built through the
official ImmortalWrt 25.12 apk-SDK and published as per-arch rolling releases
apk-latest-<arch> / apk-<tag>-<arch> (x86_64, aarch64_cortex-a53).
- ci/sdk-build-apk.sh: drives the 25.12 SDK inside debian:bookworm, compiles
.apk, then `apk mkndx --root T --keys-dir T/keys --allow-untrusted
--sign KEY --output packages.adb *.apk` — the exact form the OpenWrt 25.12
buildsystem uses (unsigned members, signed index).
- ci/build-feed-apk.sh: per-arch runner entrypoint (same --volumes-from and
artifact-order contract as ci/build-feed.sh).
- ci/gen-apk-key.sh: one-shot EC (prime256v1) keypair generator; private half
-> Gitea secret KEY_APK, public dist/shater-apk.pem committed.
- release.yml: additive build-apk / release-apk jobs; `on:` triggers untouched
(v* tags + workflow_dispatch); apk release tags deliberately non-`v*`.
- docs-shater/INSTALL.md section 6, .gitignore (out-apk/), dist/shater-apk.pem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The three dependent plaques (Keep file on flash / File size cap /
Download) only had their inner controls disabled — the rows still
looked live and hoverable. Field gains a disabled prop rendering
set-field--off: pointer-events none, opacity .45, grayscale, flattened
background — the whole plaque reads and behaves as switched off
(aria-disabled included). Wired to !logToFile on all three.
Verified live on the testbed: with the toggle off all three plaques are
inert and dimmed; flipping it back on restores them.
embed.FS carries no timestamps, so SPA responses went out with neither
Last-Modified nor ETag and browsers fell back to HEURISTIC caching — a
stale index.html kept showing the previous panel after a daemon upgrade
(user saw pre-c61cfe3a download buttons enabled with the toggle off).
index.html / SPA fallback / favicon / 404s => Cache-Control: no-cache
(revalidate every load); a HIT under assets/ (content-hashed by Vite)
=> public, max-age=31536000, immutable. Guarded by TestStaticCacheHeaders.
Verified live on the testbed: / and /settings no-cache, hashed asset
immutable, missing asset 404 no-cache; a plain reload now picks up the
new SPA.
Operator decision (supersedes ad9781bf): "Log file (downloadable)" off
must leave NO trace — delete the saved log files outright, and gray the
download buttons out while the file is off.
logsink: New and Reconfigure purge the active segment and the rotated
.1 whenever ToFile is off — at the old and new configured locations AND
both standard paths (a Persist flip must not leave a stale copy). A
daemon booting with the toggle off sweeps leftovers from a previous
life too.
panel: /api/log reverts to the pre-ad9781bf precedence (toggle off =>
syslog scrape / '# logging disabled'; segments are never served while
the file is off, even if a leftover exists). SPA: the three download
buttons are disabled when LogToFile is off; note/flash texts and the
?mock fixture say the files were deleted.
Verified on the docker-OpenWrt testbed via the panel: off+apply deletes
/var/log/shaterd.log* (and /etc/shater), buttons gray out; on+apply
starts a fresh file and downloads work again.
Flipping "Log file (downloadable)" off looked like it deleted the logs:
the file stayed on disk, but GET /api/log switched to the logread scrape
and the collected history became undownloadable (user report). The
toggle stops WRITING — it must not disown what was already collected.
New precedence: retained segments are streamed whenever they exist,
prefixed with a '# note: file logging is off …' line when the toggle is
off (even with syslog off too); the syslog-scrape and '# logging
disabled' fallbacks now speak only when nothing is retained. Settings
note/flash texts and the ?mock fixture updated to match.
Verified on the docker-OpenWrt testbed: with log_file=0 the download
returns the note + full history; re-enabling via the panel resumes
appending to the same file with nothing lost.
modernc.org/sqlite is the only pure-Go SQLite and costs ~3.5 MB in the
static shaterd link; the stats store never used anything SQL-specific —
it is a ring of two append-only streams with a monotonic seq cursor.
bbolt is already linked via experimental/cachefile, so the swap is free.
sqlitering.go -> boltring.go: buckets queries/conns keyed by 8-byte
big-endian seq (bbolt key order == cursor order), rows as JSON of the
existing LogEntry/ConnLogEntry structs, meta bucket carries the durable
per-stream HWM (same max-only monotonic semantics). The async writer
contract is untouched (writeCh 4096, drop counters, 256-row/500ms
batches, 30s retention tick). Disk cap: chunked oldest-first deletes
with the same hysteresis, then at most one bbolt Compact per pass
(sagernet/bbolt exports Compact) behind the same 110%+1MiB free-space
guard that gated VACUUM. A legacy SQLite-format stats.db (or any
unreadable file) is replaced in place with one warning; open failure
still falls back to the in-memory ring.
Zero user-visible change: the "sqlite" backend selector value and the
Snapshot.Backend string are kept verbatim. Tests ported assert-for-
assert plus new coverage: legacy-file replacement, overflow drops,
memRing parity round-trip, disk-cap convergence.
Router shaterd (linux/amd64): 28,004,478 -> 24,428,670 bytes (-3.58 MB);
modernc.org/* gone from go.mod/go.sum and the dep graph.
shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS
transport is never generated, and the slim shater/registry never
registers the transport, so the tag gated nothing in this binary.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
Upstream defect: acme.go is behind with_acme but acme_logger.go was not,
so go.uber.org/zap linked into every build even with ACME disabled. Only
acme.go references ACMELogWriter/ACMEEncoderConfig, so the twin gate is
behaviour-preserving; a with_acme build still compiles.
Marked lx:acme_logger_gate; upstream-PR candidate (drop the lx block on
rebase once merged). -94 KB on the router shaterd link.
The admin panel is shater's own web server and generate never emits a
clash_api service (shater/engine/engine.go pre-registers its own
dnstrack.Manager precisely because no api/clash_api observer exists on
the router). With include.Context gone the Clash server was already out
of the link; dropping the tag records the decision. Desktop/CLI LX_TAGS
keeps with_clash_api for external dashboards.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
The shater data plane is tproxy/redirect (netplane); generate never emits
a tun inbound, so the userspace gvisor netstack is unreachable code. With
the slim registry it was already dead-code eliminated by the linker —
dropping the tag makes the intent explicit and stops compiling ~3.6 MB of
gvisor sources into the build at all. A future tun inbound would fall
back to the system stack; re-add the tag if that ever lands.
D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
include.Context registers upstream's entire zoo — tor (bine), ssh, snell,
anytls, naive, masque, mdns/resolved, and the api service whose daemon
bridge links grpc+protobuf — none of which shater/generate ever emits.
shater/registry registers exactly what the generator can produce (tproxy/
redirect/direct/socks/http/mixed inbounds; direct/block/selector/urltest/
socks/http/ss/vmess/trojan/vless/shadowtls outbounds + hysteria2/tuic
behind with_quic; wireguard endpoint behind with_wireguard; tcp/udp/tls/
https/hosts/local/fakeip + DoQ/DoH3 DNS transports; xhttp + v2rayquic
transport blank imports), with build-tag stub twins so a tag-less
'go build ./...' stays green. Zero upstream diff.
Measured on linux/amd64 with the D9 router tag set: 47.05 MB -> 31.07 MB
raw (-34%); the unreachable gvisor stack and grpc/protobuf are dead-code
eliminated even before any tag changes. UPX --lzma artifact: 12.49 MB ->
~8.8 MB. Since a UPX-packed binary unpacks fully into anonymous pages,
the same ~16 MB comes off resident RAM on the router.
New "Daemon log" group on Settings (Faceplate): the LogLevel verbosity
select (relocated, honest note — "none" is a turn-down to panic-only, not
a true off; failures still alert), LogToFile / LogToSyslog / LogPersist
toggles, a validated LogMaxKB editor (128–8192), and three download
buttons (day / 3 days / everything) → downloadLog() fetches
GET /api/log?range=… with the session cookie, filename from
Content-Disposition, blob save. Honest warn plates: file-off = only a
slice of the syslog ring (ranges approximate); both-off = nothing is
written anywhere; flash vs tmpfs (lost on reboot, wears flash, ~33 MB
budget). Globals type gains LogToSyslog/LogToFile/LogPersist/LogMaxKB
1:1 with the backend; mock.ts mirrors the honesty contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The daemon's own log (engine + control-plane) went only to os.Stderr →
procd → the logread RAM ring: no file, no wall-clock timestamps, no size
cap, and "LogLevel=none" silenced ONLY the engine while the control-plane
kept writing at trace. So "download last day/3d/all", "limit the size" and
"fully turn it off" were all unmet.
New shater/logsink: one long-lived, atomically-reconfigurable Sink that
receives BOTH halves' byte streams, stamps every complete line with a UTC
RFC3339 wall clock (what makes date ranges real), and fans each line to a
size-capped 2-segment rotated file (ToFile) and/or the real os.Stderr
(ToSyslog). Both off = the line is dropped — the only true full silence.
Persistent path sits behind a stats-style disk-free guard (suspend+warn
once, auto-resume); tmpfs path is bounded by the cap itself. ANSI stripped
from the file copy only.
Wiring: control-plane via log.SetStdLogger over the sink; engine via a new
box.Options.DefaultLogWriter threaded into all three box.New sites
(apply/close-then-start/restore) by engine.SetDefaultLogWriter; live
reconfigure on every apply.Reconcile (SIGHUP / control socket / panel
apply) so panel changes take effect without a daemon restart.
controlLogLevel now makes the control-plane respect Globals.LogLevel
(silent vocab → panic-only; unknown → warn, mirroring generate).
Globals: LogToSyslog/LogToFile (default true), LogPersist (default false =
/var/log tmpfs; true = /etc/shater flash), LogMaxKB (default 2048, clamped
[128,8192]; 0 = default, not off — LogToFile is the off switch). UCI
parse/render/aliases + ValidateGlobals clamp-warn.
Endpoint GET /api/log?range=1d|3d|all (session-gated): streams the log line
by line, oldest segment first, filtered by the timestamp prefix; UTC
attachment filename. Honest fallbacks — file off + syslog on → a
"# note: … syslog ring only, ranges approximate" comment then a
`logread -e shater` scrape; both off → "# logging disabled". Unknown range
→ 400.
init.d: shater/shater-cron gate their `logger -t` status lines on
log_syslog so "logread off" is honest at the shell layer too.
Tests: logsink rotation-cap/timestamp/toggle-gating/engine→sink,
model round-trip + validate, endpoint session-gate/range/fallbacks.
VM-verified on QEMU (x86_64, OpenWrt 24.10): download+ranges, size-cap
rotation, file-off/full-off, persistent path, live reconfigure — all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A rule disabled in UCI but force-enabled by the active profile, with an
iface:/zone: source outside the tproxy-inbound set, got its engine route
rule but no nft divert — its traffic never entered the engine, and the
fail-closed forward drop and accept_local sysctls skipped the device too.
Root cause: generate applied profile enable/disable in its own
effectiveRules while netplane read raw Rule.Enabled. Fixed with one shared
resolver in the leaf model package (ResolveActiveProfile +
ApplyProfileRuleOverrides) that both the engine route plan and the nft
divert plan consult, so they can never disagree about which rules are in
force. applyLocked now threads a single now through generate + nft render +
sysctls, closing the schedule-boundary race between the two planes.
Verified: a profile-enabled iface rule now joins the divert set, the
per-rule tproxy emit, the fail-closed drop and the accept_local sysctl;
the inverse (profile-disabled) drops the device. Parity regression on the
existing generate profile tests stays green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The schedule evaluator called time.LoadLocation, but the router binary
embeds no tzdata and OpenWrt ships none — so LoadLocation always failed
and windows silently ran in UTC while the panel promised local time.
- Windows now anchor to SchedUTCOffset (minutes east of UTC), which the
panel captures from the editing browser on every schedule save; the
daemon evaluates now.UTC()+offset with no location database. This
sidesteps the weekly-recurring day-shift that a full local<->UTC
conversion cannot express in one window. SchedTZ is deleted (documented
in the removed-options list; old configs parse and drain it). DST is a
stated limitation (followed on re-save). generate/schedule.go collapses
from a second copy of the evaluator to a thin adapter over the model one.
- The iface-profile schedule was honored by the WAN watcher since
08d5d6cc, but generate warned "the watcher does not look at the schedule"
and the panel muted the editor with "the router ignores the schedule" —
both false. Warning and lie removed; the editor is live and labelled
"applies together with the uplink match".
- Stale fictions: FEATURES.md nftset/FakeIP-mode MVP line and the shipped
conffile's dead `option dns_mode 'nftset'` corrected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- The generator iterated dns_rules in raw slice order and never read
DNSRule.Order; first-match "top to bottom" was true only because the
panel pre-sorts. Now sorted by (Order, index) like route rules, so a
hand-edited UCI or any API client gets the declared order.
- dns_rule match_src silently dropped zone:/iface:/MAC entries (the
in-engine DNS plane matches source IPs only), which could widen a rule
to ALL clients or skip it entirely. Each dropped entry now warns, with
the consequence spelled out.
- BlockDoH :443 IP list was incomplete (no NextDNS anycast, no actual
cloudflare-dns.com 104.16.x, sparse v6). Extended across all listed
providers, now accepts anycast CIDRs, with a maintenance note that the
list is manual. The hostname NXDOMAIN + canary layers already cover
resolve-by-name; UI still says "well-known providers only".
- Panel: intercept-OFF copy no longer overstates the bypass (plaintext to
external resolvers is already hijacked by the D14 catch-all); allowlist
note gains the per-device-Block-wins caveat.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Audit found the Insights numbers were real but mislabelled:
- "Blocked" counted every NXDOMAIN, upstream timeout and zero-answer as a
block. Now blocked = strictly the engine's own filter verdict
(dnstrack.SourceFiltered: D15 blocklist + BlockDoH predefined-NXDOMAIN).
Failures (timeout/SERVFAIL-reject) become their own `failed` category;
the three counters are mutually exclusive and sum to Queries. The DNS
log "block" tag follows the same signal.
- Per-minute sparkline positioned buckets evenly by index over a sparse
slice, so "60 min" could span hours. Now points sit at their real
Bucket.Minute, gaps render as gaps, and the label states the actual
span + active-minute count instead of a fictional "last N min".
- Per-device domains skipped the LAN filter every other view applies, so
the router's own urltest/sub-fetch dials appeared as a phantom WAN-IP
device. Now folded into the `router` pseudo-device like the DNS log.
- Honest labels: "Outbounds/exits" -> "DNS lookups per exit"; top
domains/hosts meta "N tracked" -> "top N shown". Overview query log
shows the real per-device attribution, not the resolver tag; stale
"DNS events have no client IP" comments removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The whitelist accepted random but the human-readable "Supported:" tail
still named only four strategies — caught live on the VM where the model
and generate warnings disagreed about the supported set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- node_down is gone from AlertEventNames: nothing ever fired it, so a
channel subscribed to it was silence dressed as monitoring. An old
config's `list event 'node_down'` now warns as an unknown event and is
dropped. The accepted and emitted sets now coincide; the reserved-event
branch of ValidateAlerts stays as the guard against future divergence.
- model.ValidateGroups + KnownGroupStrategies: a typo'd strategy is
warned at validation time (was: silently built as least_test with only
a generate-time warning). Mirrors generate's warnGroupStrategy list.
- doc-comment honesty: random is a real engine mode (api.ts), sqlite
stats backend is a real persistent store (api.ts + model.go), resolver
type list gains tcp, pages/index.ts no longer claims Placeholder pages,
failover doc says fail-back exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New balancer flag priority (option.URLTestBalancerOptions.Priority): the
pool is re-derived from CONFIG ORDER every health-check tick via
balancePoolPriority/planPriorityPool — the first live member owns slot 0,
so when the top node answers probes again traffic returns to it on the
next tick (30s failover interval). Probing walks top-down and stops at
the first live node, so the steady-state cost stays one probe per tick.
Replace-in-slot deliberately does not apply here: failover forces sticky
["none"], so relocating nodes across slots breaks no flow keys. Plain
round_robin/random paths are untouched.
failoverBalancer() now emits Priority:true; the KNOWN LIMITATION note and
the panel's "nothing brings it back" blurb are gone because the
limitation is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- random is a REAL urltest mode (lx SPEC 019 v2): uniform draw over LIVE
slots only, pool sized to every member; dead slots keep their place
(never-shrink) but are never picked, for random AND round_robin AND
sticky (degrade-to-live). All-dead pools fall back to Select.
- Globals.SweepInterval + Globals.GroupHealth master switch, resolved by
one pure function (model.SweepSchedule) shared by validator and apply;
unparseable is warned-and-ON, never silently off. ConfigureSweep no
longer resets the cursor on every cron reconcile (release blocker:
a ~6-min cycle was restarted every 60s and never completed).
- multi-WAN egress gateway: ubus netifd status -> uci static -> main
table; a gatewayless non-P2P egress warns CRITICAL instead of silently
blackholing the second uplink.
- endpoint resolver (route.default_domain_resolver): bootstrap-direct
clone of a named resolver, profile override beats globals.
- chains are composable: chain: hops flatten recursively, cycle-guarded,
entry egress lifts only at position 0 (fail-closed mid-path).
- group test publishes its scope so "measuring" lights only the cards a
run covers; health run is explicitly global (all_nodes).
- panel: biased-sample honesty (no ratio until a failure CAN be on
record), profiles auto-pin plate, sweep/GroupHealth settings UI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two threads, both from the same question: does this setting do what it says?
## Health is per-group, because a dial path is per-group
Overview reported "119 up / 179 untested" over all nodes, and Nodes showed a
per-row ping. Both measured the wrong object. A group with an egress binding does
not dial the base node outbound at all — generate materialises per-member copies
(group-<name>-m<i>-<member>) and the group balances over those. So a node can be
alive direct and dead through the tunnel a group is bound to, and the panel said
"up". The same node in two groups with different egresses is two states that were
being collapsed into one number.
No new prober was needed: the engine already keeps a process-wide
urltest.HistoryStorage keyed by outbound tag, groups already probe their own
members into it, and stats already reads it — we simply projected it onto base
tags only. The copy tag carries the member NAME, so recovery needs no change to
generate. GET /api/groups/health now reports alive/dead/untested per group, with
an opt-in member list; the same summaries ride /api/stats so Overview needs no
extra poll.
Presentation is "alive / tested" with the untested remainder as a quiet aside,
never folded into dead: groups probe lazily and only while in use, so on a fresh
boot with a 376-node subscription almost everything is legitimately unmeasured,
and calling that "down" would scream catastrophe exactly when nothing is wrong.
Three things this exposed, all fixed here:
- TestAllNodes enumerated only om.Outbounds(), which by design excludes
endpoints. Every WireGuard/AmneziaWG node read "untested" forever no matter how
often the button was pressed — on a product whose driving requirement is AWG.
- Our ProbeFailDelay sentinel is gone from the engine's history entirely. It was
safe for least_test (slowest wins last) but round_robin's pool planner treats
any entry as alive, so a dead node could occupy the single slot of a failover
group — pinning failover to a corpse, which is the one thing it exists to
prevent. Failures now live in an engine-side overlay, invalidated by timestamp
against any later success; the engine's history holds measurements only.
- Writers now measure with the probe URL of the group that owns the tag. A manual
run used the global URL and overwrote a group's own measurement, leaving
least_test comparing latencies to different servers. Where one tag is claimed by
two groups with different URLs the ambiguity is inherent to the engine's keying,
so we use the neutral global URL and say so rather than picking a silent winner.
A scheduled sweep (engine/sweep.go, on by default) fills what nobody probes:
24 measurements per 10s tick, 12 in flight, skipping anything fresher than 5
minutes — ~5 min per full cycle on the production config. The freshness gate is
load-bearing beyond cost: testNodes skips a member whose history is younger than
the group's interval, so a sweep that kept refreshing would starve a failover
group's own 30s check. It is a layer under group-local probing, never a
replacement.
## Options that did not exist are deleted, not decorated
Audited every enumerated choice the panel offers against what this fork actually
implements (constant/, option/, protocol/group/, dns/), and split the results into
works / synonym / fiction. Fictions are removed outright — pre-release, so no
legacy path is kept for values nobody has.
Deleted: Group.Strategy random and leastload (both silently became least_test);
Egress.Type proxy and block (emitted no outbound at all — every binding dangled
and the traffic left over the plain WAN with the real IP); alert event node_down
(no emitter anywhere); Globals.DNSMode, Inbound.Sniff, Profile.ProbeURL/ProbeMode,
Node.XUDPConcurrency/XUDPProxyUDP443, Egress.Target.
Repaired instead of removed, because the engine could do them all along:
LogLevel "none" (asked for silence, got default verbosity — now LogOptions.Disabled);
Subscription.Format (the hint was stored, badge-rendered and ignored — the sniffing
parser always ran); Ruleset.Format (never read; the extension decided);
Group.Strategy failover (urltest + round_robin + pool 1 / tolerance 0 is exactly
"first working node in order" — verified through box.New with a sensitivity control).
Relabelled where the words lied: Single promised "first up" but a selector never
checks liveness; inbound "http" opens Mixed and answers SOCKS5 on the same port.
An unresolvable egress binding no longer fails open. It resolves to block, so the
bound traffic stops visibly instead of leaving with the real IP. Refusing the
config was the alternative and is worse: a dead engine under a closed kill-switch
blackholes the whole LAN over one mistyped name.
Also: Egress.Port no longer defaults to 1080 for every type. The parser invented
it, render persisted it, and the new "port is ignored" warning then fired on a
correctly written config — a warning on a healthy install is how a findings list
gets ignored.
## Geo data is no longer hardwired to one publisher
sing-geoip publishes country codes and nothing else — 238 files, all two-letter.
So "route Netflix around the tunnel" meant loading geoip-us: 159,125 prefixes and
~20 MB of kernel memory for something the netflix list does in 108 prefixes and
~14 KB. Provider selection is now a chain (generate/geosource.go): country codes
still resolve to SagerNet byte-identically, everything else to Loyalsoldier, and
metacubex adds AS<number> routing. Third-party .srs was verified to load with our
own reader (v1/v2 against our v5 ceiling) before any of this was built.
The ruleset preflight reads four header bytes over a ranged GET instead of HEAD,
so a rule-set whose format version we cannot parse degrades like an unreachable
one — that case would otherwise abort engine start, which is how the LAN goes down.
The status strip carried five pips — ENGINE active, UPTIME, CONFIG enabled,
DATA PLANE installed, KILL-SWITCH — and the user had to AND three of them
together to learn whether they were protected. `plane` and `engine_running`
already encode that, and more precisely than the booleans did. UPTIME duplicated
the Engine module's "running for"; KILL-SWITCH duplicated the module directly
below it. Collapsed to one derived line phrased in terms of traffic:
Protected — traffic from your network is going through the tunnel
Traffic blocked — the tunnel is down (hold, amber)
Not protected — traffic is going out directly (none + fail-closed, crit)
Not protected — running direct (none + fail-open, amber)
hold and none stay distinct: one is the kill-switch catching it, the other is
no safety net at all. Nothing was lost — every removed value still lives in the
module that owns it.
Findings are now routed by severity instead of all landing on the front page
(panel/src/findings.ts):
critical / warning -> Overview. Something needs attention.
info -> the page that owns the setting.
An info finding is a statement about the configuration: it never clears and asks
for nothing, so a permanent front-page entry only teaches people to skim the
list — which is how a real critical finding gets missed. The untunnelable note
now renders inside the Networks "Other traffic" section, beside the control it
describes. With nothing needing attention the section renders nothing at all.
Also fixed, found while auditing the rest of the labels: the Kill-switch module
read ARMED / policy: fail-closed with a green lamp even at plane=none — a
reassuring light directly beneath a readout saying nothing is protected. A
fail-closed setting is only armed if something is installed to enforce it, so it
now reads NOT IN EFFECT with a crit lamp and a "blocking now: no — nothing
installed" row; policy -> setting.
planeState.ts became the single source of the wording, and the plane banner was
dropped from Overview — it exists to carry the alarm to pages with no status
readout, and stacked under the new line it just said the same thing twice.
Verified against the live daemon on the bench: healthy, critical and hold states
all render correctly, console clean, note present on Networks and absent from
Overview.
.gitignore: MemPalace per-project files, added by the tooling.
Traffic TPROXY cannot carry (ICMP, IGMP, ESP/AH, GRE) was dropped for the whole
LAN regardless of routing. A box configured to tunnel only 8.8.8.8/32 still lost
ping to the entire internet, and with the shipped config RU addresses were
unpingable even though `ru-direct` sends them out unproxied — the very path where
TCP already exposes the real IP, so the drop prevented no leak at all.
The drop is now scoped to destinations the rules actually tunnel:
iifname "br-lan" meta l4proto != { tcp, udp } ip daddr @unt_d4_1 accept
iifname "br-lan" meta nfproto ipv4 drop
Destination sets come from the engine's already-parsed rule-sets via
ExtractIPSet(), so no .srs parsing and no second read of the bbolt cache the
engine holds locked. Rules are taken from the generated route rules, not the raw
model, so preset packs, WAN-profile overrides and schedules are all included.
Domain/geosite matchers are skipped when classifying: a packet with no stream
carries no domain, so such a rule can never apply to it.
Every policy line carries `l4proto != { tcp, udp }`, so no destination decision
can ever accept TCP/UDP — fail-closed is structurally untouched. Anything the
walk cannot prove direct (list not yet fetched, logical rule, unknown action,
inverted match) falls through to the drop and says so via an info finding.
No element cap: a continent-scale list loads in full. Measured on the bench with
geoip-us — 4s apply, 1.25 MB ruleset in 29.5k lines, ~27 MB RSS growth, engine
healthy. Cost is reported, not enforced; `untunnelable=direct` loads no sets.
Also fixed here, found while building it:
- plan warnings were computed and dropped, never reaching the operator; routing
them through the netplane channel was wrong (it marks everything critical by
construction), so they get their own info-level path
- nft ran with no timeout while holding the apply flock: one wedged invocation
would have deadlocked every later apply, reconcile and teardown. 60s cap; the
ruleset commits as a single netlink transaction, so killing it is safe
- set elements were emitted as one 3.1 MB line the lexer would hold as a single
token; now wrapped at 8 per line (identical to nft, readable when debugging)
- untunnelable copy still claimed ping never works; rewritten for the new
semantics across all three modes
panel: the theme switch read as a power toggle — it reused the component that
turns features on and off and sat inside the status cluster next to the ONLINE
lamp, so in light theme it looked like a switched-off appliance. Now a two-key
sun/moon selector, both states always visible (neither theme is an "off"), the
engaged key raised and lit by shading rather than accent colour, separated from
the indicators by a groove.
Verified on the OpenWrt bench: RU addresses ping, non-RU stay blocked, TCP routes
unchanged through the tunnel, DNS filtering and Block-DoH unaffected.
Group egress — for the case where the protocols themselves are DPI-blocked:
every node in the group dials ITS OWN server through the chosen egress (an
AmneziaWG tunnel, say), so the provider sees tunnel traffic instead of a VLESS
handshake. It binds the outgoing dial, not post-proxy traffic.
The binding is per-group, and that is the whole difficulty: group members are
SHARED outbounds, so two groups built from one subscription — one bound, one not
— would either leak the binding into the unbound group or fail to apply it. The
members of a bound group are therefore materialised as per-group copies
(group-<g>-m<i>-<member>), reusing the same rebuildNode the chain builder uses
for per-hop copies. Copies are made only when Egress is set, so an unbound group
over a 331-node subscription does not double the engine config. Copy tags are
checked against the node/group/egress/copy namespaces; a collision skips the
member with a warning rather than shadowing a real node. Precedence is chain hop
-> Node.Egress -> Group.Egress: a node pinned to a particular uplink was pinned
for a reason the group cannot know. A member whose copy cannot be built is
dropped rather than falling back to its unbound tag — falling back would leak
exactly the traffic the binding exists to hide.
Group test answers "what am I exiting through, and how fast": selected member,
latency, exit IP and country, via cloudflare.com/cdn-cgi/trace (country comes
free, so no GeoIP database on the router) with api.ipify.org as fallback. The
probe is pinned to the group's own outbound and refuses the direct outbound — a
direct answer would print the ISP's address and claim the tunnel works when it
does not. Measuring latency but failing to resolve the address stays ok=true
with an empty exit_ip; that is a working tunnel, not an error.
Uptime: /api/status gains started_unix + uptime_seconds, measured from process
start over a monotonic seam so an NTP step on an RTC-less router cannot be
reported as uptime. It is the daemon's uptime, not time since the last apply.
Also fixes: renaming an egress did not rewrite Group.Egress, silently dropping
the group back to the default route.
Verified on the testbed with the real 331-node subscription: two groups over one
subscription, one bound, one not — the bound group selected
group-auto-egress-m130-IE-trojan-141 while the unbound one selected the shared
IE-trojan-141, exit IP and country resolved for both, no group warnings.
Measured cost of binding a 331-node group: engine outbounds 335 -> 666, config
50 KB -> 114 KB, daemon RSS 62 MB -> 75 MB.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
I wrote the gap up as open while reviewing an agent report I had not yet seen;
the hosts/plain/AdBlock parse-and-compile path had in fact landed in the same
commit. Records the measurement that settles the disk question: StevenBlack's
2.4 MB of text compiles to 80873 domains in a 491 KB .srs, so it ships in the
production posture instead of being traded away for the 8 KB geosite list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE