16 Commits
Author SHA1 Message Date
omarandClaude Opus 5 625942834b fix(gate): [5/7] hid the runner's exit code exactly when it explained everything
test / go + panel tests (push) Successful in 1m40s
release / test gate (push) Successful in 1m40s
release / apk aarch64_cortex-a53 (push) Successful in 5m25s
release / apk x86_64 (push) Successful in 3m16s
release / release apk (push) Successful in 8s
The line naming a nonzero `go test` status was printed only when every
privileged test had produced a verdict — on the reasoning that a named FAILED
already explains the status. The case that actually happens is the opposite
one: the run dies at package level, so it names no test, so the loop above
prints MISSING for all of them, and the one line pointing at the real cause was
the one suppressed. A reader then goes hunting for three vanished tests instead
of at the build error above.

To be exact about what was and was not broken, because the framing matters: the
exit status was never SWALLOWED. priv_bad is set by the MISSING branch, so
FAILED is set and the gate fails either way — this was a diagnosis bug, not a
correctness one. What changes is whether the log says why.

Verified on the branch a green run never reaches, by driving the edited block
with all four (priv_rc, priv_bad) combinations: the new message appears only for
(1,1), the old one only for (1,0), and priv_bad/FAILED come out 1 in both. The
full gate is green with the change in, which covers the (0,0) path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-28 08:43:08 +03:00
omarandClaude Opus 5 4869d62e02 feat(egress)!: remove byedpi — what it replaced was not weak, it was broken (D29)
test / go + panel tests (push) Successful in 1m39s
release / test gate (push) Successful in 1m39s
release / apk aarch64_cortex-a53 (push) Failing after 2m54s
release / apk x86_64 (push) Failing after 2m54s
release / release apk (push) Failing after 1m35s
The `byedpi` egress kind, the `openwrt/byedpi` package (`ciadpi`), the readiness
endpoint and the panel plate are gone. D13 is not deleted from DECISIONS.md; it
is REVERSED there, with the reason, because the reason is the whole point.

D13 adopted an external desync process on an observation: the engine's own
`tls_fragment`/`tls_record_fragment` were tried against a live ISP and did not
get through, so the method was judged too weak for anything past "just fragment
the ClientHello". The method was never tried. `common/tlsfragment` dropped a
number of labels equal to the number of DOTS in the name, and a name always has
one more label than it has dots — so the cut always landed inside the FIRST
label. `www.youtube.com` was split inside `www` and `youtube` went to the wire
in one piece, which is the word the DPI matches on. Of six blocked names exactly
one got through: `youtube.com`, the one whose first label IS the blocked word.
That defect is fixed (815011dfb, efb2177f4). With it fixed the built-in presets
do the job the external process was brought in to do, and the process is 100 KB
of binary, a second procd service, a second UCI file, a port that agreed with
our egress by hand-written comment only, a readiness prober, a five-state
service model and a panel plate — all to work around fifteen lines of ours.

So this is not "ByeDPI turned out to be bad". It is a good tool that turned out
not to be needed, and the reason we thought it was needed was ours.

A CONFIG THAT STILL SAYS `type 'byedpi'` IS THE PART THAT NEEDED WORK. Nothing
is migrated and nothing is rewritten: the kind stays unbuildable, therefore
fail-closed — no outbound, no mark, no `ip rule`, no routing table, so every
node, group and rule bound to it is blocked rather than released onto the plain
WAN. A migration to `direct` was considered and rejected: it is the only rewrite
that leaves the egress routing at all, and it would silently turn a blocked
egress into a live plain-WAN path with the router's real address — by an
upgrade, on a config nobody touched. `CurrentSchemaVersion` is therefore not
bumped either: no stored field changes meaning, and a bump would only make this
build's configs unreadable to an older daemon for no gain.

What changes is what the operator is TOLD. `model.RetiredEgressTypes` is a
closed, positive table read by BOTH `ValidateEgresses` and the generator (one
copy of the sentence, because two copies drift). It names the removal, denies
that it is a typo, says nothing is built and that the traffic is blocked rather
than leaked, names the replacement (`direct`/`interface` with `dpi 'record'`),
refuses to promise which preset defeats a given ISP, and says `apk del byedpi`.
The generic "unknown type" is still there and still says something different, on
purpose: "we took this kind away" and "you mistyped something" send an operator
to different places, and a value that was correct on the day it was written must
not be reported as a spelling mistake. The type list stays closed and positive —
`interface`, `direct`, the alias `tunnel` — and `EgressTypeKnown` does NOT admit
the retired kind: being told it was removed and having it work anyway is worse
than either alone.

`Egress.Port` goes with the kind: no surviving egress dials anything, so the
option is no longer parsed and drains out of /etc/config/shater on the next
render, the same way the deleted per-group probe_url/probe_interval did.

Tests, verified by mutation, each failing by name:
  - drop the retired branch in `ValidateEgresses` -> the retired kind is
    reported as "is not one of interface/direct" and
    TestRetiredEgressTypeIsReportedByTheValidator fails on both spellings;
  - drop it in the generator -> "unknown type \"byedpi\"" and
    TestRetiredEgressTypeIsReportedByTheGenerator fails;
  - the FAIL-OPEN mutation, which is the one that matters: let `byedpi` fall
    into the `direct` arm and be a known type -> four tests fail, including the
    two that check no outbound is emitted. A removal that quietly starts routing
    the traffic it used to block, under a reassuring message, is the failure with
    the worst consequence;
  - the panel half: empty RETIRED_EGRESS_TYPES -> two egressEdit tests fail.
Controls beside the claims: `interface`, `direct`, the `tunnel` alias and the
empty synonym must still resolve, warn about nothing and emit an outbound
(TestSupportedEgressTypesAreUntouched), and never-supported values — `proxy`,
`block`, `wireguard`, `byedpi2`, `bye dpi`, `sorcery` — must NOT draw the
removal sentence, which names a replacement for something that never existed.

CI and docs: the feed loses its fourth package everywhere the four were named —
`apk upgrade shaterd shater-core luci-app-shater`, in CLAUDE.md, both READMEs,
INSTALL.md, the release body and `shaterd`'s own diag bundle. The version
exception (byedpi carried upstream's version, ours come from the git tag) is
gone with it, so ci/version.sh and ci/sdk-build-apk.sh no longer have an
exception to remember and the "expected >=4 of OUR .apk" collect check is now 3.
INSTALL.md §5.3 gains the half a feed cannot do: dropping the package from the
feed does not take it off a router it is already on, so `apk del byedpi` is
written down, with what it removes and why it is safe.

Panel: 368 tests -> 339. Deleted with the mechanism they covered:
byedpiReady.test.ts, byedpiAge.test.ts, byedpiRefusal.test.ts (34 tests);
egressEdit.test.ts gains 5 for the retired-type sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 17:13:50 +03:00
omarandClaude Opus 5 17843be5ad fix(tests): running the suite deleted the router's own state — four packages did it
shater/stats/store_test.go ended with

    _ = os.Remove(statsFilePath())

and statsFilePath() is not a test path. It is THE product path: /etc/shater/stats.db
on every host where that directory exists, which is the testbed and the router. So
`go test ./shater/...` deleted the accumulated query and connection log of whatever
machine ran it. The test passed. It had always passed — damage done by a test is a
side effect, not a wrong answer, and no instrument in this tree could see one.

A filesystem sweep (the new scripts/check-test-fs-isolation.sh: seed a router-shaped
canary tree in a container, run the whole gated suite, diff) found it was not alone.
Six packages, by measurement, not by reading:

  shater/stats    DELETED  /etc/shater/stats.db        (the line above; also
                           TestComboBackendSwitchSequence opened and pruned the
                           live DB, which the delete had been hiding)
  shater/logsink  DELETED  /etc/shater/shaterd.log and /var/log/shaterd.log —
                           New()/Reconfigure() purge BOTH product locations when
                           the file toggle is off, so Config.Path (which every test
                           here already set) never protected them. The daemon's own
                           log, the one an operator reads after an outage.
  shater/apply    DELETED  /var/run/shater.active — the ONE token hotplug and cron
                           check before touching the data plane. Clearing it on a
                           live router makes both stand down on a box that is up.
                           holdstate_test.go's `t.Cleanup(os.Remove(ActiveFlag))`
                           was not a cleanup; it was the delete.
  shater/panel    REWROTE  /etc/shater/stats.db — stats.NewStore("sqlite") from
                           TestStatsEndpointsAcrossBackends resolves the product
                           path too.
  shater/model    CREATED  /etc/shater/config.pre-v{0,1,2}.bak, config.pre-unreadable.bak
  shater/generate REWROTE  /etc/shater/cache.db

The last two are NOT fixed here — another agent is working in those trees. Both are
one TestMain away: model already has liveConfigPath/configBackupDir as vars, and
generate already has cacheFilePersistent; what leaks is product code (backupBeforeChange,
the engine's cache_file) called from tests that do not redirect them.

THE FIX is the seam generate/cache.go and generate/ruleset.go already use — the path
becomes a package-level var that only tests assign — plus, in each case, a test that
still pins the SHIPPED value, because an isolation that leaves the real decision
untested has only moved the defect:

  stats:   statsDirPersistent/statsFilePersistent/statsFileFallback + the exported
           SetPathsForTest (exported because shater/panel needs it from outside).
           New TestStatsFilePathPrefersPersistentDir covers both branches.
  logsink: PersistPath/TmpfsPath + a TestMain, since the hazard is in New(), which
           every test calls. New TestLogPathsAreTheShippedOnes.
  apply:   ActiveFlag + the existing TestMain. New TestActiveFlagIsTheShippedPath,
           which also records WHY /var/run: tmpfs, so a reboot clears it.

TestNewStoreSelection got stronger rather than weaker. Its "sqlite" case used to
accept "sqlite" OR "memory" because the real path might not open on this host — an
expected value that depended on the machine. At a private path there is no excuse:
a writable directory MUST report "sqlite", and a new control at an unopenable path
MUST report "memory" (the honest "persistence is not active" signal) without a crash.

TWO GUARDS, because one of them cannot see half of it:

  shater/testguard/fsisolation_test.go — parses every _test.go under shater/ and
  fails BY NAME when a filesystem-mutating call gets a path that is not PROVABLY
  temp-rooted. Positive and closed: what it cannot prove is a failure, not a
  default, which is the only rule that catches a path built by a function call.
  It follows local vars, closures, filepath.Join/Sprintf/+, helper parameters via
  their call sites, helper return values, and the save/override/restore idiom.
  Four waivers, each keyed on file+function+callee, each with the reason printed on
  every run, each a struct field traced by hand; a waiver that stops matching fails
  the test as STALE. Runs inside [2/7] and [4/7] — no new gate step, no new minute.
  Blind spot, stated: damage done by PRODUCT code a test merely calls (which is
  exactly logsink, model and generate above).

  scripts/check-test-fs-isolation.sh — the dynamic half, for that blind spot. It
  refuses to run outside a container unless told twice, because its method is to
  let the damage happen and then look, and it seeds/unseeds only what was missing.

Verified:
  - mutation, task 1: statsFilePath forced to the fallback -> the new path test
    fails ("with ... present = .../fallback-stats.db, want the persistent ...");
    newPersistent forced to memory -> "Backend = \"memory\", want \"sqlite\"";
    the fallback made to report "sqlite" -> "Backend = \"sqlite\", want \"memory\"".
    Green again after each revert.
  - mutation, the guard: the original os.Remove(statsFilePath()) put back -> named
    at store_test.go:154 with "the path comes out of statsFilePath(), which this
    check cannot follow"; a planted test writing "/etc/config/network" -> named as
    a literal path; the walk pointed at one package -> its own <150-file control
    fires ("reading a blank page"); a waiver matching nothing -> STALE WAIVER.
  - control, the sweep: with a planted violator it reports DELETED /etc/shater/stats.db
    and MODIFIED /etc/config/network; without it, those are gone and only the two
    foreign packages remain. Its bisect named shater/stats.TestComboBackendSwitchSequence
    on its own.
  - counts, declared vs executed (go test -list against top-level verdicts):
    stats 112/112, panel 121/121, apply 122/122, logsink 26/26, testguard 1/1,
    0 skips, 0 failures.
  - scripts/run-tests.sh: [1/7][2/7][3/7][5/7][6/7][7/7] green. [4/7] -race fails on
    four TestByeDPI* in shater/panel — a data race between byedpi.go's background
    probe and byedpi_test.go's forceByeDPIBinary cleanup, in another agent's
    uncommitted work (shater/panel/byedpi_cache_test.go is untracked). Proven not
    ours: a pristine HEAD tree carrying ONLY this commit's files passes -race over
    all 35 packages, and the same run with -skip ^TestByeDPI is green on the live
    tree too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 12:33:45 +03:00
omarandClaude Opus 5 56c9ea56e6 fix(gate): [4/7] reported a failure that did not exist — 664 s of pure sleep
The -race step failed with

    FAIL shater/netplane 600.019s
    panic: test timed out after 10m0s
      running tests: TestApplyIfaceSysctlsCoversRuleDivertedIface

over code that was neither hung nor wrong. Measured (golang:1.26, 32 cores):
shater/netplane is 1.971 s without -race and 663.762 s with it. A 337x factor
is not "-race is slower".

Nine of netplane's test files intercept nft/ip/ubus/uci/sysctl by re-exec'ing
the test binary as a no-op helper — the standard os/exec trick. Under -race
that child is ThreadSanitizer-instrumented, and TSan's atexit_sleep_ms DEFAULTS
TO 1000: every -race process sleeps a flat second before exiting, on no CPU.
~660 intercepted commands, one second each. The per-test times said so out
loud — 12.17 / 13.18 / 14.17 / 129.62 s — they were counting, not measuring.

Isolated, five runs each, of a `func main() {}` with nothing in it:

    built plain                       0.0014 s/run
    built with -race                  1.010  s/run
    built with -race, sleep disabled   0.008  s/run

So the children now run with GORACE=atexit_sleep_ms=0, set once in a package
TestMain rather than in each of the nine fakes (they all build the child env as
append(os.Environ(), ...), so one assignment covers the ones written later too).
TSan reads GORACE at process init, long before TestMain, so the detector of the
test process itself is untouched; only the children see it, and they do nothing
but write a canned string and exit. Proven, not assumed: a deliberate data race
in netplane is still reported under -race with this in place.

    shater/netplane   663.762 s -> 10.625 s   (203 === RUN and 128 top-level
                                               verdicts on both sides)
    shater/devices     28.412 s ->  0.358 s   (same disease, same cure)
    gate [4/7] end to end: was a 600 s timeout, now 56 s

WHAT THE GATE ITSELF WAS MISSING. A deadline and a failed assertion both exit
non-zero, and this script printed the same "FAILED [race]: go test exited 1"
for both — so the reader could not tell "the product is wrong" from "nobody
knows yet". [2/7]/[4/7] now name a timeout as a TIMED OUT, list the tests that
were still running, print only the goroutine dump instead of a quarter megabyte
of PASS lines, and spell out the two opposite fixes (a block, or slowness that
must be MEASURED first). Verified both ways: a sleeping test reads TIMED OUT, a
t.Fatal still reads FAILED.

The deadline stays at go test's own 10m, now written down with the measurement
beside it, and stays there as the hang detector — the slowest package under
-race is 18.9 s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 11:27:23 +03:00
omarandClaude Opus 5 e38108a7c4 test(gate): install iproute2 in the docker lane — without ip every slot is free
test / go + panel tests (push) Successful in 15m18s
release / test gate (push) Successful in 10m57s
release / apk aarch64_cortex-a53 (push) Successful in 5m52s
release / apk x86_64 (push) Failing after 28s
release / release apk (push) Successful in 6s
This change was already in the working tree when this session started; it is
committed here because it is load-bearing and an uncommitted load-bearing file
is a trap.

netplane.L3SlotFor asks the kernel through `ip link show` and reclaims through
`ip link del`. golang:1.26 ships no iproute2, so in the docker re-exec lane
every slot read as FREE, TestIntegrationL3StaleSlotIsReclaimed stood itself
down rather than pass while proving the opposite of what it claims, and [5/7]
then failed the gate — correctly, since this environment HAS root and
/dev/net/tun and the capability guard is therefore not what skipped it.

Installing it is also what made the concurrent-namespace defect visible at all
(see 06c04c157): with no `ip` on PATH, no `ip link del` was ever issued and the
two test binaries that were destroying shater/generate's TUN looked innocent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 03:25:58 +03:00
omarandClaude Opus 5 0f69880150 test(gate): a skipped test is a test that did not run — name it, or fail
Three holes, one shape: work that reads as coverage and is not.

1. shater/apply's TestApplyInstallsHoldWhenEngineFailsToStart — the only
   end-to-end test between "the engine died" and "the LAN forwards to the
   WAN in the clear" — asserted nothing. It broke the engine by pointing a
   rule-set at /nonexistent/nope.srs and stood itself down with t.Skip when
   that failed to break anything; it stopped breaking anything once
   LocalRuleSet.reloadFile began treating an unreadable file as empty.
   Measured in golang:1.26: the skip fired unconditionally and the package
   still printed `ok shater/apply`.

   It now injects the failure at the engineApply seam — the branch under
   test is applyLocked's, and a particular cause that stops causing retires
   the test silently — and COUNTS the seam calls, so applyLocked ceasing to
   go through it fails by name instead of quietly asserting something else.
   Everything else stays real: the model, generate, the kill-switch
   decision, netplane.RenderHoldNft, the latch, Status. New companion
   TestEngineApplyReallyFailsWithoutStarting is the control that the real
   engine.Apply can fail with the engine left stopped, so the simulated
   state is one this fork can be in.

   Mutation-checked both ways: drop the holdLocked call from applyLocked and
   the test fails with "0 holding planes were installed, want 1"; bypass the
   seam and it fails with "the engine-swap seam ran 0 times, want exactly 1".

2. warnings_test.go had two of the same genre. The len(genWarnings)==0
   t.Skip is now a t.Fatal — an unloadable blocklist must always warn, and a
   generate that stops saying so is the W7 regression, not a reason to stand
   down. TestStatusWarningsAlwaysNonNil pins readConfig itself: its
   "zero warnings" assertion was true on a build host only because the
   config read failed SILENTLY, so once that failure started publishing a
   critical warning the same line meant two different things in two
   environments.

3. The gate could not see any of it. It now runs the suites with -v and
   matches every `--- SKIP` against SKIP_DECLARED; an undeclared skip fails
   BY NAME, a declared one prints its reason on every run. check_skips
   proves its own instrument first (no `=== RUN` line => the check was
   reading a blank page), and it also reports on a suite that failed
   elsewhere, so a red tree cannot become a hiding place. -v costs no test
   time (38/25/24 s plain vs 38/24/24 s, warm) — only output, which is
   filtered on a green run.

Also closes the same hole one language over: [6/7] requires every non-Go
test file in the tree to be claimed by a named runner, and [7/7] runs the
ones this gate owns with a verdict by name. openwrt/luci-app-shater/tests/
status-readout.test.js — 24 assertions over the one screen an operator
reaches while the LAN is cut off — was executed by nothing at all, and
[1/7] could not report it because `go list` is its instrument. The non-Go
suites run on the HOST before the docker re-exec, so the local loop really
executes them rather than printing "did not run" every time; where there is
no node at all they are named and the notice replaces the closing banner.

Controls, all run and reverted: a planted t.Skip is caught and named; a
planted failing .test.js is caught and named; an unclaimed test file is
caught and named.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:30:39 +03:00
omarandClaude Opus 5 164b703a7d feat(l3): ping travels the tunnel by default, and every LAN zone can reach it
l3_tunnel was opt-in, and "off" had no honest win left in it. Off, a LAN ping
is decided by `untunnelable` alone and every rung is a drop (block) or a
disclosure (icmp/direct send the echo out of the WAN with the client's real
address). "Ping works" was never the state where ping was tunnelled — it was
the state where ping was leaking. On, an L3-capable outbound carries the echo
and one that is not drops it honestly: adapter.JudgeFlow returns ActionDrop for
an ICMP flow whose outbound is not a tun.Port, so no reply is forged. The price
is a standing TUN + gVisor netstack, ~2 MB RSS, and it is stated where the
option is.

The switch stays. It is a real answer on a 32/64 MB device and when bisecting
whether the L3 ingress is what broke a box — but it is now a WARNED answer:
ValidateGlobals says what the off state does to ping and names the policy that
takes over. Two combinations also changed meaning and are now reported:
untunnelable=icmp is no longer "block plus working ping" (the prerouting L3
mark claims every ICMP packet before the forward chain the echo accept lives
in, and a LAN host's ICMP errors are marked in with them and dropped in the
TUN), and the existing =direct report gains a sibling rather than standing
alone.

The fw4 seeding was the second half of the same problem. The divert set spans
every LAN inbound and every iface:/zone: rule source, but 30_shater-core seeded
a forwarding into shater_l3 for `lan` only — so on a multi-zone router ICMP
from the other zones is marked, routed, accepted by `inet shater`, and dropped
by fw4's zone policy with nothing in any log. Every zone gets a forwarding now,
guarded by a scan of the actual src/dest pairs so a re-run adds nothing. Every
zone including an uplink, because guessing which zones hold clients is wrong
somewhere and a superfluous entry authorises nothing: accept_to_shater_l3 is
`oifname "shater-l3*" accept`, and the only thing that routes a packet into
that device is our own fwmark rule.

scripts/testbed-lao.sh builds the second LAN zone this needs to be visible at
all. It is not installed by the package — that is the whole opt-in mechanism.

Verified on local_openwrt (ImmortalWrt 25.12.1 r37978): three runs of the
seeder leave exactly one forwarding per zone (lan/wan/lao) and no existing
section altered; deleting the lao forwarding removes `jump accept_to_shater_l3`
from chain forward_lao and re-seeding restores it; with the idempotency guard
disabled two runs produce nine forwardings instead of three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-27 00:26:00 +03:00
omarandClaude Opus 5 76da5134ef test(gate): the two tests that need a kernel may not skip in silence
The L3 branch adds TestIntegrationL3TunInboundStarts and
TestIntegrationL3EgressICMPIsAFlow — the only tests that prove the engine
really opens shater-l3 and that the egress outbound really is a FlowOutbound.
Both need root plus /dev/net/tun, both guard themselves with t.Skip, and the
gate could not see either: `go test` prints `ok <pkg>` whether a test ran or
skipped, so [2/5]'s per-package `ok` check is satisfied and the gate closes by
claiming it "passes every test we own". That is this script's own founding
failure (115 of 116 test files never running while CI stayed green) one level
down, and it would have shipped invisibly.

Two halves.

Where the capability CAN be granted, grant it. From a non-linux host the gate
re-execs into a container; that container now gets --cap-add NET_ADMIN and
--device /dev/net/tun, probed rather than assumed, so a plain
`scripts/run-tests.sh` on a dev box actually exercises the kernel path instead
of quietly stepping over it.

Where it cannot, say so where it cannot be missed. The act_runner is an LXC
guest whose kernel has no tun module at all (checked on 10.10.10.211:
`modprobe tun` -> "Module tun not found", /dev/net does not exist, act_runner
runs job containers with privileged:false and no container.options), so the
device cannot be handed down without reconfiguring the Proxmox host. New step
[5/5] therefore DISCOVERS every ^TestIntegration under the fork's trees — no
hand-kept list, so a privileged test written next month joins on the day it is
named — runs them with -v, and demands a verdict for each BY NAME: RAN, or
FAILED/MISSING (fatal), or SKIPPED while the environment could have run it
(fatal, because the capability guard cannot be what skipped it), or skipped for
a reason this box genuinely has — which replaces the closing banner, so the
last line of the gate can never claim coverage it does not have.
SHATER_REQUIRE_PRIVILEGED=1 makes that last case fatal for runs that can.

The discovery call carries -ldflags for the same reason every other call does:
`go test -list` links each test binary, and without -checklinkname=0 every
package pulling common/badtls fails to link. The first cut of this step omitted
it, swallowed the error, and printed "none declared" — a check against silent
skipping that was itself silently skipping. Its exit status is now inspected
and an empty list is only ever reported after a successful enumeration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BHw89tdWddzhjUc4bAH4tS
2026-07-26 18:50:46 +03:00
omarandClaude Opus 5 4078334d85 fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health
board only ever inserted — the delete exists but no path in this fork calls it —
and it lives on the engine context, so it outlives every generation. Its keys are
node tags, and providers rename nodes on each subscription refresh: about 440k
keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The
stats aggregator's server and outbound counters were the only ones with no cap,
no prune and no top-N, and one of them was handed to the panel whole on every
poll.

They are bounded now, evicting least-recently-seen, with numbers argued from this
box rather than round: the board holds 4096 against a live generation of about
1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped
silently — the same rule the log sink already follows — and a new Dropped section
in the snapshot reports all six bounded aggregates, including the three that had
been evicting without saying so.

Snapshot did O(devices × domains) under the aggregator lock, sorting five
thousand entries to show fifteen, and could read the DHCP lease file from inside
it. Meanwhile the event subscribers have 64-slot buffers that drop without a
counter, so an open Overview page cost the query log real rows. Selection is
top-K now — proven byte-identical to the old sort over 200 random trials — and
both the lease read and the row ordering happen outside the lock.

The panel server had one timeout, on headers. An unauthenticated client could
hold a goroutine, a socket and a descriptor forever by sending its body one byte
at a time; a stopped reader on the log stream held the handler, the pipe and a
child process that outlived the request. Every phase is bounded now, with the
unauthenticated route on a tighter budget than the rest, and the log stream
renewing its deadline per chunk so a slow-but-reading client is never truncated.

And the last of the detour transports: each call built a fresh one, and the alert
delivery path dropped it, pinning keep-alive sessions through the engine's own
outbounds for 90 seconds — eighteen times the budget a retiring generation gets.

The race skip is gone from the gate. The test it existed for raced in its own
clock, not in the product; that is fixed, so nothing is excluded under -race any
more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:40:21 +03:00
omarandClaude Opus 5 56a276bcc1 ci: make the tests a gate instead of a decoration
The fork had a full suite and no CI that ran it. Upstream's test workflows
trigger on stable/testing/unstable; this repo only has main. And Gitea does not
read .github/workflows at all once .gitea/workflows exists, so those files were
decoration here. 115 of the 116 test files under shater/** had never executed in
CI even once, which is how TestDNSFilterRemoteBlocklistHTTPClient stayed red
across two published releases without anyone noticing.

The gate is a job inside release.yml that build-apk needs, because a separate
workflow cannot block another one. It runs the suite under the shipped tag set,
on Linux — 6 of 7 test files in transport/wireguard and 12 in shater/generate
compile only there or only under those tags, and those are exactly the files
covering AmneziaWG.

Three guards stop it from passing by running nothing, which is the failure this
whole change is about. The tag set may only ADD test files, never remove one.
Every package go list says has tests must appear as "ok <pkg>" in the output, so
a suite that collapses to "no test files" fails instead of passing. And the
panel run counts its test files first, because node --test exits 0 with "pass 0"
when the glob matches nothing.

The publish step used to exit 0 having published nothing: its assertions all
live inside a loop over artifacts, so an empty directory ran the body zero times
and reported success. It now counts what it published and fails on zero.

Verified by extracting the shipped step text and running it against stubs: empty
artifacts gives exit 0 before and exit 10 after; the rolling-release readback
still fires its own exit 14.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 04:53:14 +03:00
omarandClaude Opus 5 1746d4d0ef fix: stop the panel and the shipped binary from lying about what works
release / apk aarch64_cortex-a53 (push) Successful in 9m13s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 7s
Four defects, all found by the owner on the live router, all of the same
family: something declared itself working while it was not.

WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave
"create WireGuard device: gVisor is not included in this build". The router
tag set carried with_wireguard and with_awg but not with_gvisor, so
sing-tun compiled its stub instead of the netstack every WireGuard device
needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving
requirement", so this was a broken promise, not a trim.

The tag itself was the small half. The tag set was the ONE build
configuration nothing in the repo tested: TestAmneziaWGEndpoint passes
because tests build with the full upstream tags. So the set now lives in
one file (scripts/router-tags.sh) and two guards hold it to the feature
list -- a static check that needs no tags, no Linux and no network (so the
next such gap fails on the developer's machine), and a behavioural one that
constructs every declared protocol through box.New UNDER THE SHIPPED TAGS,
where skipping is forbidden. Removing the tag now fails with the feature
name, the missing tag, and why: "Either add the tag back, or stop declaring
the feature -- those are the only two honest options." Cost: +2.8 MB raw,
+0.6-0.7 MB packed per arch. D23; D9 corrected.

THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came
from plane === 'full', which reports whether the data plane is installed --
nft table, policy routing, live engine -- and says nothing about where the
traffic goes. On a config with one `default -> direct` rule and no groups
the plane is fully installed and every packet leaves in the clear, so the
worst possible state rendered as the reassuring one.

The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the
moment they reach the engine, not from the model: buildRoute changes the
answer (a scheduled rule outside its window is never emitted, only the last
condition-less rule reaches Final, an unresolved target is rewritten by
ruleKillFallback), and re-deriving it anywhere else is a second
implementation that will drift -- model/reachability.go exists because two
already did. Four verdicts, not three: `blocked` is separate because under
a closed kill-switch with no catch-all nothing leaks, and calling that
"going out directly" is a lie in the alarm direction. Rider: Overview's
defaultTarget printed the highest-Order enabled rule as the default; a rule
becomes Final by having no conditions, whatever its Order.

"PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE
(B15). Once the browser suppresses dialogs, window.confirm returns false
immediately, so all 15 confirmations across 7 pages read as "cancelled" and
silently did nothing, with no way to recover from inside the panel. Replaced
with an in-app dialog the browser cannot mute: focus trapped and parked on
Cancel, Esc and veil cancel, focus returned to the opener, crit styling for
destructive commits. useConfirm() throws if the provider is missing rather
than falling back to a quiet false -- the failure mode being fixed.

HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed,
so a feed's nodes of those types vanished. The real landmine was one layer
up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so
without schemePrefixes the links were gone before any parser ran and the
fix would have looked complete. Undeliverable parameters are refused when
the node cannot work or would be less secure than the link asked (obfs,
pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it
survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot
produce a QUIC TLS config, and that fails at dial time, not at box.New.

Also: nodes added by hand can be named and renamed. The name is the
outbound tag, so a rename rewrites every reference in one PUT -- rule
targets, group members, chain hops, detours -- in the spelling each already
uses, and is refused outright when a group answers to the same bare name.
Subscription nodes state why they cannot be renamed instead of hiding the
control.

go build, go vet, go test ./shater/... (13 packages), panel npm run build
and npm test (13/13) all green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 20:08:58 +03:00
omarandClaude Opus 5 a8f2b0f068 ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.

ci/version.sh is now the single source of truth. It derives the version
from `git describe`:

    tag `vX.Y.Z`   -> PKG_VERSION=X.Y.Z  PKG_RELEASE=1
    off-tag build  -> nearest tag + PKG_RELEASE=<commits since it> + 1
    no tag/no git  -> 0.0.0-r1 (below everything ever published)

Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.

The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.

The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.

byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.

Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:32:38 +03:00
omar 53cbdc75f1 build(shaterd): drop with_dhcp from the router tag set
shater resolver types are udp/tcp/doh/dot/local/fakeip; a dhcp:// DNS
transport is never generated, and the slim shater/registry never
registers the transport, so the tag gated nothing in this binary.

D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
2026-07-23 10:36:59 +03:00
omar 0e5b1afb80 build(shaterd): drop with_clash_api from the router tag set
The admin panel is shater's own web server and generate never emits a
clash_api service (shater/engine/engine.go pre-registers its own
dnstrack.Manager precisely because no api/clash_api observer exists on
the router). With include.Context gone the Clash server was already out
of the link; dropping the tag records the decision. Desktop/CLI LX_TAGS
keeps with_clash_api for external dashboards.

D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
2026-07-23 09:31:42 +03:00
omar d291c90cab build(shaterd): drop with_gvisor from the router tag set
The shater data plane is tproxy/redirect (netplane); generate never emits
a tun inbound, so the userspace gvisor netstack is unreachable code. With
the slim registry it was already dead-code eliminated by the linker —
dropping the tag makes the intent explicit and stops compiling ~3.6 MB of
gvisor sources into the build at all. A future tun inbound would fall
back to the system stack; re-add the tag if that ever lands.

D9 in DECISIONS.md and the INSTALL.md tag block updated to match.
2026-07-23 09:31:14 +03:00
omarandClaude Opus 4.8 a5c74209e4 feat(openwrt/shaterd): release build + prebuilt shaterd package (Phase 8 ship)
Makes the whole product installable — shater-core DEPENDS +shaterd, and
this is what resolves it.

- scripts/build-shaterd.sh: the release build. Builds the panel SPA
  (npm ci && npm run build), copies panel/dist -> shater/panel/webroot
  (the go:embed dir), cross-builds shaterd for amd64 + arm64 with the D9
  router tag set (CGO_ENABLED=0, -checklinkname=0 -s -w, static ET_EXEC no
  PT_INTERP), then UPX --lzma --best (D10) and stages the .upx into
  openwrt/shaterd/files. Version from arg/SHATER_VERSION/git-describe.
  Measured: amd64 40.3MB->10.4MB, arm64 37.6MB->8.5MB.
- openwrt/shaterd: prebuilt-binary package (npm+embed+UPX don't reproduce
  cleanly in the SDK, so CI stages the artifact). Maps OpenWrt ARCH
  (x86_64->amd64, aarch64->arm64 = both BPI routers) to files/shaterd-<a>.upx,
  installs /usr/bin/shaterd. RSTRIP/STRIP disabled (the SDK strip would
  corrupt the UPX binary); DEPENDS empty (static); errors clearly when no
  artifact is staged. GPL-3.0-or-later.
- docs-shater/INSTALL.md: build + install order (shaterd -> shater-core ->
  luci-app-shater, optional byedpi) + enable/apply.
- gitignore: dist/shaterd-*, openwrt/shaterd/files/*.upx, panel webroot.

Verified on the OpenWrt musl VM: dist/shaterd-amd64.upx (10.4MB) decompresses
into RAM + runs (shaterd status OK), serves the REAL embedded Faceplate SPA
at :8088 ('SPA embedded=true', real Vite index.html + assets — not the
placeholder).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LLthkP2S8WAfxu7fcYbPfE
2026-07-16 00:18:39 +03:00