a0f6083e2822bcc5e493c5dfeff9f4fc9b12aa1f
3027
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a0f6083e28 |
fix(panel): show whether a rule is in force, not just what was saved
With two catch-all rules both enabled in UCI and a WAN profile enabling one and disabling the other, the panel drew BOTH switches on while the engine ran only one chain. GET /api/config is right to return the raw model — that is the desired state the panel PUTs back — but Routing.tsx read the row state and the active count from it too, so the interface claimed a setting was in force when it was not. Same defect class as the Protected badge. /api/rules/reachability now carries the effective flag and, where the active profile changed the outcome, its name and direction. The annotation is a DIFF of ApplyProfileRuleOverrides output against desired state rather than a second reading of the profiles name lists, so profile logic is not duplicated and cannot drift — an unmigrated rule the profile is forbidden to enable produces no diff and gets no badge, with nothing here needing to know about LegacyDst. In the UI the two states stay separate: the switch remains the only carrier of desired state and still writes UCI, while the effective state drives the dimmed row, the badge, the banner and the header count. Mirroring the effective state into the switch would be worse than the original bug — the operator would be toggling someone elses control, and the profiles decision would be written back as their own choice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGNapk-v0.2.12-aarch64_cortex-a53 apk-v0.2.12-x86_64 v0.2.12 |
||
|
|
77369aedfe |
fix(logsink): collapse repeated lines instead of erasing the routers syslog
A broken outbound makes the engine repeat one line about once a second — 370 copies in six minutes. The routers syslog ring holds ~760 lines, so within minutes it evicts the history of every other subsystem and our own startup lines with it. Diagnosing the WireGuard duplication above required restarting the service purely to catch the first seconds of a boot. Collapse runs into "last message repeated N times". The comparison key is level + text with the uptime field dropped: comparing whole lines would suppress only same-second bursts, because that counter ticks. The per connection "[id duration]" group is deliberately KEPT in the key — those ids are distinct connections, and folding "50 connections failed" into one count would be a worse lie than the flood. Window 5s, so a standing fault keeps being reported instead of looking like a frozen log. fatal/panic are never suppressed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
515ae6d1b7 |
test(generate): make the remote-blocklist test exercise the remote path
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two releases, which made the whole package exit non-zero no matter what the code did — a real regression would have drowned in the familiar red. The cause is not the packages no-network fetcher stub, as it first appears. ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so the fixture fell into the TEXT-list path: downloaded by generates own fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which every assertion below then contradicted. No stub content could fix that; the stub decides the lists contents, not the rule-sets type. Give the URL the .srs suffix the test always meant it to have, so the engine fetches the compiled set itself through the direct outbound. No assertion is weakened and the no-network stub stays in place. Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags): 338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
a2ffbb1292 |
fix(generate): one WireGuard device per private key
A node may be copied freely by this package: a per-chain hop copy and a per-group egress copy are rebuilt from the share-link so each can carry its own Detour. For vless that is right — a copy is another TCP client. For WireGuard it is not: each emitted endpoint is a real device holding the nodes private key, and a peer keeps exactly ONE session per public key. Two devices from one key evict each other continuously, and with keepalive on both the loop never settles: NEITHER passes traffic. buildOutboundsAndEndpoints emits the base endpoint for every enabled node whether or not anything references it, so a WG node used only as a chain hop always produced two devices. That is what any chain containing a WG node looks like — every such chain was permanently dead. Observed on the box: two UDP sockets from shaterd to the same peer port, the servers peer endpoint flapping between them, +32 bytes/min through the tunnel and every hop failing with "context deadline exceeded". Deduplicate once on the assembled options, which catches all three producer paths by construction. Duplicates are DELETED, not merely unreferenced: box.New starts every endpoint regardless of reachability, so a leftover would still bring its device up and still fight for the session. Dangling references go to block, never to direct — a consumer whose tunnel just disappeared must stop, not fall out onto the plain WAN. Subscription fetch detours seed the reachability walk (they are direct references like any rule), mirroring engine.ViaToTag exactly, with a tripwire test against drift. A config that genuinely needs two devices for one key keeps one and fail-closes the rest with a critical warning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
1746d4d0ef |
fix: stop the panel and the shipped binary from lying about what works
Four defects, all found by the owner on the live router, all of the same family: something declared itself working while it was not. WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave "create WireGuard device: gVisor is not included in this build". The router tag set carried with_wireguard and with_awg but not with_gvisor, so sing-tun compiled its stub instead of the netstack every WireGuard device needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving requirement", so this was a broken promise, not a trim. The tag itself was the small half. The tag set was the ONE build configuration nothing in the repo tested: TestAmneziaWGEndpoint passes because tests build with the full upstream tags. So the set now lives in one file (scripts/router-tags.sh) and two guards hold it to the feature list -- a static check that needs no tags, no Linux and no network (so the next such gap fails on the developer's machine), and a behavioural one that constructs every declared protocol through box.New UNDER THE SHIPPED TAGS, where skipping is forbidden. Removing the tag now fails with the feature name, the missing tag, and why: "Either add the tag back, or stop declaring the feature -- those are the only two honest options." Cost: +2.8 MB raw, +0.6-0.7 MB packed per arch. D23; D9 corrected. THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came from plane === 'full', which reports whether the data plane is installed -- nft table, policy routing, live engine -- and says nothing about where the traffic goes. On a config with one `default -> direct` rule and no groups the plane is fully installed and every packet leaves in the clear, so the worst possible state rendered as the reassuring one. The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the moment they reach the engine, not from the model: buildRoute changes the answer (a scheduled rule outside its window is never emitted, only the last condition-less rule reaches Final, an unresolved target is rewritten by ruleKillFallback), and re-deriving it anywhere else is a second implementation that will drift -- model/reachability.go exists because two already did. Four verdicts, not three: `blocked` is separate because under a closed kill-switch with no catch-all nothing leaks, and calling that "going out directly" is a lie in the alarm direction. Rider: Overview's defaultTarget printed the highest-Order enabled rule as the default; a rule becomes Final by having no conditions, whatever its Order. "PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE (B15). Once the browser suppresses dialogs, window.confirm returns false immediately, so all 15 confirmations across 7 pages read as "cancelled" and silently did nothing, with no way to recover from inside the panel. Replaced with an in-app dialog the browser cannot mute: focus trapped and parked on Cancel, Esc and veil cancel, focus returned to the opener, crit styling for destructive commits. useConfirm() throws if the provider is missing rather than falling back to a quiet false -- the failure mode being fixed. HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed, so a feed's nodes of those types vanished. The real landmine was one layer up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so without schemePrefixes the links were gone before any parser ran and the fix would have looked complete. Undeliverable parameters are refused when the node cannot work or would be less secure than the link asked (obfs, pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot produce a QUIC TLS config, and that fails at dial time, not at box.New. Also: nodes added by hand can be named and renamed. The name is the outbound tag, so a rename rewrites every reference in one PUT -- rule targets, group members, chain hops, detours -- in the spelling each already uses, and is refused outright when a group answers to the same bare name. Subscription nodes state why they cannot be renamed instead of hiding the control. go build, go vet, go test ./shater/... (13 packages), panel npm run build and npm test (13/13) all green. NOT yet verified on hardware. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGNapk-v0.2.11-aarch64_cortex-a53 apk-v0.2.11-x86_64 v0.2.11 |
||
|
|
f86501bf77 |
ci!: drop the opkg lane — apk only, and fix the stale rolling release
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has no `opkg` binary at all. The 24.10 lane was building and signing a feed no device could consume. Removed jobs `build` and `release` with the scripts only they called (ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh) and the usign trust anchor dist/shater-feed.pub. A committed public key is an instruction: it invites the old install path for a feed that is no longer produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD still holds the secret half, and a usign secret contains its own public half, so the identity is reconstructible if a 24.10 device ever needs serving. D7 is marked SUPERSEDED by the new D22 rather than deleted. Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from 2026-07-24 while every tag run published its versioned release correctly. The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest (workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run never touched the rolling pointer. Asset replacement was never the problem; ci/gitea-release.sh already deletes before recreating. A router pinned to the rolling URL sat on 0.2.0 while `apk update` reported success: silent staleness, the failure mode this repo keeps having to close. The rolling pointer is now published on EVERY run, tag runs included, and a new assert reads the release back over the API afterwards: our three tag-versioned packages at the built version plus the index and the key must be present (exit 13), and no package asset at any other version may survive (exit 14). Same class of check as sdk-build-apk.sh's package-version assert, added for the same reason -- the previous failure mode was silent. KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references it. Docs state plainly that mini_router is deliberately pinned to a versioned URL and that the hand-edit per release is the price of pinning. Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can no longer install our packages. Its 25.12 rebuild is in flight separately. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
eccfc6136c |
fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review ofapk-v0.2.10-aarch64_cortex-a53 apk-v0.2.10-x86_64 v0.2.10 |
||
|
|
244b7c4199 |
feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are gone from api.ts, so the Routing page loses the two controls that wrote them. The add form's Match picker (rulesets / ip / port) collapses to a plain Port(s) field beside the ruleset checkboxes — with no inline address list there was nothing left to choose between. The edit form drops its "Domain(s) — legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form shows, which is the honest shape of a rule that carries one destination mechanism. The destination picker renders even when the config has no rulesets yet, and says where to get one. Hiding it (the old behaviour when the list was empty) would leave the rule form with no destination control at all, at precisely the moment the user needs to know one exists. It is checkboxes and nothing more: creating and filling a list stays in the Rulesets panel, so a list is authored in one place and its naming and entry rules cannot drift between two editors. isCatchAll() drops the same two fields as model.IsCatchAll, so the "never applies" badge and the daemon's apply warning keep agreeing about which rule is the default; the matcher chips lose their `dns` and `ip` rows for the same reason. The mock backend's reachability shim follows. Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the removed wide inputs used, is deleted rather than left dangling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>apk-v0.2.9-aarch64_cortex-a53 apk-v0.2.9-x86_64 v0.2.9 |
||
|
|
a8ef887c56 |
feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.
`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.
THE BARE-ENTRY TRAP, and why the migration is not a copy
A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).
`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.
`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.
THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)
Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.
Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.
ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.
Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.
untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.
Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8fd5c52488 |
fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.
Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.
With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.
Fix (upstream file, lx:tproxy_writeback_connect):
* CONNECT the write-back socket to the one peer it ever talks to. The kernel's
compute_score() rejects a connected socket for any other peer, and a
connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
reuseport selection outright — so it can no longer be handed a datagram it
will not read. Nothing about the reply changes: same spoofed source, same
single peer, Write instead of WriteTo. The unconnected path is kept verbatim
for a destination that cannot be bound (domain socksaddr).
* A failed cached write now CLOSES the socket instead of only dropping the
reference (upstream left the fd to the GC finalizer).
* TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
but the cache evicts lazily, so after the inbound is gone nothing wakes the
live sessions and each strands its write-back socket. Invisible upstream
(one close at shutdown); on this fork the engine is rebuilt on every apply,
so it was one stranded generation per apply.
Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.
The netplane UDP:53 conntrack flush from
apk-v0.2.8-aarch64_cortex-a53
apk-v0.2.8-x86_64
v0.2.8
|
||
|
|
32e8f8ff0b |
fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.
Cause: `restart` is not synchronised end to end.
* procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
`stop; start`, and the `service delete` ubus call returns as soon as
SIGTERM has been SENT. `start_service` therefore re-adds the instance
(and runs `shaterd migrate`) while the outgoing `shaterd run` is still
executing its honest teardown.
* The successor's only defence was `daemonAlive()` -> exit(1), leaning on
procd's `respawn 3600 5 0` to try again five seconds later. That is a
blind retry, not synchronisation: it neither knows nor waits for the
teardown, and it turns every restart into a logged crash plus a
five-second hole with no data plane.
* `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
engine.Close of a several-hundred-outbound box flushes cache.db to
flash before the netplane teardown even starts — aborting the teardown
at an arbitrary point and leaving the plane HALF removed.
* Nothing in the tree ever touched conntrack, so flows that crossed one
of those windows kept entries formed against a plane that no longer
exists. For UDP there is no handshake to resynchronise on and every
retry merely refreshes the entry, so the flow stays wedged for as long
as the client keeps asking — a flow-scoped, permanent failure that no
ruleset rebuild can reach.
* RoutingPresent() reported "plane intact" from the ip RULE alone, while
ApplyRouting installs a rule AND a `local default dev lo` route removed
by two independent commands. A teardown interrupted between them was
therefore invisible, applyLocked's fast-path skipped ApplyRouting
forever, and no reconcile could repair it.
Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):
* init: `start_service` waits for a live predecessor pidfile to clear
before opening the instance, so restart == stop + pause + start. Zero
cost at boot. term_timeout 10 -> 30 so an honest teardown is never
killed halfway.
* daemon: the single-owner guard WAITS for the predecessor (bounded,
60s) instead of exiting 1; it still refuses if the budget expires.
* netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
a blanket flush would drop the admin's own SSH/LuCI sessions) called
on every plane transition: after a ruleset loads, after the table is
removed, and once more in applyLocked when the whole plane (table +
policy routing + sysctls) is assembled.
* netplane: RoutingPresent() now verifies both halves it installs.
Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.7-aarch64_cortex-a53
apk-v0.2.7-x86_64
v0.2.7
|
||
|
|
c562579ef3 |
docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute does not emit a condition-less rule as a match-all route rule: it sets route.Final and continues, so the LAST condition-less rule by order wins, and it can never shadow a rule that has conditions (those are emitted ahead of Final regardless of order). For the config on the router this inverts the conclusion: traffic IS going through the proxy (order=100 -> group:auto is the live default) and the dead knob is the order=20 `direct` one. Severity downgraded from high to medium accordingly — a dead setting, not a leak. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6f89acbae7 |
feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final", so a config with two of them showed two identical claims and no hint that only the last one is the default the router uses. A superseded rule now loses those marks — it keeps its real Order in the rail instead of the "·" that means final — and gains a "never applies" badge plus a line naming the rule that beat it and what to do about it: give this one a condition, or delete one of the two. Warn semantics throughout (--amber, dashed frame, dimmed target chip): orange is the ACTIVE state on this faceplate, and a rule the router ignores is the opposite of active. Verdicts come from GET /api/rules/reachability and are keyed by the rule's index in Rules, never by name — the config that prompted this had two rules both called `default`. They are re-fetched after every save, and a verdict whose echoed name/order no longer matches the row is dropped rather than shown, so the window between an optimistic edit and the refetch cannot badge a working rule. Rule rows were also keyed by name in React, which silently collapses two rows that share one; the key now carries the model index. The mock fixture gains a second condition-less rule so `?mock` renders the state, and mock.getRulesReachability derives its verdicts from the live fixture config rather than hard-coding them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88a82c7297 |
feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:
* two condition-less rules retire each other, and the LAST one by Order
wins, so an earlier "default -> direct" is dead while looking live;
* a condition-less rule can NEVER retire a rule that HAS conditions —
those are emitted ahead of Final whatever their Order.
A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.
model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.
Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.
Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.
Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.
GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.
Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a8f2b0f068 |
ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.
ci/version.sh is now the single source of truth. It derives the version
from `git describe`:
tag `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
off-tag build -> nearest tag + PKG_RELEASE=<commits since it> + 1
no tag/no git -> 0.0.0-r1 (below everything ever published)
Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.
The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.
The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.
byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.
Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0b32a6d58b |
fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:
daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
[ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...
`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.
Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:
* control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
DisableColors at its false zero value. It now comes from
controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
lives in an untagged file (same split as profilewatch.go) so it is
unit-testable off the linux target.
* engine (shater/generate): the generated option.LogOptions never set
DisableColor, so box.New built a colouring formatter over the shared
sink. logOptions() now sets it from the same TTY gate (seam:
logColorAllowed).
logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.
The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.
Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
893fdc500c |
fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as
|
||
|
|
02c266188f |
docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the previous install, install from the signed apk feed, verify the default state, restore a working config with 315 subscription nodes, then exercise the data plane, panel API, config lifecycle, resilience and DNS. 74 PASS. Findings (detailed separately): two catch-all `default` rules where the first sends all unspecific traffic direct and makes the second unreachable; `shaterd nodes` is a stub returning [] while usage promises the node list; DNS to the router LAN address dies after `service shater restart` (stop+pause+start is fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI colour codes reach syslog. Also records the four-iteration CI hunt that ended in the green apk lane. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
024e9308c9 |
fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.
The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:
config PACKAGE_kmod-mlx5-core
tristate
default m
No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.
The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.
Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.6-aarch64_cortex-a53
apk-v0.2.6-x86_64
v0.2.6
|
||
|
|
eab1db2c3f |
fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (
v0.2.5
|
||
|
|
22d7161c08 |
fix(ci/apk): disable the SDK's ALL/ALL_KMODS mass-select
release / aarch64_cortex-a53 (push) Successful in 3m25s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m3s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.
The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:
config ALL_NONSHARED ... default ALL
config ALL_KMODS ... default ALL
config ALL ... default y
In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.
Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.
Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.
Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v0.2.4
|
||
|
|
24c5a1615d |
fix(ci/apk): build .config from scratch — stop packing all 3593 kmods
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m18s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m2s
release / release (push) Successful in 7s
release / release apk (push) Successful in 5s
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn — none of which we ship) before it ever got near our four. Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the .config that ships inside the ImmortalWrt SDK tarball. That file is the buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so `make defconfig` re-selected every kernel module of the target as =m and package/kernel/linux/compile — pulled in via shater-core's nft kmod deps — packed the lot. Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds": start the .config EMPTY so kconfig can only pull in what our packages actually select. Carried over from the SDK's .config, nothing more: the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a wrong guess here means silently cross-compiling for another arch), CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel). Also adds the diagnostics this lane never had, since a failed run leaves a 27 MB log: the carried-over identity lines, the post-defconfig kmod count and target readout, a hard check that all four of our packages survived defconfig, an abort if the kmod count is back in the hundreds, and du/df after compile. opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR, CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>v0.2.3 |
||
|
|
4492f0599c |
ci: harden feed/packaging shell scripts
release / aarch64_cortex-a53 (push) Successful in 7m3s
release / x86_64 (push) Successful in 3m13s
release / apk aarch64_cortex-a53 (push) Failing after 4m34s
release / apk x86_64 (push) Failing after 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in the signed Packages index can't mask an empty SHA256. Script survives -u (all vars use :? or :- defaults). - .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a trap-based tmpdir cleanup, and set -euo pipefail. bash -n clean on both. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>v0.2.2 |
||
|
|
bc2b53069a |
fix(security): audit remediation — file perms, const-time auth, leaks, CSPRNG
Backend audit fixes (upstream-file edits wrapped in // lx: markers): - experimental/libbox oom_report.go/report.go: OOM reports + configuration.json (server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files → 0o700 / 0o600. [sec-perms] - daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare. [sec-consttime] - service/oomkiller/timer.go: network-extension cleanupTriggered logic was inverted, so FreeOSMemory was never called after a trigger; flip both assignments so a trigger schedules the deferred free and the next poll runs + clears it. [sec-oomcleanup] - transport/v2rayxhttp/client.go (lx-native file): session id used math/rand → crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface. - daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a goroutine + the ssh-agent fd on every closed session (second io.Copy blocked on an idle agent Read forever); tie both copies + the session ctx to a cancel that closes both ends. [sec-sshagent] - daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to 1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate] - route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in select while stopIdleSuspend niled it after close (race + goroutine leak on Close-during-tick); pass the stop channel to the loop by value. go build ./... (default) and the D9 shaterd linux build (tags with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp, with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/... ./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend and v2rayxhttp with with_xhttp. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c257d6c5cc |
docs(readme): rewrite root README as Russian shater product face
- README.md: new Russian product README (what/features/architecture mermaid/install both feeds/build/repo layout/CI/upstream/docs/license) - README.en.md: concise English mirror (root readme was previously English) - README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme, a competing Russian README) -> points to README.md + engine-fork docs - docs-shater/README.md: folder index Install commands copied verbatim from docs-shater/INSTALL.md; all links verified against existing files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
225cce5397 |
chore: purge working-session junk from tree; gitignore recurrence
Remove untracked-quality artifacts accidentally committed during work sessions (all authored downstream, unreferenced anywhere in code/docs/CI): - 5 session screenshots in repo root (devices-after-copy-fix.png, live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png) - tmp/gen_linux_test (29 MB throwaway traffic-gen binary) Guard against repeats: ignore /*.png (root screenshots) and /tmp/. Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
84d2766592 |
chore: remove stray committed test binary, ignore nested .idea, bump PKG_RELEASE
release / aarch64_cortex-a53 (push) Successful in 6m26s
release / x86_64 (push) Successful in 3m32s
release / apk aarch64_cortex-a53 (push) Successful in 4m58s
release / release (push) Has been cancelled
release / release apk (push) Has been cancelled
release / apk x86_64 (push) Has been cancelled
- drop c/Users/.../gen_linux_test (28MB binary accidentally committed in
v0.2.1
|
||
|
|
33f940c2fb |
ci(release-apk): publish available arch feeds even if one build fails
release-apk now runs with if: !cancelled() so an unrelated arch build failure (e.g. x86_64) does not block publishing the aarch64 apk feed. download-artifact only fetches existing artifacts and the publish loop already skips missing apkfeed-* dirs.latest |
||
|
|
a57717dabb |
health plan S7: docs, contract comments, SPEC 019 update
release / aarch64_cortex-a53 (push) Successful in 3m43s
release / x86_64 (push) Successful in 3m30s
release / apk aarch64_cortex-a53 (push) Successful in 5m11s
release / apk x86_64 (push) Failing after 5m8s
release / release apk (push) Has been skipped
release / release (push) Successful in 12s
- lx-changelog: health board + observatory + global probe + sub cache entry - DECISIONS.md: D18 board vs delete-and-overlay, D19 observatory vs sweep, D20 global probe settings - contract comments: urltest.go CheckOutbounds (fast circuit) + observatory.go loop (background circuit + freshness gate) document the two-circuit split - SPEC 019: dial-error section updated - slots still not moved, but board verdict demotes dead slot on next pick + retry (§5.B); sticky/replace-in-slot/never-shrink invariants preservedv0.2.0-healthplan |
||
|
|
129e31fbdc |
health plan wave 3: integration tests + chain unused-badge
- S6: integration tests in shater/generate (health_retry, chain_health, observatory_reach); portable tests pass -race; linux-tagged tests vet+build under GOOS=linux for OpenWrt CI - §5.E fix: unused-badge extended to chains; observatory used-set (chain exit tags) inverted via parseChainExitTag -> usedChainNames -> ChainHealth(names) -> /api/groups/health chains field -> ChainRow renders gh--unused badge (same pattern as groups) |
||
|
|
4f0618515e |
health plan wave 2: alive-only selection+retry, observatory replaces sweep
- S3: Select() alive-only by verdict; dial failure marks board + retries <=3 within ctx; ListenPacket retries to first send; balancer slot liveness reads board verdict; testNodes marks-fail instead of delete; SPEC 019 slot invariants preserved; selector untouched - S4: new observatory.go (reachability plan from rules, batch<=24/concurrency<=12/timeout 5s, freshness gate, cursor preserved on identical plan); probeplan BuildObservatoryPlan; health.go on board verdicts (TTL=max(3*interval,10min)); engine.dead overlay removed; sweep.go+probeall.go+TestAllNodes+/api/nodes/test removed (->404); GroupHealth.Used published; exit-test extended to chains; panel unused-badge + chain Test button; stats on board |
||
|
|
fd698162c9 |
health plan wave 1: urltest health board, global probe settings, sub cache out of UCI
- S1: History gains LastOK/Delay/LastFail; MarkFailed/Verdict on storage (common/urltest/board_lx.go); StoreURLTestHistory preserves LastFail - S2: per-group ProbeURL/ProbeInterval removed, globals only; sweep_interval drained as dead option; panel fields dropped - S5: subscription nodes cached in /etc/shater/subs/<name>.json; UCI keeps manual nodes only; sub update writes cache file; panel PUT split; legacy from_sub migration |
||
|
|
ffa78d67bc |
feat(panel): drop the query log from Overview
The filtered-query-log block (SegMeter + top blocked + live QueryLog) is redundant with Insights. The stats poll stays - it still feeds the DNS filtering and Groups modules. |
||
|
|
61e51495c3 |
feat(panel): prune subscription editor to essentials with a key-value header list
The subscription form now carries exactly: name, URL, update interval, fetch via (+detour when proxied), User-Agent, HWID, and extra headers. Format, device identity, regex/proto/country filters, dedup and expiry-alert knobs are gone from the form (still honoured from UCI; a save carries them through untouched). Headers are edited as key-value rows and serialize to the existing `Headers: []string` "Key: value" contract. Name is editable: a rename rewrites FromSub on the sub's cached nodes and refuses collisions. |
||
|
|
bfe71cd1dd |
feat(generate): fail closed on unresolved rule targets, never fall back to the default route
Rule.Kill ""/"default" used to drop the rule, letting its traffic fall through to the broader rules below and finally the default route - a silent leak of exactly the traffic the operator singled out. ruleKillFallback now always returns an outbound: ""/"default"/"closed"/unrecognised block the rule's traffic in place; only an explicit kill=open goes direct. The default route exists solely for traffic no rule matched. |
||
|
|
9b3644becb |
fix(alert): don't repeat the IP as the device name in new-device alerts
With no DHCP hostname the alert read "10.67.0.223 (mac) at 10.67.0.223". Lead with the MAC instead: "aa:bb:cc:dd:ee:ff at 10.67.0.223". |
||
|
|
0a8bdbaf46 |
feat(devices): merge discovered devices by MAC
One physical device with several addresses (v4+v6, multiple leases) used to show as several devices. Discover now folds addresses sharing a MAC into a single row: new `ips` field lists every address primary-first, `ip` stays the primary (most recent lease), state is the best among addresses. MAC-less hosts remain one-per-IP. The panel shows the extra addresses as secondary chips; naming keys the config entry by MAC whenever it is known. |
||
|
|
ebe2e7807b |
feat(stats): drop router-originated DNS queries from insights
The engine resolves domains for itself (node server names, urltest probes, subscription/DoH fetches). Those queries carried an invalid client address and still landed in every insights surface. Gate them out at the single ingestion point (Aggregator.handleEvent): an event with an invalid or loopback client is dropped before totals, top domains, per-server counts, the timeline, and the query-log rings. Only LAN-client traffic is collected. |
||
|
|
c1e1b17a61 |
feat(profiles): remove preset packs
The built-in block-ads / ru-bypass / private rule bundles are gone: model.Preset, Model.Presets, the `config preset` UCI section, its render, the panel Preset type, and every fixture. The generate-side expansion was already removed with the profile rewrite in the previous commit. |
||
|
|
782770f306 |
feat(profiles): drop default route/egress/schedule overrides; pick uplink from UCI interfaces
A profile is now a pure uplink-conditional rule switch: Name/Enabled/ Priority/MatchIface/Enable-DisableRules/EndpointResolver. The per-profile DefaultTarget/DefaultEgress overrides and the profile-level schedule window (SchedDays/SchedStart/SchedEnd/SchedUTCOffset) are removed from the model, UCI parse/render, the generator, the WAN watcher, and the panel. Rule-level scheduling is untouched. The panel's uplink condition is now picked from a dropdown of the router's UCI interfaces (GET /api/interfaces, same source as the egress picker); stored interfaces missing from the live list render as stale chips. generate/profile.go is rewritten here (applyProfilesAndPresets -> applyProfiles), which also drops the generate-side preset-pack expansion; the preset model/UCI/panel surface is removed in the next commit. |
||
|
|
8980a25a59 |
perf(ci): cache SDK feeds checkouts — the biggest recurring build cost
Audit of the run-51 logs showed actions/cache@v3.3.2 works on the act_runner
(cold: "Cache saved" x4; next job: "Cache restored" in ~2s, npm --fast skip,
usign/dl reused) and the sdk-cache mirror seeds correctly — but the single
biggest recurring cost was NOT cached: `scripts/feeds update -a` re-cloned
base+packages+luci+routing+telephony every run (~7.8 min warm x 4 SDK jobs on
the serial runner ≈ ~28 min/run wasted; github ~1 MB/s from this host).
Cache .cache/feeds/{opkg,apk} (workspace dir, actions/cache-persisted, visible
in the SDK container via --volumes-from) symlinked over the SDK's empty feeds/:
`feeds update` now git-fetches deltas (seconds) instead of full clones, always
checking out feeds.conf's pins. Fail-safe: any error on the cached checkouts
wipes the cache and clones fresh. Key by SDK release (feeds-opkg-24.10.4 /
feeds-apk-25.12.1) — stable across runs, invalidates on an SDK bump; both arch
jobs of a lane share one entry (identical pins, serial runner).
Steady-state warm run: ~60+ min -> ~20-22 min. Also documented in the workflow
header: never key a cache on github.sha — each cache SAVE stalls the act_runner
~3 min, so per-run-changing keys would add +3 min/entry every run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
cb983afee1 |
perf(ci): stall-proof SDK fetch + caching (SDK/dl/go/npm/apt/usign) + concurrency
Builds were dominated by re-fetching the ImmortalWrt 25.12 SDK tarball
(~300 MB) every run, and a stalled downloads.immortalwrt.org transfer wedged
the apk job for 40+ min (plain `wget -q`, no timeout — same class as the
elfutils hang).
- New ci/fetch-sdk.sh (runner-side): cache -> our durable `sdk-cache` release
mirror -> upstream with a stall-kill (curl --speed-limit 64K --speed-time 60
--max-time 1800) + 3 retries + zstd-magic/size validation; seeds the mirror
best-effort (github.token, non-fatal) so cold runs never touch upstream again.
A 40-min hang is now impossible; the in-container fallback wget also gets
--timeout=60 --tries=3.
- actions/cache@v3.3.2 (last release on the OLD cache API that Gitea act_runner
implements; v4/v3.4.x use the new GitHub cache service) for: SDK tarball, SDK
dl/ sources (hash of package Makefiles; PKG_HASH re-verified so a stale cache
can't leak a wrong source), Go mod+build (go.sum), npm node_modules
(package-lock.json) with build-shaterd.sh --fast, apt archives, built usign.
Degrades safely if the cache server is off — the SDK mirror is independent.
- concurrency group release-${github.ref} cancel-in-progress so a re-dispatch
cancels the stale run instead of piling up (tags stay isolated).
Signing (usign/apk), both keys, per-arch publish, manual triggers, LOCALMIRROR
and the scoped 4-package collection are unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
sdk-cache
|
||
|
|
6d2a729eaa |
feat(panel): gate byedpi egress on package presence
byedpi (ciadpi) ships as a separate optional package; the panel offered the
`byedpi` egress type regardless, so selecting it without the package installed
created a dead, fail-closed egress. Now GET /api/status reports
`byedpi_installed` (exec.LookPath("ciadpi"), os.Stat fallback), and the egress
type picker disables the ByeDPI option with a hint when it's absent. Existing
byedpi egresses are never hidden or rewritten (config is sacred) — shown with an
amber warning and still round-trip on save; only NEW selection is blocked.
Unknown status (older daemon / fetch fail) => no gating.
Bump shaterd PKG_RELEASE 1 -> 2 (the SPA is embedded in the daemon binary).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
4a1542b30a |
fix(shater-core): release procd flock so install can't deadlock the pkg manager
On a live `apk add` / `opkg install`, shater-core's post-install hung forever (observed on BananaWRT 25.12 at "Executing shater-core...post-install", child `flock 1000` in locks_lock_inode_wait). Root cause: a USE_PROCD init sources /lib/functions/procd.sh on every rc.common action, whose procd_lock takes a BLOCKING exclusive flock on /var/lock/procd_<svc>.lock held until the process exits. shater-cron re-execs itself as the eternal `loop`, so it held that lock forever; base-files' default_postinst then ran `/etc/init.d/shater-cron enable` synchronously inside the transaction, blocking on the flock while the package manager waited on the postinst — a permanent deadlock. Fix (two layers): - shater-cron `loop()`: `exec 1000>&-` closes fd 1000 up front so the eternal loop never holds the rc.common flock (no-op when procd_lock is absent). - 30_shater-core: defer enable/restart into a detached (setsid + bounded) background block that waits for apk/opkg to finish before touching init.d, with all fds to /dev/null (a held stdout pipe would hang apk on EOF too). Bump PKG_RELEASE 1 -> 2 so existing installs pick up the fix. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b7070ad7e8 |
fix(ci): install python3-distutils for the ImmortalWrt 25.12 apk-SDK
The apk lane runs the SDK on a bare debian:bookworm host, and the ImmortalWrt
25.12 SDK prerequisite check requires python3-distutils ("Checking
'python3-distutils'... failed. Prerequisite check failed." ->
.prereq-build Error 1), aborting before any package built. The opkg lane was
unaffected because the openwrt/sdk image ships the prereqs. Add
python3-distutils (and python3-setuptools defensively) to the host deps. apk
lane only; opkg untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ea24bad232 |
fix(ci): make $OUT writable for the SDK user and collect only our 4 packages
Two build-harness bugs surfaced once the SDK builds actually ran:
1. Permission denied writing the feed. ci/build-feed.sh creates $OUT as root on
the runner, but the openwrt/sdk container runs as the unprivileged `buildbot`
(uid 1000) — so `cp` of the .ipk into $OUT failed ("Permission denied"),
yielding 0 packages and then "usign signing failed" (nothing to sign). Set
`chmod 0777 "$OUT"` on the runner before docker run (a chmod from inside the
container, as buildbot, cannot fix a root-owned dir). The apk lane already
chmods $OUT from its root debian container, so it was unaffected.
2. Collecting the whole SDK. ci/sdk-build.sh did `find bin -name '*.ipk'`, which
swept up the hundreds of prebuilt kmod/base .ipk shipped in the SDK image —
bloating the feed and signing foreign kmods under our key. Collect strictly
our four by name (`<pkg>_*.ipk`) and require >=4. Applied the same narrowing
to ci/sdk-build-apk.sh (apk names carry no arch: `<pkg>-*.apk`), keeping the
"wrong SDK produced only .ipk" guard.
No change to the feed format/signing (usign/KEY_BUILD/shater-feed.pub, apk EC
key), the package set, or triggers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
1d285cf34f |
fix(ci): route SDK source downloads through OpenWrt CDN mirror
The apk (and opkg) SDK builds intermittently hung fetching build-time sources like elfutils-0.192.tar.bz2 from sourceware.org: curl's --connect-timeout covers only the TCP handshake, not a stalled mid-transfer, so a slow upstream hangs the whole job (no --max-time in OpenWrt download.mk). Set CONFIG_LOCALMIRROR=https://sources.cdn.openwrt.org in .config before `make defconfig` in both ci/sdk-build-apk.sh and ci/sdk-build.sh so the SDK tries the fast OpenWrt source CDN before each package's own PKG_SOURCE_URL — fixes elfutils and any other flaky upstream. Mirror verified to hold the file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5092004b40 |
fix(ci): init wireguard-go submodule in the opkg build job too
Same root cause as the apk jobs: scripts/build-shaterd.sh builds through a
go.mod `replace => ./submodules/wireguard-go` (AmneziaWG fork, bumped in
|
||
|
|
bcdea8f04c |
fix(ci): init wireguard-go submodule in apk build jobs
scripts/build-shaterd.sh builds via a go.mod `replace => ./submodules/ wireguard-go` (the AmneziaWG-patched fork), so that submodule must exist or `go build` dies with "reading submodules/wireguard-go/go.mod: no such file or directory". actions/checkout does not fetch submodules by default. Init only that one submodule (public GitHub URL; clients/apple+android are large and unused) in the additive build-apk jobs — the opkg build jobs are left untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cd598b0fe2 |
feat(ci): add apk (ImmortalWrt/BananaWRT 25.12) release lane
Additive next to the opkg/24.10 lane — nothing existing changed. The same 4 packages (shaterd, shater-core, luci-app-shater, byedpi) are built through the official ImmortalWrt 25.12 apk-SDK and published as per-arch rolling releases apk-latest-<arch> / apk-<tag>-<arch> (x86_64, aarch64_cortex-a53). - ci/sdk-build-apk.sh: drives the 25.12 SDK inside debian:bookworm, compiles .apk, then `apk mkndx --root T --keys-dir T/keys --allow-untrusted --sign KEY --output packages.adb *.apk` — the exact form the OpenWrt 25.12 buildsystem uses (unsigned members, signed index). - ci/build-feed-apk.sh: per-arch runner entrypoint (same --volumes-from and artifact-order contract as ci/build-feed.sh). - ci/gen-apk-key.sh: one-shot EC (prime256v1) keypair generator; private half -> Gitea secret KEY_APK, public dist/shater-apk.pem committed. - release.yml: additive build-apk / release-apk jobs; `on:` triggers untouched (v* tags + workflow_dispatch); apk release tags deliberately non-`v*`. - docs-shater/INSTALL.md section 6, .gitignore (out-apk/), dist/shater-apk.pem. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |