v0.2.17
3051
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6476722372 |
fix(panel): stop shipping a fabricated router in the binary
mock.ts was a static import and the mock switch was read from the query string at runtime, so the bundle that ships inside the daemon carried a complete fictional router and a link ending in ?dev rendered it: protected, 119 of 122 nodes alive, without a single request to the daemon. The only tell was a line in the footer. That is worse than any wrong number — there is no data at all and nothing says so. It is out of the production bundle now, which is 21 kB smaller for it. Unknown state stopped reading as good news in two more places. The kill-switch tile treated an absent plane as armed, because the check was "not none" and undefined satisfies it — the contract in the API types says the opposite. And the apply page announced "daemon auto-rolled back" from its own timer, while the daemon, seeing the state generation move, disarms and says it is NOT rolling back in the log only. Alerts moved to Settings. They are about the kill switch, apply failures, new devices and subscription expiry, and they lived at the bottom of the DNS page, while Settings mentioned them in prose with nothing to click. Findings truncation is visible now: the notice that says how many were suppressed arrives as info, and the attention list keeps only critical and warning, so past fifty findings the operator saw forty-nine and no hint of the rest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>apk-v0.2.17-aarch64_cortex-a53 apk-v0.2.17-x86_64 v0.2.17 |
||
|
|
4078334d85 |
fix(stats,alert,panel): put a ceiling on everything that only grew
Four maps had no bound on a box with 512 MB that runs for months. The health board only ever inserted — the delete exists but no path in this fork calls it — and it lives on the engine context, so it outlives every generation. Its keys are node tags, and providers rename nodes on each subscription refresh: about 440k keys a year, some 88 MB. Alert dedup keyed on MAC with no delete at all. The stats aggregator's server and outbound counters were the only ones with no cap, no prune and no top-N, and one of them was handed to the panel whole on every poll. They are bounded now, evicting least-recently-seen, with numbers argued from this box rather than round: the board holds 4096 against a live generation of about 1200 tags, so a rename day cannot evict a tag still in use. Nothing is dropped silently — the same rule the log sink already follows — and a new Dropped section in the snapshot reports all six bounded aggregates, including the three that had been evicting without saying so. Snapshot did O(devices × domains) under the aggregator lock, sorting five thousand entries to show fifteen, and could read the DHCP lease file from inside it. Meanwhile the event subscribers have 64-slot buffers that drop without a counter, so an open Overview page cost the query log real rows. Selection is top-K now — proven byte-identical to the old sort over 200 random trials — and both the lease read and the row ordering happen outside the lock. The panel server had one timeout, on headers. An unauthenticated client could hold a goroutine, a socket and a descriptor forever by sending its body one byte at a time; a stopped reader on the log stream held the handler, the pipe and a child process that outlived the request. Every phase is bounded now, with the unauthenticated route on a tighter budget than the rest, and the log stream renewing its deadline per chunk so a slow-but-reading client is never truncated. And the last of the detour transports: each call built a fresh one, and the alert delivery path dropped it, pinning keep-alive sessions through the engine's own outbounds for 90 seconds — eighteen times the budget a retiring generation gets. The race skip is gone from the gate. The test it existed for raced in its own clock, not in the product; that is fixed, so nothing is excluded under -race any more. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0a6689b29e |
fix(quic,v2ray): close the sockets quic-go was never going to close
DialEarly with a packet conn the caller made sets a flag that means quic-go does not own it: closing the transport only stops reading from the socket. Neither DNS transport closed it. On the QUIC one it was closed on a failed handshake and never on success, so every redial — idle timeout, retry error, engine reload — left a UDP socket for the life of the process. On the HTTP/3 one the library drives its own reconnects, so the leak compounds without anything in our code looking wrong. That is the same shape as v2rayquic's, where offerNew overwrote the raw conn on every reconnect without closing the previous one. Both are now owned by a watcher tied to the connection's own context, so the socket lives exactly as long as the connection does. This matters more than it did last week: the shipped resolvers are DoH, and DNS is intercepted by default now, so the whole network's query stream rides this path on a router with 512 MB. The same upstream commit fixes both halves. We had taken the v2ray half and not the DNS one — the third time this session a paired fix arrived half-applied, and the first of those cost a day of debugging. These two files are now byte-identical to upstream so a rebase cannot reopen it. Also from that family: websocket and httpupgrade leaked their conn on failed handshakes, and a QUIC stream's Close did not release a blocked write. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ef22167b1a |
fix(apply): report the hold when the plane was armed by someone else
Booting with the armor loaded, or restarting through the handoff, left the status saying the LAN was not being held while it was being dropped. Transient after a successful apply, but permanent on the unreadable-config path — and there the apply-failure alert words itself "traffic is NOT being blocked" at the exact moment it is. That sends the operator to fix something that is not broken, past the protection that is holding. The table cannot be identified from here — netplane exposes no read-back and nft does not keep comments — but identifying it is the wrong question. Holding does not claim the holding plane is the object in the kernel; it claims the engine is down and forwarded traffic is being dropped. A leftover full ruleset does that too: with no engine socket the tproxy statement breaks its own rule before the accept, so the packet reaches the forward chain unmarked and meets the primary drop. What decides it is whether the last applied config was enabled and fail-closed, which is exactly what the boot armor's presence already means. So it is derived at read time rather than latched. A latch set from an inference would have to be remembered in order to be cleared, which is the trap the active flag already taught us. ArmHold also stops deferring to a table it cannot inspect and installs its own render instead — the honest answer to "do not claim a foreign table blindly" is to make it ours, and a fresh render beats a snapshot that predates an interface rename. Also closes the last of the detour transports: the subscription fetch took a client and dropped it, and the exits that leak are the error ones, retried by cron forever against a broken feed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cbda0fee0a |
fix(netplane): arm the fail-closed plane before the daemon can
The plane only ever existed while the daemon did. It starts at 99, after fw4 has already loaded lan→wan ACCEPT, and only reaches ArmHold after waiting out its predecessor, migrating the schema, building the engine and reading UCI — with a UPX-compressed binary decompressing off flash first. Every boot therefore had a window with no protection at all, landing exactly when Wi-Fi comes up and every client reconnects. A restart, a reload or a package upgrade opened the same window on purpose: Teardown does not consult the kill switch, and the init script guarantees the interval is non-empty. The holding plane is now persisted to /etc/shater/boot.nft on every apply and loaded by a small service at 21, right after fw4 and netifd. Its presence is the arm token: it exists only while the last applied config was enabled AND fail-closed, and goes away the moment either stops being true. Writes are content-gated — the cron reconcile runs a minute — and atomic, because the one boot that reads this file is the boot after a power cut. The service refuses to arm four ways so it can never brick a box, and its enabled-check reads /etc/rc.d directly rather than asking rc.common, which would take a blocking flock in the middle of boot. On exit the daemon re-arms only for restart and reload, read from a snapshot of rc.common's action; anything else, including an unknown one, degrades to a real stop that also disarms. An unreadable config used to leave the router bare forever: the arm call sat in the branch that requires a successful read, and nothing downstream could recover it. It now arms from the same path. A network nobody named was neither diverted nor blocked — the divert set is built from inbounds and rule sources, and the same set scopes the fail-closed drops. It is now enumerated from the interfaces whose firewall zone the operator forwards to a WAN zone — their own statement that those clients reach the internet through this box — and reported critically, by name, with both resolutions. Deliberately not closed automatically: this router cannot know a guest SSID was meant to be off the tunnel, and guessing is an outage. A device name that resolved to nothing is reported the same way, for the same reason: there is no fail-closed action available for a device we cannot name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7234817adb |
fix(panel): flush the log before serving it
Splitting the log sink made its writes asynchronous, so a download could miss the last lines still in the queue — silently, with a successful response. Those are the lines the operator came for: a log is downloaded to find out what just happened. The panel is handed a barrier, not the sink: a func() set once at startup, the same shape as the reconfigure hook and the stats setter already in the tree. It cannot write, reconfigure or close, so it stays a consumer, and nothing about the sink's type reaches it. The wait is bounded at the sink's own control budget and enforced on the panel side, so a wedged writer cannot turn the download into the new place the daemon gets stuck — the very thing the async split was for. Past the bound the handler serves what is on disk. With no barrier installed the path behaves as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>apk-v0.2.16-aarch64_cortex-a53 apk-v0.2.16-x86_64 v0.2.16 |
||
|
|
f1c36d6eea |
fix(panel): say when the engine is down, and ask before the irreversible
The header could not render "offline": it keyed on a field the daemon pinned to true, so a dead engine behind a fail-closed plane showed a pulsing green lamp. Health now needs both signals to agree before it reads as up, and a negative from either is enough to say down — which is honest against the field that was already honest, and stays honest now that the other one is too. Deleting the last catch-all rule was described as "traffic will fall through to the next rule" on the very row the page badges as the default route. What happens instead is the kill switch: closed, the network loses the internet; open, it leaves with the real address. The dialog now says which, by reading the saved setting, and the toggle asks the same question — the generator only emits enabled rules, so switching it off is the same event. The master switch tore the whole plane down without a word, while deleting a rule-set got a confirmation. Deleting a node or a resolver claimed to remove it "from the config" without mentioning what still points at it, though the reference finder was already there and used for renames. Every Apply button armed the auto-rollback, and only one page said so. The window is now recorded where all of them pass through, carried in a band under the nav on every route, and persisted — so the countdown and the keep button survive a reload, which is what made the window unconfirmable before. Overview's Confirm button is gone rather than gated: Confirm cannot fail, so a permanently live button could only ever report success. Blocklists printed "filtering" from two config checkboxes without asking whether the list had ever loaded — while the daemon grades a failed load critical. They now show what the rule-set rows already showed, and say "not loaded — nothing blocked" when that is the truth. Also: the clock read UTC while every timestamp rendered in the browser's zone, so the router appeared to have started in the future; the rule counter on Overview counted saved rules rather than the ones in force, unlike the routing page; and the hop badge counted the entry egress the rail below it does not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bb21ceb7f5 |
fix(lifecycle): a busy file is not a broken one, and three ways to lose state
Changing any stats knob on the persistent backend deleted months of history. The replacement store was opened before the outgoing one was closed, so it hit the first one's flock, timed out — and the open path treated ANY error as corruption and unlinked the file. Unlink of an open file succeeds on Linux, so the new ring opened an empty database while the panel was still told the backend had not changed. The store now hands its resources over before asking for them again, and deletion is gated on an allow-list of real corruption signals; a busy, unreadable or read-only file degrades to the in-RAM ring and is left alone. The holder's reads were unguarded in a subtler way, caught only after the gate failed twice: the accessor took the read lock, returned the pointer and released it, so the call ran outside. A reader could hold a store the swap then closed and be served its empty answer — an empty page presented as data. The accessor is gone entirely, along with the possibility of handing out an unguarded reference. Readers still do not block each other; the swap now waits out reads already in flight, which is a page at most. The urltest group published its chosen node through two plain fields written by the prober and read on every dial and every panel poll — while the selector next door does the same job atomically. They are one value now, so TCP and UDP can no longer be read as a mismatched pair. Nothing had ever dialled through a group while it was probing, which is why the detector had never seen it; a test now does, and reproduces it deterministically against the old shape. Close on a group whose ticker had already stopped returned before closing its channel, and Touch would then arm a fresh loop nothing could stop. Reached by pressing Test in the panel and applying a config within the next two minutes: the orphan kept failing probes against a cancelled context and writing forged dead verdicts into the board the live generation selects from. Close is now final. The log sink held one mutex across a blocking write. Under procd stderr is a pipe, so a reader that stopped draining wedged everything that logs — engine, panel handlers, signal loop — while the process still answered a signal. It is split: a front that assembles lines and a writer that owns the destinations, joined by a bounded queue that drops and counts rather than blocking. Proven by restoring the old shape: the package deadlocks for the full ten-minute timeout, parked exactly where the field symptom said. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a8970b8ace |
fix(apply): stop the status from reporting a state the daemon is not in
Running was the constant true. The panel builds its header from it, so the "offline" branch was unreachable code: with the engine dead and the LAN behind a fail-closed hold, the operator saw a pulsing green lamp and, on the page people open to fix things, "engine: running". The honest field sat beside it, documented as the honest answer to are-we-proxying, and was read nowhere. running now means shater is running: the daemon answered and its engine has a started instance. active stays what it always was and is documented as such — the "meant to be running" latch that gates hotplug and cron, not a health signal. It is deliberately not cleared on hold, because the cron loop gates on it and clearing it would switch off the reconcile that brings the engine back. Two paths published nothing and so left the previous config's verdict standing for as long as the fault lasted. A rollback with no snapshot re-applied the engine and the plane and never touched the traffic verdict, so a router rolled back to a direct default kept reporting the tunnel. And an apply that failed in the netplane stage had already swapped the engine, then returned before every publisher, so status described the config that was no longer running — and the next reconcile, seeing an unchanged hash, failed the same way and published nothing again. Both now publish, with an unknown verdict: after a no-snapshot rollback the engine runs options this process does not hold, and guessing from UCI would describe the config we rolled away from. The severity classifier had drifted from the texts production emits. Markers were compared case-sensitively against wording that had since changed, and the entity pattern could not match a message beginning with an upper-case tag — so a blocklist that failed to load graded as a warning while a typo in its URL graded critical, and the panel's banner, which only lights for criticals, stayed dark for the outage. RULESET-NOT-APPLIED and DNS-FILTER-NOT-APPLIED are now read as the structural markers their producer documents them to be, so severity no longer depends on wording at all. Five markers that matched no living text are deleted; three protection-section texts drop to warning, because a blocklist that is stale but still blocking lights the alarm on most reconciles behind a flaky link, and an alarm that is always on is how the real one goes unread. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4996bc0984 |
fix(untunnelable): make block actually block
The block policy collapsed into direct whenever routing's final target was direct — the common "tunnel only what is blocked, everything else direct" shape. So ICMP, ESP, AH, GRE, IGMP and SCTP left with the client's real address under the setting whose own field doc promises "nothing ever leaves with the client's real IP", including a standing VPN on the real address, which is exactly what the middle rung exists to separate out. Both ends of the ladder now short-circuit before the plan is consulted and neither may consult it: direct accepts everything, block emits no line at all and lets the fail-closed drops the caller writes next do the work. A rule scoped by source could also widen the other family: emit() skipped a family whose destination list was empty but not one whose source list was, so a rule carrying only IPv6 source prefixes rendered an IPv4 line with no ip saddr clause — an accept for every IPv4 host on the LAN. The two halves now read "scoped" the same way the catch-all collapse already did. No destination plan is built for block at all now. It is the shipped default, and a geoip-backed plan is ~159 000 prefixes pushed into kernel memory and the ruleset text for a policy that cannot use them. The operator-facing texts said IPTV works. It does not, on any of the three rungs: inbound multicast is never matched by these rules and a client's outbound multicast UDP dies at the fail-closed guard regardless. Saying otherwise invited trading the ESP/GRE block away for nothing. What actually stops working under block is stated instead, and precisely: raw ESP/AH and GRE, but not IPsec through NAT or any UDP VPN, which are ordinary tunnelled traffic. TestOnlyPinnedAddressIsTunnelled is how this hid: it asserted, on the default policy, that an exception line was emitted, and read that as the feature working. It was block rendering direct. Its render assertions move to icmp, where they mean something. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0de597d69 |
feat(dns): intercept by default, and bootstrap node addresses off the tunnel
test / go + panel tests (push) Successful in 4m56s
The posture was inverted. A client using the DHCP-supplied resolver — the router itself — was NOT intercepted: dnsmasq answered and forwarded to the ISP in the clear, so the filter, the blocklists, the per-device rules and BlockDoH were all inert for exactly the clients that did nothing wrong. A client that hardcoded 8.8.8.8 to route around us WAS intercepted, by the catch-all. Meanwhile the docs promised no DNS leaks. The default now matches the promise. Turning it on crosses a threshold that was already dangerous for anyone with two resolvers. Above one transport, a node's domain server address stops being resolved by the transport directly and goes through the client DNS plane instead — so a blocklist entry, a block_doh NXDOMAIN or any dns_rule can answer your own node's hostname, and one sloppy line in an ad list stops being an ad that got through and becomes a tunnel that never comes up. So the fix is gated on having two or more transports, not on the intercept toggle: resolver_default plus resolver_fallback always reached that threshold, long before this change. When no endpoint_resolver is configured the plane now carries a bootstrap server — the default resolver cloned with its detour dropped, keeping its type, so a DoH default stays DoH and only the tunnel hop goes. An explicit endpoint_resolver still wins. This is not a restore of the previous behaviour and the comment says so: at one transport the dialer used the default resolver WITH its detour, so a lone DoH-through-the-tunnel resolver was already a bootstrap loop. It is strictly better than what came before. Existing installs keep whatever they set — the config file is a conffile and is never replaced — and an explicit dns_intercept '0' survives the render-parse round trip, which a default-true bool otherwise makes easy to lose. The no-resolver warning stays, and no default resolver is shipped to silence it: a placeholder would remove the sentence without moving a single query, and the panel would then say a resolver was configured while nothing was filtered. Its wording is corrected instead — .lan keeps working through the built-in local transport, which the old text denied. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
544da29863 |
fix(dns,netplane): close three paths that sent traffic out in the clear
A resolver whose detour no longer resolved fell back to "the default outbound", which is not a default at all — it is a plain system socket. Every other place in this generator fails such a reference closed, with an essay explaining why, and wgdedup rewrites the very same field to block when it drops an endpoint. One field, two opposite policies, and which one applied depended on whichever code noticed the breakage first. A resolver detoured through a node the operator switched off therefore handed the whole network's query stream to the ISP in the clear, while the kill switch held the traffic itself. It now fails closed, and the warning says what that means: the resolver answers nothing, and if it is the default one, name resolution stops network-wide until the target is restored. A dns_rule naming a missing resolver used to be dropped whole, sending exactly the names the operator singled out to a resolver they did not choose; it keeps its matchers and answers NXDOMAIN instead. Not a reject action — one built in Go with an unset Method panics the engine at match time. RoutingPresent never looked at per-egress rules or tables, and applyLocked skips the whole routing stage on its word. So an egress table wiped by an ifdown was never restored: the marked traffic fell through to main and left over the plain WAN, permanently, with plane full and no warnings. It now verifies each binding it installed, recording intent rather than outcome so a broken egress keeps the plane reported absent and heals when the interface returns. addEgressRouting discarded every ip error, so an egress that failed to install reported success and the panel drew it green. Failures are now critical warnings naming the egress, the device and what ip said — but still warnings, because returning would abort the apply and punish the household for one bad uplink. Also anchors the fwmark check: with a small fwmark_base the main mark is a literal prefix of the first egress mark, so a substring match could answer "the main rule is installed" while looking at an egress rule. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
daaa0fda41 |
fix(wireguard): stop holding AmneziaWG down behind a WireGuard hop
The guard refused to start an AmneziaWG endpoint whose detour chain reached a WireGuard one, and refused silently: not an error, just started=false, after which every dial failed with "WireGuard is not ready yet". A selector hook went further and suspended an already-working node the moment its group switched to a WireGuard member. It existed because AmneziaWG inside WireGuard hung the kernel on Android. We do not ship Android, upstream dropped the guard once the cause was gone, and the cure landed here yesterday — the ClientBind reserved-gate plus the submodule pin that carries its twin. So the tree held both the cure and the prohibition on using it, and the configuration simply did not come up while looking like a node that "just does not work". Also takes the two fixes that belong with it. ClientBind.conn was read on a lock-free fast path and written under a mutex; upstream found that race with the same end-to-end test we wrote yesterday, so we had taken one half of a pair again. And the outer WireGuard UDP socket forced DF, unlike direct, hysteria and tuic — with encapsulation the datagram regularly exceeds the path MTU and the kernel drops it instead of fragmenting, a symptom indistinguishable from the bug we spent yesterday on. The race needed its own test: the existing e2e run did not flag it under -race even at -count=15. Eight goroutines over both connect branches reproduce it deterministically, naming the lock-free read and the guarded write. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
56a276bcc1 |
ci: make the tests a gate instead of a decoration
The fork had a full suite and no CI that ran it. Upstream's test workflows trigger on stable/testing/unstable; this repo only has main. And Gitea does not read .github/workflows at all once .gitea/workflows exists, so those files were decoration here. 115 of the 116 test files under shater/** had never executed in CI even once, which is how TestDNSFilterRemoteBlocklistHTTPClient stayed red across two published releases without anyone noticing. The gate is a job inside release.yml that build-apk needs, because a separate workflow cannot block another one. It runs the suite under the shipped tag set, on Linux — 6 of 7 test files in transport/wireguard and 12 in shater/generate compile only there or only under those tags, and those are exactly the files covering AmneziaWG. Three guards stop it from passing by running nothing, which is the failure this whole change is about. The tag set may only ADD test files, never remove one. Every package go list says has tests must appear as "ok <pkg>" in the output, so a suite that collapses to "no test files" fails instead of passing. And the panel run counts its test files first, because node --test exits 0 with "pass 0" when the glob matches nothing. The publish step used to exit 0 having published nothing: its assertions all live inside a loop over artifacts, so an empty directory ran the body zero times and reported success. It now counts what it published and fails on zero. Verified by extracting the shipped step text and running it against stubs: empty artifacts gives exit 0 before and exit 10 after; the rolling-release readback still fires its own exit 14. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
754bbcf1fa |
fix(submodule): point .gitmodules at the line the pin is actually on
The wireguard-go submodule is pinned to 7d15f33, which lives on lx-awg2-v005. .gitmodules named `lx` — a separate line, 42 commits one way and 131 the other, with no common recent history. That is a loaded gun rather than a cosmetic mismatch. `lx` has no hasReserved() gate in conn/bind_std.go at all, so a single `git submodule update --remote` would move the pin there and silently restore the defect fixed yesterday: the bind shreds the AmneziaWG magic header of every transport packet, handshakes complete, no data moves, and no chain containing an AmneziaWG node carries traffic. It would also drop the padding-overrun fix and the v0.0.5 re-graft. Nothing about the checked-out tree changes — the pin is untouched. Only the branch a --remote update would follow. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
63e6b709f8 |
fix(health): a chain blocked at a hop is dead on the board, unknown on the card
The short-circuit left the chain's exit tag untested, because the exit is itself a hop and every hop behind the break was rewritten that way. A stand run caught it — the test lives in a file that does not compile on the dev host, so nothing local could have. That is not neutral silence. selectExcluding ranks untested ABOVE dead and says so in its own comment: with no fresh-alive member, an untested one is a better bet than a known-dead one. Leaving a provably broken path untested is therefore a positive preference for it over a path we merely know is dead. The two readings answer different questions and now differ on purpose. Is this hop's own node alive — unknown behind a break, so the card keeps untested and blocked_by. Can this chain carry traffic — known, no, because the hop in front of it was probed and did not answer. The board carries that second answer, which is the one selection, the freshness gate and the manual test all read. The exit verdict is derived, not dialled: it records the consequence of a probe that did happen one hop earlier, and it is re-derived every pass, so the moment the blocker answers the walk reaches the exit again and the next verdict there is a real measurement. Also keeps a routed group warm. Its checker used to stop on the idle timeout and nothing filled in behind it, so a rule that fires rarely would show untested while being in force and pay a cold probe on the first real request. The gate that adds this work answers false when it does not know — the mirror of the one that withholds work, so plain sing-box keeps the lifecycle it always had. And the tls-spoof suite now skips without tcpdump instead of failing sixteen times: a missing tool is not measured, not broken. The same distinction this commit is about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>apk-v0.2.15-aarch64_cortex-a53 apk-v0.2.15-x86_64 v0.2.15 |
||
|
|
8612b0a9e9 |
fix(health): one dialler per target, and it is the group's own checker
A member of a chain hop wrapper was reachable by two probers: ours, from the observatory plan, and sing-box's, from the urltest group the wrapper actually is. Two independent readings of one node can disagree, and then neither can be trusted — which is worse than the wasted dial. The group's own checker is the right owner. A hop wrapper's members are the per-chain copies, each carrying the previous hop as its detour, so that checker already travels the chain prefix — the path the traffic takes. The plan now records who dials each target and the observatory skips the ones a live checker owns, keeping only what no group covers: node hops, the AmneziaWG endpoint, selector members, and the members of groups that have been stood down. The jobs stay in the plan rather than being deleted, and that is load-bearing: the short-circuit reads the plan as the map of which tags measure which hop, so deleting a urltest hop's members would erase that hop from the map and quietly stop it blocking anything — on exactly the chains the feature exists for. The short-circuit therefore moves to the group as well, through a ProbeGate the engine implements: a scheduled check asks whether the path in front of it is up before dialling, while an explicit check is never refused. Nothing is stored — the gate recomputes from the live board every call — and Touch still arms the ticker even while blocked, because a hop that refuses to tick has nothing left to notice its own recovery. The gate answers yes whenever it does not know: refusing on missing information is how a system talks itself into silence. Two grounds now exist for a group not to probe and they must not be merged: stood down means no rule reaches it at all, blocked means the path in front is down right now. Both doc comments say so and name the chain hop wrapper as the case where the difference bites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
96d9cfaa63 |
fix(health): stop probing a chain below a hop that is already down
Hop probes were independent, so every hop was dialled whether or not the path to it existed. A hop is dialled THROUGH the hops above it, so when hop 2 had no live member left, the probe for hop 3 failed at hop 2 and hop 3 was recorded dead. Dead means "we tested this and it did not work" — but nothing was learnt about hop 3 at all. One broken hop painted the whole chain dead and pointed the operator at the wrong place, and every one of those probes was a dial with a timeout down a path already known to be broken. Chain jobs now run in path order and the walk stops at the first hop that reads dead. Hops below it are not dialled at all and are reported untested with blocked_by naming the hop that stopped the walk — the honest answer, since nothing was measured. Nothing latches. There is no blocked flag: the gate is a fresh read of the health board at every hop of every pass, and the cursor rewinds to the top each cycle, so the first dead hop is never behind a break and is always retried. The moment it answers, the rest of the chain runs in that same pass. Only a positive dead blocks; untested never does, or a cold start would never open. Blocked hops are rewritten rather than annotated, because board records do not vanish when the prober stops dialling — they age out on their own TTL, and the worst version of that is a stale dead pointing at a hop that may be fine. The exit tag is exactly what stops being dialled, so the group test would have waited out its full deadline and then reported "not reached yet" about a chain it already knew was down. It now names the blocking hop immediately, gated on the same freshness watermark so a break seen before the request cannot short-circuit a pass that may be about to find that hop alive. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3c7536dba0 |
fix(panel): show the chain hop by hop, and stop reading unused as broken
The chain card gave a single verdict, so a dead hop was invisible: the operator saw "the chain is unhealthy" and had to guess which of four hops to look at. Meanwhile a group used only inside a chain showed "unused" next to a live alive/dead count, which reads as a diagnosis when it only means nothing measures it on that path. Render the hops as a rail that severs below the first dead one, so which hop is answered before a word is read, and split the two "not routed" messages into the routing fact and the explicit non-fact. The group one names the case directly: a group used only as a hop inside a chain reads unused here on purpose, and its real health is on that chain's card. Also fixes a bug this would otherwise have shipped: the readout painted every ok:false in the critical colour, so "not routed" would have rendered as a fault — the exact lie being removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>apk-v0.2.14-aarch64_cortex-a53 apk-v0.2.14-x86_64 v0.2.14 |
||
|
|
3c92e1cbfd |
fix(health): one prober, on the path the rules actually use
A node reached only as a chain hop was being measured twice, and the reading the panel showed was the wrong one. On a router in Russia that is not a cosmetic difference: a node the chain carries fine behind a WireGuard hop is dead when dialled straight out of the WAN, so the group card read "0 of 2 alive" while that very group was carrying every packet. Two dial paths existed outside the observatory plan. URLTestGroup.PostStart warmed up every urltest group at box start whether or not any rule reached it, and the panel's Test button reached URLTest.DialContext, whose first act is Touch() — arming a ticker that re-swept those groups directly every probe interval for the next thirty minutes. Both wrote under the BASE node tag, and both dialled the base outbound, which carries no chain detour at all. The observatory was never the liar: its plan roots come from the rules, and a chain hop copy is stored only under its own tag, so no plan job could ever write under a base tag. The fix is therefore to remove the other two paths, not to touch the plan. TestGroups now asks the observatory for an out-of-turn pass and reports what it measured; a target no enabled rule routes to is not dialled at all and says so. Unused urltest groups stand down their own self-check via a new SelfCheck option (nil keeps today's behaviour, so every existing config is unchanged). The one direct dial left is the exit-address lookup, which has no other possible source — it now runs only for a target that is both routed and already read alive, so it travels the routed path and never touches an unused group. Chain hop wrappers are probed as measurements of their own and surfaced as chains[].hops[], because "which hop is dead" is the question an operator has and the chain-level verdict cannot answer it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bcc9df9282 |
test(wireguard): drive a real AmneziaWG tunnel through ClientBind
The unit tests pin the reserved-byte gate on each side in isolation, which would still pass if the two halves disagreed about when to apply it. This wires two real wireguard-go devices together over loopback UDP through ClientBind on both ends — the bind the detour path actually uses — configures ranged h1-h4 plus s4 and junk, and asserts an inner IP packet reaches the peer's TUN. It is red against the unconditional clear and green with the gate, so it covers the failure the field hit rather than the code we happened to write. Tagged with_awg, so it runs under the shipped router tag set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>apk-v0.2.13-aarch64_cortex-a53 apk-v0.2.13-x86_64 v0.2.13 |
||
|
|
ee3641fe45 |
fix(logsink): collapse interleaved floods, not just consecutive lines
The previous suppression compared each line with the one before it, which the field never obliges. A dead chain makes the engine cycle the same message across three outbound tags, so no two identical lines are adjacent: on the router it produced 854 daemon lines in a ~760-line syslog ring and exactly one summary, all while claiming "repeated 1 time". The rest of the system's log — netifd, dnsmasq, the kernel — was evicted anyway. Track a bounded table of open series keyed by the existing repeat key instead. The first copy of a key prints; further copies inside its window are counted whatever arrives in between; the window end emits one summary per key. The summary now names its message, because several can close at once and "last message" would simply be false under interleaving. The table holds 256 keys and evicts the least recently seen, never silently: an evicted series with a pending count prints its summary on the way out, marked so the truncation is visible. Close, Reconfigure and any fatal flush every open series first — a dying daemon may never reach Close. TestRepeatAlternatingNotSuppressed asserted that A B A B must never be collapsed. That assertion was the bug. It is replaced by a stronger one: the messages get separate series, separate summaries and separate counts, so distinct events still never fold into a single number. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
439f62238f |
fix(engine): retire the superseded instance instead of leaving it running
Every config apply built a new box and left the old one alive. The engine's own log gives it away: inside a single shaterd process, lines carried uptime counters half an hour apart in the same second, and a live router was found running four generations at once. A process restart cleared it, so the leak accrued purely on re-apply. That is not just wasted memory on a 512 MB box. Each surviving generation keeps its WireGuard devices up, and two devices sharing one private key evict each other at the peer — so the leak reproduced the duplicate-device defect between generations, underneath the deduplication that only reasons about one config. Retirement now has a hard budget: 5s, which is exactly sing-box's own C.StopTimeout (past which upstream already calls a stop excessive) and stays under C.FatalStopTimeout. It is paid after the replacement is serving and only on an apply that changed something, so a no-op reconcile stays free. A close that blows the budget is ABANDONED, not waited on, and the apply is still reported as the success it is — the new box is built, started and carrying traffic, and failing there would abort the netplane stage and leave a stale ruleset over a healthy engine. The stuck instance is surfaced through PendingCloses() into `shaterd status` and the panel, and clears itself if the shutdown ever completes. Repeated applies over a stuck close no longer stack: the abandoned generation is remembered, not re-created. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d971eb85ee |
fix(wireguard): stop ClientBind from shredding the AmneziaWG magic header
An AmneziaWG node worked standalone and died the moment it was placed behind an egress or a chain hop: the handshake completed, the peer answered, and then not one byte of data ever arrived. The peer never confirmed the session, so it re-handshook every 15 seconds, forever. ClientBind cleared bytes 1-3 of every datagram on receive and stamped them on send, unconditionally. Those bytes are Cloudflare's "reserved" field. They are also where AmneziaWG puts the upper three bytes of its little-endian uint32 magic header, so zeroing them collapses the value to its low byte, which falls outside every h1-h4 range and makes the peer classify the packet as an unknown type and drop it silently. Handshakes survived because s1/s2 padding pushes their magic past byte 3 — the clear only scribbled on the random junk prefix. Transport packets have s4 = 0, so their magic starts at byte 0 and took the hit. That asymmetry is the whole signature: session up locally, zero data through. Only the detour path was affected, because Endpoint.Start picks StdNetBind when the dialer exposes WireGuardControl (no detour) and ClientBind otherwise. The gate had already landed in StdNetBind; ClientBind was its untouched twin. The two implement one contract and are now commented as the pair they are, so the next fix cannot again land on one side only. Measured on the box: h4 spans 0x60728123-0x60728155, so zeroing bytes 1-3 leaves 35..85 — the captured transport packet began with 56, while a node without a detour carried a correct 0x6b039798 at the same moment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0f6083e28 |
fix(panel): show whether a rule is in force, not just what was saved
With two catch-all rules both enabled in UCI and a WAN profile enabling one and disabling the other, the panel drew BOTH switches on while the engine ran only one chain. GET /api/config is right to return the raw model — that is the desired state the panel PUTs back — but Routing.tsx read the row state and the active count from it too, so the interface claimed a setting was in force when it was not. Same defect class as the Protected badge. /api/rules/reachability now carries the effective flag and, where the active profile changed the outcome, its name and direction. The annotation is a DIFF of ApplyProfileRuleOverrides output against desired state rather than a second reading of the profiles name lists, so profile logic is not duplicated and cannot drift — an unmigrated rule the profile is forbidden to enable produces no diff and gets no badge, with nothing here needing to know about LegacyDst. In the UI the two states stay separate: the switch remains the only carrier of desired state and still writes UCI, while the effective state drives the dimmed row, the badge, the banner and the header count. Mirroring the effective state into the switch would be worse than the original bug — the operator would be toggling someone elses control, and the profiles decision would be written back as their own choice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGNapk-v0.2.12-aarch64_cortex-a53 apk-v0.2.12-x86_64 v0.2.12 |
||
|
|
77369aedfe |
fix(logsink): collapse repeated lines instead of erasing the routers syslog
A broken outbound makes the engine repeat one line about once a second — 370 copies in six minutes. The routers syslog ring holds ~760 lines, so within minutes it evicts the history of every other subsystem and our own startup lines with it. Diagnosing the WireGuard duplication above required restarting the service purely to catch the first seconds of a boot. Collapse runs into "last message repeated N times". The comparison key is level + text with the uptime field dropped: comparing whole lines would suppress only same-second bursts, because that counter ticks. The per connection "[id duration]" group is deliberately KEPT in the key — those ids are distinct connections, and folding "50 connections failed" into one count would be a worse lie than the flood. Window 5s, so a standing fault keeps being reported instead of looking like a frozen log. fatal/panic are never suppressed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
515ae6d1b7 |
test(generate): make the remote-blocklist test exercise the remote path
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two releases, which made the whole package exit non-zero no matter what the code did — a real regression would have drowned in the familiar red. The cause is not the packages no-network fetcher stub, as it first appears. ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so the fixture fell into the TEXT-list path: downloaded by generates own fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which every assertion below then contradicted. No stub content could fix that; the stub decides the lists contents, not the rule-sets type. Give the URL the .srs suffix the test always meant it to have, so the engine fetches the compiled set itself through the direct outbound. No assertion is weakened and the no-network stub stays in place. Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags): 338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
a2ffbb1292 |
fix(generate): one WireGuard device per private key
A node may be copied freely by this package: a per-chain hop copy and a per-group egress copy are rebuilt from the share-link so each can carry its own Detour. For vless that is right — a copy is another TCP client. For WireGuard it is not: each emitted endpoint is a real device holding the nodes private key, and a peer keeps exactly ONE session per public key. Two devices from one key evict each other continuously, and with keepalive on both the loop never settles: NEITHER passes traffic. buildOutboundsAndEndpoints emits the base endpoint for every enabled node whether or not anything references it, so a WG node used only as a chain hop always produced two devices. That is what any chain containing a WG node looks like — every such chain was permanently dead. Observed on the box: two UDP sockets from shaterd to the same peer port, the servers peer endpoint flapping between them, +32 bytes/min through the tunnel and every hop failing with "context deadline exceeded". Deduplicate once on the assembled options, which catches all three producer paths by construction. Duplicates are DELETED, not merely unreferenced: box.New starts every endpoint regardless of reachability, so a leftover would still bring its device up and still fight for the session. Dangling references go to block, never to direct — a consumer whose tunnel just disappeared must stop, not fall out onto the plain WAN. Subscription fetch detours seed the reachability walk (they are direct references like any rule), mirroring engine.ViaToTag exactly, with a tripwire test against drift. A config that genuinely needs two devices for one key keeps one and fail-closes the rest with a critical warning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
1746d4d0ef |
fix: stop the panel and the shipped binary from lying about what works
Four defects, all found by the owner on the live router, all of the same family: something declared itself working while it was not. WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave "create WireGuard device: gVisor is not included in this build". The router tag set carried with_wireguard and with_awg but not with_gvisor, so sing-tun compiled its stub instead of the netstack every WireGuard device needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving requirement", so this was a broken promise, not a trim. The tag itself was the small half. The tag set was the ONE build configuration nothing in the repo tested: TestAmneziaWGEndpoint passes because tests build with the full upstream tags. So the set now lives in one file (scripts/router-tags.sh) and two guards hold it to the feature list -- a static check that needs no tags, no Linux and no network (so the next such gap fails on the developer's machine), and a behavioural one that constructs every declared protocol through box.New UNDER THE SHIPPED TAGS, where skipping is forbidden. Removing the tag now fails with the feature name, the missing tag, and why: "Either add the tag back, or stop declaring the feature -- those are the only two honest options." Cost: +2.8 MB raw, +0.6-0.7 MB packed per arch. D23; D9 corrected. THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came from plane === 'full', which reports whether the data plane is installed -- nft table, policy routing, live engine -- and says nothing about where the traffic goes. On a config with one `default -> direct` rule and no groups the plane is fully installed and every packet leaves in the clear, so the worst possible state rendered as the reassuring one. The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the moment they reach the engine, not from the model: buildRoute changes the answer (a scheduled rule outside its window is never emitted, only the last condition-less rule reaches Final, an unresolved target is rewritten by ruleKillFallback), and re-deriving it anywhere else is a second implementation that will drift -- model/reachability.go exists because two already did. Four verdicts, not three: `blocked` is separate because under a closed kill-switch with no catch-all nothing leaks, and calling that "going out directly" is a lie in the alarm direction. Rider: Overview's defaultTarget printed the highest-Order enabled rule as the default; a rule becomes Final by having no conditions, whatever its Order. "PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE (B15). Once the browser suppresses dialogs, window.confirm returns false immediately, so all 15 confirmations across 7 pages read as "cancelled" and silently did nothing, with no way to recover from inside the panel. Replaced with an in-app dialog the browser cannot mute: focus trapped and parked on Cancel, Esc and veil cancel, focus returned to the opener, crit styling for destructive commits. useConfirm() throws if the provider is missing rather than falling back to a quiet false -- the failure mode being fixed. HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed, so a feed's nodes of those types vanished. The real landmine was one layer up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so without schemePrefixes the links were gone before any parser ran and the fix would have looked complete. Undeliverable parameters are refused when the node cannot work or would be less secure than the link asked (obfs, pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot produce a QUIC TLS config, and that fails at dial time, not at box.New. Also: nodes added by hand can be named and renamed. The name is the outbound tag, so a rename rewrites every reference in one PUT -- rule targets, group members, chain hops, detours -- in the spelling each already uses, and is refused outright when a group answers to the same bare name. Subscription nodes state why they cannot be renamed instead of hiding the control. go build, go vet, go test ./shater/... (13 packages), panel npm run build and npm test (13/13) all green. NOT yet verified on hardware. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGNapk-v0.2.11-aarch64_cortex-a53 apk-v0.2.11-x86_64 v0.2.11 |
||
|
|
f86501bf77 |
ci!: drop the opkg lane — apk only, and fix the stale rolling release
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has no `opkg` binary at all. The 24.10 lane was building and signing a feed no device could consume. Removed jobs `build` and `release` with the scripts only they called (ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh) and the usign trust anchor dist/shater-feed.pub. A committed public key is an instruction: it invites the old install path for a feed that is no longer produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD still holds the secret half, and a usign secret contains its own public half, so the identity is reconstructible if a 24.10 device ever needs serving. D7 is marked SUPERSEDED by the new D22 rather than deleted. Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from 2026-07-24 while every tag run published its versioned release correctly. The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest (workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run never touched the rolling pointer. Asset replacement was never the problem; ci/gitea-release.sh already deletes before recreating. A router pinned to the rolling URL sat on 0.2.0 while `apk update` reported success: silent staleness, the failure mode this repo keeps having to close. The rolling pointer is now published on EVERY run, tag runs included, and a new assert reads the release back over the API afterwards: our three tag-versioned packages at the built version plus the index and the key must be present (exit 13), and no package asset at any other version may survive (exit 14). Same class of check as sdk-build-apk.sh's package-version assert, added for the same reason -- the previous failure mode was silent. KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references it. Docs state plainly that mini_router is deliberately pinned to a versioned URL and that the hand-edit per release is the price of pinning. Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can no longer install our packages. Its 25.12 rebuild is in flight separately. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN |
||
|
|
eccfc6136c |
fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review ofapk-v0.2.10-aarch64_cortex-a53 apk-v0.2.10-x86_64 v0.2.10 |
||
|
|
244b7c4199 |
feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are gone from api.ts, so the Routing page loses the two controls that wrote them. The add form's Match picker (rulesets / ip / port) collapses to a plain Port(s) field beside the ruleset checkboxes — with no inline address list there was nothing left to choose between. The edit form drops its "Domain(s) — legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form shows, which is the honest shape of a rule that carries one destination mechanism. The destination picker renders even when the config has no rulesets yet, and says where to get one. Hiding it (the old behaviour when the list was empty) would leave the rule form with no destination control at all, at precisely the moment the user needs to know one exists. It is checkboxes and nothing more: creating and filling a list stays in the Rulesets panel, so a list is authored in one place and its naming and entry rules cannot drift between two editors. isCatchAll() drops the same two fields as model.IsCatchAll, so the "never applies" badge and the daemon's apply warning keep agreeing about which rule is the default; the matcher chips lose their `dns` and `ip` rows for the same reason. The mock backend's reachability shim follows. Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the removed wide inputs used, is deleted rather than left dangling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>apk-v0.2.9-aarch64_cortex-a53 apk-v0.2.9-x86_64 v0.2.9 |
||
|
|
a8ef887c56 |
feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.
`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.
THE BARE-ENTRY TRAP, and why the migration is not a copy
A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).
`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.
`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.
THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)
Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.
Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.
ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.
Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.
untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.
Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8fd5c52488 |
fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.
Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.
With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.
Fix (upstream file, lx:tproxy_writeback_connect):
* CONNECT the write-back socket to the one peer it ever talks to. The kernel's
compute_score() rejects a connected socket for any other peer, and a
connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
reuseport selection outright — so it can no longer be handed a datagram it
will not read. Nothing about the reply changes: same spoofed source, same
single peer, Write instead of WriteTo. The unconnected path is kept verbatim
for a destination that cannot be bound (domain socksaddr).
* A failed cached write now CLOSES the socket instead of only dropping the
reference (upstream left the fd to the GC finalizer).
* TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
but the cache evicts lazily, so after the inbound is gone nothing wakes the
live sessions and each strands its write-back socket. Invisible upstream
(one close at shutdown); on this fork the engine is rebuilt on every apply,
so it was one stranded generation per apply.
Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.
The netplane UDP:53 conntrack flush from
apk-v0.2.8-aarch64_cortex-a53
apk-v0.2.8-x86_64
v0.2.8
|
||
|
|
32e8f8ff0b |
fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.
Cause: `restart` is not synchronised end to end.
* procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
`stop; start`, and the `service delete` ubus call returns as soon as
SIGTERM has been SENT. `start_service` therefore re-adds the instance
(and runs `shaterd migrate`) while the outgoing `shaterd run` is still
executing its honest teardown.
* The successor's only defence was `daemonAlive()` -> exit(1), leaning on
procd's `respawn 3600 5 0` to try again five seconds later. That is a
blind retry, not synchronisation: it neither knows nor waits for the
teardown, and it turns every restart into a logged crash plus a
five-second hole with no data plane.
* `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
engine.Close of a several-hundred-outbound box flushes cache.db to
flash before the netplane teardown even starts — aborting the teardown
at an arbitrary point and leaving the plane HALF removed.
* Nothing in the tree ever touched conntrack, so flows that crossed one
of those windows kept entries formed against a plane that no longer
exists. For UDP there is no handshake to resynchronise on and every
retry merely refreshes the entry, so the flow stays wedged for as long
as the client keeps asking — a flow-scoped, permanent failure that no
ruleset rebuild can reach.
* RoutingPresent() reported "plane intact" from the ip RULE alone, while
ApplyRouting installs a rule AND a `local default dev lo` route removed
by two independent commands. A teardown interrupted between them was
therefore invisible, applyLocked's fast-path skipped ApplyRouting
forever, and no reconcile could repair it.
Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):
* init: `start_service` waits for a live predecessor pidfile to clear
before opening the instance, so restart == stop + pause + start. Zero
cost at boot. term_timeout 10 -> 30 so an honest teardown is never
killed halfway.
* daemon: the single-owner guard WAITS for the predecessor (bounded,
60s) instead of exiting 1; it still refuses if the budget expires.
* netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
a blanket flush would drop the admin's own SSH/LuCI sessions) called
on every plane transition: after a ruleset loads, after the table is
removed, and once more in applyLocked when the whole plane (table +
policy routing + sysctls) is assembled.
* netplane: RoutingPresent() now verifies both halves it installs.
Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.7-aarch64_cortex-a53
apk-v0.2.7-x86_64
v0.2.7
|
||
|
|
c562579ef3 |
docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute does not emit a condition-less rule as a match-all route rule: it sets route.Final and continues, so the LAST condition-less rule by order wins, and it can never shadow a rule that has conditions (those are emitted ahead of Final regardless of order). For the config on the router this inverts the conclusion: traffic IS going through the proxy (order=100 -> group:auto is the live default) and the dead knob is the order=20 `direct` one. Severity downgraded from high to medium accordingly — a dead setting, not a leak. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6f89acbae7 |
feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final", so a config with two of them showed two identical claims and no hint that only the last one is the default the router uses. A superseded rule now loses those marks — it keeps its real Order in the rail instead of the "·" that means final — and gains a "never applies" badge plus a line naming the rule that beat it and what to do about it: give this one a condition, or delete one of the two. Warn semantics throughout (--amber, dashed frame, dimmed target chip): orange is the ACTIVE state on this faceplate, and a rule the router ignores is the opposite of active. Verdicts come from GET /api/rules/reachability and are keyed by the rule's index in Rules, never by name — the config that prompted this had two rules both called `default`. They are re-fetched after every save, and a verdict whose echoed name/order no longer matches the row is dropped rather than shown, so the window between an optimistic edit and the refetch cannot badge a working rule. Rule rows were also keyed by name in React, which silently collapses two rows that share one; the key now carries the model index. The mock fixture gains a second condition-less rule so `?mock` renders the state, and mock.getRulesReachability derives its verdicts from the live fixture config rather than hard-coding them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88a82c7297 |
feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:
* two condition-less rules retire each other, and the LAST one by Order
wins, so an earlier "default -> direct" is dead while looking live;
* a condition-less rule can NEVER retire a rule that HAS conditions —
those are emitted ahead of Final whatever their Order.
A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.
model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.
Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.
Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.
Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.
GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.
Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a8f2b0f068 |
ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.
ci/version.sh is now the single source of truth. It derives the version
from `git describe`:
tag `vX.Y.Z` -> PKG_VERSION=X.Y.Z PKG_RELEASE=1
off-tag build -> nearest tag + PKG_RELEASE=<commits since it> + 1
no tag/no git -> 0.0.0-r1 (below everything ever published)
Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.
The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.
The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.
byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.
Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0b32a6d58b |
fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:
daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
[ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...
`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.
Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:
* control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
DisableColors at its false zero value. It now comes from
controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
lives in an untagged file (same split as profilewatch.go) so it is
unit-testable off the linux target.
* engine (shater/generate): the generated option.LogOptions never set
DisableColor, so box.New built a colouring formatter over the shared
sink. logOptions() now sets it from the same TTY gate (seam:
logColorAllowed).
logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.
The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.
Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
893fdc500c |
fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as
|
||
|
|
02c266188f |
docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the previous install, install from the signed apk feed, verify the default state, restore a working config with 315 subscription nodes, then exercise the data plane, panel API, config lifecycle, resilience and DNS. 74 PASS. Findings (detailed separately): two catch-all `default` rules where the first sends all unspecific traffic direct and makes the second unreachable; `shaterd nodes` is a stub returning [] while usage promises the node list; DNS to the router LAN address dies after `service shater restart` (stop+pause+start is fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI colour codes reach syslog. Also records the four-iteration CI hunt that ended in the green apk lane. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
024e9308c9 |
fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.
The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:
config PACKAGE_kmod-mlx5-core
tristate
default m
No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.
The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.
Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.6-aarch64_cortex-a53
apk-v0.2.6-x86_64
v0.2.6
|
||
|
|
eab1db2c3f |
fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (
v0.2.5
|
||
|
|
22d7161c08 |
fix(ci/apk): disable the SDK's ALL/ALL_KMODS mass-select
release / aarch64_cortex-a53 (push) Successful in 3m25s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m3s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.
The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:
config ALL_NONSHARED ... default ALL
config ALL_KMODS ... default ALL
config ALL ... default y
In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.
Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.
Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.
Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v0.2.4
|
||
|
|
24c5a1615d |
fix(ci/apk): build .config from scratch — stop packing all 3593 kmods
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m18s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m2s
release / release (push) Successful in 7s
release / release apk (push) Successful in 5s
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn — none of which we ship) before it ever got near our four. Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the .config that ships inside the ImmortalWrt SDK tarball. That file is the buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so `make defconfig` re-selected every kernel module of the target as =m and package/kernel/linux/compile — pulled in via shater-core's nft kmod deps — packed the lot. Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds": start the .config EMPTY so kconfig can only pull in what our packages actually select. Carried over from the SDK's .config, nothing more: the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a wrong guess here means silently cross-compiling for another arch), CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel). Also adds the diagnostics this lane never had, since a failed run leaves a 27 MB log: the carried-over identity lines, the post-defconfig kmod count and target readout, a hard check that all four of our packages survived defconfig, an abort if the kmod count is back in the hundreds, and du/df after compile. opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR, CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>v0.2.3 |
||
|
|
4492f0599c |
ci: harden feed/packaging shell scripts
release / aarch64_cortex-a53 (push) Successful in 7m3s
release / x86_64 (push) Successful in 3m13s
release / apk aarch64_cortex-a53 (push) Failing after 4m34s
release / apk x86_64 (push) Failing after 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in the signed Packages index can't mask an empty SHA256. Script survives -u (all vars use :? or :- defaults). - .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a trap-based tmpdir cleanup, and set -euo pipefail. bash -n clean on both. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>v0.2.2 |
||
|
|
bc2b53069a |
fix(security): audit remediation — file perms, const-time auth, leaks, CSPRNG
Backend audit fixes (upstream-file edits wrapped in // lx: markers): - experimental/libbox oom_report.go/report.go: OOM reports + configuration.json (server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files → 0o700 / 0o600. [sec-perms] - daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare. [sec-consttime] - service/oomkiller/timer.go: network-extension cleanupTriggered logic was inverted, so FreeOSMemory was never called after a trigger; flip both assignments so a trigger schedules the deferred free and the next poll runs + clears it. [sec-oomcleanup] - transport/v2rayxhttp/client.go (lx-native file): session id used math/rand → crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface. - daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a goroutine + the ssh-agent fd on every closed session (second io.Copy blocked on an idle agent Read forever); tie both copies + the session ctx to a cancel that closes both ends. [sec-sshagent] - daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to 1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate] - route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in select while stopIdleSuspend niled it after close (race + goroutine leak on Close-during-tick); pass the stop channel to the loop by value. go build ./... (default) and the D9 shaterd linux build (tags with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp, with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/... ./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend and v2rayxhttp with with_xhttp. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c257d6c5cc |
docs(readme): rewrite root README as Russian shater product face
- README.md: new Russian product README (what/features/architecture mermaid/install both feeds/build/repo layout/CI/upstream/docs/license) - README.en.md: concise English mirror (root readme was previously English) - README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme, a competing Russian README) -> points to README.md + engine-fork docs - docs-shater/README.md: folder index Install commands copied verbatim from docs-shater/INSTALL.md; all links verified against existing files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
225cce5397 |
chore: purge working-session junk from tree; gitignore recurrence
Remove untracked-quality artifacts accidentally committed during work sessions (all authored downstream, unreferenced anywhere in code/docs/CI): - 5 session screenshots in repo root (devices-after-copy-fix.png, live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png) - tmp/gen_linux_test (29 MB throwaway traffic-gen binary) Guard against repeats: ignore /*.png (root screenshots) and /tmp/. Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |