Commit Graph
3027 Commits
Author SHA1 Message Date
omarandClaude Opus 5 a0f6083e28 fix(panel): show whether a rule is in force, not just what was saved
release / apk aarch64_cortex-a53 (push) Successful in 3m7s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 8s
With two catch-all rules both enabled in UCI and a WAN profile enabling one
and disabling the other, the panel drew BOTH switches on while the engine
ran only one chain. GET /api/config is right to return the raw model — that
is the desired state the panel PUTs back — but Routing.tsx read the row
state and the active count from it too, so the interface claimed a setting
was in force when it was not. Same defect class as the Protected badge.

/api/rules/reachability now carries the effective flag and, where the active
profile changed the outcome, its name and direction. The annotation is a
DIFF of ApplyProfileRuleOverrides output against desired state rather than a
second reading of the profiles name lists, so profile logic is not
duplicated and cannot drift — an unmigrated rule the profile is forbidden to
enable produces no diff and gets no badge, with nothing here needing to know
about LegacyDst.

In the UI the two states stay separate: the switch remains the only carrier
of desired state and still writes UCI, while the effective state drives the
dimmed row, the badge, the banner and the header count. Mirroring the
effective state into the switch would be worse than the original bug — the
operator would be toggling someone elses control, and the profiles decision
would be written back as their own choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
apk-v0.2.12-aarch64_cortex-a53 apk-v0.2.12-x86_64 v0.2.12
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 77369aedfe fix(logsink): collapse repeated lines instead of erasing the routers syslog
A broken outbound makes the engine repeat one line about once a second —
370 copies in six minutes. The routers syslog ring holds ~760 lines, so
within minutes it evicts the history of every other subsystem and our own
startup lines with it. Diagnosing the WireGuard duplication above required
restarting the service purely to catch the first seconds of a boot.

Collapse runs into "last message repeated N times". The comparison key is
level + text with the uptime field dropped: comparing whole lines would
suppress only same-second bursts, because that counter ticks. The per
connection "[id duration]" group is deliberately KEPT in the key — those ids
are distinct connections, and folding "50 connections failed" into one count
would be a worse lie than the flood. Window 5s, so a standing fault keeps
being reported instead of looking like a frozen log. fatal/panic are never
suppressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 515ae6d1b7 test(generate): make the remote-blocklist test exercise the remote path
TestDNSFilterRemoteBlocklistHTTPClient has failed on every Linux run for two
releases, which made the whole package exit non-zero no matter what the code
did — a real regression would have drowned in the familiar red.

The cause is not the packages no-network fetcher stub, as it first appears.
ruleSetURLIsEngineNative decides remote-vs-compiled-local by URL EXTENSION
alone, and httptest.NewServers bare "http://127.0.0.1:<port>" has none, so
the fixture fell into the TEXT-list path: downloaded by generates own
fetcher, parsed as a hosts file, compiled into a LOCAL rule-set — which
every assertion below then contradicted. No stub content could fix that; the
stub decides the lists contents, not the rule-sets type.

Give the URL the .srs suffix the test always meant it to have, so the engine
fetches the compiled set itself through the direct outbound. No assertion is
weakened and the no-network stub stays in place.

Verified on the stand (ImmortalWrt 25.12.1 x86_64, shipped build tags):
338 PASS / 0 FAIL / 1 SKIP, exit 0 — against 327/1/1 on pristine HEAD.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 a2ffbb1292 fix(generate): one WireGuard device per private key
A node may be copied freely by this package: a per-chain hop copy and a
per-group egress copy are rebuilt from the share-link so each can carry its
own Detour. For vless that is right — a copy is another TCP client. For
WireGuard it is not: each emitted endpoint is a real device holding the
nodes private key, and a peer keeps exactly ONE session per public key.
Two devices from one key evict each other continuously, and with keepalive
on both the loop never settles: NEITHER passes traffic.

buildOutboundsAndEndpoints emits the base endpoint for every enabled node
whether or not anything references it, so a WG node used only as a chain hop
always produced two devices. That is what any chain containing a WG node
looks like — every such chain was permanently dead.

Observed on the box: two UDP sockets from shaterd to the same peer port, the
servers peer endpoint flapping between them, +32 bytes/min through the
tunnel and every hop failing with "context deadline exceeded".

Deduplicate once on the assembled options, which catches all three producer
paths by construction. Duplicates are DELETED, not merely unreferenced:
box.New starts every endpoint regardless of reachability, so a leftover
would still bring its device up and still fight for the session. Dangling
references go to block, never to direct — a consumer whose tunnel just
disappeared must stop, not fall out onto the plain WAN.

Subscription fetch detours seed the reachability walk (they are direct
references like any rule), mirroring engine.ViaToTag exactly, with a
tripwire test against drift. A config that genuinely needs two devices for
one key keeps one and fail-closes the rest with a critical warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 21:35:06 +03:00
omarandClaude Opus 5 1746d4d0ef fix: stop the panel and the shipped binary from lying about what works
release / apk aarch64_cortex-a53 (push) Successful in 9m13s
release / apk x86_64 (push) Successful in 3m4s
release / release apk (push) Successful in 7s
Four defects, all found by the owner on the live router, all of the same
family: something declared itself working while it was not.

WIREGUARD WAS DEAD IN THE SHIPPED BINARY (B17). Setting up WireGuard gave
"create WireGuard device: gVisor is not included in this build". The router
tag set carried with_wireguard and with_awg but not with_gvisor, so
sing-tun compiled its stub instead of the netstack every WireGuard device
needs. FEATURES.md marks WireGuard [MVP] and AmneziaWG "a driving
requirement", so this was a broken promise, not a trim.

The tag itself was the small half. The tag set was the ONE build
configuration nothing in the repo tested: TestAmneziaWGEndpoint passes
because tests build with the full upstream tags. So the set now lives in
one file (scripts/router-tags.sh) and two guards hold it to the feature
list -- a static check that needs no tags, no Linux and no network (so the
next such gap fails on the developer's machine), and a behavioural one that
constructs every declared protocol through box.New UNDER THE SHIPPED TAGS,
where skipping is forbidden. Removing the tag now fails with the feature
name, the missing tag, and why: "Either add the tag back, or stop declaring
the feature -- those are the only two honest options." Cost: +2.8 MB raw,
+0.6-0.7 MB packed per arch. D23; D9 corrected.

THE PANEL CALLED A DIRECT-ONLY ROUTER "PROTECTED" (B16). The headline came
from plane === 'full', which reports whether the data plane is installed --
nft table, policy routing, live engine -- and says nothing about where the
traffic goes. On a config with one `default -> direct` rule and no groups
the plane is fully installed and every packet leaves in the clear, so the
worst possible state rendered as the reassuring one.

The verdict is now computed on the daemon FROM THE GENERATED OPTIONS at the
moment they reach the engine, not from the model: buildRoute changes the
answer (a scheduled rule outside its window is never emitted, only the last
condition-less rule reaches Final, an unresolved target is rewritten by
ruleKillFallback), and re-deriving it anywhere else is a second
implementation that will drift -- model/reachability.go exists because two
already did. Four verdicts, not three: `blocked` is separate because under
a closed kill-switch with no catch-all nothing leaks, and calling that
"going out directly" is a lie in the alarm direction. Rider: Overview's
defaultTarget printed the highest-Order enabled rule as the default; a rule
becomes Final by having no conditions, whatever its Order.

"PREVENT THIS PAGE FROM CREATING ADDITIONAL DIALOGS" KILLED EVERY DELETE
(B15). Once the browser suppresses dialogs, window.confirm returns false
immediately, so all 15 confirmations across 7 pages read as "cancelled" and
silently did nothing, with no way to recover from inside the panel. Replaced
with an in-app dialog the browser cannot mute: focus trapped and parked on
Cancel, Esc and veil cancel, focus returned to the opener, crit styling for
destructive commits. useConfirm() throws if the provider is missing rather
than falling back to a quiet false -- the failure mode being fixed.

HYSTERIA2 AND TUIC NODES WERE DROPPED (B6). No share-link parser existed,
so a feed's nodes of those types vanished. The real landmine was one layer
up: ParseSubscriptionBody splits a feed by scheme prefix before parsing, so
without schemePrefixes the links were gone before any parser ran and the
fix would have looked complete. Undeliverable parameters are refused when
the node cannot work or would be less secure than the link asked (obfs,
pinSHA256, tuic v4/non-UUID) and flagged via Proxy.Warnings when it
survives -- shaterd nodes shows both. uTLS is dropped for QUIC: it cannot
produce a QUIC TLS config, and that fails at dial time, not at box.New.

Also: nodes added by hand can be named and renamed. The name is the
outbound tag, so a rename rewrites every reference in one PUT -- rule
targets, group members, chain hops, detours -- in the spelling each already
uses, and is refused outright when a group answers to the same bare name.
Subscription nodes state why they cannot be renamed instead of hiding the
control.

go build, go vet, go test ./shater/... (13 packages), panel npm run build
and npm test (13/13) all green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
apk-v0.2.11-aarch64_cortex-a53 apk-v0.2.11-x86_64 v0.2.11
2026-07-25 20:08:58 +03:00
omarandClaude Opus 5 f86501bf77 ci!: drop the opkg lane — apk only, and fix the stale rolling release
Both routers are past opkg: mini_router runs ImmortalWrt 25.12.1 and
main_router OpenWrt 25.12.0, both with apk-tools 3.0.5, and main_router has
no `opkg` binary at all. The 24.10 lane was building and signing a feed no
device could consume.

Removed jobs `build` and `release` with the scripts only they called
(ci/build-feed.sh, ci/sdk-build.sh, ci/make-index.sh, ci/install-usign.sh)
and the usign trust anchor dist/shater-feed.pub. A committed public key is
an instruction: it invites the old install path for a feed that is no longer
produced. The key is retired, not revoked -- git history keeps it, KEY_BUILD
still holds the secret half, and a usign secret contains its own public half,
so the identity is reconstructible if a 24.10 device ever needs serving.
D7 is marked SUPERSEDED by the new D22 rather than deleted.

Separately: the rolling `apk-latest-<arch>` release was frozen at 0.2.0 from
2026-07-24 while every tag run published its versioned release correctly.
The publish loop was an either/or -- `TAG=apk-latest-<arch>` when VER=latest
(workflow_dispatch only), ELSE `TAG=apk-<ver>-<arch>` -- so a `v*` tag run
never touched the rolling pointer. Asset replacement was never the problem;
ci/gitea-release.sh already deletes before recreating. A router pinned to
the rolling URL sat on 0.2.0 while `apk update` reported success: silent
staleness, the failure mode this repo keeps having to close.

The rolling pointer is now published on EVERY run, tag runs included, and a
new assert reads the release back over the API afterwards: our three
tag-versioned packages at the built version plus the index and the key must
be present (exit 13), and no package asset at any other version may survive
(exit 14). Same class of check as sdk-build-apk.sh's package-version assert,
added for the same reason -- the previous failure mode was silent.

KEY_BUILD can now be deleted from the Gitea repo secrets; nothing references
it. Docs state plainly that mini_router is deliberately pinned to a
versioned URL and that the hand-edit per release is the price of pinning.

Known consequence: the x86_64 QEMU testbed is still OpenWrt 24.10.3 and can
no longer install our packages. Its 25.12 rebuild is in flight separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
2026-07-25 18:46:21 +03:00
omarandClaude Opus 5 eccfc6136c fix(routing)!: make the v1->v2 destination migration fail safe
release / aarch64_cortex-a53 (push) Successful in 3m56s
release / x86_64 (push) Successful in 3m25s
release / apk aarch64_cortex-a53 (push) Successful in 2m44s
release / apk x86_64 (push) Successful in 2m43s
release / release (push) Successful in 9s
release / release apk (push) Successful in 7s
Code review of a8ef887c5 + 244b7c419 ("a rule's destination is a rule-set,
and nothing else") found that the change rested on a comment that was not
true. ParseUCIExport dropped dst_domain/dst_ip on the strength of "the
migration is re-run on every load"; model.Migrate() actually runs only from
`shaterd migrate`, i.e. the service init and uci-defaults. The daemon's run
path, the SIGHUP reconcile and the panel's config write never migrate.

So an uncommitted migration (a full /overlay is the documented way that
happens) turned `list dst_domain 'bank.ru'` + `target direct` into a rule
with NO matchers, which IS the spelling of a catch-all: generate points
route.Final at it and the LAST such rule wins. One failed `uci commit` sent
every packet on the router out the plain WAN, silently.

Rule.LegacyDst is the tripwire. It is non-empty exactly when the config
still carries the removed options, and three locks hang off it:
  - ParseUCIExport holds such a rule DISABLED. Chosen over "make IsCatchAll
    false" alone, which only covers matcher-less rules: `dst_domain` plus a
    `src` was never a catch-all, and routing it without its destination
    would still have sent a whole subnet direct.
  - IsCatchAll returns false for it, so it can never own route.Final even
    if something hands its Enabled bit back.
  - ApplyProfileRuleOverrides refuses to enable it (a profile with
    `list enable_rule` would otherwise have defeated the parser).
ValidateRules reports it through the existing warning channel, before the
Enabled gate, so the one message explaining the outage is not suppressed by
the fact that caused it. The init script logs a failed migration to syslog
instead of discarding its exit code and stderr.

The write path had none of this. PUT /api/config decodes a Model straight
from the request body and render.go wrote `enabled` from it, so a panel
save erased the operator's lists (as did the subscription cron, which
re-renders the whole package), and a crafted body with Enabled:true and no
LegacyDst put a live matcher-less rule on disk -- the same whole-router
leak, re-entered from the other side. WriteUCI now reads DISK state and
refuses a rule-changing write over an unmigrated config (409, not 500);
non-rule writers pass and legacyDstOpts carries the options across so cron
preserves them; withDiskLegacyDst takes the field from disk so a fabricated
one can never reach the renderer.

Migration hardening: an entry list that migrates to nothing no longer has
its legacy option deleted (that made "matches nothing" silently become
"matches everything"); a hand-written rule-set whose name collides is no
longer allowed to swallow the entries; delete failures propagate instead of
bumping schema_version past them forever; every error path reverts the
staged uci delta so another process's commit cannot flush a half-migration.

untunnelable.go follows the destination out of the rule: a rule whose
rule-sets are known to match by name is still skipped by the ping/IPTV/VPN
plan, as its v1 form was. D21 documents the AND->OR widening for the
engine's TCP/UDP path; it does not follow that a leak-guard should widen
itself during an upgrade, and with target=direct that meant previously
tunnelled ICMP leaving with the client's real address. Inline rule-sets are
now read from the options, so an engine that has not started yet no longer
costs the operator their ping.

Rule-set vocabulary: `full:`/`suffix:`/`keyword:`/`regexp:` in a text list
fetched by URL were dropped with no diagnostic at all (normaliseListDomain
rejects any token with a colon) -- not "reported as an unknown prefix".
Unifying was rejected: published filter lists are full of colon-bearing
syntax, and a third-party `regexp:` is compiled into the router's matcher
and run per query. The difference stands and is paid for in diagnostics,
per list, on every generate. D21 gains the source/vocabulary table.

Panel: the add form warns about a matcher-less rule exactly as the edit
form does, from one shared predicate; its isCatchAll matches the daemon's
new one; an unmigrated rule reads as held-off rather than merely switched
off. The comment promising a "New list" button that D21 rejected is gone.

go build ./..., go vet ./shater/..., go test ./shater/... (13 packages) and
panel `npm run build` are green. NOT yet verified on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4PcWfrBRyg4eWN58axaGN
apk-v0.2.10-aarch64_cortex-a53 apk-v0.2.10-x86_64 v0.2.10
2026-07-25 18:04:21 +03:00
omarandClaude Opus 5 244b7c4199 feat(panel): a rule's destination is a ruleset picker, nothing else
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m19s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 9s
release / release apk (push) Successful in 6s
Follows the schema-v2 model change: `Rule.DstDomain` and `Rule.DstIP` are
gone from api.ts, so the Routing page loses the two controls that wrote them.

The add form's Match picker (rulesets / ip / port) collapses to a plain
Port(s) field beside the ruleset checkboxes — with no inline address list
there was nothing left to choose between. The edit form drops its "Domain(s)
— legacy" and "IP / CIDR(s)" fields; it now shows exactly what the add form
shows, which is the honest shape of a rule that carries one destination
mechanism.

The destination picker renders even when the config has no rulesets yet, and
says where to get one. Hiding it (the old behaviour when the list was empty)
would leave the rule form with no destination control at all, at precisely
the moment the user needs to know one exists. It is checkboxes and nothing
more: creating and filling a list stays in the Rulesets panel, so a list is
authored in one place and its naming and entry rules cannot drift between two
editors.

isCatchAll() drops the same two fields as model.IsCatchAll, so the "never
applies" badge and the daemon's apply warning keep agreeing about which rule
is the default; the matcher chips lose their `dns` and `ip` rows for the same
reason. The mock backend's reachability shim follows.

Rendered against `?mock` in both themes; `.rt-field-wide`, the only rule the
removed wide inputs used, is deleted rather than left dangling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.9-aarch64_cortex-a53 apk-v0.2.9-x86_64 v0.2.9
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 a8ef887c56 feat(routing)!: a rule's destination is a rule-set, and nothing else
`config rule` carried THREE ways to say where traffic is going: `dst_domain`
(an inline domain list), `dst_ip` (an inline CIDR list) and `dst_ruleset` (a
reference to a `config ruleset`). Three mechanisms meant three sets of
semantics to keep straight, and the inline pair was the worse half of the
trade: re-parsed per rule instead of compiled once into a .srs, unshareable
between rules, and — invisibly — already disagreeing with the rule-set
vocabulary about what a bare entry means.

`dst_domain` and `dst_ip` are removed (schema v2). `dst_ruleset` is the only
destination matcher. `Src` (the client side), `dst_port` and `proto` are
untouched: they are not lists of destinations and have no rule-set form.

THE BARE-ENTRY TRAP, and why the migration is not a copy

A bare `example.com` was an EXACT host in a routing rule (classified with
bareIsSuffix=false) and is the host AND its subdomains inside a rule-set
(bareIsSuffix=true). Copying entries across verbatim would silently widen
every such rule to every subdomain, so migrate1to2 rewrites a bare entry as
`full:example.com`. Everything else already means the same on both sides and
is copied byte-for-byte: `full:`, `suffix:`, `keyword:`, `regexp:` and a
leading dot (a synonym of `suffix:`).

`geosite:`/`geoip:` entries are copied UNCHANGED rather than promoted to a
`source=geosite` rule-set. They have been inert since the engine dropped the
route-rule geosite/geoip fields, and an unrecognised marker is equally inert
inside a rule-set — so their meaning is preserved exactly, and a dead matcher
does not start routing traffic because someone upgraded. The text is kept so
the operator can see it and convert it deliberately.

`regexp:` had no rule-set form at all, which would have made the move lossy,
so inline rule-sets learn it: peelDomainRegexes validates each pattern with
regexp.Compile before it reaches DomainRegex, because
route/rule.NewDomainRegexItem errors on an uncompilable one and that aborts
box.New for the whole config. A bare `regexp:` is dropped too — it compiles
fine and matches every host.

THE MIGRATION (schema v1 -> v2, run by `shaterd migrate` on service start and
at package install)

Per rule still carrying a legacy list: create an inline `config ruleset`
named `rule-<rule name>` (domains) and/or `rule-<rule name>-ip` (addresses),
move the entries across with the conversion above, append the new name to
`dst_ruleset`, delete the old option LAST. It is idempotent; it resumes an
interrupted run by reusing a rule-set the rule already references; and it
never overwrites a hand-written list that owns the generated name (it takes
`rule-<name>-2`). The uci sequence — `uci add` capturing the section id, then
set/add_list/delete — was verified against BananaWRT 25.12.1 in a throwaway
package.

Verified against the live router's config (4 rules, 26 entries, all
`suffix:`): every entry lands in its rule-set, every rule gains exactly one
reference, the `default` rule stays condition-less so B1's RuleReachability
still reads it as the catch-all.

ONE DELIBERATE SEMANTIC CHANGE, stated out loud: a rule that used BOTH lists
matched them with AND (an engine route rule ANDs its matcher fields), which
is almost never what "these sites and these networks" meant. The two
generated rule-sets are ORed, because `rule_set: [a, b]` matches when either
matches. Only configs that used both fields at once are affected.

Also fixed here, because schema v2 routes EVERY destination list through
inlineRulesetRule and the gap widens accordingly: a marker-only entry (".",
"full:", "keyword:") was dropped by the shared classifier SILENTLY on that
path, where the routing rule used to warn. An empty domain token aborts
box.New and an empty keyword is strings.Contains(host, "") — every host — so
the drop is right and the silence was not.

untunnelable stays honest: buildUntunnelablePlan already resolves `rule_set`
addresses through the running engine (inline sets are LocalRuleSets and
implement ExtractIPSet), and apply runs eng.Apply before building the plan.
A migrated `dst_ip` therefore resolves exactly as before; with the engine
down the walk truncates and denies, which is the conservative direction and
the state in which the netplane is fail-closed anyway.

Tests: migration coverage (real-router fixture, mixed prefixes, CIDRs,
idempotence, interrupted-run resume, name collision, geo markers stay inert,
absent config), and every matcher-classification test that used to live on
`dst_domain`/`dst_ip` moved to the inline rule-set rather than deleted —
including the new `regexp:` path and the inverted bare-entry convention. The
model tests grow a real in-memory uci emulator so a second migration run
actually sees its own writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 13:57:16 +03:00
omarandClaude Opus 5 8fd5c52488 fix(tproxy): connect the UDP write-back socket + release NAT sessions on close
release / aarch64_cortex-a53 (push) Successful in 3m29s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 2m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
B3, real root cause. On the live BPi-R3 Mini `netstat -lnup` showed shaterd
holding 33 sockets on the router's own LAN address 10.67.0.1:53, next to
dnsmasq's single socket, several with a growing Recv-Q. Reproduced read-only on
the box: 5 host queries to 10.67.0.1 -> 0 answers and total Recv-Q on those
sockets 0 -> 19200 (5 x 3840, one datagram parked in each, never read); 3
control queries to 127.0.0.1 -> all answered.

Where they come from: protocol/redirect/tproxy.go, tproxyPacketWriter.
WritePacket. The TPROXY UDP write-back socket must carry the ORIGINAL
DESTINATION as its source address, so upstream binds it there — but leaves it
UNCONNECTED (net.ListenPacket + WriteToUDPAddrPort) and sets SO_REUSEADDR AND
SO_REUSEPORT (sing's control.ReuseAddr sets both). An unconnected bound socket
is a RECEIVER as far as the kernel is concerned, so each one silently joins the
UDP demultiplex/reuseport set for that address:port. Nothing ever reads them —
this writer only sends.

With dns_intercept the original destination IS the router's LAN address, so
every intercepted DNS session parks another silent receiver on <lan-ip>:53. The
host's own queries to that address take the loopback path, are never diverted by
the nft plane (iifname is scoped to LAN devices), and are therefore spread across
that set by the reuseport 4-tuple hash: they land in a silent socket at random
and time out. Hence "2 restarts of 3 fine, the third dead", and hence a failure
that no ruleset rebuild or reconcile can touch. The stale [UNREPLIED] conntrack
entry seen alongside is a CONSEQUENCE of the unanswered query, not the cause.

Fix (upstream file, lx:tproxy_writeback_connect):
  * CONNECT the write-back socket to the one peer it ever talks to. The kernel's
    compute_score() rejects a connected socket for any other peer, and a
    connected UDP socket (sk_state == TCP_ESTABLISHED) is excluded from
    reuseport selection outright — so it can no longer be handed a datagram it
    will not read. Nothing about the reply changes: same spoofed source, same
    single peer, Write instead of WriteTo. The unconnected path is kept verbatim
    for a destination that cannot be bound (domain socksaddr).
  * A failed cached write now CLOSES the socket instead of only dropping the
    reference (upstream left the fd to the GC finalizer).
  * TProxy.Close() purges the UDP NAT cache. Closing the listener stops ingress
    but the cache evicts lazily, so after the inbound is gone nothing wakes the
    live sessions and each strands its write-back socket. Invisible upstream
    (one close at shutdown); on this fork the engine is rebuilt on every apply,
    so it was one stranded generation per apply.

Measured on the live box: the socket count is steady-state (22-40, fds 55-66),
i.e. bounded by the udpnat session lifetime rather than an unbounded leak — the
count itself is inherent to per-session write-back sockets and is harmless once
they are connected. The Close() purge removes the per-apply generations on top
of it.

The netplane UDP:53 conntrack flush from 32e8f8ff0 is KEPT, with its comment
corrected: it is hygiene on plane transitions, not the cure for B3.

Regression tests fail on the pre-fix code (verified by reverting each half):
TestWriteBackUsesConnectedSocket / TestWriteBackReusesOneSocket /
TestWriteBackClosesSocketOnWriteFailure ("use of WriteTo with pre-connected
connection") and TestTProxyCloseReleasesNatSessions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.8-aarch64_cortex-a53 apk-v0.2.8-x86_64 v0.2.8
2026-07-25 13:41:02 +03:00
omarandClaude Opus 5 32e8f8ff0b fix(restart): serialise stop->start and flush stale DNS conntrack (B3)
release / aarch64_cortex-a53 (push) Successful in 6m27s
release / x86_64 (push) Successful in 3m21s
release / apk aarch64_cortex-a53 (push) Successful in 5m38s
release / apk x86_64 (push) Successful in 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
`/etc/init.d/shater restart` left DNS to the router's own LAN address dead
and never recovering, while `stop` + pause + `start` was fine — with the
status still reporting plane=full / engine_running=true and `shaterd
reconcile` fixing nothing.

Cause: `restart` is not synchronised end to end.

  * procd's `stop` is ASYNCHRONOUS. rc.common's `restart` is literally
    `stop; start`, and the `service delete` ubus call returns as soon as
    SIGTERM has been SENT. `start_service` therefore re-adds the instance
    (and runs `shaterd migrate`) while the outgoing `shaterd run` is still
    executing its honest teardown.
  * The successor's only defence was `daemonAlive()` -> exit(1), leaning on
    procd's `respawn 3600 5 0` to try again five seconds later. That is a
    blind retry, not synchronisation: it neither knows nor waits for the
    teardown, and it turns every restart into a logged crash plus a
    five-second hole with no data plane.
  * `term_timeout 10` SIGKILLs a predecessor whose teardown outlives it —
    engine.Close of a several-hundred-outbound box flushes cache.db to
    flash before the netplane teardown even starts — aborting the teardown
    at an arbitrary point and leaving the plane HALF removed.
  * Nothing in the tree ever touched conntrack, so flows that crossed one
    of those windows kept entries formed against a plane that no longer
    exists. For UDP there is no handshake to resynchronise on and every
    retry merely refreshes the entry, so the flow stays wedged for as long
    as the client keeps asking — a flow-scoped, permanent failure that no
    ruleset rebuild can reach.
  * RoutingPresent() reported "plane intact" from the ip RULE alone, while
    ApplyRouting installs a rule AND a `local default dev lo` route removed
    by two independent commands. A teardown interrupted between them was
    therefore invisible, applyLocked's fast-path skipped ApplyRouting
    forever, and no reconcile could repair it.

Fix (fail-closed posture unchanged — no new window in which LAN traffic can
reach the WAN; teardown still removes the table LAST and the forward-chain
drop is untouched):

  * init: `start_service` waits for a live predecessor pidfile to clear
    before opening the instance, so restart == stop + pause + start. Zero
    cost at boot. term_timeout 10 -> 30 so an honest teardown is never
    killed halfway.
  * daemon: the single-owner guard WAITS for the predecessor (bounded,
    60s) instead of exiting 1; it still refuses if the budget expires.
  * netplane: new FlushDNSConntrack() (ctnetlink, UDP orig-dport 53 only —
    a blanket flush would drop the admin's own SSH/LuCI sessions) called
    on every plane transition: after a ruleset loads, after the table is
    removed, and once more in applyLocked when the whole plane (table +
    policy routing + sysctls) is assembled.
  * netplane: RoutingPresent() now verifies both halves it installs.

Regression tests fail on the pre-fix code (verified by reverting each fix):
TestApplyNftFlushesDNSConntrack, TestTeardownNftFlushesDNSConntrack,
TestRoutingPresentRequiresLocalDefaultRoute, TestWaitForPredecessor*.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.7-aarch64_cortex-a53 apk-v0.2.7-x86_64 v0.2.7
2026-07-25 12:51:19 +03:00
omarandClaude Opus 5 c562579ef3 docs(report): correct the B1 diagnosis — last catch-all wins, not the first
The report claimed the order=20 `default` shadowed the order=100 one and sent all
unspecific traffic past the proxy. That is wrong. generate/route.go:buildRoute
does not emit a condition-less rule as a match-all route rule: it sets
route.Final and continues, so the LAST condition-less rule by order wins, and it
can never shadow a rule that has conditions (those are emitted ahead of Final
regardless of order).

For the config on the router this inverts the conclusion: traffic IS going
through the proxy (order=100 -> group:auto is the live default) and the dead knob
is the order=20 `direct` one. Severity downgraded from high to medium
accordingly — a dead setting, not a leak.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:41:58 +03:00
omarandClaude Opus 5 6f89acbae7 feat(panel): badge routing rules that never apply
The Routing page drew every condition-less rule as "default route · final",
so a config with two of them showed two identical claims and no hint that
only the last one is the default the router uses.

A superseded rule now loses those marks — it keeps its real Order in the rail
instead of the "·" that means final — and gains a "never applies" badge plus
a line naming the rule that beat it and what to do about it: give this one a
condition, or delete one of the two. Warn semantics throughout (--amber,
dashed frame, dimmed target chip): orange is the ACTIVE state on this
faceplate, and a rule the router ignores is the opposite of active.

Verdicts come from GET /api/rules/reachability and are keyed by the rule's
index in Rules, never by name — the config that prompted this had two rules
both called `default`. They are re-fetched after every save, and a verdict
whose echoed name/order no longer matches the row is dropped rather than
shown, so the window between an optimistic edit and the refetch cannot badge
a working rule.

Rule rows were also keyed by name in React, which silently collapses two rows
that share one; the key now carries the model index.

The mock fixture gains a second condition-less rule so `?mock` renders the
state, and mock.getRulesReachability derives its verdicts from the live
fixture config rather than hard-coding them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 88a82c7297 feat(routing): report rules that can never fire (B1)
A routing rule with no conditions at all is not matched in sequence — it
becomes the engine's route Final (generate/route.go buildRoute points Final
at it and moves on). Two consequences were invisible everywhere:

  * two condition-less rules retire each other, and the LAST one by Order
    wins, so an earlier "default -> direct" is dead while looking live;
  * a condition-less rule can NEVER retire a rule that HAS conditions —
    those are emitted ahead of Final whatever their Order.

A config in the field had two rules both named `default`, both with zero
conditions, order 20 -> direct and order 100 -> group:auto. One of the two
did nothing, the log was clean, and the panel drew both rows with the same
"default route · final" badge.

model.RuleReachability is the one implementation of the verdict, in the
stdlib-only leaf both consumers import, so the warning and the panel badge
cannot drift. generate.isCatchAll / effectiveRuleTarget / sortedRuleIndices
now delegate to it — three copies of "what is a default and what order do
rules run in" was how this would come back.

Scope is deliberately narrow: only condition-less over condition-less, which
is certain from the config. Whether one conditional rule's matchers subsume
another's is not decidable here, and a false "never fires" badge on a working
rule is worse than no badge.

Profiles are honoured: the analysis runs on the EFFECTIVE rules
(Model.EffectiveRules applies the active WAN profile's enable/disable), so a
rule the profile switched off is not blamed for retiring anything, and one it
switched on is. A SCHEDULED default never retires anything — outside its
window the rule above it is the default again — but can itself be retired by
an unscheduled one below it, which makes its schedule pure decoration.

Apply-time this reaches the operator through the existing status warnings,
graded by consequence rather than by "a setting is dead": critical when the
surviving default is `direct` while the retired one asked for a tunnel or a
block (the operator's default policy is not in effect and everything
unmatched leaves on the plain WAN); warning otherwise. The field config's own
shape — a dead `direct` under a live tunnel — is the warning case.

GET /api/rules/reachability serves the same verdict to the panel, the routing
analogue of the per-chain `used` flag on /api/groups/health. Keyed by index
into Rules, not by name: this config has two rules called `default`.

Diagnosis only — nothing is renamed, reordered, disabled or dropped, and
apply keeps working on a config that already has two defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:36:36 +03:00
omarandClaude Opus 5 a8f2b0f068 ci: derive package versions from the git tag (B4)
PKG_VERSION/PKG_RELEASE were hand-written literals nobody bumped, so
v0.2.2 … v0.2.6 all shipped as `shaterd 0.2.0-r3` with different binaries
inside (v0.2.6's ELF is 5 491 616 B against r2's 5 488 336 B). Both opkg
and apk offer an upgrade only when the feed's version string differs from
the installed one, so `apk update` saw nothing new and the routers could
not be updated through the normal path at all.

ci/version.sh is now the single source of truth. It derives the version
from `git describe`:

    tag `vX.Y.Z`   -> PKG_VERSION=X.Y.Z  PKG_RELEASE=1
    off-tag build  -> nearest tag + PKG_RELEASE=<commits since it> + 1
    no tag/no git  -> 0.0.0-r1 (below everything ever published)

Ordering verified with the real tools, not from memory — apk-tools 3.0.3
(`apk version -t`) and opkg 38eccbb1 (`opkg compare-versions`) agree that
0.2.0-r3 < 0.2.6-r2 < 0.2.6-r10 < 0.2.6-r12 < 0.2.7-r1 < 0.3.0-r1, so a
release always outranks the rolling builds that preceded it and rolling
builds grow monotonically between releases.

The value travels as SHATER_PKG_VERSION/SHATER_PKG_RELEASE in the SDK
build environment of BOTH lanes; the Makefiles keep a literal fallback so
a manual/offline build still works with no CI and no git. Because the
hand-off crosses docker, `su` and make's env import, ci/sdk-build.sh and
ci/sdk-build-apk.sh now ASSERT that the produced .ipk/.apk really carries
that version — the B4 failure mode was a stale version shipping silently,
and that can no longer happen quietly.

The binary agrees with the package: scripts/build-shaterd.sh takes
constant.Version from the same ci/version.sh (vX.Y.Z-rR[-g<sha>]) instead
of its own `git describe`, and the workflow computes it once per job.
Both build jobs now check out with fetch-depth: 0 — `git describe` needs
tags and ancestry, which the default shallow checkout has neither of.

byedpi is deliberately left alone: PKG_VERSION:=0.17.3 is upstream
ByeDPI's own version, what PKG_HASH pins and what tells an operator which
ByeDPI is installed. Stamping our tag on it would also be a downgrade —
every comparator reads 0.2.7 < 0.17.3 (component-wise, 2 < 17), verified.

Docs: INSTALL.md gains §2.1 (the scheme + the ordering evidence), and the
update sections of §5/§6 now explicitly warn against a bare `opkg upgrade`
/ `apk upgrade` and give the targeted form instead, quoting apk-tools 3:
"If list of packages is provided, only those packages are upgraded along
with needed dependencies". README.md and the release bodies match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:32:38 +03:00
omarandClaude Opus 5 0b32a6d58b fix(log): no ANSI colour outside a TTY — syslog and the log file stay grep-clean (B5)
Every log line the daemon produced carried aurora escapes, and under procd
stderr is not a screen, it is syslog:

  daemon.err shaterd[27540]: ...Z ESC[31mERRORESC[0m[0026]
  [ESC[38;5;193m1728741629ESC[0m 70ms] dns: exchange failed ...

`logread | grep ERROR` misses that line — the level word has invisible
bytes inside it — external collectors store the escapes forever, and a
captured log reads as mojibake.

Both producers defaulted to colour, and both are fixed at the producer,
because colour is a property of the DESTINATION and should never be
generated for a destination that cannot render it:

  * control plane (cmd/shaterd): log.Formatter{BaseTime: ...} left
    DisableColors at its false zero value. It now comes from
    controlLogFormatter(), gated on logsink.IsTTY(os.Stderr). The helper
    lives in an untagged file (same split as profilewatch.go) so it is
    unit-testable off the linux target.
  * engine (shater/generate): the generated option.LogOptions never set
    DisableColor, so box.New built a colouring formatter over the shared
    sink. logOptions() now sets it from the same TTY gate (seam:
    logColorAllowed).

logsink.IsTTY is the single source of the decision: a character-device
check, so no cgo, no termios and no new dependency on a CGO_ENABLED=0
musl-static binary. Under procd stderr is a pipe => no colour; an
interactive `shaterd run` from a shell keeps it.

The file half already stripped ANSI on the way out (emitLocked ->
stripANSI); that stays as the belt to this new braces, and the leak it
never covered — the syslog half — is now closed at the source.

Tests: the syslog half of the sink carries no 0x1b for any level with a
context ID set (the connection id is coloured by a separate branch of
log/format.go, so a level-only fix would still leak); the same for the
control-plane formatter and for a factory built from the REAL generated
log block. Each has a teeth check that a colouring formatter does emit
0x1b, so the guards cannot rot into passing for the wrong reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:31:01 +03:00
omarandClaude Opus 5 893fdc500c fix(shaterd): nodes reports the real inventory instead of an empty list (B2)
`shaterd --help` promised "print nodes as JSON"; the verb answered `[]`
unconditionally — `cmdReadStub("nodes", "[]")` on the CLI side and a
hard-coded `writeLine(conn, "[]")` in the daemon's control-socket handler.
The data was never missing: on the live router /etc/shater/subs/*.json
held 315 subscription nodes and GET /api/config reported 340. An empty
array is indistinguishable from a truthful "nothing is configured", so
the verb did not fail loudly, it lied quietly — the same inverted-lie
class as 9dc954029 / aec82d444.

`nodes` now reads model.ReadUCI() — `uci export shater` merged with the
per-subscription JSON caches — which is literally the call GET
/api/config serves and generate builds the engine from, so the verb
cannot drift from the panel or from the running engine: there is no
second assembly here to drift. Both ends use the same nodesJSON():
the daemon answers over the control socket (like `stats`), and the CLI
falls back to reading the same on-disk state when no daemon is running
(like `status`). A read failure goes to stderr with a non-zero exit
instead of printing `[]`, so an empty list on stdout now means one thing.

Output is a purpose-built view rather than raw model.Node: the share-link
URI is a credential and CLI output ends up in tickets and cron mail, so
the view reports what the link decodes to (protocol/server/port) plus the
model's own facts (enabled/sub/egress/stale/fingerprint). Nodes whose URI
does not parse are still listed, with the reason in `parse_error` — the
engine skips exactly those, and hiding them would be the same lie smaller.

cmdReadStub keeps `stats`, where the default IS the truth (nothing was
counted without an engine), and now says so in its doc comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:30:27 +03:00
omarandClaude Opus 5 02c266188f docs: live test report for v0.2.6 on mini_router (79 checks, 5 findings)
Full cycle on real hardware (BPi-R3 Mini, ImmortalWrt 25.12-linkup): purge the
previous install, install from the signed apk feed, verify the default state,
restore a working config with 315 subscription nodes, then exercise the data
plane, panel API, config lifecycle, resilience and DNS.

74 PASS. Findings (detailed separately): two catch-all `default` rules where the
first sends all unspecific traffic direct and makes the second unreachable;
`shaterd nodes` is a stub returning [] while usage promises the node list; DNS to
the router LAN address dies after `service shater restart` (stop+pause+start is
fine); PKG_RELEASE unchanged since v0.2.1 so v0.2.2..v0.2.6 all ship as r3; ANSI
colour codes reach syslog.

Also records the four-iteration CI hunt that ended in the green apk lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 12:09:52 +03:00
omarandClaude Opus 5 024e9308c9 fix(ci/apk): strip the SDK's generated per-package default m blocks
release / aarch64_cortex-a53 (push) Successful in 3m21s
release / x86_64 (push) Successful in 3m12s
release / apk aarch64_cortex-a53 (push) Successful in 5m32s
release / apk x86_64 (push) Successful in 2m32s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Run 60 settled what runs 58/59 left open. The second pass wrote an explicit
`# CONFIG_PACKAGE_kmod-x is not set` for all 1126 selected kmods and re-ran
defconfig; the count came back 1078, unchanged. The same explicit form DID hold
for CONFIG_ALL/ALL_KMODS/ALL_NONSHARED in the same run.

The difference is prompts. kconfig honours a user value only for symbols that
have one — sym_calc_value ignores S_DEF_USER for a promptless symbol and falls
back to its `default`. ALL* carry prompts in the SDK's Config.in; the blocks
convert-config.pl generates are bare:

    config PACKAGE_kmod-mlx5-core
            tristate
            default m

No value written into .config can turn those off, so remove the `default m`
itself: drop every generated `config PACKAGE_*` block from Config-build.in
before the first defconfig. Nothing is lost — those blocks only replay which
packages the buildbot built. The packages stay declared, with prompts, by the
package tree (tmp/.config-package.in), which is what makes our four selectable
and what `select` acts on; KERNEL_*/LIBC/TOOLCHAIN blocks are untouched, so the
SDK still reproduces its own toolchain settings.

The .config second pass is kept as a cheap backstop (it no-ops once the count
is 0), as are both tripwires.

Verified: bash -n on the file and on the extracted INNER body; the paragraph
delete tested on a synthetic Config-build.in (3 PACKAGE blocks -> 0, KERNEL_*,
LIBC and TOOLCHAINOPTS preserved); the missing-file path exercised under set -eu.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
apk-v0.2.6-aarch64_cortex-a53 apk-v0.2.6-x86_64 v0.2.6
2026-07-25 01:27:37 +03:00
omarandClaude Opus 5 eab1db2c3f fix(ci/apk): second defconfig pass — deselect the SDK's per-kmod default m
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
Turning ALL/ALL_KMODS/ALL_NONSHARED off (22d7161c0) provably worked — run 59
logs all three as `is not set` after defconfig — and changed the kmod count by
exactly zero, 1078 both times. The kmods never came from ALL_KMODS.

They come from the SDK itself. target/sdk/Makefile generates the SDK's
Config-build.in by running convert-config.pl over the BUILDBOT's .config, in
which ALL_KMODS=y had already expanded into one `CONFIG_PACKAGE_kmod-*=m` line
per module. convert-config.pl turns every `CONFIG_X=<val>` line into a symbol
with an unconditional `default <val>`; its `next if /^(# )?CONFIG_PACKAGE/`
filter sits in the `else` branch, which a line containing `=` never reaches.
The SDK therefore ships ~1078 verbatim blocks of `config PACKAGE_kmod-x /
tristate / default m`, none of which consult ALL_KMODS.

Fix: a second pass. The names only exist after kconfig has expanded the tree,
so after the first defconfig rewrite every selected kmod to `is not set` and
re-run defconfig. Two documented kconfig rules make this exact:
  - an explicit value in .config beats a `default` (same rule that kept our
    `# CONFIG_ALL* is not set` lines alive in run 59) -> the ~1078 stay off;
  - `select` is OR-ed in after the user value, so shater-core's
    `DEPENDS:=+kmod-nft-tproxy +kmod-nft-socket` brings those (and their
    transitive kmods) back on their own.

Also correct the tripwire message, which still blamed CONFIG_ALL_KMODS: it now
prints the ALL* state AND the first few surviving kmods, so the two failure
modes are distinguishable at a glance.

Verified: bash -n on the file and on the extracted INNER heredoc body; the
rewrite simulated against a run-59-shaped .config (1078 -> 0 selected, our 4
packages, LOCALMIRROR and the ALL* lines untouched).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v0.2.5
2026-07-25 01:16:05 +03:00
omarandClaude Opus 5 22d7161c08 fix(ci/apk): disable the SDK's ALL/ALL_KMODS mass-select
release / aarch64_cortex-a53 (push) Successful in 3m25s
release / x86_64 (push) Successful in 3m14s
release / apk aarch64_cortex-a53 (push) Failing after 1m3s
release / apk x86_64 (push) Failing after 1m3s
release / release (push) Successful in 8s
release / release apk (push) Successful in 6s
Run 58 proved the previous commit aimed at the wrong thing, and the
diagnostics it added are what showed it: "0 lines carried over" plus a
`grep: .config: No such file or directory`, then 1078 kmods selected
anyway (1109 on x86_64). So an SDK tarball ships no top-level .config at
all — there was never a buildbot config for us to be appending to.

The real source is the SDK's OWN top-level Config.in, target/sdk/files/
Config.in, which it carries instead of the main tree's:

    config ALL_NONSHARED ... default ALL
    config ALL_KMODS     ... default ALL
    config ALL           ... default y

In the main tree all three default to n; the SDK flips ALL to y so that
`make world` in a bare SDK builds something. `make defconfig` therefore
selects the whole kernel from ANY .config, empty or not. This is stock
OpenWrt rather than an ImmortalWrt quirk — openwrt/openwrt's copy is
identical, which also means the awg-openwrt reference builds every kmod
too; it just never meets a disk quota on GitHub's runners.

Fix: write all three out as `# CONFIG_X is not set` before defconfig.
They have prompts in the SDK's Config.in, so they are user-settable and
an explicit value beats the default; `CONFIG_X=n` is not reliably
honoured for bools, hence the `is not set` form. Setting all three, not
just the root ALL, keeps this working whichever symbol roots the chain
in a future SDK.

Drops the hand-rolled CONFIG_TARGET_*/CONFIG_KERNEL_* carry-over as
redundant: target/sdk/convert-config.pl bakes the buildbot's non-package
settings into the SDK's generated Config-build.in as kconfig defaults,
so defconfig reproduces them by itself. A soft branch keeps target
identity and CONFIG_USE_APK if some future SDK does ship a .config.

Diagnostics gain a post-defconfig readout of the three mass-select
symbols and, while the list is short, the actual kmods selected — a
count of 0 is not fatal (the router's base feed carries them) but is
worth seeing. Guards and the 200 threshold are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v0.2.4
2026-07-25 00:57:19 +03:00
omarandClaude Opus 5 24c5a1615d fix(ci/apk): build .config from scratch — stop packing all 3593 kmods
release / aarch64_cortex-a53 (push) Successful in 3m19s
release / x86_64 (push) Successful in 3m18s
release / apk aarch64_cortex-a53 (push) Failing after 1m4s
release / apk x86_64 (push) Failing after 1m2s
release / release (push) Successful in 7s
release / release apk (push) Successful in 5s
Both apk jobs of v0.2.2 died with `Disk quota exceeded`. The SDK was
running `apk mkpkg` on 3593 kmod-* packages (mlx5, amdgpu, ata, isdn —
none of which we ship) before it ever got near our four.

Root cause: ci/sdk-build-apk.sh APPENDED our package selections to the
.config that ships inside the ImmortalWrt SDK tarball. That file is the
buildbot's fully-expanded config and carries CONFIG_ALL_KMODS=y plus
CONFIG_ALL_NONSHARED=y (see config.buildinfo next to the SDK), so
`make defconfig` re-selected every kernel module of the target as =m and
package/kernel/linux/compile — pulled in via shater-core's nft kmod
deps — packed the lot.

Fix, modelled on Slava-Shchipunov/awg-openwrt's "Setup SDK and feeds":
start the .config EMPTY so kconfig can only pull in what our packages
actually select. Carried over from the SDK's .config, nothing more:
the target choice and its BOARD/SUBTARGET/ARCH_PACKAGES identities (a
wrong guess here means silently cross-compiling for another arch),
CONFIG_USE_APK (decides .apk vs .ipk — the point of this lane), and
CONFIG_KERNEL_* verbatim (they generate the kernel .config; dropping one
makes the buildsystem reconfigure and rebuild the SDK's prebuilt kernel).

Also adds the diagnostics this lane never had, since a failed run leaves
a 27 MB log: the carried-over identity lines, the post-defconfig kmod
count and target readout, a hard check that all four of our packages
survived defconfig, an abort if the kmod count is back in the hundreds,
and du/df after compile.

opkg lane (ci/sdk-build.sh, ci/make-index.sh) untouched. LOCALMIRROR,
CONFIG_DOWNLOAD_FOLDER and every cache path are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v0.2.3
2026-07-25 00:31:28 +03:00
omarandClaude Fable 5 4492f0599c ci: harden feed/packaging shell scripts
release / aarch64_cortex-a53 (push) Successful in 7m3s
release / x86_64 (push) Successful in 3m13s
release / apk aarch64_cortex-a53 (push) Failing after 4m34s
release / apk x86_64 (push) Failing after 2m35s
release / release (push) Successful in 8s
release / release apk (push) Successful in 5s
- ci/make-index.sh: set -e → set -euo pipefail so a failing sha256sum|cut in
  the signed Packages index can't mask an empty SHA256. Script survives -u
  (all vars use :? or :- defaults).
- .github/deb2ipk.sh: quote $2/$DEB_NAME/output, derive the deb name from the
  copied file via basename instead of parsing `ls *.deb` (glob-fragile), add a
  trap-based tmpdir cleanup, and set -euo pipefail.

bash -n clean on both.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v0.2.2
2026-07-24 23:34:50 +03:00
omarandClaude Fable 5 bc2b53069a fix(security): audit remediation — file perms, const-time auth, leaks, CSPRNG
Backend audit fixes (upstream-file edits wrapped in // lx: markers):

- experimental/libbox oom_report.go/report.go: OOM reports + configuration.json
  (server secrets/keys) were written world-writable — 0o777 dirs / 0o666 files
  → 0o700 / 0o600. [sec-perms]
- daemon/server.go + experimental/libbox/command_server.go: gRPC auth secret
  compared with != (timing oracle) → crypto/subtle.ConstantTimeCompare.
  [sec-consttime]
- service/oomkiller/timer.go: network-extension cleanupTriggered logic was
  inverted, so FreeOSMemory was never called after a trigger; flip both
  assignments so a trigger schedules the deferred free and the next poll runs +
  clears it. [sec-oomcleanup]
- transport/v2rayxhttp/client.go (lx-native file): session id used math/rand →
  crypto/rand, matching Xray's uuid.New() entropy and removing the spoof surface.
- daemon/started_service_tailscale_ssh.go: forwardSSHAgentChannel leaked a
  goroutine + the ssh-agent fd on every closed session (second io.Copy blocked
  on an idle agent Read forever); tie both copies + the session ctx to a
  cancel that closes both ends. [sec-sshagent]
- daemon/managed_service.go: TriggerOOMReport had no gate — rate-limit to
  1/min so an authenticated client can't spin secret-bearing dumps. [sec-oomgate]
- route/reachability_lx.go (lx idle-suspend file): idle tick read r.idleStop in
  select while stopIdleSuspend niled it after close (race + goroutine leak on
  Close-during-tick); pass the stop channel to the loop by value.

go build ./... (default) and the D9 shaterd linux build (tags
with_quic,with_wireguard,with_utls,badlinkname,tfogo_checklinkname0,with_xhttp,
with_awg,with_lx_command) are green; go vet clean (2 pre-existing unsafe.Pointer
warnings in TriggerDebugCrash/debug.go, untouched); go test ./route/...
./daemon/... ./service/oomkiller/... green incl. -race with with_lx_idle_suspend
and v2rayxhttp with with_xhttp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:34:50 +03:00
omarandClaude Fable 5 c257d6c5cc docs(readme): rewrite root README as Russian shater product face
- README.md: new Russian product README (what/features/architecture
  mermaid/install both feeds/build/repo layout/CI/upstream/docs/license)
- README.en.md: concise English mirror (root readme was previously English)
- README.ru.md: demoted to a pointer stub (was the sing-box-lx fork readme,
  a competing Russian README) -> points to README.md + engine-fork docs
- docs-shater/README.md: folder index

Install commands copied verbatim from docs-shater/INSTALL.md; all links
verified against existing files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:31:06 +03:00
omarandClaude Fable 5 225cce5397 chore: purge working-session junk from tree; gitignore recurrence
Remove untracked-quality artifacts accidentally committed during work
sessions (all authored downstream, unreferenced anywhere in code/docs/CI):
- 5 session screenshots in repo root (devices-after-copy-fix.png,
  live-final-groups.png, profiles-*-active.png, profiles-final-vm-wan0.png)
- tmp/gen_linux_test (29 MB throwaway traffic-gen binary)

Guard against repeats: ignore /*.png (root screenshots) and /tmp/.
Upstream files (mkdocs.yml, .fpm_*) and the SPECS-020 research .log are
left untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 23:23:04 +03:00
omarandClaude Fable 5 84d2766592 chore: remove stray committed test binary, ignore nested .idea, bump PKG_RELEASE
release / aarch64_cortex-a53 (push) Successful in 6m26s
release / x86_64 (push) Successful in 3m32s
release / apk aarch64_cortex-a53 (push) Successful in 4m58s
release / release (push) Has been cancelled
release / release apk (push) Has been cancelled
release / apk x86_64 (push) Has been cancelled
- drop c/Users/.../gen_linux_test (28MB binary accidentally committed in 129e31fbd)
- .gitignore: ignore .idea/ at any depth (shater/.idea from IDE)
- CLAUDE.md: orchestrator delegates to model fable
- bump shaterd/shater-core r2->r3, luci-app-shater r1->r2 for v0.2.1 release

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
v0.2.1
2026-07-24 23:18:01 +03:00
omar 33f940c2fb ci(release-apk): publish available arch feeds even if one build fails
release-apk now runs with if: !cancelled() so an unrelated arch build
failure (e.g. x86_64) does not block publishing the aarch64 apk feed.
download-artifact only fetches existing artifacts and the publish loop
already skips missing apkfeed-* dirs.
latest
2026-07-24 22:52:17 +03:00
omar a57717dabb health plan S7: docs, contract comments, SPEC 019 update
release / aarch64_cortex-a53 (push) Successful in 3m43s
release / x86_64 (push) Successful in 3m30s
release / apk aarch64_cortex-a53 (push) Successful in 5m11s
release / apk x86_64 (push) Failing after 5m8s
release / release apk (push) Has been skipped
release / release (push) Successful in 12s
- lx-changelog: health board + observatory + global probe + sub cache entry
- DECISIONS.md: D18 board vs delete-and-overlay, D19 observatory vs sweep, D20 global probe settings
- contract comments: urltest.go CheckOutbounds (fast circuit) + observatory.go loop (background circuit + freshness gate) document the two-circuit split
- SPEC 019: dial-error section updated - slots still not moved, but board verdict demotes dead slot on next pick + retry (§5.B); sticky/replace-in-slot/never-shrink invariants preserved
v0.2.0-healthplan
2026-07-24 18:30:48 +03:00
omar 129e31fbdc health plan wave 3: integration tests + chain unused-badge
- S6: integration tests in shater/generate (health_retry, chain_health, observatory_reach); portable tests pass -race; linux-tagged tests vet+build under GOOS=linux for OpenWrt CI
- §5.E fix: unused-badge extended to chains; observatory used-set (chain exit tags) inverted via parseChainExitTag -> usedChainNames -> ChainHealth(names) -> /api/groups/health chains field -> ChainRow renders gh--unused badge (same pattern as groups)
2026-07-24 18:24:05 +03:00
omar 4f0618515e health plan wave 2: alive-only selection+retry, observatory replaces sweep
- S3: Select() alive-only by verdict; dial failure marks board + retries <=3 within ctx; ListenPacket retries to first send; balancer slot liveness reads board verdict; testNodes marks-fail instead of delete; SPEC 019 slot invariants preserved; selector untouched
- S4: new observatory.go (reachability plan from rules, batch<=24/concurrency<=12/timeout 5s, freshness gate, cursor preserved on identical plan); probeplan BuildObservatoryPlan; health.go on board verdicts (TTL=max(3*interval,10min)); engine.dead overlay removed; sweep.go+probeall.go+TestAllNodes+/api/nodes/test removed (->404); GroupHealth.Used published; exit-test extended to chains; panel unused-badge + chain Test button; stats on board
2026-07-24 16:33:28 +03:00
omar fd698162c9 health plan wave 1: urltest health board, global probe settings, sub cache out of UCI
- S1: History gains LastOK/Delay/LastFail; MarkFailed/Verdict on storage (common/urltest/board_lx.go); StoreURLTestHistory preserves LastFail
- S2: per-group ProbeURL/ProbeInterval removed, globals only; sweep_interval drained as dead option; panel fields dropped
- S5: subscription nodes cached in /etc/shater/subs/<name>.json; UCI keeps manual nodes only; sub update writes cache file; panel PUT split; legacy from_sub migration
2026-07-24 13:33:23 +03:00
omar ffa78d67bc feat(panel): drop the query log from Overview
The filtered-query-log block (SegMeter + top blocked + live QueryLog) is
redundant with Insights. The stats poll stays - it still feeds the DNS
filtering and Groups modules.
2026-07-24 01:09:48 +03:00
omar 61e51495c3 feat(panel): prune subscription editor to essentials with a key-value header list
The subscription form now carries exactly: name, URL, update interval,
fetch via (+detour when proxied), User-Agent, HWID, and extra headers.
Format, device identity, regex/proto/country filters, dedup and expiry-alert
knobs are gone from the form (still honoured from UCI; a save carries them
through untouched). Headers are edited as key-value rows and serialize to
the existing `Headers: []string` "Key: value" contract. Name is editable:
a rename rewrites FromSub on the sub's cached nodes and refuses collisions.
2026-07-24 01:09:48 +03:00
omar bfe71cd1dd feat(generate): fail closed on unresolved rule targets, never fall back to the default route
Rule.Kill ""/"default" used to drop the rule, letting its traffic fall
through to the broader rules below and finally the default route - a silent
leak of exactly the traffic the operator singled out. ruleKillFallback now
always returns an outbound: ""/"default"/"closed"/unrecognised block the
rule's traffic in place; only an explicit kill=open goes direct. The default
route exists solely for traffic no rule matched.
2026-07-24 01:09:48 +03:00
omar 9b3644becb fix(alert): don't repeat the IP as the device name in new-device alerts
With no DHCP hostname the alert read "10.67.0.223 (mac) at 10.67.0.223".
Lead with the MAC instead: "aa:bb:cc:dd:ee:ff at 10.67.0.223".
2026-07-24 01:09:47 +03:00
omar 0a8bdbaf46 feat(devices): merge discovered devices by MAC
One physical device with several addresses (v4+v6, multiple leases) used to
show as several devices. Discover now folds addresses sharing a MAC into a
single row: new `ips` field lists every address primary-first, `ip` stays
the primary (most recent lease), state is the best among addresses.
MAC-less hosts remain one-per-IP. The panel shows the extra addresses as
secondary chips; naming keys the config entry by MAC whenever it is known.
2026-07-24 01:09:47 +03:00
omar ebe2e7807b feat(stats): drop router-originated DNS queries from insights
The engine resolves domains for itself (node server names, urltest probes,
subscription/DoH fetches). Those queries carried an invalid client address
and still landed in every insights surface. Gate them out at the single
ingestion point (Aggregator.handleEvent): an event with an invalid or
loopback client is dropped before totals, top domains, per-server counts,
the timeline, and the query-log rings. Only LAN-client traffic is collected.
2026-07-24 01:09:47 +03:00
omar c1e1b17a61 feat(profiles): remove preset packs
The built-in block-ads / ru-bypass / private rule bundles are gone:
model.Preset, Model.Presets, the `config preset` UCI section, its render,
the panel Preset type, and every fixture. The generate-side expansion was
already removed with the profile rewrite in the previous commit.
2026-07-24 01:09:47 +03:00
omar 782770f306 feat(profiles): drop default route/egress/schedule overrides; pick uplink from UCI interfaces
A profile is now a pure uplink-conditional rule switch: Name/Enabled/
Priority/MatchIface/Enable-DisableRules/EndpointResolver. The per-profile
DefaultTarget/DefaultEgress overrides and the profile-level schedule window
(SchedDays/SchedStart/SchedEnd/SchedUTCOffset) are removed from the model,
UCI parse/render, the generator, the WAN watcher, and the panel. Rule-level
scheduling is untouched.

The panel's uplink condition is now picked from a dropdown of the router's
UCI interfaces (GET /api/interfaces, same source as the egress picker);
stored interfaces missing from the live list render as stale chips.

generate/profile.go is rewritten here (applyProfilesAndPresets ->
applyProfiles), which also drops the generate-side preset-pack expansion;
the preset model/UCI/panel surface is removed in the next commit.
2026-07-24 01:09:46 +03:00
omarandClaude Opus 4.8 8980a25a59 perf(ci): cache SDK feeds checkouts — the biggest recurring build cost
Audit of the run-51 logs showed actions/cache@v3.3.2 works on the act_runner
(cold: "Cache saved" x4; next job: "Cache restored" in ~2s, npm --fast skip,
usign/dl reused) and the sdk-cache mirror seeds correctly — but the single
biggest recurring cost was NOT cached: `scripts/feeds update -a` re-cloned
base+packages+luci+routing+telephony every run (~7.8 min warm x 4 SDK jobs on
the serial runner ≈ ~28 min/run wasted; github ~1 MB/s from this host).

Cache .cache/feeds/{opkg,apk} (workspace dir, actions/cache-persisted, visible
in the SDK container via --volumes-from) symlinked over the SDK's empty feeds/:
`feeds update` now git-fetches deltas (seconds) instead of full clones, always
checking out feeds.conf's pins. Fail-safe: any error on the cached checkouts
wipes the cache and clones fresh. Key by SDK release (feeds-opkg-24.10.4 /
feeds-apk-25.12.1) — stable across runs, invalidates on an SDK bump; both arch
jobs of a lane share one entry (identical pins, serial runner).

Steady-state warm run: ~60+ min -> ~20-22 min. Also documented in the workflow
header: never key a cache on github.sha — each cache SAVE stalls the act_runner
~3 min, so per-run-changing keys would add +3 min/entry every run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 22:41:53 +03:00
omarandClaude Opus 4.8 cb983afee1 perf(ci): stall-proof SDK fetch + caching (SDK/dl/go/npm/apt/usign) + concurrency
Builds were dominated by re-fetching the ImmortalWrt 25.12 SDK tarball
(~300 MB) every run, and a stalled downloads.immortalwrt.org transfer wedged
the apk job for 40+ min (plain `wget -q`, no timeout — same class as the
elfutils hang).

- New ci/fetch-sdk.sh (runner-side): cache -> our durable `sdk-cache` release
  mirror -> upstream with a stall-kill (curl --speed-limit 64K --speed-time 60
  --max-time 1800) + 3 retries + zstd-magic/size validation; seeds the mirror
  best-effort (github.token, non-fatal) so cold runs never touch upstream again.
  A 40-min hang is now impossible; the in-container fallback wget also gets
  --timeout=60 --tries=3.
- actions/cache@v3.3.2 (last release on the OLD cache API that Gitea act_runner
  implements; v4/v3.4.x use the new GitHub cache service) for: SDK tarball, SDK
  dl/ sources (hash of package Makefiles; PKG_HASH re-verified so a stale cache
  can't leak a wrong source), Go mod+build (go.sum), npm node_modules
  (package-lock.json) with build-shaterd.sh --fast, apt archives, built usign.
  Degrades safely if the cache server is off — the SDK mirror is independent.
- concurrency group release-${github.ref} cancel-in-progress so a re-dispatch
  cancels the stale run instead of piling up (tags stay isolated).

Signing (usign/apk), both keys, per-arch publish, manual triggers, LOCALMIRROR
and the scoped 4-package collection are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
sdk-cache
2026-07-23 21:24:49 +03:00
omarandClaude Opus 4.8 6d2a729eaa feat(panel): gate byedpi egress on package presence
byedpi (ciadpi) ships as a separate optional package; the panel offered the
`byedpi` egress type regardless, so selecting it without the package installed
created a dead, fail-closed egress. Now GET /api/status reports
`byedpi_installed` (exec.LookPath("ciadpi"), os.Stat fallback), and the egress
type picker disables the ByeDPI option with a hint when it's absent. Existing
byedpi egresses are never hidden or rewritten (config is sacred) — shown with an
amber warning and still round-trip on save; only NEW selection is blocked.
Unknown status (older daemon / fetch fail) => no gating.

Bump shaterd PKG_RELEASE 1 -> 2 (the SPA is embedded in the daemon binary).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 20:26:23 +03:00
omarandClaude Opus 4.8 4a1542b30a fix(shater-core): release procd flock so install can't deadlock the pkg manager
On a live `apk add` / `opkg install`, shater-core's post-install hung forever
(observed on BananaWRT 25.12 at "Executing shater-core...post-install", child
`flock 1000` in locks_lock_inode_wait). Root cause: a USE_PROCD init sources
/lib/functions/procd.sh on every rc.common action, whose procd_lock takes a
BLOCKING exclusive flock on /var/lock/procd_<svc>.lock held until the process
exits. shater-cron re-execs itself as the eternal `loop`, so it held that lock
forever; base-files' default_postinst then ran `/etc/init.d/shater-cron enable`
synchronously inside the transaction, blocking on the flock while the package
manager waited on the postinst — a permanent deadlock.

Fix (two layers):
- shater-cron `loop()`: `exec 1000>&-` closes fd 1000 up front so the eternal
  loop never holds the rc.common flock (no-op when procd_lock is absent).
- 30_shater-core: defer enable/restart into a detached (setsid + bounded)
  background block that waits for apk/opkg to finish before touching init.d,
  with all fds to /dev/null (a held stdout pipe would hang apk on EOF too).

Bump PKG_RELEASE 1 -> 2 so existing installs pick up the fix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 20:26:22 +03:00
omarandClaude Opus 4.8 b7070ad7e8 fix(ci): install python3-distutils for the ImmortalWrt 25.12 apk-SDK
The apk lane runs the SDK on a bare debian:bookworm host, and the ImmortalWrt
25.12 SDK prerequisite check requires python3-distutils ("Checking
'python3-distutils'... failed. Prerequisite check failed." ->
.prereq-build Error 1), aborting before any package built. The opkg lane was
unaffected because the openwrt/sdk image ships the prereqs. Add
python3-distutils (and python3-setuptools defensively) to the host deps. apk
lane only; opkg untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 17:31:15 +03:00
omarandClaude Opus 4.8 ea24bad232 fix(ci): make $OUT writable for the SDK user and collect only our 4 packages
Two build-harness bugs surfaced once the SDK builds actually ran:

1. Permission denied writing the feed. ci/build-feed.sh creates $OUT as root on
   the runner, but the openwrt/sdk container runs as the unprivileged `buildbot`
   (uid 1000) — so `cp` of the .ipk into $OUT failed ("Permission denied"),
   yielding 0 packages and then "usign signing failed" (nothing to sign). Set
   `chmod 0777 "$OUT"` on the runner before docker run (a chmod from inside the
   container, as buildbot, cannot fix a root-owned dir). The apk lane already
   chmods $OUT from its root debian container, so it was unaffected.

2. Collecting the whole SDK. ci/sdk-build.sh did `find bin -name '*.ipk'`, which
   swept up the hundreds of prebuilt kmod/base .ipk shipped in the SDK image —
   bloating the feed and signing foreign kmods under our key. Collect strictly
   our four by name (`<pkg>_*.ipk`) and require >=4. Applied the same narrowing
   to ci/sdk-build-apk.sh (apk names carry no arch: `<pkg>-*.apk`), keeping the
   "wrong SDK produced only .ipk" guard.

No change to the feed format/signing (usign/KEY_BUILD/shater-feed.pub, apk EC
key), the package set, or triggers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 17:09:07 +03:00
omarandClaude Opus 4.8 1d285cf34f fix(ci): route SDK source downloads through OpenWrt CDN mirror
The apk (and opkg) SDK builds intermittently hung fetching build-time
sources like elfutils-0.192.tar.bz2 from sourceware.org: curl's
--connect-timeout covers only the TCP handshake, not a stalled mid-transfer,
so a slow upstream hangs the whole job (no --max-time in OpenWrt download.mk).

Set CONFIG_LOCALMIRROR=https://sources.cdn.openwrt.org in .config before
`make defconfig` in both ci/sdk-build-apk.sh and ci/sdk-build.sh so the SDK
tries the fast OpenWrt source CDN before each package's own PKG_SOURCE_URL —
fixes elfutils and any other flaky upstream. Mirror verified to hold the file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:50:42 +03:00
omarandClaude Opus 4.8 5092004b40 fix(ci): init wireguard-go submodule in the opkg build job too
Same root cause as the apk jobs: scripts/build-shaterd.sh builds through a
go.mod `replace => ./submodules/wireguard-go` (AmneziaWG fork, bumped in
16a47b596), and actions/checkout does not fetch submodules by default, so
`go build` died with "reading submodules/wireguard-go/go.mod: no such file or
directory" in the opkg build jobs (x86_64 + aarch64_cortex-a53) as well. Init
only that one submodule — build-harness only, no change to the opkg feed
format/signing (usign/KEY_BUILD/shater-feed.pub) or package set.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:22:55 +03:00
omarandClaude Opus 4.8 bcdea8f04c fix(ci): init wireguard-go submodule in apk build jobs
scripts/build-shaterd.sh builds via a go.mod `replace => ./submodules/
wireguard-go` (the AmneziaWG-patched fork), so that submodule must exist or
`go build` dies with "reading submodules/wireguard-go/go.mod: no such file or
directory". actions/checkout does not fetch submodules by default. Init only
that one submodule (public GitHub URL; clients/apple+android are large and
unused) in the additive build-apk jobs — the opkg build jobs are left untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:16:23 +03:00
omarandClaude Opus 4.8 cd598b0fe2 feat(ci): add apk (ImmortalWrt/BananaWRT 25.12) release lane
Additive next to the opkg/24.10 lane — nothing existing changed. The same 4
packages (shaterd, shater-core, luci-app-shater, byedpi) are built through the
official ImmortalWrt 25.12 apk-SDK and published as per-arch rolling releases
apk-latest-<arch> / apk-<tag>-<arch> (x86_64, aarch64_cortex-a53).

- ci/sdk-build-apk.sh: drives the 25.12 SDK inside debian:bookworm, compiles
  .apk, then `apk mkndx --root T --keys-dir T/keys --allow-untrusted
  --sign KEY --output packages.adb *.apk` — the exact form the OpenWrt 25.12
  buildsystem uses (unsigned members, signed index).
- ci/build-feed-apk.sh: per-arch runner entrypoint (same --volumes-from and
  artifact-order contract as ci/build-feed.sh).
- ci/gen-apk-key.sh: one-shot EC (prime256v1) keypair generator; private half
  -> Gitea secret KEY_APK, public dist/shater-apk.pem committed.
- release.yml: additive build-apk / release-apk jobs; `on:` triggers untouched
  (v* tags + workflow_dispatch); apk release tags deliberately non-`v*`.
- docs-shater/INSTALL.md section 6, .gitignore (out-apk/), dist/shater-apk.pem.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:11:23 +03:00